Method for generating training data for software agent, method for training software agent, method for assisting driver when controlling motor vehicle, and motor vehicle

By generating training data, using software agents to learn driver reactions and adjust the actuator system of the driver assistance function, the problem of adapting the driver assistance system to individual drivers is solved, and the driver's acceptance and usage frequency are improved.

CN120659734APending Publication Date: 2025-09-16VOLKSWAGEN AG
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202480011767.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-09
Filing Date
2024-01-09
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing driver assistance systems have difficulty effectively adapting to individual drivers' driving styles and decisions, leading to negative reviews or system shutdowns.

Method used

By generating training data, software agents are used to learn the driver's reactions and preferences, and the actuator system of the driver assistance function is adjusted to reduce the difference between the driver and the system. Reinforcement learning and supervised learning methods are used to optimize the adaptation of the driver assistance function.

Benefits of technology

It achieves efficient adaptation of driver assistance functions to individual drivers, and improves drivers' acceptance and frequency of use of assistance functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120659734A_ABST
    Figure CN120659734A_ABST
Patent Text Reader

Abstract

The invention relates to a method for generating training data (1) for a software agent (2) which sets up a driver assistance function (3) for controlling a motor vehicle, in which method a trajectory (16) for the motor vehicle is planned for a determined context (5) by means of the driver assistance function (3), according to the invention, a longitudinal control and / or a transverse control for guiding settings of the motor vehicle along a trajectory (16) is planned, and an actuator system of the motor vehicle is set for implementing the planned longitudinal control and / or transverse control, a reaction of the driver (4) to the set actuator system is determined, and the driver (4) is ascertained as a function of the determined reaction. The difference (11) between the longitudinal control and / or lateral control of the motor vehicle along the trajectory (16), said longitudinal control and / or lateral control being planned by the driver, said longitudinal control and / or lateral control being characterized by the response, and said longitudinal control and / or lateral control being planned by the driver assistance function (3), said longitudinal control and / or lateral control being characterized by the affected actuator system, said longitudinal control and / or lateral control being characterized by the difference value. And providing the context together with the determined assigned difference value as training data (1) for the software agent (2).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for generating training data for a software agent (or software intelligent body, i.e., Software agent), a method for training a software agent, a method for assisting a driver in controlling a motor vehicle with the aid of a driver assistance function, and a motor vehicle with a driver assistance system. Background Art

[0002] DE 10 2018 202 146 A1 discloses a method for selecting a driving configuration for a motor vehicle, a driver assistance system therefor, and a motor vehicle equipped therewith. The method is configured to detect, based on the driver's operating behavior, a deviation between the driving configuration currently desired by the driver and one of a plurality of predefined driving configurations used up to that time. After detecting the deviation, a query is output to the driver as to whether the driver assistance system should learn the deviation for the current situation. After the driver confirms the query, the driver assistance system adapts at least one parameter that influences the selection of the driving configuration to be used to the current situation, such that the deviation is taken into account when automatically selecting the driving configuration to be used in future situations corresponding to the current situation.

[0003] Furthermore, CN 109 927 725 A discloses an adaptive speed regulation system with the capability of learning driving style.

[0004] Furthermore, EP 3 690 769 A1 discloses a learning method for obtaining at least one personalized reward function for executing a reinforcement learning algorithm and which corresponds to a personalized optimal strategy for the driver.

[0005] Driver assistance systems and automated driving functions are increasingly taking over human driving tasks. These systems and functions also bring their own pre-programmed driving style and driving decisions to the driving task. In this context, driving style primarily describes aspects that, while also influenced by external factors such as traffic conditions, weather, time of day, or the driver's health, are largely attributable to individual characteristics. These include speed and route selection, acceleration and braking behavior, approach to roundabouts, preferred distance to the edge, coasting at red lights or on country roads or at town entrances. In this context, driving decisions primarily describe behaviors such as overtaking, lane changes, merging, general interactions in road traffic, and how closely a vehicle is followed before taking action. If these pre-programmed characteristics differ from the driver's driving style and the driver's personal or subjective decisions, this can lead to negative reviews or even the deactivation or non-use of these driver assistance systems. Therefore, one of the greatest challenges, particularly with driver assistance systems, but also with fully automated driving, is for customers to positively evaluate the system and demonstrate a readiness to truly use it. There are various control options here, the core of which is the discrepancy, or ideally, the absence of a discrepancy, between the driver's driving style and the situational decisions of the driver assistance system. Summary of the Invention

[0006] The object of the present invention is to provide a solution which allows a driver assistance function to be particularly well adapted to the driver of the motor vehicle.

[0007] This object is achieved by the subject matter of the independent claims. Further possible embodiments of the invention are disclosed in the dependent claims, the description, and the drawings. The features, advantages, and possible embodiments described within the scope of the description with respect to one of the subject matter of an independent claim are to be considered at least analogously to the features, advantages, and possible embodiments of the respective subject matter of the other independent claims, as well as to any possible combination of the subject matter of the independent claims, where appropriate in conjunction with one or more of the dependent claims.

[0008] The present invention relates to a method for generating training data for a software agent that is designed to control a driver assistance function of a motor vehicle. With the aid of this driver assistance function, the driver of the motor vehicle can be supported in specific driving situations, for example by controlling actuator systems of the motor vehicle. With the aid of this driver assistance function, the driver can be supported in the lateral and / or longitudinal guidance of the motor vehicle.

[0009] A software agent, also known as an agent or soft robot, is a computer program capable of performing specific, independent, and autonomous behaviors. This means that, depending on the state, it executes specific processes without requiring external initiation signals or external control intervention during the process. Software can be defined as an agent if it is autonomous, cognitive, communicative, modally adaptive, proactive, reactive, robust, and / or social. These characteristics describe the degree of autonomy of a computer program. Autonomy means that the software agent operates independently of user intervention. Cognitive means that the software agent is learnable and learns based on previous decisions or observations. Communication means that the software agent communicates its state to the surrounding environment as a function of its state. Modal adaptation means that the software agent changes its settings, particularly parameters and / or structure, based on its own state and the state of its surrounding environment. Proactive means that the software agent performs actions based on its own initiative. Reactive means that the software agent reacts to changes in its surrounding environment. Robust means that the software agent compensates for external and internal disturbances. Social is understood to mean that software agents communicate with other agents.

[0010] The method further provides for planning a trajectory for the motor vehicle using the driver assistance function for a determined scenario. Furthermore, it provides for planning longitudinal and / or lateral control arrangements for guiding the motor vehicle along the trajectory and for adjusting the actuator systems of the motor vehicle to implement the planned longitudinal and / or lateral control arrangements. In other words, the driver assistance function performs path planning for the motor vehicle and correspondingly adjusts the actuator systems influencing the steering, acceleration, and / or braking of the motor vehicle to control the motor vehicle along the planned trajectory or to assist in guiding the motor vehicle along the planned trajectory. The method further provides for determining the driver's reaction to the configured actuator systems. Thus, it is determined how the driver reacts to the steering, acceleration, and / or braking influenced by the actuator systems. Based on the determined driver reaction, it is determined to what extent the longitudinal and / or lateral control of the motor vehicle along the trajectory, as characterized by the reaction, differs from the longitudinal and / or lateral control planned by the driver assistance function, as characterized by the affected actuator systems, as characterized by a difference value. Therefore, based on the driver's reaction to the actuator system set by the driver assistance function, it is determined whether the driver agrees with the longitudinal and / or lateral control planned by the driver assistance function, as represented by the set actuator system. The smaller the difference value, the smaller the deviation between the longitudinal and / or lateral control of the vehicle planned by the driver, as represented by the driver's reaction, and the longitudinal and / or lateral control of the vehicle planned by the driver assistance function, as represented by the set actuator system. Therefore, when the difference value is small, it can be assumed that the driver of the vehicle agrees with the longitudinal and / or lateral control planned by the driver assistance function, as represented by the set actuator system. When the difference value is high, it is assumed that the driver disagrees with the longitudinal and / or lateral control of the vehicle planned by the driver assistance function, as represented by the set actuator system. In this case, the longitudinal and / or lateral control of the vehicle planned by the driver, as represented by the driver's reaction, deviates significantly from the longitudinal and / or lateral control of the vehicle planned by the driver assistance function, as represented by the set actuator system. The difference value thus indicates the extent to which the driver and the driver assistance function differ with respect to the planned control of the vehicle.

[0011] In this method, the driver's reaction to vehicle control assisted by a driver assistance function is determined. This means that the actuator systems of the vehicle are adapted by the driver assistance function in each case. A check of so-called "ground truth data" (in which the vehicle is controlled purely manually by the driver) is not absolutely necessary.

[0012] The method further provides for providing the determined scenarios and the determined assigned difference values ​​together as training data for the software agent. In the method, the driver of a motor vehicle is assisted in controlling the motor vehicle using a driver assistance function within a scenario by planning a trajectory for the vehicle and setting actuator systems for the lateral and / or longitudinal control of the vehicle based on the planned trajectory. The method also provides for assigning the determined difference values ​​in each determined scenario to the training data as so-called labels, allowing the software agent to be particularly well trained for identical or similar scenarios using each determined scenario. This generated training data enables the software agent to be trained so that the differences represented by the difference values ​​are particularly small. This means that the driver's assistance using the driver assistance function is particularly well adapted to the driver's individual wishes or expectations, and the driver's consent to the individual actions of the driver assistance function is selected as a target value. The software agent may, for example, include an artificial neural network that can be trained using the generated training data.

[0013] Furthermore, the present invention may include an electronic computing device configured to execute a method for generating training data for a software agent. The electronic computing device may be configured to control the driver assistance function and / or trigger settings of the actuator system. Furthermore, the electronic computing device may be configured to determine a driver's reaction and, depending on the determined reaction, to determine a difference represented by a difference value. Furthermore, the electronic computing device may be configured to generate training data from the determined situations and the determined assigned difference values ​​and to provide these data to the software agent.

[0014] In a possible refinement of the present invention, the actuator system is configured based on a preset intervention dominance of the driver assistance function. This intervention dominance characterizes the degree to which the driver assistance function implements its planned longitudinal and / or lateral control for the trajectory in response to the driver's goals and actions. This intervention dominance describes the intensity of the software agent's intervention. The higher the intervention dominance, the more strongly the driver assistance function implements the longitudinal and / or lateral control planned by the driver assistance function along the trajectory. At an intervention dominance of 0%, the driver manually controls the vehicle. At an intervention dominance of 100%, the driver has no influence on the control of the vehicle, so the vehicle is purely controlled by the driver assistance function. In this case, the driver is overridden by the driver assistance function. At a medium intervention dominance, the driver assistance function assists the driver to varying degrees. Currently, the preset intervention dominance is greater than 0%. The settings of the actuator system are adapted by the driver assistance function based on the preset intervention dominance to more or less strongly implement the longitudinal and / or lateral control planned by the driver assistance function along the trajectory in response to the driver's actions. For example, when the intervention dominance is low, the actuator system can assist the driver with steering by simply simplifying the application of steering torque to the steering system in one direction and making it more difficult in another. With high intervention dominance, the driver assistance function can directly influence the steering of the vehicle by applying a steering torque. By setting the actuator system based on the preset intervention dominance, it can be ensured that the driver assistance function adheres to the degree of assistance set by the driver.

[0015] In another possible embodiment of the present invention, the driver's reaction is determined based on the operation of the accelerator pedal and / or brake pedal and / or steering system. The driver's reaction is thus determined based on the influence on the driving direction and / or vehicle acceleration, and thus on the driver's influence on the longitudinal or lateral guidance of the vehicle. Depending on the extent to which the driver influences the longitudinal and / or lateral guidance of the vehicle, it is determined how the driver reacts to the adaptation of the driver assistance function actuator system. The driver's reaction can be determined particularly easily and reliably based on the operation of the accelerator pedal and / or brake pedal and / or steering system.

[0016] In another possible embodiment of the present invention, the driver's reaction is determined based on error-related brain potentials measured in the driver's brain. When a person recognizes an error during a task, an error-related brain potential can be measured in the person's brain as a reaction. These error-related brain potentials can be measured, for example, using an electrode cap worn by the driver. This means that, to determine the driver's reaction, these brain potentials are measured in the driver's brain when a driver assistance function assists in controlling the vehicle in a specific scenario. In other words, in this method, the actuator system of the vehicle is configured using the driver assistance function for the specific scenario to assist the driver in controlling the vehicle. Brain potentials in the driver's brain are also determined, and based on these potentials, the driver's reaction to the assistance provided by the driver assistance function is determined. Based on these brain potentials, it is particularly easy to identify whether the driver agrees or disagrees with the driver assistance function's support for lateral and / or longitudinal guidance, without the driver having to explicitly influence the lateral and / or longitudinal guidance of the vehicle. Therefore, the driver's unconscious consent or unconscious rejection of the driver's support of the longitudinal and / or lateral control of the motor vehicle by the driver assistance function can be determined particularly well based on these brain potentials.

[0017] In another possible embodiment of the present invention, it is provided that the difference value is compared with a preset threshold value. It is also provided that if it is determined that the difference value is less than the preset threshold value, the situation is assigned in the training data that there is no difference in the situation. If it is determined that the difference value is greater than or equal to the preset threshold value, the situation is assigned in the training data that there is a difference. This means that a label can be assigned to each situation in the training data, i.e., there is a difference or there is no difference. Therefore, each situation can be very simply marked as "Difference: Yes" or "Difference: No" in the training data. This makes it particularly easy to train the software agent with the help of the training data.

[0018] The present invention further relates to a method for training a software agent, particularly comprising an artificial neural network, using training data generated using the method described in conjunction with the present invention for generating training data. The training of the software agent allows the agent to control driver assistance functions in a particularly effective manner, tailored to the individual driver's wishes of the vehicle. This ensures a particularly high degree of driver consent to the driver assistance functions assisting with longitudinal and / or lateral control of the vehicle. Consequently, the driver of the vehicle is particularly likely to use the driver assistance functions when controlling the vehicle.

[0019] According to the present invention, an electronic computing device can also be provided that is configured to execute the method for training the software agent using the training data. The electronic computing device can thus be configured to receive the generated training data and train the software agent. The electronic computing device can, for example, be a hardware component on which the software agent can be implemented.

[0020] In a possible refinement of the method for training a software agent, the software agent is trained using reinforcement learning (also known as so-called reinforcement learning) by minimizing the frequency of identified differences or their value as a reward. This means that the reinforcement learning objective is to ensure that the value of the difference determined by the driver's reaction to the actuator system set by the driver assistance function is particularly low. Alternatively or additionally, the reinforcement learning objective can be such that, in a plurality of processes executed in different scenarios in which the actuator system of the motor vehicle is controlled by the driver assistance function controlled by the software agent, the frequency of individual scenarios in which differences are identified is particularly low compared to the frequency of scenarios in which no differences are identified. This means that the reinforcement learning objective is to ensure that, in a plurality of scenarios in which the driver of the motor vehicle is supported in controlling the motor vehicle by the driver assistance function controlled by the software agent, the proportion of scenarios in which differences are identified is particularly low.

[0021] Reinforcement learning, also known as RL, represents a family of machine learning methods in which software agents autonomously learn strategies to maximize the reward they receive. Currently, the reward is maximized when particularly low variance values ​​are reached or when there is a particularly small proportion of situations in which the corresponding variance has been determined. In reinforcement learning, the software agent is not shown which actions and situations are best. Instead, it receives a reward at a specific point in time through its interaction with its environment, which can also be negative. Currently, if the software agent is trained while controlling a motor vehicle, it can be continuously trained with newly generated training data, thereby further improving. Reinforcement learning methods enable driver assistance functions controlled by the software agent to be adapted particularly quickly and effectively to the individual wishes of the motor vehicle driver.

[0022] The present invention further relates to a method for assisting a driver in controlling a motor vehicle using a driver assistance function. In this method, a software agent trained using the method described in conjunction with the present invention for training a software agent is adapted to the driver assistance function. The adapted driver assistance function assists the driver in controlling the motor vehicle by configuring the vehicle's actuator systems for a specific situation. This is an end-to-end machine learning approach, in which situational camera data is provided to the software agent, which then directly outputs commands for the vehicle's actuator systems. This method enables the vehicle to be controlled using driver assistance functions that are particularly well adapted to the driver's wishes.

[0023] In a possible refinement of the method, the software agent is used to adapt the path planning and / or intervention control of the driver assistance function. The adapted driver assistance function assists the driver in controlling the motor vehicle by planning a trajectory for the motor vehicle within the scope of an adapted path planning for a specific situation, planning longitudinal and / or lateral control arrangements for guiding the motor vehicle along the trajectory, and setting actuator systems of the motor vehicle as a function of the adapted intervention control to achieve the planned longitudinal and / or lateral control. The software agent can perform the driver assistance function itself, or the driver assistance function can be performed by a driver assistance system that is distinct from the software agent and can be influenced by the software agent to adapt the path planning and / or intervention control, thereby enabling modification. Adapting the path planning and / or intervention control performed by the driver assistance function using a trained software agent allows the driver assistance function, or the support the driver receives from the driver assistance function in controlling the motor vehicle, to be particularly well adapted to the driver's wishes.

[0024] In this context, it can be provided, in particular, that the criticality of the identified situation is determined and that the intervention dominance is selected depending on the determined criticality of the situation. This means that the more critical the situation is assessed, the greater the intervention dominance of the driver assistance function is selected. Thus, if a critical situation is determined, the driver assistance function can intervene particularly strongly in the longitudinal and / or transverse guidance of the vehicle by setting the actuator system to avoid or at least reduce danger to the vehicle and its occupants. If the identified situation is determined to be particularly noncritical, a particularly low intervention dominance can be selected so that the transverse and / or longitudinal guidance of the vehicle is particularly less affected by the driver assistance function. As a result, in such particularly noncritical situations, the driver is given particularly great freedom in the longitudinal and / or transverse guidance of the vehicle. This can give the driver the feeling that he or she can particularly well influence the control of the vehicle and, therefore, has a particularly high degree of control over the vehicle's operation.

[0025] The present invention further relates to a motor vehicle with a driver assistance system configured to assist a driver in controlling the motor vehicle using a driver assistance function. The driver assistance function is adapted using the method described in conjunction with the present invention for assisting a driver in controlling a motor vehicle using a driver assistance function. The driver assistance system can be adapted by the trained software agent. Alternatively, the software agent can be the driver assistance system or a part thereof and configured to directly control the driver assistance function. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Further features of the invention can be derived from the following description of the drawings and from the drawings. The features and feature combinations mentioned above in the description and the features and feature combinations shown individually in the following description of the drawings and / or in the drawings can be used not only in the respectively specified combination but also in other combinations or alone without departing from the scope of the invention.

[0027] In the attached figure:

[0028] Figure 1 A first method diagram is shown of a method for generating training data for a software agent and a method for assisting a driver in controlling a motor vehicle by means of a driver assistance function;

[0029] Figure 2 A second method diagram shows a method for generating training data for a software agent and a method for assisting a driver in controlling a motor vehicle by means of a driver assistance function;

[0030] Figure 3shows a coordinate system in which the trajectory of the motor vehicle actually traveled is represented, said trajectory being in a state in which the driver of the motor vehicle is assisted in controlling the motor vehicle by means of a driver assistance function, wherein the trajectory planned for the motor vehicle by the driver assistance function is additionally plotted; and

[0031] Figure 4 Another coordinate system is shown, in which the trajectory planned by the driver assistance function for the motor vehicle and the trajectory from Figure 3 The actual trajectory of the motor vehicle is different.

[0032] Identical or functionally identical elements are provided with the same reference symbols in the figures. DETAILED DESCRIPTION

[0033] exist Figure 1 , a method diagram for generating training data 1 for a software agent 2 that is configured for controlling a driver assistance function 3 of a motor vehicle is shown. Furthermore, this method diagram presents a method for assisting a driver 4 in controlling a motor vehicle by means of the driver assistance function 3. The training data 1 can be generated during the assistance of the driver 4 in controlling the motor vehicle. Figure 2 , a method diagram for generating training data 1 for a software agent 2 is also shown, wherein the software agent 2 executes and thus controls a driver assistance function 3. Figure 2 The method diagram in FIG. 1 further includes a method for assisting a driver 4 in controlling a motor vehicle by means of a driver assistance function 3. The software agent 2 can be configured to Figure 1 Alternatively, the software agent 2 can be adapted as shown in Figure 2 3. Figure 1 and Figure 2 As can be seen from the method diagram in FIG, on the one hand a driver 4 can be assisted in controlling a motor vehicle by means of a driver assistance function 3 and at the same time training data 1 can be generated for a software agent 2 .

[0034] In order to assist the driver 4 in controlling the motor vehicle, a situation 5 in which the motor vehicle is located is determined. This determined situation 5 is provided to the driver assistance function 3. With the aid of the driver assistance function 3, a target trajectory 16 is planned for the motor vehicle within the scope of path planning 6 for the determined situation 5. Subsequently, the vehicle dynamics control 7 is carried out with the aid of the driver assistance function 3, within the scope of which longitudinal and / or lateral control arrangements for guiding the motor vehicle along the target trajectory 16 are planned. In this case, for example, a steering angle can be calculated for the planned target trajectory 16. The driver assistance function 3 is also configured to subsequently carry out an actuator system control 8. Within the scope of the actuator system control 8, the calculated steering angle is implemented with the aid of the actuator system of the motor vehicle. Within the scope of the actuator system control 8, the actuator system of the motor vehicle is set to implement the planned longitudinal and / or lateral control. Due to the adaptation of the actuator system, the lateral and / or longitudinal control of the motor vehicle can be influenced by influencing the individual control devices 9 of the motor vehicle. At the same time, driver 4 may have an image in his head of how the trajectory should unfold, along which the vehicle should be guided in the defined scenario 5, and how the longitudinal guidance of the vehicle should unfold along this trajectory. Driver 4 can thus influence the lateral and / or longitudinal guidance of the vehicle. This influence of driver 4 on the lateral and / or longitudinal guidance of the vehicle in turn results in an actual trajectory 10 traveled by the vehicle at a defined speed profile. Based on the settings of control devices 9 generated in the vehicle, such as the accelerator pedal and / or brake pedal and / or steering, and based on the actual trajectory 10 actually traveled by the vehicle in the defined scenario 5, it can be determined whether there is a discrepancy 11 between the longitudinal and / or lateral control of the vehicle along a target trajectory 16 planned by driver assistance function 3 and the driver's desired behavior. This discrepancy 11 is determined by driver 4's reaction to actuator system adjustments 8 performed by driver assistance function 3.

[0035] Within the scope of actuator system control 8 , the driver assistance function 3 can, for example, adjust individual actuators to make it easier for the driver 4 to turn the steering wheel in a first direction, while more difficult to turn in a second direction opposite the first. Thus, the driver assistance function 3 can encourage the driver 4 to steer in the first direction so that the vehicle is guided as closely as possible along the target trajectory 16 planned by the driver assistance function 3 . To assist the driver 4 in longitudinal control, the driver assistance function 3 can, for example, adjust at least one actuator to make it easier for the driver 4 to press the accelerator pedal, while resisting when pressing the brake pedal. This encourages the driver 4 to accelerate the vehicle and follow the trajectory more quickly. Through actuator system control 8 , the driver assistance function 3 can thus encourage the driver 4 to control the vehicle according to the longitudinal and / or lateral guidance planned by the driver assistance function 3 . By appropriately operating the steering wheel, accelerator pedal, or brake pedal, the driver 4 can control the vehicle according to the longitudinal and / or lateral guidance planned by the driver assistance function 3 , or counteract the longitudinal and / or lateral guidance planned by the driver assistance function 3 .

[0036] To determine the difference 11, the driver's 4 reaction to the actuator system set by the driver assistance function 3 is determined. In this case, the driver's 4 reaction is determined based on the operation of the control device 9 by the driver 4 and / or based on the actual trajectory 10 traveled by the motor vehicle in the determined scenario 5. The difference 11 can be determined, for example, based on the deviation of the actual trajectory 10 actually traveled by the motor vehicle from the target trajectory 16 planned by the driver assistance function 3. The difference 11 can be determined, in particular, in the form of a difference value that characterizes the difference 11. The difference 11 describes the deviation between the longitudinal and / or lateral control planned by the driver assistance function 3 and the longitudinal and / or lateral control planned by the driver 4.

[0037] In this method, the extent of the discrepancy 11, represented by the discrepancy value between the longitudinal and / or lateral control planned by the driver 4 and characterized by the reaction, and the longitudinal and lateral control of the motor vehicle planned by the driver assistance function 3 along the planned target trajectory 16 and characterized by the affected actuator system, is determined based on the determined reaction of the driver 4. The scenario 5 and the determined assigned discrepancy value are provided together as training data 1 for the software agent 2. The software agent 2 can in turn be trained using the training data 1, in particular using reinforcement learning. In this training method, the predetermined goal is that the individual discrepancy values ​​determined for the further scenarios 5 should be particularly low, or that the presence of discrepancies 11 should be determined as infrequently as possible for the further examined scenarios 5.

[0038] Currently, it is set up so that the reaction of the driver 4 is additionally determined based on the measured error-related brain potential 12. This brain potential 12 can be determined with the help of electrodes fixed to the head of the driver 4. In particular, it is determined how the driver's 4 brain potential 12 reacts to the actual longitudinal guidance and / or lateral guidance of the motor vehicle in the determined situation 5. Based on the determined error-measured brain potential 12, the difference 11 characterized by the difference value can be determined. The situation 5 can be provided together with the difference value determined based on the brain potential 12 as training data 1 for the software agent 2. Error-related brain potentials, which can also be called error potentials, are automatically generated by the brain when the real world deviates from its own expectations. These brain potentials can be measured by a brain-computer interface. If the driver assistance function 3 does not do what the driver 4 expects, a passive but measurable error potential signal can be measured.

[0039] Within the scope of path planning 6 , a safety prediction 13 can be performed. Within the scope of this safety prediction 13 , the criticality of situation 5 is determined in conjunction with the target trajectory 16 determined within the scope of path planning 6 . This means that the likelihood of damage, in particular a collision of the motor vehicle with another object, in situation 5 is determined, and whether this risk can be avoided if the motor vehicle follows the target trajectory 16 created within the scope of path planning 6 . Depending on the criticality of this determined situation 5 , the intervention dominance 14 of the driver assistance function 3 is adapted. This intervention dominance 14 determines to what extent the control of the motor vehicle should be influenced by the driver assistance function 3 and to what extent the control of the motor vehicle should be influenced by the driver 4 . Depending on the set intervention dominance 14 , the actuator systems set by the actuator system control 8 can be adapted by means of an adaptation control 15 , so that the longitudinal and / or transverse guidance of the motor vehicle is more or less influenced by the driver assistance function 3 in accordance with the level of intervention dominance 14 .

[0040] exist Figure 1 In the method shown in FIG, a trained software agent 2 is used to adapt the path plan 6, the safety prediction 13, and / or the intervention dominance 14. By adapting the safety prediction 13, the software agent 2 can specify which types of situations 5 are to be rated as highly critical. By influencing the intervention dominance 14, the software agent 2 can set the dominance of the driver assistance function 3 relative to the driver 4, given the determined criticality of the respective situation 5.

[0041] The software agent 2 trained on the basis of the training data 1 is then adapted to the path planning 6 and / or intervention control 14 of the driver assistance function 3. The driver 4 is then assisted in controlling the motor vehicle by means of the adapted driver assistance function 3. In this assistance, the driver assistance function 3 plans a target trajectory 16 for the motor vehicle within the scope of the adapted path planning 6 for the determined situation 5, plans a longitudinal control and / or lateral control for guiding the motor vehicle along the target trajectory 16, and, to achieve the planned longitudinal control and / or lateral control, adjusts the actuator systems of the motor vehicle by means of the actuator system control 8 as a function of the adapted intervention control 14.

[0042] In order to assign a particularly simple label to each scenario 5 in the training data 1, it can be provided that the determined difference value characterizing the difference 11 is compared with a predefined threshold value, and if the difference value is found to be less than the predefined threshold value, the absence of the difference 11 in this scenario 5 is assigned to the scenario 5 in the training data 1. If the difference value is found to be greater than or equal to the predefined threshold value, the presence of the difference 11 is assigned to the scenario 5 in the training data 1. Therefore, in the training data 1, each scenario 5 is simply labeled "Difference: Yes" or "Difference: No."

[0043] exist Figure 2 In the embodiment of the method shown in FIG, the software agent 2 is configured to implement the driver assistance function 3. In this case, the software agent 2 is similar to the one already described in connection with FIG. Figure 1 The described method is trained on the basis of training data 1 in which the determined differences 11 belonging to the respective determined situations 5 are assigned as labels in the form of difference values ​​or in the form of "difference exists" / "difference does not exist". Figure 2 In the method shown in FIG, it is provided that path planning 6, vehicle dynamics control 7, and actuator system control 8 are performed by means of a software agent 2. In addition, a safety prediction 13 and an adaptation of an intervention initiative 14 can be performed by means of the software agent 2, and the actuator systems specified by the actuator system control 8 can be adapted by means of an adaptation control 15 as a function of the intervention initiative 14.

[0044] The method for assisting driver 4 in controlling a motor vehicle can be used, for example, within the scope of a lane keeping assist system as driver assistance function 3. Driver assistance function 3 can thus apply a steering torque, wherein driver 4 steers the motor vehicle.

[0045] exist Figure 3In FIG. 1 , the actual trajectory 10 driven by the motor vehicle is shown along the curve, compared to the target trajectory 16 planned by the driver assistance function 3 . It can be seen that the actual trajectory 10 actually driven by the motor vehicle deviates from the planned target trajectory 16 . Consequently, there is a discrepancy 11 between the lateral control of the motor vehicle planned by the driver assistance function 3 , as expressed by the target trajectory 16 , and the lateral guidance of the motor vehicle desired by the driver 4 , which, in combination with the actuator system set by the driver assistance function 3 , results in the actual trajectory 10 .

[0046] exist Figure 4 In [ ], the assistance torque y applied by the driver assistance function 3 is plotted on the ordinate, and the driver torque x applied by the driver 4 is plotted on the abscissa, both in percentages. Each torque is expressed as a percentage and normalized to the input value and, therefore, to the measurement range of the respective sensor recording it. This coordinate system contains two convergence zones 17 and two divergence zones 18. The convergence zones 17 characterize situations in which the torque applied by the driver 4 is identical or similar to the torque applied by the driver assistance function 3. In each divergence zone 18, the torque applied by the driver 4 deviates from the torque applied by the driver assistance function 3. Depending on whether the measured values ​​measured at the various sensors at a certain point in time result in a point in the diagram that lies in one of the difference zones 18 or in one of the convergence zones 17, it can be determined that a difference 11 exists if the point lies in one of the difference zones 18, and that a difference 11 does not exist if the point lies in one of the convergence zones 17. In this correlation analysis between driver torque x and assistance torque y, the deviation of actual trajectory 10 from target trajectory 16 planned by driver assistance function 3 can also be taken into account. In particular, the direction of the deviation can also be taken into account.

[0047] Discrepancies 11 between driver 4 and driver assistance function 3 arise particularly during path planning 6 and / or actuator system control 8. Within the scope of this method, conflicts that are measurable in physical signals are measured and used to adapt driver assistance function 3. Many conflicts can arise within the scope of path planning 6 and actuator system control 8, for example, due to deviating plans or unfavorable intervention strategies. These discrepancies 11 can arise, for example, if driver 4 perceives the behavior of driver assistance function 3 as undesirable or unexpected, or if driver assistance function 3 plans the same trajectory but the execution appears to driver 4 to be jarring, for example, too rigid.

[0048] As additional statistical analysis possibilities for evaluating the differences 11 , mean values ​​and covariance matrices can be used, in particular for supervised learning methods for subsequent analysis and adaptation, or as episode rewards for reinforcement learning agents.

[0049] Since the introduction of assistance systems, research has been underway to reduce the discrepancy between the driver's 4 driving style and the system's situational decisions. 11 This is often done by directly inputting the driver's wishes, or by inferring the user configuration using a larger amount of data. This has the disadvantage that manual parameterization is only feasible for a very small number of easily understood parameters (e.g., the time gap in a distance control system). Furthermore, the driver's observations used to generate the data for the user configuration are only meaningful during manual driving, as the driving styles of the driver 4 and the driver assistance function 3 are intertwined during assisted driving. It is unclear whether the data from manual driving actually corresponds to the driver's 4 desire to be assisted while driving. Just as a passenger might not want to be driven as they would if they were driving themselves, so too might the situation in assistance mode where one might not want to be assisted using their own driving style.

[0050] If the assistance intervention differs from the personal or subjective decision of the driver 4 involved, this can lead to a negative evaluation or even the deactivation or non-use of the driver assistance function 3. If the driver 4's behavior differs from that of the driver assistance function 3, their preferred actions on the steering wheel, accelerator, and brake pedal will differ. The method for generating training data 1 provides for quantifying these differences 11 between the driver 4 and the driver assistance function 3 as feature values ​​(currently, as difference values) and using these differences 11 as input for a machine learning method. Both supervised and reinforcement learning methods can use the features of these differences 11 as labels or as negative reward functions to gradually adapt the actions of the driver assistance function 3 so that they result in fewer and less frequent deviations 11 from the driver's wishes, while still encouraging a change or improvement in the driver's 4 driving style by taking into account the system wishes of the driver assistance function 3. This enables automated adaptation of the driver assistance function 3, which also eliminates the need for manual driving to acquire data and allows learning to continue even after the driver 4 has already given up the assistance function. The latter is possible because the software agent 2 does not learn from the observed driving style (which, in assistance mode, is a mixture of the driver 4 and the driver assistance function 3), but rather from the recognition of differences 11 between the driver 4 and the driver assistance function 3, or even from the absence of differences 11. It is thus possible to adapt the driver assistance function 3 step by step to the ideal assistance conception of the driver 4 using this method.

[0051] The difference 11 can be determined between specific activities on the steering wheel, accelerator pedal and brake pedal, either as a direct difference value, or as an indirect or non-causal characteristic value, as a difference 11 between the driver assistance function 3 and the perception of the driver 4 .

[0052] Through actions on the vehicle interface, driver assistance function 3 and driver 4 can perceive each other. The discrepancy 11 between driver 4's actions and the system actions of driver assistance function 3 arises primarily from different assumptions about the future. Driver assistance function 3 introduces additional interaction elements, which can be evaluated positively or negatively, due to the different control parameters implemented in actuator system control 8. An example of this is the maximum torque that driver assistance function 3 can utilize to assist driver 4.

[0053] Difference 11 can be calculated based on characteristic values ​​based on the final vehicle action. By building correlations and covariances for longer measurement sequences, additional characteristic values ​​can be calculated that can be used to evaluate each point in time for adaptation. Based on the correlation between the driver's steering torque and the auxiliary steering torque, it can be distinguished whether driver 4 is simply passively holding the steering wheel or actively counteracting the action suggested by driver assistance function 3. An actively intervening driver 4 can confirm the actions of driver assistance function 3 or counteract them.

[0054] There are different ways to calculate the difference features. These can also be called reverse acceptance features or contradiction determinations. These difference features can be adapted to any configuration of an automated system, in particular also to systems that consist entirely of artificial neural networks and no longer have the same Figure 2 . Error-related brain potentials can be introduced as additional markers in the system adaptation and thus used to determine the labels of the individual situations 5 in the training data 1. The software agent 2 can undergo pre-training based on data collected when the motor vehicle is manually driven by the driver 4. The influence of the adaptation and thus the influence of the software agent 2 on the adaptation of the driver assistance function 3 can be clearly recorded and limited in the method. In particular, it is provided that the software agent 2 only adapts the parameterization of the algorithm underlying the driver assistance function 3. For example, the software agent 2 can learn that if it never intervenes, no differences 11 will occur. Therefore, it can be predefined for the adaptation of the driver assistance function 3 or the control of the motor vehicle with the aid of the driver assistance function 3 that the intervention of the driver assistance function 3 within the scope of the actuator system regulation 8 cannot fall below a predefined limit value in order to prevent this.

[0055] This method makes it possible for the driver assistance function 3 to be continuously adapted to the driver 4 and without active programming by the driver 4 and to become increasingly better during the ongoing operation of the motor vehicle.

[0056] Overall, the present invention shows how differentiating features can be used to adapt a learning algorithm for a driver assistance function.

[0057] Reference Signs List

[0058] 1 Training Data

[0059] 2 Software Agents

[0060] 3 Driver Assistance Features

[0061] 4 Driver

[0062] 5 Situations

[0063] 6 Path Planning

[0064] 7 Driving dynamics control

[0065] 8 Actuator system adjustment

[0066] 9 Control Equipment

[0067] 10 Actual trajectory

[0068] 11 Differences

[0069] 12 Error-related brain potentials

[0070] 13 Security Predictions

[0071] 14 Intervention Dominance

[0072] 15 Adaptation

[0073] 16 Target trajectory

[0074] 17 Convergence Zone

[0075] 18 Difference Area

Claims

1. A method for generating training data (1) for a software agent (2) which is designed to control a driver assistance function (3) of a motor vehicle, wherein: planning a trajectory (16) for the motor vehicle for a determined situation (5) by means of the driver assistance function (3), planning a longitudinal control and / or a transverse control of a device for guiding the motor vehicle along the trajectory (16), and setting an actuator system of the motor vehicle to implement the planned longitudinal control and / or transverse control, determining a driver's (4) reaction to said set actuator system, Ascertaining, based on the determined reaction, how great the difference (11) is between the longitudinal and / or lateral control planned by the driver and characterized by the reaction and the longitudinal and / or lateral control planned by the driver assistance function (3) and characterized by the affected actuator system, characterized by a difference value, The context and the determined assigned difference values ​​are provided together as training data (1) for the software agent (2).

2. The method according to claim 1, wherein The actuator system is set as a function of a predefined intervention dominance (14) of the driver assistance function (3).

3. The method according to claim 1 or 2, wherein: The driver's (4) reaction is determined based on the operation of the accelerator pedal and / or the brake pedal and / or the steering device.

4. A method according to any one of the preceding claims, wherein The driver's (4) reaction is determined based on error-related brain potentials (12) measured in the driver's (4) brain.

5. A method according to any one of the preceding claims, wherein The difference value is compared with a preset threshold value, and if it is found that the difference value is less than the preset threshold value, the situation (5) is assigned in the training data (1) that there is no difference (11) in the situation (5), and if it is found that the difference value is greater than or equal to the preset threshold value, the situation (5) is assigned in the training data (1) that there is a difference (11).

6. A method for training a software agent (2) with the aid of training data (1) generated in a method according to any one of the preceding claims.

7. The method according to claim 6, wherein: The software agent (2) is trained by means of reinforcement learning in that the frequency of the determined differences (11) or the difference value is used as a reward and is to be minimized.

8. A method for assisting a driver (4) when controlling a motor vehicle by means of a driver assistance function (3), in which method a driver assistance function (3) is adapted by a software agent (2) that has been trained in the method according to claim 6 or 7, and the driver (4) is assisted when controlling the motor vehicle by means of the adapted driver assistance function (3), in that an actuator system of the motor vehicle is set for a determined situation (5) by means of the driver assistance function (3).

9. The method according to claim 8, wherein The software agent (2) adapts the path planning (6) of the driver assistance function (3) and / or the intervention dominance (14) of the driver assistance function (3), and assists the driver (4) in controlling the motor vehicle with the aid of the adapted driver assistance function (3), in that the driver assistance function (3) plans a trajectory (16) for the motor vehicle within the scope of the adapted path planning (6) for a determined situation (5), plans longitudinal control and / or lateral control of a device for guiding the motor vehicle along the trajectory (16), and sets the actuator system of the motor vehicle as a function of the adapted intervention dominance (14) for implementing the planned longitudinal control and / or lateral control.

10. A motor vehicle having a driver assistance system, which is designed to assist a driver in controlling the motor vehicle by means of a driver assistance function (3) adapted in a method according to claim 8 or 9.

Citation Information

Patent Citations

  • Self-adaptive cruise system with driving style learning capacity and implementation method

    CN109927725A

  • Method for selecting a driving profile of a motor vehicle, driver assistance system and motor vehicle

    DE102018202146A1

  • Learning method and learning device for supporting reinforcement learning by using human driving data as training data to thereby perform personalized path planning

    EP3690769A1