Selection of driving operations for a vehicle that travels at least partially automatically
The method uses a machine learning model to suppress non-permissible driving operations by setting their probabilities to zero in the probability distribution, ensuring safe and adaptive driving behaviors without the need for extensive retraining, addressing the challenge of complex boundary conditions and new regulations in automated driving.
Patent Information
- Application Number
- JP2023533896
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-03
- Filing Date
- 2021-11-30
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2041-11-30
AI Technical Summary
Existing methods for automated driving operation planning in vehicles struggle to effectively suppress non-permissible driving operations, especially when faced with complex boundary conditions or new traffic regulations, and may require extensive retraining when such conditions change.
A method that uses a trained machine learning model to create a situation representation, maps it to a probability distribution, and suppresses non-permissible driving operations by setting their probabilities to zero or modifying the distribution to ensure only permissible operations are selected, allowing for flexible adaptation to changing regulations without extensive retraining.
Ensures more realistic and safe driving behaviors by preventing non-permissible actions, adapting to new traffic rules without repeated training, and maintaining safe driving operations even when the model is updated.
Smart Images

Figure 0007705937000001 
Figure 0007705937000002
Abstract
Description
Technical Field
[0001] The present invention relates to situation-dependent driving operation planning for a vehicle that travels at least partially automatically.
Background Art
[0002] Background Art A vehicle that travels at least partially automatically continuously captures the situation in order to adapt its relatively imminent future driving operation planning to changes in the situation in which it is placed. Changes in the situation to which the vehicle must react can be caused, for example, by the vehicle moving to another location with a different given state. However, the movement of other objects, such as other road users, can also significantly change the situation, which may require a reaction. German Patent Application Publication No. 10 2018 210 280 discloses a method that can predict the trajectories of other objects, thereby enabling the trajectory of the host vehicle to be adapted accordingly.
[0003] Some methods for driving operation planning create a representation of the situation in which the vehicle is placed and map this representation, by means of a trained machine learning model, to a probability distribution representing the probabilities of basically usable driving operations. From this probability distribution, one driving operation is derived as the driving operation to be executed, and the operation of the actuator device of the vehicle is controlled accordingly.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Means for Solving the Problems
[0005] Disclosure of the Invention Within the framework of the present invention, a method has been developed for selecting driving operations to be performed by at least partially automated vehicles. The method starts with creating a representation of the situation in which the vehicle is located by using measurement data from at least one sensor mounted on the vehicle. This representation of the situation is mapped to a probability distribution by a trained machine learning model. The representation can be, in particular, a comprehensive representation of the situation created in any form and manner.
[0006] The measurement data can be, in particular, for example, image data, video data, radar data, LIDAR data and / or ultrasonic data.
[0007] A machine learning model is considered, in particular, to be a model that embodies a function parameterized with adaptable parameters having excellent generalization ability. When training the machine learning model, the parameters can be adapted, in particular, as follows. That is, when the learning representations are input into the model, they are adapted so that the pre-known target outputs corresponding to those learning inputs are reproduced as well as possible. The machine learning model can, in particular, include an artificial neural network KNN, and / or the machine learning model can be an artificial neural network KNN.
[0008] The probability distribution represents the probability that each driving operation belonging to a predetermined catalog of available driving operations is executed. From the probability distribution, one driving operation is derived as the driving operation to be executed.
[0009] Additionally, by using at least one aspect of the situation in which the vehicle is located, a subset of driving operations not allowed in this situation is determined. The execution of these non-allowed driving operations is suppressed.
[0010] The training of the machine learning model is adjusted to separate, in individual situations, driving operations that are more suitable for the purpose from those that are less suitable for the purpose. Therefore, most of the non-permissible driving operations are already regarded by the machine learning model as being less suitable for the purpose. By deriving the driving operation to be ultimately executed from a probability distribution, a more realistic driving behavior, especially one that does not surprise other road users as much, is brought about by the machine learning model compared to directly mapping the representation to exactly one driving operation. However, it is impossible to avoid that non-permissible driving operations are also assigned a probability different from zero in the probability distribution. What this means is that non-permissible driving operations are actually selected and executed with a certain probability. This probability may be higher than the acceptable residual risk specified for the automated driving operation.
[0011] Furthermore, the boundary conditions that make a specific driving operation not permissible in a particular situation can be relatively complex and / or may be characteristics that are not suitable for incorporation into the training of the machine learning model. The strength of the machine learning model lies precisely in its ability to generalize from a limited amount of training situations to an unspecified large number of situations. However, for example, if it is stipulated by the conditions determined for the automated driving operation that a specific driving operation may only be executed within a specific speed range or that an overtaking operation may only be initiated when there is a defined minimum distance to oncoming traffic approaching, the above-mentioned generalization ability is not the optimal means. Rather, this type of boundary condition is more advantageously implemented without detouring through the machine learning model.
[0012] Therefore, the suppression of unacceptable driving operations can be carried out completely independently of the machine learning model. What this means is that the machine learning model can be trained first, separated from the acceptance or rejection of individual driving operations, and the acceptance or rejection is only implemented later. Thus, even if the regulations regarding acceptance or rejection are changed later, they will no longer affect the training of the machine learning model. That is, such changes can be made multiple times without the need to completely or partially repeat the training. For the adaptation of training, usually, test drives in which the representation of specific new situations is recorded must be completed. Then, for those representations, the desired driving operations in each situation must be manually labeled.
[0013] For example, such effort is not required to adapt the speed range in which driving operations are allowed. In particular, new rules such as releasing the hard shoulder as a lane in a severely congested highway section can be easily implemented by declaring that a lane change to the hard shoulder, which is usually rejected as not allowed, is allowed when appropriately released.
[0014] In a particularly advantageous embodiment, the execution of at least one unacceptable driving operation is suppressed by setting the probability of its execution to zero in the probability distribution. This generates a modified probability distribution. In this way, the following is guaranteed. That is, when deriving from the probability distribution, only acceptable driving operations can be selected, and at the same time, if there are no available results, the derivation is not maintained as it is. Therefore, for cases where unacceptable driving operations are derived, individual "error handling" is not required.
[0015] In particular, for example, after at least one probability is set to zero, the probability distribution can be normalized such that the probabilities of driving operations that are still different from zero sum up to one. This also generates a modified probability distribution. That is, a driving operation that was discarded as unacceptable abandons its previous probability of being selected and distributes that probability to the remaining driving operations that remain as acceptable according to their shares. This is somewhat similar to the situation where, if one of multiple applicants for one residence or job fails the disqualification criteria, the chance of that applicant up to that point is redistributed to the remaining applicants. That is, even if this one applicant fails, it remains certain that (with probability 1) the residence or job will be given to someone.
[0016] Instead of or in combination with changing the probability distribution retrospectively, the execution of at least one unacceptable driving operation can be suppressed such that, in response to the unacceptable driving operation being derived from the probability distribution, a new driving operation is derived from the probability distribution. Since the probability distribution assigns higher probabilities to acceptable driving operations, it can be expected but not guaranteed that an acceptable driving operation will be selected for the new derivation. In some cases, it may be necessary to repeat the new derivation until an acceptable driving operation is obtained as a result.
[0017] The advantage of repeating the derivation is that it can capture and evaluate situations where an unacceptable driving operation was initially selected. If such situations occur frequently, this can be used as an indicator that the machine learning model no longer captures the situation appropriately and its training needs to be adapted accordingly. One possible reason for this could be that new rules regarding road traffic were introduced after the training of the machine learning model.
[0018] For example, the traffic signs newly introduced for the environmental zone are generally based on the traffic signs indicating the start of a 30 km / h zone. Here, the only difference is that the number "30" in the red circle is replaced by the word "environment". If the machine learning model can only recognize the signs of "30 km / h zones", there may be a problem with the recognition rate that is desirable for human drivers. For example, if a highway with a speed limit of 80 km / h leads into the environmental zone, the machine learning model may identify it as a speed limit of 30 km / h and accordingly recommend driving operations suitable for emergency braking. If the resulting emergency braking exceeds the maximum deceleration allowed for the automated driving operation, the brake is discarded as not allowed and not executed. In this way, if it is captured that the same driving operations have been discarded as not allowed at the same location many times during the driving operation, the vehicle user will receive feedback that something fundamental is not functioning properly and the machine learning model needs to be updated.
[0019] In yet another particularly advantageous embodiment, at least one non-permissible driving operation is determined based on information called from a spatially resolved digital map based on the current position of the vehicle. The digital map can record, in particular, for example, the course of the road, the number of lanes provided for each driving direction, speed limits, overtaking bans, and other traffic rules. For example, if it is determined from the map that there is no adjacent lane to the left or right of the current lane, or there is no adjacent drivable area at all, a lane change to the left or right can be evaluated as not allowed.
[0020] Therefore, non-permissible driving operations are, in particular, for example, the following, namely, · Deviation from the lane, and / or · Violation of general traffic rules, and / or · Damage to specific mechanisms for automated driving operations, and / or · A driving operation that poses a risk of collision between the host vehicle and another vehicle or other object can be performed.
[0021] Thus, unacceptable driving operations can include, for example, the following, namely, · A lane change that results in deviation from the lane, and / or · A lane change to a lane that is currently inaccessible, and / or · An acceleration operation and / or overtaking operation prohibited by traffic rules, and / or · Following behind another vehicle that is currently located behind the host vehicle may be included.
[0022] As described above, filtering out and excluding unacceptable driving operations can be performed separately from the machine learning model, which initially has complete freedom to propose any driving operation that is basically usable. However, knowledge of which driving operations are unacceptable can also be incorporated into the training of the machine learning model in addition to this.
[0023] Accordingly, the present invention also relates to a method for training a machine learning model. This machine learning model maps a representation of the situation in which the vehicle is placed to a probability distribution, and this probability distribution represents the probability that each driving operation belonging to a predetermined catalog of available driving operations is executed.
[0024] Within the framework of this method, a learned representation of the situation and a corresponding target probability distribution to which these learned representations are mapped by the machine learning model are prepared. The learned representations are supplied to the machine learning model and mapped to a probability distribution by this machine learning model. The agreement between this probability distribution and the individual target probability distributions is evaluated based on a predetermined cost function. The parameters characterizing the behavior of the machine learning model are optimized for the purpose of increasing the evaluation by the cost function by further processing the learned representations.
[0025] In this case, with respect to at least one driving operation that is not allowed in a situation characterized by a learned representation, an increase in the probability assigned to this driving operation is suppressed from increasing the evaluation by the cost function.
[0026] What this means is that the machine learning model cannot bring advantages in terms of the evaluation by the cost function by proposing a driving operation that is not allowed. Therefore, in order to increase the evaluation, the machine learning model has to consider increasing the probabilities of other driving operations. This does not rule out the fact that a driving operation that is not allowed is still assigned a probability different from 0 in the probability distribution. However, clearly preferably, only the probabilities of allowed driving operations are increased.
[0027] In an advantageous embodiment, the cost function is extended by a penalty term that absorbs and / or overcompensates for the gain that would be obtained in relation to the original cost function by an increase in the probability assigned to a driving operation that is not allowed. In this case, similar to criminal law where even a very severe crime cannot be completely prevented by the threat of punishment, the fact that a driving operation that is not allowed is still assigned a probability different from 0 is not excluded. However, instead of only increasing the probabilities of allowed driving operations, a strong incentive is provided.
[0028] Alternatively or in combination therewith, the probability assigned to a non - allowed driving operation supplied from a machine learning model can be set to zero before evaluation by a cost function. The increase in this probability is not explicitly penalized through this, but also no longer has any effect during optimization. During the training process, the machine learning model learns that the probability of a non - allowed driving operation no longer has an effect on the optimization of the cost function, even if it changes. In this case, the corresponding attempts are abandoned so as to be advantageous for the optimization of the probability of allowed driving operations. Such behavior is somewhat comparable to a "time - out chair", where a child who throws a tantrum to force attention is made to sit and is not given attention, thereby making the child understand reason.
[0029] In yet another advantageous embodiment, the probability distribution is regularized and / or discretized such that probabilities below a predetermined threshold are reduced to zero. In this way, it is possible to provide the guarantee that no non - allowed driving operations are brought about by derivation from the probability distribution. For this purpose, for example, the L1 norm can be used.
[0030] This method can in particular be implemented, for example, on one or more computers, and to that extent can be embodied as software. Thus, the present invention also relates to a computer program, which comprises machine - readable instructions for causing one or more computers to carry out one of the aforementioned methods when it is executed on one or more computers.
[0031] Similarly, the present invention also relates to a machine-readable data carrier and / or a download product comprising a computer program. The download product is a digital product that can be transmitted via a data network, i.e., can be downloaded by a user of the data network, and this digital product can be sold, for example, in an online store for immediate download.
[0032] Furthermore, a computer can be provided with the above computer program, machine-readable data carrier or download product.
[0033] Further measures for improving the present invention will be shown in detail below together with the description of the preferred embodiments of the present invention based on the drawings.
Brief Description of the Drawings
[0034]
Figure 1
Figure 2
Modes for Carrying Out the Invention
[0035] Examples FIG. 1 is a schematic flowchart of one embodiment of method 100. The purpose of method 100 is to select an operation 4 to be executed that is adapted to the current situation 60 of vehicle 50 from a predetermined catalog consisting of operations 3a-3f.
[0036] For this purpose, in step 110, using the measurement data 51a of at least one sensor 51 mounted on the vehicle, a representation 61 of the situation 60 in which the vehicle 50 is placed is created. This representation 61 of the situation 60 is mapped to a probability distribution 2 by the trained machine learning model 1 in step 120. The probability distribution 2 represents, for each of the driving operations 3a to 3f belonging to a predetermined catalog of available driving operations 3a to 3f, the probabilities 2a to 2f that the driving operations 3a to 3f are executed.
[0037] Already at this point, in step 130, using at least one aspect 62 of the situation 60, the driving operation 3a * ~3f * that is not allowed in this situation 60 is to be determined as a subset. Subsequently, in step 140, it is assumed that the execution of these non-allowed driving operations 3a * ~3f * is suppressed. For this purpose, in particular, according to block 141, in the probability distribution 2, the probabilities 2a to 2f that the non-allowed driving operations 3a * ~3f * are executed are set to zero. After being set to zero in this way, according to block 142, it is assumed that the probability distribution is normalized so that the probabilities 2a to 2f of the driving operations 3a to 3f that are still different from zero sum to 1. Thereafter, a modified probability distribution 2' is generated.
[0038] In step 150, one of the driving operations 3a to 3f is derived as the driving operation 4 to be executed from this modified probability distribution 2' or from the original probability distribution 2.
[0039] At this point, an intervention process may also be performed to filter out and exclude the non-allowed driving operations 3a * ~3f * For this purpose, in step 160, similar to step 130, using at least one aspect 62 of the situation 60, the non-allowed driving operations 3a * ~3f *shall be required. Subsequently, these unacceptable driving operations 3a * ~3f * can be suppressed in step 170. For this purpose, in particular, for example, in response to the unacceptable driving operation 3a * ~3f * being derived from probability distribution 2, new driving operations 3a to 3f can be derived from probability distribution 2. Thus, finally, instead of the originally derived driving operation 4, a newly derived driving operation 4' occurs.
[0040] The unacceptable driving operations 3a * ~3f * can be obtained respectively based on information called from a digitally mapped space-decomposed map based on the current position of the vehicle 50, in particular, for example, according to block 131 or 161.
[0041] Figure 2 is a schematic flowchart of one embodiment of a method 200 for training the machine learning model 1. The machine learning model 1 maps a representation 61 of the situation 60 in which the vehicle 50 is located to a probability distribution 2. This probability distribution 2 represents, for each of the driving operations 3a to 3f belonging to a predetermined catalog of available driving operations 3a to 3f, the probabilities 2a to 2f that the driving operations 3a to 3f are executed.
[0042] In step 210, a learned representation 61a of situation 60 and a corresponding target probability distribution 2a are prepared, and the machine learning model 1 is to map the learned representation 61a to the target probability distribution 2a as described above. In step 220, the learned representation 61a is supplied to the machine learning model 1 and mapped to the probability distribution 2 by the machine learning model 1. In step 230, the agreement between these probability distributions 2 and the individual target probability distributions 2a is evaluated based on a predetermined cost function 5. In step 240, the parameter 1a characterizing the behavior of the machine learning model 1 is optimized. The purpose of this optimization is to increase the evaluation 230a by the cost function 5 by further processing the learned representation 61a. The state in which the training of the parameter 1a is completed is represented by reference numeral 1a * is represented by
[0043] In this case, with respect to at least one driving operation 3a * ~3f * not allowed in the situation 60 characterized by the learned representation 61a * ~3f * the increase in the probabilities 2a~2f assigned to this driving operation 3a
[0044] is suppressed from increasing the evaluation 230a by the cost function 5. FIG. 2 exemplarily shows two methods capable of realizing this. * ~3f * For example, according to block 231, the cost function 5 may be extended by a penalty term that absorbs and / or overcompensates for the gain that would be obtained in relation to the original cost function 5 due to the increase in the probabilities 2a~2f assigned to the non - allowed driving operations 3a
[0045] Alternatively or in combination with this, according to block 221, the probabilities 2a~2f assigned to the non - allowed driving operations 3a * ~3f * supplied from the machine learning model 1 may be set to zero before the evaluation by the cost function 5.
[0046] Furthermore, according to block 222, the probability distribution 2 may be regularized and / or discretized such that probabilities 2a to 2f that are below a predetermined threshold value are reduced to zero.
Claims
1. A method (100) for selecting a driving operation (4) to be performed by a vehicle (50) that travels at least partially automatically, comprising: - creating a representation (61) of a situation (60) in which the vehicle (50) is located by using measurement data (51a) from at least one sensor (51) mounted on the vehicle in step (110); - mapping the representation (61) of the situation (60) to a probability distribution (2) by means of a trained machine learning model (1), the probability distribution (2) representing, for each of a predetermined catalog of available driving operations (3a - 3f), the probability (2a - 2f) that the driving operation (3a - 3f) is to be performed, in step (120); - deriving one driving operation (3a - 3f) as the driving operation (4) to be performed from the probability distribution (2, 2') in step (150); ・Additionally using at least one aspect (62) of the situation (60) in which the vehicle (50) is placed, an operation (3a * ~3f * ) to obtain a subset of steps (130, 160), · A step (140, 170) of suppressing execution of the inadmissible driving operation (3a * to 3f * ); and including generating a modified probability distribution (2') by setting to zero the probability (2a - 2f) that at least one non - allowed driving operation (3a* - 3f*) is to be performed in the probability distribution (2) (141), thereby suppressing the execution of the non - allowed driving operation (3a* - 3f*), method (100).
2. The method (100) according to claim 1, wherein the modified probability distribution (2') is generated by normalizing the probability distribution (2) (142) such that, after setting at least one probability (2a - 2f) to zero, the probabilities (2a - 2f) of the driving operations (3a - 3f) that are still different from zero sum up to 1.
3. executing at least one unacceptable driving operation (3a * to 3f * ) is suppressed (171) by deriving a new driving operation (3a to 3f) from the probability distribution (2) in response to the unacceptable driving operation (3a * to 3f * ) being derived from the probability distribution (2), the method (100) according to claim 1 or 2.
4. At least one non-permissible driving operation (3a * to 3f * ) is determined (131, 161) based on information called from a spatially resolved digital map based on the current position of the vehicle (50), the method (100) according to any one of claims 1 to 3.
5. At least one unacceptable driving operation (3a * to 3f * ) is - deviation from a lane, and / or - violation of general traffic rules, and / or - damage to a specific mechanism for automated driving actions, and / or - collision of the host vehicle (50) with another vehicle or other object The method (100) according to any one of claims 1 to 4, which is a driving operation that poses a risk to.
6. The unacceptable operation (3a * to 3f * ) is - a lane change leading to deviation from a lane, and / or - a lane change to a lane that is currently unreachable, and / or - an acceleration operation and / or overtaking operation prohibited by traffic rules, and / or - following behind another vehicle that is currently located behind the host vehicle (50) The method (100) according to any one of claims 1 to 5, including.
7. A computer program, which, when executed on one or more computers, includes machine-readable instructions for causing the one or more computers to perform the method (100, 200) according to any one of claims 1 to 6.
8. A machine-readable data carrier comprising the computer program according to claim 7.
9. A computer comprising the computer program according to claim 7 and / or the machine-readable data carrier according to claim 8.
Citation Information
Patent Citations
Adapting the trajectory of an ego vehicle to moving foreign objects
DE102018210280A1
Driving support information creation device, and driving support device
JP2016149013A
Automatic operation device
JP2018024286A
Learning method, learning device, and learning program for ai agent behaving like human
JP2020191022A
Control device, control method, and program
WO2019017253A1