Method for generating training data for a machine learning model for predicting a trajectory of a road user and vehicle

The method improves trajectory prediction by selecting relevant training data based on similarity measures and normalized errors, addressing inefficiencies in existing methods to enhance the accuracy and consistency of road user trajectory prediction.

DE102023004669B4Active Publication Date: 2025-10-09MERCEDES BENZ GROUP AG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE102023004669
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-10-09
Estimated Expiration
2043-11-16

AI Technical Summary

Technical Problem

Existing methods for generating training data for machine learning models to predict road user trajectories are inadequate in ensuring the multi-modality of the model, leading to potential collisions and inefficiencies in trajectory prediction.

Method used

A method involving determining real data of road users at different times and perspectives, using a similarity measure to select training targets based on predefined thresholds, and employing backpropagation with normalized errors to refine the model, thereby improving trajectory prediction accuracy.

Benefits of technology

Enhances the accuracy and consistency of trajectory predictions by selecting relevant training data and minimizing errors, reducing the likelihood of collisions and enhancing the model's predictive capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method for generating training data for a machine learning model that predicts trajectories of road users detected in the environment of a vehicle (1) by means of the machine learning model, characterized in that to determine the training data, a plurality of real data of the road users (11, 13) in a predetermined traffic environment (2) are determined at different times and perspectives and several training targets are extracted from the real data of the road users (11, 13) by means of a similarity measure, wherein the similarity measure is derived from real data comprising driving conditions between a first road user (11), for whom the model is trained to predict a trajectory, at a fixed time (t) and one or more further road users (13) at the continuous times (T) associated with the driving conditions, and if it is detected that the similarity measure lies in a range beyond a predetermined threshold value, a trajectory of the one or more further road users (13) is selected as a training target from a time of detection, wherein the trajectories used as training target are limited to a predetermined maximum number, whereby the predetermined maximum number depends in particular on the underlying traffic scene (2).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for generating training data for a machine learning model for predicting a trajectory of a road user, in which the predicted trajectories of road users detected in the surroundings of a vehicle are determined taking into account deviations, the deviations being determined by means of the machine learning model, and to a vehicle.

[0002] From DE 10 2016 215 314 A1 a means of transport and a method for predicting a traffic situation are known, in which a prediction device is used to determine a predicted trajectory of objects detected by an object detection device in the environment of a vehicle using a trained machine learning model.

[0003] US 2021 / 0 263 157 A1 discloses a method in which a neural network model is used to predict other vehicles in the vicinity of an autonomously driving vehicle. Combined sensor data from high-end lidar, camera, and radar are used to determine objects and information or metadata of objects surrounding the autonomous vehicle for data collection at each point in time during a planning cycle. The method can use the detected objects and metadata processed from the data acquired by the high-end lidar and other sensors as ground truth to label data from a similar scene acquired by a low-end lidar mounted on the vehicle.

[0004] The object of the invention is to provide a method and a vehicle for generating training data for a machine learning model for predicting a trajectory of a road user, with which the multi-modality of the machine learning model is improved.

[0005] The invention is based on the features of the independent claims. Advantageous developments and refinements are the subject of the dependent claims. Further features, possible applications, and advantages of the invention will become apparent from the following description and the explanation of exemplary embodiments of the invention illustrated in the figures.

[0006] The problem is solved by the subject matter of patent claim 1 or 6.

[0007] In the method for generating training data for a machine learning model, which, after a training phase, predicts trajectories of road users detected in the environment of a vehicle using the machine learning model, a large number of real data of the road users in a predetermined traffic environment at different times and perspectives are determined to determine the training data, and several training targets are extracted from the real data of the road users using a similarity measure. The similarity measure is derived from real data comprising driving conditions between a first road user, for which the model is trained to predict a trajectory, at a specified time (t) and another road user at all times (T) associated with the driving conditions recorded with the real data, and if it is detected that the similarity measure lies beyond a predetermined range lying within a predetermined threshold value, a trajectory of the other road user is selected as a training target alongside the trajectory of the first road user from the time at which the threshold value is undershot.

[0008] When determining real data as a basis for training the model for trajectory prediction, a predetermined traffic environment is driven through several times, whereby a large amount of real data on the road users is determined for the traffic scene. A traffic environment in the sense of this description is understood to be a traffic infrastructure, for example an intersection, a roundabout with different entry lanes, a pedestrian crossing or a road with mixed road sections, i.e. one containing vehicles, bicycles and pedestrians. A traffic scene describes the movements or trajectories of road users in the traffic environment. The real data includes the driving states, position, acceleration, speed and orientation of the road users and the associated times of data collection, whereby an associated trajectory is described from the real data.The data sets comprising trajectories each belong to many road users, with each of these trajectories intended for training being divided into input data and training target (e.g., 5 seconds of input + 6 seconds of training target). The data sets of the road users in the predetermined traffic environment can be determined at different times and from different perspectives. The machine learning model is improved by having potentially multiple training data sets available for each traffic environment. The training data relating to the first road user for whom the model is trained to predict a trajectory is determined based on the similarity measure to other road users determined from driving conditions.If the similarity measure to another road user lies within a predefined range beyond the threshold value, then from the time the similarity measure falls below the threshold value, the remaining trajectory of the other road user is selected as the training target. The predefined range can be above or below the threshold value depending on how the similarity measure is calculated. In the following, it is assumed that a threshold value below the similarity measure reflects a greater similarity between the first road user and another road user than a similarity measure above the threshold value. The similarity measure is determined from a driving state of the first vehicle at a specified fixed point in time relative to the trajectory of the first vehicle with driving states of the other vehicle at consecutive points in time relative to their recorded trajectory, one after the other, until the threshold value is undershot.The fixed point in time is the last observed driving state in the input data of the first vehicle included in the trajectory. Training data for training a machine learning model for predicting an object's trajectory can advantageously be extracted from the real data of road users.

[0009] The trajectories used as training targets are limited to a predefined maximum number, whereby the predefined maximum number can be made dependent on the underlying traffic environment. Of the trajectories of the other road users that can be used as training targets, the trajectories associated with the greatest degree of similarity to the first road user are used. This prevents a collapse of different modes during machine learning.

[0010] Advantageously, the driving state of the first road user and the detected additional road user includes their position and / or their speed and / or their acceleration and / or their orientation. These parameters allow the similarity measure to be reproduced particularly accurately, so that training data from additional road users can be determined to train the model for the first road user to predict a trajectory.

[0011] In one embodiment, the colliding trajectories of other road users are discarded as training data / targets. This ensures that only those trajectories consistent with the future movement of the other road users are retained. This eliminates the need for an expensive ensemble model for machine learning.

[0012] It is advantageous if the training of the machine learning model is continued with the non-colliding trajectories of the other road users and the trajectory of the road user to be predicted, with an error being backpropagated for the trajectory whose error relative to the ground-truth data of the first road user is the smallest. The ground-truth data is defined here by the real data of the first road user. The backpropagation represents an error feedback, by means of which the machine learning model is trained and the error relative to the ground-truth data is minimized.

[0013] In a further embodiment, the error provided for backpropagation is normalized by dividing it by a factor that includes the number of non-colliding trajectories. In other words, the error is divided by the number of available training targets, i.e., the number of training targets of the non-colliding trajectories m, preferably increased by the one trajectory of the first road user with m+1.

[0014] This compensates for the fact that in each traffic scene considered in the traffic environment, potentially a different number of training targets were considered.

[0015] A further aspect of the invention relates to a vehicle comprising a machine learning model for predicting a trajectory of the vehicle, wherein the machine learning model is designed to carry out the method according to at least one feature described in this patent application.

[0016] Further advantages, features, and details will become apparent from the following description, in which at least one exemplary embodiment is described in detail—possibly with reference to the drawings. Described and / or illustrated features may form the subject matter of the invention alone or in any meaningful combination, possibly independently of the claims, and may, in particular, also be the subject of one or more separate applications. Identical, similar, and / or functionally equivalent parts are provided with the same reference numerals.

[0017] They show: Fig. 1 an embodiment of the vehicle according to the invention, Fig. 2 an embodiment of the method according to the invention and Fig. 3 schematically illustrated model of machine learning.

[0018] In Fig. 1 shows a vehicle 1 comprising a control unit 3 connected to environmental sensors, for example an environmental camera 5, and vehicle-internal sensors, such as a GPS sensor 9. For this purpose, the control unit can have at least one computing unit, at least one memory, and at least one interface, which can be implemented in hardware and / or software. A speed sensor 7 detects the speed of the vehicle 1, while the GPS sensor 9 detects its position in x, y, and z coordinates. The environmental camera 5 detects all road users 11, 13 in a traffic environment 2, such as an intersection, and creates images or a video stream, which is transmitted to the control unit 3 as a data set for further processing.The control unit 3 determines, as far as necessary, real data comprising driving conditions of a first vehicle 11 and other vehicles 13 involved in the traffic scene 2 based on its own GPS position and speed.

[0019] It should be noted that the traffic environment is further traversed by vehicle 1 at different times and perspectives and that with each subsequent passage the real data of vehicles 13 shown as dashed lines is determined.

[0020] These real-time data, summarized into a scenario, are used in the control unit 3 to train a machine learning model that is configured to precess a trajectory of the first vehicle 11.

[0021] An embodiment of the method according to the invention is shown in Fig. 2. A model located in the control unit 3 is to be trained for the first vehicle 1 to predict a driving trajectory.

[0022] After the method has started in block 100, the driving state X of the first road user 11 is detected in block 110 at a fixed time t. The driving state X is characterized by the position x, y of the road user 11, its speed vx, vy, its acceleration ax, ay and its orientation δ. The state X is mapped, for example, in a two-dimensional coordinate system (e.g., UTM). Based on the driving state X of the first road user at the fixed time t, a similarity measure is continuously calculated in block 120 for each driving data of one or more further road users 13 present in the data set predetermined by the scene. The data set consists of many traffic scenarios, with each of these traffic scenarios being divided in time into input data and training target for training, e.g., 5 seconds input + 6 seconds training target.

[0023] As a similarity measure, an error measure is used which is based on how similar the driving state X of the first road user 11 at the fixed time t and the driving state of the one or more further road users 13 at the time T are to each other.

[0024] The following applies: alpha*(x_A,t−x_B,T)2+alpha*(y_A,t−y_B,T)2+beta*(vx_A,t−vx_B,T)2+beta*(vy_A,t−vy_B,T)2+…, where the time t is fixed as the last observed time in the trajectory of the first vehicle 11 and T describes the continuous times associated with the determination of the driving conditions of the one or more further road users 13 from the beginning to the end, where A stands for the first vehicle 11 and B for each of the further vehicles 13.

[0025] Each time point T and the associated driving states of one or more other road users 13 from the real data are determined to calculate the similarity measure.

[0026] The similarity measure describes how similar the driving state X of the first road user at a fixed time t is to the driving state of one or more other road users at different times T. The more similar the driving state of the first road user is to the other vehicle(s), the smaller the calculated similarity value. Alpha, beta, and similar factors represent weighting factors that can be determined manually, for example.

[0027] Using the similarity measure, several training targets are extracted from real data.

[0028] In block 130, trajectories of road users 13 are extracted for which the similarity measure falls below a certain threshold S. As soon as the similarity measure of the respective road user 13 is below the threshold S at a detection time, the trajectory from the detection time of the respective road user is selected as the training target in block 140.

[0029] If the detection time is so late that the entire length of the training target of the trajectory of one of the other road users is no longer available, then the similarity measure is calculated and averaged during training only for the available time steps.

[0030] If too many extracted trajectories N fall below the threshold, their number is limited to a predefined maximum value M (block 150), whereby M trajectories are selected that exhibit the greatest similarity to the trajectory of the first vehicle. Subsequently, for each of the selected scenes, the future trajectory of the other road users from the time of detection is extracted in block 160. These trajectories, along with that of the first vehicle, are assumed to be training targets, so-called "ground truths." In a machine learning process, these represent the data that allows the quality of the model to be verified and the model to be further trained.

[0031] In block 170, collision-based pruning (simplified shortening of a decision tree) is performed for the M trajectories. Each of the M trajectories is inserted sequentially into the traffic scene specified in block 110. If a collision occurs with a "ground truth" of another road user 13, the trajectory is discarded. This ensures that only the M trajectories that are consistent with the future movement of other road users 13 are reused.

[0032] Subsequently, the M' trajectories determined after pruning are used in block 180 for training a model according to Fig.3 is used to predict the trajectory of the first road user 11. The model comprises input encoder 200, output decoder 210, and an interaction module 215 arranged between them, also referred to as a hidden layer. Decoder modules 220 arranged in parallel in output decoder 210 are referred to as prediction heads. Initially, for road user 11, the prediction is based on its own historical movements, the historical movements of other road users 13, and information about the static infrastructure. The result is multiple predictions of the trajectory (depending on the number of prediction heads) for the first road user 11.

[0033] In the following, the M'+1 training targets of the other road user 11 and the first road user 11 are used to determine the quality and thus the error measure for the multiple predictions of the first road user 11 generated by backpropagation. For each of the M'+1 training targets, the prediction head is determined that best covers the training target defined by the trajectory of the first road user 11, e.g., the lowest average Euclidean distance. An error between the prediction of the prediction head and this ground truth is then backpropagated using the prediction head selected in this way, which is performed for each of the M'+1 training targets.

[0034] In other words, the proposed method provides not only one ground truth for the road user 11 but M'+1 ground truths for training the model.

[0035] Before the error is fed back into the model, which is also known as backpropagation, a normalization step of the error measure is performed. This is necessary because potentially multiple traffic scenes can be processed in parallel in a forward pass. The normalization is performed to give each traffic scene equal weight. To do this, the error between the prediction of the prediction head that best covers the training target and the actual training target (prediction of the trajectory of vehicle 1) of each traffic scene is divided by the factor M'+1. This compensates for the fact that each traffic scene potentially contains a different number of training targets.

Claims

[1] Method for generating training data for a machine learning model that predicts trajectories of road users detected in the environment of a vehicle (1) by means of the machine learning model, characterized by , that to determine the training data, a plurality of real data of the road users (11, 13) in a predetermined traffic environment (2) are determined at different times and perspectives and several training targets are extracted from the real data of the road users (11, 13) by means of a similarity measure, wherein the similarity measure is derived from real data comprising driving conditions between a first road user (11), for whom the model is trained to predict a trajectory, at a fixed time (t) and one or more further road users (13) at the continuous times (T) associated with the driving conditions, and if it is detected that the similarity measure lies in a range beyond a predetermined threshold value, a trajectory of the one or more further road users (13) is selected as a training target from a time of detection, wherein the trajectories used as training target are limited to a predetermined maximum number, whereby the predetermined maximum number depends in particular on the underlying traffic scene (2). [2] Method according to claim 1, characterized by that the driving state of the first road user (11) and the detected further road user (13) includes their position and / or their speed and / or their acceleration and / or their orientation. [3] Method according to at least one of the preceding claims, characterized bythat selected trajectories of the other road users (13) that collide with each other are discarded as training target. [4] Method according to claim 3, characterized by that the training of the machine learning model is continued with the non-colliding trajectories of the other road users (13) and the trajectory of the first road user (11), wherein for the training an error is backpropagated for the trajectory whose error to ground truth data of the first vehicle (11) is the smallest. [5] Method according to claim 4, characterized by that the error provided for backpropagation is normalized by dividing it by a factor that includes the number of non-colliding trajectories. [6] Vehicle (1) comprising a control unit (3) with a machine learning model for predicting a trajectory of road users (11, 13), characterized bythat the machine learning model is trained using a method of the training data determined according to the preceding claims.

Citation Information

Patent Citations

  • Driver assistance system, means of transport and methods for predicting a traffic situation

    DE102016215314A1

  • Automated labeling system for autonomous driving vehicle lidar data

    US20210263157A1