Method for determining and evaluating trajectory of road user
By extracting the characteristics of road user trajectory and determining the trajectory end point, and evaluating the reliability of the trajectory in combination with machine learning algorithms, the problem of difficult to predict and evaluate the future trajectory of road users in the prior art is solved, and efficient trajectory prediction and reliability evaluation are achieved.
Patent Information
- Application Number
- CN202411199919.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-08-29
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to effectively predict the future trajectory of road users, especially in complex traffic scenarios, and it is difficult to evaluate the reliability of the trajectory.
Relevant characteristics are extracted and used to determine one or more trajectory end points by receiving input data associated with the movement state and environment of the road user. Then, for each trajectory end point, the corresponding trajectory is determined and the reliability of the trajectory is evaluated by classification, providing a confidence score.
This method can effectively predict the trajectory of road users while reducing the computational workload, and provide trajectory reliability assessment, supporting safe and convenient path planning of autonomous vehicles.
Smart Images

Figure CN120196693A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a computer-implemented method for determining and evaluating the trajectories of road users. Background Art
[0002] For autonomous driving, understanding the entire traffic scene around an autonomous vehicle is an important and challenging task. For example, understanding the road structure, identifying the intentions of other road users (such as other vehicles and pedestrians), and predicting the future movements of all road users. Planning a safe and convenient future trajectory for an autonomous vehicle depends to a large extent on the understanding of the traffic scene in the external environment of the autonomous vehicle and the prediction of its dynamics.
[0003] To accurately predict the future trajectories of surrounding road users, it is necessary to consider and model the effects of static environments such as lanes and road structures, traffic signs, etc., as well as the interactions between road users. The interactions between road users have different time spans and various distances, which lead to a high degree of complexity.
[0004] Existing methods for trajectory prediction may generate multiple future trajectories for road users, and each of these future trajectories may be feasible. Therefore, one of the challenges of the trajectory prediction task is to provide some ranking of the generated future trajectories, that is, which predicted future trajectories are more likely than others.
[0005] For example, existing methods for trajectory prediction may rely on entity-based models that encode each road user and different parts of an abstract representation of the road around an autonomous vehicle, for example. However, such entity-based models typically scale with the number of road users in the region of interest (i.e., the external environment of the autonomous vehicle).
[0006] Other known methods use grid maps in a so-called bird's-eye view, which may contain very detailed information about the environment of the autonomous vehicle. Such grid-based maps or models can be regarded as projections of the three-dimensional world and are usually independent of the number of road users at runtime.
[0007] This method using a grid map typically predicts a so-called heat map that includes the possible future positions of road users in the traffic scene of the environment of the autonomous vehicle. Therefore, the resulting grid map can provide corresponding probabilities for multiple future trajectories within the corresponding cells of the grid map, and these multiple future trajectories can be learned via special convolutional kernels that are incorporated into the machine learning algorithms of the corresponding methods. In addition, such grid-map-based methods can easily combine inputs from different sensors.
[0008] However, it may not be easy to obtain the predicted probability distribution in the resulting heatmap, and thus appropriate post - processing and sampling may be required to extract the actual coordinates and further dynamic attributes of the road users along the predicted future trajectories.
[0009] Therefore, a method is needed that can predict at least one trajectory of a road user with reduced effort and allows the reliability of the trajectory to be evaluated. SUMMARY OF THE INVENTION
[0010] The present disclosure provides a computer - implemented method, a computer system, and a non - transitory computer - readable medium according to the independent claims. Embodiments are given in the dependent claims, the detailed description, and the drawings.
[0011] In one aspect, the present disclosure relates to a computer - implemented method for determining and evaluating a trajectory of a road user. According to the method, input data associated with the movement state and environment of the road user is received. Characteristics associated with the road user are extracted from the input data. One or more trajectory endpoints are determined for the road user by using the extracted characteristics. For each of the trajectory endpoints, a corresponding trajectory associated with one of the trajectory endpoints is determined by using the associated trajectory endpoint and the extracted characteristics; and the corresponding trajectory of the road user is evaluated by using a classification that depends on the extracted characteristics to provide a confidence score for each trajectory associated with one of the trajectory endpoints.
[0012] The input data associated with the movement state of the road user may be provided by another system, which may be regarded as a so - called backbone system and may be capable of providing information related to the trajectory of the road user. The movement state of the road user may include the respective position, speed, and acceleration of the road user within a predetermined number of time steps.
[0013] In addition to the dynamics of the road user, the "dynamic context" and "static context" of the road user are also considered in the form of input data associated with the environment of the road user. The dynamic context may include dynamic items such as other road users, e.g., other vehicles and pedestrians, while the static context may include static items such as lane markings or traffic signs for indicating the road route. In addition to the characteristics of the static entities in the environment that may be related to the trajectory of the road user (i.e., the static context), the characteristics or dynamics of other road users (i.e., the dynamic context) may also be provided by the backbone system.
[0014] However, this information can be provided in an encoded and abstract form, i.e., by a prediction algorithm representing the backbone system. For example, the input data can be related to a so-called feature map (i.e., a grid map for a corresponding future time step), which includes relevant information about road users within different cells of the grid map. In addition, the input data can also include information about the static and dynamic situations of road users in an encoded and abstract format.
[0015] Therefore, the presence of input data received, for example, from the backbone system is a prerequisite for performing the method according to the present disclosure. That is, the input data can depend on data provided by one or more sensors of a perception system that can be installed, for example, on an autonomous vehicle, and the data detected by such a perception system may have been preprocessed for the movement state of road users to provide the input data received by the method according to the present disclosure.
[0016] Steps are required to extract features related to road users from the input data, because the input data can be provided in an encoded and abstract form as described above, for example, by an encoder of the backbone system. The backbone system may have a multi-scale structure, such that the output provided by the system (e.g., the feature map) may need to be further processed in order to provide data suitable for determining one or more trajectory endpoints in the next method step. For example, a convolutional neural network can be used for the step of extracting features from the feature map.
[0017] In most practical situations, multiple trajectory endpoints (i.e., more than one trajectory endpoint) can be determined for a road user by using the features extracted from the input data, and for each of the multiple trajectory endpoints, a single trajectory of the road user can be determined by using the corresponding trajectory endpoint and the features extracted from the input data. Each trajectory of the road user can be evaluated by using classification.
[0018] Therefore, the task of generating and evaluating multiple trajectories for a certain road user can be transformed into determining multiple endpoints and determining a single trajectory for each of the endpoints. This can significantly reduce the required computational workload.
[0019] In order to evaluate the corresponding trajectories determined for the corresponding endpoints, classification can be provided or learned by applying ground truth data including known trajectories, which are associated with the perception data received for the corresponding traffic scenario. The perception data can also be processed by the backbone system, which can generate corresponding input data, and the steps of extracting features, determining one or more trajectory endpoints, and determining the predicted trajectories can be applied to these input data in order to associate the predicted trajectory with one of the known trajectories, for example, by using the framework of a generative adversarial network (GAN) to learn the classification of the trajectories of road users.
[0020] The confidence score of each trajectory associated with one of the trajectory endpoints can also be provided as an output of the method, e.g., provided to a further application or control system of a vehicle to which the method is applied. Thus, information about the reliability and feasibility of the trajectories determined by the method can be provided to further applications and systems in order to select the best one of these trajectories for controlling the vehicle.
[0021] By evaluating each of the trajectories to provide a confidence score, a ranking indicating the most feasible or realistic trajectories of the road user can be provided. The confidence score of the corresponding trajectory can be expressed as a probability. That is, the confidence score can be provided as a single number from 0 to 1.
[0022] The output of the method, i.e., one or more evaluated trajectories for the road user, can be provided for a further method or a further system, which may be able to select the final trajectory of the road user based on the evaluated trajectories. In other words, the last method step of evaluating the corresponding trajectories of the road user can be regarded as a discriminator step, which can provide a single number or value for each trajectory in order to evaluate the corresponding trajectory, i.e., to quantify the realism of the corresponding trajectory. For example, a further system or method can select the most likely trajectory with the best evaluation.
[0023] One advantage of the method according to the present disclosure is that by determining only one or more trajectory endpoints instead of determining one or more complete trajectories, the possible multimodality provided by the backbone or prediction system is handled, i.e., more than one feasible trajectory of the road user that may be generated due to specific input data. Thus, the necessary post-processing and computational effort required to perform the method is reduced.
[0024] Furthermore, the extracted features are utilized when determining one or more trajectory endpoints, when determining the corresponding trajectories for the associated endpoints, and when providing the classification of the corresponding trajectories. Thus, the method steps are repeatedly related to the extracted features reflecting the traffic scenario in which the road user is currently located. Thus, the reliability of the output of the method can be improved.
[0025] Furthermore, the method provides an evaluation of the feasibility of the trajectories based on the results of the evaluation step. This can support the planning of future trajectories (e.g., of an autonomous vehicle), since a planning system or method can rely on the evaluation of one or more trajectories of the road user in future time steps.
[0026] According to an embodiment, the method may rely on at least one machine learning algorithm, and the input data may include embedding features that may be provided by a prediction algorithm and may include information about other road users and information about the static environment of the road users, i.e., information other than information about the movement state of the considered road user. A feature extraction algorithm may be applied to the embedding features to provide the extracted features.
[0027] The feature extraction algorithm of the machine learning algorithm on which the method relies may thus be capable of extracting relevant information about the considered road user, about other road users, and about the static environment of the road users, such that the extracted information is suitable for determining one or more trajectory end points. The machine learning algorithm may be configured to learn to extract the relevant information and determine one or more trajectory end points based on this information, e.g., by using a training process with actual trajectory end points.
[0028] As described above, the data provided by the feature extraction algorithm is additionally used to determine and evaluate the corresponding trajectory of the road user. Thus, the input data provided as embedding features and the features extracted therefrom can be used not only to determine one or more trajectory end points, but also to support the steps of determining the entire trajectory and evaluating the corresponding trajectory.
[0029] Determining one or more trajectory end points for a road user may include: applying a machine learning algorithm that is configured to learn to encode the extracted features and the ground truth data of the true trajectory into at least one latent variable. The machine learning algorithm may further be configured to learn to decode the at least one latent variable into a corresponding predicted trajectory end point associated with a corresponding one of the true trajectories in the true trajectory. Additionally, determining at least one trajectory end point for the road user by using the extracted features may further include: sampling at least one latent variable from a prior distribution provided by the machine learning algorithm when learning to encode the extracted features and the ground truth data of the true trajectory; and decoding the at least one latent variable into one or more trajectory end points.
[0030] That is to say, the separate encoding-decoding process can be used for the step of determining one or more trajectory endpoints. Therefore, the machine learning algorithm for determining the trajectory endpoints can be trained independently of other parts of the machine learning algorithm on which the entire method depends. Since the machine learning algorithm can be trained with real or known trajectories with known endpoints within a given time frame (i.e., a predetermined time step), the reliability of the prediction of the trajectory endpoints can be improved. As a result of the training, a prior distribution can be provided such that the method can be capable of sampling latent variables from the prior distribution during the inference process (i.e., during the operation phase of the method). In addition, at least one latent variable can be a continuous latent variable, and the continuous latent variable can be decoded for providing one or more trajectory endpoints.
[0031] Determining the corresponding trajectory of a road user can include: determining the position of the road user from the road user's current position to the position of the associated trajectory endpoint, and determining the dynamic parameters of the road user for each position among the positions. For example, the dynamic parameters can include the speed, acceleration, and / or yaw rate of the road user under consideration. Therefore, the determined trajectory of the road user can be regarded not only as the "future route" of the road user's movement, but also include the dynamic parameters of the road user for each position among the positions. A method or system for planning the future trajectory of a road user (e.g., of an autonomous vehicle) can rely on the entire set of information including the future position and future dynamic parameters of the road user.
[0032] In addition, determining the corresponding trajectory of a road user can include applying a machine learning algorithm, which can be configured to learn the parameters of a bicycle model. The parameters of the bicycle model can include, for example, the acceleration and yaw rate of the road user. Although such a machine learning algorithm can include, for example, a neural network that can learn to determine the acceleration and yaw rate of the road user, the underlying bicycle model may require a lower computational workload.
[0033] According to another embodiment, the machine learning algorithm can be trained for the evaluation step regarding the trajectory by the following steps: i) providing the ground truth trajectory and the predicted trajectory for a plurality of road users, wherein the predicted trajectory can be determined by applying the extracted features derived from the training input data associated with the respective movement states and respective environments of each of the plurality of road users; and ii) learning the classification of the road users by applying a generative adversarial network to the ground truth trajectory and the predicted trajectory.
[0034] In other words, the method can learn to evaluate or classify predicted trajectories with data related to feasible actual trajectories, and corresponding predictions for these feasible true trajectories can be provided by the method (e.g., by the generator of a generative adversarial network (GAN)). Trajectory prediction can be evaluated by the discriminator of the generative adversarial network during the training process based on the ground truth trajectory.
[0035] In other words, supervised learning can be performed by applying a generative adversarial network. When performing supervised learning, the ground truth trajectory can be marked with "true trajectory", while the predicted trajectory can be marked with "false trajectory". The goal is to train a classifier that provides classification in a way that attempts to distinguish between these two types of trajectories. At the same time, the prediction algorithm (e.g., the trajectory generator) may try to predict the trajectory as "real" as possible. At the end of the training, the classifier may thus be unable to distinguish the difference between the predicted trajectory and the ground truth trajectory.
[0036] Through this training of the machine learning algorithm based on the true trajectory, a reliable evaluation or classification of the trajectories provided by the method can be achieved. As described above, during the inference process, when classifying or evaluating the predicted trajectory, the characteristics extracted that reflect the dynamics of the road user and their dynamic and static situations can be considered.
[0037] In addition, the generative adversarial network can be used to provide a classification for evaluating the corresponding trajectories of road users. For example, the discriminator of the generative adversarial network can not only evaluate the corresponding trajectory to improve the trajectory prediction provided by the trajectory generator, but also provide a classification of the corresponding trajectory based on which the confidence score is determined.
[0038] The general training process of the machine learning algorithm can be performed for the steps of determining and evaluating the corresponding trajectories of road users. Therefore, the step of evaluating or assessing the trajectory can interact with the step of determining one or more trajectories by applying the previously determined corresponding end points. Due to this interaction of these two method steps in the general training process (this interaction can also be provided by the generative adversarial network), better trajectories can be determined in the first of these two method steps, i.e., trajectories improved in terms of safety, reliability, and feasibility. The training of the machine learning algorithm can rely on a combination of mean squared error loss and min-max loss, which ensures the convergence of the training.
[0039] On the other hand, the present disclosure relates to a computer system configured to perform several or all of the steps of the computer-implemented method described herein.
[0040] The computer system may include a processing unit, at least one storage unit, and at least one non-transitory data memory. The non-transitory data storage and / or memory unit may include a computer program for instructing the computer to execute several or all steps or aspects of the computer-implemented method described herein.
[0041] As used herein, terms such as processing units and modules may refer to, be, an application specific integrated circuit (ASIC), an electronic circuit, a combinational logic circuit, a field programmable gate array (FPGA), a processor (shared, dedicated, or group) executing code, other suitable components providing the functions, or a part of a combination of some or all of the above, or include an application specific integrated circuit (ASIC), an electronic circuit, a combinational logic circuit, a field programmable gate array (FPGA), a processor (shared, dedicated, or group) executing code, other suitable components providing the functions, or a combination of some or all of the above, such as in a system on a chip. The processing unit may include a memory (shared, dedicated, or group) storing code executed by the processor.
[0042] According to an embodiment, the computer system may further include: an end-point generator configured to determine, for a road user, one or more trajectory end-points by using features extracted from input data; a trajectory generator configured to determine, for each of the trajectory end-points, a corresponding trajectory associated with one of the trajectory end-points by using the associated trajectory end-point and the extracted features; and a discriminator configured to evaluate each trajectory associated with one of the trajectory end-points.
[0043] In addition, the end-point generator, the trajectory generator, and the discriminator may be implemented as a generative adversarial network, and the discriminator may be configured to evaluate each trajectory by using the classification provided by the generative adversarial network.
[0044] In another aspect, the present disclosure relates to a vehicle including the perception system and the computer system described herein.
[0045] According to an embodiment, the vehicle may further include a control system configured to control the actual trajectory of the vehicle. The computer system may be configured to transmit at least one evaluated trajectory of the road user to the control system such that the control system can apply the at least one evaluated trajectory of the road user when controlling the actual trajectory of the vehicle.
[0046] In another aspect, the present disclosure relates to a non-transitory computer-readable medium comprising instructions for performing several or all steps or aspects of the computer-implemented methods described herein. The computer-readable medium may be configured as: an optical medium such as a compact disc (CD) or a digital versatile disc (DVD); a magnetic medium such as a hard disk drive (HDD); a solid-state drive (SSD); a read-only memory (ROM); a flash memory; or the like. Additionally, the computer-readable medium may be configured as a data memory accessible via a data connection such as an Internet connection. The computer-readable medium may be, for example, an online data repository or a cloud memory.
[0047] The present disclosure also relates to a computer program for instructing a computer to perform several or all steps or aspects of the computer-implemented methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Exemplary embodiments and functions of the present disclosure are described herein in connection with the following drawings, which schematically show:
[0049] Figure 1 a diagram of a vehicle and the vehicle's surrounding environment including a computer system according to the present disclosure,
[0050] Figure 2 a diagram of an example of a network architecture for providing input data to a method according to the present disclosure,
[0051] Figure 3 an overview of a method according to the present disclosure,
[0052] Figure 4 a detailed diagram of a method according to the present disclosure,
[0053] Figure 5 a flowchart illustrating a method for determining the trajectory of a road user according to various embodiments,
[0054] Figure 6 a diagram of a system according to various embodiments, and
[0055] Figure 7 a computer system having a plurality of computer hardware components configured to perform the steps of a computer-implemented method as described herein. DETAILED DESCRIPTION
[0056] Figure 1Depicts a schematic illustration of a vehicle 100 and objects that may surround the vehicle 100 in a traffic scenario. The vehicle 100 includes a perception system 110 having an instrument field of view indicated by line 115. The vehicle 100 further includes a computer system 120 that includes a processing unit 121 and a data storage system 122, and the data storage system 122 includes, for example, a memory and a database. The processing unit 121 is configured to receive data from the perception system 110 and store the data in the data storage system 122. The vehicle 100 further includes a control system 124 that is configured to set and control the actual trajectory of the vehicle 100.
[0057] The perception system 110 may include a radar system, a lidar system, and / or one or more cameras to monitor the external environment or the surrounding environment of the vehicle 100. Thus, the perception system 110 is configured to monitor the dynamic situation 125 of the vehicle 100, and the dynamic situation 125 includes a plurality of road users 130 that can move in the external environment of the vehicle 100. The road users 130 may include, for example, other vehicles 140 and / or pedestrians 150.
[0058] The perception system 110 is further configured to monitor the static situation 160 of the vehicle 100. The static situation 160 may include, for example, traffic signs 170 and lane markings 180.
[0059] The perception system 110 is configured to determine the trajectory characteristics of the road users 130. The trajectory characteristics include the current position, the current speed, and the object category of each road user 130. The current position and the current speed are determined by the perception system 110 relative to the vehicle 100, that is, relative to a coordinate system having its origin (for example, the origin is located at the centroid of the vehicle 100, its x-axis is along the longitudinal direction of the vehicle 100, and its y-axis is along the transverse direction of the vehicle 100). In addition, the perception system 100 determines the trajectory characteristics of the road users 130 within a predetermined number of time steps (for example, each step is 0.5 seconds).
[0060] The computer system 120 transmits the results or outputs of the method according to the present disclosure (that is, at least one predicted trajectory 432 and the evaluation of at least one predicted trajectory 432 in the form of a confidence score 444 (see Figure 4 )) to the control system 124 so that the control system 124 can incorporate at least one evaluated trajectory into the control of the actual trajectory of the vehicle 100.
[0061] Figure 2 Depicts being included in the computer system 120 of the vehicle 100 (see Figure 1) Details of the processing unit 121. The processing unit 121 includes a backbone system for the method of the present disclosure in the form of a deep neural network 210, which is provided with different inputs. The inputs include a dynamic situation 125 (i.e., the trajectory characteristics of the road user 130 as described above), a static situation 160, and the ego dynamics 220 of the vehicle 100. The deep neural network 210 is used to generate a backbone output 230. When training the deep neural network 210, the backbone output 230 and the ground truth (GT) 240 are provided to a loss function 250 for optimizing the deep neural network 210.
[0062] The static situation 160 includes static environment data, which includes the respective positions and respective dimensions of static entities in the environment of the vehicle 100, e.g., the positions and dimensions of traffic signs 170 and lane markings 180. The static situation 160 (i.e., the static environment data of the vehicle 100) is determined via the perception system 110 of the vehicle 100 and additionally or alternatively determined from a predefined map available in the surrounding environment of the vehicle 100.
[0063] The static situation 160 is represented by one or more of the following:
[0064] - A rasterized image from an HD (high definition) map (see, e.g., Figure 5 its visualization at 532 therein), where the high definition map covers, for example, the exact positions of lane markings such that when the vehicle can be accurately positioned in the HD map, accurate information about its surrounding environment is provided to the vehicle,
[0065] - A drivable area determined via the perception system 110 of the vehicle 100, e.g., a grid map or an image data structure, where each pixel of such a map or image represents the drivability of a specific area in the instrument field of view of the perception system 110,
[0066] - Lane / road detection via the sensors of the perception system 110, where, using the detected lane markings, road boundaries, guardrails, etc. from the sensors, the perception system 110 can be configured to construct a grid map or image-like data similar to a rasterized map for describing the static situation 160,
[0067] - A static occupancy grid map.
[0068] The ego dynamics 220 can also be represented as one of the road users 130 and can thus be included in the dynamic situation input.
[0069] As will be described in detail below, the deep neural network 210 is used as a backbone for providing the embedding features 310 (see Figure 3 ) for the input data in the form of. The backbone output 230 at least partially includes the embedding features 310. In an abstract form, the backbone output 230 thus includes information related to the road user 130 and the dynamic and static situations 125, 160.
[0070] The ground truth 240 defines the task of the deep neural network 210. It covers, for example, position information such as occupancy probability and offset within the grid, as well as other attributes such as speed and acceleration, and / or other regression and classification tasks, such as the future position, speed, maneuvers, etc. of the road user 130 monitored in the current traffic scene.
[0071] Figure 3 An overview of the most important steps and items related to the method according to the present disclosure is depicted. The method relies on the specific backbone outputs 230, 310 provided by the above-mentioned deep neural network 210 as the input data for the method. Specifically, within the backbone system 210 for the method according to the present disclosure, a machine learning algorithm as described by M. Schaefer et al. in "Context-Aware Scene Prediction Network" (arXiv:2201.06933v1, January 18, 2022) is used. Specifically, the encoder of the system described in this reference is only used to represent the backbone system 210 that provides the input data for the method according to the present disclosure. However, as long as the input data is provided in a suitable form and contains appropriate information for generating and evaluating one or more trajectories of a specific road user 130 (see Figure 1 ), other methods and algorithms can also be used to provide the input data for the method according to the present disclosure.
[0072] As the backbone outputs 230, 310, the deep neural network 210 or another backbone system provides grid maps, which can also be referred to as feature maps, and includes encoded data on the movement states of a specific road user 130 over multiple time steps. Each cell of such a grid map includes different channels, which include encoded information on the probability of the position or location of the road user 130, as well as additional channels for encoded information on different dynamic attributes of the road user 130.
[0073] To derive one or more trajectories of a road user 130 from the backbone outputs 230, 310 of a deep neural network 210 (i.e., from the feature maps provided for multiple future time steps), post - processing is typically required, which can be computationally expensive. Additionally, the same may be true for the refinement of the identified or generated trajectories.
[0074] Furthermore, multimodality is a general property of the trajectory prediction task. That is, given a specific traffic scenario (including the static context 160 and the dynamic context 125 of the road user 130), there may be multiple possible or feasible future trajectories for the road user 130. Multimodality is a challenge for machine - learning algorithms for predicting trajectories because for each training sample, there is only one given ground - truth (GT) item. However, a prediction system needs to include the ability to predict other alternatives or possibilities of trajectories that are not given as ground - truth items. In existing machine - learning algorithms for trajectory prediction, the possibility of multimodality is inherently included in the feature maps output by the machine - learning algorithm because such feature maps can model all possible future positions of the road user that reflect the environment of the road user 130 within the feature map.
[0075] To avoid or at least reduce the computational effort required for post - processing the backbone outputs 230, 310 of existing prediction methods, a method according to the present disclosure is provided. As Figure 3 depicted, the backbone outputs 230, 310 of a deep neural network 210 (i.e., the encoder of a system representing such a deep neural network 210) are provided as a set of embedding features that are associated with the feature maps as described above. That is, the embedding feature 310 includes information about the movement state of the road user 130 for each grid cell, but this information is represented in an abstract form that cannot be directly interpreted.
[0076] Thus, the method includes a feature extraction step 320 that extracts information or characteristics about the road user 130 from the embedding feature 310 in such a way that the output of the feature extraction 320 is suitable for use in a trajectory generation step 330. The output of the trajectory generation step 330 is evaluated or assessed by a trajectory evaluation step 340. Thus, the output of the method according to the present disclosure is one or more trajectories of the road user 130 and a corresponding evaluation for each of these trajectories. The evaluation of each trajectory in the trajectories is provided as a confidence or probability score for each trajectory, i.e., a number from 0 to 1.
[0077] Figure 4Depicts the details of the machine learning algorithm on which the method according to the present disclosure depends. The backbone outputs 230, 310 of the deep neural network 210 are provided as embedding features 310 to the feature extraction network 410. The feature extraction network 410 provides information or characteristics about the movement state of the road user 130 to other parts of the entire machine learning algorithm in a suitable manner, that is, the target end point generator 420, the target trajectory generator 430, and the target trajectory discriminator 440.
[0078] The task of trajectory generation 330 (see Figure 3 ) includes two stages implemented by the target end point generator 420 and the target trajectory generator 430. For the first stage, the target end point generator 420 is configured to predict one or more end points 422 of a specific trajectory of the road user 130. Different from predicting one or more complete trajectories of the road user 130 according to the above-mentioned multimodality, the target end point generator 420 predicts one or more end points 422 reflecting multimodality in a simplified manner. Due to the dynamic context 125 and static context 160 of the road user, only the multimodality required for the target end point generator 420 to process the trajectory of the road user 130 is required.
[0079] For the second stage of the target trajectory generation task 330, the target trajectory generator 430 predicts a single trajectory for each of the predicted end points 422, that is, the corresponding trajectory reaching the specific end point 422. For the actual trajectory generation of the target trajectory generator 430, multimodality does not need to be considered because the "target" of the corresponding trajectory has been given as a conditional input.
[0080] The target end point generator 420 is implemented as a so-called conditional variational autoencoder (CVAE). In addition, other solutions for implementing the target end point generator 420 via machine learning algorithms are also possible. During the training process of the conditional variational autoencoder, the underlying model learns to encode the provided ground truth (GT) end points and information about the dynamic context 125 and static context 160 of the road user 130, that is, information about other road users 130 and, for example, about the road map, where the contexts 125, 160 are provided by the backbone in the form of a deep neural network 210. This entire information is encoded into continuous latent variables, and as a result of the training process, a prior distribution is provided. During the training process, the predicted trajectory end points are compared with the GT end points.
[0081] During the inference process, first, continuous latent variables are sampled from the prior distribution, and then decoded into predicted end points 422, that is, decoded into the context-conditioned multimodal target positions or end points of the road user 130. The target end point generator 420 can be trained independently of other parts of the entire machine learning algorithm on which the method according to the present disclosure depends.
[0082] The predicted end point 422 is transmitted to the target trajectory generator 430, which also receives the output of the feature extraction network 410, the output including information or characteristics regarding the dynamic situation 125 and the static situation 160 of the road user 130. The target trajectory generator 430 uses the predicted end point 422 to generate one or more kinematically feasible trajectories for the road user 130, i.e., a single trajectory 432 for each predicted end point 422.
[0083] The target trajectory generator 430 applies a bicycle model that constrains the corresponding trajectory generated based on the predicted end point 422 to the reality boundary. The target trajectory generator 430 uses a bicycle model, where a machine learning algorithm configured to learn the parameters of the bicycle model is applied. The parameters of the bicycle model include, for example, the acceleration and yaw rate of the corresponding road user 130. The output of the target trajectory generator 430 includes one or more predicted trajectories 432, the one or more predicted trajectories 432 including the corresponding positions of the road user 130 along the trajectory and information regarding the dynamic attributes of the road user 130. Although the target trajectory generator 430 may also predict a probability score for a corresponding one of the predicted trajectories 432 in the predicted trajectories 432, it may not be able to correctly identify which one of the trajectories 432 in the trajectory 432 is realistic and consistent with the dynamic situation 125 and the static situation 160 of the road user 130 (i.e., consistent with the situation of the corresponding traffic scene).
[0084] The predicted trajectories 432, together with the output of the above-mentioned feature extraction network 410, are transmitted to the target trajectory discriminator 440 to evaluate each of the predicted trajectories 432 in the predicted trajectories 432. In other words, the trajectory evaluation step 340 is performed by the target trajectory discriminator 440. For training the evaluation, the target trajectory discriminator 440 also receives information regarding the true trajectory as the ground truth for the considered road user 130 (e.g., see Figure 1 ). That is, features related to the true trajectory from the backbone system (i.e., the deep neural network 210) are used to provide corresponding predicted trajectories, which are compared with the true or ground truth trajectories to learn to classify the predicted trajectories as true or false. During the inference process, the predicted trajectories 432 are affected by the classification learned by the target trajectory discriminator 440 in order to evaluate the corresponding predicted trajectories 432.
[0085] Specifically, the target end point generator 420, the target trajectory generator 430, and the target trajectory discriminator 440 are implemented as a generative adversarial network (GAN). The target trajectory discriminator 440 is configured to evaluate each predicted trajectory 432 by using the classification provided by the GAN.
[0086] For each predicted trajectory 432 in the predicted trajectories 432, the target trajectory discriminator 440 outputs a normalized confidence score 444. The confidence score 444 includes a number from 0 to 1 for each predicted trajectory 432 in the predicted trajectories 432 and describes the degree of feasibility or reality of the corresponding predicted trajectory 432.
[0087] The entire machine learning algorithm and the model relied on by the method according to the present disclosure are trained by a combination of mean squared error loss and Minmax loss, and this combination ensures the convergence of the training. The entire output of the method (i.e., the predicted trajectories 432 and the corresponding confidence scores 444 for each trajectory 432 in the trajectories 432) can be used to support the host planning task. For example, for the planning of the main vehicle 100 as Figure 1 schematically shown. That is, multiple feasible trajectories are generated for the main vehicle 100, and the likelihoods of these trajectories are evaluated (e.g., in a tree search-based method).
[0088] In summary, the multi-modal trajectory generation 330 (see Figure 3 ) uses a two-stage method to convert the multi-modal trajectory prediction problem into two smaller problems: multi-modal target prediction or end-point prediction and single-modal trajectory prediction, that is, the prediction of a single trajectory for each predicted end-point in the predicted end-points. In addition, a confidence score 444 is provided for each predicted trajectory 432 in the predicted trajectories 432 to provide information about the reliability of the predicted trajectory 432.
[0089] Figure 5 A flowchart 500 is shown. The flowchart 500 illustrates a method for determining and evaluating the trajectory of a road user. At 502, input data associated with the movement state and environment of the road user can be received. At 504, features associated with the road user can be extracted from the input data. At 506, one or more trajectory end-points can be determined for the road user by using the extracted features. At 508, for each trajectory end-point, a corresponding trajectory associated with one of the trajectory end-points can be determined by using the associated trajectory end-point and the extracted features. At 510, the corresponding trajectories of the road user can be evaluated by using a classification that depends on the extracted features to provide a confidence score for each trajectory associated with one of the trajectory end-points
[0090] According to various embodiments, the method can rely on at least one machine learning algorithm; the input data can include embedded features provided by a prediction algorithm and can include information about other road users and information about the static environment of the road user; and a feature extraction algorithm can be applied to the embedded features to provide the extracted features.
[0091] According to various embodiments, determining one or more trajectory endpoints for a road user may include: applying a machine learning algorithm configured to learn to encode the extracted features and ground truth data of the true trajectory into at least one latent variable.
[0092] According to various embodiments, determining one or more trajectory endpoints for a road user by using the extracted features may further include: sampling at least one latent variable from a prior distribution provided by the machine learning algorithm when learning to encode the extracted features and the ground truth data of the true trajectory; and decoding the at least one latent variable into one or more trajectory endpoints.
[0093] According to various embodiments, determining the corresponding trajectory of a road user may include: determining the position of the road user from the road user's current position to the position of the associated trajectory endpoint, and determining the dynamic parameters of the road user for each of the positions.
[0094] According to various embodiments, determining the corresponding trajectory of a road user may include: applying a machine learning algorithm that may be configured to learn the parameters of a bicycle model associated with the position of the road user along the trajectory.
[0095] According to various embodiments, the machine learning algorithm may be trained for the evaluation step regarding the trajectory by: providing ground truth trajectories and predicted trajectories for a plurality of road users, where the predicted trajectories may be determined by applying the extracted features derived from training input data associated with the respective movement states and respective environments of each of the plurality of road users; and learning the classification of road users by applying a generative adversarial network to the ground truth trajectories and the predicted trajectories.
[0096] According to various embodiments, a generative adversarial network may be used to provide a classification for evaluating the corresponding trajectories of road users.
[0097] According to various embodiments, a general training process of the machine learning algorithm may be executed for the steps of determining and evaluating the corresponding trajectories of road users.
[0098] Each of steps 502, 504, 506, 508, and 510 and the further steps described above may be performed by computer hardware components.
[0099] Figure 6Illustrates a trajectory determination and evaluation system 600 according to various embodiments. The trajectory determination and evaluation system 600 may include an input data receiving circuit 602, an extraction circuit 604, an end point determination circuit 606, a trajectory determination circuit 608, and a trajectory evaluation circuit 610.
[0100] The input data receiving circuit 602 may be configured to: receive input data associated with the movement state and environment of a road user. The extraction circuit 604 may be configured to: extract characteristics related to the road user from the input data. The end point determination circuit 606 may be configured to: determine one or more trajectory end points for the road user by using the extracted characteristics. The trajectory determination circuit 608 may be configured to: for each of the trajectory end points, determine a corresponding trajectory associated with one of the trajectory end points by using the associated trajectory end point and the extracted characteristics. The trajectory evaluation circuit 610 may be configured to: evaluate the corresponding trajectory of the road user by using a classification that depends on the extracted characteristics to provide a confidence score for each trajectory associated with one of the trajectory end points.
[0101] The input data receiving circuit 602, the extraction circuit 604, the end point determination circuit 606, the trajectory determination circuit 608, and the trajectory evaluation circuit 610 may be coupled to each other, for example, via an electrical connection 612 (such as a cable or a computer bus) or via any other suitable electrical connection to exchange electrical signals.
[0102] "Circuit" may be understood as any type of logical implementation entity, which may be a dedicated circuit or a processor that executes a program stored in a memory, firmware, or any combination thereof.
[0103] Figure 7 Illustrates a computer system 700 having a plurality of computer hardware components configured to perform steps of a computer-implemented method for integrating a radar sensor in a vehicle according to various embodiments. The computer system 700 may include a processor 702, a memory 704, and a non-transitory data storage 706.
[0104] The processor 702 may execute instructions provided in the memory 704. The non-transitory data storage 706 may store a computer program that includes instructions that may be transferred to the memory 704 and then executed by the processor 702.
[0105] The processor 702, the memory 704, and the non-transitory data storage 706 may be coupled to each other, for example, via an electrical connection 708 (such as, for example, a cable or a computer bus) or via any other suitable electrical connection to exchange electrical signals.
[0106] Thus, the processor 702, the memory 704, and the non-transitory data storage 706 can represent the input data receiving circuit 602, the extraction circuit 604, the end point determination circuit 606, the trajectory determination circuit 608, and the trajectory evaluation circuit 610 as described above.
[0107] The terms “coupled” or “connected” are respectively intended to include direct “coupling” (e.g., via a physical link) or direct “connection” as well as indirect “coupling” or indirect “connection” (e.g., via a logical link).
[0108] It should be understood that what is described above for one of the methods can be similarly applied to the trajectory determination and evaluation system 600 and / or the computer system 700.
[0109] List of Reference Numerals
[0110] 100 Vehicle
[0111] 110 Sensing System
[0112] 115 Field of View
[0113] 120 Computer System
[0114] 121 Processing Unit
[0115] 122 Memory, Database
[0116] 124 Control System
[0117] 125 Dynamic Situation
[0118] 130 Road User
[0119] 140 Vehicle
[0120] 150 Pedestrian
[0121] 160 Static Situation
[0122] 170 Traffic Sign
[0123] 180 Lane Marking
[0124] 210 Deep Neural Network
[0125] 220 Ego Dynamics of the Host Vehicle
[0126] 230 Trunk
[0127] 240 Ground Truth
[0128] 250 Loss Function
[0129] 310 Embedded Feature
[0130] 320 Feature extraction
[0131] 330 Trajectory generation
[0132] 340 Trajectory evaluation
[0133] 410 Feature extraction network
[0134] 420 Target end - point generator
[0135] 422 Predicted end - point
[0136] 430 Target trajectory generator
[0137] 432 Predicted trajectory
[0138] 440 Target trajectory discriminator
[0139] 444 Confidence score
[0140] 500 Flowchart showing a method for determining and evaluating the trajectory of a road user
[0141] 502 Step: Receive input data associated with the movement state and environment of a road user
[0142] 504 Step: Extract features related to the road user from the input data
[0143] 506 Step: Determine one or more trajectory end - points for the road user by using the extracted features
[0144] 508 Step: For each of the trajectory end - points, determine the corresponding trajectory associated with one of the trajectory end - points by using the associated trajectory end - point and the extracted features
[0145] 510 Step: Evaluate the corresponding trajectory of the road user by using a classification that depends on the extracted features to provide a confidence score for each trajectory associated with one of the trajectory end - points 600 Trajectory determination and evaluation system
[0146] 602 Input data receiving circuit
[0147] 604 Extraction circuit
[0148] 606 End - point determination circuit
[0149] 608 Trajectory determination circuit
[0150] 610 Trajectory evaluation circuit
[0151] 612 Connection
[0152] 700 Computer system
[0153] 702 Processor
[0154] 704 Memory
[0155] 706 Non-Transitory Data Storage
[0156] 708 Connection
Claims
1. A computer-implemented method for determining and evaluating a trajectory of a road user (130), the method comprising: receiving input data (230, 310) associated with a road user's (130) mobility state and environment; extracting characteristics related to the road user from the input data (230, 310); determining one or more trajectory endpoints (422) for the road user (130) by using the extracted characteristics; For each of the trajectory endpoints (422), determining a corresponding trajectory (432) associated with one of the trajectory endpoints (422) by using the associated trajectory endpoint (422) and the extracted characteristics; The respective trajectories (432) of the road users (130) are evaluated using a classification dependent on the extracted characteristics to provide a confidence score (444) for each trajectory (432) associated with one of the trajectory endpoints (422).
2. The method according to claim 1, characterized in that: The method relies on at least one machine learning algorithm; The input data (230, 310) includes embedded features (310) provided by a prediction algorithm (210) and including information about other road users (130) and information about the road users' static environment (160); and A feature extraction algorithm (410) is applied to the embedded features (310) to provide the extracted characteristics.
3. The method according to claim 1, characterized in that: Determining one or more trajectory endpoints (422) for the road user (130) includes applying a machine learning algorithm configured to learn to encode the extracted characteristics and ground truth data of the real trajectory into at least one latent variable.
4. The method according to claim 3, characterized in that: Determining one or more trajectory endpoints (422) for the road user (130) by using the extracted characteristics further comprises: In learning to encode the extracted features and the ground truth data of the real trajectory, sampling the at least one latent variable from a prior distribution provided by the machine learning algorithm; and The at least one latent variable is decoded into one or more trajectory endpoints (422).
5. The method according to claim 2, characterized in that: Determining the corresponding trajectory (432) of the road user (130) includes: determining the position of the road user (130) starting from the current position of the road user (130) to the associated trajectory end point (422), and determining a dynamic parameter of the road user (130) for each of the positions.
6. The method according to claim 5, characterized in that: Determining the corresponding trajectory (432) of the road user (130) includes applying a machine learning algorithm configured to learn parameters of a bicycle model associated with the position of the road user (130) along the trajectory (432).
7. The method according to any one of claims 2 to 6, characterized in that: The machine learning algorithm is trained for the evaluation step of the trajectory (432) by the following steps: providing ground truth trajectories and predicted trajectories for a plurality of road users, wherein the predicted trajectories are determined by applying extracted features derived from training input data associated with a respective movement state and a respective environment for each of the plurality of road users; and The classification of the road user (130) is learned by applying a generative adversarial network to the ground truth trajectory and the predicted trajectory.
8. The method according to any one of claims 2 to 6, characterized in that: The machine learning algorithm includes a generative adversarial network for providing the classification for evaluating the corresponding trajectory (432) of the road user (130).
9. The method according to any one of claims 2 to 6, characterized in that: A general training process of the machine learning algorithm is performed for the steps of determining and evaluating the respective trajectories of the road users (130).
10. A computer system (700) configured to perform the computer-implemented method of at least one of claims 1 to 9.
11. The computer system (700) of claim 10, comprising: an endpoint generator (420) configured to: determine one or more trajectory endpoints (422) for a road user (130) by using characteristics extracted from the input data (230, 310); a trajectory generator (430) configured to: for each of the trajectory endpoints (422), determine a corresponding trajectory (432) associated with one of the trajectory endpoints (422) by using the associated trajectory endpoint (422) and the extracted characteristic; An identifier (440) is configured to evaluate each trajectory (432) associated with one of the trajectory endpoints (422).
12. The computer system (700) of claim 11, wherein: The endpoint generator (420), the trajectory generator (430) and the discriminator (440) are implemented as a generative adversarial network; The discriminator (440) is configured to evaluate each trajectory (432) by using the classification provided by the generative adversarial network.
13. A vehicle (100) comprising a perception system (110) and a computer system (700) as claimed in any one of claims 10 to 12.
14. The vehicle (100) of claim 13, further comprising a control system (124) configured to control an actual trajectory of the vehicle (100), in, The computer system is configured to transmit at least one evaluated trajectory (432) of the road user (130) to the control system (124) so that the control system (124) can apply the evaluated trajectory (432) of the road user (130) when controlling the actual trajectory of the vehicle (100).
15. A non-transitory computer-readable medium comprising instructions for implementing the computer-implemented method of at least one of claims 1 to 10.