Autopilot Trajectory Prediction and Planning Method Based on Time Series Autoregressive Model

Through the trajectory prediction method based on the timing autoregression model, discrete the agent trajectory and use the Transformer architecture and reinforcement learning, the consistency and OOD problems of trajectory planning in multi-agent scenarios are solved, and trajectory prediction and planning with higher accuracy and safe trajectory prediction and planning are achieved.

CN119781485BActive Publication Date: 2025-07-08ZHEJIANG YOULU ROBOT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510282830.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-08
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The existing trajectory prediction and planning models are difficult to generate consistent trajectories in multi-agent scenarios, and cannot effectively predict and plan the interactive behavior between agents. There are external distribution (OOD) problems, which affect the accuracy and reliability of trajectory planning.

Method used

The trajectory prediction method based on the timing autoregression model is adopted to construct a motion marker vocabulary library through discrete agent trajectory, and the Transformer architecture and reinforcement learning method are used to generate the future trajectory of the agent, alleviate the OOD problem, and improve the trajectory prediction accuracy and stability.

Benefits of technology

It improves the accuracy and stability of trajectory prediction in multi-agent scenarios, significantly improves the interactive modeling capabilities of the agent, reduces the risk of collision, and enhances the safety and efficiency of trajectory planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119781485B_ABST
    Figure CN119781485B_ABST
Patent Text Reader

Abstract

The present invention discloses an autonomous driving trajectory prediction and planning method based on a time series autoregressive model. The method includes: acquiring data of a current vehicle agent and preprocessing the data; discretizing the agent trajectory and constructing a motion marker vocabulary library; inputting the agent trajectory and map information into a time series autoregressive trajectory generation model based on the Transformer architecture and performing trajectory prediction in an autoregressive manner; fine-tuning the model using a reinforcement learning method; and generating the future trajectory of the agent. The beneficial effects of the present invention: Compared with the prior art, the beneficial effects of the present invention are: it can significantly improve the prediction accuracy and stability, can significantly improve the modeling ability of the interaction between agents, and through continuous interaction with the environment, reinforcement learning uses a reward and punishment mechanism to dynamically adjust the behavior strategy of the agent, encouraging the agent to choose a safer and more effective path and avoiding potential dangerous behaviors caused by the incompleteness of training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and specifically to an autonomous driving trajectory prediction and planning method based on a time series autoregressive model. Background Art

[0002] With the continuous progress of deep learning technology, data-driven autonomous driving trajectory prediction and planning methods have shown great potential in solving planning problems in complex scenarios. Compared with traditional algorithms, data-driven trajectory prediction and planning methods can generate intelligent decision-making paths by learning the driving habits and styles of human drivers from a large amount of data, thus eliminating the need for manual judgment of complex algorithm logics in practical applications. However, existing trajectory prediction and planning models usually have difficulty generating consistent trajectories in multi-agent scenarios and cannot effectively predict and plan the interaction behaviors between agents. In addition, data-driven models based on supervised learning have the "Out of Distribution" (hereinafter referred to as OOD) problem, that is, the model may gradually accumulate small errors in data or neural network outputs during the iterative planning process, resulting in significant deviations in the planning results and even eventually leading to collision accidents. Existing supervised learning methods have not been able to effectively solve this problem, thus affecting the accuracy and reliability of trajectory planning. Summary of the Invention

[0003] The purpose of the present invention is to provide an autonomous driving trajectory prediction and planning method based on a time series autoregressive model to solve the problems raised in the above background art.

[0004] To achieve the above purpose, the present invention provides the following technical solutions:

[0005] In a first aspect, the present invention provides an autonomous driving trajectory prediction and planning method based on a time series autoregressive model, and the method includes:

[0006] Obtain data of the current vehicle agent and preprocess the data;

[0007] Discretize the agent trajectory and construct a motion label vocabulary;

[0008] Input the agent trajectory and map information into a time series autoregressive trajectory generation model based on the Transformer architecture and perform trajectory prediction in an autoregressive manner;

[0009] Fine-tune the model using a reinforcement learning method;

[0010] Generate the future trajectory of the agent.

[0011] Preferably, the obtaining of the data of the current vehicle agent at least includes: agent trajectory information in the scenario, map lane centerline information, map lane boundary line information, and traffic light information.

[0012] Preferably, the discretization of the multi-agent trajectories and the construction of the motion token vocabulary include:

[0013] Normalize the trajectories of each agent so that they are standardized with respect to the position and orientation at the initial time;

[0014] Discretize the incremental actions of the trajectories into a series of predefined discrete values, and construct a discretized action vocabulary by uniformly quantizing each coordinate axis;

[0015] Map each continuous trajectory increment to a corresponding index, and convert the continuous actions of the trajectory into a discrete token sequence;

[0016] Merge the incremental actions of each coordinate into a single index value as the motion identifier.

[0017] Preferably, the input of the agent trajectory and the map information into the temporal autoregressive trajectory generation model based on the Transformer architecture and the use of the autoregressive method for trajectory prediction include:

[0018] Receive the initial state information of the scenario through the encoder and generate a shared scenario embedding;

[0019] Predict the behaviors of all agents at the current time step based on the output of the previous time step and the fixed scenario embedding;

[0020] The decoder generates a distribution of N output identifiers of the agents in the scenario by inputting a set of motion identifiers and the scenario embedding at each prediction step.

[0021] Preferably, the decoder generates a distribution of N output identifiers of the agents in the scenario by inputting a set of motion identifiers and the scenario embedding at each prediction step, including:

[0022] Divide the decoder into multiple layers, apply the self-attention mechanism between the input identifiers in each layer, and apply the cross-attention mechanism to the scenario embedding;

[0023] Set all N motion identifiers in step t to be able to cross-attention, and can attention to all previous identifiers, where each row represents a query identifier, each column represents a key identifier, and the green blocks represent the key identifiers that the query can attention to;

[0024] In the running After the prediction step, obtain , the output motion identifier forms the complete trajectories of N agents.

[0025] Preferably, the method for fine-tuning the model using the reinforcement learning method includes:

[0026] Modeling the generation of the complete trajectories of N agents as a multi-agent Markov decision process;

[0027] Designing the reward function as whether each agent collides with the map boundary line or other agents at each time step at each moment, and its calculation formula is:

[0028] where a and ß are the weight coefficients for collision with the road boundary and collision with other agents respectively,

[0029] Setting the model to autoregressively unfold actions within T_pred time steps. For step t and each agent i, the return is: where is the reward of agent i at step t', and γ is the discount factor used to weight future rewards;

[0030] Using the REINFORCE algorithm to calculate the policy gradient and performing reinforcement learning training to further optimize the decision-making ability of the model.

[0031] In a second aspect, the present invention provides an automatic driving trajectory prediction and planning device for implementing the method described in any of the above embodiments. The device includes:

[0032] An acquisition module, which is used to acquire data of the current vehicle agent and preprocess the data;

[0033] A construction module, which is used to discretize the agent trajectories and construct a motion label vocabulary library;

[0034] A prediction module, which inputs the agent trajectories and map information into a temporal autoregressive trajectory generation model based on the Transformer architecture and performs trajectory prediction in an autoregressive manner;

[0035] A fine-tuning module, which fine-tunes the model using the reinforcement learning method;

[0036] A generation module, which is used to generate future trajectories of the agent.

[0037] In a third aspect, the present invention provides an electronic device including at least one processor, and the processor is communicatively connected to at least one memory. Among them, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any of the above embodiments.

[0038] In a fourth aspect, the present invention provides a computer-readable storage medium storing computer instructions for causing a processor to execute the method according to any one of the above embodiments when executed.

[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0040] 1. Multi-agent trajectory tokenization: The present invention innovatively adopts a multi-agent joint trajectory tokenization method, representing the trajectory of each agent as a set of discretized Tokens. By discretizing the continuous movement of the trajectory into a series of quantified actions, this method not only improves the accuracy of trajectory representation but also transforms the trajectory prediction and planning problems into a relatively simple sequence prediction problem, i.e., predicting the next Token. This Token-sequence representation provides a theoretical basis for the application of time series autoregressive models and can effectively capture the interaction relationships between agents. Compared with traditional trajectory prediction methods, the tokenization method of the present invention can use deep learning models to more efficiently model trajectories, especially in complex multi-agent scenarios, significantly improving prediction accuracy and stability. Specifically, after representing the trajectory as a combination of Tokens, the model can utilize the characteristics of time series autoregression to sequentially predict the state of each agent at future times, thus achieving more accurate and consistent multi-agent trajectory planning.

[0041] 2. Modeling the interaction process of agents using a time series autoregressive model: The Transformer architecture included in the time series autoregressive model can efficiently capture the interaction process between agents using the attention mechanism, thereby explicitly modeling the interaction game of agents. By transforming the trajectory prediction problem into a time series prediction problem, the time series autoregressive model can significantly improve the modeling ability of the interaction between agents using the self-attention mechanism of the Transformer. At each time step, the model dynamically adjusts the weights between agents using the attention mechanism based on the current state of the agent and the states of other agents, thereby accurately capturing the role and influence of each agent in complex interactions. In this way, the model can not only predict the future trajectory of a single agent but also accurately simulate the interaction game process between multiple agents.

[0042] 3. Alleviating the OOD problem: By introducing a reinforcement learning algorithm, the present invention effectively alleviates the problem of performance degradation of traditional models when facing "out-of-distribution" (OOD) data. In the context of autonomous driving, models typically only encounter a limited number of environmental samples during training, resulting in potential prediction failures and decision-making errors when facing unseen environments in actual applications. Reinforcement learning dynamically adjusts the agent's behavior strategy through continuous interaction with the environment and using a reward and punishment mechanism, encouraging the agent to choose safer and more effective paths and avoiding potential dangerous behaviors caused by the incompleteness of training data. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a flowchart of the method according to an embodiment of the present invention;

[0044] Figure 2 It is a sub-flowchart of step S200 of the present invention;

[0045] Figure 3 It is a sub-flowchart of step S300 of the present invention;

[0046] Figure 4 It is a sub-flowchart of step S330 of the present invention;

[0047] Figure 5 It is a sub-flowchart of step S400 of the present invention;

[0048] Figure 6 It is a diagram of the device modules according to an embodiment of the present invention;

[0049] Figure 7 An example diagram of the operating scenario of the present invention;

[0050] Figure 8 It is a diagram of the electronic device modules of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0052] Please refer to Figure 1 , the present invention provides a technical solution: A method for autonomous driving trajectory prediction and planning based on a time series autoregressive model, including:

[0053] S100: Obtain the data of the current vehicle agent and preprocess the data.

[0054] In this embodiment, intelligent agent trajectory information, map lane centerline information, map lane boundary line information, traffic light information, etc. in the scenario are collected and parsed. These raw data are processed, for example, it may include data cleaning (removing noise, outliers, etc.), data format conversion (converting data from different sensors or data sources into a unified format), coordinate system unification (making the intelligent agent trajectory information and map information in the same coordinate system), etc. operations to ensure the quality and usability of the data. Provide accurate, complete and appropriate input data for subsequent trajectory prediction and planning.

[0055] S200: Discretize the intelligent agent trajectory and construct a motion label vocabulary.

[0056] In this embodiment, the specific operation of trajectory discretization: The continuous trajectory of the intelligent agent is segmented according to certain rules (such as time interval or spatial distance), and it is converted into discrete small segments. For example, if segmented by time interval, a trajectory point can be taken every certain time (such as 0.5 seconds), thus obtaining a series of discrete trajectory point sequences.

[0057] The specific operation of constructing the motion label vocabulary includes: Defining a series of labels that can describe the motion state of the intelligent agent, such as forward, backward, left turn, right turn, acceleration, deceleration, uniform speed, etc. Then, according to the characteristics of the discrete trajectory segments, each segment is represented by the corresponding motion label. For example, if the coordinates of the intelligent agent in a trajectory segment gradually increase in the horizontal direction and the speed is stable, it can be labeled as "uniform forward". In this way, all possible motion state labels are summarized to form a motion label vocabulary, preparing for representing the trajectory as an identifier (Token) and inputting it into the model later.

[0058] S300: Input the intelligent agent trajectory and map information into a temporal autoregressive trajectory generation model based on the Transformer architecture and perform trajectory prediction using the autoregressive method.

[0059] In this embodiment, specifically, it includes data input, encoding, and decoding processes, where:

[0060] Data input: Input the preprocessed map information (such as lane centerline, boundary line, traffic light status, etc.) and intelligent agent motion information (discretized trajectory and related states) into a temporal autoregressive trajectory generation model based on the Transformer architecture.

[0061] Encoding process: The encoder of the model takes the representation of the initial state of the scenario (including traffic light status, map topology information, and historical trajectory data of all traffic participants, etc.) as input and generates a scene embedding. This scene embedding can capture the key information in the scenario and provide a context environment for subsequent trajectory prediction.

[0062] Autoregressive decoding prediction: In the decoder stage, an autoregressive model based on Transformer is used to iteratively generate the distribution of the joint trajectories of agents at future time steps to be predicted. At each prediction step, the decoder takes as input a set of motion tokens (from the previous step or the initial input) and the scene embedding, and generates the distribution of all agent tokens in the scene. In an autoregressive manner, that is, the prediction at each time step is based on the output of the previous time step, the model can gradually predict the trajectory states of agents at future time steps and can take into account the interaction relationships between agents. For example, when predicting the trajectory of agent A at a certain moment, the previous trajectories and current states of other agents (such as agent B, C, etc.) will be considered to more accurately predict agent A.

[0063] S400: Fine-tune the model using reinforcement learning methods.

[0064] In this embodiment, specifically, first model the problem: Model the trajectory prediction and planning problem in the autonomous driving scenario as a multi-agent Markov decision (MDP) process, and clarify elements such as actions, states, dynamic transition processes, and observations.

[0065] Actions: The action space of the agent is defined as an incremental action space, and the action of each agent is represented as its acceleration in the X and Y directions at the current moment, and these accelerations are discretized into a uniformly quantized grid and represented by discrete indices.

[0066] States: The state space includes all information describing the agent and its surrounding environment, such as map features, the joint state of agents, and traffic light states.

[0067] Dynamic transition process: Based on the motion laws of the physical model, according to the current state (position and speed) of the agent and the actions taken (acceleration), calculate its position at the next moment, thereby defining the rule for the agent to transfer from one state to the next state.

[0068] Observations: The observation space contains historical context (all objects, traffic light states, and map features at historical time steps), the historical actions of all objects, and the identity identifier of the current agent, helping the agent understand the current scene and environmental state.

[0069] Furthermore, the reward function is designed according to whether each agent collides with the map boundary line or other agents at each moment. By setting an indicator function representing the collision of the agent with the boundary and other agents at the moment, and \(\omega_i\) as the corresponding weight coefficient, encourage the agent to choose a safer path and avoid collisions.

[0070] Furthermore, the REINFORCE algorithm with a baseline, a policy gradient method, is used to further optimize the decision-making ability of the model. The REINFORCE algorithm with a baseline reduces the variance of policy updates by introducing the advantage function, thus accelerating the learning process. While training the model, a state-value function network is trained to evaluate the value of the current state and serves as a baseline for calculating the advantage function, which represents the advantage of the agent taking a certain action in that state. Then, based on the calculated gradient information, the policy network is updated to gradually optimize the agent's behavioral decision-making, enabling the model to learn a better trajectory generation policy and make more reasonable decisions when facing complex and changing environments, thereby improving the safety and efficiency of trajectory planning.

[0071] S500: Generate the future trajectories of the agent.

[0072] In this embodiment, the output of the temporal autoregressive trajectory generation model is the probability distribution of the future trajectories of each agent. The specified-length future trajectory Tokens are obtained from the probability distribution through sampling or search methods. The sampling method can randomly sample according to the probability distribution and iteratively generate the future trajectory Tokens; the search method selects appropriate Tokens according to the probability distribution and also iteratively generates the specified-length future trajectory Tokens. Finally, the Tokens are restored to the agent trajectories using the trajectory motion marking modeling method, and the number of generated scenarios can be customized according to task requirements. In this way, the specific future trajectories available for autonomous vehicles are obtained, and the vehicle can make driving decisions based on these trajectories, such as choosing the appropriate lane, controlling the vehicle speed, planning turns, etc., to achieve safe and efficient autonomous driving.

[0073] In an embodiment of the present invention, the data obtained in step S100 for the current vehicle agent at least includes: the agent trajectory information in the scenario, the map lane centerline information, the map lane boundary line information, and the traffic light information.

[0074] In this embodiment, when autonomous driving trajectory prediction and planning are required, the data needs to be preprocessed first to parse the input data required by the model. For example, different processing for different data is as follows:

[0075] Agent trajectory information: Collect the trajectory data of agents (such as vehicles) in the scenario. These data record the position and other information of the agents at different time points and are the basis for subsequent processing and prediction.

[0076] Map lane centerline information: Obtain the centerline data of the lanes on the map, which helps the model understand the direction and position of the lanes, thereby better planning the driving trajectory of the vehicle within the lanes.

[0077] Map lane boundary line information: Clearly define the boundary range of the lane, enabling the model to know the boundary of the area where the vehicle can travel and avoiding the planning of unreasonable trajectories that exceed the lane range.

[0078] Traffic light information: Includes the position, status (such as red light, green light, yellow light, etc.) and switching time of traffic lights. This information is crucial for trajectory planning of vehicles in traffic scenarios such as intersections. For example, a vehicle needs to decide whether to stop and wait or continue to drive based on the traffic light status.

[0079] In one embodiment of the present invention, please refer to Figure 2 , step S200 specifically includes the following steps:

[0080] S210: Normalize the trajectory of each agent so that it is standardized with respect to the position and orientation at the initial moment.

[0081] In this embodiment, assume that in a two-dimensional plane coordinate system, for the trajectory of each agent, with its position at the initial moment (for example, the initial coordinates ( )) as the origin, perform relative transformation on the coordinates of subsequent trajectory points. At the same time, consider the initial orientation of the agent (which can be represented by an angle, such as the angle between the initial direction of the vehicle's head and the positive direction of the axis ), and perform transformations such as rotation and translation on the coordinates of the trajectory points according to this initial orientation, so that the trajectories of all agents seem to start moving from the same position and orientation. The advantage of this is that when performing subsequent trajectory increment calculation, quantization, and discretization, the trajectory changes between different agents are comparable, and data processing deviation will not be caused due to differences in initial position and orientation. For example, in a multi-vehicle autonomous driving scenario, different vehicles may start at different positions and directions, but after normalization, their trajectories can be analyzed and processed in a unified manner, providing a consistent basis for accurately constructing a motion label vocabulary. This step ensures that the trajectories of all agents start from a unified reference point, laying the foundation for subsequent quantization and discretization operations.

[0082] S220: Discretize the incremental actions of the trajectory into a series of predefined discrete values, and the discrete values construct a discrete action vocabulary by uniformly quantizing each coordinate axis.

[0083] In this embodiment, for the incremental actions of the trajectory on each coordinate axis (such as the x-axis and y-axis), that is, the coordinate differences Δx and Δy of two adjacent trajectory points on the coordinate axis, a certain quantization interval is set. For example, for the increment in the x-axis direction, if the quantization intervals are set as [-5,3), [-3,-1), [-1,1), [1,3), [3,5], then when a trajectory increment Δx = 2, it will be mapped to the discrete value 3 (corresponding to the interval [1,3)). Similarly, a similar quantization operation is performed on the y-axis. In this way, the possible continuous increment values on each coordinate axis are converted into a finite number of discrete values, and then these discrete values are combined to construct a discretized action vocabulary. For example, in the above example, if the quantization result of the y-axis is the discrete value 2, then a combination (3,2) can be used to represent the incremental action of this trajectory point, and this combination becomes an element in the action vocabulary. In this way, the originally continuously changing trajectory increment can be represented by a finite number of discrete combinations (labels), enabling the trajectory movement to be described in a more concise and regular manner, which provides convenience for subsequent trajectory modeling and prediction.

[0084] S230: Map each continuous trajectory increment to a corresponding index, converting the continuous actions of the trajectory into a discrete token sequence.

[0085] In this embodiment, based on the discretized action vocabulary constructed in the foregoing steps, a unique index is assigned to each possible discrete value combination (label). For example, in the case where the trajectory increment is represented by (3,2) as mentioned above, if this combination is defined as the 5th element in the action vocabulary, then its corresponding index is 5. For the entire trajectory, the incremental actions of each trajectory point are mapped to indexes in this way, and an index sequence is obtained, and this index sequence is the discrete token sequence. For example, the trajectory of an agent consists of a series of trajectory points. After quantization and index mapping of the incremental actions of each trajectory point, an index sequence such as [2,4,1,3,…] is obtained, and this sequence represents the trajectory movement of the agent, converting the continuous actions of the trajectory into a discrete token sequence that the model can process in a concise and orderly manner. In subsequent model inputs, this token sequence can be used as input data, and the model can learn and predict the future trajectory of the agent based on this sequence because the pattern and regular information of the agent's trajectory movement are contained in this sequence.

[0086] In this embodiment, in order to improve the accuracy of trajectory reconstruction, a greedy search strategy is used to select the quantization action that can minimize the error at each time step. To further simplify dynamic modeling, the Verlet step method is adopted, where when the action is zero, it means that the current increment is the same as the previous step, which helps to reduce the size of the vocabulary space and maintain the smoothness of the speed change.

[0087] S240: Combine the incremental actions of each coordinate into a single index value as the motion identifier. This reduces the representation complexity of the model. In the previous steps, the incremental actions of the trajectory on each coordinate axis (such as the x-axis and y-axis) have been discretized and index mapped separately. In step S240, the indexes corresponding to the incremental actions of each coordinate are combined. For example, assume that for a certain trajectory point, the index after mapping the incremental action of its x-axis is , and the index after mapping the incremental action of the y-axis is . These two indexes can be combined into a single index value i through a certain predefined combination rule. A simple combination method can be to splice the two indexes (such as i = i x× N + i y , where N is the number of discrete values of the y-axis, which can ensure the uniqueness of the combined index value), or other coding methods can be used as long as the information of the two dimensions can be effectively integrated into one value.

[0088] Through this combination operation, the model that originally needed to process the indexes in two coordinate directions separately now only needs to process a single index value. For the entire trajectory, the incremental actions of all trajectory points are combined into a single index value sequence in this way, which greatly reduces the dimension and complexity of the model input data. For example, in a scenario containing multiple agent trajectories, if no combination is performed, the model needs to process two index sequences of the x-axis and y-axis of each agent trajectory, with a large amount of data and a relatively complex structure; after combination, the model only needs to process a single index sequence, which is more convenient and fast to process, can improve the training and prediction speed of the model, and at the same time reduces the demand for storage resources, helping to optimize the performance of the entire autonomous driving trajectory prediction and planning system.

[0089] In an embodiment of the present invention, please refer to Figure 3 , S300 specifically includes:

[0090] S310: Receive the initial state information of the scene through an encoder and generate a shared scene embedding;

[0091] S320: Predict the behaviors of all agents at the current time step based on the output of the previous time step and the fixed scene embedding;

[0092] S330: The decoder generates a distribution of N output identifiers of the agents in the scene by inputting a set of motion identifiers and the scene embedding in each prediction step.

[0093] In an embodiment of the present invention, please refer to Figure 4 , S330 further includes:

[0094] S3310: Divide the decoder into multiple layers, apply the self-attention mechanism between the input identifiers for each layer, and apply the cross-attention mechanism to the scene embedding;

[0095] S3320: Set that all N motion identifiers in step t can pay attention to each other and can pay attention to all previous identifiers, where each row represents a query identifier and each column represents a key identifier, and the green blocks represent the key identifiers that the query can pay attention to;

[0096] S3330: After running times of prediction steps, obtain , and the output motion identifiers form the complete trajectories of N agents.

[0097] In a specific embodiment of the present invention, the proposed temporal autoregressive trajectory generation model is based on the Transformer architecture and uses the autoregressive method for trajectory prediction. The encoder of the model receives a set of identifiers (tokens) representing the initial state of the scene and generates a shared scene embedding. These initial states include the states of traffic lights, map topology information, and the historical trajectory data of all traffic participants. During model inference, the decoder predicts the behaviors of all agents at the current time step in an autoregressive manner based on the output of the previous time step and the fixed scene embedding. In each prediction step, the decoder generates the distribution of N output identifiers by inputting a set of motion identifiers and the scene embedding.

[0098] In each prediction step, the decoder receives a set of motion identifiers and the scene embedding as input and generates the distribution of N output identifiers. The decoder consists of multiple layers, and each layer applies self-attention between the input identifiers and cross-attention to the scene embedding. All N motion tokens in step t can pay attention to each other and can pay attention to all previous tokens, where each row represents a query token and each column represents a key identifier (token), and the green blocks represent the key identifiers (tokens) that the query can pay attention to. After running times of prediction steps, the obtained , and the output motion tokens form the complete trajectories of N agents. This autoregressive method ensures that the actions of each agent are based on the temporal causal relationship with the previous actions of all traffic participants, thereby improving the modeling effect of the interaction between agents within the prediction time range.

[0099] Compared with traditional trajectory generation methods, the sequential autoregressive trajectory generation model of the present invention can accurately predict the trajectories of each agent in a scenario of multi-agent interaction, and can fully consider the mutual influence between agents, thereby improving the prediction accuracy and robustness of the model. Through the autoregressive decoding process, the model can not only generate trajectories that conform to actual traffic behaviors, but also better handle the dynamically changing traffic environment, achieving more efficient trajectory planning and prediction.

[0100] In one embodiment of the present invention, refer to Figure 5 , S400 includes:

[0101] S410: Model the complete trajectory generation of N agents as a multi-agent Markov decision process (MDP).

[0102] In one embodiment of the present invention, the components of the multi-agent Markov decision include actions, states, dynamic transition processes, and observations. Specifically, the functions of each component are as follows:

[0103] Actions: The action space of the agent is defined as the incremental action space. The action of each agent is represented by its acceleration in the X and Y directions at the current moment. In actual operation, these accelerations are discretized into a uniformly quantized grid, where each action is represented by a discrete index.

[0104] States: The state space is all the information describing the agent and its surrounding environment. In this model, the state space includes map features, the joint state of the agents, and traffic light states, etc.

[0105] Dynamic transition process: The dynamic transition matrix defines the rules for the agent to transfer from one state to the next state. In this model, the state transition of the agent calculates the position of the agent at the next moment based on the current state (i.e., the current position and speed of the agent) and the action taken (acceleration). This transition process is based on the motion laws of the physical model and simulates the motion of the agent through discretized acceleration actions.

[0106] Observations: The observation space defines the information that the agent can perceive at each time step. In our method, the observations include: historical context, the historical actions of all objects, and the identity identifier of the current agent. The historical context includes all objects, traffic light states, and map features of historical time steps, helping the agent understand the current scenario and environmental state.

[0107] S420: Design the reward function as whether each agent collides with the map boundary line or other agents at each time step at each moment, and its calculation formula is:

[0108] Where a and ß are the weight coefficients for collisions with road boundaries and other agents respectively,

[0109] The model is set to autoregressively unfold actions within \(T_{pred}\) time steps. For step \(t\) and each agent \(i\), the reward is: where \(r_{i,t'}\) is the reward of agent \(i\) at step \(t'\), and \(\gamma\) is the discount factor used to weight future rewards;

[0110] S430: Use the REINFORCE algorithm to calculate the policy gradient and perform reinforcement learning training to further optimize the model's decision-making ability.

[0111] In the embodiment, the REINFORCE algorithm with a baseline reduces the variance of policy updates by introducing the advantage function, accelerating the learning process. While training the model, a state-value function network is trained to evaluate the value of the current state and serves as a baseline to calculate the advantage function, representing the advantage of an agent taking a certain action in that state , and the advantage of an agent taking a certain action in that state is calculated through the advantage function:

[0112] where \(Q(s,a)\) is the action value of the agent taking action \(a\) in state \(s\), and \(V(s)\) is the value function of state where, \(\pi(a|s)\) is the policy of agent \(i\) choosing action \(a\) at time step \(t\), and

[0113] \(A(s,a)\)

[0114] is the advantage function, representing the advantage of the agent currently taking a certain action. Based on this, by maximizing the following objective function, the REINFORCE algorithm is used to calculate the policy gradient:

[0115] Table 1: Test results of the method of the present invention on the Google self-driving open dataset (Waymo Open Dataset)

[0116]

[0117] The experimental results in Table 1 show that in the prediction task, the method provided by the present application can effectively generate the trajectories of all agents in the scenario, and the minimum average distance error of trajectory prediction is 1.39 meters. In the task of simulating the behavior of human drivers, the model provided by the present application can simulate the interactive movement of vehicles and conform to the characteristics of human drivers. In the planning task, compared with the basic model, the model fine-tuned by reinforcement learning has a 60% reduction in the trajectory collision rate of agents, significantly reducing the collision risk of agents.

[0118] Please refer to Figure 6 , Figure 6 which shows an automatic driving trajectory prediction and planning device provided by the present invention for implementing the method described in any one of the above embodiments. The device includes:

[0119] An acquisition module 100, which is used to acquire the data of the current vehicle agent and preprocess the data;

[0120] A construction module 200, which is used to discretize the agent trajectory and construct a motion marker vocabulary library;

[0121] A prediction module 300, which inputs the agent trajectory and map information into a time series autoregressive trajectory generation model based on the Transformer architecture and uses the autoregressive method for trajectory prediction;

[0122] A fine-tuning module 400, which fine-tunes the model using the reinforcement learning method;

[0123] A generation module 500, which is used to generate the future trajectory of the agent.

[0124] Through the above modules, an autoregressive prediction result can be generated for the driving scenario. Please refer to Figure 7 , Figure 7 which shows the future trajectory of the agent generated by autoregressive generation in time series by an automatic driving trajectory prediction and planning device provided by an embodiment of the present invention at a complex intersection, where the arrows on each line represent the future trajectory points of agents with different numbers and the corresponding orientation predictions.

[0125] Please refer to Figure 8 , Figure 8 which shows a schematic diagram of the mechanism of an electronic device 20 that can implement the embodiments of the present invention. The electronic device is intended to represent various forms of control devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described herein and / or claimed.

[0126] The electronic device 20 includes at least one processor 21, and a memory communicatively connected to the at least one processor 21, such as a read-only memory (ROM) 22, a random access memory (RAM) 23, etc. The memory stores computer programs executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 22 or the computer programs loaded from the storage unit 28 into the random access memory (RAM) 13. In the RAM 23, various programs and data required for the operation of the electronic device 20 can also be stored. The processor 21, the ROM 22, and the RAM 23 are connected to each other through a bus 24. An input / output (I / O) interface 25 is also connected to the bus 24.

[0127] Multiple components in the electronic device 20 are connected to the I / O interface 25, including: an input unit 26, such as a keyboard, a mouse, etc.; an output unit 27, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 29, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 29 allows the electronic device 20 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0128] The processor 21 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 21 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 21 executes the various methods and processes described above.

[0129] In some embodiments, the method of the above embodiments can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 28. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 20 via the ROM 22 and / or the communication unit 29. When the computer program is loaded into the RAM 23 and executed by the processor 21, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the processor 21 can be configured to execute the method of the above embodiments in any other appropriate manner (e.g., by means of firmware).

[0130] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0131] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0132] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0133] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0134] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0135] The computing system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The relationship between the client and the server is created by computer programs that run on the respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0136] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for autonomous driving trajectory prediction and planning based on a time series autoregressive model, characterized in that The method includes: Obtaining data of the current vehicle agent and preprocessing the data; Discretizing the agent trajectory and constructing a motion token vocabulary, specifically including: Normalizing the trajectory of each agent to standardize it with respect to the position and orientation at the initial moment; Taking the coordinate difference between two adjacent points of the trajectory on the coordinate axis as an incremental action, and by uniformly quantizing each coordinate axis, discretizing the incremental action into a series of predefined discrete values, and these discrete values are combined to form a discretized action vocabulary; Mapping each continuous trajectory increment to a corresponding index, thereby converting the continuous actions of the trajectory into a discrete token sequence; Combining the indices corresponding to the incremental actions of each coordinate into a single index value as the motion identifier; Inputting the agent trajectory and map information into a temporal autoregressive trajectory generation model based on the Transformer architecture and using the autoregressive method for trajectory prediction, specifically including: Receiving the initial state information of the scene through the encoder, which includes traffic signal states, map topology information, and historical trajectory data of all traffic participants, etc., and the encoder encodes this information to generate a shared scene embedding; Predicting the behaviors of all agents at the current time step based on the output of the previous moment and the fixed scene embedding part; The decoder generates a distribution of N output identifiers of the agents in the scene by inputting a set of motion identifiers and the above-generated scene embedding at each prediction step; Fine-tuning the model using the reinforcement learning method; Generating the future trajectory of the agent.

2. The method for predicting and planning an autonomous driving trajectory based on a time series autoregressive model according to claim 1, wherein The obtaining of the data of the current vehicle agent at least includes: agent trajectory information in the scene, map lane centerline information, map lane boundary line information, and traffic light information.

3. The method for predicting and planning an autonomous driving trajectory based on a time series autoregressive model according to claim 2, wherein: The decoder generates a distribution of N output identifiers of the agents in the scene by inputting a set of motion identifiers and the scene embedding at each prediction step, including: Dividing the decoder into multiple layers, applying the self-attention mechanism between the input identifiers in each layer, and applying the cross-attention mechanism to the scene embedding; Setting all N motion identifiers in step t to be able to pay attention to each other and be able to pay attention to all previous identifiers, where each row represents a query identifier, each column represents a key identifier, and the green blocks represent the key identifiers that the query can pay attention to; After running the prediction step T times, T pred × N is obtained, and the output motion identifiers form the complete trajectories of N agents.

4. The method for predicting and planning an autonomous driving trajectory based on a time series autoregressive model according to claim 3, wherein: The using of the reinforcement learning method to fine-tune the model includes: Modeling the generation of the complete trajectories of N agents as a multi-agent Markov decision process; Designing the reward function as whether each agent collides with the map boundary line or other agents at each time step at each moment, and its calculation formula is: r t,i = -α·I boundary (t, i) + β·I collision (t, i) where a and β are the weight coefficients for collision with the road boundary and collision with other agents respectively, Setting the model to autoregressively unfold actions within T_pred time steps. For step t and each agent i, the return is: where is the reward of agent i at step t', and γ is the discount factor used to weight future rewards; Using the REINFORCE algorithm to calculate the policy gradient for reinforcement learning training.

5. An automatic driving trajectory prediction and planning device for implementing the method according to any one of claims 1 to 4, characterized in that, The device includes: An acquisition module, which is used to acquire the data of the current vehicle agent and preprocess the data; A construction module, which is used to discretize the agent trajectory and construct a motion label vocabulary library; A prediction module, which inputs the agent trajectory and map information into a temporal autoregressive trajectory generation model based on the Transformer architecture and uses the autoregressive method for trajectory prediction; A fine-tuning module, which fine-tunes the model using a reinforcement learning method; A generation module, which is used to generate the future trajectory of the agent.

6. An electronic device, characterized in that, Comprising at least one processor, the processor is communicatively connected to at least one memory, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to implement the method according to any one of claims 1 to 4 when executed.

Citation Information

Patent Citations

  • Cooperative guidance control method for mixed vehicle fleet at entrance lane of intersection

    CN119169818A

  • Methods and systems for trajectory forecasting with recurrent neural networks using inertial behavioral rollout

    US20200379461A1