Adaptive Iterative Trajectory Prediction Method and System Based on Attention Mechanism
By building an intelligent network and a lane map network, combining a multi-layer perceptron and scene adaptive decoder, the problem of redundant computing in the existing technology is solved, efficient trajectory prediction is achieved, and the safety and comfort of autonomous driving vehicles are improved.
Patent Information
- Application Number
- CN202510317275.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-18
AI Technical Summary
In the prior art, the agent-centered trajectory prediction scheme needs to re-normalize and re-encode the input when sliding the observation window forward, resulting in redundant calculations during online prediction.
Adaptive iterative trajectory prediction method based on attention mechanism is adopted, and the historical trajectory and high-definition map of road participants are obtained through sensor modules, and the intelligent network and lane map network are constructed, and relative position embedding is combined with a multi-layer perceptron, and data decoding is used to generate motion trajectories at future moments.
It effectively reduces redundant calculations, improves the accuracy and calculation efficiency of trajectory prediction, and enhances the ride safety and comfort of autonomous driving vehicles.
Smart Images

Figure CN119840643B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of trajectory prediction, and specifically relates to an adaptive iterative trajectory prediction method and system based on an attention mechanism. Background Art
[0002] Due to the rapid development of current science and technology, driverless vehicles are gradually becoming the main players in the industry competition. As a key part of intelligent driving vehicles, the trajectory prediction module plays a connecting role, so it is very important to improve its prediction accuracy.
[0003] Currently, trajectory prediction in the research field is mainly based on vectorized scene representation, and on this basis, agent-centered trajectory prediction is proposed. However, the agent-centered modeling scheme requires re-normalization and re-encoding of the input when the observation window slides forward, resulting in redundant calculations during online prediction. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide an adaptive iterative trajectory prediction method and system based on an attention mechanism, aiming to solve the problem of redundant calculations during online prediction in the prior art that the agent-centered modeling scheme in the prior art requires re-normalization and re-encoding of the input when the observation window slides forward.
[0005] The first aspect of the present invention lies in providing an adaptive iterative trajectory prediction method based on an attention mechanism, and the method includes:
[0006] When the vehicle is driving on the road, obtain the historical trajectories of a number of road participants through a preset sensor module and perform historical trajectory encoding to obtain an agent network;
[0007] Obtain the high-definition map around during the driving process through the sensor module and perform encoding to obtain a lane map network;
[0008] Aggregate the agent network and the lane map network to obtain a symmetric scene network, calculate the relative positions between the vehicle and any road participant, perform encoding through a multi-layer perceptron to obtain relative position embeddings, and aggregate them into the symmetric scene network;
[0009] Perform data decoding through a scene-adaptive decoder to obtain the motion trajectories of the road participants at future moments.
[0010] According to one aspect of the above technical solution, the step of obtaining the historical trajectories of a number of road participants through a preset sensor module and performing historical trajectory encoding to obtain an agent network when the vehicle is driving on the road includes:
[0011] During the process of a vehicle traveling on a road, historical trajectories of a number of road participants are obtained through a preset sensor module on the vehicle; each road participant is an agent;
[0012] The historical trajectories of a single agent at different historical moments are feature-encoded by a historical trajectory encoder and linearly aggregated to obtain an agent network.
[0013] According to one aspect of the above technical solution, the steps of obtaining a high-definition map of the periphery during driving through the sensor module and encoding it to obtain a lane map network include:
[0014] During the process of a vehicle traveling on a road, a high-definition map of the periphery during the vehicle's travel is obtained through a preset sensor module on the vehicle;
[0015] According to the high-definition map, each road is segmented into multiple lanes, and each lane is divided into multiple lane segments;
[0016] The multiple lane segments in the same road are feature-encoded by a node encoder and linearly aggregated to obtain a lane map network.
[0017] According to one aspect of the above technical solution, the steps of aggregating the agent network and the lane map network to obtain a symmetric scenario network, calculating the relative position between the vehicle and any road participant, encoding it through a multi-layer perceptron to obtain a relative position embedding, and aggregating it into the symmetric scenario network include:
[0018] Obtain the heading difference, relative azimuth angle, and distance between the vehicle and the road participant, and determine the relative position between the vehicle and the road participant according to the heading difference, relative azimuth angle, and distance;
[0019] Encode the relative position through a multi-layer perceptron to obtain a relative position embedding; where the relative azimuth angle is represented by sine and cosine values;
[0020] Output agent features according to the agent network, and output lane environment features according to the lane map network;
[0021] Feature fusion of the agent features, lane environment features, and relative position embedding is performed in a vectorized manner to obtain a symmetric scenario network.
[0022] According to one aspect of the above technical solution, the steps of performing feature fusion of the agent features, lane environment features, and relative position embedding in a vectorized manner to obtain a symmetric scenario network include:
[0023] Using the context features as keys and values, taking the extended input agent as a query, and based on the multi-head attention mechanism, information is passed from the context features to the agent features for information update.
[0024] According to one aspect of the above technical solution, the step of obtaining the motion trajectory of the road participant at a future moment by performing data decoding through a scene-adaptive decoder includes:
[0025] Generating a number of adaptive trajectory anchors in a data-driven manner through a scene-adaptive decoder;
[0026] Selecting the trajectory anchors through an adaptive optimization iteration module to obtain the motion trajectory of the road participant at a future moment.
[0027] According to one aspect of the above technical solution, the method further includes:
[0028] Using the motion trajectory to adaptively select the trajectory anchors, retrieving key context elements, generating the offset of each trajectory segment in the motion trajectory through encoding the context, and optimizing each trajectory segment to obtain the final motion trajectory.
[0029] The second aspect of the present invention lies in providing an adaptive iterative trajectory prediction system based on the attention mechanism, which is applied to the method described in the above technical solution. The system includes:
[0030] The first network construction module is used to obtain the historical trajectories of a number of road participants through a preset sensor module and perform historical trajectory encoding to obtain an agent network when the vehicle is driving on the road;
[0031] The second network construction module is used to obtain the high-definition map around during driving through the sensor module and perform encoding to obtain a lane map network;
[0032] The aggregation module is used to aggregate the agent network and the lane map network to obtain a symmetric scene network, calculate the relative position between the vehicle and any road participant, perform encoding through a multi-layer perceptron to obtain a relative position embedding, and aggregate it into the symmetric scene network;
[0033] The motion trajectory prediction module is used to perform data decoding through a scene-adaptive decoder to obtain the motion trajectory of the road participant at a future moment.
[0034] The third aspect of the present invention lies in providing a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in the above technical solution is implemented.
[0035] The fourth aspect of the present invention is to provide an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in the above technical solution is implemented.
[0036] Compared with the prior art, the beneficial effects of adopting the adaptive iterative trajectory prediction method and system based on the attention mechanism shown in the present invention are as follows:
[0037] By building an agent network, a map network, and a scene-adaptive decoder, the present invention effectively embeds map topological structure information into the attention mechanism through a lane graph network and achieves excellent results. At the same time, a symmetric scene network is used, combined with global attention and relative position embedding updates, making the design more compact. The scene-adaptive decoder combined with adaptive optimization iteration enables a dynamic and reasonable allocation between accuracy and computational consumption. In the actual vehicle deployment work, it can greatly improve the utilization rate of each module, achieve efficient output of multi-modal trajectory prediction, and thus improve the riding safety and comfort of autonomous vehicles. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The above and / or additional aspects and advantages of the present invention will become apparent and easier to understand from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0039] Figure 1 is a schematic flow chart of an adaptive iterative trajectory prediction method based on the attention mechanism in an embodiment of the present invention;
[0040] Figure 2 is a structural block diagram of an adaptive iterative trajectory prediction system based on the attention mechanism in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] To make the objectives, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention is provided in conjunction with the accompanying drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0042] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right", and similar expressions used herein are for illustrative purposes only.
[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs. The terms used in the description of the present invention herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0044] Embodiment 1
[0045] Please refer to Figure 1 , the first embodiment of the present invention provides an adaptive iterative trajectory prediction method based on an attention mechanism, and the method includes steps S10 - S40:
[0046] Step S10, when the vehicle is driving on the road, obtain the historical trajectories of a number of road participants through a preset sensor module and perform historical trajectory encoding to obtain an agent network.
[0047] In this embodiment, the vehicle refers to the host vehicle, which is equipped with a sensor module including a camera, lidar, millimeter-wave radar, etc. During the driving process of the vehicle, obtain the historical trajectories of the surrounding road participants through the sensor module, that is, obtain the historical movement trajectories of the road participants at historical moments.
[0048] Among them, road participants include vehicles, cyclists, pedestrians on the road and on the periphery of the road, as well as movable obstacles, which can be stationary or moving at historical moments. For example, the vehicle as a road participant and the host vehicle are driving in the same direction of adjacent vehicles.
[0049] Further, the step of obtaining the historical trajectories of a number of road participants through a preset sensor module and performing historical trajectory encoding to obtain an agent network when the vehicle is driving on the road includes:
[0050] During the driving process of the vehicle on the road, obtain the historical trajectories of a number of road participants through the preset sensor module on the vehicle; each road participant is an agent;
[0051] Perform feature encoding on the historical trajectories of a single agent at different historical moments through a historical trajectory encoder and perform linear aggregation to obtain an agent network.
[0052] Specifically, during the process of the host vehicle driving on the road, the sensor modules on the host vehicle, such as cameras, lidars, and millimeter-wave radars, will identify the surrounding road participants and obtain the historical trajectories of the road participants at historical moments. Each road participant is an independent agent. Then, the historical trajectory encoder encodes the historical trajectories of a single agent at different moments and performs linear aggregation to obtain the agent network. Finally, the interaction between agents in the current driving scenario is managed through this agent network.
[0053] Step S20: Obtain the high-definition map of the surrounding area during driving through the sensor module and perform encoding to obtain the lane graph network.
[0054] In this embodiment, during the driving process of the vehicle, the high-definition map of the surrounding area during driving will also be obtained through the above-mentioned sensor module. Then, the roads, lanes, or lane segments are encoded and output as a vector map to obtain the lane graph network.
[0055] Further, the step of obtaining the high-definition map of the surrounding area during driving through the sensor module and performing encoding to obtain the lane graph network includes:
[0056] During the process of the vehicle driving on the road, the high-definition map of the surrounding area during vehicle driving is obtained through the preset sensor module on the vehicle;
[0057] According to the high-definition map, each road is divided into multiple lanes, and each lane is divided into multiple lane segments;
[0058] The node encoder encodes the feature of multiple lane segments in the same road and performs linear aggregation to obtain the lane graph network.
[0059] Specifically, during the driving process of the vehicle, the sensor module will simultaneously identify and obtain the historical trajectories of road participants and the high-definition map of the surrounding environment, and then construct the agent network and the lane graph network respectively. When constructing the lane graph network, specifically, each road is divided into multiple lanes according to the high-definition map, and each lane is divided into multiple lane segments. Then, the node encoder divides and encodes the features in units of lane segments, and linear aggregation is performed to obtain the lane graph network.
[0060] Step S30: Aggregate the agent network and the lane graph network to obtain the symmetric scenario network, calculate the relative position between the vehicle and any road participant, and obtain the relative position embedding through encoding by a multi-layer perceptron and aggregate it into the symmetric scenario network.
[0061] In this embodiment, aggregating the agent network and the lane map network to obtain a symmetric scenario network, calculating the relative position between the vehicle and any road participant, and encoding it through a multi-layer perceptron to obtain a relative position embedding, and aggregating it into the symmetric scenario network includes the following steps:
[0062] Obtain the heading difference, relative azimuth angle, and distance between the vehicle and the road participant, and determine the relative position between the vehicle and the road participant according to the heading difference, relative azimuth angle, and distance;
[0063] Encode the relative position through a multi-layer perceptron to obtain a relative position embedding; wherein, the relative azimuth angle is represented by sine and cosine values;
[0064] Output agent features according to the agent network, and output lane environment features according to the lane map network;
[0065] Perform feature fusion on the agent features, lane environment features, and relative position embedding in a vectorized manner to obtain a symmetric scenario network.
[0066] Among them, in the step of performing feature fusion on the agent features, lane environment features, and relative position embedding in a vectorized manner to obtain a symmetric scenario network, it includes:
[0067] Use context features as keys and values, use the extended input agent as a query, and transfer information from the context features to the agent features based on the multi-head attention mechanism for information update.
[0068] Specifically, in this embodiment, the relative position of the agent is calculated pairwise using the heading difference, relative azimuth angle, and distance, and further encoded through a multi-layer perceptron to obtain a relative position embedding. To enhance numerical stability, the angle is represented by its sine and cosine values.
[0069] Then, perform effective feature fusion on the agent features, surrounding features, and RPE in a vectorized manner, use context features as keys and values, and use the extended input agent features as queries. Finally, apply the multi-head attention mechanism to transfer messages from the context features to the agent features for real-time information update.
[0070] Step S40, perform data decoding through a scenario-adaptive decoder to obtain the motion trajectory of the road participant at a future moment.
[0071] In this embodiment, the step of performing data decoding through a scenario-adaptive decoder to obtain the motion trajectory of the road participant at a future moment includes:
[0072] Generate a number of adaptive trajectory anchors in a data-driven manner through a scene-adaptive decoder;
[0073] Select the trajectory anchors through an adaptive optimization iteration module to obtain the motion trajectory of the road participant at a future moment.
[0074] Wherein, the method further includes:
[0075] Use the motion trajectory to adaptively select the trajectory anchors, retrieve key context elements, generate the offset of each trajectory segment in the motion trajectory through encoding the context, so as to optimize each trajectory segment and obtain the final motion trajectory.
[0076] Specifically, the scene-adaptive decoder uses a cyclic, anchor-free proposal module to generate K adaptive trajectory anchors in a data-driven manner, and then uses an adaptive optimization iteration module to select the final predicted trajectory. In the pattern and scene attention, a cross-attention layer is used to update the pattern queries with multiple contexts, including the historical encoding of the target agent, the map encoding, and the encoding of adjacent agents. After the pattern and scene attention, the multi-pattern queries query each other through pattern-to-pattern self-attention to improve the diversity of multiple patterns.
[0077] In order to reduce the context retrieval burden of the queries and improve the quality of the trajectory anchors, the decoder of DETR in object detection is generalized to a cyclic manner. During the cyclic process, the context-aware pattern queries only decode the future trajectory through a multi-layer perceptron at the end of each cyclic step. In the subsequent process, these queries become the input again and retrieve the scene context related to the prediction of the next few waypoints.
[0078] In addition, in this embodiment, the initial predicted trajectory is also used to adaptively select the trajectory anchors and retrieve key context elements, and then each trajectory segment is optimized by generating the offset of each trajectory segment using the encoded context. Moreover, the model will predict a quality score measuring the prediction quality and adaptively determine the number of refinement iterations.
[0079] More specifically, in this embodiment, the positions of each agent in the current driving scene are projected from the BEV perspective. For the i historical state of the th agent , the goal of this embodiment is to predict the future state at p time steps and the related probability , where T represents the state of the agent at time step , and its composition is Represents the coordinates of the agent center, Indicates whether the agent lacks historical trajectories (filled with 0 if not present), Represents the quality of the trajectory (from 0 to 3), where a higher value indicates a more comprehensive historical trajectory, Indicates the type of the target agent, Represents the heading angle related to the agent, Represents the speed of the agent on the X - axis and Y - axis.
[0080] Among them, the target agent will interact with context elements including the historical states of surrounding agents and the high - definition map m . For the high - definition map m , a vectorized representation is adopted, where each lane is defined as a sequence of points along its centerline, and each point is represented by coordinates and semantic information. Therefore, the classic motion prediction task is represented as , where f represents the prediction model.
[0081] In this embodiment, first, an initial trajectory and trajectory features are generated according to a backbone network , so the classic trajectory prediction method is represented as . The initialized trajectory is used to select trajectory anchors and retrieve key context elements , and then they are fed into the refinement model along with the trajectory features for trajectory refinement. And since the trajectory refinement is iterative, the formula for the trajectory refinement in the i th iteration is , and finally the generated offset is added to the input trajectory for trajectory refinement.
[0082] In this embodiment, a historical trajectory encoder HTE is used to encode the temporal features of the historical trajectories in a single agent. The historical state of the agent is represented as , where represents the number of input features. First, the input features are converted to using an MLP, that is , where represents the multi - layer perceptron network, represents the dimension of the transformed features, and then a stacked Transformer structure is adopted for the TEncode the features within a time step. The attention mechanism of the j th layer of the Transformer structure is denoted as , where . Then, use the MLP-based aggregator to construct a singular feature vector for each trajectory, thereby obtaining the final output feature of the historical trajectory encoder HTE. Identify the differences between agent trajectories through the agent network AIN and effectively collect information about surrounding agents. AIN also adopts the Transformer structure, which is denoted as , where , and the final output feature is .
[0083] In this embodiment, use the node encoder NE to encode the node features on the same lane and further aggregate these nodes into lane-level features. NE takes the feature as input and calculates the aggregated feature , where represents the number of lanes in the scene, represents the number of nodes in each lane, represents the input dimension of the map, and represents the output dimension generated by the NE network.
[0084] Among them, the node encoder NE and the historical trajectory encoder THE have the same structure and data processing method, but have different hyperparameters and weights. Compared with one-dimensional data, map data shows a complex connection pattern, including differences in spatial distances and various types of node connections. Traditional coding methods have difficulty capturing these complex relationships, so a lane graph network LGT is introduced in this embodiment.
[0085] The attention mechanism of the lane graph network LGT model is designed as:
[0086] ;
[0087] where , , , , respectively represent the matrices of the connection types of the forward connection, successor connection, left connection, right connection, and map topology, and W is the corresponding weight matrix of each matrix.
[0088] The structure of the LGT model consists of stacked Transformer layers that embed topological information from the map. All layers in this structure share the RPE matrix, take the output of the NE as the input of the LGT, and obtain map information , where D denotes the dimension of the output . Among them, the attention matrix of the j -th layer is denoted as .
[0089] In this embodiment, a bias matrix B is adopted to represent the spatial relativity of the map topology. Since the longitudinal speed of the vehicle is usually higher than its lateral speed, the agent shows more dependence in the front-back direction.
[0090] To capture the differences in traffic scene movement, four connection relationships are defined in this embodiment: former, successor, left neighbor, and right neighbor, and four PRE matrices are constructed. In addition, the lateral connectivity may affect lane changes, so a connection type matrix is constructed, and matrix multiplication is performed to represent the lateral connectivity between adjacent channels, denoted by . For static map elements, such as lane segments, the centroid of the polyline is used as the trajectory anchor point, and the displacement vector between the endpoints is used as the heading angle.
[0091] Specifically, the anchor position of the element i in the global coordinate system can be represented by its position and the heading vector . Three quantities are used to describe the relative position between the element i and the element j , including the heading difference , the relative azimuth angle , and the distance .
[0092] To enhance numerical stability, the angle is represented by its sine and cosine values.
[0093] In this embodiment, the heading difference is represented as , the relative azimuth angle (the displacement vector is ) is represented as , and the relative spatial information is represented as a 5D vector . In addition, the relative position encoding is further encoded by the MLP to generate the relative position embedding RPE with the shape of [N, N, D].
[0094] In this embodiment, once the agent features, map features, and relative position embedding (RPE) are obtained, the symmetric scene network is used to update the features in a viewpoint-invariant manner. The structure of the symmetric scene network, which contains multiple stacked SFT layers, is similar to the standard Transformer. Denote the features of the i th and j th agents as , respectively. The RPE vectors associated with the edges from to are denoted as . The tuple contains all the information planned to be passed from node i to node j . Therefore, a simple MLP can be used to encode these features and obtain the j th context vector of node i , denoted as , where represents the concatenation operator, and indicates that the MLP consists of a linear layer, layer normalization, and ReLU activation function. Then, cross-attention is performed on the target node and its context, denoted as , where is the standard multi-head attention function, and is the set of context vectors of feature j . Meanwhile, also contains , indicating that there is a self-loop for each node. Similar to the standard Transformer, a pointwise feed-forward fully connected layer is integrated after the attention mechanism. In addition, in each layer, is updated by re-encoding the context vector using another MLP and then added to the input RPE through a residual connection.
[0095] In this embodiment, the last step of prediction is to use the scene encoding output by the encoder to decode the K future trajectories of each target agent, which is very important because the encoder only returns a set of feature embeddings. Inspired by object detection research, a decoder similar to DETR is adopted to handle such a one-to-many problem, that is, multiple learnable queries interact with the scene encoding and decode the trajectories. The query-based decoder uses a cyclic, traceless anchor proposal module to generate adaptive trajectory anchors, followed by a scene-adaptive decoder, which adaptively iterates between retrieving key context elements and predicting more accurate trajectories to improve the performance of the trajectory prediction model with limited additional computation.
[0096] In the pattern and scene attention module, an interactive attention layer is used to update the pattern queries with multiple contexts, including the historical trajectory of the target agent, the map encoding, and the encoding of adjacent agents. Thanks to the cross-attention layer, the pattern queries can retrieve the scene context and quickly narrow down the search space of the trajectory anchors. Then, the K pattern queries pay attention to each other through the pattern-to-pattern self-attention module to improve the diversity of multiple patterns. To predict the trajectories of multiple agents in parallel, a set of scene encodings is shared among all target agents in the scene. Since these encodings are retrieved from their local spatio-temporal coordinate systems, they need to be projected into the current viewpoints of each agent to achieve the same effect as agent-centered modeling.
[0097] Trajectory refinement requires information-rich trajectory anchor selection and context retrieval. After generating the initial trajectory and trajectory features in the previous step decoder, the trajectory refinement module first selects a certain number of trajectory anchors along the trajectory, and then retrieves the context near the trajectory anchors with an adaptive radius. To balance context richness and computational efficiency, the trajectory is divided into N lane segments, and the end point of each lane segment is selected as the trajectory anchor. When retrieving the context elements c around the trajectory anchor, previous methods usually used a fixed radius or area around the trajectory anchor to extract local context, but this method is very poor especially when applied to different datasets and scenarios that may require different retrieval ranges. Therefore, in this embodiment, a method for retrieving trajectory anchors with an adaptive radius is proposed, which can be achieved by the following principle, that is, the number of trajectory refinement iterations and the speed of the target agent v , which is expressed as , where will decrease the retrieval range of the radius as the number of trajectory refinement iterations increases, making the trajectory more accurate.
[0098] Previous methods usually fuse all embeddings at once and refine the entire future trajectory, while there are long-term cumulative errors in the future trajectory. To alleviate this problem, a cyclic refinement strategy is adopted, which divides the trajectory into N trajectory segments and refines the trajectory N times each time. Here, N is equal to the number of trajectory anchors, that is, each trajectory corresponds to one trajectory anchor. For each trajectory segment, only the context retrieved by the corresponding trajectory anchor is used for trajectory refinement to enhance local context fusion. In addition, all N trajectory segments are refined, that is, one trajectory refinement iteration is completed, and this embodiment will perform multiple trajectory refinement iterations.
[0099] Specifically, in each cyclic refinement, in this embodiment, by refining one trajectory segment, a set of scene contexts a is embedded around the trajectory anchor corresponding to the trajectory segment, as well as the future trajectory embedding of the target agent In this embodiment, a cross-attention mechanism is adopted to fuse them, where the trajectory embedding is used as the query to process the keys and values from the context embedding. The fused trajectory embedding will be used for prediction. , that is, the offset of the waypoints in the trajectory segment, to subtly adjust the original trajectory segment. The updated trajectory embedding will also be used as a new query to refine the next trajectory segment. After all time steps within one iteration are completed, the entire trajectory will be adjusted, and the trajectory embedding will also be updated using rich context elements. Since multiple possible trajectories are predicted, the trajectory embedding will also be used to predict the probability of each trajectory. N After one iteration, the updated trajectory and trajectory embedding will be used to initiate another refinement iteration. This embodiment first repeats the mentioned adaptive trajectory anchor point selection and context retrieval, and then performs additional
[0100] cyclic trajectory refinement steps. In the first refinement iteration, a compression network operation is adopted in this embodiment to reduce the hidden dimension of the trajectory embedding of the initial trajectory prediction module for efficient refinement, and this incurs little performance loss. As subsequent iterations are entered, the predicted trajectories become more accurate. However, after multiple trajectory refinement iterations, the accuracy of trajectory prediction may also decrease. N To balance the performance gain and the workload of additional computations, this embodiment proposes an adaptive trajectory refinement strategy, where the number of refinement iterations is dynamically adjusted according to the current prediction quality and the remaining prediction refinement improvement. Specifically, this embodiment proposes a quality score to quantify the current prediction quality, which is expressed as
[0101] , where represents the maximum prediction error among all iterations, and represents the minimum prediction error among all iterations. To enable the model to predict the quality score, this embodiment first uses a GRU to repeatedly process the trajectory embeddings of all previous iterations, and then uses an MLP to generate the quality score, generating a total of three types of outputs: the predicted trajectory, the probability of the predicted trajectory, and the quality score. And the quality score at each refinement iteration will also be predicted. Therefore, this embodiment proposes a simple and effective strategy to dynamically determine whether another trajectory refinement iteration is needed, and its inputs include: the initial trajectory prediction model
[0102] , the trajectory refinement model , the quality score decoder , the agent's historical trajectory , the scene context , the quality score threshold c , and the maximum number of refinement iterations in the trajectory refinement stage ; Its output includes: the future trajectory of the target agent and the corresponding trajectory probability p .
[0103] Compared with the prior art, the adaptive iterative trajectory prediction method based on the attention mechanism shown in this embodiment has the beneficial effects that:
[0104] In this embodiment, by building an agent network, a map network, and a scene-adaptive decoder, the map topological structure information is effectively embedded into the attention mechanism through the lane graph network, and excellent results are produced. At the same time, the symmetric scene network is used, combined with global attention and relative position embedding update, so that the design is more compact, and the scene-adaptive decoder combined with adaptive optimization iteration enables a dynamic and reasonable allocation between the accuracy and the computational cost consumption. In the actual vehicle deployment work, the utilization rate of each module can be greatly improved, the efficient output of multi-modal trajectory prediction can be realized, and thus the riding safety and comfort of the autonomous vehicle can be improved.
[0105] Embodiment 2
[0106] Please refer to Figure 2 , the second embodiment of the present invention provides an adaptive iterative trajectory prediction system based on the attention mechanism, which is applied to the method described in the above embodiment. The system includes:
[0107] The first network construction module 10 is used to obtain the historical trajectories of a number of road participants through a preset sensor module and perform historical trajectory encoding to obtain an agent network when the vehicle is driving on the road;
[0108] The second network construction module 20 is used to obtain the high-definition map around during driving through the sensor module and perform encoding to obtain a lane graph network;
[0109] The aggregation module 30 is used to aggregate the agent network and the lane graph network to obtain a symmetric scene network, calculate the relative position between the vehicle and any road participant, perform encoding through a multi-layer perceptron to obtain a relative position embedding, and aggregate it into the symmetric scene network;
[0110] The motion trajectory prediction module 40 is used to perform data decoding through a scene-adaptive decoder to obtain the motion trajectory of the road participant at a future moment.
[0111] Compared with the prior art, the adaptive iterative trajectory prediction system based on the attention mechanism shown in this embodiment has the beneficial effects that:
[0112] In this embodiment, by building an agent network, a map network, and a scene-adaptive decoder, the map topology information is effectively embedded into the attention mechanism through the lane graph network, and excellent results are achieved. At the same time, the symmetric scene network is utilized, combined with global attention and relative position embedding updates, making the design more compact. The scene-adaptive decoder combines adaptive optimization iterations, enabling a dynamic and reasonable allocation between accuracy and computational consumption. In the actual vehicle deployment work, the utilization rate of each module can be greatly improved, achieving efficient output of multi-modal trajectory prediction, thereby enhancing the riding safety and comfort of autonomous vehicles.
[0113] Embodiment 3
[0114] The third embodiment of the present invention provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in the above embodiments is implemented.
[0115] Embodiment 4
[0116] The fourth embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in the above embodiments is implemented.
[0117] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0118] The above-described embodiments merely represent several implementation manners of the present invention. The descriptions are relatively specific and detailed, but should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. An adaptive iterative trajectory prediction method based on the attention mechanism, characterized in that, The method includes: When the vehicle is driving on the road, obtaining the historical trajectories of a number of road participants through a preset sensor module and performing historical trajectory encoding to obtain an agent network; Obtaining the high-definition map around during driving through the sensor module and performing encoding, dividing each road into multiple lane segments, using the centroid of the polyline of the lane segment as a trajectory anchor point, and calculating the heading angle through the endpoint displacement vector to obtain a lane map network; Aggregating the agent network and the lane map network to obtain a symmetric scenario network, calculating the heading difference, relative azimuth angle, and distance between the vehicle and any road participant, and performing encoding through a multi-layer perceptron to obtain a relative position embedding and fusing it into the symmetric scenario network; Generating initial trajectory anchor points and trajectory features through a scene-adaptive decoder, and selecting a certain number of trajectory anchor points along the trajectory features based on a cyclic refinement strategy, retrieving the context near the trajectory anchor points with an adaptive radius, specifically including dividing the trajectory into N lane segments and selecting the end point of each lane segment as the trajectory anchor point, and iteratively optimizing the trajectory to output the optimized motion trajectory of the road participant at a future moment; specifically including: Dynamically adjust the retrieval radius of the trajectory anchor point according to the real-time speed of the target agent, satisfying the relational expression , where Reduce the radius retrieval range as the number of iterations increases to make the trajectory more accurate, is the speed of the target agent; Process the iterative trajectory embedding through the recurrent neural network GRU to generate a predicted quality score When reaches a preset threshold or the number of iterations exceeds the maximum value, terminate the refinement; Among them, the process of iteratively optimizing the trajectory includes: Dividing the trajectory into N trajectory segments, that is, N lane segments, and performing N times of trajectory refinement iterations. N is equal to the number of trajectory anchor points, that is, each trajectory segment corresponds to a trajectory anchor point. When optimizing each trajectory segment, only the context retrieved by the corresponding trajectory anchor point is used for trajectory refinement to enhance local context fusion; after one iteration ends, the updated trajectory and trajectory embedding will be used to initiate another refinement iteration; Among them, the number of refinement iterations is dynamically adjusted according to the current prediction quality score and the remaining prediction refinement improvement. The current prediction quality is quantified by the prediction quality score, and the expression is: , where represents the maximum prediction error among all iterations, represents the minimum prediction error among all iterations; Among them, determining whether another trajectory refinement iteration is required, the input includes: the initial trajectory prediction model , the trajectory refinement model , the mass fraction decoder , the agent's historical trajectory , the scene context c , the mass fraction threshold , the maximum number of refinement iterations in the trajectory refinement stage ; the output includes: the future trajectory of the target agent and the corresponding trajectory probability p .
2. The adaptive iterative trajectory prediction method based on the attention mechanism according to claim 1, wherein When the vehicle is driving on the road, the step of obtaining the historical trajectories of a number of road participants through a preset sensor module and performing historical trajectory encoding to obtain an agent network includes: When the vehicle is driving on the road, obtaining the historical trajectories of a number of road participants through the preset sensor module on the vehicle; each road participant is an agent; Performing feature encoding on the historical trajectories of a single agent at different historical moments through a historical trajectory encoder and performing linear aggregation to obtain an agent network.
3. The adaptive iterative trajectory prediction method based on the attention mechanism according to claim 1, wherein The step of obtaining the high-definition map around during driving through the sensor module and performing encoding, dividing each road into multiple lane segments, using the centroid of the polyline of the lane segment as a trajectory anchor point, and calculating the heading angle through the endpoint displacement vector to obtain a lane map network includes: When the vehicle is driving on the road, obtaining the high-definition map around during the vehicle's driving process through the preset sensor module on the vehicle; According to the high-definition map, dividing each road into multiple lanes and dividing each lane into multiple lane segments; Performing feature encoding on multiple lane segments in the same road through a node encoder and performing linear aggregation to obtain a lane map network.
4. The adaptive iterative trajectory prediction method based on the attention mechanism according to claim 1, wherein Aggregate the agent network and the lane graph network to obtain a symmetric scenario network, and calculate the heading difference, relative azimuth angle, and distance between the vehicle and any road participant, and encode them through a multi-layer perceptron to obtain relative position embeddings and fuse them into the symmetric scenario network. The steps include: Obtain the heading difference, relative azimuth angle, and distance between the vehicle and the road participant, and determine the relative position between the vehicle and the road participant based on the heading difference, relative azimuth angle, and distance; Encode the relative position through a multi-layer perceptron to obtain relative position embeddings; wherein, the relative azimuth angle is represented by sine and cosine values; Output agent features according to the agent network, and output lane environment features according to the lane graph network; Perform feature fusion on the agent features, lane environment features, and relative position embeddings in a vectorized manner to obtain a symmetric scenario network.
5. The adaptive iterative trajectory prediction method based on the attention mechanism according to claim 4, wherein The step of performing feature fusion on the agent features, lane environment features, and relative position embeddings in a vectorized manner to obtain a symmetric scenario network includes: Use context features as keys and values, use the extended input agent as a query, and transfer information from the context features to the agent features based on the multi-head attention mechanism for information update.
6. An adaptive iterative trajectory prediction system based on an attention mechanism, characterized in that, Applied to the method according to any one of claims 1-5, the system includes: A first network construction module, configured to obtain historical trajectories of a plurality of road participants through a preset sensor module and perform historical trajectory encoding to obtain an agent network when the vehicle is driving on the road; A second network construction module, configured to obtain a high-definition map of the periphery during driving through the sensor module and perform encoding to obtain a lane graph network; An aggregation module, configured to aggregate the agent network and the lane graph network to obtain a symmetric scenario network, calculate the relative position between the vehicle and any road participant, encode it through a multi-layer perceptron to obtain relative position embeddings, and aggregate them into the symmetric scenario network; A motion trajectory prediction module, configured to generate an initial trajectory anchor point through a scene-adaptive decoder, and iteratively optimize the trajectory based on a cyclic refinement strategy, and output the optimized motion trajectory of the road participant at a future moment.
7. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1-5.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1-5.
Citation Information
Patent Citations
Multi-vehicle joint trajectory prediction method based on intent learning of surrounding vehicles
CN118968441A