Automatic driving track prediction method and device based on scene guide gating
By introducing a scenario-guided gating mechanism and a dynamic diffusion graph convolution module in the prediction of autonomous driving trajectory, the existing models' shortcomings in scene heterogeneity and social interaction between agents are solved, and a more accurate and robust trajectory prediction is achieved.
Patent Information
- Application Number
- CN202510102080.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The existing autonomous driving trajectory prediction model fails to fully utilize the impact of scenario heterogeneity on driving behavior, and it is difficult to effectively describe highly random social interactions between agents, resulting in a decline in prediction performance in diverse scenarios.
A method of autonomous driving trajectory prediction based on scene-guided gating is proposed. Through the gated loop unit and the scene-guided gating unit, the intelligent body encoding is filtered and updated, and combined with the self-attention mechanism and the dynamic diffusion graph convolution module, the lane attention map and path encoding are generated to achieve a more socially conscious trajectory prediction.
By selectively filtering information using scene-specific non-shared parameters, a scenario-guided information bottleneck is established, and the dynamics and randomness of social interactions between agents are effectively captured, and the accuracy and robustness of trajectory prediction are improved.
Smart Images

Figure CN119928912A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of autonomous driving trajectory prediction, and in particular, to an autonomous driving trajectory prediction method and device based on scene-guided gating. Background Art
[0002] Accurate trajectory prediction is crucial to improving the safety of autonomous driving. It predicts the future trajectory of the target vehicle by analyzing the historical trajectories of surrounding agents and complex lane maps. Accurate and reliable trajectory prediction can support decision-making and planning tasks downstream of autonomous driving to avoid potential dangers. However, the uncertain intentions of agents and the highly random social interactions between agents pose severe challenges to this task.
[0003] To address these challenges, some studies have proposed deep learning-based models for trajectory prediction. The models are usually based on an encoder-decoder structure, where the encoder is responsible for converting historical trajectories and high-definition map information into vector representations, while the decoder is responsible for generating multiple credible future trajectories. However, most existing studies ignore the impact of scene heterogeneity on driving behavior. Scene heterogeneity can affect driving behavior. In relatively simple scenarios, drivers may adopt more aggressive driving strategies, such as changing lanes and overtaking. Correspondingly, in relatively complex scenarios, such as when there are many vehicles or pedestrians around, drivers are more inclined to adopt more cautious driving strategies. Existing models usually fail to fully utilize the impact of scene heterogeneity on driving behavior, making it difficult for the model to achieve more socially aware trajectory prediction, thereby affecting the prediction performance of the model in diverse scenarios.
[0004] In addition, the highly random social interactions between agents have not been fully explored, reducing the reliability of predictions. The future trajectories of agents are largely affected by the social interactions between agents, which are dynamic and uncertain in nature and propagate randomly between agents. For example, the social interactions between agents may change due to changes in the scene or changes in the direction of the agents. Even in almost the same scene, such social interactions may change due to different driving habits and intentions of drivers. However, existing models mainly learn fixed or single-hop social interactions, which makes it difficult to fully describe the highly random social interactions between agents. Summary of the invention
[0005] The purpose of this application is to provide a method and device for autonomous driving trajectory prediction based on scene-guided gating to overcome the impact of scene heterogeneity on driving behavior and the impact of highly random social interactions between intelligent agents, thereby achieving more accurate trajectory prediction.
[0006] In order to achieve the above purpose, the technical solution of this application is as follows:
[0007] A method for predicting trajectory of autonomous driving based on scene-guided gating, comprising:
[0008] Use gated recurrent units to encode the target vehicle’s past trajectory, lane map, and surrounding agent’s past trajectory to obtain target encoding, lane encoding, and agent encoding.
[0009] The scene-guided gating unit is used to filter the agent code to obtain a filtered agent code, and the filtered agent code is used to update the lane code;
[0010] The lane map adjacency matrix and the updated lane encoding are used to generate a lane attention map through the self-attention mechanism;
[0011] Perform a diffusion convolution operation on the updated lane code and lane attention map, and then add them to the updated lane code after random dropout operation to obtain the lane node feature;
[0012] Input the target code, lane node features and the adjacency matrix of the lane graph into the strategy head to obtain the lane node set;
[0013] Get the lane node features corresponding to the lane node set, form a path code, and use the path code to update the target code;
[0014] The position of the target vehicle, the latent variables following the multivariate Gaussian distribution, and the updated target encoding are passed through a linear layer and then filtered by a scene-guided gating unit. The filtered features are passed through a linear layer to obtain the predicted trajectory.
[0015] Furthermore, the scene guides the gating unit to perform the following operations:
[0016] Construct the embedding vectors corresponding to the surrounding agents through the embedding layer;
[0017] Concatenate the embedding vectors corresponding to the surrounding agents with the position of the target vehicle to obtain heterogeneous scene information;
[0018] Input the heterogeneous scene information into the meta-learner to learn the weight parameters of the meta-linear layer;
[0019] The original code to be filtered is input into the meta-linear layer, and then after the activation operation, the Hadamard product is performed with the original code, and after the Hadamard product, it is added with the original code to obtain the filtered code.
[0020] Further, the updating of the lane code by using the filtered agent code includes:
[0021] Linearly map the filtered agent encoding to obtain the key vector and value vector;
[0022] Linearly map the lane encoding to obtain the query vector;
[0023] Perform a scaled dot product attention operation to obtain a lane encoding with agent context;
[0024] Then the lane coding with the agent context is concatenated with the lane coding and a linear transformation is performed to obtain the updated lane coding.
[0025] Furthermore, the step of generating a lane attention map by using the adjacency matrix of the lane map and the updated lane coding through a self-attention mechanism includes:
[0026] Linearly map the updated lane coding to obtain the query vector, key vector and value vector;
[0027] Perform a self-attention operation and then perform a Hadamard product with the adjacency matrix of the lane map to obtain the lane attention map.
[0028] Further, the updating of the target code using the path code includes:
[0029] Linearly map the path encoding to obtain the key vector and value vector;
[0030] Linearly map the target encoding to obtain the query vector;
[0031] Perform a scaled dot product attention operation to obtain the target encoding with path context;
[0032] Then the target encoding with the path context is concatenated with the target encoding to obtain the updated target encoding.
[0033] The present application also proposes an autonomous driving trajectory prediction device based on scene-guided gating, comprising a processor and a memory storing a plurality of computer instructions, wherein the computer instructions, when executed by the processor, perform the steps of the above-mentioned autonomous driving trajectory prediction method based on scene-guided gating.
[0034] This application proposes a method and device for autonomous driving trajectory prediction based on scene-guided gating, and proposes a new scene-guided gating mechanism to utilize the impact of scene heterogeneity on driving behavior. This mechanism selectively filters information by generating scene-specific non-shared parameters, thereby establishing a scene-guided information bottleneck, which is conducive to achieving more socially aware trajectory prediction. Secondly, a dynamic diffusion graph convolution module is proposed to deeply explore the highly random social interactions between agents. This module regards social interactions as a diffusion process on the lane map, and uses bidirectional random walks on the lane attention map to simulate the randomness of social interactions. Compared with the existing technical solutions, the technical solution of this application has better effectiveness and robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is the structural diagram of the dynamic diffusion graph convolutional network based on scene-guided gating for this application.
[0036] Figure 2 This is a flow chart of the autonomous driving trajectory prediction method based on scene-guided gating in this application.
[0037] Figure 3 The structure diagram of the gating unit is guided by this application scenario.
[0038] Figure 4 This is the structural diagram of the dynamic diffusion graph convolution module for this application. DETAILED DESCRIPTION
[0039] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0040] For the trajectory prediction task, given the past trajectories of the target vehicle and surrounding agents (vehicles or pedestrians) and the lane map representation of the scene, the goal of trajectory prediction is to learn a mapping function f that can predict multiple possible trajectories of the target vehicle and the associated confidence scores. Each predicted trajectory includes the target vehicle from time step 1 to t f An ordered pair of two-dimensional coordinates of .
[0041] In each scenario, the past trajectory of agent i (vehicle or pedestrian) can be expressed as past t p The trajectory vector of time steps Each in are the two-dimensional coordinate pair, velocity, acceleration, and yaw rate of agent i at time step t; I i Represents the type of agent i: vehicle (I i =0) or pedestrians (I i =1). In addition, i=0 represents the target vehicle and t=0 represents the current time step.
[0042] The lane graph is represented as a directed graph G = (V, E), where V is a set of nodes representing lane centerlines, and E is a set of edges (including successor edges and adjacent edges). Each lane centerline v is divided into fixed-length segments represented by N points. The nth point is represented by a vector c. n =[x n ,y n ,θ n ,I n ] means, where (x n ,y n ) and θ nare the two-dimensional coordinate pair and the yaw rate; I n is a two-dimensional binary vector indicating whether the point is on a stop line or a crosswalk.
[0043] This application proposes a new deep learning model, namely, the scene-guided gating-based dynamic diffusion graph convolutional network (SGG-DDGCN). The proposed SGG-DDGCN model aims to learn the uncertain future relationships between agents and selectively filter information. Figure 1 As shown in the figure, it includes an encoder module, a policy head, and a decoder module. The encoder module encodes the past trajectory of the target vehicle, the lane map, and the past trajectory of the surrounding agents, and outputs a learned representation of the lane map containing the context of the surrounding agents. The module contains a scene-guided gating unit SGG for selectively filtering information and a dynamic diffusion graph convolution module DDGCM for describing the uncertain future relationship between agents. Subsequently, the policy head outputs the probability distribution of each edge of the lane map to support path sampling on the lane map. Finally, based on the position of the target vehicle, the randomly generated latent variables, and the path traversed by the policy head, the decoder module realizes multimodal trajectory prediction at multiple future time steps.
[0044] In one embodiment, Figure 2 As shown, a method for automatic driving trajectory prediction based on scene-guided gating is provided, including:
[0045] Step S1: Use a gated recurrent unit to encode the target vehicle's past trajectory, lane map, and surrounding agent's past trajectory to obtain a target code, a lane code, and an agent code.
[0046] like Figure 1 As shown, the encoder module includes a gated recurrent unit (GRU), a scene-guided gating unit (SGG), and a dynamic diffusion graph convolution module (DDGCM).
[0047] This embodiment first uses a gated recurrent unit (GRU) to encode past trajectories, lane maps, and past trajectories of surrounding agents to obtain target encoding, lane encoding, and agent encoding.
[0048] Vehicle drivers tend to pay more attention to scene information in nearby time steps rather than scene information in distant time steps. Based on this assumption, the encoder module of this embodiment directly uses the gated recurrent unit (GRU) to encode the target vehicle's past trajectory, lane map, and surrounding agent's past trajectory, and outputs the final hidden state. Specifically, it outputs the target encoding Lane Coding and agent encoding Where D is the encoding feature dimension.
[0049] It should be noted that, considering that the number of agents around the target vehicle may be different in different scenarios, the GRU configured for each agent in this embodiment is independent, but its parameters are shared among the agents. Similarly, the GRU used in each lane also follows the same principle.
[0050] Step S2: Use the scene-guided gating unit to perform scene filtering on the agent code to obtain the filtered agent code, and use the filtered agent code to update the lane code.
[0051] This step uses the scene-guided gating unit (SGG) to filter the agent encoding The key information in the filter is obtained to obtain the filtered agent encoding. Figure 3 As shown, the SGG proposed in this embodiment extracts meta-knowledge from scene information, generates a set of scene-specific non-shared parameters, and uses these parameters to guide the information filtering process to fully utilize the impact of scene heterogeneity on driving behavior. SGG aims to assign similar model parameters to similar scenes, thereby constructing a scene-guided information bottleneck, which is conducive to achieving more socially aware trajectory prediction.
[0052] Scene Guided Gating Unit (SGG), the original code to be filtered is the agent code, and the following operations are performed:
[0053] Step 2.1: Construct the embedding vector corresponding to the surrounding agents through the embedding layer.
[0054] In this embodiment, in the embedding layer, the number information of the surrounding intelligent agents is mapped to the vector space to obtain an embedding vector. In this embodiment, the corresponding embedding vector can be generated according to the total number of surrounding intelligent agents, or the embedding vectors corresponding to vehicles and pedestrians can be generated according to the number of vehicles and pedestrians in the surrounding intelligent agents.
[0055] Preferably, considering that pedestrians and vehicles have different impact logics on the target vehicle, for example, surrounding vehicles mostly have a direct impact on the target vehicle, while pedestrians on the sidewalk have a smaller impact on the target vehicle, and pedestrians on zebra crossings or intersections have a direct impact on the target vehicle. Therefore, a technical solution is adopted to generate embedding vectors corresponding to vehicles and pedestrians respectively according to the number of vehicles and pedestrians in the surrounding intelligent agents.
[0056] Specifically, in the embedding layer, the vector space is learnable based on the parameters and The number of surrounding vehicles and pedestrians in each scene (divided by 5 to facilitate model learning) can be matched to the corresponding vector representation and Among them, d (set to D / 2) is the vector dimension of the vector space. and They represent the maximum number of vehicles and pedestrians around the target vehicle respectively.
[0057] Step 2.2: Concatenate the embedding vectors corresponding to the surrounding intelligent agents with the position of the target vehicle to obtain heterogeneous scene information.
[0058] In this embodiment, E vehicle 、E pedestrian and the position of the target vehicle Perform stitching to obtain heterogeneous scene information:
[0059]
[0060] in Represents heterogeneous scene information.
[0061] Step 2.3: Input the heterogeneous scene information into the meta-learner to learn the weight parameters of the meta-linear layer.
[0062] In this embodiment, HS is input into the meta-learner to extract meta-knowledge and obtain the parameters of the meta-linear layer. and
[0063] The meta-learner is responsible for generating the parameters of the meta-linear layer and A common practice is to use a single linear layer to map HS to W meta and b meta However, this will result in a large number of parameters that need to be learned by the model. For example, (D+5)×(D·D) parameters are required to map HS to W meta To solve this problem, the meta-learner in this embodiment adopts a two-stage design: first, HS is mapped to a smaller dimension sd (set to 2), and then mapped to dimension D·D. This two-stage design only requires (D+5)×sd+sd×(D·D) parameters to map HS to D·D dimensions, as shown in equations (2) and (3).
[0064] W meta =σ(HSW w1 +b w1 )W w2 +b w2 (2)
[0065] b meta =σ(HSW b1 +b b1 )W b2 +b b2 (3)
[0066] in and All are learnable parameters; σ(.) represents the Sigmoid activation function.
[0067] Step 2.4: Input the agent code into the meta-linear layer, and then after the activation operation, perform the Hadamard product with the agent code, and then add it to the agent code after the Hadamard product to obtain the filtered agent code.
[0068] SGG uses scene information as a guide to filter the surrounding agent codes. For example, the process can be expressed as equation (4):
[0069]
[0070] in is the agent code after SGG filtering; ⊙ represents the Hadamard product.
[0071] Subsequently, the filtered agent encoding is used to update the lane encoding. The social interactions between agents largely follow the lane information. Based on this view, the encoder module uses the agent encoding filtered by SGG to update the lane encoding, thereby supporting the DDGCM of this embodiment to model the uncertain future relationships between agents.
[0072] In a specific embodiment, the lane code is updated using the filtered agent code, including:
[0073] Linearly map the filtered agent encoding to obtain the key vector and value vector;
[0074] Linearly map the lane encoding to obtain the query vector;
[0075] Perform a scaled dot product attention operation to obtain a lane encoding with agent context;
[0076] Then the lane coding with the agent context is concatenated with the lane coding and a linear transformation is performed to obtain the updated lane coding.
[0077] Formally, this process can be expressed as:
[0078]
[0079]
[0080] in, and are the key vector and value vector obtained by linearly mapping the filtered agent encoding; is the query vector obtained by linearly mapping the lane encoding; Attention(.) is the scaled dot product attention operation (agent-lane attention); is the lane encoding with agent context; and is a learnable parameter; || is a concatenation operation; a linear transformation is performed on the concatenated encoding to obtain the updated lane encoding, is the updated lane encoding. It should be noted that the scaled dot product attention operation only considers agents within a certain distance threshold to the lane node.
[0081] The future trajectory of the agent is inherently uncertain and depends on the scene information as well as the agent's intentions. Typically, drivers may make different decisions in different situations. For example, when there are more agents around (whether vehicles or pedestrians), drivers tend to adopt more cautious driving strategies. Although such scene heterogeneity affects driving behavior, this has not been fully utilized. Therefore, this embodiment proposes SGG to selectively filter information based on scene-specific non-shared parameters. SGG will assign similar model parameters to similar scenes, thereby establishing a scene-guided information bottleneck, which in turn promotes the model to achieve more socially aware trajectory prediction.
[0082] Step S3: Generate a lane attention map using the adjacency matrix of the lane map and the updated lane encoding through a self-attention mechanism.
[0083] Capturing social interactions between agents is crucial to achieving accurate and reliable trajectory prediction, and graph neural networks (GNNs) are very suitable for handling this task. Most current GNN-based methods use graph convolutional networks and graph attention networks to capture social interactions between agents. However, these methods usually only focus on fixed or single-hop social interactions, and it is difficult to effectively model the randomly propagated social relationships between agents. Therefore, this embodiment proposes a dynamic diffusion graph convolution module DDGCM, which aims to model highly random social interactions between agents. Specifically, this module models the social interactions between agents as a diffusion process on the lane map, and constructs a lane attention map to dynamically calculate the diffusion ratio.
[0084] like Figure 4 As shown, DDGCM uses the adjacency matrix A of the lane graph lane and updated lane coding First, it uses the self-attention mechanism to generate the lane attention map, as shown in formula (7).
[0085]
[0086] in, They are respectively the query vector, key vector, and value vector obtained by linearly mapping all updated lane encodings in the current scene; is the lane attention map under the lane map adjacency matrix mask; N V is the number of lane nodes, which varies in different scenarios; ⊙ represents the Hadamard product; Attention(.) here is the self-attention operation.
[0087] Step S4: Perform a diffusion convolution operation on the updated lane coding and lane attention map, and then add them to the updated lane coding after random dropout operation to obtain lane node features.
[0088] The dynamic diffusion graph convolution module DDGCM proposed in this embodiment implements diffusion convolution to explicitly capture the dynamics and randomness of social interactions between agents. Formally, this process can be expressed as:
[0089]
[0090]
[0091] Among them, D K represents the number of diffusion steps (set to 2); A f =A / rowsum(A) and A b =A T / rowsum(A T ) represent the forward and backward transfer probability matrices of the lane attention map respectively; is the set of all updated lane codes, size N V ×D; and is a learnable parameter; is the lane encoding containing social interaction information; Dropout(.) represents the random inactivation operation (dropout); are the lane node features output by the proposed DDGCM.
[0092] It should be noted that, in this embodiment, steps S1 to S4 are descriptions of the entire encoder module, wherein steps 2.1 to 2.4 describe the operations performed by the scene-guided gating unit (SGG), and steps S3 and S4 are descriptions of the dynamic diffusion graph convolution module DDGCM. The dynamic diffusion graph convolution module DDGCM proposed in this embodiment aims to comprehensively describe the highly random social interactions between agents. To achieve this goal, the module uses the dynamic diffusion process on the lane map to simulate the random social interactions between agents on the lane. Based on the bidirectional random walk on the lane attention map, the module explicitly captures the dynamics, uncertainty, and random propagation of social interactions between agents.
[0093] Step S5: Input the target code, lane node features, and the adjacency matrix of the lane graph into the strategy header to obtain a lane node set.
[0094] In order to find the lane segment most relevant to the future trajectory of the target vehicle, the strategy head aims to predict the discrete probability distribution from each lane node to the next lane node. Based on this probability distribution, this embodiment can determine the most likely future driving path of the target vehicle through graph traversal. Specifically, the strategy head uses the target encoding E target , Lane node features The adjacency matrix of the lane graph is used as input, and the probability distribution is predicted through multilinear mapping, as shown in formula (10):
[0095]
[0096] in, and yes A subset of; (i, j) is the successor edge in the lane graph structure; φ[.] represents a three-layer multilayer perceptron; N E is the number of successor edges of lane node i; S i,j is the probability that the vehicle at lane node i will travel to lane node j. i,j , the strategy head can get a set of lane nodes [v 1 ,v 2 ,…,v M ], these nodes represent the most likely future driving paths of the target vehicle.
[0097] It should be noted that the strategy head predicts probability distribution through multiple linear mappings and performs training through behavioral cloning. This is a relatively mature technology in this field and will not be elaborated here.
[0098] Step S6: Obtain lane node features corresponding to the lane node set, form a path code, and use the path code to update the target code.
[0099] Based on the lane node set provided by the policy header [v 1 ,v 2 ,…,v M ], the decoder module can obtain the path encoding This example uses E path to update the target encoding to be consistent with the most likely future driving path.
[0100] In a specific embodiment, the updating of the target code using the path code includes:
[0101] Linearly map the path encoding to obtain the key vector and value vector;
[0102] Linearly map the target encoding to obtain the query vector;
[0103] Perform a scaled dot product attention operation to obtain the target encoding with path context;
[0104] Then the target encoding with the path context is concatenated with the target encoding to obtain the updated target encoding.
[0105] The specific process is shown in formulas (11) and (12).
[0106] E target ′=Attention(Q target ,K path ,V path ) (11)
[0107] H target =E target ||E target ′ (12)
[0108] Among them, K path and V path are the key vector and value vector obtained by linear mapping path encoding; Q target is the query vector obtained by linearly mapping the target encoding; is the target encoding with path context; is the updated target encoding, which contains the most likely future driving path information. Here, Attention(.) is the scaled dot product attention operation (target vehicle-lane attention).
[0109] Step S7: The position of the target vehicle, the latent variables following the multivariate Gaussian distribution, and the updated target encoding are passed through a linear layer and then filtered through a scene-guided gating unit. The filtered features are passed through a linear layer to obtain a predicted trajectory.
[0110] In this embodiment, the decoder module aggregates information to generate the future trajectory of the target vehicle. In the decoder module, the position of the target vehicle, the latent variables following the multivariate Gaussian distribution, and the updated target encoding are passed through the linear layer and then filtered by the scene guided gating unit. The filtered features are passed through the linear layer to obtain the predicted trajectory.
[0111] The process can be expressed as the following formula:
[0112]
[0113] H hidden ′=f SGG (H hidden ) (14)
[0114]
[0115] in, is the position of the target vehicle; is a latent variable following a multivariate Gaussian distribution, used to enhance uncertainty modeling; and is a learnable parameter; D h is the dimension of the hidden feature; is a hidden feature; f SGG (.) is the operation set of the proposed SGG; is the hidden feature after filtering; is the predicted trajectory. LeakyReLU[.] is the activation function.
[0116] To achieve multimodal trajectory prediction, the proposed model samples various credible future driving paths in the policy head. It obtains different path encodings and generates K predicted trajectories
[0117] It should be noted that the scene guided gating unit SGG in the decoder module needs to filter the original code Steps S6 and S7 describe the operations performed by the decoder module.
[0118] In one embodiment of the present application, a multi-task loss function is used to train the proposed SGG-DDGCN model:
[0119]
[0120] Among them, λ 1 and λ 2 is a hyperparameter; E GT is the set of edges visited by the true trajectory; is the true trajectory at time step t.
[0121] Another embodiment of the present application further provides an autonomous driving trajectory prediction device based on scene-guided gating, comprising a processor and a memory storing a plurality of computer instructions, wherein the computer instructions, when executed by the processor, perform the steps of the above-mentioned autonomous driving trajectory prediction method based on scene-guided gating.
[0122] For the specific definition of the automatic driving trajectory prediction device based on scene-guided gating, please refer to the definition of the automatic driving trajectory prediction method based on scene-guided gating above, which will not be repeated here. The above-mentioned automatic driving trajectory prediction device based on scene-guided gating can be implemented in whole or in part by software, hardware and a combination thereof. It can be embedded in or independent of the processor in the computer device in hardware form, or it can be stored in the memory of the computer device in software form, so that the processor can call and execute the above corresponding operations.
[0123] The memory and the processor are electrically connected directly or indirectly to achieve data transmission or interaction. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines. The memory stores a computer program that can be run on the processor, and the processor implements the method in the embodiment of the present invention by running the computer program stored in the memory.
[0124] The memory may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), etc. The memory is used to store a program, and the processor executes the program after receiving an execution instruction.
[0125] The processor may be an integrated circuit chip with data processing capabilities. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc. The methods, steps and logic block diagrams disclosed in the embodiments of the present invention may be implemented or executed. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0126] In order to evaluate the technical solution of this application, a challenging and widely used trajectory prediction dataset is used to evaluate the proposed SGG-DDGCN. The dataset provides 2 seconds of historical trajectories and 6 seconds of future trajectories with a data interval of 0.5 seconds (i.e., 2Hz), and also includes a high-definition map of the scene. The dataset is divided into three parts: training set, validation set, and test set, containing 32186, 8560, and 9041 samples, respectively. Two widely used standard indicators are used for evaluation: (1) the minimum average displacement error (MinADE) of the top k predictions k ) and (2) miss rate (MissRate k,2 ), the former measures the minimum point-to-point L2 distance between the predicted trajectory and the true trajectory, while the latter indicates the proportion of the predicted results that are more than 2 meters away from the true trajectory.
[0127] The performance of the proposed SGG-DDGCN is compared with the state-of-the-art baseline models on the nuScenes benchmark, and the results are shown in Table 1. Overall, the traversal-based models (PGP, XHGP, and the proposed SGG-DDGCN) achieve more accurate trajectory predictions compared to other models, validating the effectiveness of the traversal-based prediction framework. In addition, driven by the proposed SGG and DDGCM, the proposed SGG-DDGCN outperforms the others in four metrics (MinADE 5 、MinADE 10 、MissRate 5,2 and MissRate 10,2 ) achieves the best performance among traversal-based models.
[0128] Table 1
[0129]
[0130] In order to study the contribution of each key component of the proposed SGG-DDGCN, this application conducts ablation experiments on the nuScenes benchmark. Specifically, the learning rate strategy (LRS), the proposed SGG and DDGCM are evaluated and the results are shown in Table 2. The experimental results show that compared with the base model (decay rate of 0.5 every 10 cycles) (row 1), the use of the proposed LRS (decay rate of 0.97 per epoch) (row 2) can better promote model training, thereby improving the accuracy of trajectory prediction. In addition, after adding the proposed SGG (row 3) and DDGCM (row 6), all four indicators of the model are improved. It is worth noting that adding GCN (row 4) or GAT (row 5) produces ambiguous results: some indicators become worse and some indicators become better. This may be due to the fact that the decoding module aggregates information about the agent context along the traversal path, making fixed or single-hop social interactions redundant.
[0131] Table 2
[0132]
[0133] In order to analyze the computational performance of the proposed SGG-DDGCN, this application compares the running efficiency of different models on the nuScenes benchmark. Specifically, according to the source code and recommended hyperparameters provided by the authors of PGP and XHGP, the traversal-based models (PGP, XHGP, and the proposed SGG-DDGCN) were tested on the same server. Table 3 shows the number of parameters of these models and the time required to complete a training or validation cycle. It can be observed that the proposed SGG-DDGCN does not significantly increase the number of parameters. In addition, the proposed SGG-DDGCN has higher training and reasoning efficiency while achieving excellent prediction performance.
[0134] Table 3
[0135]
[0136] Through extensive experiments on the nuScenes benchmark, this paper verifies the effectiveness and robustness of the proposed SGG-DDGCN.
[0137] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A method for predicting trajectory of autonomous driving based on scene-guided gating, characterized in that: The method for predicting the trajectory of an autonomous driving system based on scene-guided gating includes: Use gated recurrent units to encode the target vehicle’s past trajectory, lane map, and surrounding agent’s past trajectory to obtain target encoding, lane encoding, and agent encoding. The scene-guided gating unit is used to filter the agent code to obtain a filtered agent code, and the filtered agent code is used to update the lane code; The lane map adjacency matrix and the updated lane encoding are used to generate a lane attention map through the self-attention mechanism; Perform a diffusion convolution operation on the updated lane code and lane attention map, and then add them to the updated lane code after random dropout operation to obtain the lane node feature; Input the target code, lane node features and the adjacency matrix of the lane graph into the strategy head to obtain the lane node set; Get the lane node features corresponding to the lane node set, form a path code, and use the path code to update the target code; The position of the target vehicle, the latent variables following the multivariate Gaussian distribution, and the updated target encoding are passed through a linear layer and then filtered by a scene-guided gating unit. The filtered features are passed through a linear layer to obtain the predicted trajectory.
2. The method for automatic driving trajectory prediction based on scene-guided gating according to claim 1, characterized in that: The scene guides the gating unit to perform the following operations: Construct the embedding vectors corresponding to the surrounding agents through the embedding layer; Concatenate the embedding vectors corresponding to the surrounding agents with the position of the target vehicle to obtain heterogeneous scene information; Input the heterogeneous scene information into the meta-learner to learn the weight parameters of the meta-linear layer; The original code to be filtered is input into the meta-linear layer, and then after the activation operation, the Hadamard product is performed with the original code, and after the Hadamard product, it is added with the original code to obtain the filtered code.
3. The method for automatic driving trajectory prediction based on scene-guided gating according to claim 1, characterized in that: The method of using the filtered agent code to update the lane code includes: Linearly map the filtered agent encoding to obtain the key vector and value vector; Linearly map the lane encoding to obtain the query vector; Perform a scaled dot product attention operation to obtain a lane encoding with agent context; Then the lane coding with the agent context is concatenated with the lane coding and a linear transformation is performed to obtain the updated lane coding.
4. The method for automatic driving trajectory prediction based on scene-guided gating according to claim 1, characterized in that: The process of generating a lane attention map by using the adjacency matrix of the lane map and the updated lane coding through a self-attention mechanism includes: Linearly map the updated lane coding to obtain the query vector, key vector and value vector; Perform a self-attention operation and then perform a Hadamard product with the adjacency matrix of the lane map to obtain the lane attention map.
5. The method for automatic driving trajectory prediction based on scene-guided gating according to claim 1, characterized in that: The method of updating the target encoding by using the path encoding includes: Linearly map the path encoding to obtain the key vector and value vector; Linearly map the target encoding to obtain the query vector; Perform a scaled dot product attention operation to obtain the target encoding with path context; Then the target encoding with the path context is concatenated with the target encoding to obtain the updated target encoding.
6. An automatic driving trajectory prediction device based on scene-guided gating, comprising a processor and a memory storing a plurality of computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Obstacle trajectory prediction method and device
CN111079721A
Trajectory prediction method, device and equipment and storage medium
CN111523643A
Intelligent vehicle track prediction system and method fusing peripheral vehicle interaction information
CN113954864A
Multi-agent position prediction method and device, electronic equipment and storage medium
CN114239974A
Trajectory prediction method, device and equipment
CN116654012A