Training Method for Trajectory Prediction Model and Trajectory Prediction Method

Through the combination of embedded layer, encoder and decoder, combined with vehicle trajectory and map information, the trajectory prediction model is optimized, which solves the problem of insufficient trajectory prediction accuracy in the existing model in the autonomous driving system, and improves prediction accuracy and storage performance.

CN119918587BActive Publication Date: 2025-07-11NEOLITHIC HUITONG TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510415133.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-11
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing trajectory prediction model is difficult to effectively combine vehicle historical trajectory and local map information in the autonomous driving system, resulting in insufficient accuracy of path planning and vehicle control.

Method used

Using a combination of embedded layer, encoder and decoder, through feature extraction, encoding and decoding, combining vehicle trajectory and map information, a multi-head attention mechanism and a hybrid Gaussian model are used to predict trajectory, and the model training process is optimized to improve prediction accuracy.

Benefits of technology

It improves the accuracy and storage performance of the trajectory prediction model, can better maintain local structures, enhances the prediction ability of obstacle vehicles, and improves the safety of the autonomous driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119918587B_ABST
    Figure CN119918587B_ABST
Patent Text Reader

Abstract

The present disclosure provides a training method for a trajectory prediction model and a trajectory prediction method, which relate to the field of artificial intelligence technologies such as autonomous driving and intelligent transportation. Among them, the initial trajectory prediction model includes: an embedding layer, an encoder, and a decoder. The method includes: using the embedding layer to extract features from the historical trajectory information of a sample vehicle and the local map information associated with the historical trajectory information to obtain vehicle trajectory features and map features; using the encoder to encode the vehicle trajectory features and map features to obtain vehicle trajectory encodings and map element encodings; using the decoder to decode the vehicle trajectory encodings and map element encodings to obtain the target prediction trajectory of the target vehicle; training the initial trajectory prediction model according to the loss between the target prediction trajectory and the true driving trajectory of the target vehicle to obtain a trajectory prediction model. The method provided by the present disclosure improves the prediction accuracy of the trajectory prediction model and the storage performance of the trajectory prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technologies, specifically to the fields of autonomous driving and intelligent transportation technologies, and particularly to a method for training a trajectory prediction model and a trajectory prediction method. Background Art

[0002] In an autonomous driving system, the vehicle trajectory prediction result determines the path planning of the vehicle and the vehicle's own control. A current series of trajectory prediction schemes focus on the modeling and interaction of the scenario, that is, the encoder architecture, such as VectorNet (Vectorized Scene Representation Network for Trajectory Prediction), PointNet encoding method, LaneGCN, etc. Summary of the Invention

[0003] The present disclosure provides a method for training a trajectory prediction model and a trajectory prediction method.

[0004] According to a first aspect of the present disclosure, there is provided a method for training a trajectory prediction model, wherein the initial trajectory prediction model includes: an embedding layer, an encoder, and a decoder. The method includes: using the embedding layer to extract features from the historical trajectory information of a sample vehicle and the local map information associated with the historical trajectory information to obtain vehicle trajectory features and map features, wherein the sample vehicle includes a target vehicle and an obstacle vehicle; using the encoder to encode the vehicle trajectory features and map features to obtain vehicle trajectory encodings and map element encodings; using the decoder to decode the vehicle trajectory encodings and map element encodings to obtain the target prediction trajectory of the target vehicle; and training the initial trajectory prediction model according to the loss between the target prediction trajectory and the true driving trajectory of the target vehicle to obtain the trajectory prediction model.

[0005] According to a second aspect of the present disclosure, there is provided a trajectory prediction method, including: obtaining a first current driving trajectory of a target vehicle and a second current driving trajectory of an obstacle vehicle; inputting the first current driving trajectory and the second current driving trajectory into the trajectory prediction model, and outputting the predicted trajectory of the target vehicle, wherein the trajectory prediction model is trained by using the method described in any implementation manner of the first aspect.

[0006] According to a third aspect of the present disclosure, there is provided a training device for a trajectory prediction model. The initial trajectory prediction model includes: an embedding layer, an encoder, and a decoder. The device includes: an embedding module configured to use the embedding layer to extract features from the historical trajectory information of a sample vehicle and the local map information associated with the historical trajectory information, to obtain vehicle trajectory features and map features, where the sample vehicle includes a target vehicle and an obstacle vehicle; an encoding module configured to use the encoder to encode the vehicle trajectory features and the map features to obtain vehicle trajectory encodings and map element encodings; a decoding module configured to use the decoder to decode the vehicle trajectory encodings and the map element encodings to obtain a target prediction trajectory of the target vehicle; and a training module configured to train the initial trajectory prediction model according to the loss between the target prediction trajectory and the actual driving trajectory of the target vehicle to obtain a trajectory prediction model.

[0007] According to a fourth aspect of the present disclosure, there is provided a trajectory prediction method, including: an acquisition module configured to acquire a first current driving trajectory of a target vehicle and a second current driving trajectory of an obstacle vehicle; and a prediction module configured to input the first current driving trajectory and the second current driving trajectory into a trajectory prediction model and output a predicted trajectory of the target vehicle, where the trajectory prediction model is trained by using the method described in any implementation manner of the first aspect.

[0008] According to a fifth aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any implementation manner of the first aspect or the second aspect.

[0009] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in any implementation manner of the first aspect or the second aspect.

[0010] According to a seventh aspect of the present disclosure, there is provided a computer program product including a computer program, where the computer program, when executed by a processor, executes the method described in any implementation manner of the first aspect or the second aspect.

[0011] According to an eighth aspect of the present disclosure, there is provided a self-driving vehicle including the electronic device described in the fifth aspect.

[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0013] The accompanying drawings are used to better understand the present solution and do not constitute a limitation to the present disclosure. Among them:

[0014] Figure 1 is a flowchart of the first embodiment of the training method of the trajectory prediction model according to the present disclosure;

[0015] Figure 2 is a flowchart of the second embodiment of the training method of the trajectory prediction model according to the present disclosure;

[0016] Figure 3 is a flowchart of the third embodiment of the training method of the trajectory prediction model according to the present disclosure;

[0017] Figure 4 is a flowchart of the fourth embodiment of the training method of the trajectory prediction model according to the present disclosure;

[0018] Figure 5-1 is an application block diagram of the training method of the trajectory prediction model according to the present disclosure;

[0019] Figure 5-2 is a schematic structural diagram of another decoder;

[0020] Figure 5-3 is a schematic structural diagram of yet another decoder;

[0021] Figure 6 is a flowchart of an embodiment of the trajectory prediction method according to the present disclosure;

[0022] Figure 7 is a schematic structural diagram of an embodiment of the training device of the trajectory prediction model according to the present disclosure;

[0023] Figure 8 is a schematic structural diagram of an embodiment of the trajectory prediction device according to the present disclosure;

[0024] Figure 9 is a block diagram of an electronic device for implementing the training method or the trajectory prediction method of the trajectory prediction model in the embodiments of the present disclosure. Specific Embodiments

[0025] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0026] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments may be combined with each other. The following will detail the present disclosure with reference to the drawings and in conjunction with the embodiments.

[0027] An exemplary system architecture for implementing the training method or trajectory prediction method of the trajectory prediction model provided by the present disclosure may include a terminal device, a network, and a server. The network is used to provide a communication link between the terminal device and the server and may include various connection types, such as wired communication links, wireless communication links, or fiber optic cables, etc.

[0028] Users can use the terminal device to interact with the server through the network to receive or send information, etc. Various client applications can be installed on the terminal device, such as map-based, navigation-based, entertainment-based, etc. client applications.

[0029] The terminal device may be, for example, the in-vehicle system of vehicles such as autonomous driving vehicles and delivery robots. This system can be implemented in a hardware manner, or in a software manner, or in a manner combining hardware and software.

[0030] The server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules (such as those used to provide distributed services) or as a single software or software module. No specific limitation is made here.

[0031] It should be noted that the execution subject (hereinafter simply referred to as the "execution subject") of the training method or trajectory prediction method of the trajectory prediction model provided by the present disclosure can be executed by the server in the above system architecture or can be implemented through the above terminal device. For example, when the server executes, the server sends a control command to the client, such as the in-vehicle system of a vehicle, and then the in-vehicle system predicts the trajectory of the vehicle according to the control command, such as predicting the driving trajectory of the vehicle, and controls the driving of the vehicle according to the predicted trajectory. Another example is that when the terminal device (such as the in-vehicle system of a vehicle) executes, the terminal device predicts the trajectory of the vehicle according to the control command, such as predicting the driving trajectory of the vehicle, and controls the driving of the vehicle according to the predicted trajectory.

[0032] Figure 1 Flow 100 of the first embodiment of the training method of the trajectory prediction model according to the present disclosure is shown. In this embodiment, the initial trajectory prediction model includes an embedding layer, an encoder, and a decoder, and the training method of the trajectory prediction model includes the following steps:

[0033] Step 101: Use the embedding layer to extract features from the historical trajectory information of the sample vehicle and the local map information associated with the historical trajectory information, to obtain vehicle trajectory features and map features.

[0034] In this embodiment, the execution entity will use the embedding layer to extract features from the historical trajectory information of the sample vehicle and the local map information associated with the historical trajectory information, to obtain vehicle trajectory features and map features. Here, the sample vehicle includes the target vehicle and obstacle vehicles.

[0035] Here, the target vehicle is the ego vehicle, and the obstacle vehicles are the obstacles around the ego vehicle. The number of obstacle vehicles can be multiple. Since the sample vehicle includes the ego vehicle and obstacle vehicles, the historical trajectory information of the sample vehicle includes the historical trajectory information of the ego vehicle and the historical trajectory information of the obstacle vehicles. The historical trajectory information can be the trajectory information within a preset time range, such as the trajectory information within one hour from the current moment; the historical trajectory information can be the trajectory information within a preset distance, such as the trajectory information within five kilometers from the current position.

[0036] In addition, the above execution entity will obtain the local map information associated with the historical trajectory information from the pre-constructed high-precision map. For example, the above execution entity can first determine a position range according to the historical trajectory information, and then obtain the map corresponding to this position range from the high-precision map, so as to obtain the local map.

[0037] It should be noted that the Embedding layer is a special layer in the neural network, which is used to map discrete input data into a continuous vector space.

[0038] In this embodiment, the historical trajectory of the sample vehicle and the obtained local map corresponding to the historical trajectory are input into the embedding layer. The embedding layer will extract features from the historical trajectory to obtain vehicle trajectory features, and will also extract features from the local map to obtain map features.

[0039] Step 102: Use the encoder to encode the vehicle trajectory features and map features, to obtain vehicle trajectory encodings and map element encodings.

[0040] In this embodiment, the above execution entity will use the encoder to encode the vehicle trajectory features and map features, so as to obtain the encoded vehicle trajectory encodings and map element encodings.

[0041] In a neural network, the task of the encoder is to encode the input sequence into a context vector. The encoder usually consists of one or more neuron layers, and each neuron receives the input data and transforms it into an output through an activation function. The role of the encoder is to compress the input data into a lower-dimensional representation for subsequent processing by the neural network.

[0042] The encoder in this embodiment is composed of multiple layers of MHA (Multi-Head Attention). For example, the encoder can be obtained by stacking 6 layers of MHA. Here, the multi-head attention mechanism (MHA) captures different features in the input sequence by computing multiple attention heads in parallel. Each attention head has its own query (Q), key (K), and value (V) matrices, and their main functions are as follows: Query matrix Q: The query matrix is the 'question' for finding certain information. The query matrix is a projection of the input, representing the requirements of the current token (the smallest discrete unit in text or sequence data) for other tokens, which can help determine its position in the sequence and what content to pay attention to; Key matrix K: The key matrix is the 'information' or 'identifier' provided by each token. Each token has a key associated with it, which is used to compare with the query to determine its relevance to the query; Value matrix V: The value is the actual information, providing the content of the word vector. According to the matching degree of Q and K, V is finally used to generate the output vector.

[0043] In this embodiment, the vehicle trajectory features and map features are input into the encoder. The multiple layers of MHA in the encoder will capture different features in these two sequences and encode the features, thereby outputting the vehicle trajectory encoding and map element encoding.

[0044] Step 103, use the decoder to decode the vehicle trajectory encoding and map element encoding to obtain the target prediction trajectory of the target vehicle.

[0045] In this embodiment, the above-mentioned execution entity will use the decoder to perform decoding processing on the vehicle trajectory encoding and map element encoding, thereby obtaining the target prediction trajectory of the target vehicle after the decoding processing.

[0046] In a neural network, the task of the decoder is to generate an output sequence from the context vector. The decoder usually reads the context vector of the encoder first and then starts to generate the output sequence. In the trajectory prediction scenario of this embodiment, the decoder will decode the feature vector output by the encoder into specific trajectory data, thereby outputting the target prediction trajectory of the target vehicle.

[0047] The decoder in this embodiment is also composed of multiple layers of MHA (Multi-Head Attention), and the output of the previous attention layer is the input of the next attention layer.

[0048] Step 104: Train the initial trajectory prediction model according to the loss between the target prediction trajectory and the actual driving trajectory of the target vehicle to obtain a trajectory prediction model.

[0049] In this embodiment, the above-mentioned execution entity will train the initial trajectory prediction model according to the loss between the target prediction trajectory and the actual driving trajectory of the target vehicle to obtain a trajectory prediction model.

[0050] The model training samples in this embodiment include the historical trajectory information of the sample vehicle (i.e., the historical trajectory information of the target vehicle and the historical trajectory information of the obstacle vehicle) and the actual driving trajectory of the target vehicle. Here, the historical trajectory information and the actual driving trajectory are for the current moment (also called the prediction moment). That is, at the current moment, the trajectory of the target vehicle needs to be predicted. Then, the trajectory before the current moment is called the historical trajectory, and the trajectory after the current moment is called the actual driving trajectory. The above-mentioned execution entity will generate a target prediction trajectory according to the foregoing steps. The target prediction trajectory is the trajectory that the predicted target vehicle will travel after the current moment.

[0051] The above-mentioned execution entity will calculate the loss between the target prediction trajectory and the actual driving trajectory. For example, calculate the regression loss between the target prediction trajectory and the actual driving trajectory according to the regression loss function, and then use this loss to adjust the parameters of the embedding layer, encoder, and decoder of the initial trajectory prediction model until the preset conditions are met to obtain a trained trajectory prediction model.

[0052] The training method of the trajectory prediction model provided by the embodiments of the present disclosure first uses the embedding layer to extract features from the historical trajectory information of the sample vehicle and the local map information associated with the historical trajectory information to obtain vehicle trajectory features and map features; then, uses the encoder to encode the vehicle trajectory features and map features to obtain vehicle trajectory encodings and map element encodings; then, uses the decoder to decode the vehicle trajectory encodings and map element encodings to obtain the target prediction trajectory of the target vehicle; finally, trains the initial trajectory prediction model according to the loss between the target prediction trajectory and the actual driving trajectory of the target vehicle to obtain a trajectory prediction model. When training the trajectory prediction model, this method encodes based on a local connection graph (instead of a global connection graph), so that the local structure can be better maintained, and the prediction accuracy of the trajectory prediction model is improved; in addition, a larger map encoding is achieved through more efficient storage, thereby improving the storage performance of the trajectory prediction model.

[0053] In addition, in the technical solutions involved in the present disclosure, the acquisition, storage, use, processing, transportation, provision, and disclosure of the trajectory information of the vehicle and the like all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0054] Continuing to refer to Figure 2 , Figure 2 Figure 200 shows the flow of the second embodiment of the training method of the trajectory prediction model according to the present disclosure. The training method of the trajectory prediction model includes the following steps:

[0055] Step 201: Input the historical trajectory information and the local map information into the normalization layer, and output the vehicle trajectory matrix and the map element matrix.

[0056] In this embodiment, the embedding layer of this embodiment includes a normalization layer and a conversion layer. The execution subject of the training method of the trajectory prediction model first inputs the historical trajectory information (the historical trajectory information of the target vehicle and the historical trajectory information of the obstacle vehicle) and the local map information of the sample vehicle into the normalization layer to perform normalization processing on the historical trajectory information and the local map information.

[0057] Specifically, the local map information and the obstacle trajectory information (i.e., the vehicle trajectory information) can be converted into a local normalized representation centered on the target obstacle through the following formula. These are collectively referred to as polyline. After normalization processing, the vehicle trajectory matrix and the map element matrix can be obtained:

[0058]

[0059] where is the vehicle historical trajectory information, is the local map information, and MLP is a multi-layer perceptron.

[0060] After normalization processing, has a dimension of Agent[N, T, C]. Among them, N represents the number of obstacles as 128, T is the number of historical frames as 30, and C is the number of vehicle features as 15. The vehicle features include x, y, heading, vel, etc. x and y represent coordinates, heading represents the vehicle body orientation, and vel represents the velocity component.

[0061] After normalization processing, has a dimension of Map[M, N, C]. M is the number of map elements as 768, N is the number of points of each map element as 20, and C is the number of map features as 6. The map features include x, y, type, etc. x and y represent coordinates, and type represents the type of the map element.

[0062] Step 202: Input the vehicle trajectory matrix and the map element matrix into the conversion layer, and output the vehicle trajectory features and the map features.

[0063] In this embodiment, the above-mentioned execution entity inputs the vehicle trajectory matrix and the map element matrix into the conversion layer, that is, the normalized polyline is encoded and converted using MLP (Multilayer Perceptron) and MaxPooling (Max Pooling). The vehicle trajectory matrix and the map element matrix are used as the original feature encoding, and the finally encoded vehicle trajectory features and the map features have tensor dimensions of [N, D] and [M, D] respectively, where D = 256. The MLP of the agent (obstacle) consists of three layers of 256 and two layers of 256, and the MLP of the map (map) consists of a 5-layer 64-dimensional MLP and a 2-layer 256-dimensional MLP.

[0064] The vehicle trajectory information and the local map information are processed through the embedding layer, so that the relationship between the data can be captured through the embedding layer, and the data is converted into low-dimensional vectors for subsequent processing.

[0065] Step 203: Input the vehicle trajectory features and the map features into the encoding network, and output the vehicle trajectory encoding and the map element encoding.

[0066] In this embodiment, the encoder in this embodiment includes an encoding network and a regression output layer. The above-mentioned execution entity inputs the vehicle trajectory features and the map features into the encoding network first, and then outputs the vehicle trajectory encoding and the map element encoding.

[0067] The encoder in this embodiment is composed of multiple layers of MHA. For example, the encoder can be obtained by stacking 6 layers of MHA, and the output of the previous layer of MHA is the input of the next layer of MHA. The vehicle trajectory features and the map features are input into the encoding network of MHA, that is, the input is [Agent, Map], and the dimension is [N + M, D], that is, [128 + 768, 256]. It can be expressed as the formula:

[0068] MultiHeadAttn(query = +P , key = k( )+P , value = k( ))

[0069] Among them, the input of the first layer = , k() represents finding the 16 nearest Map polylines for each Agent polyline (obstacle trajectory line) using the KNN (K-Nearest Neighbor) algorithm. PE represents position encoding, and finally, the vehicle trajectory encoding A_past and the map element encoding M are output.

[0070] Step 204: Input the vehicle trajectory encoding into the regression output layer to output the initial predicted trajectory and predicted speed of the sample vehicle.

[0071] In this embodiment, the above-mentioned execution entity inputs the vehicle trajectory encoding into the regression output layer, thereby outputting the initial predicted trajectory and predicted speed of the sample vehicle, which can be expressed as:

[0072] =MLP( )

[0073] where represents the initial predicted trajectory and predicted speed of the obstacle from time 1 to T, and i represents the number of future time steps.

[0074] Step 205: Perform local normalization processing and encoding conversion on the initial predicted trajectory and predicted speed to obtain the vehicle predicted trajectory encoding.

[0075] In this embodiment, the above-mentioned execution entity will process using the method of embedding layer encoding, that is, perform local normalization processing and encoding conversion on the initial predicted trajectory and predicted speed, thereby obtaining the vehicle predicted trajectory encoding, that is, encoding as A_future.

[0076] Step 206: Concatenate the vehicle predicted trajectory encoding and the vehicle trajectory encoding to obtain the vehicle trajectory encoding.

[0077] In this embodiment, the above-mentioned execution entity will concatenate the vehicle predicted trajectory encoding A_future and the vehicle trajectory encoding A_past, and then wrap 3 layers of MLP (the dimension of the fully connected layer is 512) as the enhanced representation of A, which can be expressed as:

[0078] A = MLP( )

[0079] where A is the vehicle trajectory encoding.

[0080] Here, future trajectories (predicted) are used for modeling interactions. By leveraging the obtained scene encoding information, a simple regression output head is externally connected to output the future trajectories and speeds of all obstacles, and vehicle trajectory encodings are generated based on the future trajectories and speeds. Therefore, the vehicle trajectory encodings generated by the encoder provide additional future context information for the decoder network, enabling the trajectory prediction model to predict more future trajectories that conform to the scene when predicting target obstacles.

[0081] Step 207: Use the decoder to decode the vehicle trajectory encoding and the map element encoding to obtain the target prediction trajectory of the target vehicle.

[0082] Step 208: Train the initial trajectory prediction model based on the loss between the target prediction trajectory and the true driving trajectory of the target vehicle to obtain the trajectory prediction model.

[0083] Steps 207 - 208 are basically the same as steps 103 - 104 in the foregoing embodiment. The specific implementation manner can refer to the description of steps 103 - 104 above and will not be elaborated here.

[0084] From Figure 2 it can be seen that compared with the corresponding embodiment in Figure 1 for the training method of the trajectory prediction model in this embodiment, this method highlights the step of encoding using the encoder. By using future trajectories (predicted) for modeling interactions, that is, leveraging the obtained scene encoding information, a simple regression output head is externally connected to output the future trajectories and speeds of all obstacles, and vehicle trajectory encodings are generated based on the future trajectories and speeds. Therefore, the vehicle trajectory encodings generated by the encoder provide additional future context information for the decoder network, enabling the trajectory prediction model to predict more future trajectories that conform to the scene when predicting target obstacles.

[0085] Continuing to refer to Figure 3 , Figure 3 shows the flow 300 of the third embodiment of the training method of the trajectory prediction model according to the present disclosure. The training method of the trajectory prediction model includes the following steps:

[0086] Step 301: Use the embedding layer to extract features from the historical trajectory information of the sample vehicle and the local map information associated with the historical trajectory information to obtain vehicle trajectory features and map features.

[0087] Step 302: Use the encoder to encode the vehicle trajectory features and the map features to obtain vehicle trajectory encodings and map element encodings.

[0088] Steps 301-302 are basically the same as steps 101-102 of the foregoing embodiment. The specific implementation manner can refer to the description of steps 101-102 above, and will not be elaborated here.

[0089] Step 303: For each decoding layer, input the vehicle trajectory encoding, map element encoding, and pre-generated global intention vector representation and auxiliary vector representation into the decoding network of this decoding layer, and output the decoded feature information.

[0090] In this embodiment, the decoder includes N decoding layers, and each decoding layer includes a decoding network and a decoding output layer, where N is a positive integer. For example, the decoder can consist of 6 decoding layers.

[0091] For each of the N decoding layers, a global intention vector representation and an auxiliary vector representation are pre-generated, and the global intention vector representation (also referred to as a static intention query vector) and the auxiliary vector representation (also referred to as a dynamic search query vector) are used as an intention query vector pair, so as to model the prediction task as a joint optimization task of global intention localization and local motion optimization. Specifically, the above-mentioned execution entity will input the vehicle trajectory encoding, map element encoding, and pre-generated intention query vector pair into the decoding network of this decoding layer, so as to output the decoded feature information.

[0092] In some optional implementation manners of this embodiment, the global intention vector representation is generated through the following steps: clustering the end points of the predicted trajectory to obtain multiple cluster centers; using a fully connected layer to transform the position encodings corresponding to the multiple cluster centers to obtain the global intention vector representation.

[0093] In this implementation manner, the above-mentioned execution entity performs clustering processing on the end points of the future trajectories of the obstacles through the k-means algorithm, so as to obtain multiple cluster centers, and the obtained cluster centers are the intention representations. The clustering dimension considers both direction and speed, represented by I. Performing position encoding and fully connected layer transformation on I can obtain the global intention vector representation , which can be expressed as:

[0094] =MLP(PE(I))

[0095] Among them, PE(I) represents the position encoding corresponding to I.

[0096] Each intention query is responsible for predicting the trajectory of a specific motion pattern, thus stabilizing the training process and promoting the prediction of multi-modal trajectories, because each motion pattern has its own learnable embedding. Therefore, only 64 query vectors are needed to achieve a good effect, instead of dense target points to cover the intention targets of the obstacles.

[0097] In some alternative implementation manners of this embodiment, the auxiliary vector representation is generated through the following steps: determining a map element line corresponding to the initial prediction trajectory and meeting a preset condition from local map information, and generating an auxiliary vector representation based on the map element line.

[0098] In this implementation manner, it aims to refine the trajectory by iteratively collecting fine-grained trajectory features to supplement the global intention positioning.

[0099] Each dynamic search query vector is also the position embedding encoding of a spatial point, initialized with its corresponding intention point, and dynamically updated according to the prediction trajectory of each decoder layer. For each motion query pair, a dynamic map collection module is designed to extract fine-grained trajectory features (querying map features from the local area aligned with the trajectory, 128 polylines of the map closest to the prediction trajectory), that is, querying the closest map polylines with fine-grained trajectories, which can be specifically expressed as:

[0100] =MLP(PE( ))

[0101] where j is the number of decoder layers, is the trajectory feature corresponding to time T, and PE represents position encoding.

[0102] Since the behavior of obstacles depends to a large extent on the map, this strategy can continuously focus on the latest local context information for iterative motion optimization.

[0103] In some alternative implementation manners of this embodiment, the decoding network includes: a first decoding sub-network, a second decoding sub-network, a third decoding sub-network, and a decoding fully connected layer, and each decoding sub-network is an MHA; and step 303 further includes steps 3031 - 3034, specifically:

[0104] Step 3031, inputting the global intention vector representation into the first decoding sub-network, and outputting the first decoding information.

[0105] Here, the global intention vector representation will be input into the first decoding sub-network first, so as to output the first decoding information , and the first decoding sub-network can be expressed as:

[0106] MultiHeadAttn(query = + , key = + , value = )

[0107] Among them, j is the number of decoder layers, is the output of the previous decoder layer. It should be noted that is initialized to 0.

[0108] That is, for this MHA, the query is the output of the previous decoder layer and the global intent vector representation , the key is the output of the previous decoder layer and , the value is the output of the previous decoder layer. Through this MHA, can be obtained.

[0109] Step 3032: Input the first decoding information, the auxiliary vector representation, the vehicle trajectory encoding, and the position encoding corresponding to the vehicle trajectory encoding into the second decoding sub-network, and output the second decoding information.

[0110] Here, the above-mentioned execution entity will input , , the vehicle trajectory encoding A and the position encoding P corresponding to A into the second decoding sub-network, so as to output the second decoding information , and the second decoding sub-network can be expressed as:

[0111] MultiHeadAttn(query = , key = , value = )

[0112] That is, for this MHA, the query is and , the key is A and P , the value is A. Through this MHA, can be obtained.

[0113] Step 3033: Input the first decoding information, the auxiliary vector representation, the map element line, and the position encoding corresponding to the map element line into the third decoding sub-network, and output the third decoding information.

[0114] Here, the above-mentioned execution entity will input , , the map element line M and the position encoding P corresponding to M into the third decoding sub-network, so as to output the third decoding information , and the third decoding sub-network can be expressed as:

[0115] MultiHeadAttn(query = , key = , value = α(M))

[0116] Among them, α(M) is the output of the dynamic map collection module, that is, M.

[0117] That is, for this MHA, the query is and , key is M and P , value is M, and through this MHA, can be obtained.

[0118] Step 3034, input the second decoding information and the third decoding information into the decoding fully connected layer, and output the decoded feature information.

[0119] Finally, input and into the decoding fully connected layer, and the decoded feature information can be output , which can be expressed as:

[0120] = MLP( )

[0121] Thus, the decoding information corresponding to the encoded information is generated through each decoding sub-network and the decoding fully connected layer of the decoding layer, and the accuracy of the obtained decoding information is improved through the multi-layer MHA structure.

[0122] Step 304, input the decoded feature information into the decoding output layer of this decoding layer, and output the probability distribution of each position point in the initial predicted trajectory.

[0123] In this embodiment, the above execution subject will input the decoded feature information into the decoding output layer of this decoding layer, thereby outputting the probability distribution of each position point in the initial predicted trajectory. The decoding output layer here includes a mixture of Gaussian models, and the mixture of Gaussian models can be expressed as:

[0124] = MLP( )

[0125] =

[0126] Among them, includes N1: k ( ; ) Gaussian distributions and probability distributions, i is between 1 - T, P1: k has a total of 6 elements, P(o) is the probability that an obstacle appears at position o, and the predicted trajectory is generated by simply extracting the predicted center of the Gaussian distribution.

[0127] It should be noted that and is and the mean value of and is and the standard deviation of is the correlation coefficient, which refers to and the correlation between

[0128] Step 305: Generate a target prediction trajectory according to the probability distribution of each position point in the initial prediction trajectories output by N decoding layers.

[0129] In this embodiment, since the decoding output layer of each decoding layer outputs the probability distribution of each position point in the initial prediction trajectory, the above-mentioned execution entity comprehensively generates a target prediction trajectory according to the probability distribution of each position point in the initial prediction trajectories output by N decoding layers. For example, using NMS (Non-Maximum Suppression principle), 6 final output trajectories are selected from 64 candidate trajectories based on the trajectory end position, and then the target prediction trajectory is determined from these 6 final output trajectories. Here, a probability distribution of each coordinate value x and y is output instead of a definite trajectory coordinate value, so as to further determine the target prediction trajectory according to the probability distributions of multiple layers, improving the accuracy of the target prediction trajectory.

[0130] Step 306: Train the initial trajectory prediction model according to the loss between the target prediction trajectory and the true driving trajectory of the target vehicle to obtain a trajectory prediction model.

[0131] Step 306 is basically the same as step 104 in the foregoing embodiment, and the specific implementation manner can refer to the description of step 104 above, which will not be elaborated here.

[0132] It can be seen from Figure 3 that compared with the corresponding embodiment of Figure 1 the training method of the trajectory prediction model in this embodiment highlights the step of decoding using a decoder. Specifically, using the intention query vector pair to model the prediction task as a joint optimization task of global intention localization and local motion optimization, the learnable intention query pair not only eliminates a large number of target candidates, but also predicts different motion patterns through specific pattern motion queries, thus reducing the optimization burden of the model.

[0133] Continue to refer to Figure 4 , Figure 4Flowchart 400 shows the fourth embodiment of the training method of the trajectory prediction model according to the present disclosure. The training method of the trajectory prediction model includes the following steps:

[0134] Step 401, using an embedding layer to extract features from the historical trajectory information of the sample vehicle and the local map information associated with the historical trajectory information, to obtain vehicle trajectory features and map features.

[0135] Step 402, using an encoder to encode the vehicle trajectory features and map features, to obtain vehicle trajectory encodings and map element encodings.

[0136] Step 403, for each decoding layer, input the vehicle trajectory encodings, map element encodings, and pre-generated global intention vector representations and auxiliary vector representations into the decoding network of this decoding layer, and output the decoded feature information.

[0137] Step 404, input the decoded feature information into the decoding output layer of this decoding layer, and output the probability distribution of each position point in the initial predicted trajectory.

[0138] Step 405, generate a target predicted trajectory according to the probability distribution of each position point in the initial predicted trajectory output by N decoding layers.

[0139] Steps 401-405 are basically the same as steps 301-305 of the foregoing embodiment. The specific implementation manner can refer to the description of steps 301-305 above, and will not be elaborated here.

[0140] Step 406, calculate the regression loss between the target predicted trajectory and the actual driving trajectory.

[0141] In this embodiment, the execution subject of the training method of the trajectory prediction model can use the following L1 regression loss function to calculate the regression loss L1 between the target predicted trajectory and the actual driving trajectory:

[0142] L1 =

[0143] where n is the number of samples, is the target predicted trajectory, the actual driving trajectory.

[0144] Step 407, generate a comprehensive loss according to the regression loss and the target Gaussian loss.

[0145] In this embodiment, since each decoding layer in the decoder includes a mixture of Gaussian models, therefore, the above execution subject will also calculate the Gaussian loss in the entire decoding process, that is, the target Gaussian loss. Finally, the comprehensive loss Ltotal is calculated according to the regression loss L1 and the target Gaussian loss LG.

[0146] In some alternative implementation manners of this embodiment, the target Gaussian loss is obtained through the following steps: for each decoding layer, calculate the Gaussian regression loss of this decoding layer; add the Gaussian regression losses of N decoding layers to obtain the target Gaussian loss.

[0147] In this implementation manner, for each decoding layer, the above-mentioned execution entity calculates the Gaussian loss of this decoding layer according to the following formula:

[0148]

[0149] =log +log +0.5log(1 - ( -2 ) -

[0150] Then, add the Gaussian regression losses of N decoding layers to obtain the target Gaussian loss LG. By first calculating the process Gaussian loss of each decoding layer and then adding the Gaussian losses of N decoding layers to obtain the target Gaussian loss, the Gaussian loss of the entire decoding process is accurately calculated.

[0151] In some alternative implementation manners of this embodiment, step 407 includes: determining a first weight corresponding to the regression loss and a second weight corresponding to the target Gaussian loss; calculating a comprehensive loss according to the regression loss, the first weight, the target Gaussian loss, and the second weight.

[0152] In this implementation manner, the above-mentioned execution entity respectively sets corresponding weights for L1 and LG, that is, the first weight and the second weight, and then calculates Ltotal according to L1 and the first weight, LG and the second weight.

[0153] For example, when the first weight and the second weight are both set to 0.5, Ltotal can be calculated based on the following formula:

[0154] Ltotal = 0.5 L1 + 0.5 LG

[0155] Of course, the values of the first weight and the second weight can be set according to actual situations, and this embodiment does not make specific limitations in this regard.

[0156] By setting corresponding weights for the regression loss and the Gaussian loss, and calculating the final comprehensive loss through weighted calculation, and training the model with this comprehensive loss, the training efficiency of the trajectory prediction model and the accuracy of the trajectory prediction model are improved.

[0157] Step 408: Train the initial trajectory prediction model using the comprehensive loss to obtain the trajectory prediction model.

[0158] In this embodiment, the above-mentioned execution entity trains the embedding layer, encoder, and decoder in the initial trajectory prediction model according to the comprehensive loss between the target prediction trajectory and the actual driving trajectory of the target vehicle, so as to obtain the trajectory prediction model.

[0159] From Figure 4 it can be seen that compared with the Figure 3 corresponding embodiment, in the training method of the trajectory prediction model in this embodiment, the steps of training the initial trajectory prediction model are as follows: first, generate the regression loss between the predicted trajectory and the actual trajectory, then obtain the overall target Gaussian loss according to the Gaussian loss of each layer of the decoder, and finally generate the total loss using the regression loss and the target Gaussian loss, and adjust the parameters of the embedding layer, encoder, and decoder according to this total loss, so as to obtain the trajectory prediction model, thereby improving the prediction accuracy of the trajectory prediction model.

[0160] Figure 5-1 Fig. shows an application block diagram of the training method of the trajectory prediction model of the present disclosure. It can be seen from this block diagram that it includes the following parts:

[0161] I. First, perform Agent Embedding (vehicle feature encoding) and Map Embedding (map feature encoding). Specifically, take Agent_in as the original data, with a dimension of [N, T, C], and use MLP and MaxPooling for feature extraction and encoding conversion to obtain the vehicle trajectory feature Agent_p. Take Map_in as the original data, with a dimension of [M, N, C], and use MLP and MaxPooling for feature extraction and encoding conversion to obtain the map trajectory feature Map_p.

[0162] II. Use Agent_p and Map_p as input data, and after passing through the encoder Encoder, output G_i. G_i includes the vehicle trajectory encoding Agent Feature and the map element encoding Map Feature. The Encoder includes multiple layers of MHA. Specifically, take Agent_p and Map_p and their corresponding position encodings as the query (abbreviated as q) of MHA, take the drawn lines of the map elements found by the KNN algorithm as the key (abbreviated as k) of MHA, and take the position encoding corresponding to the drawn lines of the map elements as the value (abbreviated as v) of MHA, so as to obtain the output G_i of MHA.

[0163] Then, according to the dense future prediction task and the MLP, an enhanced representation A of the Agent Feature is generated.

[0164] III. Construct the intent query pair, namely the static query vector Q_i and the dynamic query vector Q_s. Specifically, kmeans clustering can be performed according to the end speed and end position of the predicted trajectory, so that through the corresponding position encoding and MLP, Q_i can be obtained. According to the position encoding corresponding to the intent target point and MLP, Q_s can be obtained.

[0165] IV. Output the predicted trajectory through the decoder Decoder. The decoder includes multiple layers of MHA. Each MHA in the decoder includes multiple sub-networks and an output layer. The sub-networks are C_sa, C_A, C_M, and C_j respectively, and the output layer is GMM (Gaussian Mixture Model).

[0166] Specifically, for C_sa, the query is the output C_j-1 of the previous layer decoder and Q_i, the key is C_j-1 and Q_i, and the value is C_j-1. Through this MHA, C_sa can be obtained.

[0167] For C_A, the query is C_sa and Q_s, the key is A and the position encoding corresponding to A, and the value is A. Through this MHA, C_A can be obtained.

[0168] For C_M, the query is C_sa and Q_s, and the key is the output of the dynamic map module (updated according to MapFeature) and the corresponding position encoding, and the value is Through this MHA, C_M can be obtained.

[0169] After that, C_A and C_M go through multiple layers of fully connected MLP to obtain C_j.

[0170] After that, C_j goes through GMM to obtain the probability distribution of each position point in the predicted trajectory, and then by synthesizing the outputs of N decoding layers, the final predicted trajectory can be obtained.

[0171] V. Calculate the L1 loss between this predicted trajectory and the true driving trajectory; calculate the Gaussian loss of each decoding layer to obtain the overall Gaussian loss; obtain the target loss according to the L1 loss and the overall Gaussian loss; finally, use this target loss to adjust the parameters of the initial trajectory prediction model (including Embedding, Encoder, Decoder) to obtain the trajectory prediction model.

[0172] For further reference 5-2,Figure 5-2 The structural schematic of another decoder is shown. This Decoder also includes multiple layers of MHA. Each MHA in this Decoder includes multiple sub-networks and an output layer. The sub-networks are C_sa, C_A_M, and C_j respectively, and the output layer is GMM (Gaussian Mixture Model).

[0173] Specifically, the intention query pair will also be constructed first, that is, the static query vector Q_i and the dynamic query vector Q_s. Specifically, kmeans clustering can be performed according to the end speed and end position of the predicted trajectory, so as to obtain Q_i through the corresponding position encoding and MLP. Q_s is obtained according to the corresponding position encoding and MLP of the intention target point.

[0174] Then, for C_sa, the query is the output C_j-1 of the previous layer decoder and Q_i, the key is C_j-1 and Q_i, and the value is C_j-1. Through this MHA, C_sa can be obtained.

[0175] For C_A_M, the query is C_sa and Q_s, the key is A and the corresponding position encoding of A as well as and the corresponding position encoding, and the value is A and , and through this MHA, C_A_M can be obtained.

[0176] After that, C_A_M can obtain C_j after passing through multiple layers of fully connected MLP.

[0177] Specifically, it can be expressed as:

[0178] MultiHeadAttn(query= + , key= + , value= )

[0179] MultiHeadAttn(query= , , key= , value= )

[0180] =MLP( )

[0181] Among them, the input query vector of C_A_M can be obtained by splicing C_sa and Q_s, and A and The key vector can be obtained by splicing, and A and are spliced to obtain the value vector.

[0182] For further reference to 5-3, Figure 5-3 Figure 5-3 shows the structural schematic of another decoder. This Decoder also includes multiple layers of MHA. Each MHA in this Decoder includes a C_A_M_SA sub-network, a C_j sub-network, and an output layer, and the output layer is a GMM (Gaussian Mixture Model).

[0183] Specifically, the intention query pair will also be constructed first, that is, the static query vector Q_i and the dynamic query vector Q_s. Specifically, kmeans clustering can be performed according to the end speed and end position of the predicted trajectory, so that through the corresponding position encoding and MLP, Q_i can be obtained. According to the position encoding and MLP corresponding to the intention target point, Q_s can be obtained.

[0184] Then, for C_A_M_SA, the query is the output C_j-1 of the previous layer decoder, Q_i and Q_s, the key is C_j-1, Q_i, A and the corresponding position encoding of A, and and the corresponding position encoding, and the value is C_j-1, A and , and through this MHA, C_AM_SA can be obtained.

[0185] After that, C_AM_SA can obtain C_j through multiple layers of fully connected MLP.

[0186] Specifically, it can be expressed as:

[0187] MultiHeadAttn(query= , key= , value= )

[0188] =MLP(

[0189] Among them, splicing C_j-1, Q_i and Q_s can obtain the input query vector of C_AM_SA, splicing C_j-1, Q_i, A and the corresponding position encoding of A can obtain the key vector, and splicing C_j-1, A and to obtain the value vector.

[0190] From Figure 5-2 and Figure 5-3 it can be seen that the sub-networks included in the MHA of these two Decoders are the same asFigure 5-1 The sub-networks included in the MHA in the shown Decoder are different. Therefore, the process of decoding the encoded information is also different, and the accuracy of the obtained decoded information is also different. Specifically, Figure 5-1 the decoding accuracy obtained by the Decoder in Figure 5-2 is greater than that obtained by the Decoder in Figure 5-3 is greater than that obtained by the Decoder in

[0191] Figure 6 Fig. 600 shows the flow of an embodiment of the trajectory prediction method according to the present disclosure. The trajectory prediction method includes the following steps:

[0192] Step 601, obtaining the first current driving trajectory of the target vehicle and the second current driving trajectory of the obstacle vehicle.

[0193] In this embodiment, the execution subject will obtain the current driving trajectory of the target vehicle (i.e., the host vehicle), that is, the first current driving trajectory. This current driving trajectory can be the trajectory within a preset time from the current moment or within a preset distance from the current position. For example, the first current driving trajectory can be the driving trajectory of the host vehicle within one hour before the current moment, and the first current driving trajectory can also be the driving trajectory of the host vehicle within three kilometers from the current distance.

[0194] Correspondingly, the execution subject will also obtain the current driving trajectory of the obstacle vehicle, that is, the second current driving trajectory. The second current driving trajectory can be the driving trajectory of the obstacle within one hour before the current moment, and the second current driving trajectory can also be the driving trajectory of the obstacle within three kilometers from the current distance.

[0195] Step 602, inputting the first current driving trajectory and the second current driving trajectory into the trajectory prediction model, and outputting the predicted trajectory of the target vehicle.

[0196] In this embodiment, the execution subject will input the first current driving trajectory and the second current driving trajectory into the trajectory prediction model, so as to output the predicted trajectory of the target vehicle. Among them, the trajectory prediction model can be trained by the method described in the foregoing embodiment. That is, the trajectory prediction model can predict the future driving trajectory of the host vehicle according to the driving trajectory of the host vehicle (i.e., the first current driving trajectory) and the driving trajectory of the obstacle vehicle (i.e., the second current driving trajectory), and thus output the predicted trajectory of the host vehicle.

[0197] The trajectory prediction method provided by the embodiments of the present disclosure inputs the first current driving trajectory of the target vehicle and the second current driving trajectory of the obstacle vehicle into a trajectory prediction model, so that the trajectory prediction model predicts the trajectory of the host vehicle to obtain a predicted trajectory. Thus, the trajectory information of the obstacle vehicle is integrated when predicting the trajectory, thereby improving the accuracy of trajectory prediction.

[0198] Further referring to Figure 7 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a training device for a trajectory prediction model. This device embodiment corresponds to Figure 1 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0199] As shown in Figure 7 , the training device 700 for the trajectory prediction model in this embodiment includes: an embedding module 701, an encoding module 702, a decoding module 703, and a training module 704. The embedding module 701 is configured to use an embedding layer to extract features from the historical trajectory information of the sample vehicle and the local map information associated with the historical trajectory information to obtain vehicle trajectory features and map features, where the sample vehicle includes the target vehicle and the obstacle vehicle; the encoding module 702 is configured to use an encoder to encode the vehicle trajectory features and map features to obtain vehicle trajectory encodings and map element encodings; the decoding module 703 is configured to use a decoder to decode the vehicle trajectory encodings and map element encodings to obtain the target predicted trajectory of the target vehicle; the training module 704 is configured to train an initial trajectory prediction model according to the loss between the target predicted trajectory and the true driving trajectory of the target vehicle to obtain a trajectory prediction model.

[0200] In this embodiment, for the specific processing of the embedding module 701, encoding module 702, decoding module 703, and training module 704 in the training device 700 for the trajectory prediction model and the technical effects brought by them, reference can be respectively made to Figure 1 the relevant descriptions of steps 101-104 in the corresponding embodiments, which will not be elaborated here.

[0201] In some optional implementation manners of this embodiment, the embedding layer includes a normalization layer and a conversion layer; and the embedding module 701 is further configured to: input the historical trajectory information and the local map information into the normalization layer, and output to obtain a vehicle trajectory matrix and a map element matrix; input the vehicle trajectory matrix and the map element matrix into the conversion layer, and output to obtain vehicle trajectory features and map features.

[0202] In some alternative implementation manners of this embodiment, the encoder includes an encoding network and a regression output layer; and the encoding module 702 is further configured to: input the vehicle trajectory feature and the map feature into the encoding network, and output a vehicle trajectory encoding and a map element encoding; input the vehicle trajectory encoding into the regression output layer, and output an initial predicted trajectory and a predicted speed of the sample vehicle; perform local normalization processing and encoding conversion on the initial predicted trajectory and the predicted speed to obtain a vehicle predicted trajectory encoding; and splice the vehicle predicted trajectory encoding and the vehicle trajectory encoding to obtain a vehicle trajectory encoding.

[0203] In some alternative implementation manners of this embodiment, the decoder includes N decoding layers, and each decoding layer includes a decoding network and a decoding output layer, where N is a positive integer; and the decoding module 703 includes: a first decoding sub-module, configured to, for each decoding layer, input the vehicle trajectory encoding, the map element encoding, a pre-generated global intention vector representation, and an auxiliary vector representation into the decoding network of this decoding layer, and output decoded feature information; input the decoded feature information into the decoding output layer of this decoding layer, and output a probability distribution of each position point in the initial predicted trajectory; and a second decoding sub-module, configured to generate a target predicted trajectory according to the probability distribution of each position point in the initial predicted trajectory output by the N decoding layers.

[0204] In some alternative implementation manners of this embodiment, the training device 700 of the above trajectory prediction model further includes: a clustering module, configured to cluster the endpoints of the predicted trajectories to obtain a plurality of clustering centers; a first representation module, configured to use a fully connected layer to transform the position encodings corresponding to the plurality of clustering centers to obtain a global intention vector representation; and a second representation module, configured to determine a map element line corresponding to the initial predicted trajectory and meeting a preset condition from local map information, and generate an auxiliary vector representation based on the map element line.

[0205] In some alternative implementation manners of this embodiment, the decoding network includes: a first decoding sub-network, a second decoding sub-network, a third decoding sub-network, and a decoding fully connected layer; and the first decoding sub-module is further configured to: input the global intention vector representation into the first decoding sub-network, and output first decoding information; input the first decoding information, the auxiliary vector representation, the vehicle trajectory encoding, and the position encoding corresponding to the vehicle trajectory encoding into the second decoding sub-network, and output second decoding information; input the first decoding information, the auxiliary vector representation, the map element line, and the position encoding corresponding to the map element line into the third decoding sub-network, and output third decoding information; and input the second decoding information and the third decoding information into the decoding fully connected layer, and output decoded feature information.

[0206] In some alternative implementation manners of this embodiment, the training module 704 includes: a regression loss calculation sub-module configured to calculate a regression loss between a target predicted trajectory and an actual driving trajectory; a comprehensive loss calculation sub-module configured to generate a comprehensive loss according to the regression loss and a target Gaussian loss; and a training sub-module configured to train an initial trajectory prediction model by using the comprehensive loss to obtain a trajectory prediction model.

[0207] In some alternative implementation manners of this embodiment, the above-mentioned training device 700 for the trajectory prediction model further includes: a Gaussian loss calculation module configured to calculate, for each decoding layer, a Gaussian regression loss of this decoding layer; and add the Gaussian regression losses of N decoding layers to obtain a target Gaussian loss.

[0208] In some alternative implementation manners of this embodiment, the comprehensive loss calculation sub-module is further configured to: determine a first weight corresponding to the regression loss and a second weight corresponding to the target Gaussian loss; and calculate the comprehensive loss according to the regression loss, the first weight, the target Gaussian loss, and the second weight.

[0209] Further referring to Figure 8 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a trajectory prediction device. This device embodiment corresponds to Figure 6 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0210] As Figure 8 shown, the trajectory prediction device 800 of this embodiment includes: an acquisition module 801 and a prediction module 802. The acquisition module 801 is configured to acquire a first current driving trajectory of a target vehicle and a second current driving trajectory of an obstacle vehicle; the prediction module 802 is configured to input the first current driving trajectory and the second current driving trajectory into the trajectory prediction model, and output a predicted trajectory of the target vehicle. The trajectory prediction model is trained by using the method described in the foregoing embodiment.

[0211] In this embodiment, for the specific processing of the acquisition module 801 and the prediction module 802 in the trajectory prediction device 800 and the technical effects brought thereby, reference can be respectively made to Figure 6 the relevant descriptions of steps 601-602 in the corresponding embodiment, which will not be elaborated herein.

[0212] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, a computer program product, and an autonomous vehicle.

[0213] Figure 9FIG. 0 shows a schematic block diagram of an exemplary electronic device 900 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers and other suitable computing devices. The electronic device can also represent various forms of mobile devices suitable for performing calculations. In particular, the electronic device can also be an electronic device integrated in a vehicle, such as a car computer.

[0214] The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0215] As Figure 9 shown, the electronic device 900 includes a computing unit 901 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0216] A plurality of components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disc, etc. It should be understood that the input and output units listed here are only examples existing in some special scenarios. In some other application scenarios, the input and output units can be in other forms. For example, when the device is a car computer device, the input unit can be a touch screen, a rotary button, etc.

[0217] A communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0218] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the training method or the trajectory prediction method of the trajectory prediction model. For example, in some embodiments, the training method or the trajectory prediction method of the trajectory prediction model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the training method or the trajectory prediction method of the trajectory prediction model described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the training method or the trajectory prediction method of the trajectory prediction model in any other suitable manner (e.g., by means of firmware).

[0219] The autonomous vehicle provided by the present disclosure may include the above-mentioned electronic device as shown in Figure 9 The above, and when its processor executes, it can implement the training method or the trajectory prediction method of the trajectory prediction model described in any of the above embodiments.

[0220] The various embodiments of the systems and technologies described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, the one or more computer programs being executable and / or interpretable on a programmable system including at least one programmable processor, the programmable processor being a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0221] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program codes may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0222] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0223] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, speech input, or tactile input).

[0224] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0225] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating blockchain.

[0226] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0227] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A training method for a trajectory prediction model, wherein, The initial trajectory prediction model includes an embedding layer, an encoder, and a decoder. The method includes: Using the embedding layer to extract features from the historical trajectory information of the sample vehicle and the local map information associated with the historical trajectory information, to obtain vehicle trajectory features and map features, where the sample vehicle includes a target vehicle and an obstacle vehicle; Using the encoder to encode the vehicle trajectory features and the map features, to obtain a vehicle trajectory encoding and a map element encoding; Using the decoder to decode the vehicle trajectory encoding and the map element encoding, to obtain the target prediction trajectory of the target vehicle, where the decoder includes N decoding layers, and each decoding layer includes a decoding network and a decoding output layer, and N is a positive integer; Training the initial trajectory prediction model according to the loss between the target prediction trajectory and the true driving trajectory of the target vehicle, to obtain a trajectory prediction model; Wherein, the step of using the decoder to decode the vehicle trajectory encoding and the map element encoding to obtain the target prediction trajectory of the target vehicle includes: For each decoding layer, inputting the vehicle trajectory encoding, the map element encoding, a pre-generated global intention vector representation, and an auxiliary vector representation into the decoding network of this decoding layer, and outputting decoded feature information; inputting the decoded feature information into the decoding output layer of this decoding layer, and outputting the probability distribution of each position point in the initial prediction trajectory of the sample vehicle; wherein, the global intention vector representation is obtained by using a fully connected layer to transform the position encodings corresponding to multiple clustering centers, and the multiple clustering centers are obtained by clustering the end points of the prediction trajectory, and the auxiliary vector representation is generated by map element lines in the local map information corresponding to the initial prediction trajectory and meeting a preset condition; Generating the target prediction trajectory according to the probability distribution of each position point in the initial prediction trajectory output by the N decoding layers.

2. The method according to claim 1, wherein, The embedding layer includes a normalization layer and a transformation layer; and The step of using the embedding layer to extract features from the historical trajectory information of the sample vehicle and the local map information associated with the historical trajectory information to obtain vehicle trajectory features and map features includes: Inputting the historical trajectory information and the local map information into the normalization layer, and outputting a vehicle trajectory matrix and a map element matrix; Inputting the vehicle trajectory matrix and the map element matrix into the transformation layer, and outputting the vehicle trajectory features and the map features.

3. The method according to claim 1, wherein, The encoder includes an encoding network and a regression output layer; And The step of using the encoder to encode the vehicle trajectory features and the map features to obtain a vehicle trajectory encoding and a map element encoding includes: Inputting the vehicle trajectory features and the map features into the encoding network, and outputting a vehicle trajectory encoding and the map element encoding; Inputting the vehicle trajectory encoding into the regression output layer, and outputting the initial prediction trajectory and the predicted speed of the sample vehicle; Perform local normalization processing and encoding conversion on the initial predicted trajectory and predicted speed to obtain a vehicle predicted trajectory encoding; Concatenate the vehicle predicted trajectory encoding and the vehicle trajectory encoding to obtain the vehicle trajectory encoding.

4. The method according to claim 1, wherein The method further includes: Cluster the end points of the predicted trajectory to obtain multiple cluster centers; Use a fully connected layer to transform the position encodings corresponding to the multiple cluster centers to obtain the global intention vector representation; Determine a map element line corresponding to the initial predicted trajectory and meeting a preset condition from the local map information, and generate the auxiliary vector representation based on the map element line.

5. The method according to claim 1, wherein The decoding network includes: a first decoding sub-network, a second decoding sub-network, a third decoding sub-network, and a decoding fully connected layer; and Input the vehicle trajectory encoding, the map element encoding, and the pre-generated global intention vector representation and auxiliary vector representation into the decoding network of this decoding layer, and output the decoded feature information, including: Input the global intention vector representation into the first decoding sub-network, and output the first decoded information; Input the first decoded information, the auxiliary vector representation, the vehicle trajectory encoding, and the position encoding corresponding to the vehicle trajectory encoding into the second decoding sub-network, and output the second decoded information; Input the first decoded information, the auxiliary vector representation, the map element line, and the position encoding corresponding to the map element line into the third decoding sub-network, and output the third decoded information; Input the second decoded information and the third decoded information into the decoding fully connected layer, and output the decoded feature information.

6. The method according to claim 1, wherein, Training the initial trajectory prediction model according to the loss between the target predicted trajectory and the true driving trajectory of the target vehicle to obtain a trajectory prediction model, including: Calculate the regression loss between the target predicted trajectory and the true driving trajectory; Generate a comprehensive loss according to the regression loss and the target Gaussian loss; Use the comprehensive loss to train the initial trajectory prediction model to obtain the trajectory prediction model.

7. The method according to claim 6, wherein The method further includes: For each decoding layer, calculate the Gaussian regression loss of this decoding layer; Add the Gaussian regression losses of the N decoding layers to obtain the target Gaussian loss.

8. The method according to claim 6, wherein The generating a comprehensive loss according to the regression loss and the target Gaussian loss includes: Determine a first weight corresponding to the regression loss and a second weight corresponding to the target Gaussian loss; Calculate the comprehensive loss according to the regression loss, the first weight, the target Gaussian loss, and the second weight.

9. A trajectory prediction method, including: Obtain the first current driving trajectory of the target vehicle and the second current driving trajectory of the obstacle vehicle; Input the first current driving trajectory and the second current driving trajectory into a trajectory prediction model, and output the predicted trajectory of the target vehicle, where the trajectory prediction model is trained by using the method according to any one of claims 1-8.

10. A training device for a trajectory prediction model, wherein, The initial trajectory prediction model includes an embedding layer, an encoder, and a decoder. The apparatus includes: An embedding module, configured to use the embedding layer to extract features from the historical trajectory information of a sample vehicle and local map information associated with the historical trajectory information, to obtain vehicle trajectory features and map features, where the sample vehicle includes a target vehicle and an obstacle vehicle; An encoding module, configured to use the encoder to encode the vehicle trajectory features and the map features, to obtain a vehicle trajectory encoding and a map element encoding; A decoding module, configured to use the decoder to decode the vehicle trajectory encoding and the map element encoding, to obtain a target prediction trajectory of the target vehicle, where the decoder includes N decoding layers, and each decoding layer includes a decoding network and a decoding output layer, and N is a positive integer; A training module, configured to train the initial trajectory prediction model according to a loss between the target prediction trajectory and the true driving trajectory of the target vehicle, to obtain a trajectory prediction model; Wherein, the decoding module is further configured to: For each decoding layer, input the vehicle trajectory encoding, the map element encoding, a pre-generated global intention vector representation, and an auxiliary vector representation into the decoding network of this decoding layer, and output decoded feature information; input the decoded feature information into the decoding output layer of this decoding layer, and output a probability distribution of each position point in the initial prediction trajectory of the sample vehicle; where the global intention vector representation is obtained by using a fully connected layer to transform position encodings corresponding to multiple clustering centers, the multiple clustering centers are obtained by clustering the end points of the prediction trajectory, and the auxiliary vector representation is generated by map element lines in the local map information corresponding to the initial prediction trajectory and meeting a preset condition; Generate the target prediction trajectory according to the probability distribution of each position point in the initial prediction trajectory output by the N decoding layers.

11. A trajectory prediction apparatus, including: An acquisition module, configured to acquire a first current driving trajectory of a target vehicle and a second current driving trajectory of an obstacle vehicle; A prediction module, configured to input the first current driving trajectory and the second current driving trajectory into a trajectory prediction model, and output a prediction trajectory of the target vehicle, where the trajectory prediction model is trained by using the method according to any one of claims 1-8.

12. An electronic device, including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the method according to any one of claims 1 to 9.

13. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.

14. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 9.

15. An autonomous vehicle comprising the electronic device according to claim 12.

Citation Information

Patent Citations

  • Vehicle track prediction model training method and device, equipment and storage medium

    CN116401549A

  • Obstacle trajectory prediction method and device and storage medium

    CN118585792A