Agent Trajectory Prediction Method Based on Heterogeneous Data Association Mining and Metric Learning
By constructing the loss function of agent trajectory prediction network and measurement learning of heterogeneous data correlation mining, the problem of heterogeneous data correlation is ignored and supervised training poor results in the prior art, and more accurate agent trajectory prediction is achieved.
Patent Information
- Application Number
- CN202211163069.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-23
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-09-23
AI Technical Summary
In the existing agent trajectory prediction methods, heterogeneous data correlation is ignored and supervised training effects are poor, resulting in a decrease in prediction accuracy.
Agent trajectory prediction network for heterogeneous data correlation mining is constructed, and the correlation between the agent's historical trajectory and the active area is mined through a high-definition map encoding subnet, and a loss function based on metric learning is designed to calculate the similarity between future trajectory prediction results and labels in the measurement space.
It improves the accuracy and robustness of trajectory prediction, avoids noise interference, and the prediction results are closer to the real trajectory of the agent.
Smart Images

Figure CN115688019B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of physical technologies, and further relates to an intelligent agent trajectory prediction method based on heterogeneous data association mining and metric learning in the field of digital processing of electrical data. The present invention can be used to predict the movement trajectories of intelligent agents in the fields of autonomous driving, human-computer interaction, and traffic control. Background Art
[0002] The intelligent agent trajectory prediction task is to predict the future movement trajectory of an intelligent agent based on its historical movement trajectory. Since there are a large number of interaction behaviors between the intelligent agent and its movement environment, such as vehicles and pedestrians need to move on the roadway and sidewalk respectively and cannot exceed the regional boundary, using only the historical movement trajectory of the intelligent agent cannot obtain an accurate future trajectory prediction result. People use a high-definition map of the intelligent agent movement area for auxiliary prediction during trajectory prediction. With the rapid development of deep learning in the fields of motion modeling, computer vision, and heterogeneous data processing, applying neural networks to the field of intelligent agent trajectory prediction for supervised learning has also become the latest trend.
[0003] Salzmann, Ivanovic and others disclosed an agent trajectory prediction method based on heterogeneous data fusion in their published paper "Trajectron++: Dynamically-Feasible Trajectory Forecasting With Heterogeneous Data" (Conference Name: European Conference on Computer Vision’2020). This method first reads in the historical trajectory data of the agent and the high-definition map of the agent's action area; then uses Long Short-Term Memory networks (LSTM) to encode the historical trajectory data in a graph structure, and at the same time uses a Convolutional Neural Network (CNN) to extract the image features of the high-definition map; then inputs the graph structure features into a Conditional Variational Auto Encoder (CVAE) to obtain the high-dimensional latent variables of the graph structure features; finally, inputs the high-dimensional latent variables and the high-definition map image features into a Gated Recurrent Unit (GRU) to decode and predict the future trajectory of the agent. There are two deficiencies in this method. First, the historical trajectory of the agent and the high-definition map are encoded independently. Such an encoding method cannot focus on the activity area that the agent is about to enter, and the areas that the agent has passed through and will not enter cannot help with trajectory prediction and may even introduce noise. Second, when this method conducts supervised training on all the networks included in the trajectory prediction, it does not make full use of the supervised information of the labels, which will reduce the prediction accuracy of the trained network. Summary of the Invention
[0004] The object of the present invention is to address the deficiencies of the above-mentioned existing technologies and propose an agent trajectory prediction method based on heterogeneous data association mining and metric learning to solve the problems of the neglect of the relevance of heterogeneous data in the trajectory prediction process and the poor effect in the supervised training stage.
[0005] The idea of achieving the object of the present invention is as follows: To solve the problem that the relevance of heterogeneous data is ignored in the process of trajectory prediction, the present invention constructs an intelligent agent trajectory prediction network for mining heterogeneous data association, which can calculate the map features of the activity area that the intelligent agent is about to enter based on the historical trajectory of the intelligent agent, so as to avoid the noise brought by the map features of the areas that the intelligent agent has passed through and the areas that it will not enter during trajectory prediction. To solve the problem of poor supervised training effect of the network in the process of trajectory prediction, the present invention designs a loss function based on metric learning, which can calculate the similarity between the future trajectory prediction result and the future trajectory label in the metric space, so as to make full use of the supervision information of the future trajectory label during supervised training and make the trajectory predicted by the trained network closer to the real trajectory of the intelligent agent.
[0006] The implementation steps of the present invention are as follows:
[0007] Step 1, construct an intelligent agent trajectory prediction network for mining heterogeneous data association:
[0008] Respectively construct a historical trajectory graph structure encoding sub-network, a high-definition map encoding sub-network, a high-dimensional distribution encoding sub-network and a feature decoding sub-network; form a parallel network by combining the historical trajectory graph structure encoding sub-network and the high-definition map encoding sub-network, and connect the parallel network in series with the high-dimensional distribution encoding sub-network and the feature decoding sub-network in sequence to form an intelligent agent trajectory prediction network for mining heterogeneous data association;
[0009] The high-definition map encoding sub-network is composed of three modules connected in series, a position matrix initialization module without parameters for calculating the position matrix, a Transformer network module for extracting map features according to the position matrix, and a first fully connected layer for calculating the position offset. Set the input dimension and output dimension of the Transformer network module to [[100×100×dis(M)],[32×2]] and 32 respectively, where dis(M) is the dimension of the high-definition map, and the input dimension and output dimension of the first fully connected layer are 1 and 2 respectively;
[0010] Step 2, generate a training set:
[0011] Select the action trajectory data of at least 7500 intelligent agents and the high-definition map data of their action areas;
[0012] The action trajectory data of the intelligent agent contains at least 2 historical trajectories and 1 future trajectory;
[0013] Set the historical trajectory of the intelligent agent and the high-definition map data of its corresponding action area as the network input data, and set the future trajectory of the intelligent agent as the label;
[0014] Combine the network input data with its corresponding label to form a training set;
[0015] Step 3, obtaining the predicted trajectory of the agent using the training set:
[0016] Input the training set into the agent trajectory prediction network for heterogeneous data association mining. The historical trajectory graph structure encoding sub-network calculates the graph structure encoding of the agent's historical trajectory. The high-definition map encoding sub-network mines the correlation between the agent's historical trajectory and its activity area and calculates the map features. The high-dimensional distribution encoding sub-network calculates the high-dimensional distribution of the historical trajectory graph structure encoding and the activity area map features. The feature decoding sub-network decodes the high-dimensional distribution, graph structure features, and map features to obtain l possible predicted trajectories of the agent.
[0017] The specific steps for the high-definition map encoding sub-network to mine the correlation between the agent's historical trajectory and its activity area and calculate the map features are as follows:
[0018] Calculate the action angle of the agent according to the following formula:
[0019]
[0020] where absc end , ordi end and absc end-1 , ordi end-1 respectively represent the abscissa and ordinate of the agent's last two historical trajectories in the activity map;
[0021] Determine the action area that the agent will enter according to its action angle and generate a position matrix in this area. Update the position matrix and use the features at the positions in the position matrix in the activity map as the map features;
[0022] Step 4, calculate the loss value between the future trajectory prediction result and the label according to the following loss function based on metric learning;
[0023]
[0024] where, is the loss value between the predicted result y of the future trajectory and the label of the future trajectory, log(·) is the logarithm operation with base 10, α is the weight parameter of the metric space distance value y between the label feature F of the future trajectory and the predicted feature of the future trajectory. The metric space uses the cosine similarity space, β is the weight parameter of the relative entropy value D t (z, z KL (z, z t ), z is the high-dimensional distribution of the agent's comprehensive feature F A , z tFor the comprehensive feature F of the agent A And the signature trajectory feature F y The high-dimensional distribution after splicing;
[0025] Step 5, optimize the agent trajectory prediction network for heterogeneous data association mining:
[0026] Use the Adam optimization algorithm to iteratively update the parameters of the agent trajectory prediction network for heterogeneous data association mining, and repeat Step 3 and Step 4 until the loss value between the future trajectory prediction result and the future trajectory label Converges to obtain the trained network;
[0027] Step 6, use the trained network for prediction:
[0028] Input the historical trajectory information of the agent to be predicted and the high-definition map of its activity area into the trained network, and output the future trajectory prediction result of the agent.
[0029] The present invention has the following advantages compared with the prior art:
[0030] First, by constructing an agent trajectory prediction network for heterogeneous data association mining, the present invention calculates the map features of the activity area that the agent will enter according to the agent's historical trajectory, and further predicts the agent's future trajectory, thus overcoming the problem that the map features of the areas that the agent has passed through and the areas that it will not enter bring noise caused by independent coding of heterogeneous data in the prior art, making the network constructed by the present invention have the advantages of strong heterogeneous data association mining ability and the predicted agent trajectory being closer to its true trajectory.
[0031] Second, by designing a loss function based on metric learning and making full use of the supervision information of the label by calculating the similarity between the future trajectory prediction result and the future trajectory label in the metric space, the present invention overcomes the problem of poor supervised training effect in the trajectory prediction process in the prior art, making the trained trajectory prediction network of the present invention have the advantages of high accuracy and strong robustness. Description of the Drawings
[0032] Figure 1 Is the flow chart of the present invention;
[0033] Figure 2 Is the explanatory diagram of the working mechanism of the initialization module of the high-definition map coding sub-network matrix of the present invention;
[0034] Figure 3 Is the result comparison diagram of the present invention and the advanced method Trajectron++ for predicting the future trajectory of agents in different scenarios. Detailed Embodiment
[0035] The technical solutions and effects of the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments.
[0036] Refer to Figure 1 and the embodiments to further describe the specific implementation steps of the present invention in detail.
[0037] Step 1: Construct an intelligent agent trajectory prediction network for heterogeneous data association mining.
[0038] Construct a historical trajectory graph structure encoding sub-network, a high-definition map encoding sub-network, a high-dimensional distribution encoding sub-network, and a feature decoding sub-network respectively; form a parallel network by combining the historical trajectory graph structure encoding sub-network and the high-definition map encoding sub-network, and connect the parallel network in series with the high-dimensional distribution encoding sub-network and the feature decoding sub-network in sequence to form an intelligent agent trajectory prediction network for heterogeneous data association mining.
[0039] The historical trajectory graph structure encoding sub-network consists of two parallel modules, the first long short-term memory network module and the second long short-term memory network module, which are respectively used for undirected graph node encoding and undirected graph edge encoding. Set the input dimension and hidden layer dimension in both the first long short-term memory network module and the second long short-term memory network module to 2 and 32.
[0040] The high-definition map encoding sub-network consists of three modules connected in series, a parameter-free position matrix initialization module for calculating the position matrix, a Transformer network module for extracting map features based on the position matrix, and a first fully connected layer for calculating the position offset. Set the input dimension and output dimension of the Transformer network module to [[100×100×dis(M)],[32×2]] and 32 respectively. In this embodiment, dis(M) is 3, so the input dimension and output dimension of the Transformer network module are [[100×100×3],[32×2]] and 32 respectively. The input dimension and output dimension of the first fully connected layer are 1 and 2 respectively.
[0041] The high-dimensional distribution encoding sub-network consists of three parallel modules, the first fully connected layer, the second fully connected layer, and the third fully connected layer, which are used to generate the high-dimensional distribution of the historical trajectory graph structure features and map features. The input dimension and output dimension of the first fully connected layer are 96 and 25 respectively. The input dimension and output dimension of the second fully connected layer are 2 and 1 respectively. The input dimension and output dimension of the third fully connected layer are 96+len(fut) and 25 respectively, where len(fut) is the length of the agent's future trajectory. In this embodiment, len(fut) is 8, so the input dimension and output dimension of the fourth fully connected layer are 104 and 25 respectively.
[0042] The feature decoding sub-network consists of three first fully connected layers, a first long short-term memory network module, and a second fully connected layer in series, and is used to decode the high-dimensional distribution and generate the predicted trajectory of the agent. The input dimension and output dimension of the first fully connected layer are set to 1 and l respectively, where l is the number of predicted future trajectories. In this embodiment, l is 200. Therefore, the input dimension and output dimension of the fifth fully connected layer are set to 1 and 200 respectively. The input dimension and hidden dimension of the first long short-term memory network module are set to 121 and 128 respectively, and the input dimension and output dimension of the second fully connected layer are set to 128 and 2 respectively.
[0043] The structure of each long short-term memory network module in the historical trajectory graph structure encoding sub-network and the feature decoding sub-network is the same. Each long short-term memory network module includes an input gate unit, a state update unit, a forgetting gate unit, and an output gate unit.
[0044] Each unit calculates the output matrix of each unit through four parameter matrices according to the following formula:
[0045] Output unit =W i ×input+b i +W h ×hidden+b h ,
[0046] where Output unit is the output matrix of the unit-th unit, input and hidden are the input state matrix and the hidden state matrix respectively, W i , W h , b i and b h are the input weight parameter matrix, the hidden weight parameter matrix, the input bias parameter matrix, and the hidden bias parameter matrix respectively.
[0047] Since W i and b i perform matrix calculations with input, and W h and b h perform matrix calculations with hidden, so the parameter matrices W i , W h , b i and b h of the long short-term memory network module have dimensions of [4×dim(hidden),dim(input)], [4×dim(hidden),dim(hidden)], [4×dim(hidden),1], and [4×dim(hidden),1] respectively, where dim represents the dimension of the matrix.
[0048] In the fully connected layer modules of the historical trajectory map structure encoding sub-network, the high-dimensional distribution encoding sub-network, and the high-feature decoding sub-network, two parameter matrices are used to calculate the output matrix of the module according to the formula Output = W × input + b. Here, input is the input state matrix, and W and b are the input weight parameter matrix and the input bias parameter matrix respectively. Therefore, the dimensions of W and b in the fully connected layer module are [dim(input), dim(output)] and [dim(input), 1] respectively.
[0049] Step 2: Generate a training set.
[0050] In the embodiment of the present invention, the action trajectory data of 7500 agents and the high-definition map data of their action areas are extracted from the publicly available autonomous driving dataset nuScenes. The types of agents include pedestrians and vehicles.
[0051] The sampling duration of the agent action trajectory data is 8 seconds, and the sampling frequency is 2 Hz. The data in the first 4 seconds and the last 4 seconds of the 8-second trajectory data of the agent are respectively set as the historical trajectory and the future trajectory.
[0052] The high-definition map data is a binary image marked with the vehicle drivable area, the road boundary of the vehicle drivable area, and the pedestrian drivable area. The pixel value of the drivable position in the vehicle drivable area image and the pedestrian drivable area image is set to 1, and the other positions are set to 0. The pixel value of the road boundary position in the vehicle drivable area road boundary image is set to 1, and the other positions are set to 0. The area centered at the position of each agent at the 4th second, with a length and width of 100 respectively, in the high-definition map data is set as the action area of each agent. The vehicle drivable area image, the vehicle drivable area road boundary image, and the pedestrian drivable area image of each agent's action area are stitched together to obtain the high-definition map data of each agent's action area, and the dimension of the map data is [100×100×3].
[0053] The historical trajectory of the agent and the high-definition map data of its corresponding action area are set as the network input data, and the future trajectory of the agent is set as the label.
[0054] The network input data and its corresponding label are combined to form a training set.
[0055] Step 3: Obtain the predicted trajectory of the agent using the training set.
[0056] The training set is input into the agent trajectory prediction network of heterogeneous data association mining. The historical trajectory graph structure encoding subnetwork calculates the graph structure encoding of the agent's historical trajectory. The high-definition map encoding subnetwork mines the correlation between the agent's historical trajectory and its activity area and calculates the map features. The high-dimensional distribution encoding subnetwork calculates the high-dimensional distribution of the historical trajectory graph structure encoding and the activity area map features. The feature decoding subnetwork decodes the high-dimensional distribution, graph structure features and map features to obtain 200 possible predicted trajectories of the agent.
[0057] Agent A i Step 3 is further explained by taking the future trajectory prediction process of as an example.
[0058] Agent A i The historical trajectory is set to an input state matrix of dimension [8×2], the zero matrix is set to a hidden state matrix, and then the input state matrix and the hidden state matrix are input into the first long short-term memory network module of the historical trajectory graph structure encoding subnetwork. The network module outputs the historical trajectory features of dimension [1×32] as the undirected graph node feature F node When two agents A i and A j In A i Map of the action area When the distance between them is less than 10 pixels, it is determined that there is interaction between them. i Agent A with interactive behavior j The historical trajectory and A i After adding the historical trajectories, the input state matrix with a dimension of [8×2] is set, and the zero matrix is set as the hidden state matrix. Then the input state matrix and the hidden state matrix are input into the second long short-term memory network module of the historical trajectory graph structure encoding subnetwork. The network module outputs the historical trajectory features with a dimension of [1×32] as the undirected graph edge feature F edge . node With F edge The concatenation results in the historical trajectory graph structure encoding feature F with a dimension of [1×64] g .
[0059] Agent A i The historical positions at 3.5 seconds and 4 seconds are input into the position matrix initialization module, and the agent A is calculated according to the following formula i The action angle,
[0060]
[0061] Among them, ordi and absc represent agent A, i exist The horizontal and vertical coordinates in the figure. Determine A according to the action anglei The action area R to be entered ent , and initialize 32 positions in R ent to form an initial position matrix B with a dimension of [32×2] according to a Gaussian distribution.
[0062] Refer to the appendix Figure 2 for a further description of the working mechanism of the matrix initialization module.
[0063] Figure 2 The red arrow in represents the action angle of agent A i at the 4th second, and the yellow area represents the area R i that A is about to enter ent , and the yellow circles represent the position points in the initial position matrix B.
[0064] Take the positions in the position matrix B after 4 updates as the positions in R ent that are most helpful for trajectory prediction. Each time the position matrix B is updated, and B are input into the high-definition map encoding sub-network Transformer network module, and this network module outputs the features at the positions in B in, obtaining an updated feature F with a dimension of [32×1] B . Then input F B into the first fully connected layer of the high-definition map encoding sub-network. This fully connected layer outputs an offset vector offset of B, and the dimension of offset is [32×2]. Finally, add offset to B to update the positions in B. Repeat the above steps to update the positions in B 4 times to obtain the final position matrix B final , and the final map features F final at the positions in B in M , F M has a dimension of [32×1].
[0065] Transpose F M and splice it with F g to obtain an integrated feature F of the agent with a dimension of [1×96] A . Input F A into the first fully connected layer of the high-dimensional distribution encoding sub-network. This fully connected layer outputs a high-dimensional distribution z, and the dimension of z is [1×25].
[0066] Splice z with F A and transpose it, then input it into the first fully connected layer of the feature decoding sub-network to obtain a feature F with a dimension of [200×121] Az . Input F AzSet it as the input state matrix, set the zero matrix as the hidden state matrix, and then input the input state matrix and the hidden state matrix into the first long short-term memory network module of the feature decoding sub-network. This network module outputs the decoded features Then Input it into the second fully connected layer of the feature decoding sub-network. This fully connected layer outputs 200 possible future trajectories at 4.5 seconds. After obtaining 200 possible future trajectories at 4.5 seconds, set F Az as the input state matrix, set it as the hidden state matrix, and then input the input state matrix and the hidden state matrix into the first long short-term memory network module of the feature decoding sub-network. This network module outputs the decoded features Then Input it into the second fully connected layer of the feature decoding sub-network. This fully connected layer outputs 200 possible future trajectories at 5 seconds. Repeat the above steps until the prediction results of 200 possible future trajectories from 4.5 seconds to 8 seconds are obtained
[0067] Step 4: Calculate the loss value between the future trajectory prediction result and the label according to the following loss function based on metric learning
[0068]
[0069] Among them, is the loss value between the prediction result y of the future trajectory and the label of the future trajectory , log(·) is the logarithm operation with base 10, α is the metric space distance value between the label feature F y of the future trajectory and the predicted feature of the future trajectory is the weight parameter of the metric space. The metric space uses the cosine similarity space, β is the relative entropy value D t between z and z KL (z, z t ), z is the high-dimensional distribution of the agent comprehensive feature F A , z t is the high-dimensional distribution after splicing the agent comprehensive feature F A and the signature trajectory feature F y
[0070] When calculating the above-mentioned F y and , input y with dimension [8×2] and into the second fully connected layer of the high-dimensional distribution encoding sub-network respectively. This fully connected layer outputs F with dimension [8×1] y and
[0071] When calculating the said z t , transpose F y and splice it with F A to obtain the training feature F with a dimension of [1×104] ay . Then, input F ay into the third fully connected layer of the high-dimensional distribution encoding sub-network. The output dimension of this fully connected layer is the high-dimensional distribution z with a dimension of [1×25] t .
[0072] Step 5, optimize the agent trajectory prediction network for heterogeneous data association mining.
[0073] Use the Adam optimization algorithm to iteratively update the parameters of the agent trajectory prediction network for heterogeneous data association mining. Repeat Step 3 and Step 4 until the loss value between the future trajectory prediction result and the future trajectory label converges, and obtain the trained network.
[0074] Step 6, use the trained network for prediction.
[0075] Input the historical trajectory information of the agent to be predicted and the high-definition map of its activity area into the trained network, and output the future trajectory prediction result of the agent.
[0076] The following further illustrates the effect of the present invention in combination with simulation experiments:
[0077] 1. Simulation experiment conditions:
[0078] The software platform for the simulation experiment of the present invention is as follows: The simulation of the present invention is carried out on the Ubuntu20.04 system with an Intel(R) Core(TM) i9-10900x CPU, an Nvidia 1070 graphics card, and 128G of memory, using Python3.7 software and the Pytorch1.5 deep learning toolkit.
[0079] The agents to be predicted used in the simulation experiment of the present invention are 1500 agents extracted from the publicly available autonomous driving dataset nuScenes. The data of the agents includes their action trajectory data and the high-definition map data of their action areas.
[0080] 2. Simulation content and its result analysis:
[0081] The simulation experiment of the present invention uses the present invention and an existing technology (the Trajectron++ method based on heterogeneous data fusion) to predict the future trajectories of the agents to be predicted respectively, and each method generates 200 possible trajectories.
[0082] In the simulation experiment, an existing technology used refers to:
[0083] The existing Trajectron++ method for trajectory prediction based on heterogeneous data fusion refers to the agent trajectory prediction method based on heterogeneous data fusion proposed by Salzmann, Ivanovic, etc. in their published paper "Trajectron++: Dynamically-Feasible Trajectory Forecasting With Heterogeneous Data" (Conference Name: European Conference on Computer Vision’2020).
[0084] The following combines Figure 3 simulation diagrams to further describe the effects of the present invention.
[0085] Figure 3 The figures showing the results of future trajectory prediction of agents by using the method of the present invention and the Trajectron++ method for trajectory prediction based on heterogeneous data fusion, where Figure 3 (a) represents the intersection scenario diagram, Figure 3 (a) represents the roundabout scenario diagram, Figure 3 (a) represents the T-junction scenario diagram and Figure 3 (a) represents the straight-ahead scenario diagram. Each scenario has the result diagrams corresponding to the two methods. Figure 3 The trajectories composed of white hollow circles represent the labels of the future trajectories of the agents, and the black solid circles represent the centers of the 200 generated trajectory points. As Figure 3 can be seen, the black solid circles of the method of the present invention are closer to the white hollow circles in the above four scenarios. Therefore, compared with the Trajectron++ method for trajectory prediction based on heterogeneous data fusion, the predicted trajectories of the method of the present invention are closer to the real trajectories of the agents in the four scenarios.
[0086] To verify the simulation results of the present invention, two evaluation metrics (the minimum average error of 200 trajectories, 200-ADE, and the minimum final error of 200 trajectories, 200-FDE) are used to evaluate the predicted trajectories of the two methods for the next 4 seconds. Using the following formulas, 200-ADE and 200-FDE, all the calculation results are plotted in Table 1:
[0087]
[0088] 200-FDE = min (the final prediction errors of 200 possible trajectories)
[0089] Table 1. Quantitative analysis table of the prediction results of the present invention and the advanced technology in the simulation experiment
[0090]
[0091] As can be seen from Table 1, the prediction accuracy of the present invention is improved compared with the advanced method Trajectron++ in all evaluation indexes at all prediction times, which proves that the present invention can obtain a prediction result closer to the real trajectory of the agent.
[0092] The above simulation experiments show that: the agent trajectory prediction method based on heterogeneous data association mining and metric learning of the present invention mines the correlation between the historical trajectory of the agent and the map information of its activity area by constructing an agent trajectory prediction network for heterogeneous data association mining, avoids the noise brought by the map information of the area that the agent has passed through and the area that the agent will not enter, and designs a loss function based on metric learning to train the network, improving the prediction accuracy and robustness of the network, and making the prediction result of the network closer to the real trajectory of the agent.
Claims
1. An intelligent agent trajectory prediction method based on heterogeneous data association mining and metric learning, characterized in that Construct an agent trajectory prediction network for heterogeneous data association mining, obtain the predicted trajectory of the agent using the training set, and train the network by designing a loss function based on metric learning. The specific steps of this method are as follows: Step 1, construct an agent trajectory prediction network for heterogeneous data association mining: Construct a historical trajectory graph structure encoding sub-network, a high-definition map encoding sub-network, a high-dimensional distribution encoding sub-network, and a feature decoding sub-network respectively; form a parallel network by combining the historical trajectory graph structure encoding sub-network and the high-definition map encoding sub-network, and connect the parallel network in series with the high-dimensional distribution encoding sub-network and the feature decoding sub-network in sequence to form an agent trajectory prediction network for heterogeneous data association mining; The high-definition map encoding sub-network consists of three modules connected in series, a parameter-free position matrix initialization module for calculating the position matrix, a Transformer network module for extracting map features based on the position matrix, and a first fully connected layer for calculating the position offset. Set the input dimension and output dimension of the Transformer network module to [[100×100×dis(M)],[32×2]] and 32 respectively, where dis(M) is the dimension of the high-definition map, and the input dimension and output dimension of the first fully connected layer are 1 and 2 respectively; Step 2, generate a training set: Select the action trajectory data of at least 7,500 agents and the high-definition map data of their action areas; The agent action trajectory data contains at least 2 historical trajectories and 1 future trajectory; Set the agent's historical trajectory and the high-definition map data of its corresponding action area as the network input data, and set the agent's future trajectory as the label; Combine the network input data with its corresponding label to form a training set; Step 3, obtain the predicted trajectory of the agent using the training set: Input the training set into the agent trajectory prediction network for heterogeneous data association mining. The historical trajectory graph structure encoding sub-network calculates the graph structure encoding of the agent's historical trajectory. The high-definition map encoding sub-network mines the relevance between the agent's historical trajectory and its activity area and calculates the map features. The high-dimensional distribution encoding sub-network calculates the high-dimensional distribution of the historical trajectory graph structure encoding and the activity area map features. The feature decoding sub-network decodes the high-dimensional distribution, graph structure features, and map features to obtain l possible predicted trajectories of the agent; The specific steps for the high-definition map encoding sub-network to mine the relevance between the agent's historical trajectory and its activity area and calculate the map features are as follows: Calculate the action angle of the agent according to the following formula: Among them, absc end , ordi end and absc end-1 , ordi end-1 respectively represent the abscissa and ordinate of the last two historical trajectories of the agent in the activity map; Determine the action area that the agent is about to enter based on the action angle and generate a position matrix in this area. Update the position matrix and use the features at the positions in the position matrix of the activity map as the map features; Step 4, calculate the loss value between the future trajectory prediction result and the label according to the following loss function based on metric learning; Among them, is the loss value between the predicted result y of the future trajectory and the label of the future trajectory , log(·) is the logarithmic operation with base 10, α is the label feature F of the future trajectory y and the predicted feature of the future trajectory of the metric space distance value of the weight parameter, the metric space uses the cosine similarity space, β is the relative entropy value D of z and z t ; KL (z, z t ) of the weight parameter, z is the high-dimensional distribution of the agent comprehensive feature F A , z t is the high-dimensional distribution after splicing the agent comprehensive feature F A and the signature trajectory feature F y ; Step 5, optimize the agent trajectory prediction network for heterogeneous data association mining; Use the Adam optimization algorithm to iteratively update the parameters of the intelligent agent trajectory prediction network for heterogeneous data association mining, and repeat Step 3 and Step 4 until the loss value between the future trajectory prediction result and the future trajectory label converges, and the trained network is obtained; Step 6, use the trained network for prediction: Input the historical trajectory information of the agent to be predicted and the high-definition map of its activity area into the trained network, and output the future trajectory prediction result of the agent.
2. The intelligent agent trajectory prediction method based on heterogeneous data association mining and metric learning according to claim 1, characterized in that: The historical trajectory map structure encoding sub-network described in step 1 is as follows: The historical trajectory map structure encoding sub-network consists of two parallel modules, the first long short-term memory network module and the second long short-term memory network module, which are respectively used for undirected graph node encoding and undirected graph edge encoding; the input dimension and hidden layer dimension in the first long short-term memory network module and the second long short-term memory network module are both set to 2 and 32.
3. The intelligent agent trajectory prediction method based on heterogeneous data association mining and metric learning according to claim 1, characterized in that: The high-dimensional distribution encoding sub-network described in step 1 is as follows: The high-dimensional distribution encoding sub-network consists of three parallel modules, the first fully connected layer, the second fully connected layer and the third fully connected layer, which are used to generate the high-dimensional distribution of the historical trajectory map structure features and map features. The input dimension and output dimension of the first fully connected layer are respectively set to 96 and 25, the input dimension and output dimension of the second fully connected layer are respectively set to 2 and 1, and the input dimension and output dimension of the third fully connected layer are respectively set to 96 + len(fut) and 25, where len(fut) is the length of the agent's future trajectory.
4. The intelligent agent trajectory prediction method based on heterogeneous data association mining and metric learning according to claim 1, characterized in that: The feature decoding sub-network described in step 1 is as follows: The feature decoding sub-network consists of three serially connected first fully connected layers, a first long short-term memory network module and a second fully connected layer, which are used to decode the high-dimensional distribution and generate the predicted trajectory of the agent. The input dimension and output dimension of the first fully connected layer are respectively set to 1 and l, where l is the number of predicted future trajectories. The input dimension and hidden dimension of the first long short-term memory network module are respectively set to 121 and 128, and the input dimension and output dimension of the second fully connected layer are respectively set to 128 and 2.
Citation Information
Patent Citations
Trajectory data mining method based on deep semi-supervised neural network
CN111368879A
Intelligent agent trajectory prediction method, system and device and storage medium
CN114022847A