A pedestrian trajectory prediction system and method considering interaction
By introducing a framework of attention mechanism in the pedestrian trajectory prediction system, considering the interaction relationship between pedestrians and vehicles, the problem that the prior art is difficult to accurately predict pedestrian trajectory in complex traffic environments is solved, and higher prediction accuracy and effective description of interaction relationships are achieved.
Patent Information
- Application Number
- CN202211724137.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-12-30
AI Technical Summary
The existing pedestrian trajectory prediction technology is difficult to accurately predict pedestrian movement trajectory in complex traffic environments, especially in road traffic scenarios where people and vehicles are mixed, ignoring the interaction between pedestrians and pedestrians and vehicles.
The pedestrian trajectory prediction framework based on attention mechanism is adopted, and the position information of pedestrians and vehicles is obtained through the sensor unit. The encoder calculates the historical trajectory, the attention module calculates the importance of time and spatial characteristics, and the decoder performs multimodal trajectory prediction and selects the optimal prediction trajectory.
It improves the accuracy of pedestrian trajectory prediction in complex traffic environments, can effectively describe the interaction between people-person and people-vehicle, enhances the simulation of uncertainty, and improves the accuracy of trajectory prediction.
Smart Images

Figure CN116129637B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of automobile intelligent driving, and in particular relates to a pedestrian trajectory prediction method considering interaction in an intelligent traffic environment. Background Art
[0002] Most of the existing intelligent driving is studied in a relatively closed environment. However, with the advancement of changing levels of intelligent driving, it is inevitable to face more complex traffic environments, especially in road traffic scenarios with mixed pedestrians and vehicles. Accurately predicting the movement trajectory of pedestrians is crucial to vehicle behavior decision-making.
[0003] However, most of the existing pedestrian trajectory prediction technologies are aimed at specific scenarios such as traffic intersections and crosswalks, and almost all of them only consider the characteristics of pedestrians themselves, ignoring the possible interactions between pedestrians and vehicles and other pedestrians, as well as the impact of the interaction on pedestrian trajectory changes. At the same time, the existing technology is difficult to accurately measure the importance of the impact of different factors on the trajectory in the interactive situation by considering only time or space characteristics. Therefore, studying the movement trend of pedestrians in the human-vehicle interaction relationship will help to more accurately obtain the future movement trajectory of pedestrians, which is of great value to the promotion of intelligent driving in all scenarios. Summary of the invention
[0004] In view of the above-mentioned deficiencies in the prior art, the purpose of the present invention is to provide a pedestrian trajectory prediction system and method that takes interaction into consideration, and to improve the accuracy of pedestrian trajectory prediction on roads with mixed pedestrian and vehicle traffic by constructing a pedestrian trajectory prediction framework based on an attention mechanism.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] A pedestrian trajectory prediction system considering interaction of the present invention comprises: a sensor unit, a pedestrian history trajectory database, a vehicle history trajectory database, and a vehicle-mounted computing unit;
[0007] Sensor units, including but not limited to vehicle-mounted cameras and laser radars, roadside sensor units, and pedestrian position sensors, are used to obtain position coordinate information of vehicles and pedestrians, and store the information in a vehicle history trajectory database and a pedestrian history trajectory database;
[0008] Vehicle history track database, used to store the historical location coordinate information of the vehicle;
[0009] Pedestrian historical trajectory database, used to store pedestrian historical location coordinate information;
[0010] An onboard computing unit, which includes an encoder, an attention module, and a decoder connected in sequence;
[0011] The encoder is used to calculate and store the historical trajectories of pedestrians and vehicles. The attention module includes a temporal attention module and a spatial attention module. The temporal attention module is used to calculate historical temporal features, and the spatial attention module is used to calculate global spatial features. The decoder is used to calculate the predicted trajectory of pedestrians taking into account interactive relationships.
[0012] A pedestrian trajectory prediction method considering interaction of the present invention is based on the above system, and the steps are as follows:
[0013] 1) The encoder encodes the motion characteristics in the pedestrian historical trajectory database and the vehicle historical trajectory database to generate hidden states and output tensors;
[0014] 2) The attention module encodes the temporal and spatial information of the motion characteristics of surrounding pedestrians and vehicles, assigns weights to the influence of different temporal and spatial information, determines the importance of each object's spatial information to pedestrians, and the importance of historical moment information to the current moment, and concatenates the calculated temporal feature vector and the global spatial feature vector to obtain a feature vector;
[0015] 3) The feature vector containing the surrounding pedestrians and vehicles and Gaussian sampling noise are input into the LSTM-based decoder for multimodal trajectory prediction to obtain a pedestrian prediction trajectory set, compare and distinguish the actual trajectory and the predicted trajectory, and select the optimal pedestrian prediction trajectory from the trajectory set through loss function calculation.
[0016] Furthermore, in step 1), the encoder uses an LSTM network to calculate hidden states and output tensors, as follows:
[0017] 11) The position coordinate information of each moment in the pedestrian historical trajectory database is concatenated and used as the encoder input to obtain a single pedestrian feature vector and semantic vector;
[0018]
[0019] In the formula, X i represents a single pedestrian feature vector, Represents information at each moment, represents the semantic vector, FC represents the fully connected layer, and W fcl represents the weights of the encoder's fully connected layer;
[0020] 12) Calculate the hidden state of pedestrian i at the current moment through the LSTM network, set the pedestrians in the pedestrian history trajectory database to use the same network weight, splice the encoder output of a single pedestrian at each moment, and obtain the encoder output of a single pedestrian; splice the encoder output corresponding to all pedestrians in the scene to obtain the output tensor of all pedestrians:
[0021]
[0022] In the formula, and Respectively represent the hidden state of pedestrian i at the current moment and the previous moment; W en represents the LSTM network weight; b en Indicates the deviation, H i,en represents the output of a single pedestrian encoder, H en Output tensor representing all pedestrians.
[0023] Furthermore, in step 2), the time attention coefficient is used to measure the importance of historical moment information to the current moment, specifically including:
[0024] 21) The pedestrian encoder output is used as the input of the temporal attention module to obtain the temporal attention coefficient:
[0025]
[0026] In the formula, represents the time attention coefficient, represents the intermediate vector, and They represent the intermediate vectors at time k and time s respectively, tanh is the activation function, and b α is the network parameter;
[0027] 22) The temporal attention module concatenates the temporal attention coefficient with the hidden state in step 12) and outputs the temporal feature vector
[0028]
[0029] Furthermore, in step 2), the spatial attention coefficient is used to measure the importance of each object's spatial information to the pedestrian, and the normalized attention coefficient is used to calculate the pedestrian's global spatial features, specifically including:
[0030] 23) Calculate the distance and angle information of human-human and human-vehicle spatial interactions:
[0031]
[0032] In the formula, is the distance between pedestrian i and surrounding pedestrian j at time t; is the distance between pedestrian i and the vehicle at time t, is the angle between the velocity vector of pedestrian i and the direction vector from pedestrian i to pedestrian j; is the angle between the velocity vectors of pedestrian i and pedestrian j; is the angle between the pedestrian i and the vehicle velocity vector; are the position coordinates of pedestrian i and pedestrian j at time t, is the position coordinate of the vehicle at time t;
[0033] 24) Cascade the human-human and human-vehicle interaction information to obtain three interaction feature vectors:
[0034]
[0035] In the formula, is the human-human interaction relationship vector between pedestrian i and pedestrian j; is the human-vehicle interaction relationship vector between pedestrian i and vehicle; is the vector of the interaction relationship between pedestrian i and other road participants;
[0036] 25) Calculate the spatial attention coefficient and normalize it to obtain the global spatial feature vector of pedestrian i at time t
[0037]
[0038] In the formula, is the spatial attention coefficient of the interaction between pedestrian i and pedestrian j at time t, is the feature vector concatenated from the interactive feature vectors of pedestrian i, W fc2 are the fully connected layer parameters of the spatial attention module, is the distance between pedestrian j and the unmanned vehicle, is the distance between pedestrians other than pedestrian i and the vehicle, is the feature vector formed by the distance between vehicles and pedestrians, N is the number of pedestrians, and φ is the normalization function.
[0039] Furthermore, in step 2), the calculated time feature vector and global space feature vector are concatenated, which is expressed as:
[0040]
[0041] In the formula, is the feature vector.
[0042] Furthermore, in step 3), a multi-layer perceptron (MLP) is used to compare and distinguish the actual trajectory and the predicted trajectory, and the pedestrian predicted trajectory set is:
[0043]
[0044] Where k is the number of predicted trajectories, and is the intermediate state vector between time t and time t-1, is the feature vector of the kth trajectory, is Gaussian sampling noise, and the prediction time is t = T obs+1 , ..., T pred ; and is the hidden state of the decoder at time t and time t-1, W de is the decoder network weight parameter, To predict the trajectory, W m2 is the weight parameter of the multi-layer perceptron, and are the outputs of the observation period decoder and encoder respectively.
[0045] Furthermore, in step 3), the multimodal loss function is minimized, and an exponential term is set to control the proportion of each error in the entire loss. According to the calculation results, the trajectory with the smallest calculated value is selected from the pedestrian prediction trajectory set as the optimal pedestrian prediction trajectory. The calculation method of minimizing the multimodal loss function L is as follows:
[0046]
[0047] In the formula, T is the prediction time, is the predicted trajectory to be selected, is the true trajectory, λ is a hyperparameter, and e is a constant.
[0048] Beneficial effects of the present invention:
[0049] 1. The present invention assigns different weights to the input sequence from the time dimension, reduces the deviation of pedestrians at different historical moments, and simultaneously extracts the human-human and human-vehicle interaction information, and simulates the interaction relationship between surrounding pedestrians and vehicles from the spatial dimension. By considering the interaction effects of the time and space dimensions at the same time, the accuracy of trajectory prediction under interaction is improved. In addition, random noise is introduced to further enhance the simulation of uncertainty.
[0050] 2. The method of the present invention can effectively describe the human-human and human-vehicle interaction relationships, which helps the pedestrian trajectory prediction model to learn the pedestrian trajectory characteristics in detail, and can obtain a unified pedestrian trajectory prediction feature representation without losing time and space information. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a block diagram of the pedestrian trajectory prediction principle in the present invention. DETAILED DESCRIPTION
[0052] In order to facilitate the understanding of those skilled in the art, the present invention is further described below in conjunction with embodiments and drawings. The contents mentioned in the implementation modes are not intended to limit the present invention.
[0053] A pedestrian trajectory prediction system considering interaction of the present invention comprises: a sensor unit, a pedestrian history trajectory database, a vehicle history trajectory database, and a vehicle-mounted computing unit;
[0054] Sensor units, including but not limited to vehicle-mounted cameras and laser radars, roadside sensor units, and pedestrian position sensors, are used to obtain position coordinate information of vehicles and pedestrians, and store the information in a vehicle history trajectory database and a pedestrian history trajectory database;
[0055] Vehicle history track database, used to store the historical location coordinate information of the vehicle;
[0056] Pedestrian historical trajectory database, used to store pedestrian historical location coordinate information;
[0057] An onboard computing unit, which includes an encoder, an attention module, and a decoder connected in sequence;
[0058] The encoder is used to calculate the historical trajectories of pedestrians and vehicles and store them (including the position coordinates of pedestrians and vehicles at different times); the attention module includes a temporal attention module and a spatial attention module. The temporal attention module is used to calculate historical time features, and the spatial attention module is used to calculate global spatial features. The decoder is used to calculate the predicted trajectory of pedestrians taking into account interactive relationships.
[0059] The pedestrian trajectory prediction method considering interaction of the present invention is based on the above system, such as Figure 1 As shown, the steps are as follows:
[0060] 1) The encoder encodes the motion characteristics in the pedestrian historical trajectory database and the vehicle historical trajectory database to generate hidden states and output tensors;
[0061] Among them, the encoder uses the LSTM network to calculate the hidden state and output tensor, as follows:
[0062] 11) The position coordinate information of each moment in the pedestrian historical trajectory database is concatenated and used as the encoder input to obtain a single pedestrian feature vector and semantic vector;
[0063]
[0064] In the formula, X i represents a single pedestrian feature vector, Represents information at each moment, represents the semantic vector, FC represents the fully connected layer, and W fc1 represents the weights of the encoder's fully connected layer;
[0065] 12) Calculate the hidden state of pedestrian i at the current moment through the LSTM network, set the pedestrians in the pedestrian history trajectory database to use the same network weight, splice the encoder output of a single pedestrian at each moment, and obtain the encoder output of a single pedestrian; splice the encoder output corresponding to all pedestrians in the scene to obtain the output tensor of all pedestrians:
[0066]
[0067] In the formula, and Respectively represent the hidden state of pedestrian i at the current moment and the previous moment; W en represents the LSTM network weight; b en Indicates the deviation, H i,en represents the output of a single pedestrian encoder, H en Output tensor representing all pedestrians.
[0068] 2) The attention module encodes the temporal and spatial information of the motion characteristics of surrounding pedestrians and vehicles through the attention mechanism calculation method, assigns weights to the influence of different temporal and spatial information, determines the importance of each object's spatial information to pedestrians, and the importance of historical moment information to the current moment, and concatenates the calculated temporal feature vector and the global spatial feature vector to obtain the feature vector;
[0069] Among them, the time attention coefficient is used to measure the importance of historical moment information to the current moment, including:
[0070] 21) The pedestrian encoder output is used as the input of the temporal attention module to obtain the temporal attention coefficient:
[0071]
[0072] In the formula, represents the time attention coefficient, represents the intermediate vector, and They represent the intermediate vectors at time k and time s respectively, tanh is the activation function, and b α is the network parameter;
[0073] 22) The temporal attention module concatenates the temporal attention coefficient with the hidden state in step 12) and outputs the temporal feature vector
[0074]
[0075] In addition, the spatial attention coefficient is used to measure the importance of each object's spatial information to the pedestrian, and the normalized attention coefficient is used to calculate the pedestrian's global spatial features, including:
[0076] 23) Calculate the distance and angle information of human-human and human-vehicle spatial interactions:
[0077]
[0078] In the formula, is the distance between pedestrian i and surrounding pedestrian j at time t; is the distance between pedestrian i and the vehicle at time t, is the angle between the velocity vector of pedestrian i and the direction vector from pedestrian i to pedestrian j; is the angle between the velocity vectors of pedestrian i and pedestrian j; is the angle between the pedestrian i and the vehicle velocity vector; are the position coordinates of pedestrian i and pedestrian j at time t, is the position coordinate of the vehicle at time t;
[0079] 24) Cascade the human-human and human-vehicle interaction information to obtain three interaction feature vectors:
[0080]
[0081] In the formula, is the human-human interaction relationship vector between pedestrian i and pedestrian j; is the human-vehicle interaction relationship vector between pedestrian i and vehicle; is the vector of the interaction relationship between pedestrian i and other road participants;
[0082] 25) Calculate the spatial attention coefficient and normalize it to obtain the global spatial feature vector of pedestrian i at time t
[0083]
[0084] In the formula, is the spatial attention coefficient of the interaction between pedestrian i and pedestrian j at time t, is the feature vector concatenated from the interactive feature vectors of pedestrian i, W fc2 are the fully connected layer parameters of the spatial attention module, is the distance between pedestrian j and the driverless car, is the distance between pedestrians other than pedestrian i and the vehicle, is the feature vector formed by the distance between vehicles and pedestrians, N is the number of pedestrians, and φ is the normalization function.
[0085] The calculated time feature vector and global space feature vector are concatenated and expressed as:
[0086]
[0087] In the formula, is the feature vector.
[0088] 3) The feature vectors and Gaussian sampling noise containing surrounding pedestrians and vehicles are input into the LSTM-based decoder for multimodal trajectory prediction to obtain a pedestrian prediction trajectory set, compare and distinguish the actual trajectory and the predicted trajectory, and select the optimal pedestrian prediction trajectory from the trajectory set through loss function calculation;
[0089] Among them, a multi-layer perceptron (MLP) is used to compare and distinguish the actual trajectory and the predicted trajectory. The pedestrian prediction trajectory set is:
[0090]
[0091] Where k is the number of predicted trajectories, and is the intermediate state vector between time t and time t-1, is the feature vector of the kth trajectory, is Gaussian sampling noise, and the prediction time is t = T obs+1 , ..., T pred ; and is the hidden state of the decoder at time t and time t-1, W de is the decoder network weight parameter, To predict the trajectory, W m2 is the weight parameter of the multi-layer perceptron, and are the outputs of the observation period decoder and encoder respectively.
[0092] By minimizing the multimodal loss function, an exponential term is set to control the proportion of each error in the entire loss. According to the calculation results, the trajectory with the smallest calculated value is selected from the pedestrian prediction trajectory set as the optimal pedestrian prediction trajectory. The calculation method for minimizing the multimodal loss function L is as follows:
[0093]
[0094] In the formula, T is the prediction time, is the predicted trajectory to be selected, is the true trajectory, λ is a hyperparameter, and e is a constant.
[0095] The present invention has many specific application paths. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements can be made without departing from the principle of the present invention. These improvements should also be regarded as the protection scope of the present invention.
Claims
1. A pedestrian trajectory prediction method considering interaction, based on a pedestrian trajectory prediction system considering interaction, the system includes: Sensor unit, pedestrian historical trajectory database, vehicle historical trajectory database, vehicle-mounted computing unit; Sensor units, including but not limited to vehicle-mounted cameras and laser radars, roadside sensor units, and pedestrian position sensors, are used to obtain position coordinate information of vehicles and pedestrians, and store the information in a vehicle history trajectory database and a pedestrian history trajectory database; Vehicle history track database, used to store the historical location coordinate information of the vehicle; Pedestrian historical trajectory database, used to store pedestrian historical location coordinate information; An onboard computing unit, which includes an encoder, an attention module, and a decoder connected in sequence; The encoder is used to calculate and store the historical trajectories of pedestrians and vehicles. The attention module includes a temporal attention module and a spatial attention module. The temporal attention module is used to calculate historical temporal features, and the spatial attention module is used to calculate global spatial features. The decoder is used to calculate the predicted trajectory of pedestrians taking into account the interaction relationship; The method is characterized in that the steps are as follows: 1) The encoder encodes the motion characteristics in the pedestrian historical trajectory database and the vehicle historical trajectory database to generate hidden states and output tensors; 2) The attention module encodes the temporal and spatial information of the motion characteristics of surrounding pedestrians and vehicles, assigns weights to the influence of different temporal and spatial information, determines the importance of each object's spatial information to pedestrians, and the importance of historical moment information to the current moment, and concatenates the calculated temporal feature vector and the global spatial feature vector to obtain a feature vector; 3) The feature vectors and Gaussian sampling noise containing surrounding pedestrians and vehicles are input into the LSTM-based decoder for multimodal trajectory prediction to obtain a pedestrian prediction trajectory set, compare and distinguish the actual trajectory and the predicted trajectory, and select the optimal pedestrian prediction trajectory from the trajectory set through loss function calculation; In step 2), the time attention coefficient is used to measure the importance of historical moment information to the current moment, specifically including: 21) The pedestrian encoder output is used as the input of the temporal attention module to obtain the temporal attention coefficient: In the formula, represents the temporal attention coefficient, represents the intermediate vector, and They represent the intermediate vectors at time k and time s respectively, tanh is the activation function, and b α is the network parameter; 22) The temporal attention module concatenates the temporal attention coefficient with the hidden state in step 12) and outputs the temporal feature vector In step 2), the spatial attention coefficient is used to measure the importance of each object's spatial information to the pedestrian, and the normalized attention coefficient is used to calculate the pedestrian's global spatial features, specifically including: 23) Calculate the distance and angle information of human-human and human-vehicle spatial interactions: In the formula, is the distance between pedestrian i and surrounding pedestrian j at time t; is the distance between pedestrian i and the vehicle at time t, is the angle between the velocity vector of pedestrian i and the direction vector from pedestrian i to pedestrian j; is the angle between the velocity vectors of pedestrian i and pedestrian j; is the angle between the pedestrian i and the vehicle velocity vector; are the position coordinates of pedestrian i and pedestrian j at time t, is the position coordinate of the vehicle at time t; 24) Cascade the human-human and human-vehicle interaction information to obtain three interaction feature vectors: In the formula, is the human-human interaction relationship vector between pedestrian i and pedestrian j; is the human-vehicle interaction relationship vector between pedestrian i and vehicle; is the vector of the interaction relationship between pedestrian i and other road participants; 25) Calculate the spatial attention coefficient and normalize it to obtain the global spatial feature vector of pedestrian i at time t In the formula, is the spatial attention coefficient of the interaction between pedestrian i and pedestrian j at time t, is the feature vector concatenated from the interactive feature vectors of pedestrian i, W fc2 are the fully connected layer parameters of the spatial attention module, is the distance between pedestrian j and the unmanned vehicle, is the distance between pedestrians other than pedestrian i and the vehicle, is the feature vector formed by the distance between vehicles and pedestrians, N is the number of pedestrians, and φ is the normalization function; In step 2), the calculated time feature vector and global space feature vector are concatenated, which is expressed as: In the formula, is the feature vector; In step 3), a multi-layer perceptron is used to compare and distinguish the actual trajectory and the predicted trajectory. The pedestrian predicted trajectory set is: Where k is the number of predicted trajectories, and is the intermediate state vector between time t and time t-1, is the feature vector of the kth trajectory, is Gaussian sampling noise, and the prediction time is t = T obs+1 , ..., T pred ; and is the hidden state of the decoder at time t and time t-1, W de is the decoder network weight parameter, To predict the trajectory, W m2 is the weight parameter of the multi-layer perceptron, and are the outputs of the observation period decoder and encoder respectively; In the step 3), the multimodal loss function is minimized, and an exponential term is set to control the proportion of each error in the entire loss. According to the calculation results, the trajectory with the smallest calculated value is selected from the pedestrian prediction trajectory set as the optimal pedestrian prediction trajectory. The calculation method of minimizing the multimodal loss function L is as follows: In the formula, T is the prediction time, is the predicted trajectory to be selected, is the true trajectory, λ is a hyperparameter, and e is a constant.
2. The method for predicting pedestrian trajectories considering interaction according to claim 1, characterized in that: In step 1), the encoder uses an LSTM network to calculate hidden states and output tensors, as follows: 11) The position coordinate information of each moment in the pedestrian historical trajectory database is concatenated and used as the encoder input to obtain a single pedestrian feature vector and semantic vector; Where, X i represents a single pedestrian feature vector, Represents information at each moment, represents the semantic vector, FC represents the fully connected layer, and W fc1 represents the weights of the encoder's fully connected layer; 12) Calculate the hidden state of pedestrian i at the current moment through the LSTM network, set the pedestrians in the pedestrian history trajectory database to use the same network weight, splice the encoder output of a single pedestrian at each moment, and obtain the encoder output of a single pedestrian; splice the encoder output corresponding to all pedestrians in the scene to obtain the output tensor of all pedestrians: In the formula, and Respectively represent the hidden state of pedestrian i at the current moment and the previous moment; W en represents the LSTM network weight; b en Indicates the deviation, H i,en represents the output of a single pedestrian encoder, H en Output tensor representing all pedestrians.
Citation Information
Patent Citations
Street-crossing pedestrian trajectory prediction method based on SFM-LSTM neural network model
CN114462667A
Pedestrian trajectory prediction method based on space-time diagram attention network
CN115376103A