Track prediction method and device, readable storage medium, program product and robot
By obtaining the historical trajectors of target pedestrians and adjacent pedestrians, using the encoder-decoder model of conditional variational automatic encoder and multi-head attention mechanism, combined with map information, the problem of difficult intentions in pedestrian trajectory prediction is solved, and more accurate prediction and better robot path planning is achieved.
Patent Information
- Application Number
- CN202410008119.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-03
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, it is difficult to accurately obtain the intention of the target pedestrian, resulting in inaccurate prediction.
By obtaining the historical trajectors and adjacent pedestrians, an encoder-decoder model based on conditional variational automatic encoder and multi-head attention mechanism is used to infer the intention of the target pedestrian and improve prediction accuracy.
It improves the accuracy of target pedestrian trajectory prediction, enhances robustness in complex environments, and supports robots to better conduct path planning.
Smart Images

Figure CN120260071A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a trajectory prediction method, device, readable storage medium, program product, and robot. Background Art
[0002] Trajectory prediction is an essential direction in current robot control. Pedestrian trajectory prediction is a branch of trajectory prediction. Based on pedestrian trajectory prediction, it is convenient for the robot to plan the best planning decision for itself.
[0003] In the related technical solutions, the focus of pedestrian trajectory prediction is on solving the extraction and fusion of interaction information between the prediction target and surrounding people, making it difficult to obtain the intention of the target pedestrian and resulting in inaccurate trajectory prediction of the target pedestrian. Summary of the Invention
[0004] The present invention aims to solve at least one of the technical problems existing in the prior art or related technologies.
[0005] To this end, in the first aspect of the present invention, a trajectory prediction method is provided.
[0006] In the second aspect of the present invention, a trajectory prediction device is provided.
[0007] In the third aspect of the present invention, another trajectory prediction device is provided.
[0008] In the fourth aspect of the present invention, a non-volatile readable storage medium is provided.
[0009] In the fifth aspect of the present invention, a computer program product is provided.
[0010] In the sixth aspect of the present invention, a robot is provided.
[0011] In view of this, according to the first aspect of the present invention, a trajectory prediction method is provided, including: obtaining historical trajectories and map information, where the historical trajectories include the first historical trajectory of the target pedestrian and the second historical trajectory of the neighboring pedestrians adjacent to the target pedestrian; encoding the map information and the first historical trajectory using a first encoder to obtain a first encoding result; encoding the second historical trajectory using the first encoder to obtain a second encoding result; obtaining the predicted trajectory of the device; encoding the predicted trajectory of the device using a second encoder to obtain a third encoding result; determining a classification latent vector based on the core features in the second encoding result, the first encoding result, and the third encoding result; and decoding the classification latent vector using a decoder to obtain the predicted trajectory of the target pedestrian.
[0012] According to a second aspect of the present invention, the present invention provides a trajectory prediction device, comprising: a first acquisition unit configured to acquire historical trajectories and map information, the historical trajectories including a first historical trajectory of a target pedestrian and a second historical trajectory of a neighboring pedestrian adjacent to the target pedestrian; a first encoding unit configured to encode the map information and the first historical trajectory using a first encoder to obtain a first encoding result; a second encoding unit configured to encode the second historical trajectory using the first encoder to obtain a second encoding result; a second acquisition unit configured to acquire a prediction trajectory of the device; a third encoding unit configured to encode the prediction trajectory of the device using a second encoder to obtain a third encoding result; a determination unit configured to determine a classification latent vector based on core features in the second encoding result, the first encoding result, and the third encoding result; and a prediction unit configured to decode the classification latent vector using a decoder to obtain a prediction trajectory of the target pedestrian.
[0013] According to a third aspect of the present invention, the present invention provides another trajectory prediction device, comprising a processor and a memory, the memory storing a program or instructions that can run on the processor, and when the program or instructions are executed by the processor, the steps of the method according to any one of the above are implemented.
[0014] According to a fourth aspect of the present invention, the present invention provides a non-volatile readable storage medium, on which a program or instructions are stored, and when the program or instructions are executed by a processor, the steps of the method according to any one of the above are implemented.
[0015] According to a fifth aspect of the present invention, the present invention provides a computer program product, the computer program product being stored in a storage medium, and the computer program product being executed by at least one processor to implement the steps of the method according to any one of the above.
[0016] According to a sixth aspect of the present invention, the present invention provides a robot, comprising: any one of the above trajectory prediction devices; and / or the above non-volatile readable storage medium; and / or the above computer program product.
[0017] In the embodiments of the present application, when predicting the prediction trajectory of the target pedestrian, the historical trajectories of neighboring pedestrians adjacent to the target pedestrian are referred to. Therefore, the intention of the target pedestrian can be inferred based on the historical trajectories of the neighboring pedestrians. Based on this principle, the prediction trajectory of the target pedestrian predicted in the embodiments of the present application is more accurate, thereby improving the accuracy of the prediction trajectory of the target pedestrian.
[0018] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, in which:
[0020] Figure 1 It shows a schematic flowchart of a trajectory prediction method in an embodiment of the present invention;
[0021] Figure 2 It shows an overall block diagram of a trajectory prediction in the related art solution;
[0022] Figure 3 It shows a schematic block diagram of a trajectory prediction device in an embodiment of the present invention;
[0023] Figure 4 It shows a schematic block diagram of another trajectory prediction device in an embodiment of the present invention. Detailed implementation manners
[0024] In order to be able to more clearly understand the above aspects, features and advantages of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0025] Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0026] In one embodiment of the present application, as Figure 1 shown, a trajectory prediction method is provided, including:
[0027] Step 102, obtaining historical trajectories and map information, where the historical trajectories include the first historical trajectory of the target pedestrian and the second historical trajectory of the neighboring pedestrians adjacent to the target pedestrian;
[0028] Step 104, encoding the map information and the first historical trajectory by using a first encoder to obtain a first encoding result;
[0029] Step 106, encoding the second historical trajectory by using the first encoder to obtain a second encoding result;
[0030] Step 108, obtaining the predicted trajectory of the device;
[0031] Step 110, encoding the predicted trajectory of the device by using a second encoder to obtain a third encoding result;
[0032] Step 112, determining a classification latent vector based on the core features in the second encoding result, the first encoding result and the third encoding result;
[0033] Step 114: Use a decoder to decode the classification latent vector to obtain the predicted trajectory of the target pedestrian.
[0034] In an embodiment of the present application, a trajectory prediction method is proposed. By running the above trajectory prediction method, accurate prediction of the predicted trajectory of the target pedestrian can be achieved. Compared with related embodiments, in the embodiment of the present application, when predicting the predicted trajectory of the target pedestrian, the historical trajectories of neighboring pedestrians adjacent to the target pedestrian are referred to. Therefore, the intention of the target pedestrian can be inferred based on the historical trajectories of neighboring pedestrians. Based on this principle, the predicted trajectory of the target pedestrian predicted in the embodiment of the present application is more accurate, thereby improving the accuracy of the predicted trajectory of the target pedestrian.
[0035] Specifically, pedestrians rarely appear alone in a scene, but there will be companions or un-fixed passers-by around them. Pedestrian trajectory prediction in related embodiments usually only considers the historical trajectories and states of pedestrians themselves, or simulates the relationships between different pedestrians through spatio-temporal graphs, ignoring the different impacts of the historical states of pedestrians at different positions at different times on the prediction target, as well as the similarity of the intentions of neighboring pedestrians.
[0036] Based on this, the present invention proposes a pedestrian trajectory prediction method based on the historical behaviors of the nearest pedestrians, and fully utilizes the advantages of the attention mechanism, thereby improving the accuracy of trajectory prediction and the robustness in complex environments, and making the results of pedestrian trajectory prediction better used for the path planning of robots.
[0037] In some embodiments, optionally, the historical trajectory can be obtained by shooting or recording images and videos including the target pedestrian and neighboring pedestrians, and performing recognition and annotation on the captured images and / or recorded videos.
[0038] In some embodiments, optionally, the map information can be understood as the map of the area where the target pedestrian, neighboring pedestrians, and the device walk. It can be obtained by acquiring the map website of the area where the target pedestrian, neighboring pedestrians, and the device walk, or can be constructed by shooting or recording images and videos including the target pedestrian and neighboring pedestrians, and performing recognition on the captured images and / or recorded videos.
[0039] In some embodiments, optionally, the first encoder and the second encoder can be implemented for trajectory prediction based on a Conditional Variational Autoencoder (CVAE).
[0040] Among them, the first encoder is used to encode the historical trajectory, which can be regarded as an interactive history encoder, and the second encoder is used to encode the predicted trajectory of the device, which can be regarded as the robot's own future trajectory encoder.
[0041] In some embodiments, optionally, the first encoder includes a first neural network, the second historical trajectory includes a first sub-historical trajectory and a second sub-historical trajectory, the first sub-historical trajectory is the historical trajectory of adjacent pedestrians located in front of the target pedestrian, and the second sub-historical trajectory is the historical trajectory of adjacent pedestrians located on both sides of the target pedestrian, the second encoding result includes a first output result and a second output result, and the second historical trajectory is encoded using the first encoder to obtain a second encoding result, specifically including: encoding the first sub-historical trajectory using the first neural network to obtain a first output result; encoding the second sub-historical trajectory using the first neural network to obtain a second output result.
[0042] In this embodiment, the second historical trajectory is divided into a first sub-historical trajectory and a second sub-historical trajectory according to different nearby pedestrians, and then encoded by the first encoder, and respectively encoded by the same first neural network, thereby obtaining corresponding first output results and second output results.
[0043] In this process, on the basis of using the first neural network to extract convolutional features of historical trajectories and map information, a dual-stream method was added to extract the nearest pedestrian trajectory features, that is, to extract the second encoding result including the first output result and the second output result.
[0044] After determining the second encoding result, the multi-head attention mechanism can be used to combine the different features and the different effects on the target pedestrian at different times to re-encode it, so as to infer the intention of the target pedestrian, making the predicted trajectory of the target pedestrian predicted by the embodiment of the present application more accurate, thereby improving the accuracy of the predicted trajectory of the target pedestrian.
[0045] In some embodiments, the first neural network may be a long short-term memory network (LSTM). LSTM is a time recurrent neural network, which is specially designed to solve the long-term dependency problem of general recurrent neural networks (RNN). All RNNs have a chain form of repeated neural network modules. In standard RNN, this repeated structural module has only a very simple structure.
[0046] In this embodiment, the LSTM can be used in conjunction with the multi-head attention mechanism to extract the features of the historical trajectories of the nearest pedestrians and effectively encode these features according to their different degrees of influence on the target pedestrian. By combining the features of the nearest pedestrians, map features, and historical trajectory features of the target pedestrian obtained previously and effectively encoding them through the self-attention mechanism, the effect of accurate prediction of different trajectories can be achieved.
[0047] In some embodiments, optionally, the first encoding result includes a third output result and a fourth output result, and the first encoder further includes a second neural network. Encoding the map information and the first historical trajectory using the first encoder to obtain the first encoding result specifically includes: encoding the first historical trajectory using the first neural network to obtain the third output result; encoding the map information using the second neural network to obtain the fourth output result.
[0048] In this embodiment, the second neural network can be a Convolutional Neural Networks (CNN), which is a deep learning model or a multi-layer perceptron similar to an artificial neural network and is often used to analyze visual images.
[0049] By using the convolutional neural network, the recognition of the map information can be realized, and then its features can be extracted to obtain the corresponding features, that is, the fourth output result.
[0050] In this embodiment, encoding the first historical trajectory of the target pedestrian using the same first neural network as the historical trajectory makes the neural networks of the target pedestrian and the nearby pedestrians the same, and further makes the structures of the first output result, the second output result, and the third output result the same or similar, which is convenient for determining the similarity of the intentions of the target pedestrian and the nearby pedestrians, so as to make the predicted trajectory of the target pedestrian more accurate, and thus improve the accuracy of the predicted trajectory of the target pedestrian.
[0051] In some embodiments, optionally, the trajectory prediction method further includes: using the first output result as the key vector in the multi-head attention mechanism, using the second output result as the value vector in the multi-head attention mechanism, and using the third output result as the query vector in the multi-head attention mechanism; determining the attention score between the target pedestrian and the nearby pedestrians based on the key vector, the value vector, and the query vector; determining the core features in the second encoding result based on the attention score, the first output result, and the weight matrix.
[0052] Among them, the multi-head attention mechanism is a commonly used technology in deep learning. In the traditional attention mechanism, the model can only focus on one part of the input data, while the multi-head attention mechanism allows the model to focus on multiple parts simultaneously, thereby improving the performance of the model.
[0053] The principle of the multi - head attention mechanism is to divide the input data into multiple parts, each part has an independent attention head. Each attention head calculates the weights related to the input data, and then sums them up weighted by the weights to obtain the final output result.
[0054] In the multi - head attention mechanism, the key vector, that is, the Key vector, the value vector, that is, the Value vector, and the query vector, that is, the Query vector.
[0055] Specifically, in the multi - head attention mechanism, the Key vector and the Value vector are obtained by projecting the historical state encodings P (i,j) and A (i,j) of the pedestrians in front and on both sides. The Query vector is obtained by projecting the current action encoding H i of the predicted target. These input sequences are mapped to multiple different sub - spaces through linear transformation. Then, these components are encoded through a scaled dot - product attention block, and the formula is as follows:
[0056]
[0057] where d k is a parameter of the multi - head attention mechanism, is the transposed matrix of P (i,j) , i and j are coordinate positions, and S (i,j) is the attention score between the target pedestrian and the neighboring pedestrians.
[0058] The fused attention is obtained through weighted sum, and the formula is as follows:
[0059] R (i,j) = S (i,j) P (i,j) ;
[0060] where R (i,j) is the fused attention representation.
[0061] Each attention head captures different correlation patterns according to the attention weight S (i,j) . Finally, the outputs of multiple attention heads are connected using the weight matrix W to obtain the final multi - head attention representation R' (i,j) :
[0062] R' (i,j) = R (i,j) WR (i,j) ;
[0063] where W is the weight matrix.
[0064] In this embodiment, the self-attention mechanism is used to extract the relationships between different features, and the relationships between different inputs are utilized for fusion, thereby achieving effective fusion of diverse inputs in the encoder and laying a foundation for the subsequent decoding process.
[0065] In some embodiments, optionally, a classification latent vector is determined based on the core features in the second encoding result, the first encoding result, and the third encoding result. Specifically, it includes: constructing a discrete latent space based on the core features in the second encoding result, the first encoding result, and the third encoding result; performing random sampling on the discrete latent space to obtain the classification latent vector.
[0066] In this embodiment, the process of constructing the discrete latent space can be understood as using a fusion model to generate a discrete latent space. Specifically, a discrete latent space is generated based on the input data encoded by the first encoder and the second encoder using the fusion model, and then random sampling is performed on the discrete latent space to obtain a random sample Z, where Z is a classification latent vector encoding high-level latent behaviors, that is, the classification latent vector in the present application, so as to use the classification latent vector to generate the predicted trajectory of the target pedestrian.
[0067] In some embodiments, optionally, the decoder includes: a gated recurrent unit layer, a Gaussian mixture model component, and an integrator. The decoder is used to decode the classification latent vector to obtain the predicted trajectory of the target pedestrian. Specifically, it includes: using the gated recurrent unit layer to decode the classification latent vector to obtain a first decoding result; using the Gaussian mixture model component to decode the first decoding result to obtain a second decoding result; using the integrator to decode the second decoding result to obtain the predicted trajectory of the target pedestrian.
[0068] In this embodiment, if there are multiple gated recurrent unit layers, the decoder is composed of multiple gated recurrent unit layers, a Gaussian mixture model component, and an integrator. The above components are sequentially connected together to decode the classification latent vector to obtain the predicted trajectory of the target pedestrian.
[0069] Specifically, the gated recurrent unit layer (Gated Recurrent Unit, GRU) is a commonly used gated recurrent neural network.
[0070] The Gaussian mixture model component (Gaussian Mixture Model, GMM), that is, the Gaussian mixture model, is an extension of the single Gaussian probability density function. The GMM can smoothly approximate density distributions of arbitrary shapes.
[0071] The integrator, in the embodiment of the present application, refers to dynamic integration.
[0072] In some embodiments, optionally, the trajectory prediction method further includes: obtaining training samples, where the training samples include: historical trajectory samples, map information samples, predicted trajectory samples of a device, and predicted trajectory samples of a target pedestrian; training a preset encoder-decoder model based on the historical trajectory samples, map information samples, predicted trajectory samples of the device, and predicted trajectory samples of the target pedestrian to obtain a first encoder, a second encoder, and a decoder.
[0073] In this embodiment, training samples can be pre-constructed, and the preset encoder-decoder model can be trained based on the training samples to obtain a first encoder, a second encoder, and a decoder.
[0074] In some embodiments, optionally, the loss functions corresponding to the first encoder, the second encoder, and the decoder include the sum of the following: the loss function for the regression problem between the predicted trajectory and the actual trajectory, the loss function of the conditional variational autoencoder between the predicted trajectory and the actual trajectory, and the loss function corresponding to the mutual information between the input sample and the latent variable.
[0075] Specifically, in the method based on the Conditional Variational Auto Encoder (CVAE), the Evidence Lower Bound (ELBO) is usually used as the loss function. However, this loss function is very sensitive to outliers or noise in the training data. To balance the fitting ability and robustness of the model and improve its tolerance to outliers and noise, we introduce the loss function (Huber loss) for the regression problem between the predicted trajectory and the actual trajectory into the training process:
[0076]
[0077] where γ is a hyperparameter, L H is the Huber loss function for the predicted trajectory and the actual trajectory, y i is the actual state, and is the predicted state.
[0078] where the loss function of the Conditional Variational Auto Encoder (CVAE) is expressed as follows:
[0079]
[0080] where N is the total number of training samples. Φ, φ, and θ are the parameters of the model, which are learned through training. x i and y i are the input samples and their corresponding labels used to train the model, that is, x iis the historical state of the pedestrian, y i is the actual state, z is the latent variable sampled from the conditional probability distribution q φ (z|x i ,y i ), used to capture the representations of the input and the label, β is the hyperparameter that balances the reconstruction loss and the KL divergence term, D KL (q φ (z|x i ,y i )||p θ (z|x i ) is the KL divergence term, is the reconstruction loss.
[0081] Among them, in order to enhance the correlation between the input sample x and the latent variable z, the mutual information can be introduced into the loss function:
[0082] L MI =αI q (x;z);
[0083] Among them, L MI is the loss function corresponding to the mutual information between the input sample and the latent variable, α is the hyperparameter, I q (x;z) is the mutual information.
[0084] To maximize the mutual information, the total loss is a weighted combination of the above terms:
[0085] L total =L CVAE +L MI +L H ;
[0086] In some embodiments of the present application, as Figure 2 shown, in a scenario, we define each person as V i , and the time step is t. represents the historical state of each front pedestrian V j . represents the historical states of two adjacent side pedestrians V j . represents the historical state of the target pedestrian. M represents the high-definition map of this scenario. F t represents the future trajectory of the device. represents the future trajectory of the target pedestrian. represents the predicted future trajectory of the target pedestrian.
[0087] Among them, the interaction history encoder is also the first encoder in the present application, the robot's own future trajectory encoder is also the second encoder in the present application, and the predicted target future trajectory encoder is also the third encoder in the present application.
[0088] Among them, the second encoder and the third encoder are BI-LSTMs. Here, BI-LSTM is an abbreviation of Bi-Directional Long Short Term Memory, which is composed of a forward LSTM and a backward LSTM.
[0089] In some embodiments of the present application, exemplarily, the embodiments of the present application are developed and evaluated on the widely used benchmark dataset nuScenes. The nuScenes dataset collected 1000 driving scenarios on urban roads in Boston and Singapore. The duration of each video is 20 seconds and includes various influencing factors such as weather conditions, vehicle types, vegetation, road markings, and left and right traffic. The entire dataset was accurately 3D bounding box annotated for 23 object categories at a frequency of 2Hz and includes annotations on object visibility, mobility, and pose. The embodiments of the present application select the nearest pedestrians based on this dataset, then extract features through the network model, and finally achieve the trajectory prediction of pedestrians.
[0090] In one embodiment of the present application, as Figure 3 shown, a trajectory prediction device 300 is provided, including: a first acquisition unit 302 configured to acquire historical trajectories and map information, where the historical trajectories include the first historical trajectory of the target pedestrian and the second historical trajectory of the neighboring pedestrians adjacent to the target pedestrian; a first processing unit 304 configured to encode the map information and the first historical trajectory using a first encoder to obtain a first encoding result; a second processing unit 306 configured to encode the second historical trajectory using the first encoder to obtain a second encoding result; a second acquisition unit 308 configured to acquire the predicted trajectory of the device; a third processing unit 310 configured to encode the predicted trajectory of the device using a second encoder to obtain a third encoding result; a determination unit 312 configured to determine a classification latent vector based on the core features in the second encoding result, the first encoding result, and the third encoding result; and a prediction unit 314 configured to decode the classification latent vector using a decoder to obtain the predicted trajectory of the target pedestrian.
[0091] In the embodiments of the present application, a trajectory prediction device 300 is proposed, which can accurately predict the predicted trajectory of the target pedestrian. Compared with the related embodiments, in the embodiments of the present application, when predicting the predicted trajectory of the target pedestrian, the historical trajectories of the neighboring pedestrians adjacent to the target pedestrian are referred to. Therefore, the intention of the target pedestrian can be inferred based on the historical trajectories of the neighboring pedestrians. Based on this principle, the predicted trajectory of the target pedestrian predicted in the embodiments of the present application is more accurate, thereby improving the accuracy of the predicted trajectory of the target pedestrian.
[0092] Specifically, pedestrians rarely appear alone in a scene, but rather there are companions or random passers-by around them. Pedestrian trajectory prediction in related embodiments usually only considers the historical trajectory and state of the pedestrian itself, or simulates the relationships between different pedestrians through spatio-temporal graphs, ignoring the different impacts of the historical states of pedestrians in different positions at different times on the prediction target, as well as the similarity of the intentions of nearby pedestrians.
[0093] Based on this, the present invention proposes a pedestrian trajectory prediction method based on the historical behaviors of the nearest pedestrians, and fully utilizes the advantages of the attention mechanism, thereby improving the accuracy of trajectory prediction and the robustness in complex environments, and enabling the results of pedestrian trajectory prediction to be better used for the path planning of robots.
[0094] In some embodiments, optionally, the historical trajectory can be obtained by photographing or recording images and videos including the target pedestrian and nearby pedestrians, and performing recognition and annotation on the photographed images and / or recorded videos.
[0095] In some embodiments, optionally, the map information can be understood as the map of the area where the target pedestrian, nearby pedestrians, and the device are walking. It can be obtained by acquiring the map website of the area where the target pedestrian, nearby pedestrians, and the device are walking, or can be constructed by photographing or recording images and videos including the target pedestrian and nearby pedestrians, and performing recognition on the photographed images and / or recorded videos.
[0096] In some embodiments, optionally, the first encoder and the second encoder can implement trajectory prediction based on a Conditional Variational Autoencoder (CVAE).
[0097] Among them, the first encoder is used to encode the historical trajectory, which can be regarded as an interactive historical encoder, and the second encoder is used to encode the predicted trajectory of the device, which can be regarded as the future trajectory encoder of the robot itself.
[0098] In some embodiments, optionally, the first encoder includes a first neural network, the second historical trajectory includes a first sub-historical trajectory and a second sub-historical trajectory, the first sub-historical trajectory is the historical trajectory of the nearby pedestrians in front of the target pedestrian, the second sub-historical trajectory is the historical trajectory of the nearby pedestrians on both sides of the target pedestrian, the second encoding result includes a first output result and a second output result, and the second processing unit 306 is specifically configured to: encode the first sub-historical trajectory using the first neural network to obtain the first output result; encode the second sub-historical trajectory using the first neural network to obtain the second output result.
[0099] In this embodiment, according to different nearby pedestrians, the second historical trajectory is divided into a first sub-historical trajectory and a second sub-historical trajectory. Then, in the process of encoding them using a first encoder, the same first neural network is used to encode them respectively, and thus corresponding first and second output results are obtained.
[0100] In this process, on the basis of realizing the convolutional feature extraction of the historical trajectory and map information using the first neural network, a two-stream method is added to extract the nearest pedestrian trajectory features, that is, a second encoding result including the first output result and the second output result is obtained.
[0101] Furthermore, after determining the second encoding result, the multi-head attention mechanism can be used to re-encode it by combining the different influences of different features on the target pedestrian at different times, so as to infer the intention of the target pedestrian, making the predicted trajectory of the target pedestrian predicted in the embodiment of the present application more accurate, and thus improving the accuracy of the predicted trajectory of the target pedestrian.
[0102] In some embodiments, the first neural network may be a Long Short Term Memory (LSTM) network. LSTM is a type of recurrent neural network designed specifically to solve the long-term dependence problem existing in general recurrent neural networks (RNNs). All RNNs have a chain form of repeating neural network modules. In a standard RNN, this repeating structural module has a very simple structure.
[0103] In this embodiment, LSTM can be used in conjunction with the multi-head attention mechanism to extract the features of the nearest pedestrian's historical trajectory and effectively encode these features according to their different degrees of influence on the target pedestrian. By combining the previously obtained nearest pedestrian features, map features, and the historical trajectory features of the target pedestrian and effectively encoding them through the self-attention mechanism, the effect of accurate prediction of different trajectories can be achieved.
[0104] In some embodiments, optionally, the first encoding result includes a third output result and a fourth output result. The first encoder further includes a second neural network and a first processing unit 304, which is specifically configured to: encode the first historical trajectory using the first neural network to obtain a third output result; encode the map information using the second neural network to obtain a fourth output result.
[0105] In this embodiment, the second neural network may be a Convolutional Neural Networks (CNN), which is a deep learning model or a multi-layer perceptron similar to an artificial neural network and is often used to analyze visual images.
[0106] By using a convolutional neural network, the recognition of map information can be achieved, and then its feature extraction can be carried out to obtain corresponding features, that is, the fourth output result.
[0107] In this embodiment, the first neural network identical to the historical trajectory is used to encode the first historical trajectory of the target pedestrian, so that the neural networks for the target pedestrian and the neighboring pedestrians are the same, and further, the structures of the first output result, the second output result, and the third output result are the same or similar, which is convenient for determining the similarity of the intentions of the target pedestrian and the neighboring pedestrians, thereby making the predicted trajectory of the target pedestrian more accurate, and further improving the accuracy of the predicted trajectory of the target pedestrian.
[0108] In some embodiments, optionally, the first output result is used as the key vector in the multi-head attention mechanism, the second output result is used as the value vector in the multi-head attention mechanism, and the third output result is used as the query vector in the multi-head attention mechanism; based on the key vector, the value vector, and the query vector, the attention score between the target pedestrian and the neighboring pedestrians is determined; based on the attention score, the first output result, and the weight matrix, the core features in the second encoding result are determined.
[0109] In the multi-head attention mechanism, the key vector, that is, the Key vector, the value vector, that is, the Value vector, and the query vector, that is, the Query vector.
[0110] Specifically, the Key vector and the Value vector in the multi-head attention mechanism are obtained by projecting the historical states of the pedestrians in front and on both sides, P (i,j) and A (i,j) The Query vector is obtained by projecting the current action encoding of the predicted target, H i These input sequences are mapped to multiple different subspaces through linear transformation. Then, these components are encoded through a scaled dot-product attention block, and the formula is as follows:
[0111]
[0112] where d k is a parameter of the multi-head attention mechanism, is the transpose matrix of P (i,j) , i and j are coordinate positions, and S (i,j) is the attention score between the target pedestrian and the neighboring pedestrians.
[0113] The fused attention is obtained through weighted sum, and the formula is as follows:
[0114] R (i,j) = S (i,j) P (i,j) ;
[0115] Among them, R (i,j) is the fused attention representation.
[0116] Each attention head captures different correlation patterns according to the attention weight S (i,j) Finally, the outputs of multiple attention heads are connected using the weight matrix W to obtain the final multi-head attention representation R′ (i,j) :
[0117] R′ (i,j) = R (i,j) WR (i,j) ;
[0118] Among them, W is the weight matrix.
[0119] In this embodiment, the relationship between different features is extracted through the self-attention mechanism, and the relationship between different inputs is used for fusion, so as to effectively fuse diverse inputs in the encoder and lay a foundation for the subsequent decoding process.
[0120] In some embodiments, optionally, the determining unit 312 is specifically configured to: construct a discrete latent space based on the core features in the second encoding result, the first encoding result, and the third encoding result; perform random sampling based on the discrete latent space to obtain a classification latent vector.
[0121] In this embodiment, the process of constructing the discrete latent space can be understood as using a fusion model to generate a discrete latent space. Specifically, a discrete latent space is generated based on the input data encoded by the first encoder and the second encoder, and then random sampling is performed in the discrete latent space to obtain a random sample Z, where Z is a classification latent vector encoding high-level latent behaviors, that is, the classification latent vector in the present application, so as to use the classification latent vector to generate the predicted trajectory of the target pedestrian.
[0122] In some embodiments, optionally, the decoder includes: a gated recurrent unit layer, a Gaussian mixture model component, and an integrator. The prediction unit 314 is specifically configured to: decode the classification latent vector using the gated recurrent unit layer to obtain a first decoding result; decode the first decoding result using the Gaussian mixture model component to obtain a second decoding result; decode the second decoding result using the integrator to obtain the predicted trajectory of the target pedestrian.
[0123] In this embodiment, if there are multiple gated recurrent unit layers, the decoder consists of multiple gated recurrent unit layers, a Gaussian mixture model component, and an integrator. The above components are connected in sequence to decode the classification latent vector to obtain the predicted trajectory of the target pedestrian.
[0124] Specifically, the Gated Recurrent Unit (GRU) layer is a commonly used gated recurrent neural network.
[0125] The Gaussian Mixture Model (GMM) component, that is, the Gaussian mixture model, is an extension of a single Gaussian probability density function. The GMM can smoothly approximate density distributions of arbitrary shapes.
[0126] The integrator, in the embodiments of the present application, refers to dynamic integration.
[0127] In some embodiments, optionally, the prediction unit 314 is further configured to: obtain training samples, where the training samples include: historical trajectory samples, map information samples, predicted trajectory samples of the device, and predicted trajectory samples of the target pedestrian; and train a preset encoder-decoder model based on the historical trajectory samples, map information samples, predicted trajectory samples of the device, and predicted trajectory samples of the target pedestrian, so as to obtain a first encoder, a second encoder, and a decoder.
[0128] In this embodiment, training samples can be pre-constructed and the preset encoder-decoder model can be trained based on the training samples, so as to obtain a first encoder, a second encoder, and a decoder.
[0129] In some embodiments, optionally, the loss functions corresponding to the first encoder, the second encoder, and the decoder include the sum of the following: the loss function for the regression problem between the predicted trajectory and the actual trajectory, the loss function of the conditional variational autoencoder between the predicted trajectory and the actual trajectory, and the loss function corresponding to the mutual information between the input sample and the latent variable.
[0130] In one of the embodiments, as Figure 4 shown, the present invention provides another trajectory prediction device 400, including a processor 402 and a memory 404. The memory 404 stores a program or instructions that can run on the processor 402. When the program or instructions are executed by the processor, the steps of any of the methods described above are implemented.
[0131] Among them, the memory 404 can be used to store software programs and various data. The memory 404 mainly includes a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 404 can include a volatile memory or a non-volatile memory, or the memory can include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memory.
[0132] In one embodiment, the present invention provides a non-volatile readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method according to any one of the above are implemented.
[0133] In one embodiment, the present invention provides a computer program product, which is stored in a storage medium. The computer program product is executed by at least one processor to implement the steps of the method according to any one of the above.
[0134] In one embodiment, the present invention provides a robot, including: any one of the above trajectory prediction devices; and / or the above non-volatile readable storage medium; and / or the above computer program product.
[0135] In some embodiments, the robot can be a cleaning robot, such as a floor cleaning robot, or a walking robot.
[0136] The terms "first", "second" in the description and claims of this application may explicitly or implicitly include one or more of such features. In the written description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more. Additionally, "and / or" in the description and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.
[0137] In the written description of the present invention, it can be understood that, except for explicit regulations and limitations, the terms "installed", "connected", "coupled" should be understood in a broad sense. For example, it can be fixedly connected, detachably connected, or integrally connected; it can be a mechanical structure connection, or an electrical connection; it can be a direct connection between the two, or an indirect connection between the two through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0138] In the claims, description, and drawings of the present invention, the descriptions of terms such as "one embodiment", "some embodiments", "specific embodiments", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In the claims, description, and drawings of the present invention, the schematic representations of the above terms do not necessarily refer to the same embodiment or instance. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0139] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various modifications and changes. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A trajectory prediction method, characterized in that, Including: Obtain historical trajectory and map information, where the historical trajectory includes the first historical trajectory of the target pedestrian and the second historical trajectory of the neighboring pedestrians adjacent to the target pedestrian; Encode the map information and the first historical trajectory using a first encoder to obtain a first encoding result; Encode the second historical trajectory using the first encoder to obtain a second encoding result; Obtain the predicted trajectory of the device; Encode the predicted trajectory of the device using a second encoder to obtain a third encoding result; Determine a classification latent vector based on the core features in the second encoding result, the first encoding result, and the third encoding result; Decode the classification latent vector using a decoder to obtain the predicted trajectory of the target pedestrian.
2. The trajectory prediction method according to claim 1, wherein The first encoder includes a first neural network. The second historical trajectory includes a first sub-historical trajectory and a second sub-historical trajectory. The first sub-historical trajectory is the historical trajectory of the neighboring pedestrian in front of the target pedestrian, and the second sub-historical trajectory is the historical trajectory of the neighboring pedestrians on both sides of the target pedestrian. The second encoding result includes a first output result and a second output result. The step of encoding the second historical trajectory using the first encoder to obtain the second encoding result specifically includes: Encode the first sub-historical trajectory using the first neural network to obtain the first output result; Encode the second sub-historical trajectory using the first neural network to obtain the second output result.
3. The trajectory prediction method according to claim 2, wherein The first encoding result includes a third output result and a fourth output result. The first encoder further includes a second neural network. The step of encoding the map information and the first historical trajectory using the first encoder to obtain the first encoding result specifically includes: Encode the first historical trajectory using the first neural network to obtain the third output result; Encode the map information using the second neural network to obtain the fourth output result.
4. The trajectory prediction method according to claim 3, characterized in that The trajectory prediction method further includes: Use the first output result as the key vector in the multi-head attention mechanism, use the second output result as the value vector in the multi-head attention mechanism, and use the third output result as the query vector in the multi-head attention mechanism; Determine the attention score between the target pedestrian and the neighboring pedestrians based on the key vector, the value vector, and the query vector; Determine the core features in the second encoding result based on the attention score, the first output result, and the weight matrix.
5. The trajectory prediction method according to any one of claims 1 to 4, characterized in that The step of determining the classification latent vector based on the core features in the second encoding result, the first encoding result, and the third encoding result specifically includes: Construct a discrete latent space based on the core features in the second encoding result, the first encoding result, and the third encoding result; Randomly sample based on the discrete latent space to obtain the classification latent vector.
6. The trajectory prediction method according to any one of claims 1 to 4, characterized in that, The decoder includes: a gated recurrent unit layer, a Gaussian mixture model component, and an integrator. The step of decoding the classification latent vector using the decoder to obtain the predicted trajectory of the target pedestrian specifically includes: Decode the classification latent vector using the gated recurrent unit layer to obtain a first decoding result; Decode the first decoding result using the Gaussian mixture model component to obtain a second decoding result; Decode the second decoding result using the integrator to obtain the predicted trajectory of the target pedestrian.
7. The trajectory prediction method according to any one of claims 1 to 4, characterized in that, The trajectory prediction method further includes: Obtain training samples, where the training samples include: historical trajectory samples, map information samples, predicted trajectory samples of the device, and predicted trajectory samples of the target pedestrian; Train a preset encoder-decoder model based on the historical trajectory samples, the map information samples, the predicted trajectory samples of the device, and the predicted trajectory samples of the target pedestrian to obtain the first encoder, the second encoder, and the decoder.
8. The trajectory prediction method according to any one of claims 1 to 4, characterized in that The loss functions corresponding to the first encoder, the second encoder, and the decoder include the sum of the following: The loss function of the regression problem between the predicted trajectory and the actual trajectory, the loss function of the conditional variational autoencoder between the predicted trajectory and the actual trajectory, and the loss function corresponding to the mutual information between the input sample and the latent variable.
9. A trajectory prediction device, characterized in that, Include: A first acquisition unit for acquiring historical trajectory and map information, where the historical trajectory includes the first historical trajectory of the target pedestrian and the second historical trajectory of the neighboring pedestrians adjacent to the target pedestrian; A first processing unit for encoding the map information and the first historical trajectory using the first encoder to obtain a first encoding result; A second processing unit for encoding the second historical trajectory using the first encoder to obtain a second encoding result; A second acquisition unit for acquiring the predicted trajectory of the device; A third processing unit for encoding the predicted trajectory of the device using the second encoder to obtain a third encoding result; A determination unit for determining the classification latent vector based on the core features in the second encoding result, the first encoding result, and the third encoding result; A prediction unit for decoding the classification latent vector using the decoder to obtain the predicted trajectory of the target pedestrian.
10. A trajectory prediction device, characterized in that, Comprising a processor and a memory, the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.
11. A non-volatile readable storage medium, characterized in that, A program or instruction is stored on the non-volatile readable storage medium, and when the program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
12. A computer program product, characterized in that, The computer program product is stored in a storage medium, and the computer program product is executed by at least one processor to implement the steps of the method according to any one of claims 1 to 8.
13. A robot, characterized in that, Include: The trajectory prediction device according to claim 9 or 10; And / or The non-volatile readable storage medium according to claim 11; and / or The computer program product according to claim 12.