A Transformer-based millimeter-wave radar pedestrian trajectory prediction method

By using Transformer network and pedestrian dynamic map in millimeter-wave radar pedestrian trajectory prediction, the spatial interaction relationship between pedestrians is captured, and the problem of poor trajectory prediction accuracy in multi-peer scenes in the prior art is solved, and more efficient and accurate pedestrian trajectory prediction is achieved.

CN115690157BActive Publication Date: 2025-06-27NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211371915.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-03
Publication Date
2025-06-27
Estimated Expiration
2042-11-03

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict pedestrian trajectories in multi-peer scenes, especially when the spatial interaction between pedestrians is ignored, and the sequential structure of the LSTM model makes it difficult to improve the computing speed and performance.

Method used

A millimeter-wave radar pedestrian trajectory prediction method based on Transformer is proposed. By constructing a pedestrian dynamic map, a pedestrian trajectory prediction model is established using the Transformer network structure, taking into account the historical trajectory of the target pedestrian and the historical trajectory of adjacent pedestrians, the impact of pedestrian future trajectory on a longer future moment is modeled using the future trajectory encoder.

Benefits of technology

It improves the accuracy of pedestrian trajectory prediction in multi-peer scenes, can better capture the spatial interaction between pedestrians, improves the calculation speed and performance of the model, and can more accurately predict the future pedestrian trajectory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115690157B_ABST
    Figure CN115690157B_ABST
Patent Text Reader

Abstract

A millimeter-wave radar pedestrian trajectory prediction method based on Transformer uses a millimeter-wave radar positioning method to complete the horizontal spatial coordinate positioning of pedestrians; then uses a historical trajectory tracking module based on bipartite graph matching to complete the tracking of pedestrians' historical trajectories; finally uses a Transformer-based Trajectories Prediction Model (TTPM) to complete the prediction of pedestrians' future trajectories. This method uses a neighboring historical trajectory encoder and a future trajectory encoder to handle the changes in pedestrian trajectories caused by others. TTPA effectively reduces the average displacement error and the final displacement error of pedestrian trajectory prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of radar positioning, and particularly relates to a millimeter-wave radar pedestrian trajectory prediction method based on Transformer. Background Art

[0002] With the development of sensor technology and machine learning technology, researchers have proposed many human behavior recognition systems, which usually use sensors such as cameras, wearable devices, and radio frequency devices to detect and analyze certain human behaviors, such as pedestrian positioning. Pedestrian positioning is mainly used to understand the number, location, walking trajectory, and traveling direction of pedestrian targets, and can be used in scenarios such as unmanned mobile platform navigation, smart home, building monitoring, and human-computer interaction.

[0003] Relatively common positioning methods include stereo vision positioning, lidar positioning, etc. Compared with the problems of poor privacy and high cost of lidar in stereo vision positioning based on depth cameras, millimeter-wave radars have a series of advantages such as strong environmental adaptability, penetration ability, high privacy security, low cost, and high positioning accuracy.

[0004] Pedestrian trajectory prediction in personnel positioning is a crucial topic for the movement of unmanned mobile platforms. Only by accurately predicting the moving direction of pedestrians can the unmanned mobile platform reasonably plan the driving route and avoid pedestrians in a timely and safe manner.

[0005] Early trajectory prediction algorithms mainly predicted the future trajectory of pedestrians based on kinematics. However, kinematics-based pedestrian trajectory prediction is difficult to perform long-term trajectory prediction. Since it has been gradually proven to be more effective in long-term sequence modeling tasks, the LSTM neural network has become the most commonly used model in the field of trajectory prediction. The network model based on LSTM has better effects on single-pedestrian trajectory prediction than the Kalman filter method. However, for multi-pedestrian scenarios, there are interaction relationships in the walking routes of pedestrians. Only modeling the trajectory based on the historical trajectory of a single pedestrian ignores the influence of the trajectories of other pedestrians. Therefore, this network model is difficult to perform multi-pedestrian trajectory prediction. At the same time, due to the sequential structure of LSTM, its calculation speed and performance are difficult to improve.

[0006] The Transformer structure was initially widely used in most natural language processing tasks and can rely on its powerful attention mechanism and parallelizability to improve the model calculation speed and performance. Therefore, the Transformer network has great potential in the field of pedestrian trajectory prediction. Based on the historical speed vector of pedestrians, the Transformer network is used to predict the future speed vector of pedestrians to obtain the pedestrian position. However, it only models a single pedestrian and lacks a certain degree of robustness for a wider range of pedestrian scenarios. Summary of the Invention

[0007] To better solve the problem of millimeter-wave radar pedestrian trajectory prediction, the present invention proposes a millimeter-wave radar pedestrian trajectory prediction method based on Transformer. Based on the objective condition that pedestrians usually move along routes to avoid collisions with other pedestrians, a pedestrian dynamic graph is constructed to capture the complex spatial interaction relationships among pedestrians, and a pedestrian trajectory prediction model is established using the Transformer network structure, obtaining a Transformer-based Trajectories Prediction Model (TTPM).

[0008] A millimeter-wave radar pedestrian trajectory prediction method based on Transformer, characterized by comprising the following steps:

[0009] Step 1: Use a binocular camera and a millimeter-wave radar to obtain image and echo data, and then obtain the horizontal spatial coordinate positioning of pedestrians;

[0010] Step 2: Combine the horizontal spatial coordinates of pedestrians and the corresponding positioning time as the state vectors of pedestrians;

[0011] Step 3: Use the state vectors of pedestrians to construct a bipartite graph with the best match;

[0012] Step 4: Use the KM algorithm to solve for the best match of the weighted bipartite graph, and continuously match the current latest positioning result to the trajectory, thereby obtaining the historical trajectory sequences of all pedestrians within a certain time period;

[0013] Step 5: Apply Kalman filtering to the obtained historical trajectories to eliminate the noise in the trajectories and obtain the true historical trajectories of pedestrians;

[0014] Step 6: Use the pedestrian motion state graph to determine the neighboring pedestrians who may affect the target pedestrian, and construct the historical trajectories of the neighboring pedestrians;

[0015] Step 7: Input the historical trajectories of the target pedestrian and the historical trajectories of the neighboring pedestrians into the pedestrian historical trajectory encoder and the neighboring historical trajectory encoder respectively, and add temporal information to the input pedestrian motion state through position encoding based on sine and cosine functions;

[0016] Step 8: The Transformer-based pedestrian trajectory prediction model TTPM uses the future trajectory encoder to model the influence of the position where the pedestrian will be in the future on the position at a much later future moment;

[0017] Step 9, the three trajectory encoders encode the trajectories based on the attention mechanism provided by the Transformer, and generate memory vectors at the same time. Then, the memory vectors generated by the pedestrian historical trajectory encoder and the neighboring historical trajectory encoder are concatenated to summarize the influence of the two trajectories on the pedestrian's future trajectory;

[0018] Step 10, TTPM uses pooling and multi-layer perceptrons to extract the distribution features of the data, generates the pedestrian position latent state based on the resampling technology, and finally uses the future trajectory decoder to generate the prediction of the target pedestrian's future trajectory according to the pedestrian position latent state.

[0019] Furthermore, step 1 includes the following steps:

[0020] Step 1-1, the binocular camera obtains the depth data matrix and RGB image matrix of the area to be measured; the millimeter-wave radar obtains the echo data in parallel;

[0021] Step 1-2, use the human pose estimation algorithm to calculate the pixel coordinates of the pedestrian's body key points from the RGB image matrix, then obtain the spatial Cartesian coordinates of the pixel coordinates of the key points from the depth image, and calculate the spatial horizontal coordinates (X, Y);

[0022] Step 1-3, use the AOA algorithm for the echo data, take the coordinates (X, Y) in step 1-2 as labels, input the obtained data into the convolutional neural network, and obtain an accurate radar echo model;

[0023] Step 1-4, after denoising the data in step 1-2 using the OS-CFAR algorithm, use the DBSCAN clustering algorithm to cluster the personnel reflection signal points, extract the center of each cluster, and obtain the coordinates (R, θ) of the pedestrian in the radar polar coordinate system through coordinate mapping;

[0024] Step 1-5, after coordinate transformation of (X, Y), obtain the corresponding polar coordinates, and use the KM weighted bipartite graph matching algorithm with the polar coordinates (R, θ) obtained by the millimeter-wave radar to obtain the final pedestrian horizontal space coordinates (x i , y i ).

[0025] Furthermore, step 2 includes the following steps:

[0026] Step 2-1, according to the obtained horizontal space coordinate positioning, when there are multiple pedestrian trajectories, the last positioning of each trajectory occurs at time t-1, and the horizontal space coordinates are (x i , y i ), the state vector of the pedestrian is expressed as u i = (t-1, x i , y i ), i = 1, 2,..., q;

[0027] Step 2-2: Suppose that at the current time t, a total of k horizontal spatial positioning results (x j , y j ) are generated. The state vector of each positioning result is expressed as v j = (t, x j , y j ), where j = 1, 2,..., k.

[0028] Furthermore, Step 3 includes the following steps:

[0029] Step 3-1: Use the last state vector of each recently appeared pedestrian as the vertex of sub-graph U that constitutes the bipartite graph; the state vectors of each pedestrian at the current time constitute the vertices of another sub-graph V of the bipartite graph;

[0030] Step 3-2: Add an undirected edge (u, v) for each pair of vertices in sub-graphs U and V constructed in Step 3-1, where u ∈ U and v ∈ V. The weight of each undirected edge is the Euclidean distance between vertices u and v;

[0031] Step 3-3: Considering the problem that the number of vertices in each sub-graph of the bipartite graph is different in the actual positioning process of the millimeter-wave radar, add virtual vertices to the sub-graph with fewer vertices.

[0032] Furthermore, Step 6 includes the following steps:

[0033] Step 6-1: Represent the motion state of a pedestrian as a 6-dimensional state vector, which includes the position vector of the pedestrian velocity vector and acceleration vector

[0034] Step 6-2: Construct a pedestrian motion state graph G = (V, E) to dynamically simulate the interaction relationship between a pedestrian and its adjacent pedestrians; represent each pedestrian as a vertex v ∈ V. When the distance between two pedestrians v i and v j is too close, it is considered that their travel trajectories will affect each other. Therefore, establish an undirected edge e = (v i , v j ) ∈ E. The weight of the undirected edge e is the Euclidean distance between the two pedestrians;

[0035] Step 6-3: For each pedestrian v i in the graph, merge the motion states of all pedestrians v that have undirected edges with v i with the motion state of v i . Add the 6 dimensions in the motion state vector dimension by dimension, and convert the variable-length adjacent pedestrian state sequence into a fixed-length adjacent historical trajectory X edge , and Xedge has the same dimension and size as the pedestrian historical trajectory X obs Further, in step 7, the pedestrian historical trajectory encoder receives the input of the pedestrian historical trajectory. After vector encoding, it performs multi-head attention mechanism, residual, and normalization operations together with the position encoding, and then performs feed-forward, residual, and normalization operations; the neighboring historical trajectory encoder receives the input of the neighboring historical trajectory. After vector encoding, it performs multi-head attention mechanism, residual, and normalization operations together with the position encoding, and then performs feed-forward, residual, and normalization operations; the outputs of the two encoders are concatenated to obtain the memory vector C.

[0036] Further, in step 7, the steps of adding temporal information to the pedestrian motion state are as follows:

[0037] Step 7-1, for a given trajectory sequence of length H, let t represent the time step of the motion state.

[0038] represents the position vector corresponding to the motion state at time step t. D is the embedding dimension, d is the current dimension, and PE is the function for generating the position vector which is defined as follows:

[0039]

[0040] where the frequency ω d is defined as follows:

[0041]

[0042] Then is a pair of sine and cosine for each frequency;

[0043]

[0044] Step 7-2, add the position encoding vector to the corresponding embedding vector E to obtain a new embedding vector E' with position information:

[0045]

[0046] Using the sine and cosine functions for position encoding of the trajectory sequence ensures that for two trajectory sequences with different time lengths, the distance between any two motion states is also consistent, enabling the model to have generalization ability when facing input trajectory sequences of different lengths.

[0047] ​Further, the future trajectory encoder in step 8 receives the input of the future trajectory, and through vector encoding, together with position encoding, performs multi-head attention mechanism, residual, and normalization operations, then performs feed-forward, residual, and normalization operations, and then performs multi-head attention mechanism, residual, and normalization operations again with the memory vector C input through key-value pairs, and finally performs feed-forward operation and outputs.

[0048] Further, step 8 includes the following sub-steps:

[0049] Step 8-1, the future trajectory encoder models the future trajectory probability distribution p(Y|X of the pedestrian based on the pedestrian's own historical trajectory and the historical trajectories of neighboring pedestrians obs ,X edge ).

[0050] Step 8-2, define the pedestrian's latent state as Z, and the pedestrian's future trajectory probability distribution can be defined by the following formula:

[0051] p(Y|X obs ,X edge ) = ∫p(Y|X obs ,X edge ,Z)p(Z|X obs ,X edge )dZ

[0052] Among them, p(Z|X obs ,X edge ) is a Gaussian prior distribution inferred from the pedestrian's historical trajectory X obs and the historical trajectories of neighboring pedestrians X edge .

[0053] Step 8-3, in order to approximately estimate the probability distributions of p(Z|X obs ,X edge ) and p(Y|X obs ,X edge ,Z), use an encoder-decoder network composed of four parts: a pedestrian historical trajectory encoder, a neighboring historical trajectory encoder, a future trajectory encoder, and a future trajectory decoder.

[0054] Further, the future trajectory decoder in step 10 first receives the input of the pedestrian's latent state Z, and through vector encoding, together with position encoding, performs multi-head attention mechanism, residual, and normalization operations, then performs multi-head attention mechanism, residual, and normalization operations again with the memory vector C input through key-value pairs, and finally outputs the future position of the pedestrian through feed-forward, residual, and normalization operations, and then performs a loop process from vector encoding to residual and normalization operations.

[0055] Advantages of the present invention:

[0056] (1) By using the bipartite graph matching historical trajectory scheme mentioned in this method, prior information about the number of personnel is not required, nor is it necessary to fix the number of targets to be located. It has good problem-solving ability for pedestrian trajectory tracking in scenarios with large personnel mobility.

[0057] (2) Compared with the Transformer model, TTPM uses a neighboring state sequence encoder to capture the spatial interaction relationships of pedestrians, and has higher accuracy in predicting pedestrian trajectories in multi-pedestrian scenarios.

[0058] (3) Compared with other commonly used models, TTPM incorporates the historical trajectories of target pedestrians and neighboring historical trajectories into the modeling consideration, and takes into account the impact of pedestrians' future trajectories on their more distant future moments during the prediction process, thus predicting pedestrians' future trajectories more accurately. Description of the Drawings

[0059] Figure 1 is the algorithm framework diagram of the pedestrian trajectory prediction method in the embodiment of the present invention.

[0060] Figure 2 is the pedestrian trajectory prediction diagram in the algorithm framework diagram of the embodiment of the present invention.

[0061] Figure 3 is the structural diagram of the pedestrian historical trajectory encoder and the neighboring historical trajectory encoder in the embodiment of the present invention.

[0062] Figure 4 is the structural diagram of the future trajectory encoder in the embodiment of the present invention.

[0063] Figure 5 is the structural diagram of the future trajectory decoder in the embodiment of the present invention. Detailed Embodiment

[0064] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings of the specification.

[0065] The present invention is a millimeter-wave radar pedestrian trajectory prediction method based on Transformer.

[0066] Step 1, use the personnel positioning method jointly acted by a binocular camera and a millimeter-wave radar to complete the horizontal spatial coordinate positioning of pedestrians.

[0067] The positioning method is as follows:

[0068] Step 1-1, the binocular camera obtains the depth data matrix and RGB image matrix of the area to be measured; the millimeter-wave radar obtains echo data in parallel.

[0069] Step 1-2: Use the human pose estimation algorithm to calculate the pixel coordinates of human key points from the RGB image matrix, then obtain the spatial Cartesian coordinates of the pixel coordinates of the key points from the depth image, and calculate the spatial horizontal coordinates (X, Y).

[0070] Step 1-3: Apply the AOA algorithm to the echo data, use the coordinates (X, Y) in 1-2 as labels, input the obtained data into the convolutional neural network, and obtain an accurate radar echo model.

[0071] Step 1-4: After denoising the data in Step 1-2 using the OS-CFAR algorithm, use the DBSCAN clustering algorithm to cluster the personnel reflection signal points, extract the center of each cluster, and obtain the coordinates (R, θ) of the personnel in the radar polar coordinate system through coordinate mapping.

[0072] Step 1-5: After coordinate transformation of (X, Y), obtain the corresponding polar coordinates, and use the KM weighted bipartite graph matching algorithm with the polar coordinates (R, θ) obtained by the millimeter-wave radar to obtain the final horizontal spatial coordinates (x i , y i ).

[0073] Step 2: Obtain the state vector of the relevant target according to the horizontal spatial coordinate positioning information.

[0074] The method for obtaining the state vector is as follows:

[0075] Step 2-1: According to the obtained horizontal spatial coordinate positioning, when there are multiple pedestrian trajectories and the last positioning of each trajectory occurs at time t-1, and the spatial coordinates are (x i , y i ), the state vector of the pedestrian is expressed as u i =(t-1, x i , y i ), i = 1, 2,..., q.

[0076] Step 2-2: Assume that at the current time t, a total of k horizontal spatial positioning results (x j , y j ) are generated, and the state vector of each positioning result is expressed as v j =(t, x j , y j ), j = 1, 2,..., k.

[0077] Step 3: Construct a bipartite graph with the best match using the state vector.

[0078] The construction process of the bipartite graph with the best match is as follows:

[0079] Step 3-1: The last state vector u of each recently appeared pedestriani (i = 1, 2, ..., q) serves as the vertices of the subgraph U that constitutes the bipartite graph. The state vectors of each pedestrian at the current moment form the vertices of another subgraph V of the bipartite graph. Then, an undirected edge (u, v) is added for each pair of vertices in U and V, where u ∈ U and v ∈ V, and the weight of each undirected edge is the Euclidean distance between vertices u and v.

[0080] Step 3-2: Considering the problem that the number of vertices in each subgraph of the bipartite graph is different during the actual positioning process of the millimeter-wave radar, virtual vertices are added to the subgraph with fewer vertices. Suppose a target walks out of the positioning range at the current moment or is missed due to noise or occlusion. A virtual vertex v' can be added to subgraph V. Similarly, if a new target appears, causing the number of vertices in subgraph U to be less than that in V, a new vertex u' will also be added to subgraph U. For these two cases, the weights of all undirected edges associated with the virtual vertices can be defined as d0. After adding the virtual vertices, the bipartite graph can have an optimal matching.

[0081] Step 4: Use the KM algorithm to find the optimal matching of the bipartite graph. By continuously matching the current latest positioning results to the existing or newly created trajectories, the historical trajectory sequences of all pedestrians within a certain time period can be obtained.

[0082] Step 5: Apply the Kalman filter to the historical trajectory sequences to eliminate the noise in the trajectories and restore the true historical trajectories of the pedestrians.

[0083] Step 6: After restoring the historical trajectories of the pedestrians, use the pedestrian motion state graph to determine the neighboring pedestrians who may affect the target pedestrian. Represent the motion state of a pedestrian as a 6-dimensional state vector, which includes the position vector of the pedestrian velocity vector and acceleration vector Construct a pedestrian motion state graph G = (V, E) to dynamically simulate the interaction relationship between a pedestrian and its neighboring pedestrians. Represent each pedestrian as a vertex v ∈ V. When the distance between two pedestrians v i and v j is too close, it is considered that their travel trajectories will affect each other. Therefore, an undirected edge e = (v i , v j ) ∈ E is established, and the weight of the undirected edge e is the Euclidean distance between the two pedestrians. For each pedestrian v i in the graph, merge the motion states of all pedestrians v that have undirected edges with v i . Add the 6 dimensions in the motion state vector dimension by dimension, and convert the variable-length adjacent pedestrian state sequences into a fixed-length neighboring historical trajectory X i , and X edge , and X edgeHas the same dimension and size as the pedestrian historical trajectory X obs Has the same dimension and size.

[0084] Step 7: Use the embedding vector encoding module and the position encoding module to map the input trajectory sequence into vectors that are convenient for the model to learn, and use a position encoding method based on sine and cosine functions to add temporal information to the input pedestrian motion state.

[0085] The steps to add temporal information to the pedestrian motion state are as follows:

[0086] Step 7-1: For a given trajectory sequence of length H, let t represent the time step of the motion state represent the position vector corresponding to the motion state at time step t, D is the embedding dimension, d is the current dimension, and PE is the function for generating the position vector which is defined as follows:

[0087]

[0088] where the frequency ω d is defined as follows:

[0089]

[0090] Then is a pair of sine and cosine for each frequency.

[0091]

[0092] Step 7-2: Add the position encoding vector to the corresponding embedding vector E to obtain a new embedding vector E' with position information:[[]]

[0093]

[0094] Using the sine and cosine functions to perform position encoding on the trajectory sequence ensures that for two trajectory sequences with different time lengths, the distance between any two motion states is also consistent, enabling the model to have generalization ability when facing input trajectory sequences of different lengths.

[0095] Step 8: Construct adjacent historical trajectories. Use the Transformer-based pedestrian historical trajectory encoder and adjacent historical trajectory encoder to assign different attentions to adjacent pedestrian trajectories and the target pedestrian's own historical trajectories, and model the influence on the target pedestrian's future trajectory. The two encoders are as Figure 3 shown.

[0096] The pedestrian historical trajectory encoder receives the input of the pedestrian historical trajectory. After vector encoding, it performs multi-head attention mechanism, residual, and normalization operations together with the position encoding, and then performs feed-forward, residual, and normalization operations. The neighboring historical trajectory encoder is similar to the pedestrian historical trajectory encoder and receives the input of the neighboring historical trajectory and operates through a similar process. Finally, the outputs of the two are concatenated to obtain the memory vector C.

[0097] In addition to the historical trajectory of the pedestrian itself and the historical trajectories of neighboring pedestrians, the future position of the pedestrian will also affect the position at a much later future moment. TTPM uses a future trajectory encoder to model this influencing factor.

[0098] The principle of the future trajectory encoder is as follows:

[0099] Step 8-1, the future trajectory encoder models the future trajectory of the pedestrian based on the pedestrian's own historical trajectory and the historical trajectories of neighboring pedestrians. The future trajectory of the pedestrian is defined as Y, and its probability is p(Y|X obs ,X edge ).

[0100] Step 8-2, define the pedestrian's latent state as Z. The probability distribution of the pedestrian's future trajectory can be defined by the following formula:

[0101] p(Y|X obs ,X edge ) = ∫p(Y|X obs ,X edge ,Z)p(Z|X obs ,X edge )dZ

[0102] Among them, p(Z|X obs ,X edge ) is the Gaussian prior distribution inferred from the pedestrian's historical trajectory X obs and the historical trajectories of neighboring pedestrians X edge .

[0103] Step 8-3, in order to approximately estimate the probability distributions of p(Z|X obs ,X edge ) and p(Y|X obs ,X edge ,Z), an encoder-decoder network composed of four parts: the pedestrian historical trajectory encoder, the neighboring historical trajectory encoder, the future trajectory encoder, and the future trajectory decoder is used. Among them, the structure of the future trajectory encoder is as Figure 4 , and the structure of the future trajectory decoder is as Figure 5 .

[0104] The future trajectory encoder receives the input of the future trajectory, performs multi-head attention mechanism, residual and normalization operations together with vector encoding and positional encoding, then performs feed-forward and residual and normalization operations, and then performs multi-head attention mechanism, residual and normalization operations again with the memory vector C input through key-value pairs, and finally performs feed-forward operations and outputs.

[0105] The future trajectory decoder first receives the input of the pedestrian latent state Z, performs multi-head attention mechanism, residual and normalization operations together with vector encoding and positional encoding, then performs multi-head attention mechanism, residual and normalization operations again with the memory vector C input through key-value pairs, and finally outputs the future position of the pedestrian through feed-forward and residual and normalization operations, and then performs a loop process from vector encoding to residual and normalization operations.

[0106] Step 9, The three trajectory encoders encode the trajectories based on the attention mechanism provided by Transformer, generate memory vectors at the same time, and then splice the memory vectors generated by the pedestrian historical trajectory encoder and the neighboring historical trajectory encoder to summarize the influence of the two trajectories on the future trajectory of the pedestrian.

[0107] The pedestrian historical trajectory and the neighboring historical trajectory are encoded into embedding vectors with timestamp information through vector encoding and positional encoding. The embedding vectors after positional encoding are respectively input into the pedestrian historical trajectory encoder and the neighboring historical trajectory encoder. After the encoding of the embedding vectors is completed respectively, the two output vectors are spliced into a memory vector, and the memory vector summarizes the influence of the pedestrian historical trajectory and the neighboring historical trajectory. A mean pooling layer is used to extract features from all historical trajectories. Then, a multi-layer perceptron (MLP) is used to map to a Gaussian prior probability distribution and obtain Gaussian parameters. Through the Gumbel-Softmax reparameterization trick, the sampled value Z of the latent state can be obtained. p 。

[0108] Similar to the method of obtaining the prior probability distribution, a mean pooling layer is used to extract future trajectory features from the future trajectory, and then an MLP is used to map the future trajectory features to an approximate posterior distribution q(Z|Y,X obs ,X edge ) and obtain Gaussian parameters (μ q ,σ q ). Finally, the Gumbel-Softmax reparameterization trick is used to obtain the sampled value Z of the latent state. q 。

[0109] Step 10, According to the method in the previous step, pooling and multi-layer perceptron are used to extract the distribution features of the data, and the latent state Z of the pedestrian position is generated based on the resampling technique. p and Zq , during training, reduce Z through backpropagation p and Z q 's difference, and finally use the future trajectory decoder to generate predictions of the target pedestrian's future trajectory based on the pedestrian position latent state.

[0110] The future trajectory prediction method is as follows:

[0111] Step 10-1, the input sequence of the decoder can be expressed as where the predicted future position of the pedestrian by the model 's initial value is assigned by the state feature of the last time step of the pedestrian historical state sequence X obs . Add timestamps to each f through positional encoding t to obtain an embedding vector, and input the embedding vector into the first Multi-HeadAttention and output a query vector.

[0112] Step 10-2, input the key-value pair encoding of the query vector and the memory vector C into the second Multi-HeadAttention, and then the feed-forward network outputs the future state of the next time step.

[0113] Step 10-3, by minimizing the mean square error between the predicted trajectory and the future trajectory, approximate the conditional likelihood distribution of p(Y|X obs ,X edge ) according to the posterior probability distribution of q(Z|Y,X obs ,X edge ,Z).

[0114] In the millimeter-wave radar positioning task, there are false alarm and missed detection phenomena in millimeter-wave radar positioning, which are likely to cause discontinuities in pedestrian trajectory tracking or mis-match the positioning results with pedestrian trajectories. The discontinuous or mis-matched pedestrian historical trajectories will have an unpredictable impact on the performance of the trajectory prediction model. To verify the effectiveness of this method, using the gymnastic room test data in the Njupt-radar dataset, the ability to track pedestrian trajectories in scenarios of 1 to 5 people was tested respectively, and all experimental personnel walked uniformly along a preset trajectory during the test.

[0115] To quantitatively analyze the tracking effect of this method, the lost tracking rate (the ratio of the number of lost tracking times to the total tracking times), identity switching rate (the ratio of the number of times different identity pedestrian trajectories appear on each preset trajectory to the total tracking times), and average tracking error (the error between the tracking result and the preset trajectory) of the tracking results in scenarios with different numbers of people were statistically analyzed.

[0116] The results are shown in Table 1. In the single-person scenario, the tracking effect of this method is the best, without any cases of lost tracking or identity switching, and the average tracking error is relatively low. Although more cases of lost tracking and identity switching occur as the number of people increases, within the scenario of 5 people or less, this method still maintains a relatively low lost tracking rate, identity switching rate, and average tracking error for pedestrian trajectory tracking.

[0117] Table 1 Tracking Effect

[0118] Number of people Loss tracking rate Identity switching rate Average tracking error 1 person 0% 0% 14.35 cm 2 people 1.25% 1.11% 16.68 cm 3 people 2.28% 1.55% 19.62 cm 4 people 3.21% 2.12% 21.37 cm 5 people 4.17% 2.68% 24.74 cm

[0119] To quantitatively analyze the performance of this method, the trajectory prediction evaluation metrics include:

[0120] Mean Average Displacement (MAD): Within the next T time steps, it is the average of the Euclidean distance errors between the predicted position and the true position of the pedestrian at each time step. The mathematical definition of MAD for the i-th person is as follows:

[0121]

[0122] where obs is the current time step, n is the total number of pedestrians, T is the number of prediction time steps, is the true position of the i-th pedestrian at time t, is the predicted position of the i-th pedestrian at time t.

[0123] Final Average Displacement (FAD): In the last time step, it is the average of the Euclidean distance between the predicted trajectory and the true trajectory of the pedestrian. The mathematical definition of FAD for the i-th person is as follows:

[0124]

[0125] The dataset used for testing is Njupt-radar. To enhance the dataset for fully training the TTPM network model, after obtaining the positioning results for each data frame in Njupt-radar using the millimeter-wave radar personnel positioning method, the positioning results are respectively subjected to translation and rotation coordinate transformations, and then the pedestrian positioning results are made into historical trajectory sequences.

[0126] The Njupt-radar dataset contains 6 different scenarios. After dataset augmentation, a total of 133,758 pedestrian trajectory sequences are obtained. These trajectory sequences on average contain 154 consecutive positioning results of pedestrians, and the time interval between each positioning result is 0.2 seconds (Δt = 0.2). In this experimental verification, 70% of the pedestrian trajectory sequences are used for training, and 30% of the data is used for testing.

[0127] The results of the MAD and FAD metrics of the TTPM proposed in the present invention based on the prediction time on the Njupt-radar dataset are shown in Table 2.

[0128] Table 2 MAD and FAD metrics of TTPM on the Njupt-radar dataset

[0129]

[0130] In this experiment, the walking trajectories of pedestrians in the past 4 seconds were used to predict the trajectories of pedestrians in the future 1 second, 2 seconds, and 3 seconds respectively. The experimental results show that the TTPM proposed in the present invention has lower MDE and FDE errors. Although both the MAD and FAD errors increase with the increase of the prediction time, TTPM shows a slower error growth rate.

[0131] Compared with the Transformer model, TTPM uses a neighboring state sequence encoder to capture the spatial interaction relationship of pedestrians, so it has a higher accuracy in predicting pedestrian trajectories in multi-pedestrian scenarios.

[0132] To verify the robustness of TTPM, in addition to the self-built Njupt-radar millimeter-wave radar dataset mentioned above, this experiment also tests the TTPM algorithm based on the following public datasets.

[0133] GC dataset: 6001 RGB images were sampled from the surveillance video of New York's Grand Central Terminal, about one hour of video images, with a frame interval of 0.8 seconds, containing the walking trajectories of 12,684 manually marked pedestrians. The pedestrian coordinates are based on the RGB image pixel coordinate system.

[0134] ETH dataset: RGB images of two scenarios (ETH scenario and Hotel scenario), with a total of 750 annotated trajectories of different pedestrians, and a frame interval of 0.4 seconds.

[0135] UCY dataset: It contains two scenarios, ZARA scenario and UCY scenario. The ZARA scenario contains two parts, ZARA-01 and ZARA-02, with a total of 786 annotated trajectories of different pedestrians and a frame interval of 0.4 seconds.

[0136] Since the sampling rates of pedestrian trajectories in each dataset are different, in order to uniformly use the historical 4-second trajectory to predict the future 4-second trajectory, the settings of each dataset in this experiment are also different. For the GC dataset, this experiment predicts the future 5-frame trajectory based on the historical 5-frame trajectory. For the ETH and UCY datasets, this experiment predicts the future 10-frame trajectory based on the historical 10-frame trajectory. Compared with the campus scene of Njupt-radar, the datasets of the GC dataset, ETH dataset, and UCY dataset are collected in squares or intersections, where pedestrians are more dense and the spatial interactions between pedestrians are more frequent. The test results are shown in Table 3.

[0137] Table 3 MAD and FAD indicators of TTPM in public datasets

[0138]

[0139] Experimental results show that TTPM can well model the spatial position relationship of pedestrians through the neighboring state sequence encoder and the future state sequence encoder, and has excellent performance in predicting the future trajectory of pedestrians in scenes with dense pedestrians.

[0140] The above description is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiment. Any equivalent modifications or changes made by ordinary technicians in this field based on the contents disclosed by the present invention should be included in the protection scope recorded in the claims.

Claims

1. A millimeter-wave radar pedestrian trajectory prediction method based on Transformer, characterized in that: It includes the following steps: Step 1: Using a binocular camera and a millimeter-wave radar, obtain image and echo data, and then obtain the horizontal spatial coordinate positioning of pedestrians; Step 2: Combine the horizontal spatial coordinates of pedestrians and the corresponding positioning time as the state vector of pedestrians; Step 3: Use the pedestrian state vector to construct a bipartite graph with the best match; Step 4: Use the KM algorithm to solve the best match of the weighted bipartite graph, and continuously match the current latest positioning result to the trajectory, so as to obtain the historical trajectory sequence of all pedestrians within a certain period of time; Step 5: Apply Kalman filtering to the obtained historical trajectory to eliminate the noise in the trajectory and obtain the real historical trajectory of pedestrians; Step 6: Use the pedestrian motion state graph to determine the neighboring pedestrians who may affect the target pedestrian and construct the historical trajectory of neighboring pedestrians; Step 7: Input the historical trajectory of the target pedestrian and the historical trajectory of neighboring pedestrians into the pedestrian historical trajectory encoder and the neighboring historical trajectory encoder respectively, and add temporal information to the input pedestrian motion state through position encoding based on sine and cosine functions; Step 8: The pedestrian trajectory prediction model TTPM based on Transformer uses the future trajectory encoder to model the influence of the position where the pedestrian will be in the future on the position at a much later future moment; Step 9: The three trajectory encoders encode the trajectories based on the attention mechanism provided by Transformer, and generate memory vectors at the same time. Then, splice the memory vectors generated by the pedestrian historical trajectory encoder and the neighboring historical trajectory encoder to summarize the influence of the two trajectories on the future trajectory of pedestrians; Step 10: TTPM uses pooling and a multi-layer perceptron to extract the distribution characteristics of the data, generates the latent state of the pedestrian position based on the resampling technique, and finally uses the future trajectory decoder to generate the prediction of the future trajectory of the target pedestrian according to the latent state of the pedestrian position.

2. The millimeter-wave radar pedestrian trajectory prediction method based on Transformer according to claim 1, wherein: Step 1 includes the following steps: Step 1-1: The binocular camera obtains the depth data matrix and RGB image matrix of the area to be measured; the millimeter-wave radar obtains echo data in parallel; Step 1-2: Use the human pose estimation algorithm to calculate the pixel coordinates of the key points of the pedestrian body from the RGB image matrix, and then obtain the spatial Cartesian coordinates of the pixel coordinates of the key points from the depth image, and calculate the spatial horizontal coordinates (X, Y); Step 1-3: Apply the AOA algorithm to the echo data, use the coordinates (X, Y) in Step 1-2 as labels, input the obtained data into a convolutional neural network, and obtain an accurate radar echo model; Step 1-4: After denoising the data in Step 1-2 using the OS-CFAR algorithm, use the DBSCAN clustering algorithm to cluster the personnel reflection signal points, extract the center of each cluster, and obtain the coordinates (R, θ) of the pedestrians in the radar polar coordinate system through coordinate mapping; Steps 1-5, after transforming (X,Y) through coordinate transformation to obtain the corresponding polar coordinates, use the KM weighted bipartite graph matching algorithm with the polar coordinates (R,θ) obtained from the millimeter-wave radar to obtain the final horizontal spatial coordinates (x i ,y i ) 3. A millimeter-wave radar pedestrian trajectory prediction method based on Transformer according to claim 1, characterized in that: Step 2 includes the following steps: Step 2-1, locate according to the obtained horizontal space coordinates. When there are multiple pedestrian trajectories and the last positioning of each trajectory occurs at time t-1, the horizontal space coordinates are (x i , y i ), and the state vector of the pedestrian is expressed as u i = (t-1, x i , y i ), where i = 1, 2, …, q; Step 2-2, assume that at the current t moment, a total of k horizontal space positioning results (x j , y j ) are generated, and the state vector of each positioning result is expressed as v j = (t, x j , y j ), where j = 1, 2,..., k.

4. A millimeter-wave radar pedestrian trajectory prediction method based on Transformer according to claim 1, characterized in that: Step 3 includes the following steps: Step 3-1: Take the last state vector of each pedestrian who has appeared recently as the vertex of subgraph U that constitutes the bipartite graph; the state vectors of each pedestrian at the current moment constitute the vertices of another subgraph V of the bipartite graph; Step 3-2: Add an undirected edge (u, v) for each pair of vertices in the subgraphs U and V constructed in Step 3-1, where u ∈ U and v ∈ V. The weight of each undirected edge is the Euclidean distance between vertices u and v. Step 3-3: Considering the problem that the number of vertices in each subgraph of the bipartite graph is different in the actual positioning process of the millimeter-wave radar, add virtual vertices to the subgraph with fewer vertices.

5. A millimeter-wave radar pedestrian trajectory prediction method based on Transformer according to claim 1, characterized in that: Step 6 includes the following steps: Step 6-1, represent the motion state of the pedestrian as a 6-dimensional state vector, including the position vector of the pedestrian velocity vector and acceleration vector Step 6-2, construct a pedestrian motion state graph G=(V, E) to dynamically simulate the interaction relationship between a pedestrian and its adjacent pedestrians; represent each pedestrian as a vertex v∈V, when the distance between two pedestrians v i and v j is too close, it is considered that it will affect each other's travel trajectories. Therefore, an undirected edge e=(v i , v j )∈E is established, and the weight of the undirected edge e is the Euclidean distance between the two pedestrians; Step 6-3, for each pedestrian v in the figure i , combine the motion states of all pedestrians v that have undirected edges with v i , add the six dimensions in the motion state vector dimension by dimension, and convert the variable-length adjacent pedestrian state sequences into a fixed-length neighboring historical trajectory X i , and X edge has the same dimension and size as the pedestrian historical trajectory X edge . obs ​ 6. A method for predicting pedestrian trajectories based on a Transformer for millimeter-wave radar according to claim 1, characterized in that: In Step 7, the pedestrian historical trajectory encoder receives the input of the pedestrian historical trajectory. After vector encoding and position encoding, it performs the multi-head attention mechanism, as well as residual and normalization operations, and then performs the feed-forward, as well as residual and normalization operations. The adjacent historical trajectory encoder receives the input of the adjacent historical trajectory. After vector encoding and position encoding, it performs the multi-head attention mechanism, as well as residual and normalization operations, and then performs the feed-forward, as well as residual and normalization operations. The outputs of the two encoders are concatenated to obtain the memory vector C.

7. A millimeter-wave radar pedestrian trajectory prediction method based on Transformer according to claim 1, characterized in that: In Step 7, the steps to add temporal information to the pedestrian motion state are as follows: Step 7-1. For a given trajectory sequence of length H, let t represent the time step of the motion state. represents the position vector corresponding to the motion state at time step t. D is the embedding dimension, d is the current dimension, and PE is a function for generating the position vector which is defined as follows: Among them, the frequency ω d is defined as follows: Then is a pair of sine and cosine for each frequency; Step 7-2, add the positional encoding vector to the corresponding embedding vector E to obtain a new embedding vector E' with positional information: Using the sine and cosine functions to perform position encoding on the trajectory sequence ensures that for two trajectory sequences with different time lengths, the distance between any two motion states is also consistent, enabling the model to have generalization ability when facing input trajectory sequences of different lengths.

8. A method for predicting pedestrian trajectories based on a Transformer for millimeter-wave radar according to claim 1, characterized in that: In Step 8, the future trajectory encoder receives the input of the future trajectory. Through vector encoding, it performs the multi-head attention mechanism, as well as residual and normalization operations together with the position encoding, then performs the feed-forward, as well as residual and normalization operations, and then performs the multi-head attention mechanism, as well as residual and normalization operations again with the memory vector C input through the key-value pair, and finally performs the feed-forward operation and outputs.

9. A millimeter-wave radar pedestrian trajectory prediction method based on Transformer according to claim 1, characterized in that: Step 8 includes the following sub-steps: Step 8-1, the future trajectory encoder models the future trajectory probability distribution p(Y|X obs ,X edge ) of a pedestrian based on the pedestrian's own historical trajectory and the historical trajectories of neighboring pedestrians. Step 8-2: Define the pedestrian latent state as Z. The probability distribution of the pedestrian future trajectory can be defined by the following formula: p(Y|X obs ,X edge ) = ∫p(Y|X obs ,X edge ,Z)p(Z|X obs ,X edge )dZ Among them, p(Z|X obs ,X edge ) is a Gaussian prior distribution inferred from the pedestrian historical trajectory X obs and the neighboring pedestrian historical trajectory X edge . Step 8-3, to approximately estimate the probability distributions of p(Z|X obs ,X edge ) and p(Y|X obs ,X edge ,Z), an encoder-decoder network composed of four parts, namely, a pedestrian historical trajectory encoder, a neighboring historical trajectory encoder, a future trajectory encoder, and a future trajectory decoder, is used.

10. A method for predicting pedestrian trajectories based on a Transformer for millimeter-wave radar according to claim 1, characterized in that: In Step 10, the future trajectory decoder first receives the input of the pedestrian latent state Z. Through vector encoding, it performs the multi-head attention mechanism, as well as residual and normalization operations together with the position encoding, and then performs the multi-head attention mechanism, as well as residual and normalization operations again with the memory vector C input through the key-value pair. Finally, it outputs the future position of the pedestrian through the feed-forward, as well as residual and normalization operations, and then performs the loop process from vector encoding to residual and normalization operations.

Citation Information

Patent Citations

  • Transform and graph convolutional network-based pedestrian trajectory prediction method

    CN114757975A

  • Pedestrian trajectory prediction method and system

    CN114898550A