A pedestrian trajectory prediction method and system based on complementary attention

By adopting a complementary attention mechanism in pedestrian trajectory prediction, the space-time characteristics of general and special situations of pedestrian movement are extracted, and the problems of difficult to predict pedestrian movement in the prior art are solved, and more accurate trajectory prediction and higher application security are achieved.

CN114429490BActive Publication Date: 2025-05-13XI AN JIAOTONG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210102800.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-27
Publication Date
2025-05-13
Estimated Expiration
2042-01-27

AI Technical Summary

Technical Problem

The existing pedestrian trajectory prediction methods are difficult to predict accurate results when facing pedestrian movements, mainly because the general attention mechanism pays too much attention to the general pattern of data and cannot effectively capture pedestrian movements.

Method used

The pedestrian trajectory prediction method based on complementary attention is adopted, and the spatial and temporal movement characteristics of the general and special situations of pedestrian movement are extracted respectively through the spatial interaction neural network and the temporal motion trend neural network, and the positive and negative characteristics are fused through the adaptive weight of the gated network to generate the distribution of the future trajectory end points.

Benefits of technology

When facing the special situation of pedestrian movement, more accurate prediction results can be obtained, which improves the application safety of trajectory prediction technology, and reveals the generality and particularity of pedestrian movement decisions through positive and negative attention mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114429490B_ABST
    Figure CN114429490B_ABST
Patent Text Reader

Abstract

The present invention discloses a pedestrian trajectory prediction method and system based on complementary attention, the method comprising the following steps: obtaining a pedestrian observation trajectory sequence to be predicted; inputting the pedestrian observation trajectory sequence into a pre-trained pedestrian trajectory prediction model, and outputting a trajectory prediction result; wherein the pedestrian trajectory prediction model comprises: a spatial interaction neural network, for inputting the pedestrian observation trajectory sequence, and outputting spatial trajectory features; a temporal motion trend neural network, for inputting the spatial trajectory features, and outputting spatiotemporal trajectory features; a first multi-layer perceptron network, for inputting the spatiotemporal trajectory features, and outputting a terminal distribution as a trajectory prediction result. The pedestrian trajectory prediction method or system provided by the present invention can obtain more accurate prediction results when facing special situations of pedestrian motion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision technology, and in particular relates to a pedestrian trajectory prediction method and system based on complementary attention. Background Art

[0002] Pedestrian trajectory prediction is crucial for autonomous driving. Accurate prediction of pedestrians' future trajectories has a profound impact on the efficiency and safety of autonomous driving technology.

[0003] Due to the high randomness of pedestrian motion, that is, different pedestrians often have completely different future trajectories in similar observable trajectories and social environments, the existing pedestrian trajectory prediction methods are difficult to predict accurate results when faced with special situations of pedestrian motion. Specifically, the existing technology uses data-driven methods, especially attention mechanisms, to extract the spatial interaction features and motion trend features of pedestrians. However, the general attention mechanism in pedestrian trajectory prediction focuses too much on the general pattern of the data, which makes the general attention mechanism technology only capture the general situation of pedestrian motion, and it is difficult to predict accurate results when faced with special situations of pedestrian motion.

[0004] In summary, a new pedestrian trajectory prediction method or system is urgently needed. Summary of the invention

[0005] The purpose of the present invention is to provide a pedestrian trajectory prediction method and system based on complementary attention to solve one or more of the above-mentioned technical problems. The pedestrian trajectory prediction method or system provided by the present invention can obtain more accurate prediction results when facing special situations of pedestrian movement.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] The present invention provides a pedestrian trajectory prediction method based on complementary attention, comprising the following steps:

[0008] Obtaining a pedestrian observation trajectory sequence to be predicted;

[0009] Inputting the pedestrian observation trajectory sequence into a pre-trained pedestrian trajectory prediction model and outputting a trajectory prediction result;

[0010] Wherein, the pedestrian trajectory prediction model includes:

[0011] A spatial interactive neural network, used for inputting the pedestrian observation trajectory sequence and outputting spatial trajectory features;

[0012] A time-series motion trend neural network, used to input the spatial trajectory features and output the spatiotemporal trajectory features;

[0013] The first multi-layer perceptron network is used to input the spatiotemporal trajectory features and output the endpoint distribution as the trajectory prediction result.

[0014] A further improvement of the method of the present invention is that the spatial interactive neural network comprises:

[0015] A first discriminator, configured to input the pedestrian observation trajectory sequence and output a general case matrix and a special case matrix of spatial interaction;

[0016] A first single-head self-attention network is used to input the pedestrian observation trajectory sequence and the general case matrix of spatial interaction, and output the general case feature of spatial interaction;

[0017] A second single-head self-attention network is used to input the pedestrian observation trajectory sequence and the special case matrix of spatial interaction, and output the special case feature of spatial interaction;

[0018] The first feature fusion network is used to input the general situation feature of the spatial interaction and the special situation feature of the spatial interaction, and output the spatial trajectory feature.

[0019] A further improvement of the method of the present invention is that the first discriminator comprises:

[0020] A first multi-head self-attention network is used to input the pedestrian observation trajectory sequence and output a multi-head attention matrix;

[0021] A first convolutional neural network, used to input the multi-head attention matrix and output a single-channel self-attention matrix;

[0022] The first symbolic function is used to input the single-channel self-attention matrix processed by the Sigmoid function and output the general case matrix and the special case matrix of spatial interaction.

[0023] A further improvement of the method of the present invention is that the time-series motion trend neural network comprises:

[0024] A second discriminator, used for inputting the spatial trajectory features and outputting a general case matrix and a special case matrix of the temporal motion trend;

[0025] A third single-head self-attention network is used to input the general situation matrix of the spatial trajectory characteristics and the temporal motion trend, and output the general situation characteristics of the spatiotemporal motion;

[0026] a fourth single-head self-attention network, configured to input the special case matrix of the spatial trajectory features and the temporal motion trend, and output the special case features of the spatiotemporal motion;

[0027] The second feature fusion network is used to input the general situation features of the space-time motion and the special situation features of the space-time motion, and output the space-time trajectory features.

[0028] A further improvement of the method of the present invention is that the second discriminator comprises:

[0029] A second multi-head self-attention network is used to input the spatial trajectory features and output a multi-head attention matrix;

[0030] A second convolutional neural network, used to input the multi-head attention matrix and output a single-channel self-attention matrix;

[0031] The second symbolic function is used to input the single-channel self-attention matrix processed by the Sigmoid function, and output the general case matrix and the special case matrix of the temporal motion trend.

[0032] A further improvement of the method of the present invention is that both the first feature fusion network and the second feature fusion network are gated networks.

[0033] A further improvement of the method of the present invention is that the endpoint distribution is a mixed Gaussian endpoint distribution.

[0034] A further improvement of the method of the present invention is that the step of obtaining the pre-trained pedestrian trajectory prediction model comprises:

[0035] Obtain a sample training set, wherein each sample includes a pedestrian sample observation trajectory sequence and a future true trajectory;

[0036] Inputting the observed trajectory sequence of pedestrian samples in the selected sample into the pedestrian trajectory prediction model, and outputting the predicted terminal distribution of the future trajectory; calculating the maximum likelihood estimation loss between the terminal distribution of the predicted future trajectory and the terminal of the future real trajectory in the selected sample, and obtaining the loss value;

[0037] Based on the loss value, the model parameters are updated, and after reaching the preset convergence condition, the pre-trained pedestrian trajectory prediction model is obtained.

[0038] A further improvement of the method of the present invention is that the pedestrian trajectory prediction model further includes:

[0039] The second multi-layer perceptron is used to input the pedestrian observation trajectory sequence and the predicted trajectory endpoint obtained by sampling from the endpoint distribution, and output a complete future predicted trajectory.

[0040] The present invention provides a pedestrian trajectory prediction system based on complementary attention, comprising:

[0041] A pedestrian observation trajectory sequence acquisition module is used to obtain a pedestrian observation trajectory sequence to be predicted;

[0042] A trajectory prediction result acquisition module is used to input the pedestrian observation trajectory sequence into a pre-trained pedestrian trajectory prediction model and output a trajectory prediction result;

[0043] Wherein, the pedestrian trajectory prediction model includes:

[0044] A spatial interactive neural network, used for inputting the pedestrian observation trajectory sequence and outputting spatial trajectory features;

[0045] A time-series motion trend neural network, used to input the spatial trajectory features and output the spatiotemporal trajectory features;

[0046] The first multi-layer perceptron network is used to input the spatiotemporal trajectory features and output the endpoint distribution as the trajectory prediction result.

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] The present invention provides a prediction method for pedestrian trajectory prediction that can capture the special conditions of pedestrian motion. It obtains a pair of complementary attention matrices for both spatial interaction and motion trend to guide the extraction of spatiotemporal features of general and special cases, respectively. The above method not only provides more accurate prediction results for pedestrian trajectory prediction, but also reveals the generality and particularity of pedestrian motion decisions through positive and negative attention mechanisms, so that trajectory prediction technology can handle more situations and improve the safety of applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art; obviously, the drawings described below are some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0050] Figure 1 is a flowchart of a pedestrian trajectory prediction method based on complementary attention according to an embodiment of the present invention;

[0051] Figure 2 is a schematic diagram of obtaining spatial trajectory features in an embodiment of the present invention;

[0052] Figure 3 It is a schematic diagram of visualization results of the method of the embodiment of the present invention under the ETH and UCY data sets in the embodiment of the present invention. DETAILED DESCRIPTION

[0053] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0054] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0055] The present invention is further described in detail below in conjunction with the accompanying drawings:

[0056] See also Figure 1 and Figure 2 , a pedestrian trajectory prediction method based on complementary attention according to an embodiment of the present invention comprises the following steps:

[0057] Obtaining a pedestrian observation trajectory sequence to be predicted;

[0058] Inputting the pedestrian observation trajectory sequence into a pre-trained pedestrian trajectory prediction model and outputting a trajectory prediction result;

[0059] The pedestrian trajectory prediction model includes:

[0060] (1) a spatial interactive neural network, which is used to input the pedestrian observation trajectory sequence and output spatial trajectory features;

[0061] The spatial interaction neural network comprises:

[0062] 1) A first discriminator, configured to input the pedestrian observation trajectory sequence and output a general case matrix and a special case matrix of spatial interaction;

[0063] The first discriminator comprises:

[0064] ① A first multi-head self-attention network, used to input the pedestrian observation trajectory sequence and output a multi-head attention matrix;

[0065] ② A first convolutional neural network, used to input the multi-head attention matrix and output a single-channel self-attention matrix;

[0066] ③ The first symbolic function is used to input the single-channel self-attention matrix processed by the Sigmoid function and output the general case matrix and special case matrix of spatial interaction;

[0067] 2) a first single-head self-attention network, configured to input the pedestrian observation trajectory sequence and a general case matrix of spatial interaction, and output a general case feature of spatial interaction;

[0068] 3) a second single-head self-attention network, configured to input the pedestrian observation trajectory sequence and the special case matrix of spatial interaction, and output the special case features of spatial interaction;

[0069] 4) A first feature fusion network is used to input the general situation feature of the spatial interaction and the special situation feature of the spatial interaction, and output the spatial trajectory feature.

[0070] (2) a temporal motion trend neural network, which is used to input the spatial trajectory features and output the spatiotemporal trajectory features;

[0071] The time series motion trend neural network comprises:

[0072] 1) A second discriminator, configured to input the spatial trajectory features and output a general case matrix and a special case matrix of the temporal motion trend;

[0073] ① A second multi-head self-attention network, used to input the spatial trajectory features and output a multi-head attention matrix;

[0074] ② A second convolutional neural network, used to input the multi-head attention matrix and output a single-channel self-attention matrix;

[0075] ③ The second symbolic function is used to input the single-channel self-attention matrix processed by the Sigmoid function and output the general case matrix and the special case matrix of the temporal motion trend;

[0076] 2) a third single-head self-attention network, used to input the spatial trajectory features and the general case matrix of the temporal motion trend, and output the general case features of the temporal motion;

[0077] 3) a fourth single-head self-attention network, configured to input the spatial trajectory features and the special case matrix of the temporal motion trend, and output the special case features of the temporal motion;

[0078] 4) A second feature fusion network is used to input the general situation features of the temporal motion and the special situation features of the temporal motion, and output the spatiotemporal trajectory features.

[0079] (3) A first multi-layer perceptron network, used to input the spatiotemporal trajectory features and output an endpoint distribution as a trajectory prediction result.

[0080] Optionally, in this embodiment of the present invention, both the first feature fusion network and the second feature fusion network are gated networks.

[0081] In a preferred embodiment of the present invention, the endpoint distribution is a mixed Gaussian endpoint distribution.

[0082] Preferably, in an embodiment of the present invention, the pedestrian trajectory prediction model further includes:

[0083] The second multi-layer perceptron is used to input the pedestrian observation trajectory sequence and the predicted trajectory endpoint obtained by sampling from the endpoint distribution, and output a complete future predicted trajectory.

[0084] The method provided by the embodiment of the present invention obtains a sequence of observed trajectories of pedestrians based on a video of pedestrian motion; based on the coordinate sequence of the observed trajectories of pedestrians, a pre-trained neural network is used to discriminate the trajectories of pedestrians in the current space-time scene to obtain the general and special cases of spatial interaction of pedestrians in the current scene; the pre-trained two-way neural network extracts features of the general and special cases of spatial interaction respectively and finally performs adaptive weight fusion through a pre-trained gated neural network to obtain spatial trajectory features; based on the obtained spatial trajectory features, a pre-trained neural network is used to discriminate the trajectories of pedestrians in the current space-time scene to obtain the current scene The general and special cases of the temporal motion trends of pedestrians are analyzed; the pre-trained two-way neural network extracts features of the general and special cases of the temporal motion trends respectively and finally performs adaptive weight fusion through the pre-trained gated neural network to obtain the spatiotemporal fusion trajectory features; for the obtained spatiotemporal fusion trajectory features, the pre-trained neural network is used to optimize the Gaussian mixture distribution parameters of the predicted trajectory endpoint, and the multimodal predicted trajectory endpoint is sampled from the obtained Gaussian mixture distribution; based on the observable trajectory and the multimodal endpoint, other multimodal predicted trajectory points are generated through the pre-trained neural network to obtain the final multimodal predicted trajectory.

[0085] In the technical solution of the embodiment of the present invention, positive and negative features are fused through adaptive weights and used to generate the distribution of future trajectory endpoints, further exploring the potential of neural networks in pedestrian trajectory prediction, allowing all parameters to be learned through the network as much as possible, enhancing the learning ability of the network, and reducing the cost of artificial features; by sampling the endpoint and further predicting based on the endpoint to obtain a complete multimodal prediction result, more accurate prediction results can be obtained when facing special circumstances of pedestrian movement. Specific embodiments

[0087] In the embodiment of the present invention, the step of obtaining the spatial trajectory features based on the observed trajectory sequence coordinates of the pedestrian specifically includes:

[0088] 1) Data preprocessing, including: for N pedestrians, the observable trajectories within a period of T consecutive moments are expressed as (X, Y) two-dimensional coordinates for each person at each moment; processing them into a tensor F that can be batch-operated input , the dimension is T*N*2; F input Represents all observable inputs.

[0089] 2) Spatial interaction judgment, including: using a multi-head self-attention mechanism, for the i-th head among the H heads, first use two multi-layer perceptrons to input F input Calculate the key vector K i and the query vector Q i , and then calculate the asymmetric attention score matrix A for each head i ; A i The calculation process is as follows: Where i represents the i-th attention head, K i represents the key vector, Q i represents the query vector, d k represents the latitude of the key vector, A i It represents the calculated attention score matrix of the i-th head, with a latitude of N*N. The value of the r-th row and c-th column in the matrix represents the influence of pedestrian r on pedestrian c.

[0090] For the obtained multi-head attention matrix A i , i=(1,2,...,H), the latitude is N*N*H, and a 1×1 convolutional neural network is used to map the multi-head channel H into a 1-channel matrix S. Then the Sigmoid function is used to map the S matrix values ​​to between 0 and 1 to indicate the strength of the correlation between the sequence elements, that is, the strength of the interaction between pedestrians at a certain moment.

[0091] Next, we use the symbolic function to operate S. The specific operation of the symbolic function is to map the values ​​greater than 0.5 in S to 1 and the values ​​less than 0.5 to 0 to obtain the matrix M normal And negate it to get M inverse .

[0092] A pair of complementary attention relationship matrices M normal and M inverse They represent the pedestrian interaction relationship in general and special cases respectively. The value of the rth row and cth column in the two matrices indicates whether pedestrian r has an impact on pedestrian c.

[0093] 3) A dual-path network extracts pedestrian spatial interaction features in general and special cases, including: the two paths of the dual-path network are the normal attention path and the inverse attention path, and both paths use the same single-head self-attention neural network structure.

[0094] For the normal attention path, given the input F input and M obtained in step 2) to represent the general case of pedestrian spatial interaction normal First, two multi-layer perceptrons are used to input F input The key vector K and query vector Q are calculated separately, and then the single-head self-attention matrix is ​​calculated The calculation process is as follows: In the formula, K represents the key vector Q represents the query vector, d k represents the latitude of the key vector, Represents the calculated attention score matrix; through M normal right After constraining and normalizing, we get Represents the normal attention matrix, and the calculation process is as follows: In the formula, ⊙ represents the multiplication of the elements at the corresponding positions of the matrix. Using a multilayer perceptron and input F input The calculated V represents the value vector, and Perform matrix multiplication to obtain the spatial features of the normal attention path

[0095] For the inverse attention path, given the input F input and M obtained in step 2) for the special case of pedestrian spatial interaction inverse First, two multi-layer perceptrons are used to input F input The key vector K and query vector Q are calculated separately, and then the single-head self-attention matrix is ​​calculated The calculation process is as follows: In the formula, K represents the key vector Q represents the query vector, d k represents the latitude of the key vector, Represents the calculated attention score matrix; through M inverse right After constraining and normalizing, we get Represents the inverse attention matrix, and the calculation process is as follows: In the formula, ⊙ represents the multiplication of the corresponding elements of the matrix. Using a multilayer perceptron and input F input The calculated V represents the value vector, and Perform matrix multiplication to obtain the spatial features of the inverse attention path

[0096] 4) Adaptive weight fusion of dual-path spatial features by the gated network, including: dual-path features for general and special cases of pedestrian spatial interaction obtained in sub-step 3) and They are used as inputs and sent to a multi-layer perceptron respectively, and then the corresponding weights are obtained through the Sigmoid function. and They have the same dimension as the input features, and then use the Softmax function to adjust the two weights. and Normalize and and After passing through a multi-layer perceptron layer, the weighted sum is performed according to the normalized weights to obtain the final extracted spatial feature F spatial .

[0097] In the embodiment of the present invention, the step of obtaining the spatiotemporal fusion trajectory features based on the spatial trajectory features of the pedestrian specifically includes:

[0098] 1) Time series motion trend determination, including: given spatial features F spatial First, exchange the time dimension T and spatial dimension N of the features, and then use the multi-head self-attention mechanism. For the i-th head among the H heads, first use two multi-layer perceptrons to input F spatial Calculate the key vector K i and the query vector Q i , and then calculate the asymmetric attention score matrix A for each head i ; A i The calculation process is as follows: Where i represents the i-th attention head, K i represents the key vector, Q i represents the query vector, d k represents the latitude of the key vector, A i It represents the calculated attention score matrix of the i-th head, with latitude T*T. The value of the r-th row and c-th column in the matrix represents the influence of the pedestrian's trajectory at the r-th moment on the trajectory at the c-th moment. For the obtained multi-head attention matrix A i , i=(1,2,...,H), the latitude is T*T*H, and a 1×1 convolutional neural network is used to map the multi-channel H into a 1-channel matrix T. Then the Sigmoid function is used to map the T matrix values ​​between 0 and 1 to indicate the strength of the correlation between the sequence elements, that is, the strength of the influence between the trajectories of a pedestrian at different times.

[0099] Next, we use the symbolic function to operate T. The specific operation of the symbolic function is to map the values ​​in T greater than 0.5 to 1 and the values ​​less than 0.5 to 0 to obtain the matrix M normalAnd negate it to get M inverse .

[0100] A pair of complementary attention relationship matrices M normal and M inverse They represent the temporal motion trend relationship of pedestrians in general and special cases respectively. The value of the rth row and cth column in the two matrices represents the influence of the pedestrian's trajectory at the rth moment on the trajectory at the cth moment.

[0101] 2) A dual-path network extracts the temporal motion trend features of pedestrians in general and special cases, including: the two paths of the dual-path network are the normal attention path and the inverse attention path, and both paths use the same single-head self-attention neural network structure.

[0102] For the normal attention path, given the input F spatial and M obtained in step 2) to represent the general situation of pedestrian temporal motion trend normal First, two multi-layer perceptrons are used to input F spatial The key vector K and query vector Q are calculated separately, and then the single-head self-attention matrix is ​​calculated The calculation process is as follows: In the formula, K represents the key vector Q represents the query vector, d k represents the latitude of the key vector, Represents the calculated attention score matrix; through M normal right After constraining and normalizing, we get Represents the normal attention matrix, and the calculation process is as follows: M normal );where ⊙ represents the multiplication of the corresponding elements of the matrix. Using a multilayer perceptron and input F spatial The calculated V represents the value vector, and Perform matrix multiplication to obtain the spatiotemporal features of the normal attention path

[0103] For the inverse attention path, given the input F spatial and M obtained in step 2) to represent the special case of pedestrian temporal motion trend inverse First, two multi-layer perceptrons are used to input F spatial The key vector K and query vector Q are calculated separately, and then the single-head self-attention matrix is ​​calculated The calculation process is as follows: In the formula, K represents the key vector Q represents the query vector, d k represents the latitude of the key vector, Represents the calculated attention score matrix; through M inverse right After constraining and normalizing, we get Represents the inverse attention matrix, and the calculation process is as follows: M inverse );where ⊙ represents the multiplication of the corresponding elements of the matrix. Using a multilayer perceptron and input F spatial The calculated V represents the value vector, and Perform matrix multiplication to obtain the spatiotemporal features of the inverse attention path

[0104] 3) Adaptive weight fusion of dual-path spatial features by the gated network, including: dual-path features of general and special cases of pedestrian spatiotemporal connections obtained in sub-step 2) and They are used as inputs and sent to a multi-layer perceptron respectively, and then the corresponding weights are obtained through the Sigmoid function. and They have the same dimension as the input features, and then use the Softmax function to adjust the two weights. and Normalize and and After passing through a multi-layer perceptron layer, the weighted sum is performed according to the normalized weights to obtain the final extracted spatiotemporal fusion trajectory feature F s-t .

[0105] In the embodiment of the present invention, the step of obtaining the predicted multimodal trajectory endpoint based on the obtained spatiotemporal fusion trajectory features of the pedestrian specifically includes: for the input feature F s-t , use a multilayer perceptron to output the mixed Gaussian distribution parameters of the predicted trajectory endpoint; construct a mixed Gaussian distribution representing the predicted trajectory endpoint through the obtained mixed Gaussian distribution parameters, and randomly sample multiple possible predicted trajectory endpoints as multimodal endpoint prediction results, combine the mixed Gaussian distribution parameters and the true endpoint of the future trajectory, and use the maximum likelihood loss of the mixed Gaussian model to train the pedestrian trajectory prediction model.

[0106] In the embodiment of the present invention, the step of obtaining a complete multimodal prediction trajectory result based on the obtained multimodal prediction trajectory endpoint specifically includes: for a given input F imput , and the obtained multimodal trajectory endpoint as input, use a multilayer perceptron to calculate and output future predicted trajectory points other than the endpoint, and combine the endpoint to obtain the entire multimodal predicted trajectory. The process of training the multilayer perceptron uses the L2 distance loss between the predicted trajectory and the future true trajectory for training.

[0107] In the pedestrian trajectory prediction method based on complementary attention provided by the embodiment of the present invention, a pair of complementary attention matrices are obtained for both spatial interaction and motion trend to guide the forward and reverse attention modules to extract the spatiotemporal features of general and special cases, respectively. The present invention provides a prediction method for pedestrian trajectory prediction that can capture the special conditions of pedestrian motion, which not only provides more accurate prediction results for pedestrian trajectory prediction, but also reveals the generality and particularity of pedestrian motion decisions through the positive and negative attention mechanisms, so that trajectory prediction technology can handle more situations and improve the safety of applications.

[0108] The present invention further explores the potential of neural networks in pedestrian trajectory prediction by fusing positive and negative features through adaptive weights obtained through a gating mechanism and using them to generate the distribution of future trajectory endpoints. All parameters are learned through the network as much as possible, which enhances the learning ability of the network and reduces the cost of artificial features. A complete multimodal prediction result is obtained by sampling the endpoint and further predicting based on the endpoint. When facing special circumstances of pedestrian motion, more accurate prediction results can be obtained.

[0109] Please refer to Table 1 and Figure 3 Table 1 shows the experimental results of our method under the ETH and UCY datasets. ETH and UCY include some sub-datasets respectively; among them, ETH includes ETH and HOTEL datasets, and UCY includes UNIV, ZARA1, and ZARA2 datasets. For these five datasets, the method of the embodiment of the present invention and other compared methods use four of them for training, and the remaining one for testing. The experiment observes the trajectory within 3.2 seconds, and selects 8 moments on average, predicts the trajectory of the next 4.8 seconds, and selects 12 moments on average.

[0110] Table 1. Experimental results on ETH and UCY datasets

[0111]

[0112]

[0113] The experiment uses the average displacement error, i.e. the average value of the displacement error at each moment in the future, and the endpoint displacement error, i.e. the displacement error at the last moment in the future, as evaluation indicators. 20 endpoints are sampled from the endpoint distribution to generate 20 possible future trajectories and the best one is selected as the error calculation standard. As can be seen from Table 1, the method of the present invention improves the average displacement error and the endpoint displacement error by 13.8% and 10.4% respectively compared with the best previous method; at the same time, in order to verify the effectiveness of the complementary attention mechanism proposed by the present invention compared with the previous attention model, the model of the present invention improves the two indicators by an average of 35.9% and 43.4% compared with the fully connected attention model STAR and the sparse attention model SGCN.

[0114] In summary, the embodiment of the present invention discloses a pedestrian trajectory prediction method based on complementary attention, which belongs to the field of computer vision. Unlike other methods that only focus on the general situation of pedestrian movement, the present invention focuses on the randomness of pedestrian movement. In terms of the spatial interaction and movement trend of pedestrians, a complementary attention mechanism including forward and reverse attention is used to extract multiple possible spatiotemporal features, so that the prediction results can cover the general and special cases of pedestrian movement. The present invention generates a pair of complementary matrices for judging the general and special cases of pedestrian movement through complementary modules, guides the forward attention module to extract the features of the general situation of spatiotemporal movement, and the reverse attention module to extract the features of the special situation of spatiotemporal movement, and adaptively selects different weights for different scenes through a gating mechanism to fuse positive and negative features. Finally, the fused spatiotemporal features are used to generate the distribution of the end points of pedestrian movement, and the multimodal future trajectory that we finally predict is obtained by sampling the end points in the distribution and generating a complete predicted trajectory based on the end points.

[0115] The following are device embodiments of the present invention, which can be used to perform method embodiments of the present invention. For details not disclosed in the device embodiments, please refer to the method embodiments of the present invention.

[0116] In yet another embodiment of the present invention, a pedestrian trajectory prediction system based on complementary attention is provided in an embodiment of the present invention, comprising:

[0117] A pedestrian observation trajectory sequence acquisition module is used to obtain a pedestrian observation trajectory sequence to be predicted;

[0118] A trajectory prediction result acquisition module is used to input the pedestrian observation trajectory sequence into a pre-trained pedestrian trajectory prediction model and output a trajectory prediction result;

[0119] Wherein, the pedestrian trajectory prediction model includes:

[0120] A spatial interactive neural network, used for inputting the pedestrian observation trajectory sequence and outputting spatial trajectory features;

[0121] A time-series motion trend neural network, used to input the spatial trajectory features and output the spatiotemporal trajectory features;

[0122] The first multi-layer perceptron network is used to input the spatiotemporal trajectory features and output the endpoint distribution as the trajectory prediction result.

[0123] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0124] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0125] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0126] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A pedestrian trajectory prediction method based on complementary attention, characterized in that: The following steps are involved: Obtaining a pedestrian observation trajectory sequence to be predicted; Inputting the pedestrian observation trajectory sequence into a pre-trained pedestrian trajectory prediction model and outputting a trajectory prediction result; Wherein, the pedestrian trajectory prediction model includes: A spatial interactive neural network, used for inputting the pedestrian observation trajectory sequence and outputting spatial trajectory features; A time-series motion trend neural network, used to input the spatial trajectory features and output the spatiotemporal trajectory features; A first multi-layer perceptron network is used to input the spatiotemporal trajectory features and output a terminal distribution as a trajectory prediction result; in, The spatial interaction neural network comprises: A first discriminator, configured to input the pedestrian observation trajectory sequence and output a general case matrix and a special case matrix of spatial interaction; A first single-head self-attention network is used to input the pedestrian observation trajectory sequence and the general case matrix of spatial interaction, and output the general case feature of spatial interaction; A second single-head self-attention network is used to input the pedestrian observation trajectory sequence and the special case matrix of spatial interaction, and output the special case feature of spatial interaction; A first feature fusion network is used to input the general situation feature of the spatial interaction and the special situation feature of the spatial interaction, and output the spatial trajectory feature; The time series motion trend neural network comprises: A second discriminator, used for inputting the spatial trajectory features and outputting a general case matrix and a special case matrix of the temporal motion trend; A third single-head self-attention network is used to input the general situation matrix of the spatial trajectory characteristics and the temporal motion trend, and output the general situation characteristics of the spatiotemporal motion; a fourth single-head self-attention network, configured to input the special case matrix of the spatial trajectory features and the temporal motion trend, and output the special case features of the spatiotemporal motion; The second feature fusion network is used to input the general situation feature of the space-time motion and the special situation feature of the space-time motion, and output the space-time trajectory feature; The steps of obtaining the pre-trained pedestrian trajectory prediction model include: Obtain a sample training set, wherein each sample includes a pedestrian sample observation trajectory sequence and a future true trajectory; Inputting the observed trajectory sequence of pedestrian samples in the selected sample into the pedestrian trajectory prediction model, and outputting the predicted terminal distribution of the future trajectory; calculating the maximum likelihood estimation loss between the terminal distribution of the predicted future trajectory and the terminal of the future real trajectory in the selected sample, and obtaining the loss value; Based on the loss value, the model parameters are updated, and after reaching the preset convergence condition, the pre-trained pedestrian trajectory prediction model is obtained.

2. The pedestrian trajectory prediction method based on complementary attention according to claim 1, characterized in that: The first feature fusion network and the second feature fusion network are both gated networks.

3. The pedestrian trajectory prediction method based on complementary attention according to claim 1, characterized in that: The endpoint distribution is a mixed Gaussian endpoint distribution.

4. The pedestrian trajectory prediction method based on complementary attention according to claim 1, characterized in that: The pedestrian trajectory prediction model also includes: The second multi-layer perceptron is used to input the pedestrian observation trajectory sequence and the predicted trajectory endpoint obtained by sampling from the endpoint distribution, and output a complete future predicted trajectory.

5. The pedestrian trajectory prediction method based on complementary attention according to claim 1, characterized in that: The first discriminator comprises: A first multi-head self-attention network is used to input the pedestrian observation trajectory sequence and output a multi-head attention matrix; A first convolutional neural network, used to input the multi-head attention matrix and output a single-channel self-attention matrix; The first symbolic function is used to input the single-channel self-attention matrix processed by the Sigmoid function and output the general case matrix and the special case matrix of spatial interaction.

6. The pedestrian trajectory prediction method based on complementary attention according to claim 1, characterized in that: The second discriminator comprises: A second multi-head self-attention network is used to input the spatial trajectory features and output a multi-head attention matrix; A second convolutional neural network, used to input the multi-head attention matrix and output a single-channel self-attention matrix; The second symbolic function is used to input the single-channel self-attention matrix processed by the Sigmoid function, and output the general case matrix and the special case matrix of the temporal motion trend.

7. A pedestrian trajectory prediction system based on complementary attention, characterized in that: include: A pedestrian observation trajectory sequence acquisition module is used to obtain a pedestrian observation trajectory sequence to be predicted; A trajectory prediction result acquisition module is used to input the pedestrian observation trajectory sequence into a pre-trained pedestrian trajectory prediction model and output a trajectory prediction result; Wherein, the pedestrian trajectory prediction model includes: A spatial interactive neural network, used for inputting the pedestrian observation trajectory sequence and outputting spatial trajectory features; A time-series motion trend neural network, used to input the spatial trajectory features and output the spatiotemporal trajectory features; A first multi-layer perceptron network is used to input the spatiotemporal trajectory features and output a terminal distribution as a trajectory prediction result; in, The spatial interaction neural network comprises: A first discriminator, configured to input the pedestrian observation trajectory sequence and output a general case matrix and a special case matrix of spatial interaction; A first single-head self-attention network is used to input the pedestrian observation trajectory sequence and the general case matrix of spatial interaction, and output the general case feature of spatial interaction; A second single-head self-attention network is used to input the pedestrian observation trajectory sequence and the special case matrix of spatial interaction, and output the special case feature of spatial interaction; A first feature fusion network is used to input the general situation feature of the spatial interaction and the special situation feature of the spatial interaction, and output the spatial trajectory feature; The time series motion trend neural network comprises: A second discriminator, used for inputting the spatial trajectory features and outputting a general case matrix and a special case matrix of the temporal motion trend; A third single-head self-attention network is used to input the general situation matrix of the spatial trajectory characteristics and the temporal motion trend, and output the general situation characteristics of the spatiotemporal motion; a fourth single-head self-attention network, configured to input the special case matrix of the spatial trajectory features and the temporal motion trend, and output the special case features of the spatiotemporal motion; The second feature fusion network is used to input the general situation feature of the space-time motion and the special situation feature of the space-time motion, and output the space-time trajectory feature; The steps of obtaining the pre-trained pedestrian trajectory prediction model include: Obtain a sample training set, wherein each sample includes a pedestrian sample observation trajectory sequence and a future true trajectory; Inputting the observed trajectory sequence of pedestrian samples in the selected sample into the pedestrian trajectory prediction model, and outputting the predicted terminal distribution of the future trajectory; calculating the maximum likelihood estimation loss between the terminal distribution of the predicted future trajectory and the terminal of the future real trajectory in the selected sample, and obtaining the loss value; Based on the loss value, the model parameters are updated, and after reaching the preset convergence condition, the pre-trained pedestrian trajectory prediction model is obtained.

Citation Information

Patent Citations

  • Pedestrian trajectory prediction method and system based on trend guidance and sparse interaction

    CN112215423A

  • Generative adversarial trajectory prediction method based on attention mechanism

    CN112766561A