Pedestrian trajectory prediction method, storage medium and computer equipment
Through STR-GGRNN network and non-negative matrix factorization, the problem of low accuracy of pedestrian trajectory prediction in unknown environments is solved, and more accurate pedestrian trajectory prediction is achieved, supporting autonomous driving and intelligent traffic management.
Patent Information
- Application Number
- CN202510320311.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-18
AI Technical Summary
The existing pedestrian trajectory prediction methods are not accurate when adjusting the graph structure in unknown environments, making it difficult to accurately predict the future trajectory of pedestrians.
The STR-GGRNN network is used to map pedestrian dynamics to the spatio-temporal graph network, and social interactions are automatically inferred through the complementary graph edges, combined with non-negative matrix factorization and self-learning social neighbor recommendation system to generate future trajectories and improve prediction accuracy.
By reducing errors, the accuracy of pedestrian trajectory prediction is improved, and the local optimal solution is provided, providing more accurate pedestrian behavior analysis for autonomous driving and intelligent traffic management.
Smart Images

Figure CN120339326A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, and particularly to a pedestrian trajectory prediction method, a storage medium, and a computer device. Background Art
[0002] In recent years, the research on autonomous driving technology has become a hot topic and trend. Autonomous driving technology includes environmental perception, positioning and navigation, path planning, and decision-making control. Autonomous driving environmental perception technology includes environmental perception and understanding. Among them, pedestrian trajectory prediction plays a crucial role in advanced driver assistance, autonomous driving, and robot navigation, and can help intelligent connected vehicles make reasonable decision-making plans and improve vehicle safety.
[0003] Motion prediction refers to the ability of a robot to predict the future state of an object, including trajectory prediction, path prediction, pose prediction, etc. Trajectory prediction is a subfield of motion prediction, which refers to the task of predicting the future position, speed, direction and other state information of an object given its past or current motion trajectory. Currently, the methods for pedestrian trajectory prediction are mainly divided into three categories:
[0004] 1. Predicting pedestrian intentions by studying the surrounding environment (Bayes' theorem).
[0005] 2. Studying people's social attributes to predict pedestrian movement (social force model);
[0006] 3. Predicting pedestrian trajectories through machine learning (deep learning).
[0007] However, existing work mostly relies on assumptions about the scene and dynamic space, which poses a great challenge to adjusting the graph structure of a linear system in an unknown environment. Therefore, the prediction accuracy of human trajectories is not high. Summary of the Invention
[0008] In view of this, the present application provides a pedestrian trajectory prediction method, a storage medium, and a computer device, which can improve the accuracy of pedestrian trajectory prediction.
[0009] According to one aspect of the present application, a pedestrian trajectory prediction method is provided. The method includes:
[0010] Obtaining a surveillance video of a target area within a preset historical time period, and obtaining a spatio-temporal graph of pedestrian trajectories and map information of the target area based on the surveillance video. The spatio-temporal graph of pedestrian trajectories includes the movement trajectories of multiple target pedestrians. For any movement trajectory, each sampling moment within the preset historical time period is used as a trajectory node, and the edge between adjacent trajectory nodes is used as the movement trajectory segment between adjacent moments. Each trajectory node stores the position coordinates and head pose of the target pedestrian;
[0011] For the spatio-temporal graph of the pedestrian trajectory, based on the position coordinates of the target pedestrian at each sampling moment within a preset historical time period, a pedestrian trajectory feature vector is extracted, and based on the head postures of the target pedestrian at each sampling moment within the preset historical time period, a pedestrian head posture feature vector is extracted. Based on the map information, a scene space feature vector is extracted;
[0012] The pedestrian trajectory feature vector, the pedestrian head posture feature vector, and the scene space feature vector are encoded and fused to obtain a pedestrian scenario spatio-temporal interaction feature vector, where the pedestrian scenario spatio-temporal interaction feature vector is used to characterize the spatio-temporal interaction features between the target pedestrian and other pedestrians, and between the target pedestrian and the scene space;
[0013] Based on the social trajectory recommender and the pedestrian scenario spatio-temporal interaction feature vector, the trajectory trend of the target pedestrian within a preset future time period is predicted.
[0014] According to another aspect of the present application, a storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned pedestrian trajectory prediction method is implemented.
[0015] According to still another aspect of the present application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor. When the processor executes the program, the above-mentioned pedestrian trajectory prediction method is implemented.
[0016] By means of the above technical solution, a pedestrian trajectory prediction method, a storage medium, and a computer device provided by the present application propose a novel STR-GGRNN network (composed of a convolutional neural network and a variant of the long short-term memory network). By means of an online framework centered on the edge, the crowd dynamics are mapped onto the spatio-temporal graph network. With minimal engineering effort, social interactions are automatically inferred by completing the graph edges, reducing the error of pedestrian trajectory prediction. A novel kernel is proposed, which can map the STR-GGRNN feature gradients and generate future trajectories accordingly as an effect of the interaction changes related to the static scene over time. Integrate non-negative matrix factorization into the work of an efficient self-learning social neighbor recommendation system to evaluate the importance of edges from a neural perspective using a compact version of node and edge features. By generating variant neighborhood suggestions and checking them according to the prediction accuracy, a locally optimal solution is provided for modeling social interactions in the spatio-temporal graph, thereby improving the accuracy of pedestrian trajectory prediction.
[0017] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically exemplified below. Description of the Drawings
[0018] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0019] Figure 1 A schematic flowchart of a pedestrian trajectory prediction method provided by an embodiment of the present application is shown;
[0020] Figure 2 A schematic flowchart of another pedestrian trajectory prediction method provided by an embodiment of the present application is shown. Detailed implementation manners
[0021] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0022] In this embodiment, a pedestrian trajectory prediction method is provided. As Figure 1 shown, the method includes:
[0023] Step 101, obtain the surveillance video of the target area within a preset historical time period, and obtain the spatio-temporal graph of pedestrian trajectories and map information of the target area based on the surveillance video. Among them, the spatio-temporal graph of pedestrian trajectories includes the movement trajectories of multiple target pedestrians. For any movement trajectory, each sampling moment within the preset historical time period is used as a trajectory node, and the edge between adjacent trajectory nodes is used as the movement trajectory segment between adjacent moments. Each trajectory node stores the position coordinates and head posture of the target pedestrian.
[0024] Existing graph neural networks do not solve the problem of online graph completion. As the neighborhood in the graph grows dynamically, it is necessary to evaluate the credibility of the generated neighborhood. The association between pedestrians requires a self-growing mechanism. By considering the visual span of pedestrians, we can be more certain about how pedestrians consider each other when moving. Therefore, the learning mechanism can learn to estimate the influence and interaction between them.
[0025] To overcome this problem, in the above embodiment of the present application, a self-learning re-logical reasoning method is proposed, which plays an important role in the growth of graph edges. At the same time, a table structure is specifically designed to model the static environment as a local neighborhood. This hybrid method aims to improve the modeling of pedestrian social influence and their perception of the surrounding static environment, thereby reducing the trajectory prediction error.
[0026] Specifically, obtain the surveillance video of the target area within a preset historical time period (for example, the historical 1.6 s), and then obtain the spatio-temporal map of the pedestrian trajectory and the map information of the target area through the surveillance video, so as to predict the trajectory in the next 2.4 s through the pedestrian trajectory within the historical 1.6 s. Regarding the sampling moment, a trajectory node can be sampled every 0.2 s, that is, there are 5 trajectory nodes per second.
[0027] Optionally, in step 101, the head pose is characterized by the orientation angle of the head of the target pedestrian in the plane, and the map information includes at least one of road information, special area information, traffic signal information, and other map information.
[0028] In the above embodiments of the present application, the head pose can be characterized by the orientation angle of the head of the target pedestrian in the plane, that is, a specific angle.
[0029] When predicting the pedestrian trajectory, the map information that can affect the pedestrian trajectory may include, but is not limited to, the following aspects:
[0030] 1. Road information:
[0031] Road position: Pedestrians usually walk on the road, and the position and orientation of the road will directly affect the walking trajectory of pedestrians.
[0032] Lane direction: Although it is mainly the basis for vehicle driving, the lane direction also indirectly affects the walking path of pedestrians, especially at crosswalks or intersections.
[0033] Sidewalk information: The position, width, and orientation of the sidewalk are the main paths for pedestrians to walk, which are crucial for predicting pedestrian trajectories.
[0034] 2. Special area information:
[0035] Crosswalk: Pedestrians usually cross the road at crosswalks, so the position and width of the crosswalk will affect the crossing trajectory of pedestrians.
[0036] No-parking areas, restricted areas: These areas will restrict the walking path of pedestrians and prevent them from entering dangerous or prohibited areas.
[0037] Other special areas: Such as transportation hubs like bus stops and subway stations, which will attract a large number of pedestrians to gather and flow, thus affecting the walking trajectory of pedestrians.
[0038] 3. Traffic signal information:
[0039] Traffic light information: The status of traffic lights (red light, green light, yellow light) will directly affect the walking decisions and trajectories of pedestrians. For example, pedestrians usually wait when the red light is on and cross the road when the green light is on.
[0040] Other traffic signals: such as pedestrian traffic lights, countdown timers, etc., also provide instructions and references for pedestrians to walk.
[0041] 4. Other map information:
[0042] Obstacle information: Fixed obstacles such as trees, flower beds, buildings, etc. will limit the walking paths of pedestrians.
[0043] Terrain information: Such as slopes, steps, etc., will affect the walking speed and trajectory of pedestrians.
[0044] Environmental semantic information: Such as road structure, traffic rules, etc. Although these information do not directly constitute the geometric shape of the map, they will affect the walking decisions and trajectories of pedestrians.
[0045] In summary, map information plays a crucial role in pedestrian trajectory prediction. By comprehensively considering road information, special area information, traffic signal information, and other map information, the walking trajectories of pedestrians can be predicted more accurately.
[0046] Step 102, for the spatio-temporal graph of the pedestrian trajectory, based on the position coordinates of the target pedestrian at each sampling moment within a preset historical time period, extract the pedestrian trajectory feature vector, and based on the head pose of the target pedestrian at each sampling moment within a preset historical time period, extract the pedestrian head pose feature vector, and based on the map information, extract the scene space feature vector.
[0047] Next, for the spatio-temporal graph of the pedestrian trajectory, based on the position coordinates of the target pedestrian at each sampling moment within a preset historical time period, extract the pedestrian trajectory feature vector, and based on the head pose of the target pedestrian at each sampling moment within a preset historical time period, extract the pedestrian head pose feature vector, and based on the map information, extract the scene space feature vector. For example, by processing the pedestrian video (the aforementioned acquired surveillance video), obtain the pedestrian 2D trajectory sequence (the pedestrian trajectory feature vector, 2D refers to the two-dimensional X and Y coordinates) and the head pose sequence of the pedestrian (the pedestrian head pose feature vector), assign the i-th target pedestrian to a trajectory node to store the ground truth trajectory and the two-dimensional head pose, and construct a pedestrian graph (the spatio-temporal graph of the pedestrian trajectory) for use as the input of the subsequent prediction model.
[0048] More specifically, given a set of target pedestrians N, their trajectories over time can be represented using a spatio-temporal graph. Each graph contains N nodes, including the 2D position sequence and the 2D head pose sequence of the target pedestrians. Each target pedestrian i is assigned to a trajectory node to store the ground truth trajectory and the 2D head pose. Temporal edges connect each trajectory node to represent the temporal relationship between pedestrians.
[0049] Step 103: Encode and fuse the pedestrian trajectory feature vector, the pedestrian head pose feature vector, and the scene space feature vector to obtain a pedestrian scenario spatio-temporal interaction feature vector, where the pedestrian scenario spatio-temporal interaction feature vector is used to characterize the spatio-temporal interaction features between the target pedestrian and other pedestrians, as well as between the target pedestrian and the scene space.
[0050] Next, a new type of gated graph recurrent neighborhood network (STR-GGRNN network) can be constructed, which can encode the static scene into a fixed local neighborhood grid, model the contextual interaction between pedestrians and the surrounding environment, generate future trajectories that help consider pedestrian "context awareness", and thus make reasonable predictions. Specifically, the gated graph recurrent neighborhood network can be used to encode and fuse the pedestrian trajectory feature vector, the pedestrian head pose feature vector, and the scene space feature vector to obtain a pedestrian scenario spatio-temporal interaction feature vector.
[0051] Optionally, in step 103, the pedestrian trajectory feature vector, the pedestrian head pose feature vector, and the scene space feature vector are encoded and fused to obtain a pedestrian scenario spatio-temporal interaction feature vector. Referring to Figure 2 as shown, the specific steps are as follows:
[0052] Step 1031: Use a single GridLSTM unit GLSTM nu to encode the pedestrian trajectory feature vector and the scene space feature vector in the pedestrian trajectory spatio-temporal graph into features and initial hidden states, and generate relative social features based on the features and the initial hidden states, where the relative social features are expressed as:
[0053] f S ,h s = GLSTM nu (T, h0),
[0054] T is the feature, h0 is the initial hidden state, f S is the relative social feature, h s is the hidden state of the relative social feature.
[0055] Step 1032: Adopt the two-dimensional convolutional layer of the convolutional neural network and introduce a grid mask to encode the relative social features to obtain static features, where the grid mask is used to discretize the static scene space neighborhood depicted by the filters in the convolutional neural network into square grids, and the square grids are used to represent spatial relationships. The static features are expressed as:
[0056] C map = CNN(f S , h s ) * M,
[0057] Cmap is a static feature, CNN is a convolutional neural network, and M is a grid mask.
[0058] Step 1033, use a single GridLSTM cell GLSTM O and static features to encode the pedestrian head pose feature vector to obtain visual spatial features, where the visual spatial features are expressed as:
[0059]
[0060] f o is the visual spatial feature, and h o is the hidden state of the visual spatial feature. is the pedestrian head pose feature vector.
[0061] Step 1034, combine the relative social features and visual spatial features to generate a neighborhood representing the pedestrian scenario spatio-temporal interaction feature vector, where the representation formula of the neighborhood is:
[0062] F = f s *f o ,
[0063] F is the neighborhood representing the pedestrian scenario spatio-temporal interaction feature vector, f s is the relative social feature, and f o is the visual spatial feature.
[0064] In the above embodiments of the present application, a new gated graph recurrent neighborhood network (STR-GG RNN network) is constructed to obtain the transformed pedestrian head pose sequence (pedestrian head pose feature vector) and trajectory sequence (pedestrian trajectory feature vector) and encode each pedestrian graph (pedestrian trajectory spatio-temporal graph) G * through STR t (the constructed new STR-GGR NN network) to generate the relative social feature f S . Among them, the CNN two-dimensional convolutional layer encodes the interaction between people and space, and takes the static scene image as the input. The grid mask M is applied to the convolutional features, and the static spatial neighborhood depicted by the filter is discretized into a square grid and uniformly initialized. Use the static feature C map to encode the pedestrian head feature V to formulate a "visual space" domain representation. The GridLSTM cell combines the visual perception state (static feature C map ) C, the scene spatial feature (pedestrian head pose feature vector) V, and the normal initial hidden state H o , guides the pedestrian's attention to the physical environment, and stores the pedestrian's interaction with the situation in f oAmong them, the GG RNN combines the outputs of the GLSTM nu and the GLSTM O to generate the final neighborhood representation F.
[0065] Specifically, first generate the relative social feature f S :
[0066] The gated graph recurrent neighborhood network obtains the transformed head pose sequence from the transformation function (the transformation function is a linear layer used to transform the head pose sequence and the trajectory sequence into the input structure required by the model) and the trajectory Then the STR * runs a single GridLSTM cell GLSTM t on each graph G nu , and the graph is encoded as the feature T and the initial hidden state h0 to generate the relative social feature f S , as shown in the following formula:
[0067] f S ,h s = GLSTM nu (T, h0),
[0068] Next, encode the interaction between people and space:
[0069] The CNN two-dimensional convolutional layer encodes the interaction between people and space, taking the static scene image as the input. The grid mask M is applied to the convolutional features, and the static spatial neighborhood depicted by the filter is discretized into a square grid and uniformly initialized, as shown in the following formula:
[0070] C map = CNN(f S , h s ) * M,
[0071] Next, encode the pedestrian head pose feature
[0072] Use the static feature C map to encode the pedestrian head feature V, so as to use a single GridLSTM cell GLSTM O to establish the "visual space" neighborhood representation, as shown in the following formula:
[0073]
[0074] Finally, the generation of the neighborhood representation is:
[0075] F = f s * f o .
[0076] The final output of the gated recurrent network is a feature vector, representing the spatio-temporal features required for pedestrian trajectory prediction. The domain represents the spatio-temporal interaction features between the target pedestrian and other surrounding pedestrians and the surrounding environment. It serves as the input to the social trajectory recommender network.
[0077] Specifically, the GridLSTM cell, short for Grid Long Short-Term Memory, is a special LSTM (Long Short-Term Memory) network structure mainly used to process multi-dimensional time series data. The GridLSTM cell is an extension of the traditional LSTM cell to meet the processing requirements of multi-dimensional time series data. Compared with the standard LSTM cell, the GridLSTM cell has one or more LSTM cells respectively set in the time-domain and frequency-domain, and these cells are interconnected in a specific way to form a grid-like structure. The GridLSTM cell has the following unique features:
[0078] 1. Multi-dimensional data processing ability: The GridLSTM cell can process information in both the time-domain and frequency-domain simultaneously, which gives it unique advantages in processing multi-dimensional time series data. For example, in fields such as signal processing and image recognition, data often has multiple dimensions, and the GridLSTM cell can effectively capture the correlations and features between these dimensions.
[0079] 2. Enhanced memory ability: Since the GridLSTM cell has LSTM cells respectively set in the time-domain and frequency-domain, and these cells are interconnected through a grid-like structure, the memory ability of the network is enhanced. This enables the GridLSTM cell to better maintain the coherence and stability of information when processing long sequence data.
[0080] 3. Flexibility: The structure of the GridLSTM cell has a certain degree of flexibility and can be adjusted according to specific application scenarios and requirements. For example, different numbers of LSTM cells can be selected for combination according to the dimensions and features of the data to achieve the best processing effect.
[0081] The GridLSTM cell has broad application prospects in multiple fields, including but not limited to:
[0082] 1. Signal processing: In the field of signal processing, the GridLSTM cell can be used to extract feature information from signals, such as speech recognition and audio classification.
[0083] 2. Image recognition: In the field of image recognition, the GridLSTM cell can be used to capture spatial and temporal features in images, such as video analysis and behavior recognition.
[0084] 3. Natural Language Processing: In the field of natural language processing, GridLSTM units can be used to process temporal information and context relationships in text data, such as text classification, sentiment analysis, etc.
[0085] Although GridLSTM units have unique advantages in processing multi-dimensional time series data, there are also some challenges and limitations:
[0086] 1. Computational Complexity: Due to the relatively complex structure of GridLSTM units, more computational resources and time are required during training. This limits the use of GridLSTM units in some application scenarios with high real-time requirements.
[0087] 2. Parameter Tuning: The performance of GridLSTM units depends to a large extent on parameter tuning. However, due to the complex structure, the difficulty of parameter tuning is relatively large. This requires developers to have rich experience and skills in the design and training process.
[0088] In summary, GridLSTM units are a multi-dimensional time series processing network structure with unique advantages. It has broad application prospects in the fields of signal processing, image recognition, natural language processing, etc., but at the same time, there are also some challenges and limitations. In future research, optimization methods and application scenarios of GridLSTM units can be further explored to give full play to their potential.
[0089] Step 104: Based on the social trajectory recommender and the pedestrian scenario spatio-temporal interaction feature vector, predict the trajectory trend of the target pedestrian within a preset future time period.
[0090] Next, based on the social trajectory recommender and the pedestrian scenario spatio-temporal interaction feature vector, the trajectory trend of the target pedestrian within a preset future time period can be predicted.
[0091] Specifically, regarding the social trajectory recommender, the social trajectory recommender is a system that makes recommendations based on users' social activities and movement trajectory data. The social trajectory recommender constructs users' social activity trajectories by collecting and analyzing users' social activity data and movement trajectory information, such as location check-ins, activity participation, social media sharing, etc. Then, using these trajectory data, the recommender can recommend people, places, or services that the user is interested in. The principle is that users' movement trajectories and social activities can often reflect their interest preferences, living habits, and social circles, thus providing a basis for personalized recommendations.
[0092] The main functions of the social trajectory recommender, for example:
[0093] 1. Trajectory Recording and Analysis: The social trajectory recommender can record users' movement trajectories, including location information, timestamps, etc., and analyze and process this trajectory data to extract users' movement characteristics and interest preferences.
[0094] 2. Personalized Recommendation: Based on users' trajectory data and social activity information, the recommender can recommend locations, activities, or services that match their interest preferences to users. For example, recommend nearby restaurants, tourist attractions, or social activities to users.
[0095] 3. Social Relationship Expansion: By analyzing users' social activity trajectories, the recommender can also recommend strangers with similar interests or activity trajectories to users, thus helping users expand their social circles.
[0096] Application scenarios of the social trajectory recommender, for example:
[0097] 1. Social Network Services: On social network platforms, the social trajectory recommender can recommend interesting people, groups, or topics to users according to their activity trajectories and interest preferences, enhancing users' social experience.
[0098] 2. Travel Recommendations: By analyzing users' travel history and movement trajectories, the recommender can recommend destinations, attractions, or travel routes that match their travel preferences to users.
[0099] 3. Local Life Services: Based on users' geographical locations and activity trajectories, the recommender can recommend local life service facilities such as nearby restaurants, shopping centers, cinemas, etc. to users.
[0100] Optionally, step 104 predicts the trajectory trend of the target pedestrian within a preset future time period based on the social trajectory recommender and the pedestrian scenario spatio-temporal interaction feature vector, specifically including:
[0101] Step 1041, predict the social interaction state and the future position of the target pedestrian based on the kernel of the social trajectory recommender and the relative social features, where the prediction formula is as follows:
[0102]
[0103] is the predicted future position of the target pedestrian, H t+1 is the social interaction state, K is the kernel of the social trajectory recommender, f s is the relative social feature, is the matrix assigned to the weighted hidden state.
[0104] Step 1042: Based on the predicted future positions of the target pedestrians, obtain the future trajectory sequences of the target pedestrians, and use the future trajectory sequences of the target pedestrians as the prediction results of the trajectory trends within a preset future time period. Among them, the future trajectory sequences of the target pedestrians are tensors of N×K×T×2, where N represents the number of target pedestrians, K represents the number of modes of the predicted trajectories, T represents the number of prediction time steps, 2 represents the two-dimensional coordinates (x, y) at each time step, and the time step is the sampling moment.
[0105] In the above embodiments of the present application, the social trajectory recommendation network can generate multiple unique edge sets to connect pedestrian nodes, and then select the edge set that can generate the minimum error from its suggestions. Specifically, the kernel K of the social trajectory recommendation network is used to generate social interaction states and future position predictions. Based on the STR-GGRNN and ST-V model variants, the kernel K calculates and uses the adjacency states of each pedestrian. On the one hand, the kernel K calculates the soft attention of social and situational interaction features, and on the other hand, it calculates their correlation with the static map. It is the final neighborhood representation, which combines the outputs of GLSTM nu and GLSTM O to ensure rich multi-modal feature representations and strengthen the intertwining between context and social interactions. The STR model can adjust the influence of each region to update the situational features and evaluate the strength of social relationships. Non-negative matrix factorization is introduced to appropriately approximate social relationships based on neural attention features and static maps, generating a more compact and efficient adjacency representation.
[0106] Optionally, after step 1041: Predicting the social interaction states and the future positions of the target pedestrians based on the kernel of the social trajectory recommender and the relative social features, the following steps are further included:
[0107] Step 1042: Calculate the soft attention weights based on the neighborhood and scaled self-attention mechanism that characterizes the pedestrian scenario spatio-temporal interaction feature vectors. Based on the calculated soft attention weights, update the pedestrian scenario spatio-temporal interaction feature vectors and re-evaluate the strength of the social interaction states. The scaled self-attention mechanism is expressed as:
[0108]
[0109] a is the soft attention weight, calculated by the Softmax function. is to exponentiate the neighborhood F.
[0110] Step 1043: Introduce non-negative matrix factorization, where the non-negative matrix factorization is:
[0111]
[0112] A P= W a / max(W a ),
[0113] W a is the matrix assigned to the attention weight a, H t is the matrix assigned to the weighted hidden state, T is the number of time steps, NMF is the non - negative matrix factorization function, which is used to factorize the attention weight a into two non - negative matrices W a and H t , A P is the normalized value of the attention weight.
[0114] In the above - mentioned embodiments of the present application, the situational features can be updated and the social relationship strength can be evaluated. Specifically, after generating the future trajectory, is used to calculate the soft attention weight, and the STR model will adjust the influence of each region to update the situational features and evaluate the strength of social interaction. It deploys a scaled self - attention mechanism as shown in the following formula:
[0115]
[0116] Non - negative matrix factorization can also be cited to appropriately approximate the social relationship based on the neural attention features and the static map, generating a more compact and efficient adjacency representation. For example, A0 is the initial A matrix (all 1s) of 10 pedestrians in the scene, with a size of [10×10], W a is the matrix assigned to the attention a, H t is the matrix assigned to the weighted hidden state:
[0117]
[0118] A P = W a / max(W a )。
[0119] Optionally, step 104 predicts the trajectory trend of the target pedestrian within a preset future time period based on the social trajectory recommender and the spatio - temporal interaction feature vector of the pedestrian's situation, and further includes the following steps:
[0120] Step 1044: Based on the social trajectory recommender and the pedestrian scenario spatio-temporal interaction feature vector, for the three trajectory influencing factors that affect the pedestrian trajectory, for each trajectory influencing factor, predict the sub-factor that has the greatest impact on the trajectory of the target pedestrian respectively, and based on the predicted sub-factors with the greatest impact, jointly determine the trajectory direction of the target pedestrian within a preset future time period, where the trajectory influencing factors include other pedestrians, scene space, and the head pose of the target pedestrian. Among other pedestrians, each sub-factor is each other person except the target pedestrian in the target area. In the scene space, each sub-factor is various scene space elements in the scene space. In the head pose of the target pedestrian, each sub-factor is the head pose of the target pedestrian at each sampling moment within a preset historical time period.
[0121] In the above embodiments of the present application, specifically, determining the trajectory direction of the target pedestrian based on the predicted sub-factors with the greatest impact is a process of comprehensive analysis and decision-making. In the framework of the social trajectory recommender and the pedestrian scenario spatio-temporal interaction feature vector, for example, the following steps can be taken:
[0122] 1. Identify and predict the sub-factor with the greatest impact
[0123] First, clarify the sub-factor with the greatest impact in each of the three influencing factors (other pedestrians, scene space, and the head pose of the target pedestrian). These sub-factors may include but are not limited to:
[0124] Other pedestrians: the speed, direction, relative distance from the target pedestrian, and behavior patterns (such as whether walking in a hurry, whether avoiding, etc.) of the pedestrians.
[0125] Scene space: road structure, obstacle position, traffic signal status, visibility, etc.
[0126] Head pose of the target pedestrian: line of sight direction, facial expression, head rotation angle, etc., which may reflect the intention or attention direction of the pedestrian.
[0127] Using the social trajectory recommender and combining with the pedestrian scenario spatio-temporal interaction feature vector, we can predict these sub-factors. The prediction results may include the specific values or states of the sub-factors at future time points.
[0128] 2. Comprehensively analyze the prediction results
[0129] Next, it is necessary to comprehensively analyze these prediction results to determine how they jointly affect the trajectory direction of the target pedestrian. This may need to consider the following aspects:
[0130] Interaction between sub-factors: For example, the behavior pattern of other pedestrians may be affected by the scene space constraints, and the head pose of the target pedestrian may reflect its reaction to other pedestrians or the scene space.
[0131] Time factor: The prediction results may cover different time points, and we need to consider the changing trends of these factors over time.
[0132] Weight assignment: The degrees of influence of different sub-factors on the target pedestrian's trajectory may vary. Therefore, appropriate weights need to be assigned to each sub-factor according to the actual situation.
[0133] 3. Determine the direction of the target pedestrian's trajectory
[0134] Based on the above analysis, a comprehensive prediction result can be obtained, that is, the possible trajectory direction of the target pedestrian at future time points. This process may involve the following steps:
[0135] Trajectory simulation: Utilize the predicted sub-factors with the greatest influence and combine with the algorithm of the social trajectory recommender to simulate the possible trajectory of the target pedestrian at future time points.
[0136] Trajectory evaluation: Evaluate the simulated trajectory, considering factors such as its rationality, feasibility, and safety.
[0137] Trajectory selection: According to the evaluation results, select the trajectory that best conforms to the actual situation as the future trajectory direction of the target pedestrian.
[0138] 4. Continuous optimization and adjustment
[0139] Finally, the prediction results need to be continuously optimized and adjusted according to the actual situation. For example, when the actual observed trajectory of the target pedestrian does not match the prediction results, we need to promptly adjust the prediction model or parameters to improve the accuracy and reliability of the prediction.
[0140] In summary, determining the direction of the target pedestrian's trajectory based on the predicted sub-factors with the greatest influence is a complex and meticulous process that requires comprehensive consideration of various factors and in-depth analysis. By continuously optimizing and adjusting the prediction model, we can improve the accuracy and reliability of the prediction, providing more powerful support for fields such as intelligent traffic management and pedestrian behavior analysis.
[0141] By applying the technical solution of this embodiment, the accuracy of pedestrian trajectory prediction can be improved.
[0142] Based on the above as Figures 1 to 2 shown method, correspondingly, an embodiment of the present application also provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the pedestrian trajectory prediction method as Figures 1 to 2 shown above.
[0143] Based on such understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, external hard drive, etc.), and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various implementation scenarios of the present application.
[0144] Based on the above method as Figures 1 to 2 shown, in order to achieve the above object, an embodiment of the present application further provides a computer device, which can specifically be a personal computer, server, network device, etc. The computer device includes a storage medium and a processor; the storage medium is used for storing a computer program; the processor is used for executing the computer program to implement the pedestrian trajectory prediction method as Figures 1 to 2 shown.
[0145] Optionally, the computer device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, and so on. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a Bluetooth interface, a WI-FI interface), etc.
[0146] Those skilled in the art can understand that the structure of a computer device provided in this embodiment does not constitute a limitation on the computer device, and it may include more or fewer components, or combine certain components, or have different component arrangements.
[0147] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing and saving the hardware and software resources of the computer device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement the communication between the components inside the storage medium, as well as the communication between the storage medium and other hardware and software in the entity device.
[0148] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware. A novel STR-GGRNN network (composed of a convolutional neural network and a variant of the long short-term memory network) is proposed. By means of an edge-centered online framework, crowd dynamics are mapped onto a spatio-temporal graph network. With minimal engineering effort, social interactions are automatically inferred by complementing graph edges, reducing the error of pedestrian trajectory prediction. A novel kernel is proposed that can map the STR-GGRNN feature gradients and generate future trajectories accordingly as an effect of the interaction changes related to the static scene over time. Integrating non-negative matrix factorization into the work of an efficient self-learning social neighbor recommendation system to evaluate the importance of edges from a neural perspective using a compact version of node and edge features, providing a locally optimal solution for modeling social interactions in spatio-temporal graphs by generating variant neighborhood suggestions and checking them according to the prediction accuracy, thereby improving the pedestrian trajectory prediction accuracy.
[0149] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present application. Those skilled in the art can understand that the modules in the devices in the embodiments can be distributed in the devices in the embodiments according to the descriptions of the embodiments, or can be correspondingly changed and located in one or more devices different from the present embodiment. The modules in the above embodiments can be combined into one module, or further split into multiple sub-modules.
[0150] The serial numbers of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments. The above disclosure is only several specific embodiments of the present application. However, the present application is not limited thereto, and any changes that can be made by those skilled in the art should fall within the protection scope of the present application.
Claims
1. A pedestrian trajectory prediction method, characterized in that, The method includes: Obtaining a surveillance video of a target area within a preset historical time period, and obtaining a spatio-temporal graph of pedestrian trajectories and map information of the target area based on the surveillance video. Among them, the spatio-temporal graph of pedestrian trajectories includes the movement trajectories of multiple target pedestrians. For any movement trajectory, each sampling moment within the preset historical time period is used as a trajectory node, and the edge between adjacent trajectory nodes is used as the movement trajectory segment between adjacent moments. Each trajectory node stores the position coordinates and head pose of the target pedestrian. For the spatio-temporal graph of pedestrian trajectories, based on the position coordinates of the target pedestrian at each sampling moment within the preset historical time period, extracting a pedestrian trajectory feature vector, and based on the head pose of the target pedestrian at each sampling moment within the preset historical time period, extracting a pedestrian head pose feature vector. Based on the map information, extracting a scene space feature vector. Encoding and fusing the pedestrian trajectory feature vector, the pedestrian head pose feature vector, and the scene space feature vector to obtain a pedestrian scenario spatio-temporal interaction feature vector. Among them, the pedestrian scenario spatio-temporal interaction feature vector is used to characterize the spatio-temporal interaction features between the target pedestrian and other pedestrians, and between the target pedestrian and the scene space. Based on the social trajectory recommender and the pedestrian scenario spatio-temporal interaction feature vector, predicting the trajectory trend of the target pedestrian within a preset future time period.
2. The method according to claim 1, wherein The encoding and fusing of the pedestrian trajectory feature vector, the pedestrian head pose feature vector, and the scene space feature vector to obtain a pedestrian scenario spatio-temporal interaction feature vector includes: Using a single GridLSTM cell GLSTM nu , encoding the pedestrian trajectory feature vector and the scene space feature vector in the spatio-temporal graph of pedestrian trajectories into features and an initial hidden state, and generating relative social features based on the features and the initial hidden state, where the relative social features are expressed as: f S , h s = GLSTM nu (T, h0), where T is a feature, h0 is the initial hidden state, and f S is the relative social feature, and h s is the hidden state of the relative social feature; Using the two-dimensional convolutional layer of the convolutional neural network and introducing a grid mask to encode the relative social features to obtain static features. Among them, the grid mask is used to discretize the static scene space neighborhood depicted by the filter in the convolutional neural network into square grids, and the square grids are used to represent spatial relationships. The static features are expressed as: C map = CNN(f S , h s ) * M, C map is a static feature, CNN is a convolutional neural network, and M is a grid mask; Using a single GridLSTM cell GLSTM O and static features to encode the pedestrian head pose feature vector to obtain visual space features, where the visual space features are represented as: f o is a visual spatial feature, h o is the hidden state of the visual spatial feature, is the pedestrian head pose feature vector; Combining the relative social features and the visual space features to generate a neighborhood representing the pedestrian scenario spatio-temporal interaction feature vector. The representation formula of the neighborhood is: F = f s *f o , F is the neighborhood characterizing the spatio-temporal interaction feature vector of the pedestrian scenario, and f s is the relative social feature, and f o is the visual space feature.
3. The method according to claim 2, characterized in that, The predicting the trajectory trend of the target pedestrian within a preset future time period based on the social trajectory recommender and the pedestrian scenario spatio-temporal interaction feature vector includes: Predicting the social interaction state and the future position of the target pedestrian based on the kernel of the social trajectory recommender and the relative social features. Based on the predicted future position of the target pedestrian, obtaining the prediction result of the trajectory trend of the target pedestrian within the preset future time period. The prediction formula is as follows: is the predicted future position of the target pedestrian, H t+1 is the social interaction state, K is the kernel of the social trajectory recommender, f s is the relative social feature, is the matrix assigned to the weighted hidden state.
4. The method according to claim 3, wherein The obtaining the prediction result of the trajectory trend of the target pedestrian within the preset future time period based on the predicted future position of the target pedestrian includes: Based on the predicted future position of the target pedestrian, obtaining a future trajectory sequence of the target pedestrian, and using the future trajectory sequence of the target pedestrian as the prediction result of the trajectory trend within the preset future time period. Among them, the future trajectory sequence of the target pedestrian is a tensor of N×K×T×2. N represents the number of target pedestrians, K represents the number of modalities of the predicted trajectory, T represents the number of prediction time steps, and 2 represents the two-dimensional coordinates (x, y) of each time step. The time step is the sampling moment.
5. The method according to claim 3, characterized in that After predicting the social interaction state and the future position of the target pedestrian based on the kernel of the social trajectory recommender and the relative social features, the method further includes: Based on the neighborhood and scaled self-attention mechanism that represents the spatio-temporal interaction feature vector of the pedestrian scenario, calculating the soft attention weights, updating the spatio-temporal interaction feature vector of the pedestrian scenario based on the calculated soft attention weights, and re-evaluating the intensity of the social interaction state, where the scaled self-attention mechanism is expressed as: a is the soft attention weight, calculated by the Softmax function, which is to exponentiate the neighborhood F.
6. The method according to claim 5, wherein When updating the spatio-temporal interaction feature vector of the pedestrian scenario and re-evaluating the intensity of the social interaction state based on the calculated soft attention weights, the method further includes: Introducing non-negative matrix factorization, where the non-negative matrix factorization is: A P = W a / max(W a ) W a is the matrix assigned to the attention weight a, H t is the matrix assigned to the weighted hidden state, T is the time step, NMF is the non - negative matrix factorization function, which is used to factorize the attention weight a into two non - negative matrices W a and H t , A P is the normalized value of the attention weight.
7. The method according to any one of claims 1 to 6, characterized in that The head pose is characterized by the orientation angle of the target pedestrian's head in the plane, and the map information includes at least one of road information, special area information, traffic signal information, and other map information.
8. The method according to claim 7, wherein The method further includes: Based on the social trajectory recommender and the spatio-temporal interaction feature vector of the pedestrian scenario, for three trajectory influencing factors that affect the pedestrian trajectory, respectively predicting the sub-factor that has the greatest impact on the trajectory of the target pedestrian in each trajectory influencing factor, and based on the predicted sub-factors with the greatest impact, jointly determining the trajectory trend of the target pedestrian within a preset future time period, where the trajectory influencing factors include other pedestrians, scene space, and the head pose of the target pedestrian. Among other pedestrians, each sub-factor is each other person except the target pedestrian in the target area. Among the scene space, each sub-factor is various scene space elements in the scene space. Among the head poses of the target pedestrian, each sub-factor is the head pose of the target pedestrian at each sampling moment within a preset historical time period.
9. A storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the method for predicting pedestrian trajectories according to any one of claims 1 to 8.
10. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for predicting pedestrian trajectories according to any one of claims 1 to 8.