Pedestrian motion trajectory prediction method, system and storage medium based on visual information
Through drones collecting visual information and combining GIN and CNN technologies, a directional topology map of global interaction is constructed, which solves the problem of underutilization of visual information in the existing technology and achieves more efficient and accurate pedestrian motion prediction.
Patent Information
- Application Number
- CN202510161124.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-13
AI Technical Summary
The prior art fails to make full use of visual information in pedestrian motion prediction, resulting in insufficient prediction accuracy in complex scenarios and dynamic interactions.
UAVs are used to collect visual information, combine graph isomorphic networks (GINs) and convolutional neural networks (CNNs), and build global interaction directed topology maps, extract interaction characteristics and time series characteristics between pedestrians to predict pedestrian movement trajectory.
It significantly improves the accuracy and real-time performance of pedestrian motion prediction, enhances the adaptability and application effect of the model in complex environments, and optimizes pedestrian flow management and public safety planning.
Smart Images

Figure CN119625842B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent transportation technology, and in particular to a method, system and storage medium for predicting pedestrian motion trajectories based on visual information. Background Art
[0002] Walking is the most sustainable and basic mode of travel in urban transportation and an indispensable part of the urban transportation system. With the rapid development of social economy and the gradual increase in urban population density, the safety and efficiency of pedestrian travel are facing increasingly complex challenges, and the safety management of public places is becoming more and more important. Therefore, how to effectively manage and optimize pedestrian traffic and ensure public safety has become one of the focal issues of social concern. Advanced pedestrian movement prediction technology can provide a scientific basis for the formulation of traffic control strategies, and is an effective means of traffic planning, design and management. It has broad value in scientific research and practical applications.
[0003] With the rapid development of computer technology, various pedestrian modeling methods have emerged, providing strong technical support for the prediction of pedestrian motion. Traditional physical rule-based models, such as social force models and cellular automaton models, have been widely used to simulate the microscopic behavior of pedestrians. In fact, the prediction of pedestrian motion is a complex process. Unlike vehicles running on fixed lanes, pedestrians, as intelligent autonomous entities, have a high degree of randomness, flexibility and complexity in their behavior. These factors make it difficult to achieve the accuracy of microscopic modeling based on physical rules. In addition, the movement speed of pedestrians presents more complex changes based on microscopic behavior, especially in complex scenes, such as when there are obstacles or irregular boundaries, pedestrian motion is more difficult to predict. With the rise of artificial intelligence and data-driven methods, pedestrian motion prediction models based on deep learning have gradually become a research hotspot. Such models can be trained with a large amount of measured data to more accurately capture the complexity and diversity of pedestrian motion, thereby achieving more accurate pedestrian motion prediction.
[0004] Although various physics-based and deep learning models have made significant progress, they have failed to fully utilize visual cues such as relative angles and distances between pedestrians, and have ignored the importance of visual information for pedestrian motion prediction. As one of the most important perception systems for pedestrians in motion, vision not only plays an important role in navigation and obstacle avoidance, but also plays a key role in the interaction between pedestrians. Vision helps pedestrians make dynamic decisions by providing real-time feedback about the environment, obstacles, and other pedestrian positions, and adjusts speed and direction according to the surrounding environment and pedestrian behavior.
[0005] Prior art patent No. 2021112172418 describes a pedestrian trajectory prediction system and method based on a convolutional neural network, which uses a convolutional network to process time series data. This patent proposes a prediction model based on historical data, which applies a convolution-based feature extraction network to predict the future trajectory of pedestrians. The core of this method is that through the previous data of the sequence, the model can have a general understanding of the pedestrian's potential behavioral awareness and subjective intentions. There are also existing models based on physical rules, such as social force models and cellular automaton models. Although they can provide reasonable predictions in some scenarios, they are all limited by mathematical formulas and fixed rules, and will inevitably produce deviations in predicting trajectories in real-world scenarios. Especially in complex environments and dynamic interactions, these models may not accurately reflect the real behavior patterns of pedestrians.
[0006] Compared with models based on physical rules, models based on deep learning can learn more valuable information from a large amount of measured data, thereby more accurately capturing the complexity and diversity of pedestrian movements. Models based on deep learning are divided into sequence learning models and structure learning models. The former mainly relies on time series data, focuses on the attributes of a single pedestrian, usually only considers the characteristics of pedestrians in the time dimension, and does not combine the interaction between pedestrians with the deep learning algorithm module. In fact, pedestrians will be affected by neighboring pedestrians during their movement, and incorporating the interaction between pedestrians is crucial for accurately predicting their movements. Compared with sequence learning models, structure learning models not only focus on the individual characteristics of each pedestrian, but also integrate the dynamic changes of the surrounding environment and other pedestrians through network structures. However, the existing structure learning methods only use coordinate data when constructing the network structure, resulting in very limited feature information that the network structure can express, ignoring the importance of pedestrian visual information, and making it difficult to capture more accurate interactions between pedestrians. Summary of the invention
[0007] In order to solve the problems existing in the prior art, the present invention provides a method, system and storage medium for predicting pedestrian motion trajectories based on visual information, integrating visual information collected by drones, and combining graph isomorphism networks (GINs) and convolutional neural networks (CNNs) to respectively extract interaction features between pedestrians and capture motion patterns in time series. Through this method, pedestrian motion in complex scenes can be predicted more efficiently, enhancing the application effect of the model in pedestrian flow management and public safety planning. The method and system of the present invention predict future motion situations by analyzing pedestrian historical trajectories, significantly optimizing pedestrian flow management and improving the safety level of public places, solving the problems mentioned in the above background technology.
[0008] To achieve the above object, the present invention provides the following technical solution: a method for predicting pedestrian motion trajectory based on visual information, comprising the following steps:
[0009] S1. UAV data collection: collect data through UAV and standardize the data;
[0010] S2, coordinate projection transformation: restore the pixel coordinates in the video to the actual ground truth coordinates through the multi-grid projection transformation algorithm;
[0011] S3, coordinate sampling and trajectory aggregation: sampling the ground truth coordinates of pedestrians after projection transformation, and aggregating the coordinate data into trajectory information of multiple pedestrians;
[0012] S4. Construction of global interactive directed topological graph based on visual information: Combining spatial visual information with temporal information to construct a global interactive directed topological graph, thereby improving the model's ability to capture spatial interaction features in complex environments;
[0013] S5. Spatial feature extraction: A spatial feature extraction module is constructed using a graph isomorphic network (GIN), which adaptively learns directed edge weights based on visual information and aggregates the features of neighboring nodes to update the current node features, capturing the spatial interaction relationship between pedestrians and their neighbors.
[0014] S6. Temporal feature extraction: Use a two-dimensional convolutional neural network (CNN) to build a temporal feature extraction module. Through a multi-layer convolution structure, the temporal series features of pedestrian motion are efficiently extracted, and finally the future motion trajectory of pedestrians is predicted.
[0015] Preferably, in step S1, it specifically includes:
[0016] S11. Hover the drone at a predetermined height to ensure panoramic coverage of all pedestrian movements in the scene;
[0017] S12, obtaining the internal parameters of the drone gimbal camera, including focal length, principal point offset coordinates, radial distortion coefficient, and tangential distortion coefficient;
[0018] S13, obtaining the external parameters of the gimbal camera, including the rotation matrix and the translation vector;
[0019] S14, measuring known ground truth coordinate points in the scene;
[0020] S15. Use automated annotation software to calibrate the camera’s internal and external parameters and annotate the pixel coordinates in the pedestrian video. ,The coordinate data is normalized, including pedestrian number, time frame, horizontal coordinate and vertical coordinate data.
[0021] Preferably, in step S2, it specifically includes:
[0022] S21. Divide the scene into Grid Area ;
[0023] S22, according to the ground truth coordinate point measured in step S14 The corresponding video pixel coordinate point , calculate each grid area The perspective transformation matrix in , The expression is as follows:
[0024] ;
[0025] Among them, the projection transformation mapping coefficient to By solving the following system of equations:
[0026] ;
[0027] S23, according to the boundary information of each grid, all the collected video pixel coordinates are Assign to the corresponding grid middle;
[0028] S24. After determining the grid to which it belongs, use the perspective transformation matrix of the corresponding grid Perform a projection transformation to obtain the ground truth coordinates of the pedestrian , the transformation is expressed as:
[0029] ;
[0030] S25. Normalize the converted pedestrian ground coordinate data, specifically including the pedestrian number, time frame, horizontal coordinate and vertical coordinate fields, to ensure the consistency of the data format and make it suitable for the input of the subsequent deep learning model.
[0031] Preferably, in step S3, it specifically includes:
[0032] S31, for the coordinate data obtained in step S25, first calculate the coordinates from the first time step to the time steps to perform observation sampling and The trajectories of pedestrians are aggregated and the observed trajectories are The representation is:
[0033] ;
[0034] in, Indicates pedestrian At the moment The coordinates of
[0035] S32, from arrive The trajectories of future time steps are sampled and aggregated to serve as the target values predicted by the deep learning model. Future Trajectories The representation is:
[0036] ;
[0037] in, Indicates pedestrian At the moment The predicted coordinates.
[0038] Preferably, in step S4, it specifically includes:
[0039] S41, according to the pedestrian at the time and The displacement difference , calculate the arc value of the pedestrian's movement direction , the formula is as follows:
[0040] ;
[0041] S42: the arc value of the pedestrian's moving direction Convert to angle and adjust to to range, get pedestrians At the moment Direction of movement , the calculation formula is as follows:
[0042] ;
[0043] S43. Calculate neighbors Relative to pedestrians Angle , through the coordinate difference between the two and Calculate the radian value of the relative angle and convert it to an angle using the following formula:
[0044] ;
[0045] S44, calculation time The Euclidean distance between each pair of pedestrians , the formula is as follows:
[0046] ;
[0047] S45. Use the multi-layer perceptron network MLP to calculate the relative angles between pedestrians. and relative distance These two visual information are fused into the directed edge weight vector , indicating pedestrians With its neighbors At the moment The interactive relationship is calculated as follows:
[0048] ;
[0049] in for The weight parameter of
[0050] S46. Directed edge weight vector Aggregate into moments The edge weight set of the global directed topological subgraph , its expression formula is:
[0051] ;
[0052] S47, pedestrians At the moment Coordinates Embedded as node vector , and aggregate it to the time The node set of the global directed topological subgraph , whose expression is:
[0053] ;
[0054] S48, node collection and the edge weight set Aggregate into moments A directed subgraph of Finally, a global interactive directed topological graph based on visual information It is composed of directed subgraphs at multiple moments, and its expression is:
[0055] .
[0056] Preferably, in step S5, the following is specifically included:
[0057] S51, the node set at time t Perform linear transformation to obtain a new set of node feature vectors , the calculation process is as follows:
[0058] ;
[0059] in is the weight matrix of linear transformation;
[0060] S52. Use node feature vector set and the set of directed edge weights Perform feature aggregation, and the calculation process is as follows:
[0061] ;
[0062] in, Represents the temporary feature vector of the node after feature aggregation, Representation Node Neighbor node set;
[0063] S53, use the multi-layer perceptron network to perform nonlinear mapping to update node features, the updating process is as follows:
[0064] ;
[0065] in, represents the updated feature vector after MLP processing, is a learnable parameter, for The weight parameter of
[0066] S54, use the activation function ELU to further update the node features to obtain the final node features:
[0067] ;
[0068] The definition of the activation function ELU is as follows:
[0069] ;
[0070] in represents the input vector of ELU, for The exponential function of .
[0071] Preferably, in step S6, the following is specifically included:
[0072] S61, reshape the final node features outputted in step S54, introduce new feature dimensions, and aggregate them into a new feature vector ,in , represents the characteristic length, It is a new dimension introduced during the reshaping process, which gradually increases with the number of convolutional layers;
[0073] S62, reshaped feature vector Use multiple two-dimensional convolution operations to extract features in the time dimension; the calculation process of each convolution layer is as follows:
[0074] ;
[0075] in is the output of the convolutional layer, represents the kernel weight, Indicates the padding size, To calculate the index;
[0076] S63. After the two-dimensional convolution operation, the batch normalization operation is used to process the output of each convolution layer to standardize the features of each channel so that its mean is close to 0 and its standard deviation is close to 1, thereby improving the convergence speed and robustness of the model. The calculation formula of the batch normalization operation is as follows:
[0077] ;
[0078] in represents the input vector of batch normalization, represents the mean value of the feature matrix, represents the standard deviation of the current feature matrix, and are the scaling and shifting parameters learned during training, respectively;
[0079] S64. Use the ELU activation function defined in step S54 to perform nonlinear activation on the feature vector after the batch normalization operation, thereby improving the expressive power of the model.
[0080] S65. In the process of time series feature extraction, the channel attention mechanism is introduced to perform weight allocation in the time dimension, so that the model can adaptively assign different weights to the trajectory features at each moment and learn the important features at a specific moment. The specific weight allocation process is as follows:
[0081] ;
[0082] ;
[0083] in Represent the input and output features of the channel attention mechanism, respectively. is the channel attention feature, It is the channel feature transformation achieved through convolution operation;
[0084] S66. After completing the node feature extraction, the fully connected layer is finally introduced to map the high-dimensional feature vector to the low-dimensional coordinate space of the target trajectory and output the predicted future motion trajectory. .
[0085] On the other hand, to achieve the above-mentioned purpose, the present invention also provides the following technical solution: a pedestrian motion trajectory prediction system based on visual information, the system comprising the following modules:
[0086] UAV data collection module: collect data through UAV and standardize the data;
[0087] Coordinate projection transformation module: restores the pixel coordinates in the video to the actual ground truth coordinates through a multi-grid projection transformation algorithm;
[0088] Coordinate sampling and trajectory aggregation module: samples the ground truth coordinates of pedestrians after projection transformation, and aggregates the coordinate data into trajectory information of multiple pedestrians;
[0089] Global interactive directed topological graph construction module: combines spatial visual information with temporal information to construct a global interactive directed topological graph, thereby improving the model's ability to capture spatial interaction features in complex environments;
[0090] Spatial feature extraction module: The spatial feature extraction module is constructed using the graph isomorphic network GIN, which adaptively learns the weights of directed edges based on visual information and aggregates the features of neighboring nodes to update the features of the current node, capturing the spatial interaction relationship between pedestrians and their neighbors;
[0091] Temporal feature extraction module: A two-dimensional convolutional neural network (CNN) is used to construct a temporal feature extraction module. The multi-layer convolution structure is used to efficiently extract the time series features of pedestrian motion and ultimately predict the future motion trajectory of pedestrians.
[0092] On the other hand, to achieve the above-mentioned purpose, the present invention also provides the following technical solution: an electronic device, the electronic device comprising: a processor; and a memory for storing one or more programs;
[0093] When the one or more programs are executed by the processor, the processor executes the pedestrian motion trajectory prediction method based on visual information.
[0094] On the other hand, to achieve the above-mentioned purpose, the present invention also provides the following technical solution: a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the pedestrian motion trajectory prediction method based on visual information.
[0095] The beneficial effects of the present invention are: the present invention combines deep learning technology to efficiently and accurately predict pedestrian movement trajectories, showing significant technical advantages. The method fully considers the capture of visual information, the modeling of pedestrian interaction relationships, and the fusion of spatiotemporal features, thereby achieving outstanding results in improving prediction accuracy and enhancing model adaptability. The specific technical advantages are as follows:
[0096] 1) Accurate pedestrian motion prediction: By combining visual information and deep learning models, this invention can accurately capture the motion trajectory and behavior patterns of pedestrians. Compared with other methods, this invention makes full use of the visual information of pedestrians and can more comprehensively consider the interactions between pedestrians, thereby providing more accurate trajectory prediction.
[0097] 2) Efficient computing performance and real-time performance: This invention optimizes the structure and algorithm of the deep learning model to ensure that the model can maintain efficient computing performance under large-scale data input. While ensuring the accuracy of prediction, the system can process massive data in real time, providing strong technical support for pedestrian motion prediction in complex environments.
[0098] 3) Real-time prediction and warning: The present invention can predict pedestrian movement in real time and promptly identify potential crowded or dangerous areas. The system can issue early warnings and safety reminders when pedestrians are overcrowded or abnormal, effectively avoiding the occurrence of safety hazards.
[0099] 4) Scalability and flexibility: The method of the present invention has good scalability and can adjust and optimize the model according to specific needs. For example, the frequency of data collection and the time step of prediction can be adjusted according to the needs of different scenarios, so that the method can be flexibly adapted to different application scenarios and has strong customization capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0100] Figure 1 Schematic diagram of the steps of a method for predicting pedestrian motion trajectory based on visual information in an embodiment of the present invention;
[0101] Figure 2 Schematic diagram of a pedestrian motion trajectory prediction module based on visual information in an embodiment of the present invention;
[0102] Figure 3 A schematic diagram of a directed topological graph and a feature extraction module in an embodiment of the present invention;
[0103] Figure 4 This is a schematic diagram of the structure of an electronic device in an embodiment of the present invention;
[0104] In the figure, 110-UAV data acquisition module; 120-coordinate projection transformation module; 130-coordinate sampling and trajectory aggregation module; 140-global interactive directed topology graph construction module; 150-spatial feature extraction module; 160-temporal feature extraction module; 210-processor; 220-storage. DETAILED DESCRIPTION
[0105] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0106] See also Figure 1 The present invention provides a technical solution: a pedestrian motion trajectory prediction method based on visual information, the method flow is as follows Figure 1 As shown, the following steps are included:
[0107] S1. UAV data collection: Collect data through UAVs and standardize the data. Specifically include:
[0108] S11. Hover the drone at a predetermined height to ensure panoramic coverage of all pedestrian movements in the scene and provide high-quality observation perspective.
[0109] S12, obtaining the internal parameters of the drone gimbal camera, including focal length, principal point offset coordinates, radial distortion coefficient and tangential distortion coefficient, for subsequent image calibration.
[0110] S13. Obtain the external parameters of the gimbal camera, including the rotation matrix and the translation vector, which are used to describe the spatial position and direction of the camera relative to the scene.
[0111] S14, measuring known ground truth coordinate points in the scene, so as to facilitate coordinate calibration using a subsequent projection transformation method.
[0112] S15. Use automated annotation software to calibrate the camera’s internal and external parameters and annotate the pixel coordinates in the pedestrian video. ,The coordinate data is normalized, including pedestrian number, time frame, horizontal coordinate and vertical coordinate data.
[0113] S2, coordinate projection transformation: restore the pixel coordinates in the video to the actual ground truth coordinates through the multi-grid projection transformation algorithm. Specifically include:
[0114] S21. Divide the scene into Grid Area ;
[0115] S22, according to the ground truth coordinate point measured in step S14 The corresponding video pixel coordinate point , calculate each grid area The perspective transformation matrix in , The expression is as follows:
[0116] ;
[0117] Among them, the projection transformation mapping coefficient to By solving the following system of equations:
[0118] ;
[0119] S23, according to the boundary information of each grid, all the collected video pixel coordinates are Assign to the corresponding grid middle;
[0120] S24. After determining the grid to which it belongs, use the perspective transformation matrix of the corresponding grid Perform a projection transformation to obtain the ground truth coordinates of the pedestrian , the transformation is expressed as:
[0121] ;
[0122] S25. Normalize the converted pedestrian ground coordinate data, specifically including the pedestrian number, time frame, horizontal coordinate and vertical coordinate fields.
[0123] S3, coordinate sampling and trajectory aggregation: sampling the ground truth coordinates of pedestrians after projection transformation, and aggregating the coordinate data into trajectory information of multiple pedestrians. Specifically including:
[0124] S31, for the coordinate data obtained in step S25, first calculate the coordinates from the first time step to the time steps to perform observation sampling and The trajectories of pedestrians are aggregated and the observed trajectories are The representation is:
[0125] ;
[0126] in, Indicates pedestrian At the moment The coordinates of
[0127] S32, from arrive The trajectories of future time steps are sampled and aggregated, and the future trajectories The representation is:
[0128] ;
[0129] in, Indicates pedestrian At the moment The predicted coordinates.
[0130] S4. Construction of a global interactive directed topology graph based on visual information: Combine spatial visual information with temporal information to construct a global interactive directed topology graph, thereby improving the model's ability to capture spatial interaction features in complex environments. Specifically, it includes:
[0131] S41, according to the pedestrian at the time and The displacement difference , calculate the arc value of the pedestrian's movement direction , the formula is as follows:
[0132] ;
[0133] S42: the arc value of the pedestrian's moving direction Convert to angle and adjust to to range, get pedestrians At the moment Direction of movement , the calculation formula is as follows:
[0134] ;
[0135] S43. Calculate neighbors Relative to pedestrians Angle , through the coordinate difference between the two and Calculate the radian value of the relative angle and convert it to an angle using the following formula:
[0136] ;
[0137] S44, calculation time The Euclidean distance between each pair of pedestrians , the formula is as follows:
[0138] ;
[0139] S45. Use the multi-layer perceptron network MLP to calculate the relative angles between pedestrians. and relative distance These two visual information are fused into the directed edge weight vector , indicating pedestrians With its neighbors At the moment The interactive relationship is calculated as follows:
[0140] ;
[0141] in for The weight parameter of
[0142] S46. Directed edge weight vector Aggregate into moments The edge weight set of the global directed topological subgraph , its expression formula is:
[0143] ;
[0144] S47, pedestrians At the moment Coordinates Embedded as node vector , and aggregate it to the time The node set of the global directed topological subgraph , whose expression is:
[0145] ;
[0146] S48, node collection and the edge weight set Aggregate into moments A directed subgraph of Finally, a global interactive directed topological graph based on visual information It is composed of directed subgraphs at multiple moments, and its expression is:
[0147] .
[0148] S5. Spatial feature extraction: Use the graph isomorphic network GIN to build a spatial feature extraction module, adaptively learn the weights of directed edges based on visual information, and aggregate the features of neighboring nodes to update the current node features, so as to more accurately capture the spatial interaction relationship between pedestrians and their neighbors. Specifically, it includes the following:
[0149] S51, the node set at time t Perform linear transformation to obtain a new set of node feature vectors , the calculation process is as follows:
[0150] ;
[0151] in is the weight matrix of linear transformation;
[0152] S52. Use node feature vector set and the set of directed edge weights Perform feature aggregation, and the calculation process is as follows:
[0153] ;
[0154] in, Represents the temporary feature vector of the node after feature aggregation, Representation Node Neighbor node set;
[0155] S53, use the multi-layer perceptron network to perform nonlinear mapping to update node features, the updating process is as follows:
[0156] ;
[0157] in, represents the updated feature vector after MLP processing, is a learnable parameter, for The weight parameter of
[0158] S54, use the activation function ELU to further update the node features to obtain the final node features:
[0159] ;
[0160] The definition of the activation function ELU is as follows:
[0161] ;
[0162] in represents the input vector of ELU, for The exponential function of .
[0163] S6. Temporal feature extraction: Use a two-dimensional convolutional neural network (CNN) to build a temporal feature extraction module, and efficiently extract the temporal series features of pedestrian motion through a multi-layer convolution structure, and ultimately predict the future motion trajectory of pedestrians. Specifically, it includes the following:
[0164] S61, reshape the final node features outputted in step S54, introduce new feature dimensions, and aggregate them into a new feature vector ,in , represents the characteristic length, It is a new dimension introduced during the reshaping process, which gradually increases with the number of convolutional layers;
[0165] S62, reshaped feature vector Use multiple two-dimensional convolution operations to extract features in the time dimension; the calculation process of each convolution layer is as follows:
[0166] ;
[0167] in is the output of the convolutional layer, represents the kernel weight, Indicates the padding size, To calculate the index;
[0168] S63. After the two-dimensional convolution operation, the batch normalization operation is used to process the output of each convolution layer to standardize the features of each channel. The calculation formula of the batch normalization operation is as follows:
[0169] ;
[0170] in represents the input vector of batch normalization, represents the mean value of the feature matrix, represents the standard deviation of the current feature matrix, and are the scaling and shifting parameters learned during training, respectively;
[0171] S64, using the ELU activation function defined in step S54 to perform nonlinear activation on the feature vector after the batch normalization operation;
[0172] S65. In the process of time series feature extraction, the channel attention mechanism is introduced to perform weight distribution in the time dimension. The specific weight distribution process is as follows:
[0173] ;
[0174] ;
[0175] in Represent the input and output features of the channel attention mechanism, respectively. is the channel attention feature, It is the channel feature transformation achieved through convolution operation;
[0176] S66. After completing the node feature extraction, the fully connected layer is finally introduced to map the high-dimensional feature vector to the low-dimensional coordinate space of the target trajectory and output the predicted future motion trajectory. .
[0177] Based on the same inventive concept as the above method embodiment, the present application embodiment also provides a system for predicting pedestrian motion trajectory based on visual information, which can implement the functions provided by the above method embodiment, such as Figure 2 As shown, the system includes the following modules:
[0178] The drone data collection module 110 collects data through the drone and performs standardized processing on the data. The drone hovers above the scene to collect the panoramic video data of pedestrians, and after completing the internal and external parameter calibration, the pixel coordinates of the pedestrians in the video are extracted using the automatic annotation software.
[0179] Coordinate projection transformation module 120: restores the pixel coordinates in the video to the actual ground truth coordinates through a multi-grid projection transformation algorithm to obtain the real pedestrian position.
[0180] Coordinate sampling and trajectory aggregation module 130: samples the ground truth coordinates of pedestrians after projection transformation, and aggregates the coordinate data into trajectory information of multiple pedestrians.
[0181] Global interactive directed topological graph construction module 140: combines spatial visual information with temporal information to construct a global interactive directed topological graph, thereby improving the model's ability to capture spatial interactive features in complex environments; providing rich high-level spatiotemporal feature information for subsequent feature extraction modules, such as Figure 3 shown.
[0182] Spatial feature extraction module 150: A spatial feature extraction module is constructed using a graph isomorphism network GIN, which adaptively learns directed edge weights based on visual information and aggregates the features of neighboring nodes to update the current node features, capturing the spatial interaction relationship between pedestrians and their neighbors.
[0183] Temporal feature extraction module 160: A temporal feature extraction module is constructed using a two-dimensional convolutional neural network (CNN). The temporal series features of pedestrian motion are efficiently extracted through a multi-layer convolution structure, and the future motion trajectory of pedestrians is finally predicted.
[0184] The trajectory features are temporally encoded through a convolutional neural network (CNN), the number of channels is increased layer by layer to enhance the expression of temporal features, and a channel attention mechanism is introduced to enable the model to adaptively assign weights to trajectory features at different moments. Finally, the high-level feature information is mapped to the future motion trajectory of the pedestrian through a fully connected layer, and the predicted pedestrian motion trajectory is output.
[0185] Based on the same inventive concept as the above method embodiment, the present application embodiment also provides an electronic device, such as Figure 4 As shown, the device includes: a processor 210; and a memory 220, for storing one or more programs;
[0186] When the one or more programs are executed by the processor 210, the processor executes the pedestrian motion trajectory prediction method based on visual information.
[0187] The pedestrian motion trajectory prediction method based on visual information includes the following:
[0188] Drone data collection: collect data through drones and standardize the data;
[0189] Coordinate projection transformation: The pixel coordinates in the video are restored to the actual ground truth coordinates through a multi-grid projection transformation algorithm;
[0190] Coordinate sampling and trajectory aggregation: Sample the ground truth coordinates of pedestrians after projection transformation, and aggregate the coordinate data into trajectory information of multiple pedestrians;
[0191] Construction of a global interactive directed topological graph based on visual information: Combining spatial visual information with temporal information to construct a global interactive directed topological graph, thereby improving the model's ability to capture spatial interaction features in complex environments;
[0192] Spatial feature extraction: A spatial feature extraction module is constructed using a graph isomorphic network (GIN), which adaptively learns directed edge weights based on visual information and aggregates the features of neighboring nodes to update the current node features, capturing the spatial interaction relationship between pedestrians and their neighbors.
[0193] Temporal feature extraction: A two-dimensional convolutional neural network (CNN) is used to construct a temporal feature extraction module. The multi-layer convolution structure is used to efficiently extract the time series features of pedestrian motion and ultimately predict the future motion trajectory of pedestrians.
[0194] Based on the same inventive concept as the above method embodiment, the embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by the processor 210, the pedestrian motion trajectory prediction method based on visual information is implemented.
[0195] The pedestrian motion trajectory prediction method based on visual information includes the following:
[0196] Drone data collection: collect data through drones and standardize the data;
[0197] Coordinate projection transformation: The pixel coordinates in the video are restored to the actual ground truth coordinates through a multi-grid projection transformation algorithm;
[0198] Coordinate sampling and trajectory aggregation: Sample the ground truth coordinates of pedestrians after projection transformation, and aggregate the coordinate data into trajectory information of multiple pedestrians;
[0199] Construction of a global interactive directed topological graph based on visual information: Combining spatial visual information with temporal information to construct a global interactive directed topological graph, thereby improving the model's ability to capture spatial interaction features in complex environments;
[0200] Spatial feature extraction: A spatial feature extraction module is constructed using a graph isomorphic network (GIN), which adaptively learns directed edge weights based on visual information and aggregates the features of neighboring nodes to update the current node features, capturing the spatial interaction relationship between pedestrians and their neighbors.
[0201] Temporal feature extraction: A two-dimensional convolutional neural network (CNN) is used to construct a temporal feature extraction module. The multi-layer convolution structure is used to efficiently extract the time series features of pedestrian motion and ultimately predict the future motion trajectory of pedestrians.
[0202] The method of the present invention can more efficiently predict pedestrian movement in complex scenes, and enhance the application effect of the model in pedestrian flow management and public safety planning. The system predicts future movement by analyzing pedestrian historical trajectories, significantly optimizing pedestrian flow management and improving the safety level of public places. It can be widely used in intelligent transportation systems, accident prevention, crowd management and other aspects.
[0203] In several embodiments provided in the embodiments of the present invention, it should be understood that the disclosed apparatus and method can also be implemented in other ways. The apparatus and method embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the apparatus, method and computer program product according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0204] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.
[0205] If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code. It should be noted that in this article, the term "include", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such process, method, article or device. Without more constraints, an element defined by the phrase "comprising a..." does not exclude the existence of other identical elements in the process, method, article or apparatus comprising the element.
[0206] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.
[0207] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0208] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0209] The "first\second" mentioned in the embodiments is only to distinguish similar objects, and does not represent a specific order for the objects. It is understandable that the "first\second" can be interchanged with the specific order or sequence where permitted. It should be understood that the objects distinguished by "first\second" can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.
[0210] Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A pedestrian motion trajectory prediction method based on visual information, characterized in that: The steps include: S1. UAV data collection: collect data through UAV and standardize the data; S2, coordinate projection transformation: restore the pixel coordinates in the video to the actual ground truth coordinates through the multi-grid projection transformation algorithm; S3, coordinate sampling and trajectory aggregation: sampling the ground truth coordinates of pedestrians after projection transformation, and aggregating the coordinate data into trajectory information of multiple pedestrians; S4. Construction of a global interactive directed topological graph based on visual information: Combine spatial visual information with temporal information to construct a global interactive directed topological graph, thereby improving the model's ability to capture spatial interaction features in complex environments; specifically, the following are included: S41, based on the displacement difference between the pedestrian at time t and t-1 and Calculate the arc value of the pedestrian's movement direction The formula is as follows: S42: the arc value of the pedestrian's moving direction Convert to angle and adjust to the range of 0° to 360° to get the movement direction of pedestrian i at time t The calculation formula is as follows: S43. Calculate the angle of neighbor j relative to pedestrian i Through the difference of the two coordinates and Calculate the radian value of the relative angle and convert it to an angle using the following formula: S44. Calculate the Euclidean distance between each pair of pedestrians at time t The formula is as follows: S45. Use the multi-layer perceptron network MLP to calculate the relative angles between pedestrians. and relative distance These two visual information are fused into the directed edge weight vector It is expressed as the interaction relationship between pedestrian i and its neighbor j at time t, and the calculation formula is as follows: Where W e is the weight parameter of MLP; S46. Directed edge weight vector Aggregate into the edge weight set E of the global directed topological subgraph at time t t , its expression formula is: S47, the coordinates of pedestrian i at time t Embedded as node vector And aggregate it to the node set V of the global directed topological subgraph at time t t , whose expression is: S48, the node set V t and the edge weight set E t Aggregate into a directed subgraph G at time t t =(V t ,E t ), finally, the global interactive directed topological graph G based on visual information is composed of directed subgraphs at multiple moments, and its expression is: G={G t |t=1,2,…,T obs }; S5. Spatial feature extraction: A spatial feature extraction module is constructed using a graph isomorphic network (GIN), which adaptively learns directed edge weights based on visual information and aggregates the features of neighboring nodes to update the current node features, capturing the spatial interaction relationship between pedestrians and their neighbors. S6. Temporal feature extraction: Use a two-dimensional convolutional neural network (CNN) to build a temporal feature extraction module. Through a multi-layer convolution structure, the temporal series features of pedestrian motion are efficiently extracted, and finally the future motion trajectory of pedestrians is predicted.
2. The method for predicting pedestrian motion trajectory based on visual information according to claim 1, characterized in that: In step S1, it specifically includes: S11. Hover the drone at a predetermined height to ensure panoramic coverage of all pedestrian movements in the scene; S12, obtaining the internal parameters of the drone gimbal camera, including focal length, principal point offset coordinates, radial distortion coefficient, and tangential distortion coefficient; S13, obtaining the external parameters of the gimbal camera, including the rotation matrix and the translation vector; S14, measuring known ground truth coordinate points in the scene; S15. Use automated labeling software to calibrate the internal and external parameters of the camera and label the pixel coordinates (x′, y′) in the pedestrian video, and normalize the coordinate data, including the pedestrian number, time frame, horizontal coordinate, and vertical coordinate data.
3. The method for predicting pedestrian motion trajectory based on visual information according to claim 1, characterized in that: In step S2, it specifically includes: S21, divide the scene into n grid areas a according to the size of the scene μ ; S22, according to the ground truth coordinate point measured in step S14 The corresponding video pixel coordinate point Calculate each grid area a μ The perspective transformation matrix M in μ , M μ The expression is as follows: Among them, the projection transformation mapping coefficient to By solving the following system of equations: S23, according to the boundary information of each grid, all the collected video pixel coordinates (x′, y′) are assigned to the corresponding grid μ; S24, after determining the grid to which it belongs, use the perspective transformation matrix M of the corresponding grid μ Perform a projection transformation to obtain the ground truth coordinates (x, y) of the pedestrian. The transformation is expressed as: S25. Normalize the transformed pedestrian ground coordinate data, specifically including the pedestrian number, time frame, horizontal coordinate and vertical coordinate fields.
4. The method for predicting pedestrian motion trajectory based on visual information according to claim 1, characterized in that: In step S3, it specifically includes: S31, for the coordinate data obtained in step S25, first calculate the coordinates from the first time step to the Tth time step. obs The observation sampling is performed in time steps, and the trajectories of N pedestrians in the same time period are aggregated to observe the trajectory. The representation is: in, represents the coordinates of pedestrian i at time t; S32, from T obs+1 to T pred The trajectories of future time steps are sampled and aggregated, and the future trajectories The representation is: in, represents the predicted coordinates of pedestrian i at time t.
5. The method for predicting pedestrian motion trajectory based on visual information according to claim 1, characterized in that: In step S5, the specific steps include: S51, the node set at time t Perform linear transformation to obtain a new set of node feature vectors The calculation process is as follows: Where W l is the weight matrix of linear transformation; S52. Use node feature vector set and the set of directed edge weights Perform feature aggregation, and the calculation process is as follows: in, Represents the temporary feature vector of the node after feature aggregation, Represents the set of neighbor nodes of node i; S53, use the multi-layer perceptron network to perform nonlinear mapping to update node features, the updating process is as follows: in, represents the updated feature vector after MLP processing, ∈ is a learnable parameter, W c For MLP c The weight parameter of S54, use the activation function ELU to further update the node features to obtain the final node features: The definition of the activation function ELU is as follows: Where χ represents the input vector of ELU, and exp(χ) is the exponential function of χ.
6. The method for predicting pedestrian motion trajectory based on visual information according to claim 5, characterized in that: In step S6, the specific steps include: S61, reshape the final node features outputted in step S54, introduce new feature dimensions, and aggregate them into a new feature vector in F represents the feature length, and D is the new dimension introduced in the reshaping process, which gradually increases with the number of convolutional layers; S62, use multiple two-dimensional convolution operations on the reshaped feature vector I to extract features in the time dimension; the calculation process of each convolution layer is as follows: in is the output of the convolutional layer, K m,n represents the kernel weight, P represents the padding size, (p, q), (m, n) are the calculation indices; S63. After the two-dimensional convolution operation, the batch normalization operation is used to process the output of each convolution layer to standardize the features of each channel. The calculation formula of the batch normalization operation is as follows: Where X represents the input vector of batch normalization, μ represents the mean value of the feature matrix, σ represents the standard deviation of the current feature matrix, δ and β are the scaling and shift parameters learned during training respectively; S64, using the ELU activation function defined in step S54 to perform nonlinear activation on the feature vector after the batch normalization operation; S65. In the process of time series feature extraction, the channel attention mechanism is introduced to perform weight distribution in the time dimension. The specific weight distribution process is as follows: in Represent the input and output features of the channel attention mechanism, w c is the channel attention feature, f cnn (x) is the channel feature transformation achieved through convolution operation; S66. After completing the node feature extraction, the fully connected layer is finally introduced to map the high-dimensional feature vector to the low-dimensional coordinate space of the target trajectory and output the predicted future motion trajectory.
7. A system according to any one of claims 1 to 6, wherein: The system includes the following modules: UAV data collection module (110): collects data through the UAV and performs standardized processing on the data; Coordinate projection transformation module (120): restores pixel coordinates in the video to actual ground truth coordinates through a multi-grid projection transformation algorithm; Coordinate sampling and trajectory aggregation module (130): sampling the ground truth coordinates of pedestrians after projection transformation, and aggregating the coordinate data into trajectory information of multiple pedestrians; Global interaction directed topology map construction module (140): combines spatial visual information with temporal information to construct a global interaction directed topology map, thereby improving the model's ability to capture spatial interaction features in complex environments; Spatial feature extraction module (150): A spatial feature extraction module is constructed using a graph isomorphic network (GIN), which adaptively learns directed edge weights based on visual information and aggregates features of neighboring nodes to update current node features, capturing the spatial interaction relationship between pedestrians and their neighbors. Temporal feature extraction module (160): A temporal feature extraction module is constructed using a two-dimensional convolutional neural network (CNN). The temporal series features of pedestrian motion are efficiently extracted through a multi-layer convolution structure, and the future motion trajectory of the pedestrian is finally predicted.
8. An electronic device, characterized in that: The electronic device comprises: a processor (210); and a memory (220) for storing one or more programs; When the one or more programs are executed by the processor (210), the processor executes the pedestrian motion trajectory prediction method based on visual information as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by the processor (210), the method for predicting pedestrian motion trajectories based on visual information as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Pedestrian trajectory time sequence prediction method
CN115034459A