Instantaneous trajectory prediction method based on bidirectional collaborative diffusion model and application
Through the two-way collaborative diffusion model, the trajectory prediction accuracy and stability are solved under data scarcity and noise interference, and high-precision trajectory prediction in emergency or complex scenarios are achieved.
Patent Information
- Application Number
- CN202510026503.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-09
AI Technical Summary
The existing multi-peeder trajectory prediction methods are difficult to maintain prediction accuracy and stability in the case of scarce data, complex interactions and noise interference. Especially in emergency or complex scenarios, the model is difficult to capture sufficient timing information and complex social interaction and environmental factors.
The instantaneous trajectory prediction method based on the two-way collaborative diffusion model is adopted to generate unobserved historical trajectories and future trajectories in both directions, and to use mutual guidance mechanisms to enable the diffusion model to generate more accurate trajectory results in each step.
In the case of data scarcity and noise interference, the accuracy and stability of trajectory prediction are significantly improved, and high-quality trajectories can be generated in a very short time, suitable for autonomous driving, intelligent transportation and other fields.
Smart Images

Figure CN119961597A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pedestrian trajectory prediction, and in particular to an instantaneous trajectory prediction method based on a bidirectional cooperative diffusion model. Background Art
[0002] At present, domestic and foreign scholars have conducted a lot of research in the field of multi-pedestrian trajectory prediction and have made significant progress, mainly focusing on the following aspects: 1. Pedestrian trajectory prediction based on deep learning: Deep learning methods have performed well in trajectory prediction, especially in dealing with complex social interactions and diverse behavior patterns. For example, the Social LSTM model effectively predicts group behavior by capturing social interactions between pedestrians. Another improved deep learning model, Social GAN, introduces a generative adversarial network (GAN) to generate more natural and diverse pedestrian trajectories through adversarial training, thereby improving the robustness and accuracy of the prediction. 2. Pedestrian trajectory prediction based on graph neural networks: Graph neural networks (GNNs) have been widely used in pedestrian trajectory prediction due to their advantages in processing graph-structured data and capturing complex relationships. For example, the Social-STGCNN model can capture the dependencies in both spatial and temporal dimensions by applying the spatial-temporal graph convolutional network (ST-GCN) to pedestrian groups, significantly improving the prediction accuracy and the interpretability of the model. 3. Pedestrian trajectory prediction based on reinforcement learning: Reinforcement learning (RL) methods optimize trajectory prediction strategies through interactive learning with the environment, and perform well in dealing with uncertain and dynamically changing scenarios. For example, the RL-LSTM model combines deep reinforcement learning and LSTM networks to effectively deal with randomness and environmental changes in multi-pedestrian trajectory prediction and enhance the adaptability of the model. 4. Pedestrian trajectory prediction based on attention mechanism: The attention mechanism shows strong advantages in dealing with complexity and uncertainty due to its ability to automatically focus on important information. For example, the SoPhie model significantly improves the ability to capture pedestrian behavior and environmental characteristics by introducing spatial and social attention mechanisms, thereby achieving more accurate trajectory prediction. These studies have significantly improved the accuracy, robustness and adaptability of multi-pedestrian trajectory prediction methods, gradually meeting the needs of practical applications such as autonomous driving, intelligent transportation and robot navigation.
[0003] However, the research on multi-pedestrian trajectory prediction still faces many challenges. First, traditional trajectory prediction methods usually rely on long-term historical trajectory data to infer future paths. However, in practical applications, only very little observation data may be obtained, which makes it difficult for the model to capture sufficient temporal information, resulting in inaccurate prediction results. For example, in emergency or occlusion scenarios, pedestrians' behaviors suddenly change, and the model is difficult to cope with these emergencies when there is insufficient data. Secondly, the social interactions between pedestrians (such as avoidance and gathering) and the environment (such as obstacles and road conditions) have a significant impact on pedestrian trajectories. Capturing these complex interactions and environmental factors is one of the difficulties in trajectory prediction; secondly, the noise and errors in the data will have a great impact on the prediction results; finally, trajectory prediction needs to grasp long-term time dependencies. However, in many real-world situations, the model lacks sufficient time for observation. For example, pedestrians suddenly appear from blind spots, resulting in inaccurate predictions and even safety risks. Therefore, many difficulties in multi-pedestrian trajectory prediction focus on how to maintain the accuracy and stability of predictions in the case of scarce data, complex interactions, and noise interference. The solution to these problems is of great significance for applications in fields such as autonomous driving and intelligent transportation. Summary of the invention
[0004] The purpose of the present invention is to provide a method for instantaneous trajectory prediction based on a bidirectional collaborative diffusion model to solve the problems raised in the above background technology. By bidirectionally generating unobserved historical trajectories and future trajectories, the technology of instantaneous trajectory prediction is realized. The core is to use a mutual guidance mechanism so that in each step, the generated unobserved historical trajectories and limited observed trajectories can guide a diffusion model to generate future trajectories, and at the same time, the generated future trajectories and observed trajectories guide another diffusion model to predict unobserved historical trajectories. Through this iterative mutual guidance generation process, future and unobserved historical trajectories are continuously refined, and accurate trajectory prediction is finally achieved.
[0005] To achieve the above object, the present invention provides the following technical solution: an instantaneous trajectory prediction method based on a bidirectional cooperative diffusion model, comprising the following steps;
[0006] This method uses the forward diffusion process to gradually perturb the data into random noise, and then gradually restores the generation model of the original data from the noise through the reverse denoising process.
[0007] S1: Process the input data to form a feature vector F obs ;
[0008] S2: For F obs Encode the observed trajectory features v obs :
[0009] S3: Generate X using the diffusion model unobs and X fut , where X unobs represents the unobserved historical trajectory, where X fut Indicates future trajectory;
[0010] S4: Diffusion models use each other’s output as input for feedback adjustment and guidance;
[0011] S5: After multiple iterations and mutual guidance, the diffusion model finally outputs the trajectory result:
[0012] Specific S1: Input and initial processing. The input data includes a very small amount of observed trajectory data X obs = {x1, x2}, these data only contain trajectory points of two time frames. At the same time, define the future trajectory and unobserved historical trajectories The latter is what the model needs to generate;
[0013] Specific S2: encoding stage. The observed trajectory data X is converted into obs Encoded as observed trajectory features v obs , while capturing the scene context information feature vector e. The encoding output of this stage provides basic information for the subsequent generation process;
[0014] Specifically, S3: The generation process of the bidirectional collaborative diffusion model. The core idea of the diffusion model is to use the forward noise addition process and the reverse denoising process of the diffusion model to generate unobserved historical trajectories X through step-by-step iteration. unobs and Future Trajectory X fut Diffusion models are a type of generative model that can recover high-quality data points from noise in a step-by-step manner.
[0015] Specifically, the forward diffusion process is defined by a Markov chain, which gradually transforms the real trajectory data into standard Gaussian noise: Among them, X t represents the trajectory state of the tth step in the diffusion process, α t Controls the intensity of the noise added, defined as where β i is a predefined noise adjustment parameter. I is the unit matrix, which represents the independent noise in each dimension. Through this process, the true trajectory X0 is finally perturbed into a standard Gaussian distribution X M ~N(0,I), where M is the maximum number of diffusion steps.
[0016] The reverse process is to gradually recover the true trajectory from the noise through the conditional probability distribution: where μ θ (X t ,t) is the mean function learned by the denoising neural network, estimating the state after denoising at each step, ∑ θ (X t ,t) is the learned variance function, which controls the accuracy of each step of recovery. The purpose of the inverse process is to gradually reconstruct the trajectory to make it close to the true trajectory X0. The mean and variance in the denoising process are dynamically adjusted by the conditional input (observed trajectory and guidance information).
[0017] In the trajectory prediction task, the diffusion model can generate high-quality trajectory points in noisy scenes by taking advantage of its gradual generation characteristics, while effectively capturing the complementary information between historical and future trajectories. This model uses two coupled diffusion models to generate unobserved historical trajectories X and X respectively. unobs and Future Trajectory X fut At each generation step, the observed trajectory data X obs and the generated unobserved historical trajectory X unobs is input into the back-diffusion model to generate the guidance matrix g unobs The forward diffusion model uses g unobs and X obs As conditional information, the next future trajectory is generated. On the other hand, the observed trajectory data X obs and the generated future trajectory X fut Is input into the forward diffusion model to generate the guidance matrix g fut The reverse diffusion model uses g fut and X obs As conditional information, generate unobserved historical trajectories for the next step;
[0018] Specific S4: Mutual guidance and noise control. In order to optimize the generated trajectory in each step, the model introduces a mutual guidance mechanism, in which the two diffusion models use each other's generated trajectory for guidance at each step to refine the generated results. In the initial stage of generation, due to the existence of noise, the model designs a gating mechanism to automatically learn and adjust the weights between different guidance information to ensure the stability and accuracy of the generation process;
[0019] Specific S5: After multiple iterations and refinements, the final result includes a more accurate future trajectory X fut and the unobserved historical trajectory X unobs These trajectory outputs can be used in trajectory prediction tasks, especially in data-scarce situations, to provide reliable prediction results.
[0020] Preferably, the two diffusion models include a forward diffusion model f fut and the reverse diffusion model funobs , track X fut represents the future trajectory, X unobs represents the unobserved historical trajectory; where the reverse diffusion model f unobs Generate future trajectory X fut ; Forward diffusion model f fut Generate unobserved historical trajectory X unobs .
[0021] Preferably, inputting data in S1 includes: using X obs ={x1,x2} represents the observed trajectory data, which contains trajectory points of two time frames. x1 and x2 are the position coordinates of the pedestrian at two different time points. These two points represent the trajectory information collected by the model in a very short time. Represents the future trajectory, which is the future trajectory that needs to be predicted, including the time point from t = 3 to t = T fut +2 a series of positions, each x i (where i ≥ 3) represents the predicted position of the pedestrian at a certain point in the future, T fut Indicates the time span of the future trajectory; Represents the unobserved historical trajectory, which is the trajectory point that needs to be generated, representing the pedestrian at the initial time point t = 1-T unobs A series of positions from t to t = 0, indicating that these historical trajectory data have not been observed in practice, T unobs represents the time span of the unobserved historical trajectories.
[0022] Preferably, the data processing in S1 includes: converting the input observed trajectory data X obs 、Future Trajectory X fut and the unobserved historical trajectory X unobs Formatting to ensure that the data is input into the model in a uniform format, including representing the trajectory points as vectors or matrices; extracting the observed trajectory data X obs The pedestrian's direction, speed, acceleration and environmental interaction characteristics are integrated into a feature vector F obs , which is of the form: Among them, the eigenvector F obs is a matrix of shape N×d; where N is the number of time steps and d is the data dimension of each time step, including the velocity direction and characteristic information related to the environment.
[0023] Preferably, in S2, F obsThe encoding includes: transforming the feature vector F obs It is processed through the LSTM (Long Short-Term Memory) network and input into the encoder module of the model to generate a higher-level feature representation v obs , and combines social interaction and scene context information; using g soc Represents social interaction characteristics; f scene Represents scene context information; e represents the scene context information feature vector; the feature representation v obs , social interaction characteristics soc Combined with the scene context feature vector e, a complete feature representation v is generated total :v total =v obs +g soc +e.
[0024] Preferably, X is generated in S3 respectively unobs and X fut The following steps are included: using the encoded observed trajectory features v obs and the generated unobserved historical trajectory X unobs As guidance information, guide the generation of future trajectory X fut , this guidance process uses an LSTM network to analyze the unobserved historical trajectory X unobs and Future Trajectory X fut Encode and generate the forward guidance matrix g fut ,in:
[0025] g fut =f Bi-LSTM (X unobs )
[0026] Specifically, f Bi-LSTM represents a function based on a bidirectional long short-term memory (Bi-LSTM) network. The Bi-LSTM network can capture both forward and backward information of sequence data, which helps to better understand the dependencies in time series. Here, the Bi-LSTM network predicts the future trajectory X. fut and observed trajectory features v obs Encode and generate the forward guidance matrix g fut .
[0027] Then, the weight γ is learned through the gating mechanism Gate fut , balance v obs With g fut The contribution between them generates the final guidance information in, represents the control parameter of the guidance information, fut represents the forward guidance information parameter, which is used to adjust the output of each step of the generated trajectory, where:
[0028]
[0029] Specifically, the gating mechanism dynamically adjusts the weights between the guidance information to ensure that the contributions of different information sources are properly balanced during the generation process. Specifically, the gating mechanism learns a weight parameter γ through a linear transformation and an activation function sigmoid. fut , which determines the current observed trajectory feature v obs And the generated forward guidance matrix g fut The relative importance of v in generating the next trajectory. obs and g fut When combined in a suitable ratio, the information required in the trajectory generation process can be captured more accurately, noise interference can be reduced, and the accuracy of the generated trajectory can be improved.
[0030] Finally, through guidance information The model uses a denoising mechanism to remove noise at each generation step and generate the next future trajectory. and historical trajectory The superscript t-1 represents the trajectory generated in the previous step of time step t. Specifically, the denoising mechanism gradually reduces the influence of noise during the generation process, making the generated trajectory gradually close to the true trajectory. The denoising process in each step of generation removes a certain amount of noise on the trajectory of the current time step, thereby generating a more accurate trajectory for the next step. Specifically, the denoising process is performed using the following formula:
[0031]
[0032] Where X (t) represents all trajectories at time step t, X (t-1) represents the trajectory at time step t-1, α t is a parameter to control the noise, and ∈ represents the noise sampled from a Gaussian distribution with a mean of 0 and a covariance of the identity matrix I. By gradually removing the noise, the generated trajectory will become more and more accurate, and eventually a more accurate future trajectory X will be obtained. fut ;
[0033] Then, use the generated future trajectory X fut and the encoded observed trajectory features v obs As guidance information, reverse guidance generates historical trajectories, and the future trajectory X is generated through the LSTM network. fut and observed trajectory features v obs Encode and generate the reverse guidance matrix:
[0034] gunobs =f Bi-LSTM (X fut ,v obs ),
[0035] Specifically, f Bi-LSTM represents a function based on a bidirectional long short-term memory (Bi-LSTM) network. The Bi-LSTM network can capture both forward and reverse information of sequence data, which helps to better understand the dependencies in time series. Here, the Bi-LSTM network predicts the future trajectory X. fut and observed trajectory features v obs Encode and generate the reverse guidance matrix.
[0036] g unobs is a future trajectory X fut and observed trajectory features v obs A vector of weights γ is preferably learned using a gating mechanism. unobs , which is used to balance the influence of future trajectory and observed trajectory characteristics:
[0037]
[0038] The weights γ learned by the gating mechanism unobs Controlled g unobs and v obs The contribution of X to ensure the generated unobserved historical trajectory unobs Can reflect trajectory dynamics more accurately.
[0039] Preferably, in S4, the feedback adjustment guidance includes: the forward diffusion model uses the reverse guidance matrix g generated by the reverse diffusion model unobs To adjust the generation of future trajectories, the reverse diffusion model uses the forward guidance matrix g generated by the forward diffusion model fut To guide the generation of unobserved historical trajectories; given the forward guidance information generated by the forward diffusion model This information is used as the reverse diffusion model f unobs The input guides the generation of historical trajectories in:
[0040]
[0041] At the same time, the reverse guidance matrix generated by the reverse diffusion model As the forward diffusion model f fut The input guides the generation of future trajectories:
[0042]
[0043] This mutual feedback process ensures the consistency of trajectory generation and gradually reduces the error by using each other's generation results. represents the unobserved historical trajectory generated at time step t+1; represents the guidance information matrix generated by the forward diffusion model at time step t, which is used to guide the backward diffusion model to generate the next historical trajectory; represents the unobserved historical trajectory at time step t; represents the guidance information matrix generated by the backward diffusion model at time step t, which is used to guide the forward diffusion model to generate the next future trajectory; represents the future trajectory generated at time step t+1.
[0044] Preferably, the influence of noise is balanced by learning the weights between different guidance information. The specific method is as follows: For each time step t, the gating mechanism calculates the gating weight γ t To adjust the ratio of generated trajectories to guidance information:
[0045] γ t =σ(W g ·[v obs ,g unobs ,g fut ])
[0046] Among them, W g is a learnable weight matrix, σ is the activation function sigmoid, which is used to control the value range of the weight between [0,1]; the gating mechanism calculates the weight γ t Act on the trajectory generated at each step to balance the contribution between observed data and generated data:
[0047] X (t+1) =γ t X (t) +(1-γ t )v obs
[0048] This process dynamically adjusts the degree of fusion between the generated trajectory and the observed trajectory to ensure that the noise is gradually weakened, making the final generated trajectory smoother and more accurate.
[0049] Preferably, the trajectory X is output again in S5 fut and track X unobs , including the following steps: iteratively generate future trajectory X fut Gradually converges to the observed trajectory X obs and historical trajectory X unobs Consistent prediction path, T represents the maximum number of iteration steps, the formula is expressed as:
[0050]
[0051] in, is the reverse guidance information matrix generated after time step t, is the final future trajectory after all iterations; the back diffusion model gradually generates the unobserved historical trajectory X by learning from the future trajectory and the observed trajectory unobs , the formula is:
[0052]
[0053] in, represents the final generated historical trajectory; the final generated future trajectory and unobserved historical trajectories It is directly used as the model output for trajectory prediction tasks. The formula is:
[0054]
[0055] Preferably, an autonomous driving vehicle comprises any of the above methods.
[0056] Preferably, an intelligent navigation robot is provided, wherein the navigation robot has any of the above methods.
[0057] Preferably, an intelligent transportation system is provided, wherein the intelligent transportation system comprises any one of the above methods.
[0058] The present invention has at least the following beneficial effects:
[0059] 1. Ability to handle small amounts of observation data: Existing trajectory prediction methods usually rely on a long observation time to accurately predict future trajectories. However, in practical applications, especially in emergency or complex scenarios, there may be only a very short observation time. The model effectively utilizes limited observation data by generating unobserved historical trajectories and future trajectories in a bidirectional manner, thereby significantly improving the prediction accuracy in a very short time;
[0060] 2. Bidirectional generation and mutual guidance mechanism: Traditional methods usually rely only on one-way time series information, while the model introduces a bidirectional generation mechanism, which generates historical and future trajectories simultaneously and optimizes both in each round of generation through a mutual guidance mechanism. This mechanism not only enhances the accuracy of trajectory prediction, but also improves the robustness of the model in the face of noise and uncertainty;
[0061] 3. Gating mechanism ensures generation stability: Since the diffusion model introduces more noise in the initial step, the traditional generation method may affect the quality of the final result. The model automatically learns and adjusts the weights in the generation process through the gating mechanism to ensure that the generated trajectory remains stable even if it is affected by noise in the early stage, thereby improving the overall prediction accuracy;
[0062] 4. Wider application scenarios: Most existing methods rely on long-term data accumulation. This method is particularly suitable for instantaneous trajectory prediction tasks, and has important application value in the fields of autonomous driving and drone navigation. It can still provide high-precision prediction results when data is scarce. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 The overall architecture of the model;
[0064] Figure 2 Bidirectional cooperative diffusion decoding process;
[0065] Figure 3 Internal structure of the bidirectional cooperative diffusion unit. DETAILED DESCRIPTION
[0066] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0067] Example
[0068] See also Figure 1-Figure 3 ,An instantaneous trajectory prediction method based on a bidirectional cooperative diffusion model includes step 1, input and initial processing.
[0069] This step is divided into two sub-steps:
[0070] Step 1-1: Input data definition: Define the input data for trajectory prediction, which includes the following three parts:
[0071] 1. Observed trajectory data X obs = {x1, x2}, which is the pedestrian trajectory data that the model has observed, and only contains trajectory points of two time frames. x1 and x2 are the position coordinates of the pedestrian at two different time points. These two points represent the trajectory information collected by the model in a very short time. Usually, these position coordinates are two-dimensional, indicating the position of the pedestrian in a plane space.
[0072] 2. Future trajectory This is the future trajectory that the model needs to predict, from time point t=3 to t=T fut +2. Each x i (where i ≥ 3) represents the predicted position of the pedestrian at a certain point in the future, T fut represents the time span of future trajectories; the goal of the model is to use the observed trajectories X obs To predict these future points.
[0073] 3. Unobserved historical trajectories It is the trajectory point that the model needs to generate, representing the pedestrian at the initial time point t = 1-T unobs to a series of positions from t = 0. This means that these historical trajectory data have not been observed in reality, but by generating these points, the model can better capture the complete information of the time dimension and improve the accuracy of future trajectory prediction.
[0074] Step 1-2, initial processing, after defining the input data, the model performs preliminary processing on this data for subsequent trajectory generation and prediction steps:
[0075] 1. Data formatting: Input observed trajectory data X obs 、Future Trajectory X fut and the unobserved historical trajectory X unobs Format them to ensure that they are input into the model in a uniform format. This includes representing the trajectory points in vector or matrix form so that the subsequent encoding and diffusion processes can effectively process these data. The specific steps are as follows:
[0076] 1) Define the data structure: (1) Each trajectory point x i Represented as a two-dimensional vector in and represent the horizontal and vertical coordinates of the trajectory points respectively. (2) For each set of trajectory data (such as X obs ), arrange all trajectory points in time order to form a matrix. Assuming there are N trajectory points, then X obs The format is an N×2 matrix;
[0077] 2) Standardization: (1) Normalize the coordinates of the trajectory points and scale all coordinates to a uniform range (such as [0, 1]) so that trajectory data of different scales can be processed in the same model. (2) If data of different time frames are used, ensure the consistency of the time frame interval and add the timestamp to the data matrix so that each trajectory point can be marked with its time information, that is, each trajectory point is represented as where t i Indicates the time frame;
[0078] 3) Data storage and input preparation: Save the processed trajectory data in standard Tensor format for subsequent model calls. Ensure that all trajectory data formats are consistent to facilitate input into the model's encoder or other processing modules.
[0079] 2. Preliminary data analysis: By analyzing the observed trajectory data X obs The preliminary analysis extracts basic features such as the direction and speed of pedestrian movement. These features will be used in the subsequent encoder module to help the model better understand the dynamic characteristics of trajectory data. The specific steps are as follows:
[0080] 1) Speed calculation: For the observed trajectory data X obs Each pair of adjacent trajectory points (x i-1 ,x i ), calculate the pedestrian's velocity vector v i : where v i is a two-dimensional vector, indicating that i-1 to i Speed between time frames;
[0081] 2) Direction analysis: Based on the calculated velocity vector v i , extract the pedestrian's movement direction θ i , which is calculated as: in and are the components of the velocity vector on the x-axis and y-axis, θ i is the direction angle of pedestrian movement;
[0082] 3) Acceleration analysis: For a more detailed dynamic analysis, the acceleration of the pedestrian can be further calculated, that is, the velocity vector v i Take the derivative to get the acceleration vector a i : This analysis can help identify if pedestrians make sudden changes in speed or turn maneuvers.
[0083] 3. Feature initialization: Initialize the features of the observed trajectory data. This step usually includes calculating the average speed, direction, and interaction with the environment of the pedestrian. The initialized features will be input into the encoder of the model to further capture dynamic information in time and space. The specific steps are as follows:
[0084] 1) Average speed and direction initialization: (1) Calculate the average of all speed vectors in the observed trajectory and the average of the direction angles These averages will be used as the basis for feature initialization to provide the model with overall motion trend information:
[0085] 2) Initialization of interaction with the environment: If there is environmental information (such as obstacle location, road structure, etc.), combine this information with the observed trajectory to calculate the motion characteristics of the pedestrian relative to the environment, such as the distance to the nearest obstacle, the deviation angle between the direction of travel and the road direction, etc. These characteristics can help the model understand how pedestrians interact with the environment;
[0086] 3) Feature vector construction: (1) Integrate the above-calculated speed, direction, acceleration and environmental interaction characteristics into a feature vector F obs , which is of the form: This feature vector is passed as input to the encoder module of the model, capturing dynamic information in time and space.
[0087] Through the above steps, the input data is not only formatted and standardized, but also deeply analyzed and key dynamic features are extracted, thus providing rich initial information and a good feature foundation for the subsequent trajectory generation and prediction process.
[0088] Step 2: Data encoding.
[0089] In step 1, we have analyzed different trajectory data X obs , X fut and X unobs The formatting and standardization were performed, and the speed, direction, acceleration and environmental features were extracted to form a complete feature vector F obs , which is a matrix of shape N×d, where N represents the number of time steps and d represents the data dimension of each time step, including the velocity direction These feature vectors are processed by the LSTM network and input into the encoder module of the model to generate higher-level feature representation v obs , and combined with the scene context information feature vector e, it provides a basis for the subsequent diffusion process.
[0090] Step 2-1: Encoder for feature vector F obs The encoding:
[0091] 1. Input features: The feature vector input to the encoder is This includes the average speed of pedestrians Average direction of movement As well as the interaction characteristics with the environment. This feature vector already contains the key information that has been analyzed and initialized, representing the global dynamic characteristics of pedestrian motion.
[0092] 2. Encoder structure: The encoder uses LSTM (Long Short-Term Memory Network) to process the input time series data and extract the time-dependent features in the trajectory. in, represents the feature vector input at time t, h t It is the hidden state of LSTM, which contains the high-dimensional feature representation of the current trajectory.
[0093] 3. Output features: the final output v obs is the hidden state of the LSTM encoder after time series processing: v obs =h T .v obs As a feature representation, it not only contains the motion information of the trajectory, but also includes the historical dependencies in the time series.
[0094] Step 2-2: Combination of social interaction information and scene context information:
[0095] 1. Social interaction information soc : Use graph neural network (GNN) to model the social interaction information between multiple pedestrians. Each pedestrian node contains its trajectory feature vector F obs , and transmit and update information based on the trajectory information of neighboring nodes (i.e. other pedestrians): g soc =GNN(F obs , neighbor information). The graph neural network aggregates the information of all pedestrian nodes and outputs the social interaction feature g soc , represents the mutual influence between pedestrians and other pedestrians;
[0096] 2. Scene context information f scene :Use convolutional neural network (CNN) to extract environmental information in the scene, such as obstacles, road structure, etc. scene =CNN(scene image). The scene features are combined with the trajectory features to obtain the final feature vector e containing the scene context;
[0097] 3. Feature fusion: The trajectory feature v obs , social interaction characteristics soc Combined with the scene context feature e, a complete feature representation v is generated total :v total =v obs +g soc +e. This feature representation contains the dynamic information of trajectory data, social interactions, and environmental influences, providing input for subsequent diffusion models.
[0098] Step 3: The generation process of the two-way collaborative diffusion model.
[0099] The bidirectional cooperative diffusion model is one of the core technologies of the present invention. It mainly generates unobserved historical trajectories X through two coupled diffusion models. unobs and Future Trajectory X fut The overall structure is as follows Figure 2 The decoding process of the bidirectional cooperative diffusion model is shown in the figure. The specific steps are as follows:
[0100] Step 3-1: Forward diffusion model generates future trajectories
[0101] 1. Input data: Use the encoded observed trajectory features v obs and randomly initialized future trajectories X fut , the initial point is initialized by a random noise matrix ∈, whose dimension is (T fut , 2), where T fut is the number of time steps of the future trajectory;
[0102] 2. Definition of diffusion process: The diffusion process is a positive Markov chain generation process, which perturbs the initial data by adding noise and gradually recovers the true trajectory. At each time step t, given the current trajectory representation X (t) , the model generates the next step trajectory representation X (t+1) , the formula is as follows:
[0103]
[0104] Among them, α t is a gradually decreasing noise scaling factor that controls the size of the noise perturbation.
[0105] 3. Guidance information: using the encoded observed trajectory features v obs and the generated unobserved historical trajectory X unobs As guidance information, guide the generation of future trajectory X fut This guidance process uses an LSTM network to analyze the unobserved historical trajectory X unobs and Future Trajectory X fut Encode and generate the forward guidance matrix g fut ,in:
[0106] g fut =f Bi-LSTM (X unobs )
[0107] Specifically, f Bi-LSTMrepresents a function based on a bidirectional long short-term memory (Bi-LSTM) network. The Bi-LSTM network can capture both forward and backward information of sequence data, which helps to better understand the dependencies in time series. Here, the Bi-LSTM network predicts the future trajectory X. fut and observed trajectory features v obs Encode and generate the forward guidance matrix g fut .
[0108] Then, the weight γ is learned through the gating mechanism Gate fut , balance v obs With g fut The contribution between them generates the final guidance information in, represents the control parameter of the guidance information, and fut represents the forward guidance information parameter, which is used to adjust the output of each step of the generated trajectory.
[0109]
[0110] Specifically, the gating mechanism dynamically adjusts the weights between the guidance information to ensure that the contributions of different information sources are properly balanced during the generation process. Specifically, the gating mechanism learns a weight parameter γ through a linear transformation and an activation function sigmoid. fut , which determines the current observed trajectory feature v obs And the generated forward guidance matrix g fut The relative importance of v in generating the next trajectory. obs and g fut When combined in a suitable ratio, the information required in the trajectory generation process can be captured more accurately, noise interference can be reduced, and the accuracy of the generated trajectory can be improved.
[0111] 4. Denoising process: Finally, through the guidance information The model uses a denoising mechanism to remove noise at each generation step and generate the next future trajectory. and historical trajectory The superscript t-1 represents the trajectory generated in the previous step of time step t. Specifically, the denoising mechanism gradually reduces the influence of noise during the generation process, making the generated trajectory gradually close to the true trajectory. The denoising process in each generation step removes a certain amount of noise on the trajectory of the current time step, thereby generating a more accurate trajectory for the next step.
[0112] Step 3-2: Generate historical trajectories using the reverse diffusion model.
[0113] 1. Input data: The initial input is the future trajectory Xfut and the encoded observed trajectory features v obs . Future Trajectory X fut Also initialized from a random noise matrix ∈ with dimension (T unobs ,2), where T unobs is the number of time steps of the unobserved historical trajectory;
[0114] 2. Define the back diffusion process: Similar to forward diffusion, back diffusion generates unobserved historical trajectories by denoising. Given the current trajectory representation X (t) , the next historical trajectory is generated by the following formula:
[0115]
[0116] This process gradually generates historical trajectory points X unobs ;
[0117] 3. Guidance information: Use the generated future trajectory X fut and the encoded observed trajectory features v obs As guidance information, reverse guidance generates historical trajectories. The future trajectory and observation data are encoded through the LSTM network to generate the reverse guidance matrix:
[0118] g unobs =f Bi-LSTM (X fut ,v obs ),
[0119] g unobs is a future trajectory X fut and observed trajectory features v obs The vector of can provide guidance information for the subsequent generation of unobserved historical trajectories. In order to dynamically adjust the contribution of this information during the generation process, a gating mechanism is used to learn the weight γ unobs , which is used to balance the influence of future trajectory and observed trajectory characteristics:
[0120]
[0121] The weights γ learned by the gating mechanism unobs Controlled g unobs and v obs The contribution of X to ensure the generated unobserved historical trajectory unobs Can reflect trajectory dynamics more accurately.
[0122] 4. Denoising process: Also after each generation step, a gating mechanism is used to balance the weight between the generated trajectory and the guidance information:
[0123] X (t-1) =Gating(X(t-1) , g unobs )
[0124] Through the above two-way collaborative diffusion model, the model uses the observed trajectory data to generate future and historical trajectories, enabling the model to accurately predict pedestrian trajectories when data is scarce.
[0125] Step 4: Mutual guidance and noise control.
[0126] In the bidirectional cooperative diffusion model of the present invention, mutual guidance and noise control are key steps to ensure the stability and accuracy of trajectory generation. Through the mutual guidance mechanism, the forward and reverse diffusion models use each other's generation results for refinement. At the same time, noise control effectively suppresses the uncertainty caused by noise in the generation process through the gating mechanism, ensuring the stability of the model. The whole mutual guidance and noise control process is as follows Figure 3 shown.
[0127] Step 4-1: Mutual guidance mechanism.
[0128] 1. Mutual guidance core idea: In each generation step, the forward diffusion model f fut and the reverse diffusion model f unobs The generated guidance matrix uses each other's output as input for feedback adjustment. Specifically, the forward diffusion model uses the guidance matrix g generated by the reverse diffusion model unobs To adjust the generation of future trajectories. The reverse diffusion model uses the g generated by the forward diffusion model fut to guide the generation of unobserved historical trajectories;
[0129] 2. Mathematical modeling:
[0130] 1) Given the forward guidance information generated by the forward diffusion model This information is used as the reverse diffusion model f unobs The input guides the generation of historical trajectories
[0131]
[0132] 2) At the same time, the reverse guidance matrix generated by the reverse diffusion model As the forward diffusion model f fut The input guides the generation of future trajectories:
[0133]
[0134] This mutual feedback process ensures the coherence of trajectory generation and gradually reduces the error by leveraging each other's generation results.
[0135] Step 4-2, design of gating mechanism and noise control. In the initial stage of the generation process, due to the introduction of noise, the generated trajectory may deviate from the target trajectory. In order to overcome this problem, the present invention introduces a gating mechanism to balance the influence of noise by learning the weights between different guidance information. The specific method is as follows:
[0136] 1. Mathematical modeling: For each time step t, the gating mechanism calculates the gating weight γ t To adjust the ratio of generated trajectories to guidance information:
[0137] γ t =σ(W g ·[v obs ,g unobs ,g fut ])
[0138] Among them, W g is a learnable weight matrix, σ is the activation function sigmoid, which is used to control the value range of the weight between [0,1];
[0139] 2. Denoising process: The gating mechanism calculates the weight γ t Act on the trajectory generated at each step to balance the contribution between observed data and generated data:
[0140] X (t+1) =γ t X (t) +(1-γ t )v obs
[0141] This process dynamically adjusts the degree of fusion between the generated trajectory and the observed trajectory to ensure that the noise is gradually weakened, making the final generated trajectory smoother and more accurate.
[0142] Step 4-3: Iterative optimization and trajectory refinement.
[0143] 1. Iterative optimization: Through the mutual guidance mechanism of the forward and reverse diffusion models, the model is continuously iteratively optimized. In each iteration, the forward diffusion model uses the reverse guidance information g generated by the historical trajectory unobs To optimize future trajectory X fut The reverse diffusion model generates forward guidance information g through future trajectories. fut To correct the unobserved historical trajectory X unobs
[0144] 2. Trajectory refinement: After multiple iterations, the trajectory generated by the model gradually removes the initial noise interference and obtains a more accurate future trajectory X fut and the unobserved historical trajectory X unobs .
[0145] Through mutual guidance and noise control, the bidirectional cooperative diffusion model can effectively overcome the noise interference problem in the generation process and gradually refine the generated trajectory through mutual guidance at each step. This mechanism significantly improves the accuracy and stability of trajectory prediction, enabling the model to generate high-quality trajectories even with less observation data.
[0146] Step 5. Final output generated by iteration.
[0147] In the bidirectional collaborative diffusion model of the present invention, after multiple iterations and mutual guidance, the final output generated includes two parts: a more accurate future trajectory X fut and the unobserved historical trajectory X unobs These trajectory outputs can be directly used in trajectory prediction tasks, especially in scenarios where data is scarce, to provide reliable prediction results. The specific process is as follows:
[0148] Step 5-1: Iteratively generate future trajectory X fut .
[0149] 1. Convergence of future trajectories: After multiple rounds of iterations, the future trajectory X of the forward diffusion model fut Gradually remove uncertainty from the initial noise and gradually converge to the observed trajectory X obs and historical trajectory X unobs The consistent prediction path is expressed as:
[0150]
[0151] Where T represents the maximum number of iteration steps, is the final future trajectory after all iterations.
[0152] 2. Dynamic adjustment of future trajectories: The model uses historical trajectory information to continuously correct future trajectories in each iteration until it converges to a set of trajectory points with the minimum error. The correction and optimization of future trajectories is based on minimizing the error loss function. The loss function is:
[0153]
[0154] is the future trajectory point generated by the model at time step t, indicating the predicted position coordinates; is the true trajectory point at time step t. Through back propagation, the loss function L is minimized fut , and obtain the optimal future trajectory.
[0155] Step 5-2: Unobserved historical trajectory X unobs Generation of.
[0156] 1. Gradual generation of historical trajectories: The back-diffusion model gradually generates unobserved historical trajectories X by learning from future trajectories and observed trajectories unobs In this process, the model is adjusted according to the generated future trajectory to ensure that the generated historical trajectory is consistent with the future trajectory. The formula is expressed as:
[0157]
[0158] in Represents the final generated historical trajectory.
[0159] 2. Optimize historical trajectories: Through multiple iterations, the model gradually optimizes historical trajectories using the guidance of future trajectories until the historical trajectories can be seamlessly connected with future trajectories in the model's time series. Its loss function is:
[0160]
[0161] in, is the unobserved historical trajectory generated by the model at time step t, is the true historical trajectory at time step t, indicating the actual position of the historical movement. Back propagation minimizes L unobs Optimizing unobserved historical trajectories.
[0162] Step 5-3: Generate application of trajectory.
[0163] The resulting future trajectory and unobserved historical trajectories Directly used as model output for trajectory prediction tasks, as shown below:
[0164]
[0165] These generated trajectories are particularly suitable for use in situations where observation data is scarce, providing reliable prediction results. For example, in scenarios such as autonomous driving, robot navigation, and intelligent transportation, future path directions and behaviors can be accurately predicted based on a small number of observed trajectories.
[0166] The implementation of the present invention is based on the Ubuntu 22.04 operating system, the GPU uses NVIDIA GTX4090TI, and all models and environmental simulations are implemented using PyTorch. In order to better test the performance of the trajectory prediction method of the bidirectional collaborative diffusion model proposed in the present invention, three simulation environments commonly used in baseline algorithms are selected: the ETH / UCY dataset, the nuScene dataset, and the Stanford Drone dataset. The present invention mainly conducts comparative analysis on the ETH / UCY dataset, including comparison with a variety of baseline methods (such as Social-STGCNN, Trajectron++, etc.) to evaluate the accuracy and generation performance of trajectory prediction. In addition, the prediction performance was tested in the nuScenes and Stanford Drone datasets to verify the generalization ability of the method in different scenarios.
[0167] This paper mainly uses the following indicators to evaluate the effects of different methods on trajectory prediction tasks:
[0168] 1. Average Displacement Error (ADE), defined as the average deviation distance of all entities within the prediction time window: 2. Final Displacement Error (FDE), defined as the deviation distance of all entities at the last prediction time step:
[0169] Table 1 summarizes the comparison of different methods on the ETH / UCY dataset. These indicators are expressed as ADE / FDE(m).
[0170]
[0171]
[0172] From the data in Table 1, it can be seen that the proposed method is better than other baseline models as a whole and improves the performance of pedestrian trajectory prediction. Among them, the HOTEL and AVG datasets are more difficult, and all models perform relatively poorly in these datasets, indicating that the dataset attributes affect the results.
[0173] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0174] While the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that many changes, modifications, substitutions and variations can be made to the embodiments without departing from the principles and spirit of the invention.
Claims
1. A method for instantaneous trajectory prediction based on a bidirectional cooperative diffusion model, which uses a forward diffusion process to gradually perturb data into random noise, and then gradually restores the original data from the noise through a reverse denoising process. The method is characterized by: S1: Process the input data to form a feature vector F obs ; S2: For F obs Encode the observed trajectory features v obs : S3: Generate X using the diffusion model unobs and X fut , where X unobs represents the unobserved historical trajectory, where X fut Indicates future trajectory; S4: Diffusion models use each other’s output as input for feedback adjustment and guidance; S5: After multiple iterations and mutual guidance, the diffusion model finally outputs the trajectory result.
2. The instantaneous trajectory prediction method based on the bidirectional cooperative diffusion model according to claim 1 is characterized in that: The diffusion model includes f fut and f unobs , where f fut represents the forward diffusion model, where f unobs represents the reverse diffusion model; Among them, using f unobs Generate X fut ; Among them, using f fut Generate X unobs .
3. The instantaneous trajectory prediction method based on the bidirectional cooperative diffusion model according to claim 2 is characterized in that: The input data in S1 include: Use X obs ={x1,x2} represents the observed trajectory data, which contains trajectory points of two time frames. x1 and x2 are the position coordinates of the pedestrian at two different time points. These two points represent the trajectory information collected by the model in a very short time. in, Contains the time from t=3 to t=T fut +2 a series of positions, each x i Where i≥3 represents the predicted position of the pedestrian at a certain point in the future, T fut represents the time span of the future trajectory; in, This is the trajectory point that needs to be generated, representing the pedestrian at the initial time point t = 1-T unobs A series of positions from t to t = 0, indicating that these historical trajectory data have not been observed in practice, T unobs represents the time span of the unobserved historical trajectories.
4. The instantaneous trajectory prediction method based on the bidirectional cooperative diffusion model according to claim 2 is characterized in that: Data processing in S1 includes: Enter the X obs , X fut and X unobs Formatting to ensure uniform input to the model, including representing trajectory points as vectors or matrices; Extract X obs The direction, speed, acceleration and environmental interaction characteristics of pedestrian movement; The calculated pedestrian movement direction, speed, acceleration and environmental interaction characteristics are integrated into a feature vector F obs , which is of the form: Among them, F obs is a matrix of shape N×d; Among them, N represents the number of time steps, d represents the data dimension of each time step, including the speed direction and characteristic information related to the environment.
5. The instantaneous trajectory prediction method based on the bidirectional cooperative diffusion model according to claim 2 is characterized in that: In S2, F obs The encoding includes: F obs Processed by the LSTM network and input into the encoder module of the model to generate high-level features and combine social interaction and scene context information; Use v obs Indicates the observed trajectory features; Use g soc Represents social interaction characteristics; Use f scene Represents scene context information; Let e represent the scene context information feature vector; v obs , g soc Combined with e, a complete feature representation v is generated total :v total =v obs +g soc +e.
6. The instantaneous trajectory prediction method based on the bidirectional cooperative diffusion model according to claim 2 is characterized in that: Generate X in S3 respectively fut and X unobs The following steps are involved: Using the encoded v obs and X unobs As a guide, guide the generation of X fut , this guidance process is through an LSTM network to unobs and X fut Encode and generate the forward guidance matrix g fut ,in: g fut =f Bi-LSTM (X unobs ) Among them, f Bi-LSTM Represents a function based on a bidirectional long short-term memory (Bi-LSTM) network. The Bi-LSTM network can capture both forward and backward information of sequence data, which helps to better understand the dependencies in time series. Then, the weight γ is learned through the gating mechanism Gate fut , balance v obs With g fut The contribution between them generates the final guidance information in, represents the control parameter of the guidance information, fut represents the forward guidance information parameter, which is used to adjust the output of each step of the generated trajectory, where: Finally, through The model uses a denoising mechanism to remove noise at each generation step; and generates the future trajectory for the next step and the next unobserved historical trajectory Where t represents the time step, and the superscript t-1 represents the trajectory generated in the previous step. The denoising process is performed by the following formula: Where X (t) represents all trajectories at t, X (t-1) represents the trajectory at t-1, α t is the parameter controlling the noise, ∈ represents the noise sampled from a Gaussian distribution with a mean of 0 and a covariance of the identity matrix I; Then, use X fut and the encoded v obs As guidance information, reverse guidance generates historical trajectories, and then the LSTM network is used to train X fut and v obs Encode and generate the reverse guidance matrix: g unobs =f Bi-LSTM (X fut ,v obs ), Among them, f Bi-LSTM Represents a function based on a bidirectional long short-term memory (Bi-LSTM) network. The Bi-LSTM network can capture both forward and reverse information of sequence data, which helps to better understand the dependencies in time series. g unobs is a string containing X fut and v obs vector, using a gating mechanism to learn the weights γ unobs , which is used to balance the influence of future trajectory and observed trajectory characteristics: The weights γ learned by the gating mechanism unobs Controlled g unobs and v obs contribution to ensure that the generated X unobs Can reflect trajectory dynamics more accurately.
7. The instantaneous trajectory prediction method based on the bidirectional cooperative diffusion model according to claim 2 is characterized in that: Feedback adjustment guidance in S4 includes: The forward diffusion model uses the backward diffusion model to generate g unobs , to adjust the generation of future trajectories, where g unobs Represents the reverse guidance matrix, and the reverse diffusion model uses g generated by the forward diffusion model fut , to guide the generation of unobserved historical trajectories, where g fut represents the forward guidance matrix; Given the forward diffusion model generated in represents the forward guidance information, which is used as f unobs The input guides the generation of: in: At the same time, the reverse diffusion model generates As f fut The input guides the generation of future trajectories, where Represents the reverse guidance matrix information: This mutual feedback process ensures the coherence of trajectory generation and gradually reduces the error by leveraging each other’s generation results, where t represents the time step.
8. The instantaneous trajectory prediction method based on the bidirectional cooperative diffusion model according to claim 7 is characterized in that: The influence of noise is balanced by learning the weights between different guidance information. The specific method is as follows: For each t, use γ t To adjust the ratio of generated trajectory to guidance information, where γ t The gating mechanism is expressed by calculating the gating weight: c t =σ(W g ·[v obs ,g unobs ,g fut ]) Among them, W g is a learnable weight matrix, σ is the activation function sigmoid, which is used to control the value range of the weight between [0,1]; γ t Act on the trajectory generated at each step to balance the contribution between observed data and generated data: X (t+1) =c t X (t) +(1-c t )v obs This process dynamically adjusts the degree of fusion between the generated trajectory and the observed trajectory to ensure that the noise is gradually weakened, making the final generated trajectory smoother and more accurate.
9. The instantaneous trajectory prediction method based on the bidirectional cooperative diffusion model according to claim 2 is characterized in that: Output X again in S5 fut and X unobs , including the following steps: Iteratively generate X fut Gradually converges to X obs and X unobs Consistent prediction path, T represents the maximum number of iteration steps, the formula is expressed as: Where t represents the time step, is the reverse guidance information matrix generated after t, is the final future trajectory after all iterations; The reverse diffusion model gradually generates the final historical trajectory by learning from future trajectories and observed trajectories. The formula is expressed as: in, Represents the final generated historical trajectory; and It is directly used as the model output result for trajectory prediction tasks. The formula is:
10. An automatic driving car, characterized in that: The automobile application comprises the method according to any one of claims 1 to 9.
Citation Information
Cited By
Pedestrian trajectory prediction method, system and device based on continuous learning and medium
CN120633467A
Pedestrian trajectory prediction method based on interactive perception diffusion model
CN121438409A