Traffic participant uncertainty trajectory prediction method based on condition fusion
By adopting a trajectory prediction method based on conditional fusion in the autonomous driving system, multiple trajectory branches are generated using Perceiver-style structure and diffusion model, the problem of difficult to characterize multiple motion patterns of traffic participants in the prior art is solved, and safer autonomous driving decisions are achieved.
Patent Information
- Application Number
- CN202510195470.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The prior art is difficult to effectively characterize the various movement modes of traffic participants, resulting in potential safety threats in autonomous driving decision planning.
Using a traffic participant uncertain trajectory prediction method based on conditional fusion, multiple uncertain trajectory branches are generated by collecting state information and lane information of each participant in the traffic scene.
It realizes effective representation of various movement modes of traffic participants, improves the ability to express uncertainty in real scenarios, and provides safer autonomous driving decisions and path planning.
Smart Images

Figure CN120164320A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous vehicle trajectory prediction, and particularly to a method for predicting uncertain trajectories of traffic participants based on conditional fusion. Background Art
[0002] In the fields of traffic and autonomous driving, trajectory prediction is a key task, aiming to infer the potential positions and movement trajectories of traffic participants (such as vehicles, pedestrians, cyclists, etc.) at future moments based on historical observations and environmental information. High-precision trajectory prediction is an important prerequisite for achieving safe and efficient decision-making (decision-making planning, collision avoidance, speed control, etc.).
[0003] Prediction methods based on physical models, such as Kalman filtering, can model and predict future trajectories within a short-term time domain through the state information of a vehicle at short historical moments, but the prediction results in a long-term time domain are unreliable.
[0004] In recent years, data-driven methods have increasingly become the mainstream of research. By collecting a large number of datasets for training neural networks to capture behavioral characteristics in the traffic environment for trajectory prediction, the prediction accuracy has been greatly improved, especially in the modeling of interactions between traffic participants within a scene. The application of models such as RNN, LSTM, Transformer, and graph neural networks in the field of time series data has enabled significant progress in capturing dynamic interactions in long-term time series and complex scenarios.
[0005] In real traffic scenarios, the same historical features may often lead to multiple future directions, such as continuing straight or choosing to merge. This makes it difficult to fully model the uncertainty within the future time domain by outputting a single predicted trajectory. Although many existing solutions can predict position coordinates in multiple modalities to a certain extent, it is difficult to represent various motion patterns of different traffic participants, posing a potential safety threat to autonomous driving decision-making and planning. Summary of the Invention
[0006] Objective of the Invention: To propose a method for predicting uncertain trajectories of traffic participants based on conditional fusion to solve the above problems existing in the prior art.
[0007] The method for predicting uncertain trajectories of traffic participants based on conditional fusion proposed by the present invention is as follows:
[0008] Collect the state information and lane information of each participant in the traffic scene within a predetermined past duration;
[0009] Pack the state information and lane information into an embedded vector sequence X, map the embedded vector sequence X to a latent array Z, and input the latent array Z and the embedded vector sequence X into an encoding network to obtain high-dimensional depth representation information C; the encoding network is stacked by several layers of encoders, and each layer of encoder contains a cross-attention module and a latent self-attention module;
[0010] Select a diffusion model, input the high-dimensional depth representation information C into it, and complete the training of the diffusion model;
[0011] Use the trained diffusion model to generate the predicted trajectory of the target, and obtain the trajectory prediction result considering uncertainty through multiple samplings.
[0012] In a further embodiment, the state information is represented in the following form:
[0013] X traj_i =[x,y,vx,vy,ax,ay]
[0014] In the formula, X traj_i represents the i-th traffic participant in the scene, and x, y, vx, vy, ax, ay respectively represent the coordinate, speed, and acceleration information of the participant in the horizontal and vertical directions;
[0015] The lane information is represented in the following form:
[0016] X lane_i =[x lane ,y lane ,point]
[0017] In the formula, X lane_i represents the information of the i-th lane line, and x lane ,y lane ,point respectively represent the abscissa, ordinate, and orientation of the lane line sampling point.
[0018] In a further embodiment, packing the state information and lane information into an embedded vector sequence X specifically includes:
[0019] Calculate the embedded features of the ego vehicle, neighboring vehicles, and lane information:
[0020]
[0021] In the formula, e ego , e nbrs,i ,, e lane,i respectively represent the embedded features of the ego vehicle, neighboring vehicles, and lane information; W ego , W nbr , W lane are the weight coefficients of each embedding layer; b ego , bnbr , b lane is the offset; the size of each embedded feature dimension is d x ;
[0022] Merge the embedded features of the ego vehicle, neighboring vehicles, and lane information into an embedded vector sequence X:
[0023] X = {e ego , e lane,1 , …, e nbr,1 , e nbr,2 , …, e nbr,n}.
[0024] In a further embodiment, map the embedded vector sequence X to a latent array Z, and input the latent array Z and the embedded vector sequence X into an encoding network to obtain high-dimensional depth representation information C, which specifically includes:
[0025] Initialize the latent array
[0026]
[0027] where M is the size of the latent array, and d z is the dimension size of the latent vector;
[0028] Input the initialized latent array and the embedded vector sequence X into the encoding network, and obtain the high-order representation feature Z of the multi-moment historical information in the traffic scene through the n-layer stacking of the encoding network n , and obtain the conditional information C for input to the diffusion model through non-linear mapping.
[0029] In a further embodiment, input the initialized latent array and the embedded vector sequence X into the encoding network, and the encoding process of the l-th layer is as follows:
[0030] The cross-attention mechanism uses the latent vector of the (l - 1)-th layer as the query vector, uses the input embedded vector sequence X as the key and value, and allows the latent vector to absorb multi-modal data information through the attention mechanism to obtain the output of the first layer.
[0031] In a further embodiment, obtain the cross-attention output result of the i-th layer through the cross-attention mechanism and the residual connection
[0032]
[0033] In the formula, represents the query vector of the m-th latent vector in the cross-attention of the i-th layer, respectively represent the key and value vectors of the i-th token of the input in the i-th layer; is a trainable coefficient matrix, and layernorm(*) is a layer normalization operation.
[0034] In a further embodiment, self-attention is performed on the latent vectors to allow the latent vectors to fuse with each other to form a final output encoding result
[0035]
[0036] In the formula, are respectively the query, key, and value of the latent vector in the self-attention mechanism.
[0037] In a further embodiment, after stacking n layers, the finally obtained latent vectors are concatenated to obtain a high-order representation feature Z of multi-moment historical information in the traffic scenario n , and conditional information C for inputting into the diffusion model is obtained through a non-linear mapping:
[0038]
[0039] C = MLP(Z n )
[0040] In the formula, MLP(*) is a non-linear mapping operation.
[0041] In a further embodiment, the high-dimensional depth representation information C is input into an open-source diffusion model to complete the training of the diffusion model, specifically including:
[0042] Noise addition stage: The target future trajectory Y 0 is gradually added with a predetermined level of Gaussian noise ε for forward noise addition. After k rounds of iteration, a noisy trajectory Y k ;
[0043] Training stage: In each round of iteration, let the denoising network ∈ θ simultaneously read the noisy trajectory Y k , the high-dimensional depth representation information C, and the current iteration round number k, and output the predicted noise By calculating the mean square error between the predicted noise and the actual Gaussian noise ε, the network is trained by backpropagation;
[0044] Inference stage: Initialize pure Gaussian noise and input it into the trained denoising network ∈ θ to obtain a predicted trajectory through reverse denoising.
[0045] In a further embodiment, the inference stage specifically includes the following process:
[0046] Call the denoising network:
[0047]
[0048] Wherein, represents the model's estimate of the noise in the iteration, and m is the current step number;
[0049] Reverse diffusion update:
[0050]
[0051] Wherein, represents the noisy trajectory at the current step, is the noisy trajectory at the (m - 1)-th step after update, β m is the noise coefficient, α m is used to determine the scaling of the variance during the inference process, n m ~Ν(0,Ι) indicates that random sampling can still be performed during the denoising process;
[0052] Recursively perform directional diffusion until m - 1 = 0 to obtain the final future prediction trajectory;
[0053] By repeatedly sampling different initial noises and n at each step m and performing the reverse diffusion update described above, different prediction trajectory results are generated.
[0054] In addition, the present invention also provides an electronic device, which includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the above-mentioned traffic participant uncertainty trajectory prediction method based on conditional fusion is implemented.
[0055] In addition, the present invention also provides a computer-readable storage medium, in which at least one executable instruction is stored, and when the executable instruction runs on an electronic device, the electronic device is caused to execute the above-mentioned traffic participant uncertainty trajectory prediction method based on conditional fusion.
[0056] Compared with the prior art, the present invention has at least the following beneficial effects:
[0057] (1) Adopt a Perceiver-like structure (latent array + cross-attention + latent self-attention) to efficiently represent the states of the ego vehicle and neighboring vehicles at several past moments, rather than a conventional time-series network structure such as a single RNN. The RNN network may have the problem of gradient disappearance when processing data in a long time series; while when using the self-attention mechanism on a large scale for a sequence of N tokens, it will generate N 2The computational complexity of the level. And this solution can aggregate the time-series information of a large number of moments and multiple vehicles to form a latent vector with strong condensation and expression ability. Moreover, through a latent array of fixed size (M << N), it can achieve the characteristics of controllable computational complexity and strong expression ability.
[0058] (2) Combine the diffusion model to generate future trajectories. This solution can capture various behavioral patterns in the driving scenario through large-scale data training, and obtain multiple uncertain trajectory branches from the sampling of CNG random noise. Conventional multi-modal trajectory prediction only realizes a rough multi-mode branch network division through discrete intention classification and cannot capture more possibilities. This solution can generate multiple trajectory distributions in a continuous probability space through random initialization and randomness between steps.
[0059] (3) Inject the Perceiver condition into the denoising network to make the trajectory generation strongly related to the historical scenario. In the denoising network of the conventional diffusion model, little in-depth context combination is carried out for the time-series model. This solution explicitly combines the historical features of multiple vehicles into the denoising network, performs conditional generation for the driving scenario, and enables the denoising network to combine historical information at each step for more accurate denoising restoration. Brief Description of the Drawings
[0060] Figure 1 It is a schematic diagram of the model structure constructed based on the prediction method proposed in the embodiment of the present invention.
[0061] Figure 2 It is a schematic diagram of the traffic participant uncertainty trajectory prediction system proposed in the embodiment of the present invention. Detailed Embodiments
[0062] In the following description, a large number of specific details are given to provide a more thorough understanding of the present invention. However, it is obvious to those skilled in the art that the present invention can be implemented without one or more of these details. In other examples, in order to avoid confusion with the present invention, some well-known technical features in the art are not described.
[0063] To make up for the deficiencies of the prior art, the embodiment of the present invention is based on the idea of "uncertainty trajectory prediction", integrates the high-dimensional representation of scene information by the deep neural network and the probability characteristics generated by multiple samplings of the diffusion model (Diffusion Model), and focuses on depicting multiple possible future trajectory branches of vehicles under the same historical conditions. The distribution of future trajectories is learned through end-to-end training. During inference, only by sampling the initial noise multiple times can we obtain prediction results that are closely related to the historical environment but different from each other, improving the expression ability of uncertainty in the real scenario and providing a reliable basis for decision-making and path planning in autonomous driving and complex traffic scenarios.
[0064] This embodiment discloses a method for predicting the uncertain trajectories of traffic participants based on condition fusion. The data flow is shown in Figure 1 .
[0065] The steps are as follows:
[0066] S1. Collect the state information of each participant in the traffic scene in the past 3 seconds, including position coordinates, speed, acceleration, and the information of traffic road lines:
[0067] X traj_i = [x, y, vx, vy, ax, ay]
[0068] X lane_i = [x lane , y lane , point]
[0069] In the formula, X traj_i represents the i-th traffic participant in the scene, and x, y, vx, vy, ax, ay respectively represent the coordinate, speed, and acceleration information of the participant in the horizontal and vertical directions. X lane_i represents the information of the i-th lane line, and x lane , y lane , point respectively represent the coordinate and orientation of the lane line sampling point.
[0070] S2. Pack the information obtained in S1 into a sequence, and through the input token mapping method, map it into the latent array Z of the Perciever encoder. After experiencing the multi-layer stacking process of cross-attention and latent self-attention, finally obtain the high-dimensional depth representation information C for subsequent multi-modal trajectory generation.
[0071] S3. Based on the encoded information C obtained in S2, use the diffusion model to generate the predicted trajectory of the target by denoising, and the trajectory prediction result considering uncertainty can be obtained through multiple samplings.
[0072] As a preferred embodiment, in step S1, the trajectory information, speed and acceleration information of the vehicle can be obtained by installing on-vehicle sensors, vehicle buses and vehicle body positioning modules on the vehicle. The state parameters of surrounding vehicles can be sensed and tracked in real time by using millimeter-wave radars and cameras. The lane line pixels can be identified through the front camera combined with the lane line detection algorithm and converted into road coordinates.
[0073] As a preferred embodiment, the diffusion model selects a mainstream open-source diffusion model, such as Stable Diffusion developed by StabilityAI. It supports functions such as text-to-image generation, image restoration, and super-resolution. Its core is based on the DDPM framework and the generation efficiency is optimized, and it can run on consumer-grade graphics cards.
[0074] For another example, DDPM (Denoising Diffusion Probabilistic Models) can be selected. This model supports custom noise schedulers (such as linear and cosine noise decay).
[0075] For another example, the Diffusers library (Hugging Face) can be selected. The Diffusers library is an open-source library launched by Hugging Face, integrating various diffusion models (such as Stable Diffusion, DALL·E mini, etc.) and providing a unified API interface.
[0076] As a preferred embodiment, the encoding process of the mapping and the model in step S2 is disclosed:
[0077] S2.1. Perform embedding representation and merging on the data obtained in S1:
[0078]
[0079] In the formula, e ego , e nbrs,i , e lane,i represent the embedding feature representations of the ego vehicle, adjacent vehicle, and lane information, and W ego , W nbr , W lane are the weight coefficients of the embedding layer. b ego , b nbr , b lane are the bias terms. All the embedding features are integrated into the set X, where the dimension size of each embedding feature is d x .
[0080] S2.2. Initialize the latent array
[0081]
[0082] where M is the size of the latent array and d z is the dimension size of the latent vector, which can be randomly initialized at the beginning of training and used as learnable parameters.
[0083] S2.3. Input the initialized latent array and the embedding vector sequence X into the encoding network. The encoder is stacked with n layers, and each layer consists of two steps: cross-attention and latent self-attention. After experiencing the multi-layer stacking process of the cross-attention and latent self-attention mechanisms, high-dimensional deep features integrating the modal information in the historical time domain of the scene are finally obtained.
[0084] As a preferred embodiment, a feasible solution detail of step S2.3 is disclosed:
[0085] S2.3.1. Taking the encoding process of the l-th layer as an example: First, the cross-attention mechanism uses the latent vector of the (l - 1)-th layer as the query vector, takes the input token sequence X as the key and value, and enables the latent vector to absorb multi-modal data information through the attention mechanism to obtain the output of the first layer. The specific process is as follows:
[0086] First, through the cross-attention mechanism and residual connection, the cross-attention output result of the i-th layer is obtained
[0087]
[0088] In the formula, represents the query vector of the m-th latent vector in the cross-attention of the i-th layer, respectively represent the key and value vectors of the i-th input token in the i-th layer. is a trainable coefficient matrix, and layernorm is a layer normalization operation.
[0089] Next, perform self-attention on the latent vectors to enable the latent vectors to fuse with each other to form a higher-order latent representation:
[0090]
[0091] Among them are respectively the query, key, and value of the latent vector in the self-attention mechanism. is the final output encoding result of this layer.
[0092] S2.3.2. After stacking n layers, the finally obtained latent vectors are concatenated to obtain a high-order representation feature Z of the multi-moment historical information in the traffic scenario n , and the conditional information C for inputting into the diffusion model is obtained through a non-linear mapping:
[0093]
[0094] C = MLP(Z n )
[0095] In the formula, MLP is a non-linear mapping operation.
[0096] As a preferred embodiment, a feasible solution detail of step S3 is disclosed:
[0097] S3.1. According to the conventional diffusion model training method, first add noise: During the training process, the target future trajectory Y 0Forward add noise by gradually adding a small amount of Gaussian noise ε. After k rounds of iteration, the trajectory almost becomes pure noise Y k .
[0098] S3.2. Train the denoising network according to the conventional diffusion model training method: At each diffusion step k, let the denoising network ∈ θ Read the noisy trajectory Y k , the Perceiver encoder condition C, and the current time k, and output the prediction of the noise amount
[0099] Calculate the predicted noise and the mean square error of the actual noise ε, and perform backpropagation training on the network
[0100] In the inference stage, there is no longer a real target future trajectory Y 0 , initialize pure Gaussian noise Input it into the trained denoising network ∈ θ to obtain the predicted trajectory through reverse denoising
[0101] S3.3.1. Call the denoising network
[0102]
[0103] In the formula represents the estimated amount of noise by the model in the iteration, and m is the current step number
[0104] S3.3.2. Reverse diffusion update
[0105]
[0106] In the formula represents the noisy trajectory at the current step is the noisy trajectory at the (m - 1)-th step after update, β m is the noise coefficient, α m is used to determine the scaling of the variance during the inference process, n m ~Ν(0,Ι) means that random sampling can still be performed during the denoising process to make the result more random
[0107] S3.3.3. Recursively perform forward diffusion in the above way until m - 1 = 0 to obtain the final future predicted trajectory
[0108] S3.3.4. By repeatedly sampling different initial noises and n at each step m , and performing the above reverse denoising process, different predicted trajectory results can be generated
[0109] As a preferred embodiment, seeFigure 2 As shown, a traffic participant uncertainty trajectory prediction system 400 is disclosed. The trajectory prediction system 400 at least includes a collection module 401, a first data processing module 402, a second data processing module 403, a model training module 404, and a prediction result output module 405.
[0110] The collection module 401 is used to collect the state information and lane information of each participant in the traffic scene in the past predetermined time period.
[0111] The first data processing module 402 is used to pack the collected state information and lane information into an embedded vector sequence X.
[0112] The second data processing module 403 embeds an encoding network, and the encoding network is stacked by several layers of encoders. Each layer of encoder contains a cross-attention module and a latent self-attention module respectively. The second data processing module 403 maps the embedded vector sequence X into a latent array Z, inputs the latent array Z and the embedded vector sequence X into the encoding network, and obtains high-dimensional depth representation information C.
[0113] The model training module 404 is used to input the high-dimensional depth representation information C into an open-source diffusion model to complete the training of the diffusion model.
[0114] Use the trained diffusion model to generate the predicted trajectory of the target, and obtain the trajectory prediction result considering uncertainty through multiple samplings, and output it through the prediction result output module 405.
[0115] As a preferred embodiment, an electronic device is disclosed. The electronic device includes a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus. At least one computer-readable storage medium is encapsulated in the memory, and at least one executable instruction is stored in the storage medium. When the computer instruction or computer program is loaded or executed on the electronic device, the processes or functions described in the embodiments of the present application are generated in whole or in part, which will not be elaborated here.
[0116] The electronic device can also communicate with one or more external devices (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device, and / or communicate with any device that enables the electronic device to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface. Moreover, the electronic device can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter. The network adapter communicates with other modules of the electronic device through a bus. It should be understood that although not shown in the figures, other hardware and / or software modules can be used in combination with the electronic device, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0117] It should be noted that the above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains a set of one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0118] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0119] As described above, although the present invention has been shown and described with reference to specific preferred embodiments, it should not be construed as a limitation on the present invention itself. Various changes may be made in form and detail without departing from the spirit and scope of the present invention as defined by the appended claims.
Claims
1. A method for predicting uncertain trajectories of traffic participants based on conditional fusion, characterized in that: The steps include: Collect status information and lane information of each participant in the traffic scene over the past predetermined period of time; The state information and lane information are packaged into an embedded vector sequence X, the embedded vector sequence X is mapped into a potential array Z, and the potential array Z and the embedded vector sequence X are input into an encoding network to obtain high-dimensional deep representation information C; the encoding network is stacked by several layers of encoders, and each layer of encoders includes a cross-attention module and a potential self-attention module; Inputting the high-dimensional deep representation information C into an open-source diffusion model to complete the training of the diffusion model; The trained diffusion model is used to generate the predicted trajectory of the target, and the trajectory prediction result considering uncertainty is obtained through multiple sampling.
2. The method for predicting uncertain trajectories of traffic participants based on conditional fusion according to claim 1 is characterized in that: The status information is expressed in the following form: X traj_i =[x,y,vx,vy,ax,ay] Where, X traj_i represents the i-th traffic participant in the scene, x, y, vx, vy, ax, ay represent the coordinates, speed, and acceleration information of the participant in the horizontal and vertical directions respectively; The lane information is expressed in the following form: X lane_i =[x lane ,y lane ,point] Where, X lane_i Indicates the information of the i-th lane line, x lane ,y lane ,point respectively represents the horizontal coordinate, vertical coordinate and direction of the lane line sampling point.
3. The method for predicting uncertain trajectories of traffic participants based on conditional fusion according to claim 2 is characterized in that: The state information and lane information are packaged into an embedded vector sequence X, specifically including: Calculate the embedded features of the vehicle, neighboring vehicles, and lane information: In the formula, e ego 、e nbrs,i ,e lane,i Respectively represent the embedded features of the vehicle, neighboring vehicle, and lane information; W ego , W nbr , W lane is the weight coefficient of each embedding layer; b ego , b nbr , b lane is the bias; each embedding feature dimension is d x ; Combine the embedded features of the vehicle, neighboring vehicles, and lane information into an embedded vector sequence X: X={and ego ,And lane,1 ,…,And nbr,1 ,And nbr,2 ,…,And nbr,n }。 4. The method for predicting uncertain trajectories of traffic participants based on conditional fusion according to claim 1, characterized in that: The embedded vector sequence X is mapped to a potential array Z, and the potential array Z and the embedded vector sequence X are input into the encoding network to obtain high-dimensional deep representation information C, which specifically includes: Initialize the potential array Where M is the size of the potential array, d z is the dimension size of the potential vector; The potential array that will be initialized The embedding vector sequence X is input into the encoding network, and the high-order representation feature Z of the multi-time historical information in the traffic scene is obtained through the n-layer stacking of the encoding network. n , and obtain the conditional information C used to input the diffusion model through nonlinear mapping.
5. The method for predicting uncertain trajectories of traffic participants based on conditional fusion according to claim 1 or 4, characterized in that: The potential array that will be initialized And the embedding vector sequence X is input into the encoding network, and the encoding process of the lth layer is as follows: The cross attention mechanism converts the latent vector of the l-1 layer As the query vector, the input embedding vector sequence X is used as the key and value, and the potential vector is allowed to absorb multimodal data information through the attention mechanism to obtain the output of the first layer.
6. The method for predicting uncertain trajectories of traffic participants based on conditional fusion according to claim 5, characterized in that: The cross attention output result of the i-th layer is obtained through the cross attention mechanism and residual connection In the formula, represents the query vector of the mth latent vector in the criss-cross attention layer i, Represent the key and value vectors of the input i-th token in the i-th layer respectively; is a trainable coefficient matrix, and layernorm(*) is a layer normalization operation.
7. The method for predicting uncertain trajectories of traffic participants based on conditional fusion according to claim 6, characterized in that: Perform self-attention between latent vectors to allow the latent vectors to merge with each other to form the final output encoding result In the formula, The query, key, and value of the latent vector in the self-attention mechanism, respectively.
8. The method for predicting uncertain trajectories of traffic participants based on conditional fusion according to claim 7, characterized in that: After n layers of stacking, the final latent vectors are concatenated to obtain the high-order representation feature Z of multi-time historical information in the traffic scene n , and obtain the conditional information C used to input the diffusion model through nonlinear mapping: C=MLP(Z n ) Where MLP(*) is a nonlinear mapping operation.
9. The method for predicting uncertain trajectories of traffic participants based on conditional fusion according to claim 1 or 8, characterized in that: The high-dimensional deep representation information C is input into the open-source diffusion model to complete the training of the diffusion model, which specifically includes: Noise adding stage: the target future trajectory Y 0 Gradually add Gaussian noise ε of a predetermined magnitude for forward noise addition. After k rounds of iterations, the noisy trajectory Y is obtained. k ; Training phase: In each round of iteration, let the denoising network ∈ θ At the same time, read the noisy track Y k , high-dimensional deep representation information C and the current iteration number k, output prediction noise Predicting noise by calculation The mean square error with the actual Gaussian noise ε is used to train the network through back propagation; Inference stage: Initialize pure Gaussian noise Input to the trained denoising network ∈ θ In , the predicted trajectory is obtained by reverse denoising.
10. The method for predicting uncertain trajectories of traffic participants based on conditional fusion according to claim 9, characterized in that: The reasoning stage specifically includes: Call the denoising network: In the formula, It represents the model's estimate of the noise in the iteration, and m is the current number of steps; Backward Diffusion Update: In the formula, represents the noisy trajectory of the current step, is the noisy trajectory of the m-1th step after update, β m is the noise coefficient, α m In the inference process, it is used to determine the variance scaling, n m ~Ν(0,Ι) means that random sampling is still possible during the denoising process; The recursive direction diffuses until m-1=0, and the final future prediction trajectory is obtained; By repeatedly sampling different initial noise and n at each step m And perform the reverse diffusion update to generate different prediction trajectory results.
Citation Information
Patent Citations
Scene-level multi-agent track generation method and device based on consistent diffusion
CN117473032A
Object recommendation method and device, medium and computing equipment
CN118170992A
Automatic driving track prediction method and device based on diffusion model
CN118636913A
Remote sensing image generation method and device based on multi-condition controllable diffusion model
CN118982597A
Universal physics transformers for efficient scaling of neural operators
DE202024104633U1
Cited By
Accident scene generation method based on conditional flow matching model
CN122088721A
An accident scene generation method based on a conditional flow matching model
CN122088721B
Vehicle trajectory prediction method and system based on diffusion model and transformer
CN122510851A