Scene-level Multi-agent Trajectory Generation Method and Device Based on Consistent Diffusion

Through a scene-level multi-agent trajectory generation method based on consistent diffusion, the scene encoder and denoising network are used to process the agent trajectory, and the problem of interaction and consistency of multiple types of agent trajectories in the prior art is solved. The generated trajectory is real, smooth and has scene consistency.

CN117473032BActive Publication Date: 2025-05-30SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311548038.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-30
Estimated Expiration
2043-11-20

AI Technical Summary

Technical Problem

The prior art cannot effectively handle the interaction and consistency between trajectories of multiple types of agents, resulting in the generated trajectories being unreal and lacking local smoothing characteristics.

Method used

Using a scene-level multi-agent trajectory generation method based on consistent diffusion, the historical trajectory and context information of the agent are extracted through the scene encoder, combined with the denoising network and the Transformer model, the noise is gradually predicted and the trajectory consistency is ensured, and the joint future trajectory of multiple agents is generated.

Benefits of technology

Effective processing of interaction and consistency between trajectories of various types of agents is realized. The generated trajectories have local smoothing characteristics, and the authenticity and scene consistency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117473032B_ABST
    Figure CN117473032B_ABST
Patent Text Reader

Abstract

This application relates to the field of autonomous driving technology, and discloses a scenario-level multi-agent trajectory generation method and device based on consistent diffusion. In the training stage, the method performs shifting and splicing processing on the future trajectory sequence to construct an enhanced trajectory sequence. Then, the same method is used to process the noise sequence, and it is added to the enhanced trajectory sequence according to Gaussian transfer, so that the overlapping parts of adjacent elements in the enhanced trajectory sequence are added with the same noise. In this way, the smoothness of the generated agent trajectory can be improved. In the generation stage, a noise sequence is sampled from a Gaussian distribution and subjected to shifting and splicing processing, and the obtained enhanced and noisy trajectory sequence is input into a denoising network for denoising processing to generate the joint future trajectories of multiple agents. Among them, the denoising network ensures the consistency of information of the same state in the enhanced and noisy trajectory sequence through temporal consistency guidance during the step-by-step Gaussian state transition process, so as to improve the local smoothness of the generated joint future trajectories.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of autonomous driving, and specifically relates to a method and device for generating scenario-level multi-agent trajectories based on consistent diffusion. Background Art

[0002] Realistic scenario-level multi-agent motion simulation is crucial for the development and evaluation of autonomous driving algorithms. Traffic simulation, as a supplement to real-world recorded traffic scenarios, provides an economical and safe way to evaluate autonomous driving systems before they are deployed into the real world. However, there is diversity in agent types (including vehicles, pedestrians, bicycles, etc.), and there are complex interactions between multiple types of agents; and the agent motion has problems of uncertainty and multi-modal nature. Therefore, the generation of such scenarios is not simple.

[0003] In existing technical solutions, the solution of generating trajectories by defining the motion rules of agents is difficult to provide complex and realistic traffic scenarios.

[0004] Using the results of motion prediction tasks to generate multi-modal traffic scenarios, most are single-type trajectory generation methods and cannot be simply and directly extended to scenario-level multi-agent trajectory generation because they cannot handle the different interactions and consistencies between the generated trajectories of multiple types of agents. For example, the driving process of a vehicle is modeled as a Markov process, and a deep neural network is used to implement the state distribution and transition function. A multi-context gating module is used to process the interactions in the observed data, and a Gaussian mixture model is used to characterize the diversity of the generated trajectories. A Gated Recurrent Unit (GRU) and Convolutional Neural Networks (CNN) are used to learn multi-agent behavior from real-world data, and a latent variable model is used to formulate the joint motion strategy of agents. A collision mitigation strategy is introduced to improve the trajectories generated by the motion prediction model.

[0005] Using a generative model to learn the probability distribution of trajectory data and generate new trajectory samples, recent work has started to use methods based on diffusion models. For example, in a related technology, a framework based on a diffusion model is used to implement pedestrian trajectory prediction, where scene diffusion uses latent diffusion in an end-to-end differentiable architecture to generate a position sequence for agents. Some related technologies also propose a conditional diffusion model to achieve controllable vehicle trajectory generation, making the generated trajectories have desired attributes such as speed limits. However, these methods based on diffusion models only consider the trajectory generation of a single type of agent and cannot guarantee the local smoothness characteristics of the generated trajectories like those of agents in the real world, thus prone to generating unrealistic agent motion trajectories. Summary of the Invention

[0006] The embodiment of the present application provides a scenario-level multi-agent trajectory generation method based on consistent diffusion to solve the problems in the prior art that the interaction and consistency between the generated trajectories of multiple types of agents cannot be processed, and the local smoothness characteristics of the generated trajectories cannot be guaranteed to be the same as those of the agent trajectories in the real world, thus easily generating unrealistic agent movement trajectories.

[0007] Correspondingly, the embodiment of the present application also provides a scenario-level multi-agent trajectory generation device based on consistent diffusion and an electronic device to ensure the implementation and application of the above method.

[0008] To solve the above technical problems, the embodiment of the present application discloses a scenario-level multi-agent trajectory generation method based on consistent diffusion, and the method includes:

[0009] In the training stage:

[0010] Perform shifting and splicing processing on the real future trajectory sequence in the preset training set to obtain an enhanced trajectory sequence;

[0011] Perform shifting and splicing processing on the noise sequence sampled from the Gaussian distribution to obtain an enhanced noise sequence, and add the enhanced noise sequence to the enhanced trajectory sequence according to the Gaussian transition to obtain an enhanced noisy trajectory sequence;

[0012] Use the scenario encoder to extract the historical trajectory and context information of the agents in the training set to obtain the vector representation of the agents, and input the vector representation of the agents, the enhanced noisy trajectory sequence, and the diffusion step into the denoising network based on Transformer to gradually predict the added noise through the Gaussian state transition to obtain the future trajectory;

[0013] Optimize the scenario encoder and the denoising network using the preset loss function;

[0014] In the generation stage:

[0015] Perform shifting and splicing processing on the noise sequence sampled from the Gaussian distribution to obtain an enhanced noisy trajectory sequence;

[0016] Use the optimized scenario encoder to extract the historical trajectory and context information of the agents in the target historical scenario to obtain the vector representation of the agents, and input the vector representation of the agents, the enhanced noisy trajectory sequence, and the diffusion step into the optimized denoising network to output the joint future trajectories of multiple agents;

[0017] Wherein, in the generation stage, the denoising network ensures the consistency of the information of the same state in the enhanced noisy trajectory sequence through temporal consistency guidance during the gradual Gaussian state transition process.

[0018] The embodiments of the present application also disclose a scene-level multi-agent trajectory generation device for consistent diffusion, which includes a training module and a generation module;

[0019] The training module includes:

[0020] A trajectory enhancement module, configured to perform shifting and splicing processing on the true future trajectory sequence in the preset training set to obtain an enhanced trajectory sequence;

[0021] A noise addition module, configured to perform shifting and splicing processing on the noise sequence sampled from the Gaussian distribution to obtain an enhanced noise sequence, and add the enhanced noise sequence to the enhanced trajectory sequence according to the Gaussian transition to obtain an enhanced noise-added trajectory sequence;

[0022] A denoising module, configured to use the scene encoder to extract the historical trajectory and context information of the agents in the training set to obtain the vector representation of the agents, and input the vector representation of the agents, the enhanced noise-added trajectory sequence, and the diffusion step into the denoising network based on Transformer to gradually predict the added noise through the Gaussian state transition to obtain the future trajectory;

[0023] An optimization module, configured to optimize the scene encoder and the denoising network by using a preset loss function;

[0024] The generation module includes:

[0025] A noise sequence acquisition module, configured to perform shifting and splicing processing on the noise sequence sampled from the Gaussian distribution to obtain an enhanced noise-added trajectory sequence;

[0026] A trajectory prediction module, configured to use the optimized scene encoder to extract the historical trajectory and context information of the agents in the target historical scene to obtain the vector representation of the agents, and input the vector representation of the agents, the enhanced noise-added trajectory sequence, and the diffusion step into the optimized denoising network to output the joint future trajectories of multiple agents;

[0027] Wherein, in the generation module, during the gradual Gaussian state transition process, the denoising network ensures the consistency of the information of the same state in the enhanced noise-added trajectory sequence through temporal consistency guidance.

[0028] The embodiments of the present application also disclose an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements one or more of the methods in the embodiments of the present application.

[0029] In the embodiments of the present application, in the training stage, the true future trajectory sequences in the preset training set are shifted and spliced to obtain enhanced trajectory sequences, such that there is partial overlap between adjacent elements in the enhanced trajectory sequences. Then, the same method is used to shift and splice the noise sequences sampled from the Gaussian distribution to obtain enhanced noise sequences, and the enhanced noise sequences are added to the enhanced trajectory sequences according to Gaussian transitions, so that the overlapping parts between adjacent elements in the enhanced trajectory sequences are added with the same noise. By constructing the enhanced trajectory sequences and the specific noise addition method above, the smoothness of the agent trajectories generated in the diffusion model can be improved. Then, the scene encoder is used to extract the historical trajectories and context information of the agents in the training set to obtain the vector representations of the agents, and the vector representations of the agents, the enhanced noisy trajectory sequences, and the diffusion steps are input into the denoising network based on Transformer to gradually predict the added noise through Gaussian state transitions to obtain future trajectories; and a preset loss function is used to optimize the scene encoder and the denoising network. In the generation stage, the noise sequences sampled from the Gaussian distribution are directly shifted and spliced to obtain enhanced noisy trajectory sequences, and then the vector representations of the agents extracted by the scene encoder, the enhanced noisy trajectory sequences, and the diffusion steps are input into the optimized denoising network, and the joint future trajectories of multiple agents can be output to ensure scene consistency. Among them, in the generation stage, the denoising network ensures the consistency of the information of the same state in the enhanced noisy trajectory sequences through temporal consistency guidance during the gradual Gaussian state transition process, ensuring the temporal consistency during the denoising process to improve the local smoothness of the generated joint future trajectories.

[0030] Additional aspects and advantages of the embodiments of the present application will be given in the following description section, which will become apparent from the following description or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, wherein:

[0032] Figure 1 is a flowchart of the method for generating scene-level multi-agent trajectories based on consistent diffusion in the training stage provided by the embodiments of the present application;

[0033] Figure 2 is a flowchart of the method for generating scene-level multi-agent trajectories based on consistent diffusion in the generation stage provided by the embodiments of the present application;

[0034] Figure 3 is a schematic diagram of the framework structure of the SceneDM model provided by the embodiments of the present application;

[0035] Figure 4Schematic structural diagram of the scenario-level multi-agent trajectory generation device based on consensus diffusion in the training phase provided by the embodiments of the present application;

[0036] Figure 5 Schematic structural diagram of the scenario-level multi-agent trajectory generation device based on consensus diffusion in the generation phase provided by the embodiments of the present application;

[0037] Figure 6 Schematic structural diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0038] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation to the present application.

[0039] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their combinations. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0040] Those skilled in the art can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as here.

[0041] The solution provided by the embodiments of the present application can be executed by any electronic device, such as a terminal device or a server. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions in this regard. For the technical problems existing in the prior art, the method and device for generating scene-level multi-agent trajectories based on consistent diffusion provided by the present application aim to solve at least one of the technical problems in the prior art.

[0042] The following uses specific embodiments to elaborate in detail on the technical solution of the present application and how the technical solution of the present application solves the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of the present application in conjunction with the drawings.

[0043] An embodiment of the present application provides a possible implementation manner, providing a flowchart of a method for generating scene-level multi-agent trajectories based on consistent diffusion. This solution can be executed by any electronic device, and optionally, can be executed on the server side or the terminal device.

[0044] This method may include the following steps:

[0045] As Figure 1 shown in, in the training stage:

[0046] Step 101, perform shifting and splicing processing on the real future trajectory sequence in the preset training set to obtain an enhanced trajectory sequence.

[0047] Step 102, perform shifting and splicing processing on the noise sequence sampled from the Gaussian distribution to obtain an enhanced noise sequence, and add the enhanced noise sequence to the enhanced trajectory sequence according to the Gaussian transfer to obtain an enhanced noisy trajectory sequence.

[0048] In the real scenario, when an agent moves in the environment, it shows a smooth and continuous motion pattern. Therefore, local smoothness is an important factor for evaluating the authenticity of the generated trajectory. In the embodiments of the present application, the SceneDM model is proposed to implement the method in the embodiments of the present application. The model framework includes a scene encoder for learning the vectorized representation of the dynamic scene and a Transformer-based denoising network for the reverse diffusion process. As Figure 3As shown in the figure, the method proposed in the embodiment of the present application includes a training stage and a generation stage. In the training stage, training and learning are carried out based on a diffusion model. The diffusion model consists of two parts: a diffusion process that gradually converts the data distribution into unstructured noise and an inverse process that restores the data distribution.

[0049] During the diffusion process, the true future trajectory sequence in the preset training set is shifted and spliced to obtain an enhanced trajectory sequence, so that there is partial overlap in the states between adjacent frames in the enhanced trajectory sequence, providing rich information input for subsequent processing. Then, Gaussian noise is gradually added to the enhanced trajectory sequence. The Gaussian noise in the embodiment of the present application is a noise sequence randomly sampled from the Gaussian distribution N(0,1). In order to induce the generated sequence of the diffusion model to have similarity between adjacent elements, the embodiment of the present application applies the same enhancement method to the sampled noise sequence as the future trajectory sequence, and gradually adds the enhanced noise sequence obtained after enhancement to the enhanced trajectory sequence according to the Gaussian transfer to obtain an enhanced noisy trajectory sequence. Through the above method, the similarity between adjacent elements of the enhanced trajectory sequence can be improved, and correspondingly, the local smoothness of the original trajectory sequence can be improved, thereby improving the authenticity of the generated trajectory.

[0050] Gradually adding the enhanced noise sequence to the enhanced trajectory sequence according to the Gaussian transfer can be achieved in the following way:

[0051] For the enhanced trajectory sequence s 0 and the enhanced noisy trajectory sequence S k . In the forward diffusion process, Gaussian noise is gradually added to the enhanced trajectory sequence s 0 to obtain latent variables s 1 , s 2 , …, s K , where K represents the maximum number of diffusion steps. According to the diffusion probability model (DDPM, Denoising Diffusion Probabilistic Models), the forward diffusion process is parameterized as a Markov chain, and the final latent variable can be expressed as

[0052]

[0053] where is a positive constant representing the noise level, ∈ represents the noise sampled from the Gaussian distribution . When K is large enough, s K converges to the Gaussian distribution.

[0054] Step 103: Use the scenario encoder to extract the historical trajectory and context information of the agent in the training set to obtain the vector representation of the agent, and input the vector representation of the agent, the augmented noisy trajectory sequence, and the diffusion step into the Transformer-based denoising network to gradually predict the added noise through Gaussian state transition to obtain the future trajectory.

[0055] In the embodiment of the present application, the inverse process of diffusion is based on Figure 3 the Transformer-based denoising network in Figure 3 . In the encoding stage, use the scenario encoder to extract the historical trajectory and context information of the agent in the training set to obtain the vector representation of the agent, and use the vector representation of the agent as the condition c. By embedding the condition c into the Transformer-based denoising network, the noise in the augmented noisy trajectory sequence is eliminated. Specifically, input the vector representation of the agent (condition c), the augmented noisy trajectory sequence, and the diffusion step into the Transformer-based denoising network to gradually predict the added noise through Gaussian state transition to obtain the future trajectory.

[0056] Given diffusion steps 1, 2, …, K and condition c, the diffusion model represents the inverse process as follows:

[0057]

[0058]

[0059] where θ represents the parameters of the entire model.

[0060] The diffusion model is optimized to approximate p θ (s k-1 |s k , c, k) or equivalently predict the noise ∈ added during the diffusion process according to the objective function.

[0061] Step 104: Use the preset loss function to optimize the scenario encoder and the denoising network.

[0062] As Figure 2 shown in

[0063] In the generation stage:

[0064] Step 201: Perform shifting and splicing processing on the noise sequence sampled from the Gaussian distribution to obtain the augmented noisy trajectory sequence.

[0065] Step 202: Use the optimized scene encoder to extract the historical trajectory and context information of the agent in the target historical scene to obtain the vector representation of the agent, and input the vector representation of the agent, the enhanced noisy trajectory sequence, and the diffusion step into the optimized denoising network to output the joint future trajectories of multiple agents.

[0066] Among them, in the generation stage, the denoising network ensures the consistency of information of the same state in the enhanced noisy trajectory sequence through temporal consistency guidance in the step-by-step Gaussian state transition process.

[0067] In the generation stage, the present application embodiment proposes a temporal consistency guidance strategy for the consistency diffusion method, where the temporal consistency guidance strategy is extended from the denoising diffusion implicit model (DDIM). During the denoising process, the consistency of information of the same state in the enhanced noisy trajectory sequence is ensured through temporal consistency guidance. The optimized denoising network can generate joint future trajectories with scene consistency for various types of agents in the target historical scene.

[0068] In the embodiment of the present application, the current time is denoted as t = 0. The future trajectory of the agent is represented as where y t is an H-dimensional vector including 3-D coordinates and heading, and T represents the length of the generated future trajectory.

[0069] The method in the embodiment of the present application can generate trajectories with scene consistency for various types of agents in the scene, including but not limited to vehicles, pedestrians, and bicycles.

[0070] In the embodiments of the present application, during the training phase, the true future trajectory sequence in the preset training set is shifted and spliced to obtain an enhanced trajectory sequence, such that there is partial overlap between adjacent elements in the enhanced trajectory sequence. Then, the same method is used to shift and splice the noise sequence sampled from the Gaussian distribution to obtain an enhanced noise sequence, and the enhanced noise sequence is added to the enhanced trajectory sequence according to the Gaussian transition, so that the overlapping parts of adjacent elements in the enhanced trajectory sequence are added with the same noise. By constructing the enhanced trajectory sequence and the specific noise addition method as described above, the smoothness of the generated agent trajectories in the diffusion model can be improved. Then, the scene encoder is used to extract the historical trajectories and context information of the agents in the training set to obtain the vector representation of the agents, and the vector representation of the agents, the enhanced and noise-added trajectory sequence, and the diffusion step are input into the denoising network based on Transformer to gradually predict the added noise through the Gaussian state transition to obtain the future trajectory; and the preset loss function is used to optimize the scene encoder and the denoising network. During the generation phase, the noise sequence sampled from the Gaussian distribution is directly shifted and spliced to obtain an enhanced and noise-added trajectory sequence, and then the vector representation of the agents extracted by the scene encoder, the enhanced and noise-added trajectory sequence, and the diffusion step are input into the optimized denoising network, and the joint future trajectories of multiple agents can be output to ensure scene consistency. Among them, during the generation phase, the denoising network ensures the consistency of the information of the same state in the enhanced and noise-added trajectory sequence through temporal consistency guidance during the gradual Gaussian state transition process, ensuring the temporal consistency during the denoising process, so as to improve the local smoothness of the generated joint future trajectories.

[0071] In an alternative embodiment, shifting and splicing the true future trajectory sequence in the preset training set to obtain an enhanced trajectory sequence includes:

[0072] Shifting the elements at each position in the future trajectory sequence, and splicing the elements after shifting at each position with the elements before shifting to obtain the enhanced trajectory sequence.

[0073] In the embodiments of the present application, the future trajectory sequence is enhanced by splicing the motion states of adjacent positions, thereby achieving overlapping parts among the elements at adjacent positions, and applying the same noise to the overlapping parts. Then, the model is trained to predict the added noise.

[0074] As Figure 3 shown, first, the future trajectory sequence s 0 = [y 1 , y 2 , …, y T is enhanced such that there are overlapping parts among the elements at adjacent positions. Specifically, for the future trajectory sequence s 0 = [y 1 , y 2 , …, yT each state (element) y in t is concatenated with the state y at the next moment t+1 By concatenating the states of adjacent frames, an augmented variable is created Augmented variable combines information from y t and y t+1 providing a rich input for subsequent processing. Most importantly, overlaps with and After augmenting the states at each moment in the future trajectory sequence using the above method, an augmented trajectory sequence of the trajectory states is obtained, denoted as

[0075] In an alternative embodiment, a noise sequence sampled from a Gaussian distribution is shifted and concatenated to obtain an augmented noise sequence, including:

[0076] The elements at each position in the noise sequence are shifted, and the shifted elements at each position are concatenated with the unshifted elements to obtain the augmented noise sequence.

[0077] After augmenting the future trajectory sequence, Gaussian noise is gradually added to the augmented trajectory sequence S 0 to obtain the augmented noisy trajectory sequence S k . To maintain the consistency of the state information at the same timestamp within the augmented noisy trajectory sequence S k , the same augmentation method is applied to the sampled noise sequence in the embodiments of the present application. Specifically, first sample the noise sequence from then the noise sequence ∈ 0 is shifted, and the shifted elements at each position are concatenated with the unshifted elements to obtain the augmented noise sequence of the augmented trajectory sequence S 0 i.e., namely

[0078]

[0079]

[0080] In an alternative embodiment, an optimized scene encoder is used to extract the historical trajectory and context information of the agent in the target historical scene to obtain the vector representation of the agent, and the vector representation of the agent, the augmented noisy trajectory sequence, and the diffusion step are input into the optimized denoising network to output the joint future trajectories of multiple agents, including:

[0081] The optimized scene encoder is used to extract the historical trajectory and context information of the agent in the target historical scene to obtain the vector representation of the agent;

[0082] The vector representation of the agent, the enhanced feature, and the diffusion step are fused to obtain the fused feature;

[0083] The fused feature is input into the Transformer module to predict the joint future trajectory; wherein, the Transformer module includes alternating temporal self-attention mechanism and agent self-attention mechanism.

[0084] In the embodiment of the present application, the joint future trajectories of multiple agents are predicted through the Transformer module in the denoising network. The attention mechanism is adopted in the Transformer module to process the interaction between multiple agents and the temporal correlation of the trajectories. The temporal attention mechanism enables the model to learn to extract the changes of the trajectories over time. At the same time, the spatial attention mechanism (i.e., the agent self-attention mechanism along the agent dimension) drives the model to depict the interaction between agents and generate consistent trajectories, that is, to avoid collisions between agents. Specifically, in the embodiment of the present application, the alternating temporal attention mechanism and spatial attention mechanism are used as the basic modules of the decoder network. Multiple such modules are stacked together to process the complex interactions in the denoising process.

[0085] In an optional embodiment, fusing the vector representation of the agent, the enhanced feature, and the diffusion step to obtain the fused feature includes:

[0086] Encoding the preset diffusion step and the enhanced noisy trajectory sequence through a multi-layer perceptron and fusing them with the vector representation of the agent to obtain the fused feature.

[0087] As Figure 3 shown, the vector representation of the agent obtained by the scene encoder is used as the embedding condition c of the denoising network. For each diffusion step k, first perform positional encoding on the diffusion step and the noise variable s k through a multi-layer perceptron (MLP), and then further combine it with the condition c to form the fused feature. The vector representation obtained by the scene encoder passes through a 1×1 convolutional layer and a multi-layer perceptron in sequence to obtain the condition c.

[0088] To highlight the sequence data s 0 =[y 1 ,y 2 ,…,y TRegarding the positional relationship of [], in the embodiments of the present application, positional encoding is further applied to the fused features. Then, the features are input into a Transformer module composed of alternating temporal and agent attention layers. Finally, a multi-layer perceptron outputs the noise to be eliminated based on the features obtained from the Transformer module, and then a joint future trajectory can be obtained.

[0089] In the embodiments of the present application, the denoising network is optimized according to the following formula:

[0090]

[0091] In an optional embodiment, after inputting the fused features into the Transformer module and predicting the joint future trajectory, the method further includes:

[0092] Regularize the differences between adjacent position elements of each trajectory in the joint future trajectory.

[0093] In the embodiments of the present application, regularization is used to further improve the smoothness of the joint future trajectory. The smoothness is calculated as the difference between adjacent position elements of the original sequence data (the future trajectory without augmentation and noise addition processing), which can represent adjacent motion states, and equivalently, the difference between the first half and the second half of the elements in the augmented and noisy trajectory sequence. For example, when the motion state represents the linear velocity of an agent, the adjacent state difference reflects the linear acceleration. The smoothness loss term regularizes the differences between adjacent states of the generated joint future trajectory to be close to the differences of the real sequence trajectories observed in the real world. Mathematically,

[0094]

[0095] where, represents the predicted noise, and represent its first part and second part respectively.

[0096] By combining the above two loss functions, a hybrid optimization objective is obtained:

[0097] L hybrid = L mse + λL smooth .

[0098] By introducing a smooth regularization term into the overall loss function, the model is driven to generate trajectories with real and smooth motion patterns. The parameters of the scene encoder and the Transformer-based denoising network are trained simultaneously. The hyperparameter λ is used to adjust the balance between these two losses.

[0099] In an optional embodiment, the temporal consistency guidance makes the information of the same state in the enhanced noisy trajectory sequence consistent through averaging operations.

[0100] During the denoising process, for the information of the same state in the enhanced noisy trajectory sequence corresponding to the original sequence data (such as and ), the temporal consistency guidance strategy ensures the consistency of their information during the denoising process through averaging operations:

[0101]

[0102] During the sampling process, the model SceneDM provided in the embodiment of the present application optimizes and generates joint future trajectories by iteratively calculating Gaussian transitions from k = K to k = 0, as shown in the following formula:

[0103]

[0104] where represents the noise level at diffusion step k. By iteratively applying Gaussian transitions, the sampling process gradually eliminates noise and generates real joint future trajectories.

[0105] In an optional embodiment, after using the optimized scene encoder to extract the historical trajectories and context information of the agents in the target historical scene to obtain the vector representations of the agents, and inputting the vector representations of the agents, the enhanced noisy trajectory sequence, and the diffusion steps into the optimized denoising network to output and obtain the joint future trajectories of multiple agents, the method further includes:

[0106] Calculating the joint future trajectories according to the scene scoring module to obtain the scores of the joint future trajectories.

[0107] In an optional embodiment, calculating the joint future trajectories according to the scene scoring module to obtain the scores of the joint future trajectories includes:

[0108] Calculating the number of collisions of the agents by calculating the overlapping area between the bounding boxes of the agents in the joint future trajectories to obtain a safety verification metric value;

[0109] Calculating the number of times the agents cross the road boundary in the joint future trajectories to obtain a road compliance metric value;

[0110] Obtaining the scores of the joint future trajectories according to the safety verification metric value and the road compliance metric value.

[0111] The models in the embodiments of this application may generate scenarios that do not conform to reality or violate traffic rules, such as agents colliding or exceeding the road boundaries. Such data may have an adverse impact on subsequent autonomous driving simulation tasks. To address this issue, a scenario-level scoring module is proposed in the embodiments of this application for scenario-level scoring. Specifically, it evaluates multiple joint future trajectories corresponding to the generated multiple scenarios, and the aspects of evaluation include safety verification and road compliance metrics.

[0112] Specifically, the safety verification metric value is calculated by computing the overlapping area between agent bounding boxes. In the embodiments of this application, the penetration depth is calculated to determine the maximum overlapping distance between any two agents. A positive value indicates a collision between the two. For each generated candidate trajectory s i , the collision with other agents in the scenario is computed in parallel, and the number of collisions is denoted as r 1 (s i ).

[0113] Similarly, the road compliance metric value is measured by r 2 (s i ), which represents the number of times an agent crosses the road boundary. The final trajectory scoring function of s i is obtained as follows:

[0114] F(s i ) = r 1 (s i ) + r 2 (s i ).

[0115] Among them, the lower F(s i ), the more compliant the generated trajectory s i is with traffic rules. By calculating the average value of F(s i ), i = 1, 2, …, N, the score of the joint future trajectory can be obtained.

[0116] In addition, in the embodiments of this application, the generated joint future trajectories can be sorted and filtered based on the scores of the joint future trajectories to further improve the authenticity of the generated trajectories.

[0117] There have been relevant numerical and simulation experiments to verify that, compared with existing agent trajectory generation methods, the method proposed in the embodiments of this application has achieved significant improvements in multi-agent trajectory generation results. In the Waymo SimAgents benchmark test, the proposed method has achieved leading performance in kinematic metrics (linear velocity, linear acceleration, angular velocity, angular acceleration, etc.), interaction metrics (collision, time to collision, etc.), map-based metrics (out of bounds, etc.), and overall realism metrics.

[0118] Based on the same principle as the method provided in the embodiments of the present application, the embodiments of the present application also provide a scenario-level multi-agent trajectory generation device based on consistent diffusion. The device includes a training module and a generation module;

[0119] As Figure 4 shown, the training module includes:

[0120] A trajectory enhancement module 401, configured to perform shifting and splicing processing on the real future trajectory sequence in the preset training set to obtain an enhanced trajectory sequence;

[0121] A noise addition module 402, configured to perform shifting and splicing processing on the noise sequence sampled from the Gaussian distribution to obtain an enhanced noise sequence, and add the enhanced noise sequence to the enhanced trajectory sequence according to the Gaussian transition to obtain an enhanced noisy trajectory sequence;

[0122] A denoising module 403, configured to use the scene encoder to extract the historical trajectory and context information of the agent in the training set to obtain the vector representation of the agent, and input the vector representation of the agent, the enhanced noisy trajectory sequence, and the diffusion step into the denoising network based on Transformer to gradually predict the added noise through the Gaussian state transition to obtain the future trajectory;

[0123] An optimization module 404, configured to optimize the scene encoder and the denoising network using a preset loss function;

[0124] As Figure 5 shown, the generation module includes:

[0125] A noise sequence acquisition module 501, configured to perform shifting and splicing processing on the noise sequence sampled from the Gaussian distribution to obtain an enhanced noisy trajectory sequence;

[0126] A trajectory prediction module 502, configured to use the optimized scene encoder to extract the historical trajectory and context information of the agent in the target historical scene to obtain the vector representation of the agent, and input the vector representation of the agent, the enhanced noisy trajectory sequence, and the diffusion step into the optimized denoising network to output the joint future trajectories of multiple agents;

[0127] Wherein, in the generation module, the denoising network ensures the consistency of the information of the same state in the enhanced noisy trajectory sequence through temporal consistency guidance during the gradual Gaussian state transition process.

[0128] In the embodiment of the present application, in the training module, the true future trajectory sequence in the preset training set is shifted and spliced to obtain an enhanced trajectory sequence, so that there is partial overlap between adjacent elements in the enhanced trajectory sequence. Then, the same method is used to shift and splice the noise sequence sampled from the Gaussian distribution to obtain an enhanced noise sequence, and the enhanced noise sequence is added to the enhanced trajectory sequence according to the Gaussian transition, so that the overlapping parts of adjacent elements in the enhanced trajectory sequence are added with the same noise. By constructing the enhanced trajectory sequence and the specific noise addition method above, the smoothness of the agent trajectory generated in the diffusion model can be improved. Then, the scene encoder is used to extract the historical trajectory and context information of the agent in the training set to obtain the vector representation of the agent, and the vector representation of the agent, the enhanced and noisy trajectory sequence, and the diffusion step are input into the denoising network based on Transformer to gradually predict the added noise through the Gaussian state transition to obtain the future trajectory; and the preset loss function is used to optimize the scene encoder and the denoising network. In the generation module, the noise sequence sampled from the Gaussian distribution is directly shifted and spliced to obtain an enhanced and noisy trajectory sequence, and then the vector representation of the agent extracted by the scene encoder, the enhanced and noisy trajectory sequence, and the diffusion step are input into the optimized denoising network, and the joint future trajectories of multiple agents can be output to ensure scene consistency. Among them, in the generation module, the denoising network ensures the consistency of the information of the same state in the enhanced and noisy trajectory sequence through temporal consistency guidance during the gradual Gaussian state transition, ensuring the temporal consistency during the denoising process, so as to improve the local smoothness of the generated joint future trajectory.

[0129] The multi-agent trajectory generation device at the scene level based on consistent diffusion provided by the embodiment of the present application can implement Figures 1 to 3 each process implemented in the method embodiment. To avoid repetition, it will not be elaborated here.

[0130] The multi-agent trajectory generation device at the scene level based on consistent diffusion in the embodiment of the present application can execute the multi-agent trajectory generation method at the scene level based on consistent diffusion provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module and unit in the multi-agent trajectory generation device at the scene level based on consistent diffusion in each embodiment of the present application correspond to the steps in the multi-agent trajectory generation method at the scene level based on consistent diffusion in each embodiment of the present application. For the detailed function description of each module of the multi-agent trajectory generation device at the scene level based on consistent diffusion, reference can be specifically made to the description in the corresponding multi-agent trajectory generation method at the scene level based on consistent diffusion shown above, which will not be elaborated here.

[0131] Based on the same principle as the method shown in the embodiments of the present application, the embodiments of the present application also provide an electronic device, which may include but is not limited to: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the method for generating multi-agent trajectories at the scenario level based on consistent diffusion shown in any optional embodiment of the present application by calling the computer program. Compared with the prior art, in the training stage of the method for generating multi-agent trajectories at the scenario level based on consistent diffusion provided by the present application, the true future trajectory sequence in the preset training set is shifted and spliced to obtain an enhanced trajectory sequence, so that there is partial overlap between adjacent elements in the enhanced trajectory sequence. Then, the same method is used to shift and splice the noise sequence sampled from the Gaussian distribution to obtain an enhanced noise sequence, and the enhanced noise sequence is added to the enhanced trajectory sequence according to the Gaussian transition, so that the overlapping parts of adjacent elements in the enhanced trajectory sequence are added with the same noise. By constructing the enhanced trajectory sequence and the specific noise addition method as described above, the smoothness of the agent trajectories generated in the diffusion model can be improved. Then, the scenario encoder is used to extract the historical trajectories and context information of the agents in the training set to obtain the vector representation of the agents, and the vector representation of the agents, the enhanced and noisy trajectory sequence, and the diffusion step are input into the denoising network based on Transformer to gradually predict the added noise through the Gaussian state transition to obtain the future trajectories; and the preset loss function is used to optimize the scenario encoder and the denoising network. In the generation stage, the noise sequence sampled from the Gaussian distribution is directly shifted and spliced to obtain an enhanced and noisy trajectory sequence, and then the vector representation of the agents extracted by the scenario encoder, the enhanced and noisy trajectory sequence, and the diffusion step are input into the optimized denoising network, and the joint future trajectories of multiple agents can be output to ensure scenario consistency. Among them, in the generation stage, the denoising network ensures the consistency of the information of the same state in the enhanced and noisy trajectory sequence through temporal consistency guidance during the gradual Gaussian state transition process, ensuring the temporal consistency during the denoising process, so as to improve the local smoothness of the generated joint future trajectories.

[0132] In an optional embodiment, an electronic device is further provided, as Figure 6 shown Figure 6 The electronic device 600 shown may be a server, including: a processor 601 and a memory 603. Among them, the processor 601 and the memory 603 are connected, such as connected through a bus 602. Optionally, the electronic device 600 may further include a transceiver 604. It should be noted that in practical applications, the transceiver 604 is not limited to one, and the structure of the electronic device 600 does not constitute a limitation to the embodiments of the present application.

[0133] The processor 601 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 601 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0134] The bus 602 may include a path for transmitting information between the above components. The bus 602 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 602 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0135] The memory 603 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0136] The memory 603 is used to store the application program code for executing the solution of this application, and is controlled by the processor 601 for execution. The processor 601 is used to execute the application program code stored in the memory 603 to implement the content shown in the foregoing method embodiments.

[0137] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.

[0138] The server provided by this application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.

[0139] The above description is only a preferred embodiment of this application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in this application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A method for generating scene-level multi-agent trajectories based on consistent diffusion, characterized in that, the method includes: In the training stage: Perform shifting and splicing operations on the real future trajectory sequences in the preset training set to obtain enhanced trajectory sequences; Perform shifting and splicing operations on the noise sequences sampled from the Gaussian distribution to obtain enhanced noise sequences, and add the enhanced noise sequences to the enhanced trajectory sequences according to Gaussian transitions to obtain enhanced noisy trajectory sequences; Use a scene encoder to extract the historical trajectories and context information of agents in the training set to obtain vector representations of the agents, and input the vector representations of the agents, the enhanced noisy trajectory sequences, and the diffusion steps into a denoising network based on Transformer to gradually predict the added noise through Gaussian state transitions to obtain future trajectories; Optimize the scene encoder and the denoising network using a preset loss function; In the generation stage: Perform shifting and splicing operations on the noise sequences sampled from the Gaussian distribution to obtain enhanced noisy trajectory sequences; Use the optimized scene encoder to extract the historical trajectories and context information of agents in the target historical scene to obtain vector representations of the agents, and input the vector representations of the agents, the enhanced noisy trajectory sequences, and the diffusion steps into the optimized denoising network to output the joint future trajectories of multiple agents; wherein, in the generation stage, the denoising network ensures the consistency of information in the same state in the enhanced noisy trajectory sequences through temporal consistency guidance during the gradual Gaussian state transition process.

2. The method for generating scene-level multi-agent trajectories based on consistent diffusion according to claim 1, characterized in that, after using the optimized scene encoder to extract the historical trajectories and context information of agents in the target historical scene to obtain vector representations of the agents, inputting the vector representations of the agents, the enhanced noisy trajectory sequences, and the diffusion steps into the optimized denoising network, and outputting the joint future trajectories of multiple agents, the method further includes: Calculating the joint future trajectories according to a scene scoring module to obtain the scores of the joint future trajectories.

3. The method for generating scene-level multi-agent trajectories based on consistent diffusion according to claim 1, characterized in that, the performing shifting and splicing operations on the real future trajectory sequences in the preset training set to obtain enhanced trajectory sequences includes: Shifting the elements at each position in the future trajectory sequence, and splicing the elements after shifting at each position with the elements before shifting to obtain the enhanced trajectory sequence.

4. The method for generating scene-level multi-agent trajectories based on consistent diffusion according to claim 1, characterized in that, the performing shifting and splicing operations on the noise sequences sampled from the Gaussian distribution to obtain enhanced noise sequences includes: Shifting the elements at each position in the noise sequence, and splicing the elements after shifting at each position with the elements before shifting to obtain the enhanced noise sequence.

5. The method for generating scene-level multi-agent trajectories based on consistent diffusion according to claim 1, characterized in that, Using the optimized scene encoder to extract the historical trajectory and context information of the agent in the target historical scene to obtain the vector representation of the agent, and inputting the vector representation of the agent, the enhanced noisy trajectory sequence, and the diffusion step into the optimized denoising network to output the joint future trajectories of multiple agents, including: Using the optimized scene encoder to extract the historical trajectory and context information of the agent in the target historical scene to obtain the vector representation of the agent; Fusing the vector representation of the agent, the enhanced noisy trajectory sequence, and the diffusion step to obtain a fused feature; Inputting the fused feature into the Transformer module to predict the joint future trajectories; wherein, the Transformer module includes alternating temporal self-attention mechanism and agent self-attention mechanism.

6. The method for generating scene-level multi-agent trajectories based on consistent diffusion according to claim 5, wherein, After inputting the fused feature into the Transformer module to predict the joint future trajectories, the method further includes: Regularizing the differences between adjacent position elements of each trajectory in the joint future trajectories.

7. The method for generating scene-level multi-agent trajectories based on consistent diffusion according to claim 1, wherein, The temporal consistency guidance makes the information of the same state in the enhanced noisy trajectory sequence consistent through average operation.

8. The method for generating scene-level multi-agent trajectories based on consistent diffusion according to claim 2, wherein, Calculating the score of the joint future trajectories according to the scene scoring module, including: Calculating the number of collisions of the agents by calculating the overlapping area between the agent bounding boxes in the joint future trajectories to obtain a safety verification metric value; Calculating the number of times the agents cross the road boundary in the joint future trajectories to obtain a road compliance metric value; Obtaining the score of the joint future trajectories according to the safety verification metric value and the road compliance metric value.

9. A device for generating scene-level multi-agent trajectories based on consistent diffusion, wherein, The device includes a training module and a generation module; The training module includes: A trajectory enhancement module for performing shifting and splicing processing on the real future trajectory sequence in the preset training set to obtain an enhanced trajectory sequence; A noise addition module for performing shifting and splicing processing on the noise sequence sampled from the Gaussian distribution to obtain an enhanced noise sequence, and adding the enhanced noise sequence to the enhanced trajectory sequence according to the Gaussian transition to obtain an enhanced noisy trajectory sequence; A denoising module for using the scene encoder to extract the historical trajectory and context information of the agents in the training set to obtain the vector representation of the agents, and inputting the vector representation of the agents, the enhanced noisy trajectory sequence, and the diffusion step into the denoising network based on Transformer to gradually predict the added noise through Gaussian state transition to obtain the future trajectories; An optimization module for optimizing the scenario encoder and the denoising network by using a preset loss function; The generation module includes: A noise sequence acquisition module for performing shifting and splicing processing on a noise sequence sampled from a Gaussian distribution to obtain an enhanced noisy trajectory sequence; A trajectory prediction module for using the optimized scenario encoder to extract the historical trajectory and context information of the agent in the target historical scenario to obtain a vector representation of the agent, and inputting the vector representation of the agent, the enhanced noisy trajectory sequence, and the diffusion step into the optimized denoising network to output the joint future trajectories of multiple agents; Wherein, in the generation module, the denoising network ensures the consistency of information of the same state in the enhanced noisy trajectory sequence through temporal consistency guidance during the step-by-step Gaussian state transition.

10. An electronic device, Characterized in that It includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-agent collaborative exploration method and device based on low-order Gaussian distribution

    CN112215333A

  • Cooperative hunting method based on multi-agent generative adversarial imitation safety learning

    CN113723012A