Automatic driving test scene simulation generalization generation method and system based on knowledge distillation

By combining knowledge distillation and diffusion models, a multi-agent trajectory generation framework is constructed, which solves the cross-scenario generalization problem in the generation of autonomous driving simulation scenarios in existing technologies, realizes more realistic and stable traffic simulation scenario generation, and improves the ability to capture multimodal features and handle uncertainties of traffic flow.

CN121744955AActive Publication Date: 2026-03-27TONGJI UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-02
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing methods for generating autonomous driving simulation scenarios lack the ability to generalize across datasets and scenarios, making it difficult to realistically reflect multi-agent interactions and emergencies in complex traffic environments. Furthermore, they are insufficient in terms of multimodal features and uncertainties, which limits the stability and application value of the generated scenarios.

Method used

We adopt a method that combines knowledge distillation and diffusion models. By constructing a multi-agent trajectory generation framework, we utilize the self-distillation mechanism of the teacher-student network and combine it with simulated annealing strategy to dynamically balance task loss and distillation loss, thereby improving the model's cross-scene adaptability and generation stability and achieving high-fidelity scene synthesis.

Benefits of technology

It significantly improves traffic rule compliance and interaction coordination, reduces collision and wrong-way traffic rates, enhances the model's generalization ability across datasets and scenarios, and generates agent trajectories that are more consistent with real traffic patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121744955A_ABST
    Figure CN121744955A_ABST
Patent Text Reader

Abstract

The invention relates to an automatic driving test scene simulation generalization generation method and system based on knowledge distillation. The method comprises the following steps: acquiring real trajectory data, constructing a scene-level multi-agent trajectory generation framework based on a denoising diffusion probability model, and generating diversified and vivid agent trajectories through an iterative denoising process so as to realize modeling of multi-agent interaction behaviors in a complex traffic environment; a mixed knowledge self-distillation framework is further proposed, multi-level knowledge of a teacher model is migrated in a student model, and task loss and distillation loss are dynamically balanced in combination with a simulated annealing strategy, so that the cross-scene adaptability and generation stability of the model are improved; simulation verification shows that the generated intelligent agent track can effectively reduce the collision event rate and the retrograde event rate, and interaction coordination and traffic rule constraints are ensured. According to the system, the authenticity, diversity and cross-scene adaptive capacity of an automatic driving simulation scene can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer software technology and intelligent transportation, and in particular to a method and system for generalizing the generation of autonomous driving test scenario simulations based on a knowledge distillation and diffusion model. Background Technology

[0002] With the rapid development of autonomous driving technology, virtual simulation testing has gradually become a core means of verifying system safety and reliability. Compared with real road testing, simulation testing can reproduce complex traffic environments in a low-cost, highly controllable, and highly repeatable manner, avoiding the safety hazards and uncertainties present on real roads. Therefore, it plays an irreplaceable role in the research and development and implementation of next-generation intelligent transportation systems. By constructing diverse and realistic simulation scenarios, researchers and engineers can systematically test and optimize the perception, decision-making, and control modules of autonomous driving systems in a virtual environment.

[0003] Existing scene generation methods mainly fall into two categories: rule-driven and data-driven. Rule-driven methods typically rely on manually set traffic rules and vehicle dynamics models, which can better ensure the interpretability and parameter controllability of the model, and were widely used in early traffic simulation platforms (such as VISSIM and SUMO). However, these methods rely too heavily on fixed parameters and prior assumptions, making it difficult to realistically reflect multi-agent interactions and unexpected situations in complex traffic environments. The generated scenes also suffer from significant deficiencies in terms of diversity and cross-scene adaptability. In contrast, data-driven methods utilize large-scale traffic trajectory data and deep learning models, enabling them to automatically learn the complex spatiotemporal dependencies between traffic participants, exhibiting greater flexibility and scalability in trajectory generation and interaction modeling. In recent years, diffusion models have shown significant advantages in high-fidelity and diverse generation, and are considered an important direction for promoting the development of autonomous driving simulation.

[0004] However, while data-driven methods demonstrate significant potential, several technical bottlenecks remain to be addressed. First, existing methods generally rely on single datasets for training, lacking generalization capabilities across datasets and scenarios. When applied to unfamiliar traffic environments, this can easily lead to unreasonable generated trajectories or distorted interactions. Second, in multi-agent interaction modeling, current research largely focuses on maximizing the likelihood of a behavior or adversarial scenarios as optimization objectives, neglecting the multimodal characteristics and uncertainties of actual driving behavior. This results in generated scenarios that fail to fully reflect the complexity and diversity of real traffic flow. Finally, in closed-loop simulations, data-driven methods often suffer from accumulated trajectory deviations and high collision rates, impacting system stability and application value.

[0005] In recent years, knowledge distillation technology has been widely used in autonomous driving-related tasks such as perception, prediction, and planning. By transferring deep knowledge from the teacher model to the student model, it can not only effectively alleviate the overfitting of the student model to the training data, but also improve the robustness and adaptability of the model in unknown environments. This provides a new approach to solving the lack of generalization in existing simulation generation methods. However, in current technologies, the two are always applied independently, failing to form a synergistic optimization effect. Currently, no research has deeply integrated knowledge distillation and diffusion generation models to improve the realism, diversity, and cross-scene adaptability of simulation scenarios. Therefore, how to construct a unified framework that combines the advantages of knowledge distillation and data-driven generation to achieve efficient, realistic, and generalizable autonomous driving simulation scenario generation has become a key problem urgently needing to be solved in this field. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and system for generalizing the generation of autonomous driving test scenarios based on a combination of knowledge distillation and diffusion models, which can improve cross-scenario adaptability and robustness while ensuring realism and diversity.

[0007] The objective of this invention can be achieved through the following technical solutions: One aspect of the present invention provides a method for generalizing and generating simulation of autonomous driving test scenarios based on knowledge distillation, comprising the following steps: Collect real-world autonomous driving datasets (such as the Waymo dataset), including trajectory information of various traffic participants such as motor vehicles, non-motor vehicles, and pedestrians, establish a unified observation-action feature matrix, and construct traffic simulation scenarios to support multi-agent trajectory generation and behavior modeling. Using each traffic participant as an intelligent agent, a multi-agent trajectory generation framework is constructed based on a diffusion model. The trajectory data is gradually subjected to noise perturbation and reverse denoising recovery to achieve high-fidelity scene synthesis. A multi-agent trajectory generation framework is constructed based on the diffusion model. The trajectory of the agent is generated through an iterative denoising process to model the interactive behavior of multiple agents in complex traffic environments. A self-distillation mechanism for the teacher-student network is proposed to transfer the structured knowledge of the teacher model in the source dataset (such as the Waymo dataset) to the student model. The simulated annealing strategy is combined to dynamically balance the task loss and distillation loss, thereby improving the cross-scene adaptability and generation stability of the model. Zero-shot closed-loop simulation validation was performed on other independent autonomous driving datasets (such as the INTERACTION dataset) that were not used in the training.

[0008] The generated agent trajectory can effectively reduce the collision event rate and the wrong-way event rate, significantly improve traffic rule compliance and interaction coordination, and verify the generalization ability of the method under cross-dataset and cross-scenario conditions.

[0009] As a preferred technical solution, each traffic participant is an intelligent agent, and the traffic simulation scenario is formally defined as follows: in, express An intelligent agent in A set of trajectories at each time step. The dimension of the state vector for each agent; Corresponding to the The trajectory sequence of each agent, and the state at each time step. Including location ,speed (respectively) and (velocity component in direction), heading angle and border dimensions ( These represent the length, width, and height attributes respectively; while For the joint control sequence driving trajectory evolution, where To control the dimension of the input vector; Corresponding to the The control sequence of an agent, with its control input at each time step. Includes acceleration (Acceleration along the vehicle's longitudinal axis) and rate of change of heading angle This determines the agent's movement behavior.

[0010] Meanwhile, traffic simulation scenario context include: Road topology ,for A collection of multiple lane segments, each lane consisting of... It consists of 1 key points, each of which is 1 key point. Two-dimensional coordinates describe the lane centerline; Traffic signal status ,include The status of each traffic signal, among which The dimension representing each signal state, typically These correspond to the encoding of the signal color (red, yellow, green) and the remaining time sequence information of the current phase, respectively. Historical status ,for An agent in the past The trajectory state of the time step provides the interaction history state; among which The dimension of the trajectory state vector includes position, velocity, heading angle, and size attributes. In this embodiment... ; Initial joint state .

[0011] The generation and evolution of traffic simulation scenarios follow a discrete-time dynamic model: In the formula, It is a joint dynamic function, which is based on a particle kinematics model and controls the input. Achieve joint trajectory state of each intelligent agent Forward update. Specifically, for the first... Individual agents ( ), in the Time step ( ) state From the state of the previous moment With control input Decide: in For location, for and The velocity component in the direction, For heading angle, and These are the corresponding control inputs. The longitudinal acceleration (along the current heading direction) and the rate of change of the heading angle in the middle. For time step.

[0012] The joint dynamic function integrates the independent motion updates of all agents and can incorporate interaction terms into the modeling, thereby enabling the tracking of the joint trajectory state of multiple agents. Update.

[0013] As a preferred technical solution, based on the aforementioned definition of traffic simulation scenarios, the task of generating traffic simulation scenarios is formalized into the following trajectory optimization problem: , in, For parameterized objective functions, measure the control sequence and corresponding trajectory In context The rationality, For the model parameters, the optimization objective is to minimize Generate control sequences that conform to real traffic behavior and environmental constraints. .

[0014] It should be noted that this invention uses a denoised diffusion probability model (DDPM) for trajectory optimization training, minimizing... It is an "implicit" rule learned from large-scale real traffic data in a data-driven manner.

[0015] Furthermore, the multi-agent trajectory generation framework includes: a scene encoder module, a trajectory generation framework, and a behavior predictor.

[0016] As a preferred technical solution, a scene encoder module based on a query-controlled multi-head attention (QCMHA) architecture is used to capture the interaction relationships between multiple agents and traffic scene constraints. The process is as follows: Taking road topology, traffic signal states, and multi-agent historical states as inputs, global information is mapped to a local coordinate system centered on the target agent through coordinate system normalization, enhancing relative spatial modeling capabilities. Subsequently, a gated recurrent unit (GRU) is used to extract temporal features, and lane-level geometric features are obtained by combining a multilayer perceptron (MLP) and max pooling, which are then embedded with the signal states. After unifying the feature dimensions through linear projection, the data is input into a Transformer encoder, which models the dynamic interaction between multi-agents and the environment through self-attention and cross-attention mechanisms, generating a high-dimensional latent representation of the scene. A history dropout strategy is introduced during training to randomly discard some historical information, strengthening the model's dependence on current semantics, thereby improving the generalization and causal consistency of closed-loop simulation and providing robust contextual support for trajectory generation.

[0017] As a preferred technical solution, a trajectory generation framework based on the Denoising Diffusion Probability Model (DDPM) is adopted, in which the forward diffusion process and the reverse denoising process together constitute the core mechanism for multi-agent control sequence generation, as follows: In the forward diffusion phase, the model gradually injects Gaussian noise into the real control sequence through a Markov chain, making the distribution approximate a standard Gaussian. A log noise schedule is employed to balance the signal-to-noise ratio, preserving structural information in the early stages and adding noise later to improve generalization ability. In the inverse denoising phase, the model starts from noisy samples, combining scene encoding and diffusion step information, and gradually recovers the control sequence that conforms to physical and traffic constraints through a Transformer denoising network. A smoothed L1 loss is used to optimize the generation continuity. To achieve controllable generation, a guided generation mechanism is introduced. This mechanism dynamically adjusts the denoising mean through constraints such as speed limits, collision avoidance, and lane keeping, achieving task-oriented and rule-compliant trajectory generation that balances realism and controllability.

[0018] The behavior predictor, while maintaining the stability of independent optimization of the noise reduction generator and the scene encoder, achieves efficient modeling of the agent's multimodal motion patterns, as detailed below: This module takes scene encoding and a set of multi-endpoint anchors as input, and predicts multiple possible control sequences and their corresponding modal probability weights for each agent at future time through a lightweight multilayer perceptron network.

[0019] Here, "multimodal" refers to the various discrete behavioral intentions that an agent may take in a given scenario (e.g., go straight, turn left, turn right, stop and wait), with each modality corresponding to a complete control sequence and a candidate endpoint. "Multi-endpoint anchor points" refer to a set of discrete, representative candidate endpoint coordinates obtained by clustering endpoint positions in real trajectory datasets or by sampling from road topology, used to provide structured intent priors for behavior prediction.

[0020] The model weights the outputs of each modality in a hybrid distribution, with the modal weights reflecting their behavioral probabilities in a specific scenario. By introducing this explicit behavioral prior constraint diffusion process, the noise reduction generator is prevented from over-relying on noise structure and neglecting semantic features, thereby improving the diversity, interpretability, and scenario consistency of trajectory generation. This design enables the diffusion generation model to simultaneously learn diverse control strategies and physically reasonable dynamic features during the generation process, significantly enhancing its generalization and stability in complex traffic scenarios.

[0021] Furthermore, the self-distillation mechanism of the teacher-student network includes feature self-distillation, response self-distillation, and relation self-distillation.

[0022] As a preferred technical solution, the Feature Self-Distillation (FSD) mechanism is used to align the feature representations of the teacher model and the student model in the intermediate layer of the scene encoder, as follows: The scene encoder generates feature maps based on QCMHA. The teacher model originates from a diffusion-generative model pre-trained on diverse traffic data using the aforementioned modules, possessing rich structured semantics. The student model has the same structure as the teacher model, with parameters randomly initialized. FSD aligns the feature distributions of both models, enabling the student model to inherit the teacher model's ability to represent complex scenes in a structured manner. Features are L2 normalized before computation to focus on semantic consistency rather than numerical differences, thereby effectively transferring knowledge from the teacher model, mitigating overfitting, and enhancing the stability and rule compliance of the student model in unseen scenarios.

[0023] As a preferred technical solution, this invention proposes a Response Self-Distillation (RSD) mechanism to align the action distributions of the teacher model and the student model during the denoising generation process, as follows: The described Response Self-Distillation (RSD) mechanism uses the denoised actions output by the teacher model as a soft objective. By minimizing the difference in probability distributions between the two models, it guides the student model to learn more robust control strategies. A temperature parameter is introduced to adjust the smoothness of the distribution; higher temperatures highlight the relative relationships between actions, while lower temperatures strengthen the learning of key actions. Through distribution alignment based on KL divergence, the student model can inherit the behavioral priors and rule constraints of the teacher model, generating smoother, more realistic, and physically consistent trajectories in complex scenarios such as intersections and roundabouts, thereby significantly improving generation consistency and generalization ability.

[0024] The relational self-distillation mechanism (RlSD) is used to align the knowledge representations of the teacher model and the student model in multi-agent interaction relationship modeling, as follows: The Relational Self-Distillation (RlSD) mechanism constructs an agent interaction graph, using action differences as edge weights to describe mutual influences, and calculates the interaction differences between the two models during the denoising process. Cosine similarity loss is used to align the action direction consistency between agent pairs, thereby enhancing the learning of interaction structures. This method can capture relatively coordinated patterns such as following, avoidance, and cooperative lane changing, enabling the student model to inherit the teacher model's structured understanding of complex interactions. This improves the accuracy and diversity of trajectory generation in high-density, multi-agent scenarios, providing more realistic and generalizable social interaction modeling capabilities for autonomous driving simulation.

[0025] As a preferred technical solution, the following are also included: Trajectory accuracy assessment: The deviation between the generated trajectory and the real trajectory is quantified from two aspects: long-term prediction and multimodal trajectory generation, to evaluate the model's accuracy and spatiotemporal consistency in trajectory generation.

[0026] Closed-loop interaction authenticity assessment: The safety and robustness of the generated trajectory are assessed from multiple aspects, including driving behavior safety, traffic rule compliance and authenticity.

[0027] Another aspect of the present invention provides a knowledge distillation-based simulation generalization generation system for autonomous driving test scenarios, used to implement the aforementioned data-driven simulation method for complex interactive behavior strategies in hybrid traffic flows. The simulation system includes a model building module and an application simulation module. The model building module includes: The data acquisition and preprocessing unit is used to collect multi-source traffic flow data, including trajectory information of motor vehicles, non-motor vehicles and pedestrians, and to perform spatiotemporal alignment, feature extraction and normalization processing. Multi-agent modeling unit to realize multi-agent trajectory generation framework: Construct agent control strategy based on diffusion generation framework to realize interactive modeling of multiple types of traffic participants; The knowledge self-distillation generalization optimization unit realizes the self-distillation mechanism of the teacher-student network: by aligning the feature distribution of the teacher and student models through feature distillation and relation distillation, the generalization performance and stability of the model in complex traffic scenarios are improved.

[0028] The application simulation module includes: The scene configuration unit is used to generate an initial simulation scene with road topology, signal control, and agent distribution information. The behavior prediction and scene generation unit generates multimodal and controllable agent behavior trajectories based on the trained model, achieving highly realistic reproduction of the dynamic evolution process of traffic flow.

[0029] Compared with the prior art, the present invention has the following beneficial effects: (1) Achieve unified modeling and interactive generation of multiple types of traffic participants, significantly improving the consistency and realism of mixed traffic flow simulation; (2) Introducing a diffusion generation mechanism and using the probability generation process of gradually denoising noise to reconstruct the trajectory distribution of multiple agents can effectively capture the multimodal characteristics and uncertainties of traffic flow and generate dynamic scenes that are more in line with real traffic patterns. (3) Through various knowledge self-distillation mechanisms, the internal features and policy distribution of the model are adaptively aligned, thereby improving the stability of cross-scene migration and complex interactions. Attached Figure Description

[0030] Figure 1 This is a schematic diagram of the overall framework of the method of the present invention; Figure 2 This is a flowchart of the method for generalizing and generating autonomous driving test simulation scenarios in an embodiment of the present invention; Figure 3 Comparison of closed-loop simulation scenario generation for typical scenarios in embodiments of the present invention Figure 1 ; Figure 4 Comparison of closed-loop simulation scenario generation for typical scenarios in embodiments of the present invention Figure 2 ; Figure 5 Comparison of closed-loop simulation scenario generation for typical scenarios in embodiments of the present invention Figure 3 ; Figure 6 This invention provides a comparison chart of closed-loop simulation generalization across datasets for zero-shot scenarios in typical scenarios in this embodiment. Figure 7 This is a schematic diagram of the knowledge distillation-based autonomous driving test scenario simulation generalization generation system according to an embodiment of the present invention; Figure 8 This is a schematic diagram illustrating the implementation process of the knowledge distillation-based autonomous driving test scenario simulation generalization generation system according to an embodiment of the present invention. Detailed Implementation

[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0032] Example 1 To address the problems existing in the prior art, this embodiment provides a generalized generation method for autonomous driving test scenario simulation based on knowledge distillation, and develops a high-precision simulation model HySD for the dynamic expectations of multiple types of traffic participants in complex interactive environments. To comprehensively improve the model's generalization ability, this application introduces a two-stage modeling mechanism based on diffusion generation and knowledge self-distillation, introduces a behavior predictor and diffusion generation mechanism to enhance the diversity and realism of multimodal trajectories, and combines a closed-loop simulation evaluation system to comprehensively verify the model's generalization, stability, and physical feasibility.

[0033] Figure 1 This paper presents the overall framework of a knowledge distillation-based generalization generation method for autonomous driving test scenario simulation, using multi-agent simulation scenario generation in urban road traffic as an example. Trained on real-world interaction trajectory data, the model generates two main outputs: a joint control sequence for multiple types of traffic participants and a multimodal trajectory distribution. The control sequence, through a denoising process, recovers actions that conform to scenario constraints, reflecting the dynamic decision-making tendencies of traffic participants in a specific context, as well as specific trajectory information. The multimodal trajectory distribution generates a probability-weighted trajectory set and comprehensively considers the interaction relationships between agents, thereby simulating the trajectory generation and generalization optimization of different types of traffic participants in complex traffic flow interaction environments, obtaining the time-varying trajectories of multiple types of traffic participants in complex interaction environments.

[0034] Figure 2 This is a flowchart illustrating the generalization generation method for autonomous driving test simulation scenarios in this embodiment. The method mainly includes input, training, output, and simulation phases. The specific processes of each phase are as follows: Step S1, Model Encoding Input Stage. Acquire real-world road topology, traffic signal status, and historical trajectory data of the agents. Analyze the multi-agent interaction behavior to clarify the state and context of each time step, where each agent corresponds to a traffic participant. This step specifically includes steps S101-S103.

[0035] Step S101, coordinate system normalization. Using the last recorded agent state as a reference point, the historical trajectory is normalized to the local coordinate system through translation and rotation operations to ensure data consistency.

[0036] In this embodiment, coordinate system normalization is centered on the last observation position of the agent in the urban road traffic scene. A rotation matrix is ​​used to align all historical trajectories to the positive x-axis, eliminating global coordinate differences. After processing, the temporal resolution of the trajectory data is 0.1 s. First, spatial matching is performed on the trajectory data and road topology to filter out the agent trajector trajector trajector trajector trajectories of interest. Then, combined with traffic signal data, the trajectories of multiple agents are time-aligned according to the same time step, and trajectory data within the effective interaction period is extracted to ensure that the generated behavior occurs within the constraints. Considering that there is not interaction at every moment in the complete sequence, the scene is reasonably split through manual intervention to reduce the proportion of invalid noise while ensuring the integrity of the trajectory. Finally, to ensure the continuity and analyzability of the trajectory, the total duration of the scene is further filtered, retaining only scenes with an interaction time ≥ 8 s to improve data quality. Hundreds of effective scenes are ultimately extracted, and data augmentation techniques are used to expand the dataset. Most of these are used for subsequent model training, and the remainder for inference and validation. All scenes are used together for generalization analysis. It should be noted that this dataset is only a typical example, and other autonomous driving datasets can also be easily applied.

[0037] Step S102, Initial Feature Extraction. A GRU network is used to temporally encode the agent's historical states, and an MLP is used to statically encode the map topology and traffic signals to extract preliminary semantic features.

[0038] In this embodiment, the GRU algorithm, which does not require pre-specifying the dimension, is employed. This method, based on the principle of sequence dependency, can effectively capture temporal dynamics and accurately distinguish noise. The hidden layer dimension and activation function are determined based on domain knowledge. Then, all input sequences are traversed, and sequences containing sufficient temporal information are marked as core features. Next, adjacent sequences are recursively merged to form encoding vectors, and boundary features are assigned to the corresponding vectors. Remaining unclassified features are marked as noise.

[0039] In this embodiment, the average length of historical trajectories is selected as the input dimension of the GRU. The hidden layer for vehicle agents is set to 128, and for non-vehicle agents, it is set to 64. The activation function is set to ReLU (meaning at least multiple time steps are required to form an effective encoding). During feature extraction, traffic element types are distinguished: road topology is fused only with static elements, and signal state is fused only with dynamic elements. Furthermore, feature extraction is performed at each time step based on state data. The parameter settings of this method have clear physical meaning and can dynamically adapt to state changes in different scenarios.

[0040] Step S103, Unified Feature Encoding. The scene encoder is based on the Transformer encoder, and each layer integrates a query-centric multi-head attention mechanism (QCMHA) and a cross-attention module to fuse road, signal, and historical states into a high-dimensional latent representation. It provides rich scene context information.

[0041] In this embodiment, the observation space consists of scene topology and the encoded information of other elements. The scene topology includes the road centerline, lane boundaries, signal status, trajectory of the previous moment, speed of the previous moment, and agent type (if the agent is a vehicle, the type is a category code; if the agent belongs to a multimodal group, the observation information of other elements includes the relative position with the interaction object, speed difference, and interaction object type); the encoding space is a high-dimensional vector (without specific dimensions) and attention weights.

[0042] Step S2 involves training and generating a teacher model based on a diffusion simulation scenario. This includes both the forward diffusion process and the reverse denoising process.

[0043] Step S201, Forward Diffusion Process. The goal of the forward diffusion process is to construct a Markov chain that gradually transforms the real control sequence into a noise sequence that approximates a Gaussian distribution. According to the theory of DDPM (Denoising Diffusion Probability Model), the diffusion process can optimize the log-likelihood of the data through variational lower bounds. The probability distribution of forward diffusion is defined as a Gaussian distribution (denoised by the symbol...). express): Among them, symbols It is defined as; Indicates the first Noise control sequence after step diffusion; For the first The noise scale of the step, This represents the total number of diffusion steps. Using a reparameterization method, any number of diffusion steps... The sampling formula is: in, ,in , is the cumulative noise attenuation coefficient; To obtain from the standard Gaussian distribution The noise vector sampled in the middle, It is an identity matrix.

[0044] This embodiment employs a log noise schedule strategy, defining a Gaussian distribution with probabilistic form to ensure low noise in early steps to preserve data structure, while later steps are dominated by noise to approximate a standard Gaussian distribution. Log noise scheduling diffuses the number of steps. Based on, through The parameters increase linearly from 1e-4 to 0.02, ensuring a balanced signal-to-noise ratio (SNR). The specific formula for calculating logarithmic noise scheduling is as follows: In the formula, The scheduling parameters control the noise growth rate. The sigmoid form of logarithmic noise scheduling ensures that the early steps ( The noise is relatively low, and the data structure is preserved; the number of steps in the later stage ( Noise dominates, making The distribution is close to the standard Gaussian distribution. All sequences are used together for optimization analysis. It should be noted that this strategy is only a typical example; other scheduling methods can also be applied.

[0045] Step S202, inverse denoising process. The control sequence conforming to scene constraints is iteratively recovered from Gaussian noise and incorporated into the dynamic model and controllable guidance. This process consists of two main modules: a denoising generator and a decoder. The denoising generator is responsible for fusing scene information and noise state, outputting intermediate features; the decoder then maps these features to the final control sequence prediction.

[0046] In this embodiment, the noise reduction generator is based on the Transformer architecture and consists of four stacked Transformer decoder blocks. Its inputs include: 1) contextual features output by the scene encoder. ;2) Noise control sequence for the current diffusion step 3) Number of diffusion steps The embedding representation is used. The noise reduction generator fuses scene semantics and the current noise state through self-attention and encoder-decoder cross-attention mechanisms, and outputs a high-dimensional intermediate feature.

[0047] This feature is then fed into a decoder, which is a three-layer multilayer perceptron (MLP) whose last layer linearly maps the output to the true control sequence. The prediction.

[0048] The noise reduction generator is designed to allow the model to progressively remove noise and recover a control sequence that conforms to physics and traffic rules. Backdiffusion is defined as a conditional Markov chain. in, The parameter is The conditional probability distribution; This is a noise-reducing mean function; For the first The variance of the step is usually set to Or it may be determined based on the scheduling; This is the scene context feature representation output by the scene encoder.

[0049] The optimization objective of the noise reduction generator is to minimize the smooth L1 loss between the predicted control sequence and the true control sequence. (This is indicated to ensure the fidelity and stability of the generated sequence): in, Represents the mathematical expectation operator. Represents the loss function. These are the model parameters.

[0050] To further improve the performance of the reverse denoising process, a controllable guidance mechanism is introduced in the inference phase, which adjusts the denoising mean through additional optimization objectives (such as speed limiting, collision avoidance, etc.). .

[0051] Specifically, in this embodiment, the behavior predictor takes scene encoding and noise samples at the current moment as input, and generates a prediction distribution containing multimodal trajectories and their probabilities for each agent through a lightweight spatiotemporal attention network. Its optimization objective employs a multi-task loss function, aiming to simultaneously ensure the accuracy of trajectory prediction and the reliability of modality prediction. The optimization objective of the behavior predictor employs a hybrid supervision strategy: on the one hand, based on the predicted multimodal results, the optimal mode is selected and its smoothing L1 loss between its trajectory and the true endpoint is minimized to ensure the accuracy of the target estimation; on the other hand, the cross-entropy loss of the mode prediction is introduced to encourage the target prediction distribution to cover more possible behavioral final states. The optimization objective can be expressed as: Among them, symbols Indicates the self-distribution of the samples Scene Find the expected value; This represents the smoothed L1 loss function; Represents the cross-entropy loss function; Represents intelligent agents In the optimal mode The predicted trajectory below, The predicted modal probabilities; To balance the hyperparameters of regression and classification task weights, this embodiment takes... .

[0052] Step S203, Optimization of Teacher Model Training Loss. To achieve collaborative optimization of the denoising generator and the behavior predictor, the teacher model adopts a joint training strategy, comprehensively considering both denoising loss and behavior prediction loss. The total training loss function is defined as: In the formula, Indicates the total training loss; These are hyperparameters used to balance the contributions of the denoising and behavior prediction tasks. In this embodiment, they are selected through a grid search. Through joint optimization, the model can simultaneously learn realistic control sequence generation and diverse behavioral patterns, thereby generating diverse and high-quality traffic scenarios in closed-loop simulations.

[0053] Step S3: Self-distillation generalization optimization of the student model. Specifically, this step includes steps S301-S302.

[0054] Step S301, Teacher Model Knowledge Sampling Stage. Utilizing the global generalization ability gained by the teacher model during source domain training, high-confidence behavioral samples and mid-level semantic features are generated, enabling the student model to extract general patterns from multimodal behavioral distributions. Specifically, the teacher model outputs three pieces of knowledge: the intermediate feature map of the scene encoder. (Used for feature self-distillation, aligning semantic representations to improve robust representation), action probability distribution during denoising. (Used to respond to self-distillation, aligning softened behavior generation strategies to enhance consistency), action difference vectors of agent pairs. (Used for relational self-distillation, aligning interaction relationships to capture dynamic coordination patterns).

[0055] In this embodiment, the feature map of the teacher model The output of the scene encoder feature map is based on a query-controlled multi-head attention mechanism. The feature map generation process can be formalized as follows: in, Indicates the embedding layer. Indicates the first Output of the QCMHA module. This refers to the set of fixed parameters that have been trained in the teacher model.

[0056] Denoising action sequences for the teacher model Through noise reduction generator The generated continuous vector is first converted into a probability distribution through normalization. To enhance the stability of RSD, a temperature parameter is introduced. The output distribution is softened to smooth the probability distribution of the teacher model, thereby highlighting the relative relationships between actions. The formula for calculating the softened probability distribution is: Temperature parameters It is a hyperparameter greater than 0, used to control the smoothness of the distribution. The larger the value, the more uniform (smooth) the distribution; the smaller the value, the sharper the distribution (approaching one-hot encoding). In this embodiment, after verification through grid search, the value is set... This value can provide sufficient smooth gradients for conveying the knowledge of the teacher model while preserving the discriminative information of the teacher model.

[0057] Step S302, Student Model Self-Distillation Stage. The student model achieves knowledge transfer by aligning the feature representations, denoised action distributions, and agent interaction relationships of the teacher model.

[0058] Specifically, feature self-distillation minimizes the mean squared error loss after L2 norm normalization to align the scene encoder feature map; response self-distillation uses KL divergence to align the softened denoised action probability distribution; relation self-distillation uses cosine similarity loss to align the action difference vectors of agent pairs, thereby improving the robustness of the student model in feature representation, behavior generation, and interaction modeling. The specific formulas are as follows: in, These represent the losses from feature self-distillation, response self-distillation, and relation self-distillation, respectively. Feature maps for student models; The action probability distribution output by the student model; The intelligent agent output for the student model and Action difference vectors between them; Indicates the KL divergence; For agents logarithms This represents the Euclidean norm.

[0059] Step S303, Student Model Self-Distillation Training Phase. Training is performed by iteratively optimizing the total loss function, which includes denoising generation loss, behavior prediction loss, and weighted distillation loss (FSD, RSD, R1SD). A simulated annealing strategy is introduced to dynamically adjust the distillation weights. In the early stages of training, emphasis is placed on distillation loss to inherit teacher knowledge, while in the later stages, emphasis is placed on task loss to develop independent generation capabilities, ensuring that the student model generates realistic and rule-compliant trajectories in unseen scenarios. The specific formula is as follows: Simulated annealing weights Its update formula is as follows: (This formula decays over time.) In the formula, The loss of the denoising generator for the student model. For the loss of the behavior predictor, The weight parameters for each loss term are used to balance the contributions of different objectives. In this embodiment, they are set to 1.0, 0.5, 0.5, and 0.5 respectively through grid search. The initial weight for simulating annealing is set to 0.9; This is the current training round number; The annealing cycle is set to 50 rounds. Task loss mainly refers to the error in the model's completion of core tasks such as trajectory generation and behavior prediction, while distillation loss represents the alignment error in knowledge transfer from the teacher model.

[0060] Step S4: Multi-agent scenario generation, simulation, and evaluation. Specifically, this step includes steps S401-S404.

[0061] Step S401, Trajectory Accuracy Assessment. The similarity between the generated trajectory and the real expert trajectory is assessed based on the micro-trajectory accuracy.

[0062] In this embodiment, the minimum average distance error of the location is used ( ) and minimum final displacement error ( To quantify precision. Used to evaluate the average deviation between the generated trajectory and the true trajectory over the entire time span; Emphasizing the predictive ability of the trajectory endpoint is particularly important for long-term prediction tasks. This indicator calculates the positional deviation between the generated trajectory and the actual trajectory at each moment. The calculation method for a single scenario is as follows: in, ( t )and Representing intelligent agents respectively In modality The generated trajectory and the true trajectory in time The position vector; The total number of agents in the scene; The number of modes used to generate the trajectory; This represents the total time step of the trajectory; This represents the Euclidean norm.

[0063] The HySD model and other benchmark models were evaluated on the Waymo Sim Agents autonomous driving dataset, and the specific results are shown in Table 1. The benchmark models compared include: 1) Baseline teacher model: a raw diffusion generation model without any distillation; 2) FSD model: a student model using only feature self-distillation; 3) RSD model: a student model using only response self-distillation; and 4) RlSD model: a student model using only relation self-distillation. Overall, the HySD model significantly outperforms other benchmark models in multi-agent trajectory generation accuracy, fully demonstrating that the model can reproduce the realistic characteristics of mixed traffic flow in traffic scenarios.

[0064] Table 1 Comparison results between HySD model and benchmark model Step S402, Closed-loop simulation realism assessment. The safety and robustness of the generated trajectory are evaluated from multiple aspects, including driving behavior safety, traffic rule compliance, and realism.

[0065] In this embodiment, collision event rate, road deviation event rate, wrong-way driving event rate, kinematic infeasibility, and trajectory similarity indicators are used to quantify realism. Specifically, the collision event rate is the proportion of collisions between agents in the generated trajectory, used to assess the safety of driving behavior; the road deviation event rate is the proportion of events where agents deviate from lanes or road boundaries, reflecting compliance with traffic rules; the wrong-way driving event rate is the proportion of events that violate traffic direction (e.g., driving in the wrong direction), further examining rule compliance; kinematic infeasibility assesses whether the generated trajectory violates physical constraints (e.g., speed, steering angle, or acceleration exceeding vehicle capabilities); trajectory similarity is measured by calculating the dynamic time warping (DTW) distance between the generated trajectory and the real trajectory, used to assess the anthropomorphism and realism of the trajectory. The collision event rate for a single scene is calculated as follows: In the formula, Represents intelligent agents Or static objects in the environment over time Location, It is a small distance threshold used to determine whether a collision has occurred. A lower collision rate indicates that the model can generate safer interactive behavior.

[0066] The road deviation event rate for a single scenario is calculated as follows: In the formula, This is the total number of simulation scenarios. It is a scene The number of agents in the system Representing a scene Legal road areas It is an indicator function, if Not here The value is 1 if the road is within the acceptable range, and 0 otherwise. A lower proportion of road deviation events indicates that the model follows road constraints better.

[0067] The reverse incident rate for a single scenario is calculated as follows: In the formula, It is an intelligent agent In time The velocity vector, It is a scene In time The legal road direction vector, This represents the angle between two vectors. This is the threshold angle for determining whether a vehicle is driving in the wrong direction. A lower proportion of vehicles driving in the wrong direction indicates that the model is better able to comply with traffic rules.

[0068] The motion infeasibility calculation method for a single scene is as follows: In the formula, and They are intelligent agents In time acceleration and velocity vector, and These are the vehicle's maximum acceleration and speed limits. A lower infeasibility rate indicates that the generated trajectory is more physically realistic.

[0069] The trajectory similarity calculation method for a single scene is as follows: In the formula, It is an intelligent agent In time The generated trajectory position, This represents the actual trajectory location. A lower logarithmic divergence value indicates that the generated trajectory is closer to the actual trajectory.

[0070] The model (HySD model) and various benchmark models were evaluated on a test dataset, without distinguishing between motor vehicles, non-motor vehicles, and pedestrians. Specific results are shown in Table 2. Overall, the trajectories generated by HySD better adhere to road constraints and reduce interaction conflicts, achieving comprehensive improvements in safety and realism across multiple dimensions, demonstrating its superior generalization ability in complex traffic scenario simulations.

[0071] Table 2 Comparison of closed-loop simulation performance of the models on the WOMD dataset Figure 3 This paper presents a dynamic intersection scenario generated from the Waymo dataset, where vehicle trajectories visually reflect speed changes through color gradients (from blue to yellow, corresponding to speeds from 0 to 20 m / s). The teacher model exhibits frequent traffic rule violations during simulation, such as failing to stop at designated stop lines and disobeying yield rules at intersections. This behavior not only leads to additional unsafe maneuvers but also significantly increases the risk of collisions with other vehicles or pedestrians. For example, in scenarios at t = 2.0s and t = 4.0s, vehicles generated by the teacher model fail to stop at red lights or stop lines, directly entering the intersection without prioritizing oncoming or intersecting traffic flows. In contrast, the HySD model provided in this invention demonstrates significantly more realistic and rule-compliant driving behavior in the same test scenario. Vehicles generated by HySD strictly follow traffic signal instructions, maintaining a consistent stop at the stop line at t = 2.0s and t = 4.0s, without unauthorized entry into the intersection. Furthermore, at t = 8.0s, the vehicle generated by the HySD model was able to smoothly and appropriately execute a right-turn trajectory. The trajectory was smooth and fully adapted to the surrounding traffic flow and road conditions, avoiding potential conflicts with nearby vehicles or pedestrians. These qualitative improvements demonstrate that the HySD model has significantly enhanced its ability to perceive the environment and integrate decision-making, thereby achieving more natural, compliant, and safe driving behavior in complex multi-agent traffic scenarios.

[0072] like Figure 4As shown, the scenario depicts vehicle 11 merging into a main road, generated from the Waymo dataset, involving multi-lane interaction and dynamic traffic flow. In the teacher model, vehicle 11 failed to interact reasonably with surrounding vehicles during its merging, exhibiting unsafe behavior by crossing multiple lanes at once, resulting in collisions with vehicles 0 and 7 during the merging process. Furthermore, in the scenario generated by the teacher model, vehicle 2 exhibited significant abnormal behavior, deviating drastically from its designated lane when approaching a building, ultimately colliding with the environment (building). This abnormal behavior not only reflects the instability of trajectory generation but also exposes the limitations of the teacher model in handling vehicle interactions in complex scenarios. In contrast, the scenario generated by the HySD model provided in this invention demonstrates significant improvement. Vehicle 11 did not collide with any vehicles during its merging into the main road and achieved good dynamic interaction with vehicle 0, maintaining a stable following distance after successful merging, with a smooth trajectory that complies with traffic regulations. Similarly, vehicle 2 did not exhibit the anomalous deviations seen in the teacher model within the HySD model; its trajectory remained stable, successfully entering the building area without any scraping or collision with the environment. These results demonstrate that the HySD model significantly improves its ability to generate anthropomorphic, realistic trajectories while effectively reducing anomalous behavior, exhibiting stronger robustness and adaptability to complex traffic environments.

[0073] Figure 5 This paper focuses on the simulation behavior of vehicle 5 at an unsignalized T-junction, based on an intersection scenario generated using the Waymo dataset, involving the dynamic interaction between oncoming and intersecting vehicles. In the teacher model, vehicle 5 failed to take effective action even when safety was confirmed, exhibiting an unnecessary state of stillness, even though both oncoming and intersecting vehicles (vehicle 1) had slowed down, providing a safe left turn. Furthermore, in the teacher model's scenario, vehicle 2 exhibited severe aberrant behavior, specifically a sudden lane change and entry into a reverse driving state, further exacerbating the instability of the traffic scene. This aberrant behavior not only violates traffic rules but may also lead to additional collision risks. In contrast, the HySD model provided in this invention, which generates vehicle 5, demonstrates more human-like driving behavior in the same scenario. After confirming that both oncoming and intersecting vehicles (vehicle 1) had slowed down and the intersection was safe, vehicle 5 performed a smooth and expected left turn, with a continuous trajectory and no interference with surrounding traffic. Simultaneously, vehicle 2, generated by the HySD model, strictly followed the specified route, maintaining a stable driving state without exhibiting the aberrant behaviors such as reverse driving or sudden lane changes seen in the teacher model. These results further validate the HySD model's ability to generate stable, realistic, and compliant trajectories in complex unsignalized intersection scenarios, highlighting its advantages in perception and decision-making integration in multi-agent interactions.

[0074] Table 3 compares the generalization performance of the models in closed-loop simulations on the INTERACTION dataset. The HySD model and other benchmark models were evaluated in unseen scenarios (INTERACTION dataset), and the specific results are shown in Table 3. Overall, by integrating multiple self-distillation mechanisms, HySD significantly reduces the metrics for road departure events, collision events, and wrong-way events, and also demonstrates higher realism and stability. It achieves a comprehensive improvement in unseen scenarios, validating its generalization advantage in complex traffic simulations.

[0075] Figure 6This presentation showcases closed-loop simulation results of the HySD model on the INTERACTION dataset for several typical scenarios, including the ZS0 merging area in China (scenario number DR_CHN_Merging_ZS0), the GL intersection in the United States (scenario number DR_USA_Intersection_GL), the EP1 intersection in the United States (scenario number DR_USA_Intersection_EP1), and the LN roundabout in China (scenario number DR_CHN_Roundabout_LN). In the ZS0 merging area scenario in China, the agent generated by the teacher model (baseline model) experiences collisions during merging, resulting in a high collision rate (4.272%) and noticeable trajectory jitter. In the GL intersection scenario in the United States, the agent generated by the teacher model frequently crosses road boundaries, exhibiting high instability. For example, in the simulation from t = 2.0 s to 8.0 s, vehicles fail to follow lane lines, leading to a high road deviation event rate (1.702%). In the US EP1 intersection scenario, the teacher model's agent frequently enters the oncoming lane, violating traffic rules and increasing the risk of collisions. For example, at t = 4.0 seconds, a vehicle mistakenly enters the oncoming lane, triggering a potential conflict. In the Chinese LN roundabout scenario, the trajectory generated by the teacher model deviates significantly from the expected lane, fails to conform to road markings, and exhibits high logarithmic divergence (2.419 meters). In contrast, the HySD model provided by this invention demonstrates significant improvements in these scenarios: In the Chinese ZS0 merging area scenario, the trajectory generated by the HySD model is smoother and significantly reduces collisions, lowering the collision rate to 2.365%. At the US GL intersection, the trajectory generated by the HySD model strictly follows lane boundaries, significantly reducing out-of-bounds behavior, and lowering the road deviation event rate to 1.284%. At the US EP1 intersection, the trajectory generated by the HySD model completely avoids entering the oncoming lane, reducing the wrong-way event rate to 5.765%, which is better than the baseline model (7.567%). In the LN roundabout scenario in China, the trajectory generated by the HySD model closely follows the lane markings, exhibiting smoothness and stability, with the logarithmic divergence reduced to 2.383 meters. These qualitative results further confirm the high stability and generalization ability of the HySD model in uncomplicated interactive scenarios, highlighting its superior performance in generating realistic and safe trajectories.

[0076] In summary, the HySD model can efficiently generate autonomous driving simulation test scenarios, dynamically reproduce the behavioral interactions of multiple agents in complex traffic environments, accurately capture the multimodal uncertainty of trajectories, demonstrate the model's adaptability and robustness to unknown scenarios, and successfully achieve high-fidelity cross-domain transfer optimization.

[0077] This method has the following characteristics: (1) Achieve unified modeling and interactive generation of multiple types of traffic participants. By collecting real traffic datasets and establishing observation-action feature matrices, motor vehicles, non-motor vehicles and pedestrians are included in the same framework, thereby significantly improving the consistency, realism and accuracy of multi-agent interaction in mixed traffic flow simulation.

[0078] (2) Introducing a diffusion generation mechanism, using the probability generation process of gradually denoising noise to reconstruct the trajectory distribution of multiple agents, and combining it with the behavior predictor module, can effectively capture the multimodal characteristics, uncertainties and spatiotemporal dependencies of traffic flow, and generate dynamic scenes that are more in line with real traffic patterns.

[0079] (3) Through various knowledge self-distillation mechanisms (such as feature, response and relation self-distillation), the adaptive alignment of feature and policy distribution within the model is achieved, and the teacher-student network transfers structured knowledge, thereby improving the stability and robustness of cross-scene transfer and complex interaction.

[0080] Example 2 Based on the foregoing embodiments, see Figure 7 This embodiment provides a knowledge distillation-based autonomous driving test scenario simulation generalization generation system, which is used to implement the knowledge distillation-based autonomous driving test scenario simulation generalization generation method as described in Embodiment 1. The simulation system includes a model building module and an application simulation module.

[0081] The model building module includes: The data acquisition and preprocessing unit is used to collect multi-source traffic flow data, including trajectory information of motor vehicles, non-motor vehicles and pedestrians, and to perform spatiotemporal alignment, feature extraction and normalization processing. The multi-agent modeling unit constructs agent control strategies based on a diffusion generation framework to achieve interactive modeling of multiple types of traffic participants; The knowledge self-distillation generalization optimization unit aligns the feature distributions of teacher and student models through feature distillation and relation distillation, thereby improving the generalization performance and stability of the model in complex traffic scenarios.

[0082] The application simulation module includes: The scene configuration unit is used to generate an initial simulation scene with road topology, signal control, and agent distribution information. The behavior prediction and scene generation unit generates multimodal and controllable agent behavior trajectories based on the trained model, achieving highly realistic reproduction of the dynamic evolution process of traffic flow.

[0083] See Figure 8To illustrate the simulation process using the established simulation system, the following steps are taken: First, the model building module acquires real road topology, traffic signal status, and historical trajectory data of the agents to analyze the multi-agent interaction behavior and clarify the state and context of each time step. Then, a teacher model is built based on the diffusion model and trained through forward diffusion and reverse denoising processes. Next, the student model undergoes self-distillation generalization optimization to achieve knowledge transfer and dynamically adjust the loss weights. Finally, the application simulation module is used to evaluate the trajectory accuracy and the realism of the closed-loop simulation based on the obtained model, generating complex and diverse highly interactive scenarios.

[0084] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A generalized generation method for autonomous driving test scenario simulation based on knowledge distillation, characterized in that, Includes the following steps: Collect real-world autonomous driving datasets, including trajectory information of various traffic participants such as motor vehicles, non-motor vehicles, and pedestrians, establish a unified observation-action feature matrix, and construct traffic simulation scenarios to support multi-agent trajectory generation and behavior modeling. Using each traffic participant as an intelligent agent, a multi-agent trajectory generation framework is constructed based on a diffusion model. The trajectory data is gradually subjected to noise perturbation and reverse denoising recovery to achieve high-fidelity scene synthesis. A multi-agent trajectory generation framework is constructed based on the diffusion model. The trajectory of the agent is generated through an iterative denoising process to model the interactive behavior of multiple agents in complex traffic environments. A self-distillation mechanism of teacher-student network is adopted to transfer the structured knowledge of the teacher model in the source dataset to the student model. The simulated annealing strategy is combined to dynamically balance the task loss and distillation loss, thereby improving the cross-scene adaptability and generation stability of the model. Zero-shot closed-loop simulation verification was performed on other independent autonomous driving datasets that were not used in the training.

2. The method for generalizing and generating autonomous driving test scenario simulations based on knowledge distillation according to claim 1, characterized in that, The traffic simulation scenario is defined as follows: in, express An intelligent agent in A set of trajectories at each time step. The dimension of the state vector for each agent; Corresponding to the The trajectory sequence of each agent, and the state at each time step. Including location ,speed Heading angle and border dimensions property; , They are respectively and The velocity component in the direction, These represent length, width, and height, respectively. and For the joint control sequence driving trajectory evolution, where To control the dimension of the input vector; Corresponding to the The control sequence of an agent, with its control input at each time step. Includes acceleration and rate of change of heading angle This determines the agent's movement behavior; Meanwhile, traffic simulation scenario context include: Road topology ,for A collection of multiple lane segments, each lane consisting of... It consists of 1 key points, each of which is 1 key point. Two-dimensional coordinates describe the lane centerline; Traffic signal status ,include The status of each traffic signal The dimension representing each signal state; Historical status ,for An agent in the past The trajectory state of the time step provides the interaction history state. The dimension of the trajectory state vector; Initial joint state ; The generation and evolution of traffic simulation scenarios follow a discrete-time dynamic model: In the formula, It is a joint dynamic function, which is based on a particle kinematics model and controls the input. Achieve joint trajectory state of each intelligent agent Forward update; For the Individual agents ( ), in the Time step ( ) state From the state of the previous moment With control input Decide: in For location, for and The velocity component in the direction, For heading angle, and These are the corresponding control inputs. The longitudinal acceleration and the rate of change of the heading angle in the middle. For time step; The joint dynamic function integrates the independent motion updates of all agents and can incorporate interaction terms into the modeling, thereby enabling the tracking of the joint trajectory state of multiple agents. Update; The task of generating the traffic simulation scenario is formalized as the following trajectory optimization problem: , in, For parameterized objective functions, measure the control sequence and corresponding trajectory In context The rationality, For the model parameters, the optimization objective is to minimize Generate control sequences that conform to real traffic behavior and environmental constraints. .

3. The method for generalizing and generating autonomous driving test scenario simulations based on knowledge distillation according to claim 1, characterized in that, The multi-agent trajectory generation framework includes: a scene encoder module; The scene encoder module adopts a query-controlled multi-head attention (QCMHA) architecture to capture the interaction relationships between multiple agents and traffic scene constraints. The process is as follows: Taking road topology, traffic signal status, and multi-agent historical states as inputs, global information is mapped to a local coordinate system centered on the target agent through coordinate system normalization, enhancing the relative space modeling capability. Subsequently, a gated recurrent unit (GRU) is used to extract temporal features, and lane-level geometric features are obtained by combining a multilayer perceptron and max pooling, which are then embedded into the signal states. After unifying the feature dimensions through linear projection, the data is input into a Transformer encoder, which models the dynamic interaction between multi-agents and the environment through self-attention and cross-attention mechanisms, generating a high-dimensional latent representation of the scene. A history discarding strategy is introduced during training, randomly discarding some historical information to strengthen the model's dependence on the current semantics, thereby improving the generalization and causal consistency of closed-loop simulation and providing robust contextual support for trajectory generation.

4. The method for generalizing and generating autonomous driving test scenario simulations based on knowledge distillation according to claim 1, characterized in that, The multi-agent trajectory generation framework includes: a trajectory generation framework; The trajectory generation framework is based on a denoised diffusion probability model, in which the forward diffusion process and the reverse denoising process together constitute the core mechanism for generating multi-agent control sequences, as follows: In the forward diffusion phase, the model gradually injects Gaussian noise into the real control sequence through a Markov chain, making the distribution approach a standard Gaussian. Logarithmic noise scheduling is used to balance the signal-to-noise ratio, preserving structural information in the early stage and adding noise in the later stage to improve generalization ability. In the reverse denoising phase, the model starts from the noise sample, combines scene encoding and diffusion step information, and gradually recovers the control sequence that conforms to physical and traffic constraints through a Transformer denoising network, and optimizes the generation continuity with smooth L1 loss. To achieve controllable generation, a controllable guidance mechanism is introduced to dynamically adjust the denoising mean through constraints such as speed limit, collision avoidance, and lane keeping, so as to achieve task-oriented and rule-compliant trajectory generation.

5. The method for generalizing and generating autonomous driving test scenario simulations based on knowledge distillation according to claim 1, characterized in that, The multi-agent trajectory generation framework includes: a behavior predictor; The behavior predictor, while maintaining the stability of independent optimization of the noise reduction generator and the scene encoder, achieves the modeling of the agent's multimodal motion patterns, as detailed below: Using scene encoding and a set of multi-endpoint anchors as input, a lightweight multilayer perceptron network is used to predict multiple possible control sequences for each agent at future time points and their corresponding modal probability weights. The model weights the outputs of each modality in a hybrid distribution, with the modality weights reflecting the probability of their behavior in a specific scenario. By introducing this explicit behavioral prior constraint diffusion process, the noise reduction generator is prevented from relying too much on the noise structure and ignoring semantic features, thereby improving the diversity, interpretability and scenario consistency of trajectory generation.

6. The method for generalizing and generating autonomous driving test scenario simulations based on knowledge distillation according to claim 1, characterized in that, The self-distillation mechanism of the teacher-student network includes a feature self-distillation mechanism; The Feature Self-Distillation (FSD) mechanism is used to align the feature representations of the teacher model and the student model in the intermediate layer of the scene encoder, as follows: The scene encoder generates feature maps based on QCMHA. The teacher model originates from a diffusion generation model pre-trained on diverse traffic data based on the aforementioned modules, possessing rich structured semantics. The student model has the same structure as the teacher model, with parameters randomly initialized. FSD aligns the feature distributions of the two models, enabling the student model to inherit the teacher model's ability to represent complex scenes in a structured manner. Features are L2 normalized before computation to focus on semantic consistency rather than numerical differences, thereby effectively transferring knowledge from the teacher model, mitigating overfitting, and enhancing the stability and rule compliance of the student model in unseen scenarios.

7. The method for generalizing and generating autonomous driving test scenario simulations based on knowledge distillation according to claim 1, characterized in that, The self-distillation mechanism of the teacher-student network includes a response self-distillation mechanism; The Response Self-Distillation (RSD) mechanism is used to align the action distributions of the teacher and student models during the denoising generation process, as follows: The self-distillation mechanism uses the denoised action output by the teacher model as a soft objective. By minimizing the difference in probability distribution between the two, it guides the student model to learn a more robust control strategy. A temperature parameter is introduced to adjust the smoothness of the distribution. Higher temperatures highlight the relative relationships between actions, while lower temperatures strengthen the learning of key actions. Through distribution alignment based on KL divergence, the student model inherits the behavioral priors and rule constraints of the teacher model, generating smooth, realistic, and physically consistent trajectories in complex scenarios such as intersections and roundabouts, thereby significantly improving generation consistency and generalization ability.

8. The method for generalizing and generating autonomous driving test scenario simulations based on knowledge distillation according to claim 1, characterized in that, The self-distillation mechanism of the teacher-student network includes a relational self-distillation mechanism; The relational self-distillation mechanism RlSD is used to align the knowledge representations of the teacher model and the student model in multi-agent interaction relationship modeling, as follows: The relational self-distillation mechanism constructs an agent interaction graph, uses action differences as edge weights to describe mutual influence, and calculates the interaction differences between the two models during the denoising process. Cosine similarity loss is used to align the consistency of action directions between agent pairs, thereby enhancing the learning of interaction structures.

9. The method for generalizing and generating autonomous driving test scenario simulations based on knowledge distillation according to claim 1, characterized in that, Also includes: Trajectory accuracy assessment: The deviation between the generated trajectory and the real trajectory is quantified from two aspects: long-term prediction and multimodal trajectory generation, to evaluate the model's accuracy and spatiotemporal consistency in trajectory generation; Closed-loop interaction authenticity assessment: The safety and robustness of the generated trajectory are assessed from multiple aspects, including driving behavior safety, traffic rule compliance and authenticity.

10. A generalized generation system for autonomous driving test scenario simulation based on knowledge distillation, characterized in that, The system for implementing the method as described in any one of claims 1-9 includes: a model building module and an application simulation module; The model building module includes: The data acquisition and preprocessing unit is used to collect multi-source traffic flow data, including trajectory information of motor vehicles, non-motor vehicles and pedestrians, and to perform spatiotemporal alignment, feature extraction and normalization processing. Multi-agent modeling unit to realize multi-agent trajectory generation framework: Construct agent control strategy based on diffusion generation framework to realize interactive modeling of multiple types of traffic participants; The knowledge self-distillation generalization optimization unit realizes the self-distillation mechanism of the teacher-student network: by aligning the feature distribution of the teacher and student models through feature distillation and relation distillation, the generalization performance and stability of the model in complex traffic scenarios are improved. The application simulation module includes: The scene configuration unit is used to generate an initial simulation scene with road topology, signal control, and agent distribution information. The behavior prediction and scene generation unit generates multimodal and controllable agent behavior trajectories based on the trained model, achieving highly realistic reproduction of the dynamic evolution process of traffic flow.

Citation Information

Patent Citations

  • Implementation method of lightweight target detection neural network

    CN118133905A

  • Knowledge-driven end-to-end automatic driving method based on sparse expert mechanism and diffusion model

    CN120716776A

  • High-fidelity lightweight world model construction method for end-to-end automatic driving test

    CN120909949A

  • Multi-vehicle cooperative controllable confrontation test method based on diffusion model

    CN121389817A

  • Self-adaptive diffusion trajectory planning method based on reinforcement learning guidance

    CN121432938A

Cited By

  • A data-driven autonomous driving simulation scenario generation method and test system

    CN122285529A

  • An automatic driving method based on hierarchical world cognition and intention constraint generation

    CN122390087A