Diffusion-based safety key scene generation method and system

By collecting and labeling real driving scenario data, constructing and training a scenario diffusion model, and using guidance conditions to generate safety-critical scenarios, the problem of limited generation scale and complexity in existing technologies has been solved, and efficient, diverse and realistic generation of safety-critical scenarios has been achieved.

CN121786472APending Publication Date: 2026-04-03WESTERN CHINA SCI CITY INNOVATION CENT OF INTELLIGENT & CONNECTED VEHICLES (CHONGQING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing methods for generating safety-critical scenarios cannot effectively utilize real-world traffic network data and traffic flow information, limiting the scale and complexity of generation, and resulting in insufficient diversity and realism of generated scenarios.

Method used

By collecting and labeling a large amount of real driving scenario data, a scenario diffusion model is constructed and trained. The future trajectories of surrounding traffic participants are generated using guidance conditions, thus generating safety-critical scenarios.

Benefits of technology

It significantly expanded the scale of generating safety-critical scenarios, improved the diversity and realism of scenarios, and enhanced generation efficiency and scenario effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786472A_ABST
    Figure CN121786472A_ABST
Patent Text Reader

Abstract

The invention discloses a diffusion-based safety key scene generation method and system, and relates to the technical field of automatic driving, and the method comprises the steps: collecting and marking historical driving scene data; constructing a basic scene diffusion model; training the basic scene diffusion model; collecting new driving scene data and screening target driving scene data; marking the target driving scene data; screening confrontation traffic participants, and outputting index information of the confrontation traffic participants; generating a guide condition; generating future trajectories of surrounding traffic participants; and generating a security key scene file and storing the security key scene file in a file library. In a generation process, a large amount of real traffic road network data and traffic flow information are used as a generation basis, and a guiding condition is added in a denoising process of a target scene so as to generate a future trajectory with high naturalness, diversified behavior styles and key safety of surrounding traffic participants. Therefore, the generation efficiency and the authenticity and effectiveness of the generated scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous driving technology, and more specifically, to a diffusion-based method and system for generating safety-critical scenarios. Background Technology

[0002] Autonomous driving, as one of the main arenas for the practical application of artificial intelligence technology, has made significant progress in recent years. Before the actual deployment of autonomous driving systems, conducting thorough safety assessments is a crucial and challenging task. Conventional assessment methods involve deploying autonomous driving systems in the real world and conducting long-term road tests under various traffic scenarios. However, due to the long-tail nature of real-world scenario distribution, safety-critical scenarios that pose a serious threat to autonomous driving systems are extremely rare, making the time, manpower, and material costs of conventional assessment methods unacceptable. Recently, with the rapid development of deep generative models, utilizing related technologies to directly generate safety-critical scenarios to accelerate the safety testing of autonomous driving systems is a promising direction.

[0003] Currently, there are three main types of methods for generating safety-critical scenarios: the first type is the data-driven method, which trains a generative model based on real data collected from road tests to generate safety-critical scenarios; the second type is the adversarial method, which generates safety-critical scenarios by controlling surrounding traffic participants to generate adversarial attacks on the host vehicle; and the third type is the knowledge-based method, which uses predefined knowledge rules such as physical constraints or traffic rules to guide the generative model to generate safety-critical scenarios.

[0004] Data-driven methods face significant challenges in generating safety-critical scenarios due to the extreme imbalance between common and safety-critical scenarios in real-world datasets collected from road tests. This imbalance makes directly training generative models from this data highly difficult and inefficient. Ordinary adversarial methods, unconstrained by natural forces, often generate safety-critical scenarios that lack realism due to the simplistic and aggressive actions of surrounding road users. Knowledge-based methods suffer from difficulties integrating knowledge rules into the generative model, and manually created rules often have limited complexity, resulting in implementation difficulties and insufficient effectiveness of the generated safety-critical scenarios.

[0005] It is evident that current methods for generating safety-critical scenarios only support the generation of such scenarios under specific and simple settings. They cannot utilize the vast amounts of real-world traffic network data and traffic flow information, thus limiting the scale and complexity of generated safety-critical scenarios. Furthermore, in existing methods, the behavioral styles of surrounding traffic participants are often singular and fixed. This lack of diversity in their behavioral characteristics results in insufficient diversity and richness of the generated safety-critical scenarios. In conclusion, current mainstream methods for generating safety-critical scenarios fail to simultaneously achieve efficiency, realism, and effectiveness. Summary of the Invention

[0006] This invention provides a diffusion-based method and system for generating safety-critical scenarios, which overcomes at least one technical problem existing in the prior art.

[0007] On one hand, the present invention provides a diffusion-based method for generating safety-critical scenarios, including:

[0008] Collect data from multiple historical driving scenarios and save the historical driving scenario data to a storage server;

[0009] The historical driving scenario data is labeled to obtain labeled historical driving scenario data;

[0010] Construct a basic scenario diffusion model;

[0011] The basic scene diffusion model is trained using labeled historical driving scene data to obtain the trained scene diffusion model.

[0012] Collect data from multiple new driving scenarios and save the new driving scenario data to a storage server;

[0013] Filter target driving scenario data from the new driving scenario data;

[0014] The target driving scenario data is labeled to obtain labeled target driving scenario data;

[0015] Filter adversarial traffic participants from the labeled target driving scenario data and output the index information of the adversarial traffic participants;

[0016] Generate guiding conditions, wherein the guiding conditions are expressed as follows: Where, N a This represents the total number of surrounding traffic participants, including 1 opposing traffic participant and N. a -1 non-adversarial traffic participants, where J(i) represents the guiding function of the i-th traffic participant;

[0017] Using the trained scene diffusion model, the labeled target driving scene data, the index information of the adversarial traffic participants, and the guidance conditions, the future trajectories of surrounding traffic participants are generated;

[0018] Based on the target driving scenario data and the future trajectories of the surrounding traffic participants, a safety-critical scenario file is generated;

[0019] The safety-critical scenario files are stored in the safety-critical scenario file library.

[0020] Optionally, the labeled historical driving scenario data includes first high-precision map information and first traffic participant status information, the first traffic participant status information including historical status information and real future status information; the basic scenario diffusion model includes a scenario encoder and a diffusion decoder;

[0021] The basic scene diffusion model is trained using labeled historical driving scene data, specifically as follows:

[0022] The first high-precision map information and the historical state information are input into the scene encoder for encoding to obtain the encoding conditions;

[0023] The real future state information is subjected to noise processing to obtain noisy future state information;

[0024] The noisy future state information and the encoding conditions are input into the diffusion decoder for iterative decoding to obtain the generated future state information.

[0025] The base scene diffusion model is iteratively updated using the real future state information and the generated future state information to obtain the trained scene diffusion model.

[0026] Optionally, the basic scene diffusion model is iteratively updated using the real future state information and the generated future state information, specifically as follows:

[0027] Calculate the mean squared error loss based on the actual future state information and the generated future state information;

[0028] The convergence of the basic scene diffusion model is determined based on the mean squared error loss. If convergence is achieved, the current basic scene diffusion model is used as the trained scene diffusion model. If convergence is not achieved, the parameters θ of the basic scene diffusion model are updated, and the iterative decoding step of inputting the noisy future state information and the encoding conditions into the diffusion decoder is returned.

[0029] Optionally, target driving scenario data can be filtered from the new driving scenario data, specifically as follows:

[0030] Target driving scenario data is filtered from the new driving scenario data stored on the storage server according to predetermined conditions; the predetermined conditions include at least a specific road map type and the overall complexity of the driving scenario.

[0031] Optionally, the labeled target driving scenario data includes second high-precision map information and second traffic participant status information;

[0032] The process of filtering adversarial traffic participants from the labeled target driving scenario data is as follows:

[0033] Based on the second high-precision map information and the status information of the second traffic participant, the surrounding traffic participant with the highest risk to the main vehicle is selected as the opposing traffic participant.

[0034] Alternatively, for adversarial traffic participants, its guided function is represented as J adv =J col +J reg For non-confrontational traffic participants, the guided function is represented as J. non-adv =J reg J col J is the collision guidance function. reg For natural constraint guiding functions;

[0035] Among them, the collision guidance function The distance between the future trajectories of adversarial traffic participants generated by the trained scene diffusion model and the actual future trajectories of the main vehicle in the scene is measured; T represents the total number of future time steps; d(t) represents the distance between the trajectory points of the adversarial traffic participants generated by the trained scene diffusion model and the actual trajectory points of the main vehicle in the scene at future time step t; natural constraint guidance function. This represents the difference between the future trajectory of a traffic participant generated by the trained scene diffusion model and the actual future trajectory of that traffic participant; d represents the distance between the trajectory points of the traffic participant generated by the trained scene diffusion model and the actual trajectory points of that traffic participant at future time step t. thr This indicates the set distance threshold.

[0036] Optionally, the future trajectories of surrounding traffic participants are generated using the trained scene diffusion model, the labeled target driving scene data, the index information of the adversarial traffic participants, and the guidance conditions, specifically as follows:

[0037] The noisy future trajectories of surrounding traffic participants and the index information of the adversarial traffic participants are input into the trained scene diffusion model, and K-step iterative denoising is performed; in each iterative denoising process, the guiding conditions are used to optimize the results.

[0038] On the other hand, the present invention also provides a diffusion-based safety-critical scenario generation system, comprising:

[0039] The first acquisition module is used to acquire multiple historical driving scenario data and save the historical driving scenario data to a storage server;

[0040] The first annotation module is used to annotate the historical driving scenario data to obtain annotated historical driving scenario data;

[0041] Modules are used to build basic scene diffusion models;

[0042] The training module is used to train the basic scene diffusion model using labeled historical driving scene data to obtain the trained scene diffusion model.

[0043] The second acquisition module is used to acquire multiple new driving scenario data and save the new driving scenario data to the storage server;

[0044] The first filtering module is used to filter target driving scenario data from the new driving scenario data;

[0045] The second annotation module is used to annotate the target driving scene data to obtain annotated target driving scene data;

[0046] The second filtering module is used to filter adversarial traffic participants from the labeled target driving scenario data and output the index information of the adversarial traffic participants;

[0047] The first generation module is used to generate the guiding conditions, wherein the guiding conditions are expressed as follows: Where, N a This represents the total number of surrounding traffic participants, including 1 opposing traffic participant and N. a -1 non-adversarial traffic participants, where J(i) represents the guiding function of the i-th traffic participant;

[0048] The second generation module is used to generate the future trajectories of surrounding traffic participants by utilizing the trained scene diffusion model, the labeled target driving scene data, the index information of the adversarial traffic participants, and the guidance conditions.

[0049] The third generation module is used to generate a safety-critical scenario file based on the target driving scenario data and the future trajectories of the surrounding traffic participants;

[0050] The storage module is used to store the safety-critical scenario files in the safety-critical scenario file library.

[0051] Optionally, the labeled historical driving scene data includes first high-precision map information and first traffic participant state information, wherein the first traffic participant state information includes historical state information and actual future state information; the basic scene diffusion model includes a scene encoder and a diffusion decoder; the training module includes:

[0052] The first input module is used to input the first high-precision map information and the historical state information into the scene encoder for encoding to obtain encoding conditions;

[0053] A noise-adding module is used to add noise to the real future state information to obtain noisy future state information.

[0054] The second input module is used to input the noisy future state information and the encoding conditions into the diffusion decoder for iterative decoding to obtain the generated future state information.

[0055] The update module is used to iteratively update the basic scene diffusion model using the real future state information and the generated future state information to obtain the trained scene diffusion model.

[0056] Optionally, the update module is specifically used for:

[0057] Calculate the mean squared error loss based on the actual future state information and the generated future state information;

[0058] The convergence of the basic scene diffusion model is determined based on the mean squared error loss. If convergence is achieved, the current basic scene diffusion model is used as the trained scene diffusion model. If convergence is not achieved, the parameters θ of the basic scene diffusion model are updated, and the iterative decoding step of inputting the noisy future state information and the encoding conditions into the diffusion decoder is returned.

[0059] The innovative aspects of this invention include:

[0060] 1. In this embodiment, a large amount of real traffic network data and traffic flow information are used as the basis for generation during the generation process. Therefore, this method greatly expands the generation scale of safety-critical scenarios and can generate complex safety-critical scenarios, which is one of the innovations of this embodiment.

[0061] 2. In this embodiment, a multi-style scene diffusion model is trained using real driving scenario data with diverse behavioral styles. This model can randomly generate new trajectories with diverse behavioral styles of surrounding traffic participants, and these trajectory data have a high degree of naturalness, thereby fundamentally improving the diversity and realism of the generated safety-critical scenarios, which is one of the innovations of this embodiment.

[0062] 3. In this embodiment, adversarial traffic participants are extracted from the target scene used to generate safety-critical scenarios. Then, guidance conditions are generated for both adversarial and non-adversarial traffic participants. Guidance conditions are added during the denoising process of the target scene to generate future trajectories of surrounding traffic participants with high naturalness, diverse behavioral styles, and safety-criticality. This improves the generation efficiency and the realism and effectiveness of the generated scene, which is one of the innovative points of this embodiment. Attached Figure Description

[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0064] Figure 1 A flowchart of a generation method provided in an embodiment of the present invention;

[0065] Figure 2 A schematic diagram of a workflow after adding guiding conditions, provided in an embodiment of the present invention;

[0066] Figure 3 A flowchart for model training provided in an embodiment of the present invention;

[0067] Figure 4 Another flowchart for model training provided in an embodiment of the present invention;

[0068] Figure 5 A schematic diagram of the generation system provided in an embodiment of the present invention;

[0069] Figure 6 This is a schematic diagram of a training module provided in an embodiment of the present invention. Detailed Implementation

[0070] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0071] It should be noted that the terms "comprising" and "having," and any variations thereof, in the embodiments and drawings of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0072] This invention discloses a diffusion-based method and system for generating safety-critical scenarios. These will be described in detail below.

[0073] Figure 1 A flowchart of a generation method provided in an embodiment of the present invention is provided below. Figure 1 The diffusion-based method for generating security-critical scenarios provided in this invention includes:

[0074] Step 1: Collect data from multiple historical driving scenarios and save the historical driving scenario data to the storage server;

[0075] Step 2: Label the historical driving scenario data to obtain labeled historical driving scenario data;

[0076] Step 3: Construct a basic scene diffusion model;

[0077] Step 4: Train the basic scene diffusion model using the labeled historical driving scene data to obtain the trained scene diffusion model;

[0078] Step 5: Collect data from multiple new driving scenarios and save the new driving scenario data to the storage server;

[0079] Step 6: Filter the target driving scenario data from the new driving scenario data;

[0080] Step 7: Label the target driving scenario data to obtain labeled target driving scenario data;

[0081] Step 8: Filter adversarial traffic participants from the labeled target driving scenario data and output the index information of adversarial traffic participants;

[0082] Step 9: Generate bootstrap conditions, which are represented as follows: Where, N a This represents the total number of surrounding traffic participants, including 1 opposing traffic participant and N. a -1 non-adversarial traffic participants, where J(i) represents the guiding function of the i-th traffic participant;

[0083] Step 10: Using the trained scene diffusion model, labeled target driving scene data, adversarial traffic participant index information and guidance conditions, generate the future trajectories of surrounding traffic participants;

[0084] Step 11: Generate a safety-critical scenario file based on the target driving scenario data and the future trajectories of surrounding traffic participants;

[0085] Step 12: Store the safety-critical scenario files in the safety-critical scenario file library.

[0086] For details, please refer to Figure 1 The diffusion-based safety critical scenario generation method provided in this invention first collects multiple historical driving scenario data in step 1. Existing safety critical scenario generation methods only support generating safety critical scenarios under specific and simple scenario settings, failing to utilize the large amounts of real-world traffic network data and traffic flow information, thus limiting the scale and complexity of safety critical scenario generation. To overcome this problem, this invention collects a large amount of real driving scenario data from different geographical regions and environments. The collected driving scenario data exhibits diverse behavioral styles of traffic participants, specifically reflected in indicators such as acceleration during driving, relative distance between them, and headway when following other vehicles. For ease of subsequent use, after collecting the driving scenario data, these diverse behavioral style driving scenario data are stored in a storage server.

[0087] Because the data formats of the collected scene data are inconsistent, in order to ensure that the input data format is consistent during model training, this invention performs uniform format annotation processing on the historical driving scene data with diverse behavioral styles stored in the storage server in step 2. The annotation format can be similar to that of the Waymo Motion Open Dataset dataset. The content to be annotated in each driving scene includes high-precision map information of the scene, state information of traffic participants (mainly composed of historical state information, current time step state information, and future state information of traffic participants, and the state information of traffic participants at each time step includes position, speed, heading angle, etc.), and some necessary global information in the scene, such as the index of the main vehicle in the scene, the index of the current time step, traffic signals, etc. After annotation, the annotated historical driving scene data is obtained.

[0088] After annotation is completed, the annotated scene data can be used for model training. Model training first requires building a basic scene diffusion model in step 3. Then, in step 4, the basic scene diffusion model is fully trained using the annotated historical driving scene data. When the basic scene diffusion model converges, the trained scene diffusion model is obtained.

[0089] After model training is complete, step 5 collects multiple new driving scenario data and saves this data to a storage server for use in generating safety-critical scenarios. Then, in step 6, target driving scenario data for generating safety-critical scenarios is selected from the new driving scenario data. When selecting target driving scenario data from the new driving scenario data stored on the storage server, filtering can be based on user-defined criteria. These criteria can include, for example, specific road map types and the overall complexity of the driving scenario. Specific road map types include intersections, roundabouts, and highway ramps, while the overall complexity of the driving scenario includes the number of traffic participants.

[0090] In step 7, the target driving scene data is annotated in a standardized format to obtain annotated target driving scene data. Specific annotation details can be found in the annotation instructions for historical driving scene data, and will not be repeated here. The annotated target driving scene data includes second high-precision map information and second traffic participant status information, where traffic participants include the vehicle and surrounding traffic participants.

[0091] After obtaining the labeled target driving scenario data, step 8 is used to select a scenario from the labeled target driving scenario data in sequence. Then, using the second high-precision map information and the status information of the second traffic participants (the main vehicle and surrounding traffic participants) in the scenario at the current time step, the risk level of the surrounding traffic participants to the main vehicle is quantitatively evaluated. Based on the evaluation results, the surrounding traffic participant with the highest risk is selected as the adversary traffic participant, and then the index information of the adversary traffic participant in the scenario is output.

[0092] To improve the realism and effectiveness of the generated safety-critical scenarios, this invention uses step 9 to generate guiding conditions and adds guiding conditions during the denoising process of the trajectory of surrounding traffic participants using the trained scenario diffusion model, thereby guiding the trained scenario diffusion model to finally generate the desired safety-critical trajectory of surrounding traffic participants.

[0093] The guiding condition is expressed as Where, N a This represents the total number of surrounding traffic participants, including 1 opposing traffic participant and N. a -1 non-adversarial traffic participants, where J(i) represents the guiding function of the i-th traffic participant.

[0094] For adversarial traffic participants, the guiding conditions include adversarial guiding conditions and natural constraint guiding conditions; therefore, its guiding function is expressed as J. adv =J col +J regFor other surrounding traffic participants, i.e., non-adversarial traffic participants, the guiding condition only needs to include the natural constraint guiding condition; therefore, its guiding function is expressed as J. non-adv =J reg .

[0095] Among them, J col For collision guidance function, collision guidance function The distance between the future trajectories of adversarial traffic participants generated by the trained scene diffusion model and the actual future trajectories of the main vehicle in the scene is measured; T represents the total number of future time steps; d(t) represents the distance between the trajectory points of adversarial traffic participants generated by the trained scene diffusion model and the actual trajectory points of the main vehicle in the scene at future time step t.

[0096] J reg Natural constraint guiding function, natural constraint guiding function This represents the difference between the future trajectory of a traffic participant generated by the trained scene diffusion model and the actual future trajectory of that traffic participant; d represents the distance between the trajectory points of the traffic participant generated by the trained scene diffusion model and the actual trajectory points of that traffic participant at future time step t. thr This indicates the set distance threshold.

[0097] It should be noted that the true future trajectory of the main vehicle, the true trajectory point of the main vehicle, the true future trajectory of the traffic participants, and the true trajectory point of the traffic participants can all be obtained through the status information of the second traffic participants.

[0098] After obtaining the guiding conditions, step 10 can be used to generate the future trajectories critical to the safety of surrounding traffic participants for each target driving scenario. Scenarios are selected sequentially from the labeled target driving scenario data. First, the second high-precision map information corresponding to the current scenario is input into the map encoder, and the status information of the second traffic participant is input into the traffic participant encoder. The map embedding and traffic participant embedding are extracted using the map encoder and traffic participant encoder respectively. The map embedding and traffic participant embedding together form the encoding conditions.

[0099] The system obtains the real future state information of the surrounding traffic participants corresponding to the current scene, and adds noise to it to obtain the noisy future trajectory of the surrounding traffic participants. Then, the noisy future trajectory of the surrounding traffic participants, the encoding conditions, and the index information of the adversarial traffic participants corresponding to the current scene are input into the trained scene diffusion model, and K-step iterative denoising is performed on it.

[0100] Figure 2 This is a schematic diagram of a workflow after adding guiding conditions according to an embodiment of the present invention. Please refer to it. Figure 2 In each iteration of denoising, the gradient information of the guiding function is first calculated through a query process, and then this gradient information is backpropagated through a guided optimization process to optimize the generated future trajectory. After completing the K-step iterative denoising process, the future trajectories of surrounding traffic participants can be obtained, thus generating highly natural, diverse, and safety-critical future trajectories of surrounding traffic participants.

[0101] After obtaining the future trajectories of surrounding traffic participants, in step 11, a safety-critical scenario file is generated based on the target driving scenario data and the future trajectories of surrounding traffic participants. This safety-critical scenario file contains high-precision map information of the scenario and trajectory information of all traffic participants. The entire trajectory of the main vehicle (including historical trajectory, current time step trajectory, and future trajectory) follows the trajectory of the main vehicle in the target driving scenario data. The historical and current time step trajectories of surrounding traffic participants are also still provided by the target driving scenario data, but their future trajectories are replaced with new future trajectories of surrounding traffic participants generated by the trained scenario diffusion model. Finally, these trajectory data are stitched together to form complete traffic participant trajectories and integrated with the high-precision map information to obtain the final safety-critical scenario file.

[0102] To facilitate subsequent testing of the autonomous driving system, the present invention also stores the safety-critical scenario files in the safety-critical scenario file library through step 12.

[0103] The safety-critical scenario generation method based on diffusion provided by this invention utilizes a large amount of real traffic network data and traffic flow information as the basis for generation. Therefore, this method significantly expands the scale of safety-critical scenario generation and can generate complex safety-critical scenarios. By training a multi-style scenario diffusion model using real driving scenario data with diverse behavioral styles, this model can randomly generate new trajectories with diverse behavioral styles of surrounding traffic participants. These trajectory data have a high degree of naturalness, thereby fundamentally improving the diversity and realism of the generated safety-critical scenarios.

[0104] Furthermore, this invention extracts adversarial traffic participants from the target scene used to generate safety-critical scenarios, then generates guidance conditions for both adversarial and non-adversarial traffic participants, and incorporates guidance conditions during the denoising process of the target scene to generate highly natural, diverse, and safety-critical future trajectories of surrounding traffic participants, thereby improving generation efficiency and the realism and effectiveness of the generated scenarios.

[0105] Optionally, Figure 3Please refer to the flowchart of model training provided in this embodiment of the invention. Figure 1 and Figure 3 The labeled historical driving scenario data includes first high-precision map information and first traffic participant status information. The first traffic participant status information includes historical status information and real future status information. The basic scenario diffusion model includes a scenario encoder and a diffusion decoder. In step 4, the labeled historical driving scenario data is used to train the basic scenario diffusion model. Specifically, step 41, the first high-precision map information and historical status information are input into the scenario encoder for encoding to obtain encoding conditions; step 42, the real future status information is noise-added to obtain noisy future status information; step 43, the noisy future status information and encoding conditions are input into the diffusion decoder for iterative decoding to obtain generated future status information; step 44, the basic scenario diffusion model is iteratively updated using the real future status information and the generated future status information to obtain the trained scenario diffusion model.

[0106] Specifically, the labeled historical driving scenario data includes first-level high-precision map information and first-level traffic participant status information. The first-level traffic participant status information includes historical status information and current future status information. The basic scenario diffusion model includes a scenario encoder and a diffusion decoder.

[0107] Please refer to Figure 3 When training the basic scene diffusion model, encoding is first performed in step 41. Since the first high-precision map information and the historical state information of the first traffic participant need to be encoded separately, the scene encoder includes a map encoder and a traffic participant encoder. The first high-precision map information is input into the map encoder to obtain the map code, and the historical state information is input into the traffic participant encoder to obtain the historical state code.

[0108] Then, in step 42, noise is sampled and used to add noise to the real future state information, resulting in noisy future state information. In step 43, the noisy future state information and the encoding conditions are input into the diffusion decoder for iterative decoding to obtain the generated future state information. Here, the iterative decoding process takes K steps.

[0109] After obtaining the generated future state information, in step 44, the basic scene diffusion model is iteratively updated by comparing the error between the real future state information and the generated future state information until the error between the two meets the requirements. Then, the corresponding scene diffusion model is used as the trained scene diffusion model.

[0110] Optionally, Figure 4 For another flowchart of model training provided in this embodiment of the invention, please refer to... Figure 3and Figure 4 In step 44, the basic scene diffusion model is iteratively updated using real future state information and generated future state information. Specifically: Step 441, calculate the mean squared error loss based on the real future state information and generated future state information; Step 442, determine whether the basic scene diffusion model has converged based on the mean squared error loss. If it has converged, use the current basic scene diffusion model as the trained scene diffusion model; if it has not converged, update the parameters θ of the basic scene diffusion model and return to the step of inputting the noisy future state information and encoding conditions into the diffusion decoder for iterative decoding.

[0111] For details, please refer to Figure 4 In this embodiment, the convergence of the model is determined based on the mean squared error between the actual future state information and the generated future state information. Therefore, in step 441, the mean squared error loss is first calculated based on the actual future state information and the generated future state information.

[0112] After obtaining the mean squared error loss, in step 442, it is determined whether the current basic scene diffusion model has converged based on the mean squared error loss. For example, when the mean squared error loss changes very little or hardly changes anymore, it can be considered that the model has converged.

[0113] If the model converges, the current base scene diffusion model is used as the trained scene diffusion model. Otherwise, if it does not converge, the model needs to be updated. Updating the base scene diffusion model mainly involves updating its parameters θ. After updating the parameters θ, return to step 43, that is, input the noisy future state information and encoding conditions into the diffusion decoder again for K-step iterative decoding, recalculate the generated future state information, and use it to determine whether the current model has converged by comparing it with the true future state information. Repeat the above steps until the model converges.

[0114] Based on the same inventive concept, this invention also provides a diffusion-based safety-critical scenario generation system. Figure 5 This is a schematic diagram of a generation system provided in an embodiment of the present invention. Please refer to it. Figure 5 The present invention provides a diffusion-based safety critical scenario generation system 100, comprising:

[0115] The first acquisition module 101 is used to acquire multiple historical driving scenario data and save the historical driving scenario data to the storage server;

[0116] The first annotation module 102 is used to annotate historical driving scenario data to obtain annotated historical driving scenario data;

[0117] Module 103 is used to build the basic scene diffusion model;

[0118] Training module 104 is used to train the basic scene diffusion model using labeled historical driving scene data to obtain the trained scene diffusion model.

[0119] The second acquisition module 105 is used to acquire multiple new driving scenario data and save the new driving scenario data to the storage server;

[0120] The first filtering module 106 is used to filter target driving scenario data from the new driving scenario data;

[0121] The second annotation module 107 is used to annotate the target driving scene data to obtain the annotated target driving scene data;

[0122] The second filtering module 108 is used to filter adversarial traffic participants from the labeled target driving scenario data and output the index information of adversarial traffic participants.

[0123] The first generation module 109 is used to generate the guiding conditions, which are expressed as follows: Where, N a This represents the total number of surrounding traffic participants, including 1 opposing traffic participant and N. a -1 non-adversarial traffic participants, where J(i) represents the guiding function of the i-th traffic participant;

[0124] The second generation module 110 is used to generate the future trajectories of surrounding traffic participants by utilizing the trained scene diffusion model, labeled target driving scene data, index information of adversarial traffic participants, and guidance conditions.

[0125] The third generation module 111 is used to generate a safety-critical scenario file based on the target driving scenario data and the future trajectories of surrounding traffic participants;

[0126] Storage module 112 is used to store safety-critical scenario files in a safety-critical scenario file library.

[0127] For details, please refer to Figure 5The diffusion-based safety critical scenario generation system 100 provided in this embodiment of the invention first collects multiple historical driving scenario data through the first acquisition module 101. Since existing safety critical scenario generation methods only support generating safety critical scenarios under certain specific and simple scenario settings, they cannot utilize the large amount of real-world traffic network data and traffic flow information, thus limiting the scale and complexity of safety critical scenario generation. To overcome this problem, this invention collects a large amount of real driving scenario data from different geographical regions and environments. The collected driving scenario data shows diverse behavioral styles of traffic participants, specifically reflected in indicators such as acceleration during driving, relative distance between them, and headway when following other vehicles. For convenient subsequent use, after collecting the driving scenario data, these diverse behavioral style driving scenario data are stored in a storage server.

[0128] Because the data formats of the collected scene data are inconsistent, in order to ensure that the input data format is consistent during model training, this invention uses the first annotation module 102 to perform unified format annotation processing on the historical driving scene data with diverse behavioral styles stored in the storage server. The annotation format can be similar to the format of the Waymo Motion Open Dataset dataset. The content to be annotated in each driving scene includes high-precision map information of the scene, state information of traffic participants (mainly composed of historical state information, current time step state information, and future state information of traffic participants, and the state information of traffic participants at each time step includes position, speed, heading angle, etc.), and some necessary global information in the scene, such as the index of the main vehicle in the scene, the index of the current time step, traffic signals, etc. After annotation, the annotated historical driving scene data is obtained.

[0129] After annotation is completed, the annotated scene data can be used for model training. Model training first requires building a basic scene diffusion model using module 103. Then, training module 104 fully trains the basic scene diffusion model using the annotated historical driving scene data. When the basic scene diffusion model converges, the trained scene diffusion model is obtained.

[0130] After model training is complete, the second acquisition module 105 collects multiple new driving scenario data and saves them to a storage server for use in generating safety-critical scenarios. Then, the first filtering module 106 filters the target driving scenario data from the new driving scenario data to generate the safety-critical scenarios. When filtering target driving scenario data from the new driving scenario data stored on the storage server, filtering can be based on user-defined conditions. These user-defined conditions can include, for example, specific road map types and the overall complexity of the driving scenario. Specific road map types include intersections, roundabouts, and highway ramps, while the overall complexity of the driving scenario includes the number of traffic participants.

[0131] The second annotation module 107 performs annotation processing on the target driving scene data in a unified format to obtain annotated target driving scene data. Specific annotation content can be found in the annotation instructions for historical driving scene data, and will not be repeated here. The annotated target driving scene data includes second high-precision map information and second traffic participant status information, where traffic participants include the main vehicle and surrounding traffic participants.

[0132] After obtaining the labeled target driving scenario data, the second filtering module 108 selects a scenario from the labeled target driving scenario data in sequence. Then, using the second high-precision map information and the status information of the second traffic participants (the main vehicle and surrounding traffic participants) in the scenario at the current time step, it quantitatively evaluates the level of risk posed by the surrounding traffic participants to the main vehicle. Based on the evaluation results, it selects the surrounding traffic participant with the highest risk as the adversary traffic participant and then outputs the index information of the adversary traffic participant in the scenario.

[0133] To improve the realism and effectiveness of the generated safety-critical scenarios, this invention utilizes the first generation module 109 to generate guiding conditions. In the process of denoising the trajectories of surrounding traffic participants using the trained scenario diffusion model, guiding conditions are added to guide the trained scenario diffusion model to finally generate the desired safety-critical trajectories of surrounding traffic participants.

[0134] The guiding condition is expressed as Where, N a This represents the total number of surrounding traffic participants, including 1 opposing traffic participant and N. a -1 non-adversarial traffic participants, where J(i) represents the guiding function of the i-th traffic participant.

[0135] For adversarial traffic participants, the guiding conditions include adversarial guiding conditions and natural constraint guiding conditions; therefore, its guiding function is expressed as J. adv =J col +Jreg For other surrounding traffic participants, i.e., non-adversarial traffic participants, the guiding condition only needs to include the natural constraint guiding condition; therefore, its guiding function is expressed as J. non-adv =J reg .

[0136] Among them, J col For collision guidance function, collision guidance function The distance between the future trajectories of adversarial traffic participants generated by the trained scene diffusion model and the actual future trajectories of the main vehicle in the scene is measured; T represents the total number of future time steps; d(t) represents the distance between the trajectory points of adversarial traffic participants generated by the trained scene diffusion model and the actual trajectory points of the main vehicle in the scene at future time step t.

[0137] J reg Natural constraint guiding function, natural constraint guiding function This represents the difference between the future trajectory of a traffic participant generated by the trained scene diffusion model and the actual future trajectory of that traffic participant; d represents the distance between the trajectory points of the traffic participant generated by the trained scene diffusion model and the actual trajectory points of that traffic participant at future time step t. thr This indicates the set distance threshold.

[0138] It should be noted that the true future trajectory of the main vehicle, the true trajectory point of the main vehicle, the true future trajectory of the traffic participants, and the true trajectory point of the traffic participants can all be obtained through the status information of the second traffic participants.

[0139] After obtaining the guiding conditions, the second generation module 110 can generate the future trajectories critical to the safety of surrounding traffic participants for each target driving scenario. Scenarios are selected sequentially from the labeled target driving scenario data. The second high-precision map information corresponding to the current scenario is input into the map encoder, and the state information of the second traffic participants is input into the traffic participant encoder. The map embedding and traffic participant embedding are extracted using the map encoder and traffic participant encoder respectively. The map embedding and traffic participant embedding together form the encoding conditions.

[0140] The system obtains the real future state information of the surrounding traffic participants corresponding to the current scene, and adds noise to it to obtain the noisy future trajectory of the surrounding traffic participants. Then, the noisy future trajectory of the surrounding traffic participants, the encoding conditions, and the index information of the adversarial traffic participants corresponding to the current scene are input into the trained scene diffusion model, and K-step iterative denoising is performed on it.

[0141] Please refer to Figure 2 In each iteration of denoising, the gradient information of the guiding function is first calculated through a query process, and then this gradient information is backpropagated through a guided optimization process to optimize the generated future trajectory. After completing the K-step iterative denoising process, the future trajectories of surrounding traffic participants can be obtained, thus generating highly natural, diverse, and safety-critical future trajectories of surrounding traffic participants.

[0142] After obtaining the future trajectories of surrounding traffic participants, the third generation module 111 generates a safety-critical scenario file based on the target driving scenario data and the future trajectories of surrounding traffic participants. This safety-critical scenario file contains high-precision map information of the scenario and trajectory information of all traffic participants. The entire trajectory of the main vehicle (including historical trajectory, current time step trajectory, and future trajectory) follows the trajectory of the main vehicle in the target driving scenario data. The historical and current time step trajectories of surrounding traffic participants are also still provided by the target driving scenario data, but their future trajectories are replaced with new future trajectories of surrounding traffic participants generated by the trained scenario diffusion model. Finally, these trajectory data are stitched together to form complete traffic participant trajectories and integrated with the high-precision map information to obtain the final safety-critical scenario file.

[0143] To facilitate subsequent testing of the autonomous driving system, the present invention also stores safety-critical scenario files in a safety-critical scenario file library through storage module 112.

[0144] The safety-critical scenario generation system based on diffusion provided by this invention utilizes a large amount of real traffic network data and traffic flow information as the basis for generation. Therefore, this method significantly expands the scale of safety-critical scenario generation and can generate complex safety-critical scenarios. By training a multi-style scenario diffusion model using real driving scenario data with diverse behavioral styles, this model can randomly generate new trajectories with diverse behavioral styles of surrounding traffic participants. These trajectory data have a high degree of naturalness, thereby fundamentally improving the diversity and realism of the generated safety-critical scenarios.

[0145] Furthermore, this invention extracts adversarial traffic participants from the target scene used to generate safety-critical scenarios, then generates guidance conditions for both adversarial and non-adversarial traffic participants, and incorporates guidance conditions during the denoising process of the target scene to generate highly natural, diverse, and safety-critical future trajectories of surrounding traffic participants, thereby improving generation efficiency and the realism and effectiveness of the generated scenarios.

[0146] Optionally, the labeled historical driving scenario data includes the first high-precision map information and the state information of the first traffic participant. The state information of the first traffic participant includes historical state information and real future state information. The basic scenario diffusion model includes a scenario encoder and a diffusion decoder. Figure 6 This is a schematic diagram of a training module 104 provided in an embodiment of the present invention. Please refer to... Figure 5 and Figure 6 The training module 104 includes: a first input module 1041, used to input the first high-precision map information and historical state information into the scene encoder for encoding to obtain encoding conditions; a noise-adding module 1042, used to add noise to the real future state information to obtain noisy future state information; a second input module 1043, used to input the noisy future state information and encoding conditions into the diffusion decoder for iterative decoding to obtain generated future state information; and an update module 1044, used to iteratively update the basic scene diffusion model using the real future state information and the generated future state information to obtain the trained scene diffusion model.

[0147] Specifically, the labeled historical driving scenario data includes first-level high-precision map information and first-level traffic participant status information. The first-level traffic participant status information includes historical status information and current future status information. The basic scenario diffusion model includes a scenario encoder and a diffusion decoder.

[0148] Please refer to Figure 6 When training the basic scene diffusion model, the scene encoder includes a map encoder and a traffic participant encoder because it is necessary to encode the first high-precision map information and the historical state information of the first traffic participant separately. First, the first high-precision map information is input into the map encoder through the first input module 1041 to obtain the map code, and the historical state information is input into the traffic participant encoder to obtain the historical state code.

[0149] Then, noise is sampled by the noise-adding module 1042, and the noise is used to add noise to the real future state information to obtain noisy future state information. The second input module 1043 inputs the noisy future state information and the encoding conditions into the diffusion decoder for iterative decoding to obtain the generated future state information. Here, the iterative decoding process is K steps.

[0150] After obtaining the generated future state information, the update module 1044 iteratively updates the basic scene diffusion model by comparing the error between the real future state information and the generated future state information until the error between the two meets the requirements, and then uses the corresponding scene diffusion model as the trained scene diffusion model.

[0151] Alternatively, please refer to Figure 5The update module 1044 is specifically used for: calculating the mean squared error loss based on the real future state information and the generated future state information; determining whether the basic scene diffusion model has converged based on the mean squared error loss; if it has converged, using the current basic scene diffusion model as the trained scene diffusion model; if it has not converged, updating the parameters θ of the basic scene diffusion model, and returning the iterative decoding steps of inputting the noisy future state information and encoding conditions into the diffusion decoder.

[0152] For details, please refer to Figure 5 In this embodiment, the convergence of the model is determined based on the mean squared error between the actual future state information and the generated future state information. Therefore, the mean squared error loss is first calculated based on the actual future state information and the generated future state information.

[0153] After obtaining the mean squared error loss, we can determine whether the current basic scenario diffusion model has converged based on the mean squared error loss. For example, when the mean squared error loss changes very little or hardly changes anymore, it can be considered that the model has converged.

[0154] If the model converges, the current base scene diffusion model is used as the trained scene diffusion model. Otherwise, if it does not converge, the model needs to be updated. Updating the base scene diffusion model mainly involves updating its parameters θ. After updating the parameters θ, the noisy future state information and encoding conditions are input into the diffusion decoder again for K-step iterative decoding. The generated future state information is recalculated and compared with the true future state information to determine whether the current model has converged. The above steps are repeated until the model converges.

[0155] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of one embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing the present invention.

[0156] Those skilled in the art will understand that the modules in the apparatus of the embodiments can be distributed in the apparatus of the embodiments as described in the embodiments, or they can be located in one or more devices different from this embodiment with corresponding changes. The modules of the above embodiments can be combined into one module, or they can be further divided into multiple sub-modules.

[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A diffusion-based method for generating safety-critical scenarios, characterized in that, include: Collect data from multiple historical driving scenarios and save the historical driving scenario data to a storage server; The historical driving scenario data is labeled to obtain labeled historical driving scenario data; Construct a basic scenario diffusion model; The basic scene diffusion model is trained using labeled historical driving scene data to obtain the trained scene diffusion model. Collect data from multiple new driving scenarios and save the new driving scenario data to a storage server; Filter target driving scenario data from the new driving scenario data; The target driving scenario data is labeled to obtain labeled target driving scenario data; Filter adversarial traffic participants from the labeled target driving scenario data and output the index information of the adversarial traffic participants; Generate guiding conditions, wherein the guiding conditions are expressed as follows: Where, N a This represents the total number of surrounding traffic participants, including 1 opposing traffic participant and N. a -1 non-adversarial traffic participants, where J(i) represents the guiding function of the i-th traffic participant; Using the trained scene diffusion model, the labeled target driving scene data, the index information of the adversarial traffic participants, and the guidance conditions, the future trajectories of surrounding traffic participants are generated; Based on the target driving scenario data and the future trajectories of the surrounding traffic participants, a safety-critical scenario file is generated; The safety-critical scenario files are stored in the safety-critical scenario file library.

2. The method for generating security-critical scenarios based on diffusion according to claim 1, characterized in that, The labeled historical driving scenario data includes first high-precision map information and first traffic participant status information, the first traffic participant status information includes historical status information and real future status information; the basic scenario diffusion model includes a scenario encoder and a diffusion decoder. The basic scene diffusion model is trained using labeled historical driving scene data, specifically as follows: The first high-precision map information and the historical state information are input into the scene encoder for encoding to obtain the encoding conditions; The real future state information is subjected to noise processing to obtain noisy future state information; The noisy future state information and the encoding conditions are input into the diffusion decoder for iterative decoding to obtain the generated future state information. The base scene diffusion model is iteratively updated using the real future state information and the generated future state information to obtain the trained scene diffusion model.

3. The method for generating security-critical scenarios based on diffusion according to claim 2, characterized in that, The basic scene diffusion model is iteratively updated using the real future state information and the generated future state information, specifically as follows: Calculate the mean squared error loss based on the actual future state information and the generated future state information; The convergence of the basic scene diffusion model is determined based on the mean squared error loss. If convergence is achieved, the current basic scene diffusion model is used as the trained scene diffusion model. If convergence fails, the parameters θ of the basic scene diffusion model are updated, and the iterative decoding step of inputting the noisy future state information and the encoding conditions into the diffusion decoder is returned.

4. The method for generating security-critical scenarios based on diffusion according to claim 1, characterized in that, Filtering target driving scenario data from the new driving scenario data specifically involves: Target driving scenario data is filtered from the new driving scenario data stored on the storage server according to predetermined conditions; the predetermined conditions include at least a specific road map type and the overall complexity of the driving scenario.

5. The method for generating security-critical scenarios based on diffusion according to claim 1, characterized in that, The labeled target driving scenario data includes second high-precision map information and second traffic participant status information; The process of filtering adversarial traffic participants from the labeled target driving scenario data is as follows: Based on the second high-precision map information and the status information of the second traffic participant, the surrounding traffic participant with the highest risk to the main vehicle is selected as the opposing traffic participant.

6. The method for generating security-critical scenarios based on diffusion according to claim 5, characterized in that, For adversarial traffic participants, its guided function is represented as J. adv =J col +J reg For non-confrontational traffic participants, the guided function is represented as J. non-adv =J reg J col J is the collision guidance function. reg For natural constraint guiding functions; Among them, the collision guidance function The distance between the future trajectories of adversarial traffic participants generated by the trained scene diffusion model and the actual future trajectories of the main vehicle in the scene is measured; T represents the total number of future time steps; d(t) represents the distance between the trajectory points of the adversarial traffic participants generated by the trained scene diffusion model and the actual trajectory points of the main vehicle in the scene at future time step t; natural constraint guidance function. This represents the difference between the future trajectory of a traffic participant generated by the trained scene diffusion model and the actual future trajectory of that traffic participant; d represents the distance between the trajectory points of the traffic participant generated by the trained scene diffusion model and the actual trajectory points of that traffic participant at future time step t. thr This indicates the set distance threshold.

7. The method for generating security-critical scenarios based on diffusion according to claim 6, characterized in that, Using the trained scene diffusion model, the labeled target driving scene data, the index information of the adversarial traffic participants, and the guidance conditions, the future trajectories of surrounding traffic participants are generated, specifically as follows: The noisy future trajectories of surrounding traffic participants and the index information of the adversarial traffic participants are input into the trained scene diffusion model, and K-step iterative denoising is performed; in each iterative denoising process, the guiding conditions are used to optimize the results.

8. A diffusion-based safety-critical scenario generation system, characterized in that, include: The first acquisition module is used to acquire multiple historical driving scenario data and save the historical driving scenario data to a storage server; The first annotation module is used to annotate the historical driving scenario data to obtain annotated historical driving scenario data; Modules are used to build basic scene diffusion models; The training module is used to train the basic scene diffusion model using labeled historical driving scene data to obtain the trained scene diffusion model. The second acquisition module is used to acquire multiple new driving scenario data and save the new driving scenario data to the storage server; The first filtering module is used to filter target driving scenario data from the new driving scenario data; The second annotation module is used to annotate the target driving scene data to obtain annotated target driving scene data; The second filtering module is used to filter adversarial traffic participants from the labeled target driving scenario data and output the index information of the adversarial traffic participants; The first generation module is used to generate boot conditions, wherein the boot conditions are expressed as follows: Where, N a This represents the total number of surrounding traffic participants, including 1 opposing traffic participant and N. a -1 non-adversarial traffic participants, where J(i) represents the guiding function of the i-th traffic participant; The second generation module is used to generate the future trajectories of surrounding traffic participants by utilizing the trained scene diffusion model, the labeled target driving scene data, the index information of the adversarial traffic participants, and the guidance conditions. The third generation module is used to generate a safety-critical scenario file based on the target driving scenario data and the future trajectories of the surrounding traffic participants; The storage module is used to store the safety-critical scenario files in the safety-critical scenario file library.

9. The diffusion-based safety-critical scenario generation system according to claim 8, characterized in that, The labeled historical driving scenario data includes the first high-precision map information and the status information of the first traffic participant. The status information of the first traffic participant includes historical status information and real future status information. The basic scene diffusion model includes a scene encoder and a diffusion decoder; The training module includes: The first input module is used to input the first high-precision map information and the historical state information into the scene encoder for encoding to obtain encoding conditions; A noise-adding module is used to add noise to the real future state information to obtain noisy future state information. The second input module is used to input the noisy future state information and the encoding conditions into the diffusion decoder for iterative decoding to obtain the generated future state information. The update module is used to iteratively update the basic scene diffusion model using the real future state information and the generated future state information to obtain the trained scene diffusion model.

10. The method for generating security-critical scenarios based on diffusion according to claim 9, characterized in that, The update module is specifically used for: Calculate the mean squared error loss based on the actual future state information and the generated future state information; The convergence of the basic scene diffusion model is determined based on the mean squared error loss. If convergence is achieved, the current basic scene diffusion model is used as the trained scene diffusion model. If convergence fails, the parameters θ of the basic scene diffusion model are updated, and the iterative decoding step of inputting the noisy future state information and the encoding conditions into the diffusion decoder is returned.