Automatic driving critical scene closed-loop generation method based on large language model
By generating highly threatening background vehicle trajectories through large language models and reinforcement learning algorithms and constructing closed-loop adversarial training scenarios, the test efficiency and robustness issues of autonomous driving systems in extreme scenarios are resolved, and efficient and reliable critical scenario generation and testing are achieved.
Patent Information
- Application Number
- CN202510640082.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-10-17
Smart Images

Figure CN120805634A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of automatic driving, in particular to an automatic driving critical scene closed-loop generation method based on a large language model. BACKGROUND
[0002] With the rapid development of automatic driving technology, ensuring the safety of the system in extreme critical scenarios has become a core challenge for practical deployment. However, such scenarios occur extremely rarely in the real world, resulting in a lack of training and testing data, making it difficult for traditional methods to fully cover potential risks. Existing technologies mainly fall into two categories: one is a generation method based on parameter combination, which quickly generates scenarios by enumerating static parameter combinations of traffic elements, but it lacks modeling of dynamic interaction behavior, making it difficult to capture complex spatiotemporal correlations between vehicles; the second is an adversarial scene generation method based on behavior disturbance, which introduces adversarial background vehicle behavior to interfere with the vehicle under test, but this method relies on high-complexity models (such as deep reinforcement learning), which is computationally expensive, and the generated scenarios are often disconnected from the training process of the automatic driving system, making it difficult to form a closed-loop optimization. In addition, existing methods have limitations in scene semantic understanding and threat target screening, and cannot effectively identify high-threat background vehicles and their behavior evolution patterns, resulting in insufficiently targeted adversarial scenarios and low testing efficiency. Therefore, a new automatic driving critical scene closed-loop generation method based on a large language model has become a problem that needs to be solved by those skilled in the art. SUMMARY
[0003] The present application provides an automatic driving critical scene closed-loop generation method based on a large language model, comprising:
[0004] Step S1, environment state extraction: based on an automatic driving simulation platform, extract the environment state data of the current scene, including road geometry, ego vehicle motion trajectory, historical behavior trajectory of background vehicles, and dynamic scene semantic labels;
[0005] Step S2, generate high-threat background vehicle identification prompts: based on the environment state data in step S1, construct structured prompt information, input it into the large language model, generate an initial threat identification function containing behavior association and potential conflict relationships, describe the behavior association, spatial layout, and potential conflict relationships between traffic participants, and improve the model's perception ability of scene semantics;
[0006] Step S3, large language model scene reasoning and threat vehicle automatic iteration screening: use the large language model to perform multiple rounds of iteration optimization on the initial identification function, combine simulation evaluation and feedback mechanisms, and screen out a set of high-threat background vehicles;
[0007] Step S4, attack trajectory closed-loop optimization and generation: based on the screened high-threat background vehicles, generate attack trajectories based on reinforcement learning algorithm to maximize the collision probability, form a closed-loop adversarial training;
[0008] Step S5, adversarial scene construction: fuse the attack trajectory with the original scene to construct a physically executable and semantically consistent adversarial traffic interaction scene for automatic driving system testing and training.
[0009] Optionally, in the step S1:
[0010] The road geometry structure includes road network topology information and three-dimensional road surface characteristics;
[0011] The ego vehicle motion trajectory includes real-time position, speed, heading angle, and historical trajectory sequence;
[0012] The historical behavior trajectory of the background vehicle includes past five seconds of historical trajectory and future prediction window data;
[0013] The dynamic scene semantic label includes scene type, weather condition, and traffic rule constraint;
[0014] The data is output in JSON or Protobuf format through the API interface of the simulation platform, and a traffic scene knowledge graph is constructed.
[0015] Optionally, in the step S2:
[0016] The structured prompt information is generated through a chain thinking prompt strategy to guide the large language model to infer the threat score function in stages;
[0017] The expression form of the threat score function is:
[0018]
[0019] Where f j (j) represents the jth sub-feature function (such as longitudinal approach speed, lane change intention, intersection conflict probability, etc.), w j is the corresponding weight parameter, Δx i (t), Δv i (t) is the relative position and relative speed to the ego vehicle, and θ i represents the vehicle orientation.
[0020] Optionally, the step S3 includes:
[0021] An initial function construction module: generates an executable threat recognition function, and the input parameters include the ego vehicle and background vehicle states;
[0022] An evaluation feedback module generates optimization suggestions based on the simulation attack success rate analysis function defects;
[0023] A behavior logic evolution module adjusts function parameters or structure according to optimization suggestions, and selects the optimal function using a multi-candidate evaluation mechanism;
[0024] The iterative process meets the following conditions:
[0025] A. If the attack success rate of the candidate function improves, update to the next round;
[0026] B. If the attack success rate reaches the threshold, only fine-tune the weight parameters;
[0027] C. When there is no gain for several rounds in a row, allow adjustment of the function structure.
[0028] Optionally, in step S4:
[0029] Set a reward function based on reinforcement learning algorithm, aiming to maximize the collision probability;
[0030] Generate an attack trajectory using a behavior disturbance strategy;
[0031] According to the performance of the trained autonomous driving system, continuously update the scene to form a closed-loop optimization.
[0032] Optionally, in step S5:
[0033] The construction of the adversarial scene meets the trajectory physical executability constraint:
[0034]
[0035] Where, represents a function that integrates road networks and vehicle trajectories into a function that inserts attacker trajectories into the simulation platform according to the specified timestamp, and controls the interface to implement behavior intervention. To ensure the physical executability of the adversarial trajectory and the consistency of the road semantics, the following constraint function is designed to effectively filter the generated trajectory:
[0036] D(Y att ,W)≤
[0037] Where D(·) represents the minimum distance function between the attack vehicle trajectory and the road center line or drivable area, and is the maximum deviation threshold that can be tolerated, ensuring that the attack trajectory does not deviate from the legal driving area.
[0038] Optionally, the chain thinking prompt strategy includes the following stages:
[0039] First stage, determine the core influencing factors of threat assessment;
[0040] The second stage is reasoning about the evolution pattern of the influencing factors over time.
[0041] The third stage is constructing a threat assessment index system by weighted normalization.
[0042] Optionally, the multiple candidate evaluation mechanism specifically comprises:
[0043] N candidate functions are generated in each iteration;
[0044] The attack success rate is calculated through small-scale simulation testing, and the optimal candidate function is selected for subsequent iteration.
[0045] Optionally, the trajectory prediction model is DenseTNT, which is used to generate multiple high-probability trajectory candidates of the background vehicle and calculate the joint collision probability with the ego vehicle trajectory.
[0046] The application also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor implements the automatic driving critical scene closed-loop generation method according to any one of claims 1-9 when executing the program.
[0047] The application has the following beneficial effects:
[0048] 1. The dynamic threat reasoning and iterative optimization mechanism based on the large language model provided by the application significantly improves the efficiency and pertinence of critical scene generation: by constructing a traffic scene knowledge graph and a chain thinking prompt strategy, the spatio-temporal correlation of vehicle behavior can be deeply analyzed, and high-threat background vehicles can be accurately identified; combined with multiple rounds of simulation feedback and function evolution module, the threat assessment model is continuously optimized, the most attack-potential target is quickly selected under limited data conditions, the blindness of traditional parameter enumeration method is avoided, and the test redundancy is greatly reduced.
[0049] 2. The closed-loop adversarial training framework driven by reinforcement learning provided by the application effectively enhances the dynamic adaptability and system robustness of the test scene: by combining attack trajectory generation and automatic driving strategy updating, an adversarial optimization cycle with real-time feedback is formed, so that the generated adversarial scene can adapt to the evolution of the defense mechanism of the system under test, breaking through the limitation that the scene generation and training process are disconnected in existing methods, and systematically improving the anti-interference ability of the automatic driving system in complex interactive environment.
[0050] 3. The physical constraint and semantic consistency guarantee mechanism provided by the application ensures the authenticity and executability of the adversarial scene: hard constraint conditions such as road boundaries and driving rules are introduced in the trajectory optimization stage, and the precise synchronization of trajectory and environment is realized through the simulation platform interface, avoiding the generation of out-of-bound or irrational driving behaviors that deviate from reality, ensuring the effectiveness and reusability of the test results, and providing high-confidence key scene data for safety verification. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a framework diagram of a closed-loop generation method for critical scenarios in autonomous driving based on a large language model in an embodiment of the present application;
[0052] Figure 2 This is a flowchart of an implementation method for closed-loop generation of critical scenarios for autonomous driving based on a large language model according to an embodiment of the present application;
[0053] Figure 3 This is a diagram of an automatic iterative screening method based on a large language model in step S3 of an embodiment of the present application;
[0054] Figure 4 This is a schematic diagram of screening high-threat background vehicles in an example scene according to the method of step S3 in an embodiment of the present application;
[0055] Figure 5 This is a schematic diagram of the implementation process of a closed-loop generation method for critical scenarios in autonomous driving based on a large language model in an embodiment of the present application. DETAILED DESCRIPTION
[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0057] The present invention provides a closed-loop generation method for critical scenarios in autonomous driving based on a large language model. By introducing the scene semantic understanding and contextual reasoning capabilities of the large language model, it realizes the automatic identification and trajectory generation of highly aggressive threat vehicles in adversarial traffic scenarios, and then effectively generates critical scenarios covering complex interactions, forming a closed-loop iterative testing framework for test scenarios and autonomous driving system training processes, thereby improving the efficiency of evaluating the stability and safety of autonomous driving systems at a low computing cost.
[0058] The following steps are involved:
[0059] Step S1, environmental state extraction: extracting environmental state data of the current scene based on the autonomous driving simulation platform, including road geometry, ego vehicle motion trajectory, historical behavior trajectory of background vehicles, and dynamic scene semantic labels;
[0060] Step S2, generating high-threat background vehicle identification prompt: according to the environmental state data in step S1, constructing structured prompt information, inputting into a large language model, generating an initial threat identification function containing behavior association and potential conflict relationship, describing the behavior association, spatial layout and potential conflict relationship between traffic participants, and improving the model's perception ability of scene semantics;
[0061] Step S3, large language model scene reasoning and automatic iterative screening of threat vehicles: using a large language model to perform multiple rounds of iterative optimization on the initial identification function, combining simulation evaluation and feedback mechanism to screen out a set of high-threat background vehicles;
[0062] Step S4, attack trajectory closed-loop optimization and generation: based on the screened high-threat background vehicles, generating an aggressive trajectory based on a reinforcement learning algorithm to maximize the collision probability, forming a closed-loop adversarial training;
[0063] Step S5, construction of adversarial scene: fusing the aggressive trajectory with the original scene to construct a physically executable and semantically consistent adversarial traffic interaction scene for automatic driving system testing and training.
[0064] With the rapid development of automatic driving technology, ensuring the safety of automatic driving systems in critical key scenarios has become a key to their actual deployment. These scenarios are extremely rare in the real world, and limited data makes it difficult to support effective training and testing. Existing methods can be mainly divided into two categories: parameter combination-based rapid generation methods and behavior disturbance-based adversarial scene generation methods. The former fails to fully capture dynamic behavior, and the latter relies on complex models, resulting in high computational cost and difficulty in integrating with the training process.
[0065] As shown in Figure 1 , the embodiment of the present application provides a closed-loop generation method for critical scenarios of automatic driving based on a large language model. In this embodiment, an open automatic driving simulation platform is used as an environment support to obtain all vehicle trajectories and road information in the scene in real time. The embodiment of the present application is divided into five steps, and the implementation process is shown in Figure 2 .
[0066] Step S1, environment state extraction: extracting the environmental state data of the current scene based on the automatic driving simulation platform, including road geometry, ego motion trajectory, historical behavior trajectory of background vehicles and dynamic scene semantic label, for subsequent semantic modeling and reasoning analysis:
[0067] The automatic driving simulation platform extracts the environmental state data of the current scene, including the following basic information:
[0068] The extraction method of road geometry is to extract the road network topology information (such as the number of lanes, intersection coordinates, traffic signal position) and three-dimensional road surface characteristics (such as slope and curvature);
[0069] The extraction method of the ego vehicle motion trajectory is to obtain the real-time position, speed, heading angle and historical trajectory sequence (time step can be configured, default 0.1 seconds / frame) of the ego vehicle;
[0070] The extraction method of background historical behavior trajectory is to collect the position, speed and acceleration data of the background vehicle, and build a time series data set (including past 5 seconds of historical trajectory and future prediction window);
[0071] Dynamic scene semantic label: synchronously record the scene type (such as crossroad, highway merging area), weather condition (such as rainy day, foggy day) and traffic rule constraint (such as speed limit sign);
[0072] The above data is output in a structured format (JSON / Protobuf) through the API interface of the simulation data platform, and is input to the prompt generation module to build a traffic scene knowledge graph containing space-time association, supporting the interaction relationship reasoning and attacker identification logic generation of the subsequent LLM agent.
[0073] Step S2, generating high-threat background vehicle identification prompt: based on the traffic scene knowledge graph extracted in step S1, structured prompt information is constructed to input into a large language model to describe the behavior association, spatial layout and potential conflict relationship between traffic participants, and to improve the model's semantic perception ability in the scene;
[0074] The prompt generation module is used to generate an initial identification function code with executable function, and the input of the prompt generation module includes the state information of the ego vehicle and the background vehicle, as well as the task description and function requirement prompt content, wherein the module structure is as follows:
[0075] P0=LLMInit(T,C req )
[0076] Wherein, P0 represents the function prototype code output by the module, LLMInit(·) is a function generation operation, T represents a task description, C req represents the specific requirements and rule set of the function logic.
[0077] In order to improve the semantic reasoning ability of the large prophetic model agent in complex traffic scenes, a chain thinking prompt strategy is introduced, which guides the model to perform multi-step reasoning and gradually understand the potential association between participant behaviors. Taking threat score function T i as an example, the reasoning structure of the following form is constructed for each background vehicle i:
[0078]
[0079] where f j denotes the jth sub-feature function (such as longitudinal closing speed, lane change intention, intersection conflict probability, etc.), w j is the corresponding weight parameter, Δx i (t), Δv i (t) is the relative position and relative speed with the ego vehicle, and θ i represents the vehicle orientation.
[0080] In the task prompt, the function generation process is explicitly divided into the following three stages: identifying the core influencing factors that should be focused on in the function (such as the distance between the target vehicle and the ego vehicle, the speed difference, and the trajectory intersection probability, etc.), reasoning the patterns of evolution of these factors over time and their influence on the threat level, and weighting and normalizing all influencing factors to construct a threat assessment index system. Chain thinking makes the model reasoning chain clear and controllable, with stronger explainability and generalization ability.
[0081] Step S3, large language model scenario reasoning and automatic iteration screening of threat vehicles: using a multi-agent large language model to perform semantic reasoning on the above prompt information, evaluating the possibility of each background vehicle causing interference or potential conflict to the measured autonomous vehicle, automatically iterating, and selecting vehicles with high attack potential as threat vehicle candidates;
[0082] The embodiment designs an automatic iteration screening process based on a large language model composed of an initial function construction module, an evaluation feedback module, and a behavior logic evolution module in the step S3, as shown in Figure 3 .
[0083] First, the initial function construction module generates an executable initial identification function through a large language model, which is used to evaluate the possibility of background vehicles causing interference to the measured vehicle. The initial function construction module includes a role setting unit, a task target unit, a background information unit, an input parameter unit, a function requirement unit, and an output format unit, wherein the input parameter unit is composed of two groups of data: ego vehicle state parameters and background vehicle state parameters.
[0084] Secondly, based on the function code generated in step S2, the identification function is executed in the scenario simulation platform, and the identified threat vehicle will trigger the attack behavior and evaluate the attack success rate. During the simulation process, depending on whether the intervention causes the mission failure of the tested vehicle (such as collision, deviation from the route, etc.), the automatic iterative screening process based on the large language model in step 3 needs to conduct targeted review and optimization of the function logic to improve its ability to judge threat targets. To this end, the automatic iterative screening process based on the large language model in step 3 sets an evaluation feedback module. The main function of the evaluation feedback module is to identify the deficiencies in the existing function execution results through the large language model, input improvement suggestions accordingly, and guide the next round of function adjustments.
[0085] Before running, the evaluation and feedback module receives the following information as input: function role settings, environment and task context, the code of the current function version, the attack success rate evaluation value of the function in the scenario simulation, and the structural requirements and output specifications used for analysis. Based on this, the module uses the large language model to perform logical backtracking and comparative analysis to generate optimization suggestions. This process can be formally expressed as follows:
[0086] S i =LLMRefl(T,C req ,M,P i-1 ,ε i-1 )
[0087] ε i =Sim(P i )
[0088] Among them S i represents the modification suggestion output by the evaluation module in the i-th iteration, P i-1 ,ε i-1 are the function code of the previous iteration and its attack success rate evaluation results. M is a memory module used to record the key feedback information of previous iterations, including the current function and its corresponding attack success rate. The simulation module Sim(·) is responsible for i Calculate its attack performance in the simulation environment and output the new performance result ε i .
[0089] During the iterative process, the evaluation and feedback module compares historical functions and performance trends recorded in M to refine effective strategy adjustments. The attack success rate serves as the primary performance metric, measuring the effectiveness of modifications in each round. This continuously drives the evolution of the recognition function towards greater efficiency and attack capabilities, ultimately forming a targeted and generalizable adversarial recognition strategy.
[0090] The behavior logic evolution module updates and rewrites the current threat identification function at the structure or parameter level using a large language model according to the optimization suggestions provided by the evaluation feedback module to obtain better threat target screening capability. To achieve this goal, the behavior logic evolution module receives prompt instructions containing functional role definitions, function input items (such as vehicle or surrounding background vehicle state information), the current version of the identification function, adjustment suggestions, output requirements, and their formats, and drives the large language model to generate an improved function with higher attack effect. This process can be formally described as follows:
[0091] P i+1 = LLMUpdate (T, C req , P i , S i )
[0092] where P i is the function code generated in the i-th iteration, and S i represents the corresponding evaluation feedback suggestions. Even if the input is the same, the large language model may produce outputs with large differences in content and quality. Considering the randomness of the large language model in generating results, this module uses a multi-candidate evaluation mechanism to stabilize the output quality. That is, in the i-th iteration, N different versions of the function candidate are generated by multiple calls, and the optimal candidate solution is selected by performing rapid verification in a simulated environment. The performance of each candidate function can be represented by the following formula:
[0093]
[0094] EvalSim (·) represents a small-scale simulation test function.
[0095] If a candidate function performs better than the current version in simulation, i.e., the success attack rate is improved, then the candidate function is used for the next round of update; if it is not significantly better than the existing version and the number of candidates has reached the upper limit N, then the best one is retained to ensure stability. When the attack performance has reached the set threshold, to avoid fluctuations caused by large modifications, this module limits the adjustment strategy to slight adjustments, such as only modifying the weight parameters without changing the function structure; only in the case where continuous fine-tuning has no significant gain, structural-level modification is allowed to ensure exploration of better solution space while controlling risk.
[0096] Finally, after Q iterations of iterative updates, a stable and high-performance identification function P Q is obtained, which can be used to automatically screen a high-threat attacker set M from the original background vehicle set B, expressed as:
[0097]
[0098] In an example scenario, the iterative screening process is as shown in FIG. 3. The process of step S3 ensures the robustness and scenario adaptability of the screening function under complex dynamic traffic environment, providing a high-quality input basis for the subsequent generation of high-threat background vehicle trajectories. Figure 4
[0099] Step S4, attack trajectory closed-loop optimization and generation: based on the identified threat vehicles, a reward function is set based on the reinforcement learning algorithm to maximize the collision probability, and an attack trajectory is generated using a behavior disturbance strategy to enhance the criticality of the adversarial test. The scenario in the above process is continuously updated according to the performance of the trained autonomous driving system, forming a closed loop;
[0100] To systematically construct an efficient adversarial scenario, based on the set M of M high-threat background vehicles identified in step S3, a simulation environment driven by reinforcement learning and a trajectory optimization mechanism driven by collision probability are combined to continuously construct a set of threat background vehicle trajectories that can cause the greatest disturbance to the tested autonomous driving system, achieving closed-loop optimization and generation of attack behavior.
[0101] To systematically model the adversarial target behavior, the autonomous driving task is formalized as a Markov decision process defined by the four-tuple . Among them, the state space S is composed of the state of the tested autonomous vehicle, the state of the high-threat background vehicle, etc.; the action space U contains the acceleration and steering control commands of the ego vehicle; the reward function is used to comprehensively measure the task completion degree, path rationality and safety factor; the state transition function describes the dynamics of the environment. For the tested vehicle, the learning goal is to maximize the cumulative expected return through the strategy π:
[0102]
[0103] To evaluate the robustness of the strategy, an adversarial disturbance δ is introduced to disturb the task execution of the tested autonomous vehicle by constructing an attack background vehicle behavior, and the goal is to minimize the expected return:
[0104]
[0105] On this basis, the system iteratively updates π through an adversarial training mechanism to improve its robustness:
[0106]
[0107] Specific to a certain traffic scenario, assume that the attack is triggered from time k, and the current state is given, where W represents the road structure, is the historical trajectory of the vehicle, are the trajectories of M background vehicles. The trajectories of the future tested vehicle and M attackers are The attack goal can be transformed into maximizing the following joint posterior probability P(Y att ,Y ego |X,C):
[0108]
[0109] Where C represents a collision event.
[0110] This process assumes that the future states of the vehicles are independent of each other, and that the collision event depends only on the trajectories of the ego vehicle and the attacker. Therefore, the collision probability maximization objective can be expanded based on the Bayesian formula to obtain the following computable form:
[0111]
[0112] in, Denotes the future trajectory of the i-th attacker. Trajectory prediction is provided by the pre-trained model DenseTNT, which generates several high-probability trajectory candidates for each threat vehicle. The trajectory of the tested vehicle is generated based on the current strategy to generate the predicted trajectory. Combining the two, further evaluation And select the combination with the highest joint collision probability to construct the attack trajectory.
[0113] This process is carried out in a closed-loop learning framework. First, according to the recognition function P generated in step S3, Q , the most threatening M adversarial vehicles are screened from the normal background vehicle set. Then, the framework enters the training phase: each round samples a sample from the normal scene, and uses the prediction model to calculate and Y ego , and the final attack trajectory is determined based on the collision probability. The attack simulation results are used to update the ego vehicle policy, and the iteration is repeated until the maximum number of training steps is reached. After the training is completed, the test phase begins. For each normal scenario, an attack trajectory is constructed and simulated in turn, and the evaluation index ξ = α·C is recorded. T ·a d . C T is the number of collisions per unit time, used to measure safety; a d It is the average maximum deceleration of the system during the test period, used to measure the strength of the system's response to threats.
[0114] This closed-loop approach ensures that the attack trajectory adapts to the policy updates of the tested vehicle in real time, thereby constructing continuously effective adversarial inputs and systematically improving attack efficiency and evaluation coverage.
[0115] Step S5, adversarial scene construction: fuse the high-threat background vehicle trajectory with the original scene to construct a complete adversarial traffic interaction scene as the input for the automatic driving system training and verification, and improve the richness and representativeness of the test sample;
[0116] Specifically, let the structured information of the original road scene be W, including map topology, lane constraints, traffic light status, intersection connection relationship, etc.; the generated high-threat background vehicle trajectory The trajectory of the automatic driving vehicle under test is Y ego . Then the constructed adversarial scene S att can be expressed as:
[0117]
[0118] wherein, represents a function of fusing the road network with the vehicle trajectory, which is used to insert the attacker trajectory into the corresponding vehicle position and path control interface of the simulation platform according to the specified timestamp, so as to realize the implantation of behavior intervention. In order to ensure the physical executability and road semantic consistency of the adversarial trajectory, the following constraint function is designed to effectively screen the generated trajectory:
[0119] D(Y att ,W)≤
[0120] wherein D(·) represents a minimum distance function between the attack vehicle trajectory and the road center line or drivable area, and is a maximum deviation threshold value, which ensures that the attack trajectory will not deviate from the legal driving area (such as driving out of the lane, crossing obstacles, etc.).
[0121] The implementation process of the embodiment according to the automatic driving critical scene closed-loop generation method based on a large language model is as follows: Figure 5The implementation process is based on the CARLA open-source automatic driving simulation platform, and a scene with road semantic consistency and physical executability is constructed. Specifically, in the simulation platform, to ensure the execution accuracy of the attack vehicle trajectory and the consistency of the simulation process, the generated attack trajectory needs to be aligned and time-synchronized according to the simulation timeline. This process discretizes the continuous trajectory sequence into time frames, and by binding the attack vehicle to the corresponding entity, it calls the control interface of the CARLA platform in each frame to synchronize the trajectory pose instructions, ensuring smooth movement of the vehicle according to the expected path. Further, this embodiment queries the lane, drivable area boundary, and other information of the attack vehicle through the map API provided by CARLA, and combines the lateral error between the trajectory point and the lane center line to calculate the trajectory legality distance metric D, and then verifies whether all trajectory points meet the constraint conditions. Finally, after completing the trajectory feasibility confirmation, the system unifies the road network information, the future trajectory of the measured vehicle, and multiple attacker trajectories into a structured scene description file (which can be in JSON format or OpenSCENARIO standard). The integrated scene is integrated into the test scene library of the automatic driving simulation platform and continuously loaded for execution by the subsequent closed-loop countermeasure test system, achieving the response capability evaluation and robustness training of the automatic driving system in critical intervention scenarios.
[0122] The above is only a few embodiments of the present application, and does not limit the present application in any form. Although the preferred embodiments are disclosed as above, they are not intended to limit the present application. Any skilled person in the art can make some changes or modifications to the disclosed technical content without departing from the scope of the technical solution of the present application, which are equivalent to equivalent embodiments and belong to the scope of the technical solution.
Claims
1. A closed-loop generation method for critical scenarios in autonomous driving based on a large language model, characterized by: include: Step S1, environmental state extraction: extracting environmental state data of the current scene based on the autonomous driving simulation platform, including road geometry, ego vehicle motion trajectory, historical behavior trajectory of background vehicles, and dynamic scene semantic labels; Step S2: Generate high-threat background vehicle identification prompts: Based on the environmental state data in step S1, construct structured prompt information and input it into the large language model to generate an initial threat identification function that includes behavioral associations and potential conflict relationships. This function describes the behavioral associations, spatial layout, and potential conflict relationships between traffic participants, thereby improving the model's ability to perceive scene semantics. Step S3: Large language model scenario reasoning and automatic iterative screening of threatening vehicles: Utilize the large language model to perform multiple rounds of iterative optimization on the initial recognition function, combined with simulation evaluation and feedback mechanisms, to screen out a set of highly threatening background vehicles; Step S4, closed-loop optimization and generation of attack trajectories: Based on the selected high-threat background vehicles, an aggressive trajectory is generated with the goal of maximizing the collision probability using a reinforcement learning algorithm, forming a closed-loop adversarial training; Step S5: Constructing an adversarial scenario: Fusing the offensive trajectory with the original scenario to construct a physically executable and semantically consistent adversarial traffic interaction scenario for testing and training the autonomous driving system.
2. The closed-loop generation method for critical scenarios in autonomous driving based on a large language model according to claim 1, characterized in that: In the step S1: The road geometry structure includes road network topology information and three-dimensional road surface features; The vehicle's motion trajectory includes real-time position, speed, heading angle, and historical trajectory sequence; The historical behavior trajectory of the background vehicle includes the historical trajectory of the past five seconds and the future prediction window data; The dynamic scene semantic labels include scene type, weather conditions and traffic rule constraints; The data is output in JSON or Protobuf format through the API interface of the simulation platform, and a traffic scenario knowledge graph is constructed.
3. The closed-loop generation method for critical scenarios in autonomous driving based on a large language model according to claim 1, characterized in that: In the step S2: The structured prompt information is generated through a chain thinking prompt strategy to guide the large language model to reason about the threat scoring function in stages; The threat score function is expressed as: Among them, f j (·) represents the jth sub-characteristic function (such as longitudinal closing speed, lane change intention, intersection conflict probability, etc.), w j is the corresponding weight parameter, Δx i (t), Δv i (t) is the relative position and relative speed of the vehicle, θ i Indicates the vehicle's direction.
4. The closed-loop generation method for critical scenarios in autonomous driving based on a large language model according to claim 1, characterized in that: The step S3 comprises: Initial function building module: Generates an executable threat identification function with input parameters including the status of the ego vehicle and background vehicles; Evaluation and feedback module: Analyzes function defects based on the success rate of simulated attacks and generates optimization suggestions; Behavioral logic evolution module: adjusts function parameters or structure based on optimization suggestions and selects the optimal function using a multi-candidate evaluation mechanism; The iterative process meets the following conditions: A. If the attack success rate of the candidate function increases, it will be updated to the next round; B. If the attack success rate reaches the threshold, only fine-tune the weight parameters; C. When there is no gain for multiple consecutive rounds, the function structure can be adjusted.
5. The closed-loop generation method for critical scenarios in autonomous driving based on a large language model according to claim 1, characterized in that: In the step S4: Set a reward function based on the reinforcement learning algorithm to maximize the collision probability; A behavioral perturbation strategy is used to generate aggressive trajectories; The scenarios are continuously updated based on the performance of the trained autonomous driving system to form a closed-loop optimization.
6. The closed-loop generation method for critical scenarios in autonomous driving based on a large language model according to claim 1, characterized in that: In the step S5: The construction of the adversarial scenario satisfies the physical executable constraints of the trajectory: Here, the structured information of the original road scene is set as W, which includes map topology, lane constraints, traffic light status, intersection connection relationship, etc. The function that represents the fusion of the road network and vehicle trajectory is used to insert the attacker's trajectory into the corresponding vehicle position and path control interface in the simulation platform at the specified timestamp, thus implementing behavioral intervention. To ensure the physical executable and road semantic consistency of the adversarial trajectory, the following constraint function is designed to screen the validity of the generated trajectory: D(Y att ,W)≤ where D(·) represents the minimum distance function between the attacking vehicle trajectory and the road centerline or the drivable area, and is the maximum deviation threshold tolerated to ensure that the attacking trajectory does not deviate from the legal driving area.
7. The closed-loop generation method for critical scenarios in autonomous driving based on a large language model according to claim 3, characterized in that: The chain thinking prompt strategy includes the following stages: The first stage is to determine the core influencing factors of threat assessment; The second stage is to infer the evolution pattern of influencing factors over time; The third stage is to construct a threat assessment indicator system through weighted normalization.
8. The closed-loop generation method for critical scenarios in autonomous driving based on a large language model according to claim 4, characterized in that: The multi-candidate evaluation mechanism is specifically as follows: Generate N candidate functions in each iteration; Then, the attack success rate is calculated through small-scale simulation tests, and the optimal candidate function is selected for subsequent iterations.
9. The closed-loop generation method for critical scenarios in autonomous driving based on a large language model according to claim 5, characterized in that: The trajectory prediction model is DenseTNT, which is used to generate multiple high-probability trajectory candidates of background vehicles and calculate their joint collision probability with the ego vehicle trajectory.
10. An electronic device, characterized in that: It includes a memory, a processor and a computer program stored in the memory, and when the processor executes the program, it implements the closed-loop generation method of the critical scene of autonomous driving as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Lane changing scene automatic driving strategy and system considering adversarial attack
CN118991827A
Safety-critical traffic simulation method based on adversarial migration of driving intention
CN119358371A
Cited By
Site test scene construction method based on feature matching and trajectory optimization
CN121031143A
Method for constructing a test site scene based on feature matching and trajectory optimization
CN121031143B
End-to-end automatic driving real confrontation scene generation and closed loop verification system and method
CN121257331A
Automatic driving scene generation method and device, equipment and storage medium
CN121636640A
An automatic driving scene generation method, device, equipment and storage medium
CN121636640B