Scene simulation optimization method and device, electronic equipment and storage medium

By building high-risk test scenarios through reinforcement learning and combining proximal policy optimization and vulnerability amplification mechanisms, the problem of insufficient robustness assessment of autonomous driving systems in complex environments is solved, efficient test scenario simulation optimization is achieved, and the safety and robustness of the system under extreme conditions are improved.

CN120597707APending Publication Date: 2025-09-05BEIHANG UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510715865.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

It is difficult to fully evaluate the robustness and safety of existing autonomous driving systems in complex open road environments, especially in sudden traffic incidents, where test coverage is limited and there is a lack of semantically reasonable adversarial test scenarios and dynamic feedback mechanisms.

Method used

Reinforcement learning is used to construct test scenarios with emergency characteristics. High-risk, controllable test scenarios are generated through the scenario modeling module, environmental disturbance module, and actor configuration module. Combined with the proximal policy optimization method and vulnerability amplification mechanism, continuous stress testing and in-depth verification are carried out.

Benefits of technology

It has enhanced the depth of robustness assessment of autonomous driving systems under extreme conditions, built a high-voltage testing process covering strategy boundaries, enhanced the aggressiveness and behavioral evolution capabilities of test scenarios, and ensured the system's decision-making robustness and extreme safety performance in complex traffic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597707A_ABST
    Figure CN120597707A_ABST
Patent Text Reader

Abstract

The invention provides a scene simulation optimization method and device, electronic equipment and a storage medium. The method comprises the following steps: constructing a test scene with emergency characteristics by adopting reinforcement learning, wherein the test scene comprises a modeling module, an environment disturbance module and a behavior body configuration module; if the input automatic driving algorithm is valid in the risk of the test scene, evolving and updating a trajectory disturbance strategy of the adversarial behavior body; and if the automatic driving algorithm fails in the test scene or the evolved and updated trajectory disturbance strategy, performing continuous pressure measurement and depth verification on the decision vulnerability of the automatic driving algorithm based on a vulnerability amplification mechanism. A good theoretical basis method is provided for testing and verifying decision robustness and limit safety performance of an automatic driving system in a complex urban traffic environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of autonomous driving testing technology, and specifically relates to a method, device, electronic device and storage medium for scene simulation optimization. Background Art

[0002] The rapid development of autonomous driving technology has led to significant progress in perception, planning, and control. However, the robustness, safety, and boundary behavior of these systems remain difficult to fully assess through conventional testing methods. In complex, open road environments, autonomous vehicles face a variety of unexpected traffic incidents, such as pedestrians suddenly crossing, vehicles forcing their way in, and illegal turns at intersections. These situations often present high-risk triggers for system failure. The construction of high-risk, controllable, and semantically sound adversarial testing scenarios has become a key issue in the safety verification of autonomous driving systems.

[0003] Therefore, a new method for scene simulation optimization is expected. Summary of the Invention

[0004] Some research has explored scenario construction and adversarial testing. Some previous studies used offline sampling or rule-based scripting to synthesize scenarios, lacking dynamic feedback mechanisms for system responses. Other studies have introduced methods such as reinforcement learning to generate high-risk trajectories, but these often suffer from difficulties in modulating aggressiveness, lacking semantic rationality in behavior, and inability to close the testing process. These issues limit test coverage and hinder systematic discovery of system policy vulnerabilities.

[0005] In response to the problems existing in the prior art, the present invention provides a method, device, electronic device and storage medium for scene simulation optimization, aiming to enhance the robustness evaluation depth of autonomous driving systems under extreme conditions.

[0006] In a first aspect, an embodiment of the present invention provides a method for scene simulation optimization, comprising:

[0007] Reinforcement learning is used to construct a test scenario with emergency characteristics, including a test scenario modeling module, an environmental disturbance module, and an actor configuration module;

[0008] If the input autonomous driving algorithm is effective in the risk of the test scenario, the trajectory perturbation strategy of the adversarial actor will be evolved and updated; if the autonomous driving algorithm fails in the test scenario or the evolutionarily updated trajectory perturbation strategy, the decision-making vulnerabilities of the autonomous driving algorithm will be continuously stress-tested and deeply verified based on the vulnerability amplification mechanism.

[0009] Furthermore, the scenario modeling module generates an editable topological structure of the road space based on the structural component library and the topological splicing generator, and splices a test layout area with induction potential based on the test target. The test layout area includes the road topology structure, as well as the traffic constraints and strategic path solution space of the road topology structure. The road topology structure includes intersections, occlusion sections, and narrow passages formed by dynamic splicing.

[0010] The environmental disturbance module is used to introduce multi-dimensional disturbance factors to simulate real-world environmental factors. The multi-dimensional disturbance factors include weather, lighting, visibility, ground adhesion, and sensor noise. During the disturbance process, the disturbance control graph is used to map the disturbance to the road topology, integrating the disturbance information and the road topology information.

[0011] The actor configuration module generates traffic participants based on the test layout area and the test target. The traffic participants include regular actors and antagonistic actors. The regular actors construct the basic traffic flow density, and the antagonistic actors trigger high-risk behaviors through strategy templates.

[0012] Furthermore, when the input autonomous driving algorithm avoids the risk of the test scenario, the proximal policy optimization method is used to update the trajectory perturbation strategy of the adversarial actor.

[0013] Furthermore, the proximal strategy optimization method includes:

[0014] Construct a policy margin function based on the main vehicle response index of the autonomous driving algorithm;

[0015] Dynamically adjusting disturbance parameters based on the policy margin function, wherein the disturbance parameters include attack rhythm, direction change, and combination behavior;

[0016] The test scenario evolves round by round, and the intensity of the strategy pressure is increased step by step until the system response boundary is approached and the critical state of stability is exposed.

[0017] Furthermore, the vulnerability amplification mechanism is used to continuously stress test and deeply verify the decision vulnerabilities of the autonomous driving algorithm, including:

[0018] Through abnormal trajectory clustering, failure variable extraction and control sensitivity analysis, key control variables and vulnerable response points in the failure path are located, and multi-round disturbance trajectory belts are constructed;

[0019] Guided by policy gradient projection, the multiple rounds of perturbation trajectory are used to construct a crash path perturbation sequence by superimposing high-frequency focused perturbation functions. Combined with the semantic consistency control mechanism, high-voltage test fragments for continuous stress testing and deep verification are generated.

[0020] Furthermore, constructing the multi-round disturbance trajectory band includes:

[0021] Structural reconstruction of the process of risk avoidance, tracing back the main vehicle's decision-making response chain, extracting the key dynamic variables that cause failure, and forming the failure path state feature vector;

[0022] A strategy failure region extraction mechanism is introduced, combined with trajectory trend analysis and anomaly detection models, to identify the most triggering time periods and state segments in the failure path state feature vector, marking them as strategy vulnerable windows.

[0023] An iterative trajectory perturbation enhancement method is used for the failure fragile window, and the perturbation is concentratedly injected into the most strategically sensitive time period and control variable dimension to obtain a multi-round perturbation trajectory band.

[0024] Furthermore, it also includes:

[0025] Performing structured collection of test samples for each round of evolutionary updates, wherein the collection objectives include perturbation strategy parameters, actor states, control response characteristics, and labeling results;

[0026] The sample coverage function is used to evaluate the results of structured collection to obtain sample quality and scene diversity. Based on the sample quality and scene diversity, a test sample library with semantic labels, structural mapping and multi-dimensional indexing is established.

[0027] In a second aspect, an embodiment of the present disclosure provides a device for scene simulation optimization, the device comprising:

[0028] A scenario construction unit is configured to construct a test scenario with emergency characteristics using reinforcement learning, including a test scenario modeling module, an environment perturbation module, and an actor configuration module;

[0029] The testing unit is configured to, if the input autonomous driving algorithm is effective in the risk of the test scenario, evolve and update the trajectory perturbation strategy of the adversarial actor; if the autonomous driving algorithm fails in the test scenario or the evolved and updated trajectory perturbation strategy, continuously stress test and deeply verify the decision vulnerabilities of the autonomous driving algorithm based on the vulnerability amplification mechanism.

[0030] In a third aspect, an embodiment of the present disclosure provides an electronic device, the electronic device comprising:

[0031] at least one processor; and,

[0032] a memory communicatively connected to the at least one processor; wherein,

[0033] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the above-mentioned scene simulation optimization method.

[0034] In a fourth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium, which stores computer instructions for causing the computer to execute the above-mentioned scene simulation optimization method.

[0035] Other optional features and technical effects of the embodiments of the present invention are partially described below, and partially can be understood by reading this document.

[0036] Compared with the prior art, the present invention has the following beneficial technical effects:

[0037] The present invention provides a method for scenario simulation optimization, which can dynamically adjust the strategies of traffic participants based on the real-time response behavior of the autonomous driving system, continuously enhance the aggressiveness and behavioral evolution capabilities of the test scenario, and construct a high-voltage test process covering the strategy boundaries. It provides a good theoretical basis method for testing and verifying the decision robustness and extreme safety performance of the autonomous driving system in a complex urban traffic environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 A schematic flow chart of a method for scene simulation optimization according to an embodiment of the present disclosure is shown;

[0039] Figure 2 A schematic diagram of the structure of a scene modeling module according to an embodiment of the present disclosure is shown;

[0040] Figure 3 A schematic diagram of the process flow of the proximal strategy optimization method according to an embodiment of the present disclosure is shown;

[0041] Figure 4 A schematic flow chart of a method for scene simulation optimization with a sample database according to an embodiment of the present disclosure is shown;

[0042] Figure 5 A device diagram for scene simulation optimization according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0043] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.

[0044] The following describes the embodiments of the present disclosure through specific examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.

[0045] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this disclosure, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement an apparatus and / or practice a method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this apparatus and / or practice this method.

[0046] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present disclosure. The illustrations only show components related to the present disclosure and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.

[0047] Additionally, in the following description, specific details are provided to provide a thorough understanding of the examples. However, one skilled in the art will appreciate that the aspects described can be practiced without these specific details.

[0048] Figure 1 A flow chart 100 of a method for scene simulation optimization according to an embodiment of the present disclosure is shown. Figure 1 Shown, including:

[0049] In step S101, reinforcement learning is used to construct a test scenario with emergency characteristics, including a test scenario modeling module, an environment disturbance module and an actor configuration module.

[0050] Specifically, in some embodiments, Figure 2 A schematic diagram of the scene modeling module structure of an embodiment of the present disclosure is shown in FIG. Figure 2 As shown, the scenario modeling module generates an editable topological structure of the road space based on the structural component library and the topological splicing generator, and splices a test layout area with induction potential based on the test target. The test layout area includes the road topology structure, as well as the traffic constraints and strategic path solution space of the road topology structure. The road topology structure includes intersections, occlusion sections, and narrow passages formed by dynamic splicing;

[0051] The environmental disturbance module is used to introduce multi-dimensional disturbance factors to simulate real-world environmental factors. The multi-dimensional disturbance factors include weather, lighting, visibility, ground adhesion, and sensor noise. During the disturbance process, the disturbance control graph is used to map the disturbance to the road topology, integrating the disturbance information and the road topology information.

[0052] The actor configuration module generates traffic participants based on the test layout area and the test target. The traffic participants include regular actors and antagonistic actors. The regular actors construct the basic traffic flow density, and the antagonistic actors trigger high-risk behaviors through strategy templates.

[0053] Furthermore, in some other embodiments, step S101 aims to test the robustness and extreme responsiveness of autonomous driving algorithms in unexpected traffic events, constructing traffic scenarios with high structural complexity, diverse behavioral strategies, and risk-inducing capabilities. By integrating multiple submodules, including road topology modeling, disturbance environment injection, traffic actor semantics construction, and reinforcement learning-driven strategy layout, standardized test input scenarios with high semantic integrity, high disturbance complexity, and physical executableness are automatically generated for systematic testing of algorithm performance under extreme circumstances.

[0054] Specifically, in the test scenario modeling stage, the editable topological structure modeling of the road space is completed based on the structural component library and the topological splicing generator. This embodiment constructs a diversified road network structure by combining high-risk inducement components such as intersections, lane change sections, sight-blocking areas, narrow passages, etc., and supplemented by semantic configurations such as traffic signs, signal lights, and traffic rules to establish all traffic constraints and strategic path solution spaces in the scene. In this embodiment, the system can automatically splice out test layout areas with induction potential according to the set tasks, so that the autonomous driving algorithm being tested faces the challenges of high uncertainty and information missing points in the path planning process, providing a structural intervention basis for subsequent tests. It should be noted that the system described in this application refers to a system for testing autonomous driving algorithms, and the system at least includes the test scenario constructed in this application.

[0055] In some embodiments, the road topology is represented as follows:

[0056] T=(V,E,Ω,Ψ,Frule ,F env )

[0057] Where V={v1,v2,…,v n} represents the node set in the traffic network. Nodes can be structural units such as road segment endpoints, intersection centers, and traffic light control points. E = {e1, e2, …, e m} represents an edge set, which indicates the road connection relationship between nodes, and each edge corresponds to an actual road section; Ω represents a static attribute set of the topological structure, including basic physical parameters such as road grade (main road / secondary road), speed limit, number of lanes, width, slope, etc.; ψ:V∪E→L represents a semantic label mapping function, which defines the traffic semantic type of each node or edge (such as "intersection", "merging section", "blocked section", "blind spot entrance and exit", etc.), which is used to guide the generation of actor strategies; F rule :E→R k It represents the traffic rule constraint function, which specifies the behavior restriction conditions on each edge e, such as traffic priority, lane change prohibited area, signal light constraint status, etc., and controls the behavior decision space of traffic participants. It is defined as:

[0058]

[0059] Where: V i (x) represents the degree of violation related to the i-th type of traffic rules (e.g., speeding, red light violation, driving out of lane, etc.). It should be noted that V i (x) is a dynamic function that changes according to the vehicle's real-time position, speed, and environmental conditions; R i (x) represents the degree of compliance with the traffic rules related to the i-th type, which is usually a measure related to the strategic decision of the actor (such as whether to obey the traffic light); α i , β i They are the weight coefficients of traffic rules and rule compliance, which can be adjusted by those skilled in the art according to the importance of specific rules; rule represents a threshold value used to determine when traffic rules are violated, ∈ rule It may be a dynamic, environment-dependent value, for example, a larger error is allowed under special conditions (such as rainy days or poor visibility).

[0060] In order to enhance the environmental impact of the test scenario, in the process of building a multimodal risk environment, F env :(x,t)→R represents the environmental disturbance mapping function, which specifies the external disturbance intensity at the time and space point (x,t), such as reduced visibility, sudden change in lighting, change in adhesion coefficient, etc., and is used to construct a multimodal risk environment. It is defined as:

[0061]

[0062] in: represents a Gaussian distribution term, simulating the distribution of environmental disturbances in space, such as the impact of traffic flow changes and meteorological conditions; P i (x, t) represents the time and space function related to the disturbance factor, which indicates the intensity of the impact of the environmental disturbance on the position x and time t; T i (x, t) represents the disturbance factor related to the time series, which represents the impact of the environment on traffic in different time periods (such as daytime and nighttime); S i (x, t) represents the impact of environmental interference (such as sensor noise, perception error, etc.) under different environmental conditions, N j (x, t) represents the dynamic obstacle disturbance, which represents the impact of real-time obstacles on the road on the traffic system, such as other vehicles, pedestrians, or temporary obstacles.

[0063] The environmental disturbance module is activated after the road modeling is completed. The system introduces multi-dimensional disturbance factors to simulate non-ideal traffic conditions in reality, including weather, lighting, visibility, ground adhesion, sensor noise and other dimensions. In some embodiments, disturbance injection adopts a regional calibration method, and the distribution of disturbances in space and time is controlled by a disturbance configuration file. It supports the setting of static disturbances (such as dense fog throughout the journey), regional disturbances (such as low lighting in the tunnel only) and temporal disturbances (such as sudden changes in lighting), and maps them to the road topology through a disturbance control diagram to achieve the integrated integration of disturbance information and structural information.

[0064] The perturbation mapping function is: D(x,t)=M env (T,E,x,t)

[0065] Where: D(x,t) represents the dynamic intensity map, E represents the perturbation parameter set, T is the topological structure, and x and t represent the spatial position and time variables of the perturbation effect, respectively.

[0066] The actor configuration module combines the road structure and test objectives to generate a set of traffic participants. Each actor includes its initial state (position, speed, direction) and multi-layer semantic labels (type, intention, rule response, strategy risk level, etc.). All traffic actors are divided into two categories: regular actors and adversarial actors. Regular actors construct the basic traffic flow density; adversarial actors trigger high-risk behaviors such as sudden lane changes, rushing, and delayed intervention through strategy templates. The deployment position of adversarial actors is jointly optimized by the strategy heat function and the overlap of the main vehicle path:

[0067]

[0068] Among them: H risk (x) represents the risk heat score of position x, which is defined as:

[0069] H risk (x)=ω1·D fail (x)+ω2·K conflict (x)+ω3·C topo (x)

[0070] Where: D fail (x) represents the triggering density of historical failure events at position x; K conflict (x) represents the density of path intersection and behavior conflict; C topo (x) Topological complexity score.

[0071] P overlap (x,π0) represents the overlap probability density between the main vehicle path and the area, which is defined as:

[0072]

[0073] Where: π0(t) is the expected trajectory point of the main vehicle at time t, σ is the Gaussian kernel width, which is used to smooth the path projection; D entropy (x) represents the uncertainty entropy of traffic strategy distribution; C rule (x) indicates the penalty for violating traffic rules at the deployment point; C reach (x) represents the spatial and temporal cost for the adversarial actor to reach that point; It is used to reflect how the actor changes over time and how to adaptively adjust its strategy according to changes in the surrounding environment. This item enables the strategy to adjust according to environmental feedback, making the actor more adaptive.

[0074] To further enhance the testing value of the initial scenario configuration, this embodiment introduces a reinforcement learning strategy generation module. This module automatically optimizes the combination of actor positions and semantic labels through a strategy network. The strategy network takes topology encoding, task constraints, and historical failure regions as input and outputs strategy actions for layout optimization. The final objective function of the strategy network is designed as a multi-objective nested structure, including:

[0075]

[0076] Where: R fail (a, s) represents the induced failure risk score under the current state s, which measures the potential challenge posed by the policy configuration to the algorithm and is defined as:

[0077]

[0078] Among them, ||a current -a ideal || 2 Indicates the degree of path deviation, considering the deviation between the vehicle path and the ideal path, and assessing the risk of deviating from the safe path; represents the perception error term, which represents the error in the vehicle’s perception of dynamic targets such as obstacles and pedestrians, which depends on the sensor accuracy; represents the failure probability caused by interaction with other traffic participants.

[0079] Among them, D sem (a) represents the semantic label space coverage index, which improves semantic diversity and is defined as:

[0080]

[0081] in, Represents the i-th semantic label of the actor a, such as driving purpose, speed, etc. Mapping and comparing through this function can improve the diversity of strategies; P i (a) represents the uncertainty of the actor’s path decision, measured by entropy; T i (a) represents the time dependence of the actor, which is used to control the decision uncertainty of the actor in a dynamic environment.

[0082] U entropy (a) represents the uncertainty entropy of the perturbation strategy, which controls the diversity of trajectory behavior distribution and is defined as:

[0083]

[0084] Where: p i Represents the probability of each strategy in the strategy space, which is used to measure the diversity of strategies; ||[a i =a previous ] represents historical policy constraints, indicating the degree of overlap between the current policy and the historical policy.

[0085] C phys (a) represents the physical constraint penalty, such as initial conflict, illegal speed, etc., and is defined as:

[0086]

[0087] Among them: Violation (a i ) represents strategy a i Whether physical constraints are violated, such as speed limits, acceleration limits, etc.; ∈ phys Represents the tolerance threshold of the physical constraint.

[0088] C history (a) represents the historical redundancy penalty, which is used to avoid duplication of test samples and is defined as:

[0089]

[0090] Where: ||[a i=a prev ] indicates that it is used to detect whether the current strategy is repeated with the previous strategy; Indicates the distance between the current policy and the historical policy, if any.

[0091] Through the above-mentioned structural modeling, environmental disturbance, actor configuration and reinforcement learning optimization mechanisms, the system can generate adversarial traffic test scenarios with reasonable structure, controllable disturbance, clear semantics and flexible strategies, providing high-pressure input close to the strategy boundary for the autonomous driving algorithm, and fully supporting subsequent behavior evolution adjustment and strategy robustness closed-loop testing tasks.

[0092] Next, go to step S102.

[0093] At step S102, if the input autonomous driving algorithm is effective in the risk of the test scenario, the trajectory perturbation strategy of the adversarial actor is evolved and updated; if the autonomous driving algorithm fails in the test scenario or the evolutionarily updated trajectory perturbation strategy, the decision-making vulnerabilities of the autonomous driving algorithm are continuously stress-tested and deeply verified based on the vulnerability amplification mechanism.

[0094] In some embodiments, when the input autonomous driving algorithm avoids the risk of the test scenario, a proximal policy optimization method is used to update the trajectory perturbation strategy of the adversarial actor.

[0095] Specifically, Figure 3 A schematic diagram 200 of a proximal strategy optimization method according to an embodiment of the present disclosure is shown. Figure 3 As shown, the proximal strategy optimization method includes:

[0096] Step S201, constructing a strategy margin function based on the main vehicle response index of the autonomous driving algorithm;

[0097] Step S202: dynamically adjusting disturbance parameters based on the strategy margin function, wherein the disturbance parameters include attack rhythm, direction change, and combination behavior;

[0098] Step S203 , evolving the test scenario in rounds and gradually increasing the policy pressure intensity until the system response boundary is approached and the critical stability state is exposed.

[0099] To be more specific, in some embodiments, when the autonomous driving algorithm under test successfully completes the risk avoidance task in the adversarial test scenario, the system automatically enters the strategy evolution and disturbance enhancement process mechanism, which is driven by the strategy evolution regulator module (Strategy Evolution Regulator, SER), and aims to continuously enhance the input test pressure based on the algorithm response behavior, so that it gradually approaches the strategy collapse boundary.

[0100] Specifically, the dynamic response data of the host vehicle is collected during the test, covering the relative position relationship between the actors, collision time estimation, control command fluctuation amplitude and response lag. The dynamic response data jointly determines the stability and risk margin of the current avoidance behavior. In order to uniformly model the characteristics of the dynamic response data, this embodiment introduces the avoidance margin function R safe , as the key trigger signal for determining whether to start strategy evolution, is expressed as follows:

[0101]

[0102] Where: d min It represents the minimum relative distance between the tested vehicle and the adversarial actor in this round of testing; TTC is the estimated time to collision; It represents the gradient of the control input (such as acceleration and direction) at continuous moments, reflecting the degree of oscillation of the control strategy; delay represents the system response delay from decision to execution; S jerk represents the second-order derivative (jerk) of the control command per unit time, which measures the control jitter and response smoothness; d0 and ε are constants introduced to prevent numerical explosion or normalization; α, β, γ, δ, and η are weight coefficients set by the system.

[0103] This margin function is used in the evolution strategy adjustment module to determine whether the current test state has entered the "challenge critical zone". safe Below the preset threshold R thr , the next round of attack actor perturbation strategy update will be automatically triggered.

[0104] In order to generate trajectory disturbance behaviors with oppression, response sensitivity and diversity, the system of the present disclosure is based on the current state of the main vehicle s t , policy response record and scene local structure to extract disturbance feature encoding function This function describes the risk status, behavioral response pattern, and strategic tension level faced by the system under test in the current test round.

[0105] In order to inject disturbance into the actor trajectory, the system introduces the perturbation trajectory evolution expression:

[0106]

[0107] Where: k (t) represents the original trajectory of the adversarial actor in the kth round of evolution; [D i (t)] is an indicator function, which represents the dynamic disturbance of the vehicle at the i-th time point. The dynamic disturbance reflects the impact of environmental factors on the behavior strategy; W represents the disturbance mapping matrix, which controls the direction and intensity of the disturbance; x i (t),xj (t) represents the position of the i-th and j-th actors at time t, Calculate its spatial distance; δ adv (t) represents the disturbance activation control function, which is used to adjust the intensity distribution of the disturbance in the time domain; ψ(t,s t ) represents the disturbance injection characteristic function, which combines vehicle dynamics, local scene complexity and host vehicle response history.

[0108] To avoid overly idealized or linear perturbations, the system introduces the following perturbation timing activation function:

[0109] δ adv (t)=σ(α1·sin(ωt+φ)+α2·G(s t ))

[0110] Where: σ(·) is the Sigmoid function, which controls the smoothness of the perturbation activation; sin(ωt+φ) is the perturbation rhythm reference term, which supports periodic shocks; G(s t ) is a dynamic risk coefficient calculated based on the current state of the system (such as control tension and environmental complexity); α1 and α2 represent rhythm weight factors.

[0111] It should be noted that the mechanism described in the embodiments of this disclosure ensures that disturbances not only exist in the spatial dimension but also have dynamic scheduling capabilities in the time domain, simulating "burst-like," "flash-like," or "delay-induced" behavior characteristics, thereby more realistically challenging the response timeliness and strategy retention capabilities of the autonomous driving algorithm. After executing the disturbance, the system performs structured collection of the main vehicle control, trajectory, and judgment labels to generate an evolution sample sequence, which is recorded as:

[0112]

[0113] Where: φ k is the perturbation input feature, χ k ,χ k+1 are the front and rear trajectories, y k The system's feedback on the disturbance (such as whether it triggers labels such as avoidance failure, offset, and delay limit).

[0114] Furthermore, to prevent random divergence of perturbations in the policy space, this embodiment introduces a perturbation direction focusing mechanism to guide perturbations to the direction where the algorithm's control capability is most vulnerable. This mechanism is based on a sensitivity analysis of the response gradient of the control loss function in the perturbation space and proposes a policy vulnerability projection function:

[0115]

[0116] in: Represents the gradient vector of the control loss function with respect to the perturbation input φ, indicating the direction of the perturbation's influence in each dimension; Represents the covariance matrix of the disturbance variable, which is used to regulate the coupling and correlation between different disturbance dimensions; R proj (φ) represents the projection intensity of the current disturbance direction on the strategy failure surface, which is used to measure the “attack effectiveness” of inducing failure.

[0117] It should be noted that in the embodiment of the present disclosure, the core function of the policy vulnerability projection function is to "project" the disturbance into the decision subspace where the policy is most sensitive and most prone to collapse, so that the disturbance is no longer uniform and tentative, but precisely directed and feedback-driven, thereby improving the efficiency of stress testing resource utilization and accelerating the generation of failure samples.

[0118] The system can use the policy vulnerability projection function as: one of the reward factors of the PPO policy network to push the learning process closer to the policy collapse boundary; the distribution weight of the perturbation direction sampler to reconstruct the perturbation sampling distribution; part of the sample annotation mechanism to construct perturbation strength-vulnerability label pairs.

[0119] Ultimately, this mechanism achieves "active selection" of disturbance direction without introducing additional control execution burden, strengthens the strategic attack focus of the disturbance actor, and provides directional constraints for the subsequent strategic optimization goals.

[0120] In order to achieve the automatic evolution of the anti-disturbance strategy, this embodiment also introduces a reinforcement learning mechanism, which uses trajectory disturbance as the output space of the strategy network, and the strategy network parameters are updated according to the adversarial test results. In this process, the system constructs a multi-objective optimization function J evolve , which is used to guide the perturbation strategy to achieve a balance among aggressiveness, diversity, feasibility and efficiency during the evolution process. The objective optimization function J evolve include:

[0121]

[0122] Where: π θ represents the perturbation strategy network, with parameter θ, generating the perturbation behavior of the actor; P fail (a, s) represents the probability that the disturbance will cause the algorithm under test to fail in the current state, which is the primary optimization goal; H(π a ) represents the path information entropy of the trajectory perturbation strategy, which is used to improve diversity and avoid mode trapping; represents the strategy sensitivity index, which is used to encourage the strategy to evolve towards control failure; C phys (a) represents the physical irrationality penalty caused by the disturbance, such as violation of traffic rules and dynamic inaccessibility; ||Δ adv || 2Represents the perturbation energy regularization term to prevent excessive perturbation from affecting the scene semantics.

[0123] It should be noted that, in the embodiment of the present disclosure, the target optimization function J evolve As the dominant signal for strategy training, it guides the perturbation strategy to learn from the test feedback. After each round of simulation test, the system uses the PPO algorithm to θ to update.

[0124] During the training process, the policy network continuously receives the evolution sample sequence S evolve The state-response pairs in the algorithm are gradually increased through failure rate-driven policy enforcement, while balancing perturbation rationality and policy generalization. Through this mechanism, the system forms an adaptive perturbation evolution process, continuously pushing towards the policy boundaries of the tested algorithm and maximizing its potential failure points.

[0125] When the perturbation strategy fails to improve the failure rate after several consecutive rounds, the system automatically triggers the "attack style upgrade mechanism", switches the actor type or strategy style, and reinitializes the strategy network structure to maintain continuous innovation in the testing process and dynamically enhance the pressure capability.

[0126] During the strategy evolution process, each module works together to form a complete feedback-driven testing system. The Behavior Strategy Evolution Regulation Module (SER) serves as the core controller and works in conjunction with the following submodules:

[0127] The disturbance generation module is configured to be responsible for the construction and timing scheduling of trajectory disturbances;

[0128] The strategy evaluation module is configured to collect response data from each round of perturbation testing and generate evolution samples;

[0129] The learning update module is configured to take the evolution samples as input and optimize the perturbation strategy network;

[0130] The sample management module is configured to record the trajectory, status, and label of each round, provide feedback, and form a long-term accumulated strategy adversarial sample library.

[0131] After each round of perturbation, the system will convert the sample triplet χ k ,χ k+1 ,y k Write to the evolution database and update the risk clustering heat map to support the localization of subsequent strategy failure areas and the generation of style transfer strategies.

[0132] This mechanism ensures that the dynamic evolution of the perturbation strategy is not simply a matter of "intensity pressure," but rather incorporates multi-layered capabilities such as structural perception, directional guidance, feedback adjustment, and strategy migration. Ultimately, this creates a test module with closed-loop evolutionary capabilities. It continuously feeds complex, realistic, and directionally challenging combinations of actor strategies into the algorithm under test, continuously improving the intensity and coverage of the test, thereby maximizing the exposure of the robustness boundaries and response flaws of the autonomous driving algorithm.

[0133] This step allows for a complete stress testing process of "self-adaptation - feedback-driven - strategy amplification," avoiding insufficient test coverage caused by "single-round circumvention," maximizing the exploration of system boundary behaviors and potential failure points, and laying a solid foundation for subsequent strategy failure area focus and vulnerability amplification mechanisms.

[0134] In still other embodiments, when the autonomous driving algorithm under test fails during the strategy evolution test, such as a collision event, trajectory deviation, response lag, or control anomaly, the system automatically triggers the strategy vulnerability amplification mechanism to perform structured extraction, key variable modeling, and high-density disturbance expansion on the currently exposed robustness defects, thereby achieving continuous tracking and limit verification of the strategy collapse path.

[0135] Specifically, in the disclosed embodiment, the continuous stress testing and in-depth verification of the decision vulnerabilities of the autonomous driving algorithm based on the vulnerability amplification mechanism includes:

[0136] Through abnormal trajectory clustering, failure variable extraction and control sensitivity analysis, key control variables and vulnerable response points in the failure path are located, and multi-round disturbance trajectory belts are constructed;

[0137] Guided by policy gradient projection, the multiple rounds of perturbation trajectory are used to construct a crash path perturbation sequence by superimposing high-frequency focused perturbation functions. Combined with the semantic consistency control mechanism, high-voltage test fragments for continuous stress testing and deep verification are generated.

[0138] Specifically, constructing the multi-round disturbance trajectory band includes the following steps:

[0139] Structural reconstruction of the process of risk avoidance, tracing back the main vehicle's decision-making response chain, extracting the key dynamic variables that cause failure, and forming the failure path state feature vector;

[0140] A strategy failure region extraction mechanism is introduced, combined with trajectory trend analysis and anomaly detection models, to identify the most triggering time periods and state segments in the failure path state feature vector, marking them as strategy vulnerable windows.

[0141] An iterative trajectory perturbation enhancement method is used for the failure fragile window, and the perturbation is concentratedly injected into the most strategically sensitive time period and control variable dimension to obtain a multi-round perturbation trajectory band.

[0142] Specifically, the system reconstructs the failure process, traces back the decision-making and response chain of the main vehicle, extracts the key dynamic variables that cause the failure, and forms the failure path state feature vector:

[0143] φ fail =[t delay ,σ a ,σ ω ,e track ,e percep ,δ ctrl ,δ plan ]

[0144] Where: t delay represents the total delay of the perception-decision-control system response; σ a ,σ ω Indicates the fluctuation amplitude of acceleration and angular velocity; e track represents the trajectory execution error; e percep represents the perceptual error rate; δ ctrl ,δ plan Indicates the control command jitter amplitude and path planning disturbance rate.

[0145] This feature vector comprehensively characterizes the dynamic response capability and behavioral stability of the algorithm under test before and after failure, and is used for subsequent failure area extraction and vulnerability focusing.

[0146] The system then introduces a strategy failure region extraction mechanism, combining trajectory trend analysis with anomaly detection models to cluster abnormal trajectories. This allows the identification of the most triggering time periods and state segments within the failure path, marking them as strategy vulnerability windows. It should be noted that these strategy vulnerability windows contain one or more vulnerable response points and emphasize focal areas of temporal continuity and state evolution.

[0147] After locating the vulnerable window of failure, the system enters the "vulnerability amplification disturbance construction phase". Through iterative trajectory disturbance enhancement, high-density disturbances are injected around the strategy failure area to continuously expand the system's response pressure. The trajectory disturbance evolution formula is defined as follows:

[0148]

[0149] Where: k (t) represents the original actor trajectory when the kth round fails; Indicates the j-th perturbation step, which is represented by the function f focus Generate; φfail Represents the failure state characteristics; ξ j represents the perturbation direction and intensity control vector; f focus (·) represents the strategic vulnerability focused perturbation function, which dynamically adjusts the perturbation direction and amplitude according to the known failure path.

[0150] This function concentrates the disturbance injection into the most strategically sensitive time periods and control variable dimensions, constructing a "disturbance belt" or "continuous failure channel", which is a disturbance trajectory belt to achieve continuous suppression and fine expansion of the original vulnerability.

[0151] In some embodiments, the multiple rounds of perturbation trajectory are guided by policy gradient projection, including:

[0152] The policy gradient sensitivity measurement function is introduced to differentiate the control loss function in the failure variable dimension to form a directional vulnerability heat map:

[0153]

[0154] This gradient vector is used to locate the sensitivity of the current control strategy to each disturbance dimension, assist in constructing the disturbance direction guidance function, and achieve precise strikes and alignment with strategy vulnerabilities.

[0155] Furthermore, the method of constructing a collapse path disturbance sequence by superimposing a high-frequency focused disturbance function and combining it with a semantic consistency control mechanism includes:

[0156] We introduce a policy vulnerability amplification optimization function and combine the four indicators of failure-induced probability, control sensitivity, perturbation cost, and semantic shift to further improve the systematicness and test effectiveness of perturbation amplification. We construct the following objective function:

[0157]

[0158] Of which: TTC min represents the shortest predicted collision time, indicating the critical proximity for policy collapse; represents the control loss gradient, which represents the local vulnerability of the control strategy; C phys (χ) represents the penalty term for the trajectory violating physical constraints or behavioral logic; Represents the overall sensitivity of the control strategy to disturbances, replacing the high-frequency focused disturbance function with the sensitive dimension focused disturbance function, which is the policy gradient sensitivity measurement function:

[0159]

[0160] D sim (χ,χ k) represents the semantic offset between the current perturbation trajectory and the original failure trajectory, ensuring that the perturbation behavior is still in the same attack path.

[0161] described and D sim (χ,χ k ) represent the sensitive dimension focused perturbation function and semantic consistency control, respectively. This optimization goal ensures that trajectory perturbations can continuously increase attack strength while maintaining consistency in the original semantics, making the vulnerability amplification process continuous, reproducible, and engineering-controllable.

[0162] Next, through continuous optimization and iteration using a vulnerability amplification optimization function, all amplified trace sequences are structured and written into the system stress testing sample set, ultimately generating a crash path perturbation trace sequence:

[0163] F vuln ={χ k ,χ k+1 ,…,χ k+N}

[0164] And associate the corresponding disturbance parameters, system responses, and vulnerability labels.

[0165] The crash path perturbation sequences generated under different environments and trajectories are combined into the system stress testing sample set. The system stress testing sample set can be used to generate a general strategy vulnerability model and support future algorithm test playback, model comparison, and failure label training tasks.

[0166] The strategy vulnerability amplification mechanism can automatically identify weak areas, extract key variables, build a perturbation closed loop, and maintain test semantic consistency and intervention executable after the algorithm fails. It greatly enhances the depth and robustness boundary detection capabilities of system testing, and provides stress testing support for the strategy collapse stage of the closed-loop adaptive testing process.

[0167] In addition, in some embodiments, an embodiment of step 103 is also included. Figure 4 As shown, at step S103, it also includes:

[0168] Performing structured collection of test samples for each round of evolutionary updates, wherein the collection objectives include perturbation strategy parameters, actor states, control response characteristics, and labeling results;

[0169] The sample coverage function is used to evaluate the results of structured collection to obtain sample quality and scene diversity. Based on the sample quality and scene diversity, a test sample library with semantic labels, structural mapping and multi-dimensional indexing is established.

[0170] Specifically, based on the test samples and system response records generated during the strategy evolution and vulnerability amplification stages, a test sample database with high-dimensional features, multi-stage labels, and iterative expansion capabilities is constructed to support strategy replay, model retraining, and cross-generation comparative analysis of subsequent autonomous driving algorithms, forming an adaptive stress testing feedback system with closed-loop characteristics.

[0171] The sample library takes the adversarial actor perturbation strategy as its input core and records the following information units in a structured manner: actor state sequence (initial trajectory, evolution trajectory, failure trajectory); perturbation parameters (perturbation direction, energy, rhythm, objective function weight); algorithm response characteristics (minimum distance, TTC, acceleration, delay, etc.); system labels (whether failure is triggered, failure type, robustness interval); historical sample mapping (semantic overlap and diversity indicators with previous samples).

[0172] The sample is organized as follows:

[0173]

[0174] Among them, χ i The trajectory of the actor in the i-th test scenario; represents the disturbance parameter; is the corresponding failure variable vector; i is the system response and label; T i Configure the topology and disturbance environment.

[0175] To further improve the system test coverage capability of the sample library, the system automatically performs multi-dimensional coverage measurement and sample redundancy control mechanisms during the sample writing phase, including:

[0176] Actor strategy space coverage: measures trajectory diversity and the breadth of perturbation distribution;

[0177] System strategy response boundary coverage: based on offset, delay, and control disturbance changes;

[0178] Risk type coverage: whether the marked sample covers different types of failures (such as decision instability, perception error, control failure, mixed type, etc.);

[0179] Typical traffic structure distribution coverage: whether it includes intersections, lane change sections, curves, roundabouts and other highly complex road sections.

[0180] The system introduces a structured sample scoring function:

[0181] C cov =λ1·D traj +λ2·D sem +λ3·D fail +λ4·D struct

[0182] The structured sample scoring function is used to comprehensively evaluate the coverage quality of the current sample set, automatically screen out low-value redundant samples, and improve test efficiency and model training data quality.

[0183] Therefore, the sample library provided by the embodiments of the present disclosure can serve as: a data source for algorithm strategy vulnerability analysis, a playback environment for adversarial strategy training and reinforcement learning stress testing models; a set of general benchmark scenarios for cross-model performance comparison; and a closed-loop feedback entry for subsequent modules to guide strategy re-evolution and test logic updates.

[0184] It should be noted that the embodiment disclosed in step 103 establishes a closed-loop sample generation and management system with a three-dimensional aggregation structure of behavior-variable-label, a coverage-driven update mechanism, and a system state feedback write-back path. It not only effectively supports the continuous testing of autonomous driving algorithms under extreme emergencies, but also provides practical data resources with high semantic density for strategy learning and optimization.

[0185] A second embodiment of the present invention further provides a scene simulation optimization device, the device comprising:

[0186] A scenario construction unit is configured to construct a test scenario with emergency characteristics using reinforcement learning, including a test scenario modeling module, an environment perturbation module, and an actor configuration module;

[0187] The testing unit is configured to, if the input autonomous driving algorithm is effective in the risk of the test scenario, evolve and update the trajectory perturbation strategy of the adversarial actor; if the autonomous driving algorithm fails in the test scenario or the evolved and updated trajectory perturbation strategy, continuously stress test and deeply verify the decision vulnerabilities of the autonomous driving algorithm based on the vulnerability amplification mechanism.

[0188] A third embodiment of the present invention further provides an electronic device, comprising:

[0189] at least one processor; and,

[0190] a memory communicatively connected to the at least one processor; wherein,

[0191] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the scene simulation optimization method of any of the aforementioned embodiments.

[0192] The fourth embodiment of the present invention further provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the scene simulation optimization method described in any of the aforementioned embodiments.

[0193] The fifth embodiment of the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, which, when executed by a computer, enable the computer to execute the scene simulation optimization method of any of the aforementioned embodiments.

[0194] Figure 5 A schematic diagram of a method or device 1000 that can implement an embodiment of the present invention is shown. In some embodiments, the method or device 1000 may include more or fewer devices than shown. In some embodiments, the method or device 1000 may be implemented using a single device or multiple devices. In some embodiments, the method or device 1000 may be implemented using cloud-based or distributed devices.

[0195] like Figure 5 As shown, device 1000 includes a processor 1001, which can perform various appropriate operations and processes according to the programs and / or data stored in a read-only memory (ROM) 1002 or the programs and / or data loaded from a storage portion 1008 into a random access memory (RAM) 1003. Processor 1001 can be a multi-core processor or can include multiple processors. In some embodiments, processor 1001 can include a general-purpose main processor and one or more special coprocessors, such as a central processing unit (CPU), a graphics processing unit (GPU), a neural network processor (NPU), a digital signal processor (DSP), etc. Various programs and data required for the operation of device 1000 are also stored in RAM 1003. Processor 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to bus 1004.

[0196] The processor and memory are used together to execute the program stored in the memory. When the program is executed by the computer, the methods, steps or functions described in the above embodiments can be implemented.

[0197] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, a touch screen, and the like; an output section 1007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1008 including a hard disk and the like; and a communication section 1009 including a network interface card such as a LAN card or a modem. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1010 as needed, so that computer programs read therefrom can be installed into the storage section 1008 as needed. Figure 5 Only some components are shown schematically, which does not mean that the device 1000 only includes Figure 5 Components shown.

[0198] The systems, devices, modules, or units described in the above embodiments may be implemented by a computer or its associated components. The computer may be, for example, a mobile terminal, a smartphone, a personal computer, a laptop computer, an in-vehicle human-computer interaction device, a personal digital assistant, a media player, a navigation device, a game console, a tablet computer, a wearable device, a smart TV, an Internet of Things system, a smart home, an industrial computer, a server, or a combination thereof.

[0199] Although not shown, in an embodiment of the present invention, a computer-readable storage medium is provided, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the method for scene simulation optimization described in the embodiment is implemented.

[0200] Storage media in embodiments of the present invention include permanent and non-permanent, removable and non-removable items that can be used to store information using any method or technology. Examples of storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0201] Although not shown, an embodiment of the present invention further provides a computer program product, including: a computer program / instruction, which implements the scene simulation optimization method described in the embodiment when the computer program / instruction is executed by a processor.

[0202] The methods, programs, systems, and apparatuses of the embodiments of the present invention may be executed or implemented in a single or multiple networked computers, or may be practiced in a distributed computing environment. In the embodiments of this specification, in these distributed computing environments, tasks may be performed by remote processing devices connected via a communication network.

[0203] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, those skilled in the art will appreciate that the functional modules / units or controllers and related method steps described in the above embodiments may be implemented using software, hardware, or a combination of software / hardware.

[0204] Unless explicitly stated, the actions or steps of the methods, procedures, and methods described in accordance with the embodiments of the present invention do not have to be performed in a specific order and can still achieve the desired results. In some implementations, multitasking and parallel processing are also possible or may be advantageous.

[0205] In this document, multiple embodiments of the present invention are described, but for the sake of brevity, the description of each embodiment is not exhaustive, and the same or similar features or parts between the embodiments may be omitted. In this document, "one embodiment", "some embodiments", "example", "specific example", or "some examples" are intended to apply to at least one embodiment or example according to the present invention, but not all embodiments. The above terms do not necessarily mean to refer to the same embodiment or example. Those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are mutually contradictory.

[0206] While the exemplary systems and methods of the present invention have been specifically shown and described with reference to the foregoing embodiments, these are merely examples of the best modes for implementing the present systems and methods. Those skilled in the art will appreciate that various changes may be made to the embodiments of the systems and methods described herein when implementing the present systems and / or methods without departing from the spirit and scope of the present invention as defined in the appended claims.

[0207] In addition, the method, system, device, and medium for scene simulation optimization according to the present invention can also be implemented in the following manner:

[0208] (1) A method for scene simulation optimization, characterized by comprising:

[0209] Reinforcement learning is used to construct a test scenario with emergency characteristics, including a test scenario modeling module, an environmental disturbance module, and an actor configuration module;

[0210] If the input autonomous driving algorithm is effective in the risk of the test scenario, the trajectory perturbation strategy of the adversarial actor will be evolved and updated; if the autonomous driving algorithm fails in the test scenario or the evolutionarily updated trajectory perturbation strategy, the decision-making vulnerabilities of the autonomous driving algorithm will be continuously stress-tested and deeply verified based on the vulnerability amplification mechanism.

[0211] (2) The method according to (1), characterized in that the scenario modeling module generates an editable topological structure of the road space based on a structural component library and a topological splicing generator, and splices a test layout area with induction potential based on the test target, wherein the test layout area includes a road topological structure, and a traffic constraint and a strategic path solution space of the road topological structure, and the road topological structure includes intersections, occlusion sections, and narrow passages formed by dynamic splicing;

[0212] The environmental disturbance module is used to introduce multi-dimensional disturbance factors to simulate real-world environmental factors. The multi-dimensional disturbance factors include weather, lighting, visibility, ground adhesion, and sensor noise. During the disturbance process, the disturbance control graph is used to map the disturbance to the road topology, integrating the disturbance information and the road topology information.

[0213] The actor configuration module generates traffic participants based on the test layout area and the test target. The traffic participants include regular actors and antagonistic actors. The regular actors construct the basic traffic flow density, and the antagonistic actors trigger high-risk behaviors through strategy templates.

[0214] (3) The method according to (1) is characterized in that when the input autonomous driving algorithm avoids the risk of the test scenario, a proximal policy optimization method is used to update the trajectory perturbation strategy of the adversarial actor.

[0215] (4) The method according to (3), characterized in that the proximal strategy optimization method includes:

[0216] Construct a policy margin function based on the main vehicle response index of the autonomous driving algorithm;

[0217] Dynamically adjusting disturbance parameters based on the policy margin function, wherein the disturbance parameters include attack rhythm, direction change, and combination behavior;

[0218] The test scenario evolves round by round, and the intensity of the strategy pressure is increased step by step until the system response boundary is approached and the critical state of stability is exposed.

[0219] (5) The method according to (1), characterized in that the decision vulnerabilities of the autonomous driving algorithm are continuously stress-tested and deeply verified based on the vulnerability amplification mechanism, including:

[0220] Through abnormal trajectory clustering, failure variable extraction and control sensitivity analysis, key control variables and vulnerable response points in the failure path are located, and multi-round disturbance trajectory belts are constructed;

[0221] Guided by policy gradient projection, the multiple rounds of perturbation trajectory are used to construct a crash path perturbation sequence by superimposing high-frequency focused perturbation functions. Combined with the semantic consistency control mechanism, high-voltage test fragments for continuous stress testing and deep verification are generated.

[0222] (6) The method according to (5), characterized in that constructing the multi-round disturbance trajectory band includes:

[0223] Structural reconstruction of the process of risk avoidance, tracing back the main vehicle's decision-making response chain, extracting the key dynamic variables that cause failure, and forming the failure path state feature vector;

[0224] A strategy failure region extraction mechanism is introduced, combined with trajectory trend analysis and anomaly detection models, to identify the most triggering time periods and state segments in the failure path state feature vector, marking them as strategy vulnerable windows.

[0225] An iterative trajectory perturbation enhancement method is used for the failure fragile window, and the perturbation is concentratedly injected into the most strategically sensitive time period and control variable dimension to obtain a multi-round perturbation trajectory band.

[0226] (7) The method according to (1), further comprising:

[0227] Performing structured collection of test samples for each round of evolutionary updates, wherein the collection objectives include perturbation strategy parameters, actor states, control response characteristics, and labeling results;

[0228] The sample coverage function is used to evaluate the results of structured collection to obtain sample quality and scene diversity. Based on the sample quality and scene diversity, a test sample library with semantic labels, structural mapping and multi-dimensional indexing is established.

[0229] (8) A scene simulation optimization device, based on the scene simulation optimization method described in any one of (1) to (7), comprising:

[0230] A scenario construction unit is configured to construct a test scenario with emergency characteristics using reinforcement learning, including a test scenario modeling module, an environment perturbation module, and an actor configuration module;

[0231] The testing unit is configured to, if the input autonomous driving algorithm is effective in the risk of the test scenario, evolve and update the trajectory perturbation strategy of the adversarial actor; if the autonomous driving algorithm fails in the test scenario or the evolved and updated trajectory perturbation strategy, continuously stress test and deeply verify the decision vulnerabilities of the autonomous driving algorithm based on the vulnerability amplification mechanism.

[0232] (9) The apparatus according to (8), wherein the scenario construction unit is further configured such that the scenario modeling module generates an editable topological structure of the road space based on the structural component library and the topological splicing generator, and splices a test layout area with induction potential based on the test target, wherein the test layout area includes a road topological structure, and a traffic constraint and a strategic path solution space of the road topological structure, wherein the road topological structure includes intersections, occlusion segments, and narrow passages formed by dynamic splicing;

[0233] The environmental disturbance module is used to introduce multi-dimensional disturbance factors to simulate real-world environmental factors. The multi-dimensional disturbance factors include weather, lighting, visibility, ground adhesion, and sensor noise. During the disturbance process, the disturbance control graph is used to map the disturbance to the road topology, integrating the disturbance information and the road topology information.

[0234] The actor configuration module generates traffic participants based on the test layout area and the test target. The traffic participants include regular actors and antagonistic actors. The regular actors construct the basic traffic flow density, and the antagonistic actors trigger high-risk behaviors through strategy templates.

[0235] (10) According to the device described in (8), the test unit is further configured to use a proximal policy optimization method to update the trajectory perturbation strategy of the adversarial actor when the input autonomous driving algorithm avoids the risk of the test scenario.

[0236] (11) According to the apparatus described in (10), the testing unit is further configured such that the proximal policy optimization method includes:

[0237] Construct a policy margin function based on the main vehicle response index of the autonomous driving algorithm;

[0238] Dynamically adjusting disturbance parameters based on the policy margin function, wherein the disturbance parameters include attack rhythm, direction change, and combination behavior;

[0239] The test scenario evolves round by round, and the intensity of the strategy pressure is increased step by step until the system response boundary is approached and the critical state of stability is exposed.

[0240] (12) According to the apparatus described in (8), the testing unit is further configured to:

[0241] The vulnerability amplification mechanism is used to continuously stress test and deeply verify the decision vulnerabilities of the autonomous driving algorithm, including:

[0242] Through abnormal trajectory clustering, failure variable extraction and control sensitivity analysis, key control variables and vulnerable response points in the failure path are located, and multi-round disturbance trajectory belts are constructed;

[0243] Guided by policy gradient projection, the multiple rounds of perturbation trajectory are used to construct a crash path perturbation sequence by superimposing high-frequency focused perturbation functions. Combined with the semantic consistency control mechanism, high-voltage test fragments for continuous stress testing and deep verification are generated.

[0244] (13) According to the apparatus described in (12), the testing unit is further configured to construct the multi-round disturbance trajectory band including:

[0245] Structural reconstruction of the process of risk avoidance, tracing back the main vehicle's decision-making response chain, extracting the key dynamic variables that cause failure, and forming the failure path state feature vector;

[0246] A strategy failure region extraction mechanism is introduced, combined with trajectory trend analysis and anomaly detection models, to identify the most triggering time periods and state segments in the failure path state feature vector, marking them as strategy vulnerable windows.

[0247] An iterative trajectory perturbation enhancement method is used for the failure fragile window, and the perturbation is concentratedly injected into the most strategically sensitive time period and control variable dimension to obtain a multi-round perturbation trajectory band.

[0248] (14) The apparatus according to (8), wherein the pair testing unit is further configured to perform structured collection of test samples for each round of evolutionary updates, wherein the collection targets include perturbation strategy parameters, actor states, control response characteristics, and labeling results;

[0249] The sample coverage function is used to evaluate the results of structured collection to obtain sample quality and scene diversity. Based on the sample quality and scene diversity, a test sample library with semantic labels, structural mapping and multi-dimensional indexing is established.

[0250] (15) An electronic device, characterized in that the electronic device comprises:

[0251] at least one processor; and,

[0252] a memory communicatively connected to the at least one processor; wherein,

[0253] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the scene simulation optimization method described in any one of (1) to (7).

[0254] (16) A non-transitory computer-readable storage medium, characterized in that the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the scene simulation optimization method described in any one of (1) to (7).

[0255] (17) A computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, wherein the computer program comprises program instructions that, when executed by a computer, cause the computer to execute the method for scene simulation optimization described in any one of (1) to (7).

Claims

1. A method for scene simulation optimization, characterized in that: include: Reinforcement learning is used to construct a test scenario with emergency characteristics, including a test scenario modeling module, an environmental disturbance module, and an actor configuration module; If the input autonomous driving algorithm is effective in the risk of the test scenario, then evolving and updating the trajectory perturbation strategy of the adversarial actor; If the autonomous driving algorithm fails in the test scenario or the trajectory perturbation strategy after evolution and update, the decision vulnerabilities of the autonomous driving algorithm will be continuously stress-tested and deeply verified based on the vulnerability amplification mechanism.

2. The method according to claim 1, characterized in that The scenario modeling module generates an editable topology of the road space based on a structural component library and a topology splicing generator, and splices a test layout area with induction potential based on the test target. The test layout area includes the road topology structure, as well as the traffic constraints and strategic path solution space of the road topology structure. The road topology structure includes intersections, occlusion segments, and narrow passages formed by dynamic splicing; The environmental disturbance module is used to introduce multi-dimensional disturbance factors to simulate real-world environmental factors. The multi-dimensional disturbance factors include weather, lighting, visibility, ground adhesion, and sensor noise. During the disturbance process, the disturbance control graph is used to map the disturbance to the road topology, integrating the disturbance information and the road topology information. The actor configuration module generates traffic participants based on the test layout area and the test target. The traffic participants include regular actors and antagonistic actors. The regular actors construct the basic traffic flow density, and the antagonistic actors trigger high-risk behaviors through strategy templates.

3. The method according to claim 1, characterized in that When the input autonomous driving algorithm avoids the risks of the test scenario, the proximal policy optimization method is used to update the trajectory perturbation strategy of the adversarial actor.

4. The method according to claim 3, characterized in that The proximal strategy optimization method includes: Construct a policy margin function based on the main vehicle response index of the autonomous driving algorithm; Dynamically adjusting disturbance parameters based on the policy margin function, wherein the disturbance parameters include attack rhythm, direction change, and combination behavior; The test scenario evolves round by round, and the intensity of the strategy pressure is increased step by step until the system response boundary is approached and the critical state of stability is exposed.

5. The method according to claim 1, characterized in that The vulnerability amplification mechanism is used to continuously stress test and deeply verify the decision vulnerabilities of the autonomous driving algorithm, including: Through abnormal trajectory clustering, failure variable extraction and control sensitivity analysis, key control variables and vulnerable response points in the failure path are located, and multi-round disturbance trajectory belts are constructed; Guided by policy gradient projection, the multiple rounds of perturbation trajectory are used to construct a crash path perturbation sequence by superimposing high-frequency focused perturbation functions. Combined with the semantic consistency control mechanism, high-voltage test fragments for continuous stress testing and deep verification are generated.

6. The method according to claim 5, characterized in that Constructing the multi-round disturbance trajectory band includes: Structural reconstruction of the process of risk avoidance, tracing back the main vehicle's decision-making response chain, extracting the key dynamic variables that cause failure, and forming the failure path state feature vector; A strategy failure region extraction mechanism is introduced, combined with trajectory trend analysis and anomaly detection models, to identify the most triggering time periods and state segments in the failure path state feature vector, marking them as strategy vulnerable windows. An iterative trajectory perturbation enhancement method is used for the failure fragile window, and the perturbation is concentratedly injected into the most strategically sensitive time period and control variable dimension to obtain a multi-round perturbation trajectory band.

7. The method according to claim 1, characterized in that Also includes: Performing structured collection of test samples for each round of evolutionary updates, wherein the collection objectives include perturbation strategy parameters, actor states, control response characteristics, and labeling results; The sample coverage function is used to evaluate the results of structured collection to obtain sample quality and scene diversity. Based on the sample quality and scene diversity, a test sample library with semantic labels, structural mapping and multi-dimensional indexing is established.

8. A device for scene simulation optimization, characterized in that: The method for scene simulation optimization according to any one of claims 1 to 7, wherein the device comprises: A scenario construction unit is configured to construct a test scenario with emergency characteristics using reinforcement learning, including a test scenario modeling module, an environment perturbation module, and an actor configuration module; The testing unit is configured to, if the input autonomous driving algorithm is effective in the risk of the test scenario, evolve and update the trajectory perturbation strategy of the adversarial actor; if the autonomous driving algorithm fails in the test scenario or the evolved and updated trajectory perturbation strategy, continuously stress test and deeply verify the decision vulnerabilities of the autonomous driving algorithm based on the vulnerability amplification mechanism.

9. An electronic device, characterized in that: The electronic device includes: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the scene simulation optimization method described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores computer instructions, which are used to enable the computer to execute the scene simulation optimization method described in any one of claims 1 to 7.

Citation Information

Cited By

  • Integrated training test method for automatic driving algorithm

    CN121029622A

  • Site test scene construction method based on feature matching and trajectory optimization

    CN121031143A

  • Driving assistance testing method and system in complex scene, medium and equipment

    CN121830075A

  • Self-adaptive decision control method for autonomous vehicle considering disturbance factors

    CN122131608A