Decision planning system-oriented simulation test scene automatic generation method
Through the combination of deep learning, reinforcement learning and Bayesian optimization, the simulation test scenarios of decision planning systems are automatically generated, which solves the problems of low efficiency and limited coverage in the existing technology, and realizes efficient and diversified scenario generation and system robustness testing.
Patent Information
- Application Number
- CN202510732304.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, the generation of simulation test scenarios of decision planning systems relies on manual design, with low efficiency, limited coverage, lack of diversity, and cannot comprehensively test the robustness of the system.
The deep learning model is used to build a scenario model, use reinforcement learning to generate constraints, combine it with Bayesian optimization algorithm to generate scene parameters that meet the constraints, and evaluate the system performance through deep reinforcement learning to automatically generate diverse and high coverage test scenarios.
It improves the efficiency of scene generation, enhances scene diversity and coverage, comprehensively tests the robustness of the system, ensures that the generated scenarios meet system needs, and provides multi-dimensional system performance evaluation.
Smart Images

Figure CN120470937A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving simulation technology, and more specifically, to a method for automatically generating simulation test scenarios for a decision-making planning system. Background Art
[0002] Decision-making and planning systems are widely used in autonomous driving, robot path planning, intelligent transportation, and other fields. To ensure the reliability and safety of these systems in practical applications, they must be fully simulated and tested.
[0003] Traditional simulation test scenario generation methods mainly rely on manual design, which has the following problems: low efficiency: Manually designing scenarios is time-consuming and labor-intensive, making it difficult to meet large-scale testing needs; limited coverage: Manually designed scenarios are often limited to known typical situations and find it difficult to cover complex edge scenarios; lack of diversity: The scenarios lack diversity and randomness, making it impossible to fully test the robustness of the system.
[0004] Therefore, there is an urgent need for an automatic generation method of simulation test scenarios for decision-making planning systems. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for automatically generating simulation test scenarios for decision-making planning systems to solve the problems in the above-mentioned prior art, reduce manual intervention through automated methods, and improve the efficiency of scenario generation; generate diversified and high-coverage test scenarios, enhance scenario diversity, and comprehensively test the robustness of the system.
[0006] The present invention provides a method for automatically generating simulation test scenarios for a decision-making planning system, which includes:
[0007] Scenario modeling: Use deep learning models to build scenario models;
[0008] Constraint generation: Based on the output of the scenario model, a reinforcement learning model is used to generate constraint conditions;
[0009] Optimization search: Based on the scenario model and the constraints, a Bayesian optimization algorithm is used to generate scenario parameters that meet the constraints;
[0010] Scenario Evaluation: Evaluate system performance using deep reinforcement learning models.
[0011] In the above-mentioned method for automatically generating simulation test scenarios for a decision-making planning system, preferably, the scenario modeling: constructing a scenario model using a deep learning model, includes:
[0012] The scenario modeling module inputs historical scenario data, uses a deep learning model to capture long-term dependencies through a self-attention mechanism, connects system requirements with latent variables, and generates scenario parameters, including:
[0013] Obtaining input data and preprocessing the input data;
[0014] Model construction and training: The network structure of the model is that the encoder uses a convolutional neural network, the decoder is symmetrical with the encoder, and the input data is used to train the deep learning model;
[0015] Scene generation: Generate new scene feature vectors using the trained deep learning model;
[0016] Post-processing: Convert the generated new scene feature vector into a visualization scene.
[0017] In the above-mentioned method for automatically generating simulation test scenarios for a decision-making planning system, preferably, the step of obtaining input data and preprocessing the input data includes:
[0018] Inputting a historical scene data set, wherein the historical scene data set includes environmental parameters, states of dynamic objects, and system requirements;
[0019] Extracting features based on the historical scene dataset to encode the historical scene dataset into structured features, wherein the structured features include a numerical representation of at least one of a spatial feature, a dynamic feature, and an environmental feature;
[0020] Data standardization: normalize the input data.
[0021] The model construction and training: The network structure of the model is that the encoder adopts a convolutional neural network, the decoder is symmetrical with the encoder, and the input data is used to train the deep learning model, including:
[0022] Construct network structure: the encoder is a convolutional neural network, the input is the scene feature vector, the output is the mean and variance of the latent space, the decoder is symmetrical with the encoder, and the input is the latent variable , the output is the reconstructed scene feature vector;
[0023] Construct a reconstruction loss function and use the mean square error to measure the difference between the original data and the reconstructed data;
[0024] The KL divergence constraint constrains the latent space distribution to be close to the standard normal distribution;
[0025] According to the reconstruction loss function and KL divergence, the total loss function is constructed.
[0026] The total loss function is expressed by the following formula:
[0027] ,
[0028] Among them, L represents the total loss function, which is an indicator for comprehensively measuring the performance of the model, and x represents the original data. Represents the data after model reconstruction, β represents the balance coefficient, which is an artificially set hyperparameter used to adjust the proportion of reconstruction loss and KL divergence loss in the total loss, KL represents KL divergence, which is used to measure the difference between two probability distributions, N(μ,σ²): represents the normal distribution with mean μ and variance σ², representing the distribution of the latent space, μ represents the mean of the normal distribution, describing the center position of the distribution, σ² represents the variance of the normal distribution, describing the degree of dispersion of the distribution, N(0, 1) represents the standard normal distribution with mean 0 and variance 1, which serves as the constraint target of the latent space distribution;
[0029] Training process: Use historical scene data to train the deep learning model, and in the training process, optimize the parameters of the deep learning model through the total loss function. In the training process, the optimizer is Adam optimizer, and the learning rate is , the batch size is 128.
[0030] In the above-mentioned method for automatically generating simulation test scenarios for decision-making planning systems, preferably, the scenario generation: generating a new scenario feature vector using a trained deep learning model, includes:
[0031] Latent space sampling: randomly sample latent variables from a standard normal distribution N(0,1) ;
[0032] Conditional injection: System requirements and environmental parameters are used as conditional inputs, and potential variables are After splicing, input the decoder;
[0033] Decoding output: The decoder generates a new scene feature vector,
[0034] The post-processing: converting the generated new scene feature vector into a visual scene, including:
[0035] Parameter denormalization: restore the generated eigenvectors to actual physical quantities;
[0036] Scene visualization: The restored actual physical quantities are converted into a visualization scene through simulation tools, wherein the simulation tools include CARLA or SUMO.
[0037] In the above-mentioned method for automatically generating simulation test scenarios for decision-making planning systems, preferably, the constraint generation: based on the output results of the scenario model, using a reinforcement learning model to generate constraint conditions, includes:
[0038] The constraint generation module inputs the scene feature vector output by the scene model, uses the reinforcement learning model network to select the optimal constraint action by maximizing the Q value, and outputs mathematical constraint conditions, specifically including:
[0039] Problem modeling includes: constructing a state space: the state space includes the feature vector of the scene model; constructing an action space: generating operations with constraints; determining a reward function: the reward function includes positive rewards and negative rewards;
[0040] Determine the network structure of the reinforcement learning model network: the input layer of the reinforcement learning model network is the state vector; the hidden layer is a 3-layer fully connected network, the number of nodes in each fully connected layer is 256, 128 and 64 nodes respectively, and the activation function is ReLU; the output layer is the Q value of the action space;
[0041] Determine the training process of the reinforcement learning model network;
[0042] Constraint generation determines the optimal constraint action based on the Q-value maximization strategy and outputs mathematical constraint conditions based on the optimal constraint action, including: online reasoning: input the feature vector of the scene model and select the action with the largest Q value as the optimal constraint condition; constraint representation: map the action with the largest Q value into a mathematical expression.
[0043] The above-mentioned method for automatically generating simulation test scenarios for decision-making planning systems, wherein, preferably,
[0044] The training process of determining the reinforcement learning model network includes:
[0045] Experience replay: transferring historical state Stored in the replay buffer, where s represents the state of the agent at a certain moment, a represents the action taken by the agent in state s, r represents the reward obtained by the agent after performing action a, and s' represents the next state the agent transfers to after performing action a;
[0046] Determine the target network: regularly synchronize the main network parameters to the target network to stabilize training;
[0047] Determine the training steps, including: sampling a batch of experience data from the playback buffer; calculating the target Q value using the following formula: , where r represents the reward obtained after executing the action, γ: discount factor, ranging from [0, 1], used to measure the importance of future rewards at the current moment. The closer γ is to 1, the more importance is attached to future rewards. Represents the Q value function calculated by the target Q network, which is used to calculate the target Q value, y represents the calculated target Q value, and Q represents the Q value function calculated by the current Q network; the minimization loss function is calculated using the following formula: , where L represents the loss function, Indicates the batch size; updates the target network parameters every preset number of steps.
[0048] In the above-mentioned method for automatically generating simulation test scenarios for a decision-making planning system, preferably, the optimization search: based on the scenario model and the constraints, using a Bayesian optimization algorithm to generate scenario parameters that meet the constraints, includes:
[0049] The parameter space and the constraints generated based on the reinforcement learning model are input into the optimization search module, and the Bayesian optimization iterative method is used to output the optimal scenario parameter combination that meets the constraints.
[0050] In the above-mentioned method for automatically generating simulation test scenarios for a decision-making planning system, preferably, the parameter space and the constraints generated based on the reinforcement learning model are input into the optimization search module, and the Bayesian optimization iterative method is used to output the optimal scenario parameter combination that satisfies the constraints, including:
[0051] Defining a parameter space includes: defining a search dimension, wherein the search dimension includes a scene parameter; defining a constraint condition, wherein the constraint condition includes a hard constraint and a soft constraint generated by a reinforcement learning model;
[0052] Design objective functions, including: single-objective optimization: maximizing scenario coverage as the objective function, where the maximization of scenario coverage is defined as the proportion of the test scenario covering the system state space; multi-objective optimization: optimizing coverage simultaneously and system response time , the objective function is to maximize the coverage C and minimize the system response time T;
[0053] performing a Bayesian optimization step based on the parameter space and the objective function;
[0054] The step of performing Bayesian optimization according to the parameter space and the objective function includes:
[0055] Gaussian process modeling, including: determining the kernel function: selecting the Marten kernel to handle nonlinear relationships; hyperparameter optimization: optimizing kernel parameters through maximum likelihood estimation;
[0056] Select acquisition functions, including: balancing exploration and exploitation through expected improvement formulas,
[0057] The expected improvement formula is as follows:
[0058] ,
[0059] in, Represents the expected improvement function value, which is used to measure the expected improvement that can be brought about by sampling at point x. It represents the expected value function, which is the probability weighted average of the random variable value. represents the independent variable, representing the point considered for sampling in the search space, It represents the function value of the objective function at point x. Represents the currently known optimal solution, that is, the value of the independent variable corresponding to the point where the objective function value is optimal among the sampled points. Indicates the current optimal solution The corresponding objective function value;
[0060] Iterative optimization includes: initialization: randomly sampling multiple sets of parameters as initial observation points; modeling: fitting the objective function with a Gaussian process; selecting the next parameter point: selecting the parameter with the greatest benefit through the EI function; evaluation: running the scenario in a simulation environment and calculating the objective function value; updating: adding new parameters and results to the observation data set; stopping the iterative optimization when the cutoff condition is met: the cutoff condition includes the number of iterations reaching a preset iteration threshold or the objective function convergence.
[0061] In the above-mentioned method for automatically generating simulation test scenarios for decision-making and planning systems, preferably, the scenario evaluation: evaluating system performance using a deep reinforcement learning model, includes:
[0062] The optimized scenario parameters are input into the scenario evaluation module, and a deep reinforcement learning model is used to output a multi-dimensional performance evaluation report. If the scenario evaluation fails, the optimization search module is returned to adjust the parameters. When the optimization search cannot meet the constraints, the constraint generation module is triggered to regenerate the constraints.
[0063] The above-mentioned method for automatically generating simulation test scenarios for decision-making planning systems, wherein preferably, the optimized scenario parameters are input into the scenario evaluation module, a deep reinforcement learning model is used, and a multi-dimensional performance evaluation report is output. If the scenario evaluation fails, the optimization search module is returned to adjust the parameters. When the optimization search cannot satisfy the constraints, the constraint generation module is triggered to regenerate the constraints, specifically including:
[0064] Constructing the simulation environment, including: Interface design: connecting the decision-making planning system and the simulation platform through API; State input: inputting scene parameters and environment parameters; Action output: outputting control instructions of the decision-making planning system;
[0065] Design a deep reinforcement learning algorithm, including: Design a policy network: The input of the policy network includes: environment state , the output includes: action probability distribution; design value network: the input of the value network includes: environment state ,The output includes: state value;
[0066] Design loss functions, including: designing strategy loss function, value loss function and entropy regularization function respectively;
[0067] Executing an evaluation process includes: training an agent: running a deep reinforcement learning algorithm in a generated scenario to optimize a decision-making strategy; determining performance indicators, wherein the performance indicators include safety, efficiency, and comfort, wherein the safety includes at least one of the number of collisions and the frequency of sudden braking; the efficiency includes at least one of the average vehicle speed and the path deviation distance; and the comfort includes the standard deviation of the lateral acceleration;
[0068] Generating an evaluation report includes generating a PDF report, wherein the PDF report includes an indicator comparison curve and a key scenario analysis compared with a baseline method.
[0069] The present invention provides an automated method for generating simulation test scenarios for decision-making and planning systems. This method uses a deep neural network (Double DQN) to generate constraints, and jointly generates constraints through VAE and Double DQN. This automated method reduces manual intervention and improves scenario generation efficiency. The method generates diverse and high-coverage test scenarios, enhances scenario diversity, and comprehensively tests the robustness of the system. The method optimizes constraint processing by efficiently handling complex constraints, ensuring that the generated scenarios meet system requirements. The method also provides a comprehensive system performance evaluation method, implements multi-dimensional evaluation, and reflects the system's performance in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described below with reference to the accompanying drawings, in which:
[0071] Figure 1 A flowchart of an embodiment of a method for automatically generating simulation test scenarios for a decision-making planning system provided by the present invention;
[0072] Figure 2 A flowchart of the overall architecture of an embodiment of the method for automatically generating simulation test scenarios for a decision-making planning system provided by the present invention;
[0073] Figure 3 Flowchart for modeling VAE scenarios;
[0074] Figure 4 Graph of the Bayesian optimization iterative process. DETAILED DESCRIPTION
[0075] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. The description of the exemplary embodiments is merely illustrative and is in no way intended to limit the present disclosure, its application, or use. The present disclosure can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to make the present disclosure thorough and complete and to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that unless otherwise specifically stated, the relative arrangement of parts and steps, the composition of materials, numerical expressions, and numerical values set forth in these embodiments should be interpreted as being merely exemplary and not as limiting.
[0076] The terms "first," "second," and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are simply used to distinguish different parts. Terms such as "include" or "comprising" mean that the elements preceding the term include the elements listed after the term, and do not exclude the possibility of also including other elements. Terms such as "upper," "lower," and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0077] In the present disclosure, when a specific component is described as being located between a first component and a second component, there may or may not be an intervening component between the specific component and the first component or the second component. When a specific component is described as being connected to another component, the specific component may be directly connected to the other component without an intervening component, or may not be directly connected to the other component but have an intervening component.
[0078] All terms (including technical or scientific terms) used in this disclosure have the same meaning as those understood by one of ordinary skill in the art to which this disclosure belongs, unless otherwise specifically defined. It should also be understood that terms defined in, for example, general dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an idealized or highly formal sense, unless explicitly defined herein.
[0079] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.
[0080] In recent years, with the development of artificial intelligence and simulation technologies, automated test scenario generation methods have become a research hotspot. However, existing technologies still suffer from the following shortcomings: low scenario generation quality: generated scenarios often lack realism and complexity; limited constraint processing capabilities: inefficient handling of complex constraints; and single evaluation methods: lacking a multi-dimensional assessment of system performance. In light of these issues, the present invention provides a method for automated simulation test scenario generation for decision-making and planning systems.
[0081] like Figure 1 and Figure 2 As shown, the method for automatically generating simulation test scenarios for a decision-making planning system provided in this embodiment specifically includes the following steps during actual execution:
[0082] Step S1: Scenario modeling: Use the deep learning model (VAE) to build a scenario model.
[0083] Specifically, historical scenario data (such as road topology, vehicle trajectory, and weather parameters) are input into the scenario modeling module. Using a deep learning model, the self-attention mechanism is used to capture long-term dependencies, and system requirements (such as "rainy days") are spliced with latent variables to generate scenario parameters (such as the vehicle's initial position and the road friction coefficient).
[0084] like Figure 3 As shown, in one embodiment of the method for automatically generating simulation test scenarios for a decision planning system of the present invention, step S1 may specifically include:
[0085] Step S11: Acquire input data and pre-process the input data.
[0086] In one embodiment of the method for automatically generating simulation test scenarios for a decision planning system of the present invention, step S11 may specifically include:
[0087] Step S111: Input a historical scene dataset (such as traffic scene data in autonomous driving, etc.), which includes environmental parameters (such as road topology, weather conditions, etc.), the status of dynamic objects (such as vehicles and pedestrians) (such as position, speed, acceleration, etc.) and system requirements (such as speed limit rules).
[0088] Among them, the historical scenario data set needs to cover a variety of edge cases, such as extreme weather and sudden obstacles.
[0089] Step S112: extract features based on the historical scene dataset to encode the historical scene dataset into structured features, wherein the structured features include a numerical representation of at least one of spatial features, dynamic features, and environmental features.
[0090] For example, spatial features can be the geometric shape of the road (lane curvature, intersection layout), dynamic features can be time series data of vehicle trajectories, and environmental features can be numerical representations of weather parameters (visibility, road friction coefficient).
[0091] Step S113: Data normalization: performing normalization processing on the input data.
[0092] Min-Max normalization or Z-Score normalization can be used to eliminate dimensional differences through normalization processing.
[0093] Step S12, model construction and training: The network structure of the model is that the encoder (Encoder) adopts a convolutional neural network (CNN), the decoder (Decoder) is symmetrical with the encoder, and the input data is used to train the deep learning model.
[0094] In one embodiment of the method for automatically generating simulation test scenarios for a decision planning system of the present invention, step S12 may specifically include:
[0095] Step S121: Construct a network structure: the encoder is a convolutional neural network (CNN), the input is the scene feature vector, the output is the mean (μ) and variance (σ) of the latent space, the decoder is symmetrical with the encoder (i.e., the CNN is symmetrical with the encoder), and the input latent variable The splicing result with the system requirements is output as the reconstructed scene feature vector.
[0096] Step S122: construct a reconstruction loss function (Reconstruction Loss), and use the mean square error (MSE) to measure the difference between the original data and the reconstructed data.
[0097] Step S123: Constrain the latent space distribution to be close to the standard normal distribution through KL divergence.
[0098] Step S124: Construct a total loss function based on the reconstruction loss function and the KL divergence.
[0099] The total loss function is expressed by the following formula:
[0100] ,
[0101] Among them, L represents the total loss function, which is an indicator for comprehensively measuring the performance of the model, and x represents the original data. Represents the data after model reconstruction, β represents the balance coefficient, which is an artificially set hyperparameter used to adjust the proportion of reconstruction loss and KL divergence loss in the total loss, KL represents KL divergence, which is used to measure the difference between two probability distributions, N(μ,σ²): represents the normal distribution with mean μ and variance σ², representing the distribution of the latent space, μ represents the mean of the normal distribution, describing the center position of the distribution, σ² represents the variance of the normal distribution, describing the degree of dispersion of the distribution, N(0, 1) represents the standard normal distribution with mean 0 and variance 1, which serves as the constraint target of the latent space distribution.
[0102] Step S125, training process: Use historical scene data to train the deep learning model (VAE), and during the training process, optimize the parameters of the deep learning model through the total loss function.
[0103] Among them, during the training process, the optimizer is Adam optimizer, and the learning rate is set to , the batch size (BatchSize) is 128.
[0104] Step S13: Scenario generation: Generate a new scenario feature vector using the trained deep learning model.
[0105] In one embodiment of the method for automatically generating simulation test scenarios for a decision planning system of the present invention, step S13 may specifically include:
[0106] Step S131, latent space sampling: randomly sampling latent variables from the standard normal distribution N(0,1) .
[0107] Step S132, condition injection: system requirements (such as "rainy day scene") and environmental parameters (such as road length) are used as condition inputs and compared with potential variables After splicing, input the decoder.
[0108] Step S133, decoding output: the decoder generates a new scene feature vector.
[0109] The scene feature vector can be, for example, a dynamic object parameter (the vehicle's initial position ,speed ) and environmental parameters (road slip coefficient ,visibility ).
[0110] Step S14, post-processing: converting the generated new scene feature vector into a visual scene.
[0111] In one embodiment of the method for automatically generating simulation test scenarios for a decision planning system of the present invention, step S14 may specifically include:
[0112] Step S141, parameter denormalization: restore the generated eigenvector to the actual physical quantity.
[0113] Step S142: scene visualization: converting the restored actual physical quantities into a visualization scene through simulation tools (such as CARLA and SUMO).
[0114] Therefore, in the VAE scene modeling process, the encoder compresses the input scene into a latent variable z, and the decoder reconstructs the new scene in combination with the conditional input (environmental parameters).
[0115] Through latent space modeling and KL divergence constraints, the deep learning model (VAE) can generate physically accurate and diverse scene data. The introduction of the self-attention mechanism further enhances the model's ability to capture long-term dependencies, ensuring temporal coherence in the generated scene parameters (such as vehicle trajectories and weather conditions).
[0116] Step S2: Constraint generation: Based on the output results of the scenario model, a reinforcement learning model (DoubleDQN) is used to generate constraint conditions.
[0117] Specifically, the scene feature vector output by the scene model is input into the constraint generation module, and the reinforcement learning model (Double DQN) network is used to select the optimal constraint action by maximizing the Q value, and output mathematical constraint conditions (such as In one embodiment of the method for automatically generating simulation test scenarios for a decision planning system of the present invention, step S2 may specifically include:
[0118] Step S21: Problem modeling.
[0119] In one embodiment of the method for automatically generating simulation test scenarios for a decision planning system of the present invention, step S21 may specifically include:
[0120] Step S211 , constructing a state space (State): the state space includes feature vectors of the scene model (such as vehicle position, speed, and road topology).
[0121] Among them, the dimension of the state space is N.
[0122] Step S212, constructing an action space (Action): generating operations of constraint conditions (such as "setting the speed limit to 30 km / h" or "prohibiting vehicles from changing lanes").
[0123] Step S213: Determine a reward function (Reward): The reward function includes positive rewards and negative rewards.
[0124] Among them, positive rewards mean: rewards are given when the constraints meet the system requirements (such as avoiding collisions); negative rewards mean penalties are given when the constraints conflict with the system requirements (such as the speed limit is too low and the vehicle cannot move).
[0125] Step S22: Determine the network structure of the reinforcement learning model network.
[0126] Specifically, the input layer of the reinforcement learning model network is the state vector (dimension is N, for example, containing 10 scene features); the hidden layer is a 3-layer fully connected network, the number of nodes in each fully connected layer is 256, 128 and 64 nodes respectively, and the activation function is ReLU; the output layer is the Q value of the action space (the Q value corresponding to each action, the dimension is the number of actions M).
[0127] Step S23: Determine the training process of the reinforcement learning model network.
[0128] In one embodiment of the method for automatically generating simulation test scenarios for a decision planning system of the present invention, step S23 may specifically include:
[0129] Step S231, Experience Replay: Transfer the historical state The data is stored in the replay buffer and randomly sampled in small batches for training. Here, s represents the state of the agent at a certain moment, a represents the action taken by the agent in state s, r represents the reward obtained by the agent after performing action a, and s' represents the next state to which the agent transfers after performing action a, also called the next state.
[0130] Step S232: Determine the target network: Regularly synchronize the main network parameters to the target network to stabilize training.
[0131] Step S233: Determine the training steps.
[0132] In one embodiment of the method for automatically generating simulation test scenarios for a decision planning system of the present invention, step S233 may specifically include:
[0133] Step S2331: Sample a batch of experience data from the playback buffer.
[0134] Step S2332: Calculate the target Q value using the following formula:
[0135]
[0136] Where r represents the reward obtained after executing the action, γ is the discount factor, which ranges from [0, 1] and is used to measure the importance of future rewards at the current moment. The closer γ is to 1, the more importance is attached to future rewards. Represents the Q value function calculated by the target Q network, which is used to calculate the target Q value. y represents the calculated target Q value (Target Q - value), and Q represents the Q value function calculated by the current Q network.
[0137] Step S2333: Calculate the minimum loss function using the following formula:
[0138] ,
[0139] Where L represents the loss function, Indicates the batch size;
[0140] Step S2334: Update the target network parameters every preset number of steps (for example, 100 steps).
[0141] Step S24: constraint generation, determining the optimal constraint action based on the Q-value maximization strategy, and outputting mathematical constraint conditions according to the optimal constraint action.
[0142] In one embodiment of the method for automatically generating simulation test scenarios for a decision planning system of the present invention, step S24 may specifically include:
[0143] Step S241, online reasoning: input the feature vector of the scene model, and select the action with the largest Q value as the optimal constraint condition.
[0144] Step S242: Constraint expression: Map the action with the largest Q value into a mathematical expression. For example, the speed constraint is: , the safety distance constraint is: .
[0145] The reinforcement learning model (Double DQN) utilizes dual deep Q networks by separating the target network and the main network, effectively addressing the Q-value overestimation problem in traditional DQN and improving the stability of the constraint generation process. Experience replay and target network synchronization mechanisms further optimize training efficiency and are suitable for generating dynamic constraints.
[0146] Step S3, optimization search: Based on the scenario model and the constraints, a Bayesian optimization algorithm is used to generate scenario parameters that meet the constraints.
[0147] Specifically, the parameter space (such as vehicle speed, pedestrian frequency, etc.) and the constraints generated based on the reinforcement learning model are input into the optimization search module, and the Bayesian optimization iterative method (Gaussian process modeling combined with EI acquisition function in this invention) is used to output the optimal scenario parameter combination that meets the constraints. Figure 4 As shown, in one embodiment of the method for automatically generating simulation test scenarios for a decision planning system of the present invention, step S3 may specifically include:
[0148] Step S31: define parameter space.
[0149] In one embodiment of the method for automatically generating simulation test scenarios for a decision planning system of the present invention, step S31 may specifically include:
[0150] Step S311: define search dimensions, where the search dimensions include scene parameters.
[0151] The scene parameters can be, for example, the initial speed of the vehicle , pedestrian appearance frequency , pedestrian density.
[0152] Step S312: defining constraint conditions: the constraint conditions include hard constraints and soft constraints generated by the reinforcement learning model (Double DQN).
[0153] Among them, hard constraints are conditions that must be met, such as speed limit, ; Soft constraints are optimization objectives, such as comfort requirements.
[0154] Step S32: Design the objective function.
[0155] In one embodiment of the method for automatically generating simulation test scenarios for a decision planning system of the present invention, step S32 may specifically include:
[0156] Step S321 , single-objective optimization: maximizing scenario coverage (Coverage) is used as the objective function, where the maximizing scenario coverage is defined as the proportion of the test scenario covering the system state space.
[0157] Step S322, multi-objective optimization: optimizing coverage simultaneously and system response time , the objective function is to maximize the coverage C and minimize the system response time T.
[0158] Among them, the objective function can be expressed as .
[0159] Step S33: Execute a Bayesian optimization step according to the parameter space and the objective function.
[0160] In one embodiment of the method for automatically generating simulation test scenarios for a decision planning system of the present invention, step S33 may specifically include:
[0161] Step S331: Gaussian process modeling.
[0162] In one embodiment of the method for automatically generating simulation test scenarios for a decision planning system of the present invention, step S331 may specifically include:
[0163] Step S3311 , determine the kernel function: select the Matérn 5 / 2 kernel to process nonlinear relationships.
[0164] Step S3312, hyperparameter optimization: optimize kernel parameters through maximum likelihood estimation (MLE).
[0165] In the specific implementation, during the hyperparameter tuning process, parameters such as β and γ are optimized through grid search or AutoML tools.
[0166] Step S332: Select an acquisition function.
[0167] Specifically, the expected improvement (EI) formula is used to balance exploration and utilization.
[0168] The expected improvement formula is as follows:
[0169] ,
[0170] in, Represents the expected improvement function value, which is used to measure the expected improvement that can be brought about by sampling at point x. It represents the expected value function, which is the probability weighted average of the random variable value. represents the independent variable, representing the point considered for sampling in the search space, It represents the function value of the objective function at point x. Represents the currently known optimal solution, that is, the value of the independent variable corresponding to the point where the objective function value is optimal among the sampled points. Indicates the current optimal solution The corresponding objective function value.
[0171] Step S333: iterative optimization.
[0172] In one embodiment of the method for automatically generating simulation test scenarios for a decision planning system of the present invention, step S333 may specifically include:
[0173] Step S3331, initialization: randomly sampling multiple groups (for example, 5 groups) of parameters as initial observation points.
[0174] Step S3332, modeling: fitting the objective function using a Gaussian process.
[0175] Step S3333, select the next parameter point: select the parameter with the maximum benefit through the EI function.
[0176] Step S3334, evaluation: run the scenario in the simulation environment and calculate the objective function value.
[0177] Step S3335, update: add new parameters and results to the observation data set.
[0178] Step S3336: stop the iterative optimization when a cutoff condition is met: the cutoff condition includes that the number of iterations reaches a preset iteration threshold (for example, 50 times) or the objective function converges (for example, the rate of change is <1%).
[0179] In the optimization search process, the objective function is modeled by Gaussian process, and the parameter selection is guided by EI function to gradually approach the optimal solution.
[0180] In the Bayesian optimization process, Gaussian process modeling combined with the expected improvement (EI) function enables efficient search for optimal solutions that satisfy complex constraints in high-dimensional parameter spaces. Its high sample efficiency and support for multi-objective optimization make it particularly suitable for parameter tuning in simulation test scenarios.
[0181] Step S4, scenario evaluation: Use the deep reinforcement learning model (PPO) to evaluate system performance.
[0182] Specifically, the optimized scenario parameters are input into the scenario evaluation module, and a deep reinforcement learning model (specifically, a policy network Actor combined with a value network Critic in the present invention) is used to output a multi-dimensional performance evaluation report (such as safety, efficiency, and robustness). If the scenario evaluation fails (such as safety, efficiency, etc. do not meet the standards), the optimization search module is returned to adjust the parameters. When the optimization search cannot meet the constraints, the constraint generation module is triggered to regenerate the constraints. In one embodiment of the method for automatically generating simulation test scenarios for decision-making planning systems of the present invention, step S4 can specifically include:
[0183] Step S41: construct a simulation environment.
[0184] In one embodiment of the method for automatically generating simulation test scenarios for a decision planning system of the present invention, step S41 may specifically include:
[0185] Step S411: Interface design: connect the decision-making planning system and the simulation platform (such as Apollo+LGSVL) through API.
[0186] Step S412, state input: input scene parameters (such as vehicle position and speed) and environmental parameters (such as weather and road topology).
[0187] Step S413, action output: output the control instructions of the decision planning system (such as steering angle δ, throttle a).
[0188] Step S42: Design a deep reinforcement learning algorithm (PPO).
[0189] In one embodiment of the method for automatically generating simulation test scenarios for a decision planning system of the present invention, step S42 may specifically include:
[0190] Step S421, designing a policy network (Actor): The input of the policy network includes: environment state (such as vehicle sensor data), the output includes: action probability distribution (Gaussian distribution mean and variance).
[0191] Step S422, designing a value network (Critic): The input of the value network includes: Input: Environmental state , the output includes: state value .
[0192] Step S423, design loss function: design the strategy loss function, value loss function and entropy regularization function respectively.
[0193] Specifically, the policy loss function is:
[0194] ,
[0195] in, Represents the policy loss, which is used to measure the quality of the policy output by the policy network; represents the expectation operator, which means to calculate the expected value of the internal expression. The present invention calculates the average value over different state-action pairs. represents the importance sampling ratio, where represents the parameters of the current policy network, represents the probability of the current strategy taking action a in state s, represents the probability that the old strategy takes action a in state s, , used to compare the differences between the new and old strategies, represents the advantage function, which measures the advantage of taking action a in state s compared to the average value. clip represents the clipping function, which is used to limit r:(0) to the interval 1-e, 1+E=0.2 is the set clipping parameter, the purpose of which is to prevent the policy update from being too large.
[0196] The value loss function is:
[0197]
[0198] in, Represents the value loss function, which is used to measure the difference between the state value output by the value network and the actual cumulative reward. represents the value estimate of state s output by the value network, Indicates the cumulative reward, which is the cumulative value of all rewards obtained from the current state.
[0199] The entropy regularization function is:
[0200] , the coefficient is 0.01,
[0201] Represents the entropy regularization term, which is used to increase the exploratory nature of the strategy. The negative sign here is to convert the entropy of the probability distribution (the larger the better, representing a more uniform distribution and stronger exploratory nature) into loss (the smaller the better). The coefficient 0.01 is used to adjust the weight of the entropy regularization term in the total loss.
[0202] Step S43: Execute the evaluation process.
[0203] The present invention verifies the simulation environment by executing an evaluation process to ensure that the generated scenario is physically reasonable in the simulation tool (e.g., the vehicle dynamics model matches the real data). In one embodiment of the method for automatically generating simulation test scenarios for a decision-making planning system of the present invention, step S43 may specifically include:
[0204] Step S431, training agent: running the deep reinforcement learning algorithm (PPO) in the generated scenario to optimize the decision-making strategy.
[0205] Step S432: Determine performance indicators, where the performance indicators include safety, efficiency, and comfort. The safety includes at least one of the number of collisions and the frequency of sudden braking. The efficiency includes at least one of the average vehicle speed and the path deviation distance. The comfort includes the standard deviation of the lateral acceleration.
[0206] Step S44: Generate an evaluation report.
[0207] Specifically, a PDF report is generated, including a comparison curve of metrics compared to the baseline method and a key scenario analysis. If the evaluation fails (e.g., the number of collisions exceeds the limit), the system returns to the optimization search step (step S3) to adjust parameters. If the constraints are still not met after optimization, the system returns to step S2 to regenerate the constraints.
[0208] The deep reinforcement learning model (PPO) adopts a proximal policy optimization algorithm. Through the collaborative training of the policy network and the value network, combined with entropy regularization to enhance exploration capabilities, it can comprehensively evaluate the system's performance in multiple dimensions such as safety, efficiency, and comfort.
[0209] In terms of computing resource allocation, GPU is used to accelerate VAE and PPO training, and distributed computing supports Bayesian optimization.
[0210] This invention can be widely applied in fields such as autonomous driving, intelligent robotics, and industrial automation, providing an efficient solution for reliability verification of complex systems. Furthermore, it is compatible with the APIs of mainstream simulation platforms (such as CARLA, SUMO, and Apollo), allowing generated scenarios to be directly imported into simulation environments for testing.
[0211] The embodiment of the present invention provides an automated method for generating simulation test scenarios for decision-making and planning systems. It uses a deep network (Double DQN) to generate constraints, and jointly generates constraints through VAE and Double DQN. This reduces manual intervention through automated methods and improves scenario generation efficiency. It generates diverse and high-coverage test scenarios, enhances scenario diversity, and comprehensively tests the robustness of the system. It optimizes constraint processing by efficiently handling complex constraints, ensuring that the generated scenarios meet system requirements. It also provides a comprehensive system performance evaluation method, implements multi-dimensional evaluation, and reflects the performance of the system in different scenarios.
[0212] Thus far, various embodiments of the present disclosure have been described in detail. To avoid obscuring the concept of the present disclosure, some details known in the art have not been described. Based on the above description, those skilled in the art can fully understand how to implement the technical solutions disclosed herein.
[0213] Although some specific embodiments of the present disclosure have been described in detail through examples, those skilled in the art will understand that the above examples are for illustration only and are not intended to limit the scope of the present disclosure. Those skilled in the art will understand that the above embodiments may be modified or some technical features may be replaced with equivalents without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. A method for automatically generating simulation test scenarios for decision-making and planning systems, characterized in that: include: Scenario modeling: Use deep learning models to build scenario models; Constraint generation: Based on the output of the scenario model, a reinforcement learning model is used to generate constraint conditions; Optimization search: Based on the scenario model and the constraints, a Bayesian optimization algorithm is used to generate scenario parameters that meet the constraints; Scenario Evaluation: Evaluate system performance using deep reinforcement learning models.
2. The method for automatically generating simulation test scenarios for a decision-making planning system according to claim 1, characterized in that: The scenario modeling: constructing a scenario model using a deep learning model, including: The scenario modeling module inputs historical scenario data, uses a deep learning model to capture long-term dependencies through a self-attention mechanism, connects system requirements with latent variables, and generates scenario parameters, including: Obtaining input data and preprocessing the input data; Model construction and training: The network structure of the model is that the encoder uses a convolutional neural network, the decoder is symmetrical with the encoder, and the input data is used to train the deep learning model; Scene generation: Generate new scene feature vectors using the trained deep learning model; Post-processing: Convert the generated new scene feature vector into a visualization scene.
3. The method for automatically generating simulation test scenarios for a decision-making planning system according to claim 2, characterized in that: The step of obtaining input data and preprocessing the input data includes: Inputting a historical scene data set, wherein the historical scene data set includes environmental parameters, states of dynamic objects, and system requirements; Extracting features based on the historical scene dataset to encode the historical scene dataset into structured features, wherein the structured features include a numerical representation of at least one of a spatial feature, a dynamic feature, and an environmental feature; Data standardization: normalize the input data. The model construction and training: The network structure of the model is that the encoder adopts a convolutional neural network, the decoder is symmetrical with the encoder, and the input data is used to train the deep learning model, including: Construct network structure: the encoder is a convolutional neural network, the input is the scene feature vector, the output is the mean and variance of the latent space, the decoder is symmetrical with the encoder, and the input is the latent variable , the output is the reconstructed scene feature vector; Construct a reconstruction loss function and use the mean square error to measure the difference between the original data and the reconstructed data; The KL divergence constraint constrains the latent space distribution to be close to the standard normal distribution; According to the reconstruction loss function and KL divergence, the total loss function is constructed. The total loss function is expressed by the following formula: , Among them, L represents the total loss function, which is an indicator for comprehensively measuring the performance of the model, and x represents the original data. Represents the data after model reconstruction, β represents the balance coefficient, which is an artificially set hyperparameter used to adjust the proportion of reconstruction loss and KL divergence loss in the total loss, KL represents KL divergence, which is used to measure the difference between two probability distributions, N(μ,σ²): represents the normal distribution with mean μ and variance σ², representing the distribution of the latent space, μ represents the mean of the normal distribution, describing the center position of the distribution, σ² represents the variance of the normal distribution, describing the degree of dispersion of the distribution, N(0, 1) represents the standard normal distribution with mean 0 and variance 1, which serves as the constraint target of the latent space distribution; Training process: Use historical scene data to train the deep learning model, and in the training process, optimize the parameters of the deep learning model through the total loss function. In the training process, the optimizer is Adam optimizer, and the learning rate is , the batch size is 128.
4. The method for automatically generating simulation test scenarios for a decision-making planning system according to claim 2, characterized in that: The scene generation method uses the trained deep learning model to generate a new scene feature vector, including: Latent space sampling: randomly sample latent variables from a standard normal distribution N(0,1) ; Conditional injection: System requirements and environmental parameters are used as conditional inputs, and potential variables are After splicing, input the decoder; Decoding output: The decoder generates a new scene feature vector, The post-processing: converting the generated new scene feature vector into a visual scene, including: Parameter denormalization: restore the generated eigenvectors to actual physical quantities; Scene visualization: The restored actual physical quantities are converted into a visualization scene through simulation tools, wherein the simulation tools include CARLA or SUMO.
5. The method for automatically generating simulation test scenarios for a decision-making planning system according to claim 1, characterized in that: The constraint generation: based on the output results of the scenario model, using the reinforcement learning model to generate constraint conditions, including: The constraint generation module inputs the scene feature vector output by the scene model, uses the reinforcement learning model network to select the optimal constraint action by maximizing the Q value, and outputs mathematical constraint conditions, specifically including: Problem modeling includes: constructing a state space: the state space includes the feature vector of the scene model; constructing an action space: generating operations with constraints; determining a reward function: the reward function includes positive rewards and negative rewards; determining the network structure of the reinforcement learning model network: the input layer of the reinforcement learning model network is the state vector; the hidden layer is a 3-layer fully connected network, with the number of nodes in each fully connected layer being 256, 128, and 64 nodes respectively, and the activation function being ReLU; and the output layer is the Q value of the action space; Determine the training process of the reinforcement learning model network; Constraint generation determines the optimal constraint action based on the Q-value maximization strategy and outputs mathematical constraint conditions based on the optimal constraint action, including: online reasoning: input the feature vector of the scene model and select the action with the largest Q value as the optimal constraint condition; constraint representation: map the action with the largest Q value into a mathematical expression.
6. The method for automatically generating simulation test scenarios for a decision-making planning system according to claim 5, characterized in that: The training process of determining the reinforcement learning model network includes: Experience replay: transferring historical state Stored in the replay buffer, where s represents the state of the agent at a certain moment, a represents the action taken by the agent in state s, r represents the reward obtained by the agent after performing action a, and s' represents the next state the agent transfers to after performing action a; Determine the target network: regularly synchronize the main network parameters to the target network to stabilize training; Determine the training steps, including: sampling a batch of experience data from the playback buffer; calculating the target Q value using the following formula: , Where r represents the reward obtained after executing the action, γ is the discount factor, which ranges from [0, 1] and is used to measure the importance of future rewards at the current moment. The closer γ is to 1, the more importance is attached to future rewards. Represents the Q value function calculated by the target Q network, which is used to calculate the target Q value, y represents the calculated target Q value, and Q represents the Q value function calculated by the current Q network; the minimization loss function is calculated using the following formula: , where L represents the loss function, Indicates the batch size; updates the target network parameters every preset number of steps.
7. The method for automatically generating simulation test scenarios for a decision-making planning system according to claim 1, characterized in that: The optimization search: based on the scenario model and the constraints, using a Bayesian optimization algorithm to generate scenario parameters that meet the constraints, including: The parameter space and the constraints generated based on the reinforcement learning model are input into the optimization search module, and the Bayesian optimization iterative method is used to output the optimal scenario parameter combination that meets the constraints.
8. The method for automatically generating simulation test scenarios for a decision-making planning system according to claim 7, characterized in that: The parameter space and the constraint conditions generated based on the reinforcement learning model are input into the optimization search module, and the optimal scenario parameter combination satisfying the constraint conditions is output using the Bayesian optimization iterative method, including: Defining a parameter space includes: defining a search dimension, wherein the search dimension includes a scene parameter; defining a constraint condition, wherein the constraint condition includes a hard constraint and a soft constraint generated by a reinforcement learning model; Design objective functions, including: single-objective optimization: maximizing scenario coverage as the objective function, where the maximization of scenario coverage is defined as the proportion of the test scenario covering the system state space; multi-objective optimization: optimizing coverage simultaneously and system response time , the objective function is to maximize the coverage C and minimize the system response time T; performing a Bayesian optimization step based on the parameter space and the objective function; The step of performing Bayesian optimization according to the parameter space and the objective function includes: Gaussian process modeling, including: determining the kernel function: selecting the Marten kernel to handle nonlinear relationships; hyperparameter optimization: optimizing kernel parameters through maximum likelihood estimation; Select acquisition functions, including: balancing exploration and exploitation through expected improvement formulas, The expected improvement formula is as follows: , in, Represents the expected improvement function value, which is used to measure the expected improvement that can be brought about by sampling at point x. It represents the expected value function, which is the probability weighted average of the random variable value. represents the independent variable, representing the point considered for sampling in the search space, It represents the function value of the objective function at point x. Represents the currently known optimal solution, that is, the value of the independent variable corresponding to the point where the objective function value is optimal among the sampled points. Indicates the current optimal solution The corresponding objective function value; Iterative optimization includes: initialization: randomly sampling multiple sets of parameters as initial observation points; modeling: fitting the objective function with a Gaussian process; selecting the next parameter point: selecting the parameter with the greatest benefit through the EI function; evaluation: running the scenario in a simulation environment and calculating the objective function value; updating: adding new parameters and results to the observation data set; stopping the iterative optimization when the cutoff condition is met: the cutoff condition includes the number of iterations reaching a preset iteration threshold or the objective function convergence.
9. The method for automatically generating simulation test scenarios for a decision-making planning system according to claim 5, characterized in that: Scenario evaluation: Use deep reinforcement learning models to evaluate system performance, including: The optimized scenario parameters are input into the scenario evaluation module, and a deep reinforcement learning model is used to output a multi-dimensional performance evaluation report. If the scenario evaluation fails, the optimization search module is returned to adjust the parameters. When the optimization search cannot meet the constraints, the constraint generation module is triggered to regenerate the constraints.
10. The method for automatically generating simulation test scenarios for a decision-making planning system according to claim 9, characterized in that: The optimized scenario parameters are input into the scenario evaluation module, and a deep reinforcement learning model is used to output a multi-dimensional performance evaluation report. If the scenario evaluation fails, the optimization search module is returned to adjust the parameters. When the optimization search cannot meet the constraints, the constraint generation module is triggered to regenerate the constraints, specifically including: Constructing the simulation environment, including: Interface design: connecting the decision-making planning system and the simulation platform through API; State input: inputting scene parameters and environment parameters; Action output: outputting control instructions of the decision-making planning system; Design a deep reinforcement learning algorithm, including: Design a policy network: The input of the policy network includes: environment state , the output includes: action probability distribution; design value network: the input of the value network includes: environment state ,The output includes: state value; Design loss functions, including: designing strategy loss function, value loss function and entropy regularization function respectively; Executing an evaluation process includes: training an agent: running a deep reinforcement learning algorithm in a generated scenario to optimize a decision-making strategy; determining performance indicators, wherein the performance indicators include safety, efficiency, and comfort, wherein the safety includes at least one of the number of collisions and the frequency of sudden braking; the efficiency includes at least one of the average vehicle speed and the path deviation distance; and the comfort includes the standard deviation of the lateral acceleration; Generating an evaluation report includes generating a PDF report, wherein the PDF report includes an indicator comparison curve and a key scenario analysis compared with a baseline method.
Citation Information
Cited By
Control algorithm test method based on intelligent driving simulator
CN121764050A
A control algorithm test method based on an intelligent driving simulator
CN121764050B