Background vehicle autonomous confrontation automatic driving function safety verification method
Patent Information
- Application Number
- CN202410609715.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2044-05-16
AI Technical Summary
[0004]本发明的目的在于提供一种融合强化学习及系统观理论的背景车自主对抗自动驾驶功能安全验证方法,以解决上述背景技术中存在的至少一项技术问题
[0052] The beneficial effects of this invention are as follows: The background vehicle is endowed with autonomous adversarial behavior, and adversarial testing is used to verify the functional safety of the autonomous driving system; by utilizing the STAMP system view theory to generate safety constraints, it is possible to start from the overall structure and function of the system, deeply explore the potential safety hazards and risk sources in the system, realize system view risk identification, and help to better identify and manage the safety risks of the system, thereby improving the safety and reliability of the autonomous driving system.
Smart Images

Figure CN118521878B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving technology, specifically to a safety verification method for autonomous driving functions of background vehicles that integrates reinforcement learning and systems theory. Background Technology
[0002] Autonomous driving systems must meet stringent safety requirements to ensure safe and reliable operation under various conditions. Safety verification has become a crucial step in the development of autonomous driving technology. Traditional verification methods, such as simulation testing and test case design, often fail to cover the behavior of autonomous driving systems in all situations. This is particularly true for reinforcement learning algorithms, whose nondeterministic and uncertain learning process makes traditional verification methods difficult to apply effectively. Reinforcement learning algorithms typically exhibit high complexity; the neural network structures, numerous hyperparameters, and interactions with the environment all contribute to the system's uncertainty and complexity. Against this backdrop, functional safety verification of reinforcement learning has become an important research direction in the field of autonomous driving technology. It aims to combine the characteristics of reinforcement learning algorithms with safety requirements to propose new verification methods and techniques to ensure the safety and reliability of autonomous driving systems.
[0003] Current technologies for verifying the safety of autonomous driving functions often neglect the proactive adversarial behavior of other road users (such as other vehicles and pedestrians), typically relying on static datasets or simple simulated environments. This passive verification approach fails to fully expose the potential defects of autonomous driving systems. Adversarial environments may include malicious attacks, sensor errors, weather changes, and other factors, all of which can interfere with autonomous driving systems, leading to performance degradation or even accidents. Furthermore, closed-course testing requires significant investment of human, material, and financial resources, and poses safety risks during the testing process. The simulated and closed-course environments differ significantly from real traffic scenarios, and the reliability of test results needs improvement. Summary of the Invention
[0004] The purpose of this invention is to provide a safety verification method for autonomous driving functions of background vehicles that integrates reinforcement learning and systems theory, so as to solve at least one of the technical problems existing in the background technology.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] In a first aspect, the present invention provides a method for safety verification of autonomous driving functions of a background vehicle, comprising:
[0007] A virtual road environment is constructed, and a background vehicle is designed; the background vehicle learns driving behavior adapted to road conditions and adversarial strategies through a deep reinforcement learning algorithm; wherein,
[0008] The background vehicles are trained using interactive data in a virtual simulation environment. The training process is accelerated by GPUs. An adversarial environment is added for model training, and a verification mechanism is established to verify the training results. The training data is based on real traffic data and adversarial scenarios generated by adversarial generative networks.
[0009] Optionally, adversarial scenarios can be generated using generative adversarial networks, including:
[0010] Initialize the network: First, the generator and discriminator networks are initialized; the generator is used to generate data, and the discriminator is used to distinguish real data from fake data generated by the generator.
[0011] Generate noise input: The generator receives a random noise vector as input;
[0012] Data generation: The generator transforms the noise vector through its network structure to generate a new data point;
[0013] Discriminator classification: The discriminator receives samples from the real dataset and data samples generated by the generator, and then outputs a probability value, which represents the probability that the input sample belongs to the real dataset;
[0014] Discriminator parameter update: Update the discriminator parameters using the backpropagation algorithm based on the loss function;
[0015] Generator parameter update: Update the generator parameters using the backpropagation algorithm.
[0016] Optionally, adversarial strategy training includes: a model network structure, including a perception module, a decision module, and an adversarial training module; the perception module is used to extract features from sensor data, and the decision module is used to make decisions; wherein, an adversarial training mechanism is introduced between the perception module and the decision module, which is achieved by adding an adversarial loss function to the model to improve the robustness of the model in adversarial scenarios.
[0017] Optionally, the STAMP method is used to analyze the structure of the background vehicle autonomous driving system, identify key control variables in the system, identify system risks, and determine safety constraints; the near-end strategy optimization algorithm is combined with the priority experience playback method to determine path planning and obstacle avoidance strategies.
[0018] Optional, longitudinal control safety constraints:
[0019]
[0020]
[0021]
[0022] In the formula, d lon_min v represents the minimum longitudinal safety distance. r v represents the speed of the following vehicle. l Represents the speed of the vehicle in front. This represents the maximum acceleration of the following vehicle. This represents the minimum deceleration of the following vehicle. ρ represents the maximum deceleration of the vehicle in front, and ρ represents the reaction time of the vehicle behind.
[0023] Lateral control safety constraints:
[0024]
[0025]
[0026]
[0027] In the formula, d lat_min This represents the minimum lateral safety distance, μ is a predetermined safety distance threshold, v1 and v2 represent the lateral velocities of the two vehicles, and β represents the minimum lateral safety distance. 1,lat,min, β 2,lat,min ρ represents the minimum lateral deceleration of the two vehicles, and ρ represents the reaction time.
[0028] Optional, the collision-free ratio safety constraint represents the percentage of test sets in which autonomous vehicles do not experience collisions; the collision-free ratio measures the safety of autonomous vehicles on the road.
[0029]
[0030] P coll N represents the collision-free ratio. coll This represents the number of episodes where a collision occurred, and N represents the total number of episodes evaluated.
[0031] The lane-sharing ratio safety constraint quantifies the time required for an autonomous vehicle to yield to an emergency vehicle while the emergency vehicle is following it; that is, the time from when the emergency vehicle begins to follow the autonomous vehicle until the autonomous vehicle successfully yields to the emergency vehicle.
[0032]
[0033] P ts N represents the collision-free ratio. ts This represents the number of episodes where a collision occurred, and N represents the total number of episodes evaluated.
[0034] Optionally, during the statistical evaluation of unobstructed proportional safety constraints, the duration for which the autonomous vehicle, in order to improve efficiency, changes lanes in front of the emergency vehicle and obstructs its movement when the emergency vehicle is following on both sides of the autonomous vehicle is:
[0035]
[0036] P bf N represents the collision-free ratio. bf This represents the number of episodes where a collision occurred, and N represents the total number of episodes evaluated.
[0037] Collision penalty constraints:
[0038]
[0039] In the formula, R near lo represents the penalty value for near collisions. sl This indicates the longitudinal distance between the autonomous vehicle and the vehicle in front. It is the minimum longitudinal safety distance.
[0040] In a second aspect, the present invention provides a non-transitory computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the background vehicle autonomous driving function safety verification method as described in the first aspect, which integrates reinforcement learning and systems theory.
[0041] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the processor and the memory communicate with each other, the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the safety verification method for autonomous driving function of background vehicle integrating reinforcement learning and system view theory as described in the first aspect.
[0042] Fourthly, the present invention provides an electronic device, comprising: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to cause the electronic device to execute instructions for implementing the background vehicle autonomous driving function safety verification method as described in the first aspect, which integrates reinforcement learning and systems theory.
[0043] Safety verification: In the field of autonomous driving, safety verification refers to the comprehensive testing and evaluation of autonomous driving systems to ensure their safety and reliability in various traffic scenarios and complex environments. The purpose of safety verification is to reduce the potential risks of autonomous driving systems when operating on real roads, ensuring that the system can make correct decisions and behaviors in response to various situations and challenges in driving tasks.
[0044] Robustness refers to the ability of a system or algorithm to resist changes in input, abnormal conditions, or attacks. A robust system can maintain good performance and reliability when faced with different, even unreasonable, inputs, without crashing or exhibiting unpredictable behavior due to changes in input.
[0045] GPU: A GPU is a processor specifically designed for processing graphics and image-related tasks. It accelerates computation by processing large amounts of data in parallel, offering higher efficiency in handling parallel computing tasks compared to traditional central processing units (CPUs). Due to its parallel computing capabilities, GPUs are widely used in machine learning, deep learning, and other fields, accelerating the training and inference of neural network models and improving computational performance and efficiency.
[0046] Ray is an open-source distributed execution framework designed to provide a simple and efficient way to build and run large-scale distributed applications. Its core features include distributed task execution, task scheduling and concurrency control, distributed state management, and a Python interface. Ray is designed to simplify the development and management of distributed applications and provide high performance and scalability in large-scale environments.
[0047] TensorFlow Distributed is a component of TensorFlow used for training and inference in a distributed environment. It offers various distributed strategies, allowing users to choose the appropriate distributed training method based on different needs, such as data parallelism or model parallelism. TensorFlow Distributed also provides cluster management and communication functions, making data exchange between nodes in a distributed computing environment more efficient. Furthermore, it supports training with distributed datasets and provides a distributed optimizer for optimizing model parameters in a distributed environment. Through TensorFlow Distributed, users can leverage distributed computing resources to accelerate model training and inference, and handle larger-scale datasets and models.
[0048] GAN is a machine learning model consisting of two neural networks: a generator and a discriminator. These two networks compete with each other during training to improve the generator's performance. The generator aims to generate new data samples that are similar to real data, while the discriminator aims to distinguish the generated data from real data. During training, the generator strives to generate more realistic data, while the discriminator strives to improve its accuracy, leading to continuous competition and performance improvements between the two networks.
[0049] Adversarial training is a technique used in machine learning to improve a model's robustness to input data and its ability to defend against adversarial attacks. Introducing adversarial training into the field of autonomous driving can help vehicle models better cope with various anomalies and potential attacks.
[0050] Adversarial reinforcement learning is a method that combines reinforcement learning and adversarial learning. In this method, an agent interacts with an opponent (the environment), optimizing its own policy to maximize rewards, while the opponent tries to minimize the agent's rewards. This method is commonly used to solve multi-agent competition, game theory problems, and in research in machine learning and artificial intelligence.
[0051] STAMP is a methodology for analyzing and preventing accidents in complex systems. This theory emphasizes a systems-wide perspective, arguing that accidents are often caused by defects in the system's control structure or processes. The STAMP systems theory provides a series of methods and tools, such as STPA (System-Theoretic Process Analysis), CAST (Causal Analysis based on STAMP), STAMP-based HazOp (Hazard and Operability Analysis), and STPA-Sec (STPA for Security), to identify potential hazards in systems and propose preventative measures. The core idea of this theory is to view accidents as system functional failures, emphasizing a holistic understanding and analysis of system safety.
[0052] The beneficial effects of this invention are as follows: The background vehicle is endowed with autonomous adversarial behavior, and adversarial testing is used to verify the functional safety of the autonomous driving system; by utilizing the STAMP system view theory to generate safety constraints, it is possible to start from the overall structure and function of the system, deeply explore the potential safety hazards and risk sources in the system, realize system view risk identification, and help to better identify and manage the safety risks of the system, thereby improving the safety and reliability of the autonomous driving system.
[0053] The advantages of additional aspects of the invention will be set forth more clearly in the following description or will be learned by practice of the invention. Attached Figure Description
[0054] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0055] Figure 1 This is a flowchart illustrating the safety verification method for autonomous driving functions of a background vehicle that integrates reinforcement learning and systems theory, as described in an embodiment of the present invention. Detailed Implementation
[0056] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0057] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0058] It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as here.
[0059] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or groups thereof.
[0060] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.
[0061] To facilitate understanding of the present invention, the present invention will be further explained and described below with reference to the accompanying drawings and specific embodiments, and the specific embodiments do not constitute a limitation on the embodiments of the present invention.
[0062] Those skilled in the art should understand that the accompanying drawings are merely schematic diagrams of embodiments, and the components in the drawings are not necessarily essential for implementing the present invention.
[0063] With the rapid development of artificial intelligence and deep learning technologies, autonomous vehicles are gradually moving from the laboratory to commercial applications. However, facing complex road environments, unpredictable pedestrian behavior, and various potential hardware failures, autonomous vehicles face enormous safety challenges. Existing technologies for functional safety verification of autonomous driving often neglect the proactive adversarial behavior of other road users (such as other vehicles and pedestrians), typically relying on static datasets or simple simulated environments. This passive verification approach struggles to fully expose the potential defects of autonomous driving systems. Adversarial environments may include malicious attacks, sensor errors, weather changes, and other factors, all of which can interfere with the autonomous driving system, leading to performance degradation or even accidents. Furthermore, closed-site testing requires significant investment of manpower, resources, and capital, and poses safety risks during the testing process. Simulated and closed-site environments differ significantly from real traffic scenarios, and the reliability of test results needs improvement. This invention utilizes reinforcement learning technology to train background vehicles, enabling them to autonomously learn and simulate various driving behaviors and adversarial strategies in real road environments, thereby improving the realism and effectiveness of the verification method. By introducing the autonomous adversarial behavior of a background vehicle, a training and verification mechanism in an adversarial environment is introduced, enabling the autonomous driving system to autonomously perform reinforcement learning in a simulated adversarial environment to improve its adversarial capabilities. Finally, adversarial testing is used to verify the functional safety of the autonomous driving system.
[0064] Example 1
[0065] like Figure 1 As shown in this embodiment, an innovative method for verifying the safety of autonomous driving functions based on reinforcement learning and autonomous adversarial behavior of the background vehicle is proposed. This method introduces reinforcement learning and autonomous adversarial behavior of the background vehicle, enabling the autonomous driving system to exhibit stronger robustness and safety when facing complex and ever-changing road environments.
[0066] In this embodiment, environmental modeling is first performed, including: constructing a highly realistic virtual road environment using computer graphics technology, comprehensively utilizing various computer graphics techniques such as terrain modeling, building design, realistic rendering, traffic behavior simulation, and virtual sensor simulation, for example, using game engines such as Unity or Unreal Engine. This is combined with user interface design and interaction design to achieve a virtual environment that is both realistic and interactive.
[0067] The background vehicle design includes: introducing deep reinforcement learning techniques, such as using Deep Q-Networks (DQN), Deep Deterministic Policy Gradient (DDPG), or other neural network-based reinforcement learning methods, to allow the vehicle to learn driving strategies through interaction with the environment. The background vehicle is designed as an intelligent agent that can make decisions and execute corresponding actions based on the state of the environment. This agent will learn driving behaviors adapted to various road conditions and adversarial strategies through deep reinforcement learning algorithms.
[0068] Algorithms such as Deep Q-Network (DQN), Dual Deep Q-Network (DDQN), and Deep Deterministic Policy Gradient (DDPG) are used to train background vehicles using a large amount of interactive data in a virtual simulation environment. During the training process, GPUs are used to accelerate the process, or distributed training frameworks (such as Ray, TensorFlow Distributed, etc.) are used for parallel training to improve efficiency.
[0069] In this embodiment, an adversarial environment is incorporated for model training, and a verification mechanism is established to validate the training results. Specifically, real traffic data is collected through aerial video and radar, and adversarial scenarios generated using GANs are used as training data. The training data undergoes cleaning and preprocessing, including noise removal, data labeling, and data augmentation.
[0070] In this embodiment, adversarial scenarios are generated using Generative Adversarial Networks (GANs), or adversarial behaviors are injected using pre-recorded real traffic data.
[0071] The process of GAN generating adversarial scenarios includes:
[0072] Network initialization: First, the generator and discriminator networks are initialized. The generator is responsible for generating data, while the discriminator is responsible for distinguishing real data from fake data generated by the generator.
[0073] Generate noise input: The generator receives a random noise vector as input, which is usually derived from a Gaussian distribution or other probability distribution.
[0074] Data generation: The generator transforms the noise vector through its network structure to generate a new data point, which should ideally have the same distribution as the real data.
[0075] Discriminator classification: The discriminator receives samples from the real dataset and data samples generated by the generator, and then outputs a probability value representing the probability that the input sample belongs to the real dataset.
[0076] Calculate the discriminator loss: The discriminator loss function is usually the binary cross-entropy loss, which is calculated based on the discriminator's classification accuracy of real and generated data.
[0077] Discriminator parameter update: Based on the loss function, the parameters of the discriminator are updated using the backpropagation algorithm to improve its ability to distinguish between real and fake data.
[0078] Calculate generator loss: The generator loss function usually adopts binary cross-entropy loss, which aims to minimize the probability that the discriminator can correctly classify the generated data, that is, to hope that the discriminator will misclassify the generated data as real data.
[0079] Generator parameter update: The generator parameters are also updated using the backpropagation algorithm, with the aim of generating more realistic data to deceive the discriminator.
[0080] Iterative training: until the quality of the data generated by the generator is high enough that the discriminator can hardly distinguish between real and fake data.
[0081] In this embodiment, an adversarial training strategy is used for training, including: introducing adversarial training methods, such as adversarial training and adversarial reinforcement learning, so that the autonomous driving system can cope with various adversarial situations through autonomous reinforcement. The training model and training process are as follows: Design a suitable network structure, including a perception module, a decision module, and an adversarial training module. The perception module uses structures such as convolutional neural networks (CNNs) to extract features from sensor data, and the decision module uses recurrent neural networks (RNNs) or deep reinforcement learning models to make decisions. An adversarial training mechanism is introduced between the perception and decision modules, which can be achieved by adding an adversarial loss function to the model to make the model more robust in adversarial scenarios. Randomly initialize the model parameters and start the training process settings. Define the loss function, including the conventional classification loss, the reward function in reinforcement learning, and the adversarial loss function. Select a suitable optimization algorithm (Adam optimizer) to minimize the loss function and update the model parameters. Introduce adversarial examples into the training data and train them together with real samples, using GANs to enhance the model's adaptability to adversarial scenarios.
[0082] In this embodiment, simulation testing or real-world scenario testing is used to perform functional safety verification of the autonomous driving system's performance in adversarial environments, such as using simulation platforms like AirSim. Large-scale simulation testing is conducted using simulation platforms (such as AirSim), or testing is performed in real-world road environments to monitor the autonomous driving system's response to real-world adversarial behavior, thereby verifying its reliability and safety in real-world scenarios.
[0083] In this embodiment, the generation of safety constraints for autonomous driving scenarios includes:
[0084] First, the STAMP method is used to analyze the structure of the autonomous driving system. This includes identifying the various components of the system, such as sensors, controllers, and actuators, and the interactions between them. The system's functions are analyzed, including data acquisition, perception, decision-making, and execution. This helps in understanding the system's behavior patterns and response mechanisms under different conditions. Key control variables in the system are identified, as these directly affect system safety. Examples include vehicle speed, direction, and acceleration. The dynamic decision-making processes involved in the system, such as path planning and obstacle avoidance strategies, are analyzed. This helps determine the system's response strategies under different conditions. Through system-wide risk identification, structural defects, interaction problems between functional modules, and abnormal situations involving control variables are identified.
[0085] In this embodiment, the Proximal Policy Optimization (PPO) algorithm is combined with the Prioritized Experience Replay (PER) method, and a decision-making method is proposed based on this combination, which can be applied to path planning, obstacle avoidance strategies, etc. The decision-making process is implemented through the following steps:
[0086] Data acquisition: Real traffic data is collected through aerial video and radar, and adversarial scenarios generated by GAN based on this are used as training data.
[0087] Data processing: The controller processes the acquired data, which involves technologies such as image recognition, object detection, and road condition analysis. By processing the data, the controller can identify obstacles, lane markings, and other important features on the road.
[0088] Decision Outcomes: Ultimately, based on data analysis and the system's built-in algorithms, the controller makes decisions such as determining the optimal driving path, adjusting vehicle speed, and avoiding collisions with other vehicles or obstacles. These decisions are transmitted to the actuators, which translate them into the vehicle's actual behavior.
[0089] The PPO algorithm belongs to the policy gradient method, which directly optimizes the model's policy (i.e., the model's behavior) in an attempt to find a policy that maximizes the expected reward. Its core principle is the introduction of a novel "truncated alternative objective function." Unlike standard policy gradient methods that update only once per data sample, this function allows for multiple optimization updates using mini-batch data.
[0090] The PPO algorithm steps include:
[0091] (1) Data collection
[0092] Collect a set of interaction data by executing the current policy in the environment. The data includes state, action, reward, and possible next state.
[0093] (2) Calculate the advantage estimate
[0094] The advantage function is calculated using time difference estimation or generalized advantage estimation to evaluate the quality of an action relative to the average level.
[0095] (3) Optimize the objective function
[0096] The objective function is usually in the form of:
[0097] L(θ)=E(min(r t (θ)A, clip(r) t (θ), 1-∈, 1+∈), A)
[0098] In the formula, A represents the estimate of the advantage function, ∈ represents a small positive number (e.g., 0.1), θ represents the policy parameter, and r t (θ) represents the probability ratio, and the clip function restricts the probability ratio r. t The range of variation of (θ) is adjusted to prevent the update steps from being too large.
[0099] For a specific state s and a specific action a, the probability ratio r t (θ) is defined as:
[0100]
[0101] In the formula, π θ Indicates the new strategy, π θold This indicates the old strategy.
[0102] (4) Update Strategy
[0103] Update the policy parameters θ using the gradient ascent method:
[0104]
[0105] In the formula, α represents the learning rate.
[0106] (5) Repeat the steps
[0107] Repeat the above steps using the new strategy parameters, and stop when the stopping criteria are met, such as when the strategy performance no longer improves or when a certain number of iterations has been reached.
[0108] The performance of the PPO algorithm is affected by several hyperparameters, including:
[0109] ①T (Horizon): The time step of each actuator.
[0110] ②K(Num.epochs): The number of optimization cycles to perform on the collected data.
[0111] ③M (Minibatch size): The amount of data used in each stochastic gradient ascent.
[0112] ④ε (Clipping parameter): The interval boundary of the cutoff probability ratio
[0113] ⑤β (KL Penalty Coefficient): KL divergence penalty coefficient
[0114] The PPO algorithm limits the policy update step size, effectively avoiding large fluctuations during training and ensuring the stability of the learning process. The PPO algorithm simplifies the computation process, reduces algorithm complexity and resource requirements, and improves computational efficiency.
[0115] Prioritized Experience Replay (PER) is based on the principle of dynamically adjusting the sampling probability of experiences during training according to their importance. Compared with traditional experience replay, this method can make more effective use of the data in the experience replay pool, thereby accelerating the learning process.
[0116] In reinforcement learning, experience typically consists of a series of tuples (state, action, reward, next state). The agent accumulates these experiences through interaction with the environment. Not all experiences are equally important to the learning process. Generally, those experiences that lead to larger prediction errors (i.e., large TD errors) are more critical for policy improvement.
[0117] Temporal difference (TD) error is the difference between the current Q-value prediction and the TD objective (the actual reward plus the Q-value of the next state). A large TD error indicates that the agent's current policy is not performing well in these states and requires more attention and correction.
[0118] TD error is a method for measuring the difference between the current strategy's prediction and the actual result. Specifically, TD error can be expressed as:
[0119] δ=R+γ·V(s′)-V(s)
[0120] Or for the TD error of the action value function:
[0121] δ=R+γ·Q(s′,a′)-Q(s,a)
[0122] Where R is the reward obtained, γ is the discount factor, V is the state value function, and Q is the action value function.
[0123] PER (Percentage Error Response) reflects the importance of each experience by assigning a priority to it. Priority is typically proportional to the absolute value of the TD (True / False) error, and this relationship can be amplified using a power-law function. This means that experiences with larger TD errors are given higher priority. Larger TD errors indicate a significant deviation between the current policy's predictions and the actual rewards, suggesting that the experience may be more important for policy improvement.
[0124] (1) Priority calculation and update:
[0125] Priority is typically calculated using a power-law function, as shown in the following formula:
[0126] p i =(∈+|δ i |) α
[0127] Where: p i Indicates the priority of experience i; δ i denoted by TD error of empirical i; ∈ is a small constant used to ensure that division by zero does not occur; α is a parameter used to control the power law of priority distribution.
[0128] After experience is used to train and update the agent's policy, the priority of that experience needs to be recalculated and updated to reflect its importance under the new policy.
[0129] As the agent's policy is updated, the TD errors of some experiences may change, therefore their priorities also need to be updated accordingly. Priority updates ensure that the experiences in the experience replay pool reflect the latest state of the current policy. The update process is roughly as follows:
[0130] Calculate new TD error: After the agent takes an action from a certain state and observes the result, the new experience TD error can be calculated;
[0131] Update priority: based on the new TD error and power-law function p i =(∈+|δ i |) α Prioritize updating this experience;
[0132] Update in the data structure: Find the corresponding node in the summation tree or Fibonacci heap and update its stored priority value.
[0133] (2) Sampling of experience:
[0134] In priority-based experience replay, experience sampling is no longer uniformly random, but rather based on the priority of each experience. This method ensures that experiences with higher priority (i.e., larger TD errors) are selected with a higher probability, thus participating in training more frequently.
[0135] When sampling experiences, the probability of each experience being selected is proportional to its priority. Specifically, the probability P of experience i being selected is... i The following formula is given:
[0136]
[0137] Here, the denominator is the sum of priority powers of all experiences, used to normalize the probability. The sampling process specifically consists of the following steps:
[0138] Calculating the probability: First, the probability of being sampled based on the priority of each experience can be calculated using the formula above. Since the priority is power-law distributed, the sampling probability P... i and It is directly proportional, where α is an adjustable parameter.
[0139] Random selection: Next, through A value is randomly selected from the range to determine which experience to sample. This random value will fall within the cumulative probability range of a certain experience.
[0140] Selecting an experience: Based on a randomly selected value, find the corresponding experience in a summation tree or Fibonacci heap. This experience will be selected as the current sample.
[0141] (3) Data Structure:
[0142] To achieve efficient priority allocation and updates, PER uses special data structures, such as summation trees or Fibonacci heaps.
[0143] Summation tree:
[0144] Structure: A summation tree is a binary tree where each node stores a value, typically the sum of the values of its child nodes. Adding Experience: New experiences are added to the tree based on their priority, possibly by inserting a new node or updating an existing one. Sampling Experience: By... The corresponding leaf node is randomly selected and searched within the range to sample experience.
[0145] Update priority: When the priority of an experience is updated, the value of the corresponding node will also be updated, and if the node is a child node of a parent node, the value of the parent node also needs to be updated.
[0146] Fibonacci piles:
[0147] Structure: A Fibonacci heap is a heap structure containing multiple trees, each of which is a Fibonacci tree, a special type of binary tree.
[0148] Advantages: Fibonacci heaps are highly efficient in inserting, finding, and deleting the minimum element, making them ideal for implementing PER.
[0149] Operations: In PER, Fibonacci heaps are used to store the priority of experiences and can quickly insert new experiences, find the experience with the highest priority, and update the priority of experiences.
[0150] By using these efficient data structures, PER can quickly process large amounts of empirical data while maintaining dynamic updates to priorities, which is crucial for the performance of reinforcement learning algorithms.
[0151] (2) Safety Constraint Generation: In this embodiment, logical reasoning is used to derive the system's safety constraints based on the analysis results in the STAMP theory. This can generate safety constraints based on known rules, restrictions, and conditions. Formal methods, such as mathematical modeling and model checking, are used to formally describe the system and extract safety constraints from it. This ensures the accuracy and rigor of the safety constraints. The system's safety requirements and specifications are considered and transformed into specific safety constraints. These may include constraints related to following traffic rules, avoiding collisions, and maintaining vehicle stability.
[0152]
[0153] Longitudinal control safety constraints:
[0154]
[0155]
[0156]
[0157] In the formula, d lon_min v represents the minimum longitudinal safety distance. r v represents the speed of the following vehicle. l Represents the speed of the vehicle in front. This represents the maximum acceleration of the following vehicle. This represents the minimum deceleration of the following vehicle. ρ represents the maximum deceleration of the vehicle in front, and ρ represents the reaction time of the vehicle behind.
[0158] Lateral control safety constraints:
[0159]
[0160]
[0161]
[0162] In the formula, d lat_min This represents the minimum lateral safety distance, μ is a predetermined safety distance threshold, v1 and v2 represent the lateral velocities of the two vehicles, and β represents the minimum lateral safety distance. 1,lat,min, β 2,lat,min ρ represents the minimum lateral deceleration of the two vehicles, and ρ represents the reaction time.
[0163] Collision-free proportional safety constraints:
[0164] This metric represents the percentage of test sets in which autonomous vehicles do not experience collisions during a series of evaluation tests. The collision-free ratio is a good measure of the safety of autonomous vehicles on the road, and it is directly related to the safety and reliability of the driving process.
[0165]
[0166] In the formula, P coll N represents the collision-free ratio. coll This represents the number of episodes where a collision occurred, and N represents the total number of episodes evaluated.
[0167] Lane sharing ratio safety constraints:
[0168] This metric quantifies the time required for an autonomous vehicle to yield to an emergency vehicle while it is following the emergency vehicle; that is, the time from when the emergency vehicle begins to follow the autonomous vehicle until the autonomous vehicle successfully yields to the emergency vehicle.
[0169]
[0170] In the formula, P ts N represents the collision-free ratio. ts This represents the number of episodes where a collision occurred, and N represents the total number of episodes evaluated.
[0171] Unobstructed proportional safety constraints:
[0172] This metric measures the length of time, during the evaluation process, when an emergency vehicle is following on either side of an autonomous vehicle, the autonomous vehicle changes lanes in front of the emergency vehicle to improve efficiency and obstructs its movement.
[0173]
[0174] In the formula, P bf N represents the collision-free ratio. bf This represents the number of episodes where a collision occurred, and N represents the total number of episodes evaluated.
[0175] Collision penalty constraints:
[0176]
[0177] In the formula, R near lo represents the penalty value for near collisions. sl This indicates the longitudinal distance between the autonomous vehicle and the vehicle in front. It is the minimum longitudinal safety distance defined in formula (3).
[0178] In this embodiment, the generated safety constraints are embedded into the design and implementation of the autonomous driving system, for example, through embedded system development or software-defined networking. The execution mechanism of the safety constraints is designed and implemented, such as using hardware protection mechanisms and software monitoring systems to ensure the effective execution of the safety constraints. Real-time monitoring: Sensor data and a monitoring system are used to monitor the execution of the safety constraints in real time, promptly detecting and handling abnormal situations.
[0179] Data-driven optimization: Continuously collect and analyze real-world scenario data, and use data-driven methods to optimize and improve the autonomous driving system. Algorithm iterative optimization: Based on experimental results and system feedback, iteratively optimize the reinforcement learning algorithm and safety constraint generation algorithm to improve the system's performance and robustness.
[0180] In summary, in this embodiment, the background vehicle is endowed with autonomous adversarial behavior. An adversarial scenario is generated using a Generative Adversarial Network (GAN), or pre-recorded real traffic data is injected to introduce adversarial behavior. Through this method, the autonomous driving system can autonomously enhance its adversarial capabilities, thereby improving its adversarial performance. Furthermore, adversarial testing verifies the functional safety of the autonomous driving system. The system-view theory STAMP is used to generate safety constraints. STAMP provides a systematic safety analysis method that can delve into the potential safety hazards and risk sources within a system, starting from its overall structure and function. By comprehensively considering the system's physical structure, function, control variables, dynamic decision-making, and other characteristics, STAMP achieves system-view risk identification. The specific methods for training a risk identification model include: collecting a large amount of real traffic scenario data and injecting adversarial behaviors or using generative adversarial networks to generate adversarial scenarios; preprocessing the data, including cleaning, denoising, labeling, and feature extraction; selecting an appropriate model structure and designing a suitable architecture for the risk identification task; inputting the prepared data into the model for training, and adjusting the model parameters through backpropagation algorithms and optimizers to minimize the difference between the predicted and actual results; evaluating the model performance using validation sets or cross-validation methods, and adjusting and improving the model based on the evaluation results; finally, applying the trained risk identification model to an autonomous driving system for real-time monitoring and identification of potential safety risks. This approach helps to better identify and manage system safety risks, thereby improving the safety and reliability of autonomous driving systems.
[0181] Example 2
[0182] This embodiment 2 provides a non-transitory computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, they implement the background vehicle autonomous driving function safety verification method described in embodiment 1 above.
[0183] Example 3
[0184] This embodiment 3 provides a computer device, including a memory and a processor. The processor and the memory communicate with each other. The memory stores program instructions that can be executed by the processor. The processor calls the program instructions to execute the background vehicle autonomous driving function safety verification method as described in embodiment 1 above.
[0185] Example 4
[0186] This embodiment 4 provides an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes instructions to implement the background vehicle autonomous driving function safety verification method as described in embodiment 1 above.
[0187] In summary, the background vehicle autonomous adversarial driving function safety verification method integrating reinforcement learning and systems theory described in this invention utilizes reinforcement learning technology to train the background vehicle, enabling it to autonomously learn and simulate various driving behaviors and adversarial strategies in real road environments, thereby improving the realism and effectiveness of the verification method. Introducing the background vehicle's autonomous adversarial behavior allows the autonomous driving system to exhibit stronger robustness and safety in complex and changing road environments. The STAMP system theory is used to generate safety constraints to ensure the system can still drive safely in adversarial environments. By introducing a training and verification mechanism in an adversarial environment, the autonomous driving system can be trained in a simulated adversarial environment and its functional safety can be verified through adversarial testing. This invention can provide practical suggestions for traffic management and planning, the development and optimization of intelligent transportation systems, and can serve as a foundation for the construction of traffic simulation and testing platforms, as well as providing a theoretical basis for adaptive system design and intelligent control system design.
[0188] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0189] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0190] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0191] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment, whereby a series of operational steps are performed to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0192] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that, based on the technical solutions disclosed in the present invention, various modifications or variations that can be made by those skilled in the art without creative effort should be included within the scope of protection of the present invention.
Claims
1. A method for safety verification of autonomous driving function of a background vehicle, characterized in that, include: Construct a virtual road environment and design background vehicles; The background vehicle learns driving behavior adapted to road conditions and adversarial strategies through a deep reinforcement learning algorithm; wherein, The model is trained on background vehicles using interactive data in a virtual simulation environment. GPU acceleration is employed during training. An adversarial environment is introduced for training, and a validation mechanism is established to verify the training results. The training data includes real traffic data and adversarial scenarios generated through a generative adversarial network (GAN). The GAN-generated adversarial scenarios include: Initialize the network: First, the generator and discriminator networks are initialized; the generator is used to generate data, and the discriminator is used to distinguish real data from fake data generated by the generator. Generate noise input: The generator receives a random noise vector as input; Data generation: The generator transforms the noise vector through its network structure to generate a new data point; Discriminator classification: The discriminator receives samples from the real dataset and data samples generated by the generator, and then outputs a probability value, which represents the probability that the input sample belongs to the real dataset; Discriminator parameter update: Update the discriminator parameters using the backpropagation algorithm based on the loss function; Generator parameter update: Update the generator parameters using the backpropagation algorithm; Adversarial strategy training includes: model network structure, including a perception module, a decision module, and an adversarial training module; the perception module is used to extract features from sensor data, and the decision module is used for decision making; an adversarial training mechanism is introduced between the perception module and the decision module, which is achieved by adding an adversarial loss function to the model to improve the robustness of the model in adversarial scenarios; The STAMP method is used to analyze the structure of the background vehicle autonomous driving system, identify key control variables, identify system risks, and determine safety constraints. Near-end policy optimization algorithms are combined with a priority experience playback method to determine path planning and obstacle avoidance strategies. Among these, time difference... Priority calculation uses a power-law function. After experience is used to train and update the agent's policy, its priority needs to be recalculated and updated to reflect its importance under the new policy. Priority update ensures that the experience in the experience replay pool reflects the latest state of the current policy. The update process is as follows: calculate the new time difference, update the priority of the experience based on the new time difference and the power-law function; find the corresponding node in the summation tree or Fibonacci heap data structure and update its stored priority value. When sampling experience, the probability of each experience being selected is proportional to its priority. The sampling process includes: Calculate the probability: Calculate the probability of each experience being sampled based on its priority; Random selection: Next, through [0, To determine which experience to sample, a value is randomly selected from the range. Selecting an experience: Based on the randomly selected value, find the corresponding experience in the summation tree or Fibonacci heap, and this experience will be selected as the current sample.
2. The background vehicle autonomous driving function safety verification method according to claim 1, characterized in that, Longitudinal control safety constraints: In the formula, Indicates the minimum longitudinal safety distance. Represents the speed of the following vehicle. Represents the speed of the vehicle in front. This represents the maximum acceleration of the following vehicle. This represents the minimum deceleration of the following vehicle. This represents the maximum deceleration of the vehicle in front. This represents the reaction time of the vehicle behind. Lateral control safety constraints: In the formula, Indicates the minimum lateral safety distance. It is a predetermined safe distance threshold. These represent the lateral speeds of the two vehicles, This represents the minimum lateral deceleration of the two vehicles. This represents the reaction time.
3. The background vehicle autonomous driving function safety verification method according to claim 1, characterized in that, The collision-free ratio safety constraint represents the percentage of test sets in which autonomous vehicles do not experience collisions; the collision-free ratio measures the safety of autonomous vehicles driving on the road. ; Indicates the collision-free ratio. Indicates the episode number where the collision occurred. Indicates the total number of evaluation episodes; The lane-sharing ratio safety constraint quantifies the time required for an autonomous vehicle to yield to an emergency vehicle while the emergency vehicle is following it; that is, the time from when the emergency vehicle begins to follow the autonomous vehicle until the autonomous vehicle successfully yields to the emergency vehicle. ; Indicates the collision-free ratio. Indicates the episode number where the collision occurred. This indicates the total number of evaluation episodes.
4. The background vehicle autonomous anti-autonomous driving function safety verification method according to claim 1, characterized in that, During the statistical evaluation of unobstructed proportional safety constraints, when an emergency vehicle is following on either side of an autonomous vehicle, the length of time the autonomous vehicle changes lanes in front of the emergency vehicle and obstructs its movement in order to improve efficiency is: ; Indicates the collision-free ratio. Indicates the episode number where the collision occurred. Indicates the total number of evaluation episodes; Collision penalty constraints: In the formula, This represents the penalty value for near collisions. This indicates the longitudinal distance between the autonomous vehicle and the vehicle in front. It is the minimum longitudinal safety distance.
5. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, which, when executed by a processor, implement the background vehicle autonomous driving function safety verification method as described in any one of claims 1-4.
6. A computer device, characterized in that, The system includes a memory and a processor, which communicate with each other. The memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the background vehicle autonomous driving function safety verification method as described in any one of claims 1-4.
7. An electronic device, characterized in that, include: The device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to cause the electronic device to execute instructions that implement the background vehicle autonomous driving function safety verification method as described in any one of claims 1-4.
Citation Information
Patent Citations
Automatic driving test scene generation method and system, and storage medium
CN116894395A
Unmanned confrontation test scene generation method and system
CN117933096A