Automatic driving decision control method and device and vehicle
By generating vehicle status information and verifying safety with responsibility-sensitive safety models, the problem that existing autonomous driving decision-making control methods may output dangerous decision results is solved, and vehicle safety is guaranteed and the reliability of the autonomous driving system is improved.
Patent Information
- Application Number
- CN202510158531.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-10
AI Technical Summary
The existing autonomous driving decision-making control methods have the probability of outputting unexplained dangerous decision-making results, resulting in wrong behavior of autonomous driving vehicles and safety problems.
By generating vehicle status information, the current action value is output through the autonomous driving decision model based on this information, and the vehicle safety is verified in combination with the responsibility sensitive safety model to ensure that the action value is only issued to the vehicle control command receiver within the safe range.
It effectively avoids the risky operation of vehicle interaction during model training and deployment, ensures the safety of the vehicle, and improves the reliability of the autonomous driving system.
Smart Images

Figure CN120122641A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of vehicle autonomous driving, and particularly to an autonomous driving decision control method, device, and vehicle. Background Art
[0002] Reinforcement learning is a learning paradigm based on trial and error, which optimizes the behavior of an agent according to the rewards obtained from the environment. In the field of autonomous driving, deep learning is usually combined with reinforcement learning, that is, deep reinforcement learning. First, a deep neural network is used to extract environmental features, and then a reinforcement learning algorithm is used to make autonomous driving decisions, and finally the optimal driving strategy is learned.
[0003] However, there are also a series of challenging problems in applying deep reinforcement learning algorithms in the field of autonomous driving, and the most prominent one is the safety problem. Reinforcement learning needs to conduct interactive exploration of the environment during the learning process, sometimes involving risky operations, resulting in task failures and reducing learning efficiency. In practical applications, due to the black-box nature of deep reinforcement learning models, there is still a certain probability of outputting unexplainable dangerous decision results, leading to incorrect behaviors of autonomous driving vehicles, requiring manual takeover, and prone to safety problems, affecting the riding experience. Summary of the Invention
[0004] The main purpose of this application is to provide an autonomous driving decision control method, device, and vehicle, aiming to solve the technical problem that existing autonomous driving decision control methods have a probability of outputting unexplainable dangerous decision results, resulting in incorrect behaviors of autonomous driving vehicles.
[0005] To achieve the above object, this application proposes an autonomous driving decision control method, and the method includes: generating vehicle state information according to vehicle environmental information and vehicle working information; based on the vehicle state information, outputting a current action value through an autonomous driving decision model; based on the vehicle state information, verifying the vehicle safety through a responsibility-sensitive safety model; when the vehicle safety is qualified, sending the current action value to a vehicle control instruction receiver.
[0006] In one embodiment, the step of generating vehicle state information according to vehicle environmental information and vehicle working information includes: obtaining vehicle environmental information and vehicle working information; converting the vehicle environmental information into environmental coordinate information in the ego-vehicle coordinate system; based on vehicle navigation information and the environmental coordinate information, generating vehicle state information in combination with the vehicle working information.
[0007] In one embodiment, the step of verifying the vehicle safety through the responsibility-sensitive safety model based on the vehicle state information includes: obtaining the minimum safety distance through the responsibility-sensitive safety model based on the vehicle state information; judging the vehicle safety based on the minimum safety distance and the vehicle working information; when the vehicle working information meets the minimum safety distance, it is determined that the vehicle safety is qualified; when the vehicle working information does not meet the minimum safety distance, it is determined that the vehicle safety is unqualified and the allowable action range is output.
[0008] In one embodiment, after the step of issuing the current action value to the vehicle control instruction receiver when the vehicle safety is qualified, the following steps are further included: when the vehicle safety is unqualified, judging the safety of the current action value; when the current action value meets the allowable action range, issuing the current action value to the vehicle control instruction receiver; when the current action value does not meet the allowable action range, the responsibility-sensitive safety model generates a safe action value based on the allowable action range and issues the safe action value to the vehicle control instruction receiver.
[0009] In one embodiment, before the step of outputting the current action value through the autonomous driving decision-making model based on the vehicle state information, the following steps are further included: building an autonomous driving decision-making model based on the vehicle state information and the vehicle action value; performing a safety verification on the autonomous driving decision-making model through the responsibility-sensitive safety model to generate interaction information; generating an experience replay pool based on the interaction information; sampling the interaction information in the experience replay pool through the autonomous driving decision-making model to update the network parameters of the autonomous driving decision-making model.
[0010] In one embodiment, the step of performing a safety verification on the autonomous driving decision-making model through the responsibility-sensitive safety model to generate interaction information includes: generating a training action value in the current training environment through the autonomous driving decision-making model; verifying the safety in the current training environment through the responsibility-sensitive safety model; when the safety in the current training environment is qualified, matching a reward value to the training action value; when the safety in the current training environment is unqualified, verifying the safety of the training action value; when the safety of the training action value is unqualified, performing the step of generating a training action value in the current training environment through the autonomous driving decision-making model until the preset number of times is completed; when the safety of the training action value generated in the last time is unqualified, replacing the training action value with the safe action value generated by the responsibility-sensitive safety model; generating interaction information based on the corresponding training environment, training action value and reward value.
[0011] In addition, to achieve the above object, the present application further provides an autonomous driving decision control device, which includes: a parameter acquisition module, a calculation module, and a verification and output module; the parameter acquisition module is configured to generate vehicle state information based on vehicle environment information and vehicle working information; the calculation module is configured to output a current action value through an autonomous driving decision model based on the vehicle state information; the verification and output module is configured to verify the vehicle safety through a responsibility-sensitive safety model based on the vehicle state information; and is further configured to issue the current action value to a vehicle control instruction receiver when the vehicle safety is qualified.
[0012] In addition, to achieve the above object, the present application further provides a vehicle, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the autonomous driving decision control method as described above.
[0013] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the autonomous driving decision control method as described above.
[0014] In addition, to achieve the above object, the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the autonomous driving decision control method as described above.
[0015] One or more technical solutions proposed by the present application have at least the following technical effects:
[0016] For the autonomous driving scenario, the current state of the vehicle is obtained by collecting vehicle working information and environment information, and the action value is output through a deep reinforcement learning model, that is, an autonomous driving decision model. The responsibility-sensitive safety model is combined with the training process and deployment process of the deep reinforcement learning model to verify the safety, avoiding risk operations involved in vehicle interaction during the model training and deployment processes, and ensuring vehicle safety. Description of the Drawings
[0017] The drawings here are incorporated into the specification and form a part of the specification, showing the embodiments in line with the present application, and are used together with the specification to explain the principles of the present application.
[0018] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a schematic flowchart provided for the first embodiment of the automatic driving decision control method of this application;
[0020] Figure 2 It is another schematic flowchart provided for the first embodiment of the automatic driving decision control method of this application;
[0021] Figure 3 It is a schematic brief flowchart of the automatic driving decision control method provided for the second embodiment of this application;
[0022] Figure 4 It is a schematic diagram of the module structure of the automatic driving decision control device according to the embodiment of this application;
[0023] Figure 5 It is a schematic diagram of the device structure of the hardware operating environment involved in the automatic driving decision control method in the embodiment of this application.
[0024] The realization of the purpose, functional characteristics and advantages of this application will be further described with reference to the embodiments and the accompanying drawings. Specific Embodiments
[0025] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.
[0026] In order to better understand the technical solutions of this application, the following will be described in detail in combination with the drawings in the specification and the specific embodiments.
[0027] At present, there are also a series of challenging problems in the application of deep reinforcement learning algorithms in the field of autonomous driving. One of the most prominent is the safety problem. In the learning process of reinforcement learning, it is necessary to interactively explore the environment, which sometimes involves risky operations, resulting in task failure and reduced learning efficiency. In practical applications, due to the black-box nature of the deep reinforcement learning model, there is still a certain probability of outputting unexplainable dangerous decision results, leading to incorrect behaviors of autonomous vehicles, requiring human takeover, and prone to safety problems, affecting the riding experience.
[0028] Based on this, the embodiments of this application provide an automatic driving decision control method, referring to Figure 1 , Figure 1 It is a schematic flowchart of the first embodiment of the automatic driving decision control method of this application.
[0029] In this embodiment, the automatic driving decision control method includes steps S10 to S40:
[0030] Step S10, generating vehicle state information according to vehicle environment information and vehicle working information.
[0031] It should be noted that vehicle environmental information is the perception information of the environment around the host vehicle, usually including information about other traffic participants, obstacle information, etc. Information about other traffic participants is a crucial part of vehicle environmental information, which covers all other traffic entities sharing road resources in the same road space as the host vehicle. Such information usually includes: information about other vehicles, pedestrian information, and non-motor vehicle information. And obstacle information is another important component of vehicle environmental information, which covers any static or dynamic object on the road that may impede the driving of the host vehicle.
[0032] For example, information about other vehicles can include the precise position of other vehicles relative to the host vehicle, including lateral and longitudinal distances; the driving speed of other vehicles, which helps to determine whether they may pose a threat or potential collision risk; the driving direction of other vehicles, including going straight, turning, changing lanes, etc.; the size (such as length, width, height) and type (such as sedan, SUV, truck, etc.) of other vehicles, which helps to evaluate the potential collision consequences.
[0033] It should be noted that vehicle operation information mainly covers relevant data generated during the operation or driving of the host vehicle. These data are crucial for understanding the vehicle's operating state, making driving decisions, and subsequent vehicle maintenance, such as position, speed, heading angle, etc.
[0034] It should be noted that the vehicle state information generated based on vehicle environmental information and vehicle operation information is crucial for the autonomous driving system, vehicle monitoring system, and driver assistance system. At the same time, the state information can also include map and navigation information, which characterizes the current state of the vehicle through the environment and the vehicle's working conditions.
[0035] Specifically, this embodiment proposes a feasible implementation method to generate vehicle state information: obtain vehicle environmental information and vehicle operation information; convert the vehicle environmental information into environmental coordinate information in the host vehicle coordinate system; and generate vehicle state information based on the vehicle navigation information and the environmental coordinate information, in combination with the vehicle operation information.
[0036] It should be noted that vehicle environmental information obtains real-time data of the surrounding environment through vehicle sensors such as radar, cameras, lidar, etc., including the positions, speeds, and attribute information of other traffic participants such as vehicles, pedestrians, and road conditions such as lane lines, traffic signs, obstacles, etc.
[0037] It should be noted that vehicle operation information obtains the state data of the vehicle itself through the vehicle's own sensors and control systems, including position GPS coordinates, speed, acceleration, heading angle, engine state, braking system state, steering system state, etc.
[0038] It should be noted that the obtained vehicle environmental information is usually converted into coordinate information in the ego-vehicle coordinate system centered on the ego-vehicle from the global coordinate system or the relative coordinate system. This usually involves coordinate transformation and data processing to ensure that all environmental information is described relative to the position and orientation of the ego-vehicle.
[0039] It should be noted that vehicle navigation information includes the planned path, target position, estimated arrival time, etc. of the vehicle. This information usually comes from the navigation system of the vehicle or the path planning module of the autonomous driving system.
[0040] It can be understood that the vehicle navigation information is combined with the converted environmental coordinate information, and at the same time, vehicle working information such as the current speed, acceleration, heading angle, etc. is considered to generate vehicle state information reflecting the current state and driving environment of the vehicle, including the real-time position, speed, heading angle, distance from surrounding obstacles, collision risk level, lane keeping state, optimal path suggestion, fault diagnosis and warning, etc.
[0041] Step S20, based on the vehicle state information, output the current action value through the autonomous driving decision-making model.
[0042] It should be noted that the autonomous driving decision-making model is a deep learning reinforcement model. The vehicle state information is input into the deep reinforcement learning model, and the policy network outputs the distribution information of actions. After sampling, the action value in the current state, that is, the current action value, is output. The action value includes the vehicle acceleration and steering angle.
[0043] It should be noted that in the deep reinforcement learning model, the policy network is responsible for outputting the distribution information of actions according to the input state information. This means that for each possible action, the policy network will give a probability value indicating the possibility of selecting this action in the current state.
[0044] It should be noted that according to the action distribution information output by the policy network, the specific action value in the current state is determined through sampling (such as random sampling, greedy sampling, etc.). The action value usually includes the vehicle acceleration and steering angle, and these two parameters directly determine the driving mode and trajectory of the vehicle in the current state.
[0045] At this time, the directly output current action value has high uncertainty due to the black-box property of the autonomous driving decision-making model. Therefore, this solution introduces a responsibility-sensitive safety model to verify the vehicle safety.
[0046] Step S30, based on the vehicle state information, verify the vehicle safety through the responsibility-sensitive safety model.
[0047] It can be understood that the Responsibility-Sensitive Safety model (RSS) is a rule-based safety model that defines a series of parameters such as safety distances and response times to ensure that a vehicle does not collide with other traffic participants during driving.
[0048] It can be understood that based on the input vehicle state information, the RSS model calculates parameters such as the safety distance and response time between the current vehicle and other traffic participants, and compares them with the safety thresholds defined by the RSS model.
[0049] Specifically, first, it is necessary to obtain the minimum safety distance through the Responsibility-Sensitive Safety model based on the vehicle state information.
[0050] Specifically, in a multi-lane scenario, the longitudinal minimum safety distance between the host vehicle and surrounding vehicles is defined as follows:
[0051]
[0052] where the symbol + at the end of the formula means that the longitudinal minimum safety distance takes a positive value, and if it is negative, it takes 0. v r is the speed of the following vehicle, v f is the speed of the leading vehicle, ρ is the vehicle operation reaction time, v r,ρ = v r + ρa max,accel a max,accel refers to the maximum vehicle acceleration, a min,brake refers to the minimum vehicle deceleration, a max,brake refers to the maximum vehicle deceleration.
[0053] Specifically, in a multi-lane scenario, the lateral minimum safety distance between the host vehicle and surrounding vehicles is defined as follows:
[0054]
[0055] It can be understood that the symbol + at the end of the formula means that the lateral minimum safety distance takes a positive value, and if it is negative, it takes 0. μ is the vehicle width, v 1 is the speed of the vehicle on the left, v 2 is the speed of the vehicle on the right, ρ is the vehicle operation reaction time, refers to the lateral minimum deceleration.
[0056] Furthermore:
[0057]
[0058] Secondly, it is possible to judge the vehicle safety based on the minimum safety distance and vehicle working information.
[0059] Specifically, based on the above definitions of longitudinal and lateral minimum safety distances, a method for determining safe and dangerous states is defined:
[0060]
[0061] Among them, 0 represents a dangerous state (both horizontally and vertically in a dangerous state), 1 represents a safe state, and d lon is the longitudinal distance between the host vehicle and the surrounding vehicles, and d lat is the lateral distance between the host vehicle and the surrounding vehicles.
[0062] It can be understood that when the vehicle working information meets the minimum safety distance, the vehicle safety is determined to be qualified; when the vehicle working information does not meet the minimum safety distance, the vehicle safety is determined to be unqualified and the allowable action range is output.
[0063] Step S40, when the vehicle safety is qualified, send the current action value to the vehicle control instruction receiver.
[0064] It can be understood that when the vehicle safety is qualified, that is, the vehicle is in a safe state, that is, when there are no potential hazards that may cause a safety accident, the calculated current action value is sent to the vehicle control instruction receiver through a communication protocol (such as CAN bus, Ethernet, etc.). The vehicle control instruction receiver is a key component in the autonomous driving system responsible for receiving and executing control instructions. It ensures that the vehicle will not cause a safety accident due to incorrect control instructions during driving.
[0065] It can be understood that as Figure 2 shown, Figure 2 is another process schematic diagram provided by Embodiment 1 of the autonomous driving decision-making control method of the present application. After step S40, it further includes steps S50 to S70:
[0066] Step S50, when the vehicle safety is unqualified, judge the safety of the current action value.
[0067] It can be understood that when the vehicle safety is unqualified, that is, the vehicle is in a dangerous state, and it may face risks of collision with other traffic participants, risks of exceeding the road boundary, or other situations that may cause a safety accident. The safety of the current action value is further evaluated. This includes comprehensively considering aspects such as whether the current action value (such as acceleration, steering angle, etc.) will cause the vehicle to get closer to the dangerous state and whether it may trigger a more serious safety accident.
[0068] The basis for safety evaluation may include but is not limited to: the distance between the vehicle and surrounding obstacles, relative speed, road conditions, traffic rules, etc. The system needs to comprehensively consider these factors to judge whether the current action value is safe. Specifically, in this embodiment, an allowable action range is given to assist the system in judgment.
[0069] Step S60: When the current action value meets the allowable action range, send the current action value to the vehicle control command receiver.
[0070] Step S70: When the current action value does not meet the allowable action range, the responsible sensitive safety model generates a safety action value based on the allowable action range and sends the safety action value to the vehicle control command receiver.
[0071] It can be understood that the responsible sensitive safety model provides appropriate longitudinal and lateral responses to delimit the allowable action range. Specifically, when it is determined that the host vehicle is in a longitudinal dangerous state, the appropriate longitudinal response is as follows (here only the case of vehicles driving in the same direction is involved): hereinafter, c 1 refers to the following vehicle, c 2 refers to the preceding vehicle, ρ refers to the response time to the dangerous state, and it is defined that is the earliest time for the preceding and following vehicles to enter the dangerous state.
[0072] Within the time period (reaction time), it is allowed for c l to run at a maximum acceleration of a max,accel at most; At and after the moment of 1 it is allowed for c min,brake to run at a minimum longitudinal deceleration of a 2 at least until reaching the longitudinal safety state. It is allowed for c max,brake to run at a maximum longitudinal deceleration of a max,brake until reaching the longitudinal safety state. In summary, -a min,brake .
[0073] Specifically, when it is determined that the host vehicle is in a lateral dangerous state, the appropriate lateral response is as follows: hereinafter, c 1 refers to the vehicle on the left, c 2 refers to the vehicle on the right, ρ refers to the response time to the dangerous state, and it is defined that is the earliest time for the laterally adjacent vehicle to enter the dangerous state.
[0074] Within the time period (reaction time), the accelerations of the two laterally adjacent vehicles need to satisfy that is, the absolute values of the lateral accelerations are both less than the lateral maximum acceleration constraint. At and after the moment of 1 and c 2 are allowed to run at a minimum lateral deceleration of at least until reaching the lateral safety state. In summary,
[0075] It should be noted that the safe action value can be randomly selected from the appropriate lateral and longitudinal action ranges provided by RSS.
[0076] In this embodiment, for the autonomous driving scenario, the current state of the vehicle is obtained by collecting vehicle working information and environmental information. The action value is output by the deep reinforcement learning model, that is, the autonomous driving decision-making model. The training process and deployment process of the responsibility-sensitive safety model and the deep reinforcement learning model are combined to verify the safety, avoiding risky operations involved in the interactive exploration of the vehicle during the model training process and ensuring the safety of the vehicle.
[0077] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3 , Figure 3 Before step S20 in the brief flowchart of the autonomous driving decision control method provided in the second embodiment of the present application, the autonomous driving decision control method further includes steps A10 to A40:
[0078] Step A10, build an autonomous driving decision model based on vehicle state information and vehicle action values.
[0079] Specifically, based on the above description, the present application uses the state space to represent vehicle state information and the action space to represent action values. The state space refers to the set of all possible states that can be observed by the environment at each time step. The size of the state space determines the complexity of the reinforcement learning problem. The action space is the set of all possible actions that the agent can execute at each time step. The action can be discrete, such as turning left, turning right, stopping, etc.; or it can be continuous, such as selecting a value within a certain range. The size of the action space also affects the complexity of the problem.
[0080] It can be understood that by defining the state space and the action space, the reinforcement learning problem can be formalized as a Markov Decision Process (MDP). MDP is a five-tuple (S, A, P, R, γ), where S is the state space, representing possible observed states. A is the action space, representing possible actions of the agent. P is the state transition probability function, representing the probability of transitioning to the next state after taking a certain action in a certain state. R is the reward function, representing the immediate reward obtained by the agent after taking a certain action in a certain state. γ is the discount factor, used to measure the importance of future rewards to immediate rewards.
[0081] Based on the definition of MDP, the reinforcement learning algorithm learns a policy to determine which action to choose in each state, so that the agent can obtain the maximum cumulative reward. The policy can be deterministic (which action to choose in each state) or probabilistic (obtaining the probability distribution of actions in different states).
[0082] Specifically, for the multi-lane autonomous driving scenario, the state space S = [S ego , S sur , S map , S navi , including vehicle working information S ego , vehicle environment information S sur , local map information S map , and navigation information S navi . The more specific attribute information is as follows:
[0083] The ego-vehicle information S ego contains the following attributes: S ego = [p x , p y , v ego , φ ego , n ego , l ego , w ego , where (p x , p y ) represents the position of the ego-vehicle, v ego represents the speed of the ego-vehicle, φ ego represents the heading angle of the ego-vehicle, n ego represents the lane where the ego-vehicle is currently located, l ego , w ego represent the length and width of the ego-vehicle respectively.
[0084] The surrounding vehicle information S sur contains the following attributes: Among them, represents the position of the surrounding vehicle j, v j represents the speed of the surrounding vehicle j, φ ′ represents the heading angle of the surrounding vehicle j, n j represents the lane where the surrounding vehicle j is currently located, l j , w j represent the length and width of the surrounding vehicle j respectively.
[0085] The local map information S map is represented by a directed graph G: S map = G(ρ, ε), where ρ represents lane information and ε represents the connection relationship between lanes.
[0086] The navigation information S naviProvides the set of lanes that the autonomous vehicle can drive on at the current moment: S navi = [Lanes that can currently be driven on].
[0087] For the multi-lane autonomous driving scenario, the action space A is defined as follows, including two variables, where a represents the longitudinal acceleration and δ represents the steering wheel angle: A = (a, δ).
[0088] Specifically, in the reinforcement learning algorithm, in the field of deep reinforcement learning, the Soft Actor-Critic (SAC) algorithm is a representative method. It combines the methods of policy gradient and value function approximation, and learns the policy by optimizing the objective function, enabling the agent to make optimal decisions based on the environmental state. It uses two neural networks: the Actor network and the Critic network. The Actor network is the policy network, which takes the current state as input and outputs the probability distribution of actions, and then the Critic network evaluates the quality of the actions to adjust the policy.
[0089] In SAC, since the output of the policy network is a probability distribution, the agent needs to sample to obtain action data and use this data to update the policy parameters. During the sampling process, the idea of reparameterization is used, that is, the mean and standard deviation of the action distribution are obtained from the policy network, and then a random variable is introduced, and an independent Gaussian distribution is constructed using the mean and standard deviation to facilitate sampling and gradient calculation. The way to obtain the action value is as follows:
[0090]
[0091] where, and are the mean and standard deviation learned from the policy network, ∈ is a random variable sampled from the standard Gaussian distribution, and tanh is used to limit the action range.
[0092] Step A20, verify the safety of the autonomous driving decision model based on the responsibility-sensitive safety model, and generate interaction information. Step A30, generate an experience replay pool based on the interaction information. Step A40, sample interaction information in the experience replay pool through the autonomous driving decision model, and update the network parameters of the autonomous driving decision model.
[0093] It can be understood that in SAC, the agent collects interaction information by interacting with the environment. The interaction information includes the state, action, reward, and the next state.
[0094] In a feasible implementation, this embodiment presents a solution for generating interaction information and giving reward values: generating training action values in the current training environment through an autonomous driving decision-making model; verifying the safety in the current training environment through a responsibility-sensitive safety model; when the safety in the current training environment is qualified, matching reward values for the training action values; when the safety in the current training environment is unqualified, verifying the safety of the training action values; when the safety of the training action values is unqualified, executing the step of generating training action values in the current training environment through the autonomous driving decision-making model until the preset number of times is completed; when the safety of the training action values generated in the last time is unqualified, replacing the training action values with the safety action values generated by the responsibility-sensitive safety model; generating interaction information based on the corresponding training environment, training action values, and reward values.
[0095] It can be understood that the preset number of times can be set by oneself, and usually 3 times is the best choice.
[0096] It can be understood that these data are stored in an experience replay buffer for subsequent training. The objective function of SAC consists of two parts: the policy optimization objective and the value function approximation objective. The policy optimization objective is to maximize the expected return and obtain a better policy by optimizing the policy parameters; the value function approximation objective is to make the state value function and state-action value function estimated by the Critic network as close as possible to the true values.
[0097] During the training process, SAC uses the method of stochastic gradient descent to update the weights of the neural network. First, randomly sample a batch of data from the experience replay buffer; then, calculate the gradient of the objective function with respect to the weights of the neural network; finally, use the gradient descent method to update the weights. This process is continuously repeated until convergence or the specified number of training epochs is reached.
[0098] The Gaussian distribution mean and standard deviation of the actions are output by the trained policy network, and a set of specific action values (a * , δ * ) are sampled.
[0099] It can be understood that first, in combination with the actual vehicle speed of the host vehicle, the surrounding vehicle speeds output by the vehicle perception module, the set upper and lower limits of the vehicle acceleration, and the response time, the longitudinal minimum safety distance d lon,min and the lateral minimum safety distance d lat,min are calculated, and then, according to the actual longitudinal distance d lon between the host vehicle and the surrounding vehicles and the actual lateral distance d lat output by the perception module, the safety of the host vehicle state is judged. If rss state = 1, it means that the host vehicle state is safe and the model output action values (a * , δ* ), if rss state = 0, it indicates that the state of the host vehicle is dangerous. When the host vehicle is in a dangerous state, obtain the appropriate longitudinal and lateral responses (a lon_res , a lat_res ) provided by the RSS module. Among them, both the longitudinal and lateral responses refer to a range rather than a specific value. To further evaluate the feasibility of the action values (a * , δ * ), convert the lateral response into the steering wheel angle range. The conversion formula is as follows (the mathematical relationship between the lateral acceleration of the vehicle and the steering wheel angle):
[0100] β = k·δ, where k is the ratio of the steering wheel to the front wheel angle, and δ is the steering wheel angle.
[0101] R = L / sin(β), where L is the wheelbase and R is the turning radius.
[0102] a x = v 2 / R, where v is the vehicle speed and a x is the lateral acceleration.
[0103] From the above formulas, it can be deduced that: a x = v 2 ·sin(k·δ) / L, that is, δ = arcsin((a x ·L) / v 2 ) / k and δ lat_res = arcsin((a lat_res ·L) / v 2 ) / k.
[0104] Finally, evaluate the feasibility of the action values (a lon_res , δ lat_res ) according to the appropriate longitudinal and lateral actions (a * , δ * ). If the action values (a * , δ * ) are all within the appropriate longitudinal and lateral action ranges, it indicates that this set of action values is feasible and allows execution. If the action values exceed the appropriate lateral or longitudinal action ranges, action values within the appropriate longitudinal and lateral action ranges provided by the RSS can be randomly selected and used as control commands to be issued.
[0105] In this embodiment, the responsibility - sensitive safety model is combined with the training process and the deployment and application process of deep reinforcement learning. The appropriate response values provided by the responsibility - sensitive safety model are used to timely evaluate whether the output actions of the reinforcement learning policy network are safe and to provide alternative safe actions, ensuring the safety of the model training process and the deployment and application process.
[0106] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the automatic driving decision control method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.
[0107] This application also provides an automatic driving decision control device. Please refer to Figure 4 , the automatic driving decision control device includes: a parameter acquisition module 10, a calculation module 20, and a verification and output module 30; the parameter acquisition module 10 is used to generate vehicle state information according to vehicle environment information and vehicle working information; the calculation module 20 is used to output a current action value through an automatic driving decision model based on the vehicle state information; the verification and output module 30 is used to verify the vehicle safety through a responsibility-sensitive safety model based on the vehicle state information; and is further used to issue the current action value to a vehicle control instruction receiver when the vehicle safety is qualified.
[0108] The automatic driving decision control device provided by this application adopts the automatic driving decision control method in the above embodiment, and can solve the technical problem that the existing automatic driving decision control method may probabilistically output unexplainable dangerous decision results, resulting in incorrect behaviors of automatic driving vehicles. Compared with the prior art, the beneficial effects of the automatic driving decision control device provided by this application are the same as those of the automatic driving decision control method provided by the above embodiment, and other technical features in the automatic driving decision control device are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.
[0109] This application provides a vehicle, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the automatic driving decision control method in the first embodiment above.
[0110] Next, refer to Figure 5 , which shows a schematic structural diagram of a vehicle suitable for implementing the embodiments of this application. The vehicle in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5The vehicle shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0111] As Figure 5 shown, the vehicle may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for vehicle operation are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the vehicle to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a vehicle having various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be alternatively implemented or had.
[0112] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart may be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the method of the embodiments disclosed in the present application are executed.
[0113] The vehicle provided by the present application adopts the automatic driving decision control method in the above embodiment, and can solve the technical problem that the existing automatic driving decision control method may probabilistically output an unexplainable dangerous decision result, resulting in incorrect behavior of the automatic driving vehicle. Compared with the prior art, the beneficial effects of the vehicle provided by the present application are the same as those of the automatic driving decision control method provided in the above embodiment, and other technical features in the vehicle are the same as those disclosed in the method of the previous embodiment, and will not be elaborated herein.
[0114] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0115] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0116] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the autonomous driving decision control method in the above embodiments.
[0117] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0118] The above computer-readable storage medium can be included in a vehicle; it can also exist separately without being assembled into the vehicle.
[0119] The above computer-readable storage medium carries one or more programs, which, when executed by a vehicle, cause the vehicle to: generate vehicle state information based on vehicle environment information and vehicle working information; output a current action value through an autonomous driving decision-making model based on the vehicle state information; verify the vehicle safety through a responsibility-sensitive safety model based on the vehicle state information; and issue the current action value to a vehicle control instruction receiver when the vehicle safety is qualified.
[0120] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the block may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0122] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.
[0123] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned autonomous driving decision control method, and can solve the technical problem that the existing autonomous driving decision control method may probabilistically output unexplainable dangerous decision results, resulting in incorrect behaviors of autonomous driving vehicles. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the autonomous driving decision control method provided by the above embodiment, and will not be elaborated here.
[0124] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the autonomous driving decision control method as described above.
[0125] The computer program product provided by this application can solve the technical problem that the existing autonomous driving decision control method may probabilistically output unexplainable dangerous decision results, resulting in incorrect behaviors of autonomous driving vehicles. Compared with the prior art, the beneficial effects of the computer program product provided by this application are the same as those of the autonomous driving decision control method provided by the above embodiment, and will not be elaborated here.
[0126] The above are only some embodiments of this application, and thus do not limit the patent scope of this application. Any equivalent structural transformation made under the technical concept of this application by using the content of the specification and drawings of this application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of this application.
Claims
1. An automatic driving decision control method, characterized in that: The method comprises: Generate vehicle status information based on vehicle environment information and vehicle operation information; Based on the vehicle state information, output a current action value through an automatic driving decision model; Based on the vehicle status information, verifying vehicle safety through a responsibility-sensitive safety model; When the vehicle safety is qualified, the current action value is sent to the vehicle control instruction receiver.
2. The automatic driving decision control method according to claim 1, characterized in that: The step of generating vehicle status information according to vehicle environment information and vehicle operation information comprises: Obtain vehicle environment information and vehicle working information; Converting the vehicle environment information into environment coordinate information in the vehicle coordinate system; The vehicle state information is generated based on the vehicle navigation information and the environmental coordinate information in combination with the vehicle operation information.
3. The automatic driving decision control method according to claim 2, characterized in that: The step of verifying the vehicle safety through a responsibility-sensitive safety model based on the vehicle status information includes: Based on the vehicle state information, obtaining a minimum safety distance through a responsibility-sensitive safety model; Determining vehicle safety based on the minimum safety distance and vehicle operating information; When the vehicle operating information meets the minimum safety distance, the vehicle safety is deemed qualified; When the vehicle operation information does not meet the minimum safety distance, the vehicle safety is deemed unqualified and the allowed action range is output.
4. The automatic driving decision control method according to claim 3, characterized in that: After the step of sending the current action value to the vehicle control instruction receiver when the vehicle safety is qualified, the step further includes: When the vehicle safety is unqualified, judging the safety of the current action value; When the current action value meets the allowable action range, sending the current action value to the vehicle control instruction receiver; When the current action value does not satisfy the allowable action range, the responsibility-sensitive safety model generates a safety action value based on the allowable action range, and sends the safety action value to the vehicle control instruction receiver.
5. The automatic driving decision control method according to any one of claims 1 to 4, characterized in that: The step of outputting the current action value through the automatic driving decision model based on the vehicle state information also includes: Build an autonomous driving decision model based on vehicle status information and vehicle action values; Performing safety verification on the autonomous driving decision model based on a responsibility-sensitive safety model to generate interaction information; Generate an experience replay pool based on the interactive information; The autonomous driving decision model samples interaction information in the experience replay pool and updates network parameters of the autonomous driving decision model.
6. The automatic driving decision control method according to claim 5, characterized in that: The step of performing safety verification on the autonomous driving decision model based on the responsibility-sensitive safety model and generating interactive information comprises: Generate training action values in the current training environment through the autonomous driving decision model; Verify the safety of the current training environment through the responsibility-sensitive safety model; When the safety is qualified under the current training environment, matching the reward value for the training action value; When the safety is unqualified under the current training environment, verifying the safety of the training action value; When the safety of the training action value is unqualified, the step of generating the training action value in the current training environment by using the autonomous driving decision model is executed until a preset number of times is completed; When the security of the last generated training action value fails to meet the requirements, the training action value is replaced by the security action value generated by the responsibility-sensitive security model; Generate interaction information based on the corresponding training environment, training action value and reward value.
7. An automatic driving decision control device, characterized in that: The device comprises: a parameter acquisition module, a calculation module and a verification output module; The parameter acquisition module is used to generate vehicle status information according to vehicle environment information and vehicle working information; The calculation module is used to output a current action value through an automatic driving decision model based on the vehicle state information; The verification output module is used to verify the vehicle safety through a responsibility-sensitive safety model based on the vehicle status information; and is also used to send the current action value to the vehicle control instruction receiver when the vehicle safety is qualified.
8. A vehicle, characterized in that: The vehicle includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the automatic driving decision control method as described in any one of claims 1 to 6.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the automatic driving decision control method as described in any one of claims 1 to 6 are implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the steps of the autonomous driving decision control method according to any one of claims 1 to 6.
Citation Information
Cited By
World model driven online-continuous learning automatic driving decision control method and system
CN122481777A