Vehicle control method and device based on deep reinforcement learning, vehicle and medium
Through the vehicle control method constructed by deep reinforcement learning, the torque request upper limit value is determined in real time, which solves the problem of cumbersome development of error acceleration protection strategies, realizes safe driving in complex traffic environments, and improves user experience.
Patent Information
- Application Number
- CN202510556844.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, the formulation of misacceleration protection strategies is cumbersome and difficult to adapt to complex traffic environments, which makes it difficult to avoid safety hazards caused by drivers accidentally stepping on the accelerator.
Based on deep reinforcement learning, the vehicle control method is used to construct state space, action space and reward functions, train protection models, determine the upper limit of torque request in real time, and use bicycle driving data, driver's intentions and environmental data to control torque to prevent collisions.
It realizes real-time collision prevention in complex traffic environments, improves user driving experience, avoids safety hazards caused by accidentally stepping on the accelerator, and adapts to changing driving scenarios.
Smart Images

Figure CN120288037A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle control, and in particular to a vehicle control method, device, vehicle and medium based on deep reinforcement learning. Background Art
[0002] With the continuous development of hardware computing power and software algorithms, the number of sensors such as cameras and radars installed in vehicles has been increasing continuously, significantly improving the environmental perception ability of automobiles and promoting the rapid development of autonomous driving technology. Moreover, this rich environmental target information also brings diverse possibilities for optimizing the manual driving experience of drivers, which has become the focus of attention in the industry, academia and research.
[0003] During manual driving, the driver may accidentally step on the accelerator due to reasons such as environmental lighting, line of sight obstruction, and inattention, resulting in dangerous situations such as collisions. To solve this problem, in related technologies, rule-based control algorithms are adopted, such as using the throttle pedal opening, the change rate of the throttle pedal opening increase, and the collision time between the vehicle and obstacles in the driving direction to design the trigger conditions for the false acceleration protection strategy to limit the output torque, etc. However, for such rule-based control algorithms, different rule conditions and parameters need to be designed for different scenarios, resulting in cumbersome work for formulating the false acceleration protection strategy and being difficult to adapt to complex traffic environments. Summary of the Invention
[0004] In view of this, the present invention provides a vehicle control method, device, vehicle and medium based on deep reinforcement learning to solve the problem in related technologies that the work of formulating the false acceleration protection strategy is cumbersome and difficult to adapt to complex traffic environments.
[0005] In a first aspect, the present invention provides a vehicle control method based on deep reinforcement learning, the method comprising:
[0006] Constructing a state space of deep reinforcement learning based on the self-vehicle driving data, the driver's speed control intention and the self-vehicle surrounding environment data;
[0007] Constructing an action space of deep reinforcement learning based on the torque request upper limit value;
[0008] Constructing a reward function of deep reinforcement learning based on safety and efficiency;
[0009] Constructing and training a protection model based on deep reinforcement learning based on the state space, action space and reward function;
[0010] Inputting the current self-vehicle driving data, the current driver's speed control intention and the self-vehicle surrounding environment data into the trained protection model to obtain the current torque request upper limit value;
[0011] Based on the magnitude relationship between the current driver torque request and the upper limit value of the current torque request, torque control is performed on the host vehicle.
[0012] The present invention constructs the state space of deep reinforcement learning by utilizing the host vehicle driving data, the driver's speed control intention, and the host vehicle surrounding environment data, constructs the action space of deep reinforcement learning based on the upper limit value of the torque request, and establishes and trains a protection model based on deep reinforcement learning by constructing a reward function of deep reinforcement learning based on safety and efficiency. During the actual driving process of the vehicle, the upper limit value of the torque request can be determined in real time according to the changes of the host vehicle and the environment and the driver's torque request, and torque control is performed on the vehicle by comparing the driver torque request and the upper limit value of the torque request to prevent the occurrence of dangerous situations such as collisions. Moreover, the upper limit value of the torque request changes dynamically in real time according to the surrounding environment information of the vehicle, is not restricted by the driving environment, can be applied to complex traffic environments, avoids potential safety hazards caused by the driver accidentally stepping on the accelerator, and improves the user experience.
[0013] In an alternative embodiment, the torque control of the host vehicle based on the magnitude relationship between the current driver torque request and the upper limit value of the current torque request includes:
[0014] Judge whether the current driver torque request is greater than the upper limit value of the current torque request;
[0015] When the current driver torque request is greater than the upper limit value of the current torque request, torque control is performed on the host vehicle according to the upper limit value of the current torque request.
[0016] When the current driver torque request is greater than the upper limit value of the current torque request, in order to avoid dangerous situations such as collisions, the present invention performs torque control on the host vehicle according to the upper limit value of the current torque request, ensuring the driving safety of the vehicle and improving the user experience.
[0017] In an alternative embodiment, the method further includes:
[0018] When the current driver torque request is not greater than the upper limit value of the current torque request, torque control is performed on the host vehicle according to the current driver torque request.
[0019] When the current driver torque request is not greater than the upper limit value of the current torque request, in order to better respond to the user's driving intention, the present invention performs torque control on the host vehicle according to the previous driver torque request, further improving the user experience while ensuring the driving safety of the user.
[0020] In an alternative embodiment, the method further includes:
[0021] After detecting a change in the current driver torque request, return to the step of inputting the current self-vehicle driving data, the current driver speed control intention, and the self-vehicle surrounding environment data into the trained protection model to obtain the current torque request upper limit value.
[0022] The present invention monitors the driver torque request and recalculates the torque request upper limit value after the change, so as to provide real-time protection against accidentally stepping on the accelerator for the user during the process of driving the vehicle, further ensuring the driving safety of the user and improving the user experience.
[0023] In an optional implementation manner, the self-vehicle driving data at least includes: the longitudinal speed, longitudinal acceleration, and heading angle of the self-vehicle, and the self-vehicle surrounding environment data at least includes: the longitudinal speed, longitudinal acceleration, and heading angle of each target object in the self-vehicle surrounding environment and the longitudinal distance between the self-vehicle and the target object.
[0024] The present invention constructs the state space of deep reinforcement learning based on parameters such as the longitudinal speed, longitudinal acceleration, heading angle of the self-vehicle, the longitudinal speed, longitudinal acceleration, and heading angle of each target object in the self-vehicle surrounding environment, and the longitudinal distance between the self-vehicle and the target object, so as to comprehensively consider the influence of factors related to the self-vehicle acceleration condition during the driving process of the vehicle, further improving the accuracy of the torque request upper limit value output by the protection model based on deep reinforcement learning, ensuring that the torque request upper limit value can prevent the occurrence of dangerous situations such as collisions and avoid restricting the normal torque demand of the driver in a safe situation, and improving the user experience.
[0025] In an optional implementation manner, constructing a protection model based on deep reinforcement learning based on the state space, action space, and reward function includes:
[0026] Constructing an online policy network, a target policy network, an online Q network, and a target Q network respectively using a neural network based on the state space, action space, and reward function;
[0027] Among them, the online policy network is used to calculate the first action to be executed in the current state at the current moment; the target policy network is used to calculate the second action to be executed in the next state transferred after executing the first action at the next moment; the online Q network is used to evaluate the value of executing the first action in the current state at the current moment; the target Q network is used to evaluate the value of executing the second action in the next state at the next moment.
[0028] The present invention constructs a protection model based on deep reinforcement learning, avoiding the dependence on rule formulation and parameter calibration. The required environmental state information is input into a neural network, and the action value is output for execution. Through continuous interaction with the environment, a reward value containing indicative information is obtained. The neural network is continuously trained and updated through the reward value, so as to realize the learning of the correct control strategy and the adaptability to complex traffic scenarios.
[0029] In an alternative embodiment, the training process of the protection model based on deep reinforcement learning includes:
[0030] Input the current state corresponding to the vehicle itself at the current moment into the online policy network, output the first action, and calculate the reward obtained at the current moment;
[0031] Input the first action into a preset simulation environment and vehicle model to obtain the next state corresponding to the vehicle itself at the next moment;
[0032] Take the current state, the first action, the reward, and the next state corresponding to each moment as an experience sample to form experience sample data;
[0033] Establish an experience pool, store the experience sample data in the experience pool, randomly extract several samples from the experience pool, and update the parameters corresponding to the online Q network and the online policy network;
[0034] Soft-update the parameters corresponding to the target policy network and the target Q network according to the updated parameters of the online Q network and the online policy network.
[0035] The present invention trains the protection model based on deep reinforcement learning by using the vehicle model of the vehicle itself and the simulation environment, realizing the simulation of the scenarios of accidentally stepping on the accelerator and normal driving in the real scenario, so that the trained protection model can be applicable to complex and changeable driving scenarios, improving the accuracy of the torque request upper limit value, further improving the control accuracy of the user's accidental acceleration protection strategy, and enhancing the user experience.
[0036] In a second aspect, the present invention provides a vehicle control device based on deep reinforcement learning, and the device includes:
[0037] A first processing module, configured to construct a state space of deep reinforcement learning based on the driving data of the vehicle itself, the driver's speed control intention, and the environmental data around the vehicle itself;
[0038] A second processing module, configured to construct an action space of deep reinforcement learning based on the torque request upper limit value;
[0039] A third processing module, configured to construct a reward function of deep reinforcement learning based on safety and efficiency;
[0040] A fourth processing module, configured to construct and train a protection model based on deep reinforcement learning based on the state space, action space, and reward function;
[0041] A fifth processing module, configured to input the current vehicle driving data, the current driver's speed control intention, and the vehicle surrounding environment data into the trained protection model to obtain the current torque request upper limit value;
[0042] A sixth processing module, configured to perform torque control on the vehicle based on the magnitude relationship between the current driver torque request and the current torque request upper limit value.
[0043] In a third aspect, the present invention provides a vehicle, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method described in the first aspect and any one of its optional embodiments.
[0044] In a fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the method provided in the first aspect or any one of its corresponding embodiments.
[0045] Advantages of the present invention:
[0046] By using the vehicle driving data, the driver's speed control intention, and the vehicle surrounding environment data, the present invention constructs a state space of deep reinforcement learning, constructs an action space of deep reinforcement learning based on the torque request upper limit value, and constructs a reward function of deep reinforcement learning based on safety and efficiency to establish and train a protection model based on deep reinforcement learning. During the actual driving process of the vehicle, the torque request upper limit value can be determined in real time according to the changes of the vehicle and the environment and the driver's torque request, and the vehicle is torque-controlled by comparing the driver torque request and the torque request upper limit value to prevent the occurrence of dangerous situations such as collisions. Moreover, the torque request upper limit value changes dynamically in real time according to the vehicle surrounding environment information, is not restricted by the driving environment, can be applied to complex traffic environments, avoids potential safety hazards caused by the driver accidentally stepping on the accelerator, and improves the user experience. Description of the Drawings
[0047] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0048] Figure 1It is a flowchart of a vehicle control method based on deep reinforcement learning according to an embodiment of the present invention;
[0049] Figure 2 It is a flowchart of another vehicle control method based on deep reinforcement learning according to an embodiment of the present invention;
[0050] Figure 3 It is a logical framework diagram of a vehicle control method based on deep reinforcement learning according to an embodiment of the present invention;
[0051] Figure 4 It is a schematic structural diagram of a vehicle control device based on deep reinforcement learning according to an embodiment of the present invention;
[0052] Figure 5 It is a schematic structural diagram of a vehicle according to an embodiment of the present invention. Detailed implementation manners
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] During the manual driving process, the driver may accidentally step on the accelerator due to reasons such as environmental lighting, line of sight obstruction, and inattention, resulting in dangerous situations such as collisions; in such cases, if the sensor information can be used to judge the misacceleration scenario and limit the output torque, the driving safety during the human-driven process can be effectively improved.
[0055] To solve the above problems, most of the related technologies are rule-based algorithms. For example, by using the accelerator pedal opening, the change rate of the increase in the accelerator pedal opening, and the collision time between the vehicle and the obstacle in the driving direction to design the triggering conditions of the misacceleration protection strategy. However, this type of algorithm requires designing different rule conditions and parameters for different scenarios, and the strategy formulation work is cumbersome.
[0056] Different from the traditional rule-based control algorithms, the learning-based algorithms can get rid of the dependence on a large number of rule formulations and parameter calibrations; among them, the deep reinforcement learning algorithm can input the required environmental state information into the neural network and output the action value for execution, and can obtain the reward value containing indicative information through continuous interaction with the environment. The neural network is continuously trained and updated through the reward value, so as to realize the learning of the correct control strategy.
[0057] Based on this, an embodiment of the present invention provides a driver mis-acceleration protection control scheme based on deep reinforcement learning. Specifically, a neural network is used to receive the information of the vehicle itself and the environment, and the upper limit value of the torque request is used as the output of the neural network. By restricting the torque request in dangerous scenarios, dangerous situations caused by the driver accidentally stepping on the accelerator pedal to accelerate can be prevented.
[0058] According to an embodiment of the present invention, there is provided an embodiment of a vehicle control method based on deep reinforcement learning. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0059] In this embodiment, a vehicle control method based on deep reinforcement learning is provided, which can be applied to vehicle controllers such as single-chip microcomputers, MCU and other control chips. Figure 1 It is a flowchart of a vehicle control method based on deep reinforcement learning according to an embodiment of the present invention, as Figure 1 shown, the process includes the following steps:
[0060] Step S101, based on the vehicle's own driving data, the driver's speed control intention, and the vehicle's surrounding environment data, construct the state space of deep reinforcement learning.
[0061] Among them, the vehicle's own driving data is vehicle data related to the risk of vehicle collision, such as the vehicle's own speed, acceleration, heading angle, etc., and the present invention is not limited thereto. The driver's speed control intention can be characterized by detecting the throttle opening and the throttle opening change rate corresponding to the driver's operation on the throttle pedal. The vehicle's surrounding environment data includes the driving data of other vehicles driving around the vehicle and the road conditions data of the driving road around the vehicle, etc., which are data related to the driving safety of the vehicle. The driving data of other vehicles driving around the vehicle includes data similar to the above-mentioned vehicle's own driving data, as well as other data associated with the vehicle, such as: the longitudinal distance and angle with the vehicle, etc. The road conditions data can include: road slope, friction performance, etc., and the present invention is not limited thereto.
[0062] Specifically, the above-mentioned vehicle's own driving data, the driver's speed control intention, and the vehicle's surrounding environment data can be collected by various sensors installed on the vehicle or obtained through communication with surrounding vehicles. The specific collection and communication processes are prior art and will not be elaborated here.
[0063] Exemplarily, constructing the state space S of deep reinforcement learning includes the longitudinal speed v of the vehicle e , the longitudinal acceleration a e , the heading angle δ e , the throttle opening T, the throttle opening change rate The longitudinal velocity v of the front target o , the longitudinal acceleration a o , the heading angle δ o , the longitudinal distance x o , the ramp acceleration a s , specifically as shown in formula (1):
[0064]
[0065] Step S102, construct the action space of deep reinforcement learning based on the torque request upper limit value.
[0066] Specifically, the torque request upper limit value is the maximum torque request allowed for the vehicle to respond. By setting this torque request upper limit value, dangerous situations such as collisions can be prevented when the vehicle driver accidentally steps on the accelerator.
[0067] Exemplarily, construct the action space A of deep reinforcement learning, including the torque request upper limit value T max , specifically as shown in formula (2):
[0068] A = {T max} (2)
[0069] This value is used to limit the driver's torque request, avoiding excessive torque requests caused by accidentally stepping on the accelerator and resulting in collisions. When in a safe situation, this value is relatively large and will not limit the driver's torque request, as shown in formula (3):
[0070] T final = min(T dri , T max ) (3)
[0071] Among them, T final is the torque request after final arbitration, and T dri is the driver's torque request.
[0072] Step S103, construct the reward function of deep reinforcement learning based on safety and efficiency.
[0073] Exemplarily, construct the reward function R of deep reinforcement learning, including the safety reward R safe and the efficiency reward R eff , specifically as shown in formula (4):
[0074] R = w1R safe + w2R eff (4)
[0075] Among them, w1 and w2 are the weight coefficients of each reward. The safety reward R safe is used to prevent the host vehicle from colliding, and the efficiency reward R effIt is used to avoid the upper limit value of the torque request in the output from restricting the normal torque demand of the driver in a safe situation.
[0076] Step S104: Construct and train a protection model based on deep reinforcement learning based on the state space, action space, and reward function.
[0077] Specifically, the construction and training methods of the deep reinforcement learning model in the existing technology can be adopted and implemented based on the state space, action space, and reward function determined in the above steps, which will not be elaborated here.
[0078] Step S105: Input the current self-vehicle driving data, the current driver's speed control intention, and the self-vehicle surrounding environment data into the trained protection model to obtain the current upper limit value of the torque request.
[0079] Specifically, during the driving process of the self-vehicle, when the driver's speed control intention is detected, that is, after the driver operates the accelerator pedal, the current throttle opening, the throttle opening change rate, together with the currently collected current self-vehicle driving data and the self-vehicle surrounding environment data, will be input into the trained protection model. The model analyzes the upper limit value of the torque request in the current driving environment, providing a limit value for subsequent torque control of the self-vehicle to ensure the driving safety of the self-vehicle.
[0080] Step S106: Perform torque control on the self-vehicle based on the magnitude relationship between the current driver's torque request and the current upper limit value of the torque request.
[0081] Specifically, the driver's torque request can be calculated by real-time monitoring of the throttle opening, throttle change rate, etc. The specific calculation process is the existing technology and will not be elaborated here. By comparing the driver's torque request with the upper limit value of the torque request output by the above protection model, the torque control of the self-vehicle can be realized, avoiding the occurrence of dangerous situations such as collisions.
[0082] In the embodiment of the present invention, by using the self-vehicle driving data, the driver's speed control intention, and the self-vehicle surrounding environment data, a state space of deep reinforcement learning is constructed, and an action space of deep reinforcement learning is constructed based on the upper limit value of the torque request. A protection model based on deep reinforcement learning is established and trained by constructing a reward function of deep reinforcement learning based on safety and efficiency. It can determine the upper limit value of the torque request in real time according to the changes of the self-vehicle and the environment and the driver's torque request during the actual driving process of the vehicle, and perform torque control on the vehicle by comparing the driver's torque request and the upper limit value of the torque request to prevent the occurrence of dangerous situations such as collisions. And the upper limit value of the torque request changes dynamically in real time according to the vehicle surrounding environment information, is not restricted by the driving environment, can be applied to complex traffic environments, and avoids potential safety hazards caused by the driver accidentally stepping on the accelerator pedal, improving the user experience.
[0083] In this embodiment, a vehicle control method based on deep reinforcement learning is also provided, which can be applied to vehicle controllers such as single-chip microcomputers, MCU and other control chips. Figure 2 FIG. is a flowchart of a vehicle control method based on deep reinforcement learning according to an embodiment of the present invention, as Figure 2 shown, the process includes the following steps:
[0084] Step S201, based on the self-vehicle driving data, the driver's speed control intention, and the self-vehicle surrounding environment data, construct the state space of deep reinforcement learning.
[0085] Specifically, in some optional embodiments, the above self-vehicle driving data at least includes: the longitudinal speed, longitudinal acceleration, heading angle, etc. of the self-vehicle, and the self-vehicle surrounding environment data at least includes: the longitudinal speed, longitudinal acceleration, heading angle of each target object in the self-vehicle surrounding environment, and the longitudinal distance between the self-vehicle and the target object. In addition, in practical applications, in order to further improve the accuracy of vehicle acceleration control, the self-vehicle surrounding environment data may further include: road slope, road friction coefficient, environmental temperature, etc. The self-vehicle driving data may also include: the driver's steering intention, etc., and the present invention is not limited thereto.
[0086] In the embodiment of the present invention, by constructing the state space of deep reinforcement learning based on data such as the longitudinal speed, longitudinal acceleration, and heading angle of the self-vehicle, the longitudinal speed, longitudinal acceleration, and heading angle of each target object in the self-vehicle surrounding environment, and the longitudinal distance between the self-vehicle and the target object, etc., to comprehensively consider the influence of relevant factors on the self-vehicle acceleration condition during the driving process of the vehicle, further improve the accuracy of the output torque request upper limit value of the protection model based on deep reinforcement learning, ensure that the torque request upper limit value can prevent the occurrence of dangerous situations such as collisions, and avoid restricting the normal torque demand of the driver in a safe situation for the output torque request upper limit value, and improve the user experience.
[0087] Step S202, construct the action space of deep reinforcement learning based on the torque request upper limit value. For detailed content, refer to the relevant description of step S102 as Figure 1 shown, and details will not be described herein again.
[0088] Step S203, construct the reward function of deep reinforcement learning based on safety and efficiency. For detailed content, refer to the relevant description of step S103 as Figure 1 shown, and details will not be described herein again.
[0089] Step S204, construct and train a protection model based on deep reinforcement learning based on the state space, action space, and reward function.
[0090] Specifically, the construction process of the protection model based on deep reinforcement learning in the above step S204 is specifically to construct an online policy network, a target policy network, an online Q network, and a target Q network respectively by using a neural network based on a state space, an action space, and a reward function; among them, the online policy network is used to calculate the first action to be executed in the current state at the current moment; the target policy network is used to calculate the second action to be executed in the next state transferred to the next moment after executing the first action; the online Q network is used to evaluate the value of executing the first action in the current state at the current moment; the target Q network is used to evaluate the value of executing the second action in the next state at the next moment. Among them, the above neural network can adopt a deep neural network, etc., and the present invention is not limited thereto.
[0091] In the embodiment of the present invention, by constructing a protection model based on deep reinforcement learning, the dependence on rule formulation and parameter calibration is avoided. The required environmental state information is input into the neural network and the action value is output for execution, and the reward value containing indicative information is obtained through continuous interaction with the environment. The neural network is continuously trained and updated through the reward value, so as to realize the learning of the correct control strategy and the adaptability to complex traffic scenarios.
[0092] Furthermore, the training process of the protection model based on deep reinforcement learning in the above step S204 includes:
[0093] Step a1: Input the current state corresponding to the host vehicle at the current moment into the online policy network, output the first action, and calculate the reward obtained at the current moment.
[0094] Step a2: Input the first action into a preset simulation environment and vehicle model to obtain the next state corresponding to the host vehicle at the next moment.
[0095] Step a3: Take the current state, the first action, the reward, and the next state corresponding to each moment as an experience sample to form experience sample data.
[0096] Step a4: Establish an experience pool, store the experience sample data in the experience pool, randomly extract several samples from the experience pool, and update the parameters corresponding to the online Q network and the online policy network.
[0097] Step a5: Soft update the parameters corresponding to the target policy network and the target Q network according to the updated parameters of the online Q network and the online policy network.
[0098] In the embodiments of the present invention, by using the vehicle model and the simulation environment to train the protection model of deep reinforcement learning, the simulation of the scenarios of accidentally stepping on the accelerator and normal driving in the real scenario is realized, so that the trained protection model can be applied to complex and changeable driving scenarios, to improve the accuracy of the torque request upper limit value, further improve the control accuracy of the user's accidental acceleration protection strategy, and enhance the user experience.
[0099] Exemplarily, as Figure 3 shown, a protection model, namely the driver accidental acceleration protection model, based on the Deep Deterministic Policy Gradient (DDPG) algorithm in deep reinforcement learning is constructed and trained, specifically including:
[0100] Use a deep neural network to construct an on-policy network with parameters θ μ to calculate the action t that should be executed at the current moment t under the state s An off-policy network with parameters θ μ′ to calculate the action t that should be executed at the next moment t + 1 after the action a is executed under the state s t+1 This value is used to calculate the target Q value and is not used for actual execution; an on-policy Q network with parameters θ to evaluate the value of the action Q executed at the current moment t under the state s t , that is, the Q value; an off-policy Q network with parameters θ to evaluate the value of the action Q′ executed at the next moment t + 1 under the state s t+1 , that is, the target Q value.
[0101] Meanwhile, training this model requires a simulation environment, a driver model, and a vehicle model; among them, the simulation environment needs to randomly generate roads and environmental vehicles and provide their position, speed, acceleration, and heading angle information, and can also provide information such as road friction coefficient and ramp acceleration; the driver model requires normal driving when far from the front target and randomly stepping on the accelerator when approaching the front target, so as to create scenarios of accidentally stepping on the accelerator and normal driving for training, and the vehicle model needs to calculate the driver's required torque according to the throttle opening and control the vehicle's driving according to the finally arbitrated torque request, and can also update the vehicle's own state according to the torque request.
[0102] Among them, the driver torque request is calculated by the vehicle model based on the accelerator pedal opening input by the driver model. In fact, the torque upper limit value output by the protection model acts on the vehicle model to change the speed of the host vehicle, and the simulation environment receives the change in the speed of the host vehicle, causing other vehicles in the environment to act according to the speed of the host vehicle. The torque limit value actually input to the vehicle model for torque arbitration is output by the online policy network. The selection of the above simulation environment, driver model and vehicle model is prior art and will not be elaborated here.
[0103] The longitudinal speed v of the front target in the state space S o , the longitudinal acceleration a o , the heading angle δ o , the longitudinal distance x o is provided by the simulation environment; the throttle opening T is provided by the driver model; the longitudinal speed v of the host vehicle e , the longitudinal acceleration a e , the heading angle δ e , and the throttle opening change rate The ramp acceleration a s is provided by the vehicle model.
[0104] The specific training process includes: inputting the state s at the current time t t into the online policy network, outputting the action a t , and calculating the reward r obtained at the current time t t , inputting the action a t into the simulation environment and the vehicle model to obtain the state s of the host vehicle at time t + 1 t+1 ; establishing an experience pool and storing the experience sample data at the current time t in the experience pool. Randomly extract N samples from the experience pool and update the parameters θ Q of the online Q network. The update method is shown in formulas (5) and (6) as follows:
[0105]
[0106] L(θ Q ) = (r + γQ′ - Q) 2 (6)
[0107] Among them, α is the learning rate, is the gradient. The negative gradient is used to represent the direction of the fastest decrease of the function to modify the network parameters, and the magnitude represents the increase rate. r is the reward, γ is the discount factor, L(θ Q ) is the temporal difference error of the online Q network, that is, the TD error. Q is the Q value output by the online Q network, which is used to evaluate the value of executing the action a t under the state s at the current time t t , and Q′ is the target Q value output by the target Q network, which is used to evaluate the execution of at The state s at the next moment t+1 to which it is transferred later t+1 The action output by the target policy network to be executed next The value of. Q represents the execution of action a t Before the current state s t The value evaluation result of executing this action below, r+γQ′ represents the execution of action a t After the state s t The value evaluation result of executing this action below, so the purpose of L(θ Q ) is to make the value evaluation result of the Q network tend to be correct. Only in this way will the value evaluation results before and after the execution of the action be close. Otherwise, under the action of the reward r, the two will conflict.
[0108] Specifically, the above learning rate is the amplitude coefficient for updating the parameter θ Q .
[0109] Update the parameter θ of the on-policy network μ as shown in formulas (7) and (8):
[0110]
[0111] J(θ μ ) = Q(8)
[0112] where α is the learning rate, is the gradient, and the purpose of J(θ μ ) is to maximize the value of the action output by the on-policy network, that is, to output the optimal action.
[0113] According to the on-policy network parameter θ μ and the online Q network parameter θ Q , perform a soft update on the target policy network parameter θ μ′ and the target Q network parameter θ Q′ as shown in formulas (9) and (10):
[0114] θ μ′ = τθ μ +(1 - τ)θ μ′ (9)
[0115] θ Q′ = τθ Q +(1 - τ)θ Q′ (10)
[0116] where τ is the soft update parameter. The stability of deep learning can be increased through the soft update method.
[0117] Step S205: Input the current vehicle driving parameters, the current driver's speed control intention, and the vehicle surrounding environment parameters into the trained protection model to obtain the current torque request upper limit value. For detailed content, please refer to the relevant description of step S105 as shown in Figure 1 and no further elaboration will be provided here.
[0118] Step S206: Perform torque control on the vehicle based on the magnitude relationship between the current driver torque request and the current torque request upper limit value.
[0119] Specifically, in some optional embodiments, the above step S206 includes:
[0120] Step S2061: Determine whether the current driver torque request is greater than the current torque request upper limit value.
[0121] Step S2062: When the current driver torque request is greater than the current torque request upper limit value, perform torque control on the vehicle according to the current torque request upper limit value.
[0122] Specifically, since the torque request upper limit value is the maximum torque request allowed for the vehicle to respond, if the current driver torque request obtained after the driver operates the accelerator pedal at this time is greater than this upper limit value, it indicates that there may be a collision risk if the vehicle is torque-controlled according to the driver's operation of the accelerator, and the driver may have an accidental acceleration problem. Therefore, in the embodiments of the present invention, by not directly responding to the driver's operation and performing torque control on the vehicle according to the current torque request upper limit value, driving safety is ensured.
[0123] In the embodiments of the present invention, when the current driver torque request is greater than the current torque request upper limit value, in order to avoid dangerous situations such as collisions, by performing torque control on the vehicle according to the current torque request upper limit value, the driving safety of the vehicle is ensured and the user experience is improved.
[0124] Step S2063: When the current driver torque request is not greater than the current torque request upper limit value, perform torque control on the vehicle according to the current driver torque request.
[0125] In the embodiments of the present invention, when the current driver torque request is not greater than the current torque request upper limit value, in order to better respond to the user's driving intention, by performing torque control on the vehicle according to the previous driver torque request, the user experience is further improved while ensuring the user's driving safety.
[0126] Specifically, in some optional embodiments, the vehicle control method based on deep reinforcement learning provided by the embodiments of the present invention further includes the following steps:
[0127] Step S207: After detecting that the current driver torque request has changed, return to execute step S205.
[0128] Specifically, since the driver's manipulation of the accelerator pedal during vehicle driving is usually constantly changing, the specific judgment conditions for the above-mentioned change in the current driver torque request can be set according to actual needs. For example, if the driver torque requests are different at adjacent moments or after a preset time interval, it is considered that the current driver torque request has changed. Or, it is also possible to determine that the current driver torque request has changed when the change rate of the driver torque request exceeds a preset change rate threshold or the increase amplitude exceeds a preset amplitude threshold, so as to trigger the vehicle torque control function only when there may be a scenario of accidentally stepping on the accelerator, reducing the consumption of computing power caused by frequent calculation of the torque request upper limit value. Only by way of example, the present invention is not limited thereto. In order to provide users with a real-time safe driving experience, by monitoring the driver torque request in real time, once it is detected that the driver torque request has changed, the output of the torque request upper limit value is re-performed, thereby realizing the real-time driver misacceleration protection function.
[0129] In the embodiment of the present invention, by monitoring the driver torque request and recalculating the torque request upper limit value after it changes, real-time protection against accidentally stepping on the accelerator is provided for users during the process of driving a vehicle, further ensuring the driving safety of users and improving the user experience.
[0130] In this embodiment, a vehicle control device based on deep reinforcement learning is also provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the device described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0131] An embodiment of the present invention provides a vehicle control device based on deep reinforcement learning, as Figure 4 shown, the device includes:
[0132] A first processing module 401, configured to construct a state space of deep reinforcement learning based on the self-vehicle driving data, the driver speed control intention, and the self-vehicle surrounding environment data;
[0133] A second processing module 402, configured to construct an action space of deep reinforcement learning based on the torque request upper limit value;
[0134] A third processing module 403, configured to construct a reward function of deep reinforcement learning based on safety and efficiency;
[0135] A fourth processing module 404, configured to construct and train a protection model based on deep reinforcement learning based on the state space, the action space, and the reward function;
[0136] The fifth processing module 405 is configured to input the current self-vehicle driving data, the current driver's speed control intention, and the self-vehicle surrounding environment data into the trained protection model to obtain the current torque request upper limit value;
[0137] The sixth processing module 406 is configured to perform torque control on the self-vehicle based on the magnitude relationship between the current driver's torque request and the current torque request upper limit value.
[0138] In some alternative embodiments, the above-mentioned sixth processing module 406 includes:
[0139] A judgment unit, configured to judge whether the current driver's torque request is greater than the current torque request upper limit value;
[0140] A first processing unit, configured to perform torque control on the self-vehicle according to the current torque request upper limit value when the current driver's torque request is greater than the current torque request upper limit value.
[0141] In some alternative embodiments, the above-mentioned sixth processing module 406 further includes:
[0142] A second processing unit, configured to perform torque control on the self-vehicle according to the current driver's torque request when the current driver's torque request is not greater than the current torque request upper limit value.
[0143] In some alternative embodiments, the vehicle control device based on deep reinforcement learning further includes:
[0144] A seventh processing module, configured to, after detecting that the current driver's torque request changes, return to the step of inputting the current self-vehicle driving data, the current driver's speed control intention, and the self-vehicle surrounding environment data into the trained protection model to obtain the current torque request upper limit value.
[0145] In some alternative embodiments, the self-vehicle driving data at least includes: the longitudinal speed, longitudinal acceleration, and heading angle of the self-vehicle, and the self-vehicle surrounding environment data at least includes: the longitudinal speed, longitudinal acceleration, and heading angle of each target object in the self-vehicle surrounding environment and the longitudinal distance between the self-vehicle and the target object.
[0146] In some alternative embodiments, the above-mentioned fourth processing module 404 is specifically configured to respectively construct an online policy network, a target policy network, an online Q-network, and a target Q-network based on a state space, an action space, and a reward function by using a neural network; wherein, the online policy network is used to calculate a first action to be executed in the current state at the current moment; the target policy network is used to calculate a second action to be executed in the next state transferred to the next moment after executing the first action; the online Q-network is used to evaluate the value of executing the first action in the current state at the current moment; the target Q-network is used to evaluate the value of executing the second action in the next state at the next moment.
[0147] In some alternative embodiments, the training process of the protection model based on deep reinforcement learning of the above-mentioned fourth processing module 404 includes:
[0148] A third processing unit, configured to input the current state corresponding to the host vehicle at the current moment into the online policy network, output a first action, and calculate the reward obtained at the current moment;
[0149] A fourth processing unit, configured to input the first action into a preset simulation environment and a vehicle model to obtain the next state corresponding to the host vehicle at the next moment;
[0150] A fifth processing unit, configured to use the current state, the first action, the reward, and the next state corresponding to each moment as an experience sample to form experience sample data;
[0151] A sixth processing unit, configured to establish an experience pool, store the experience sample data in the experience pool, randomly extract a number of samples from the experience pool, and update the parameters corresponding to the online Q-network and the online policy network;
[0152] A seventh processing unit, configured to perform soft update on the parameters corresponding to the target policy network and the target Q-network according to the updated parameters of the online Q-network and the online policy network.
[0153] The vehicle control device based on deep reinforcement learning in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0154] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding method embodiments above, and will not be repeated here.
[0155] An embodiment of the present invention further provides a vehicle, such as Figure 5As shown, the vehicle includes: one or more processors 10, a memory 20, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories if needed. Similarly, multiple computer devices can be connected, and each device provides part of the necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Figure 5 In Figure 5 , a processor 10 is taken as an example.
[0156] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device can be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.
[0157] Among them, the memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the methods shown in the above embodiments.
[0158] The memory 20 can include a storage program area and a storage data area. Among them, the storage program area can store an operating system and application programs required for at least one function; the storage data area can store data created according to the use of the computer device presented by a kind of landing page of a small program, etc. In addition, the memory 20 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 can optionally include a memory remotely set relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0159] The memory 20 can include a volatile memory, such as a random access memory; the memory can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 can also include a combination of the above types of memories.
[0160] The vehicle further includes a communication interface 30 for the vehicle to communicate with other devices or a communication network.
[0161] Embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processes on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.
[0162] Although the embodiments of the present invention are described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A vehicle control method based on deep reinforcement learning, characterized in that The method includes: Construct a state space for deep reinforcement learning based on the ego-vehicle driving data, the driver's speed control intention, and the ego-vehicle surrounding environment data; Construct an action space for deep reinforcement learning based on the torque request upper limit value; Construct a reward function for deep reinforcement learning based on safety and efficiency; Construct and train a protection model based on deep reinforcement learning based on the state space, action space, and reward function; Input the current ego-vehicle driving data, the current driver's speed control intention, and the ego-vehicle surrounding environment data into the trained protection model to obtain the current torque request upper limit value; Perform torque control on the ego-vehicle based on the magnitude relationship between the current driver torque request and the current torque request upper limit value.
2. The method according to claim 1, wherein The performing torque control on the ego-vehicle based on the magnitude relationship between the current driver torque request and the current torque request upper limit value includes: Judge whether the current driver torque request is greater than the current torque request upper limit value; When the current driver torque request is greater than the current torque request upper limit value, perform torque control on the ego-vehicle according to the current torque request upper limit value.
3. The method according to claim 2, wherein The method further includes: When the current driver torque request is not greater than the current torque request upper limit value, perform torque control on the ego-vehicle according to the current driver torque request.
4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: After detecting a change in the current driver torque request, return to the step of inputting the current ego-vehicle driving data, the current driver's speed control intention, and the ego-vehicle surrounding environment data into the trained protection model to obtain the current torque request upper limit value.
5. The method according to claim 1, wherein The ego-vehicle driving data at least includes: the longitudinal speed, longitudinal acceleration, and heading angle of the ego-vehicle, and the ego-vehicle surrounding environment data at least includes: the longitudinal speed, longitudinal acceleration, and heading angle of each target object in the ego-vehicle surrounding environment, and the longitudinal distance between the ego-vehicle and the target object.
6. The method according to claim 1, wherein Constructing a protection model based on deep reinforcement learning based on the state space, action space, and reward function includes: Respectively construct an online policy network, a target policy network, an online Q network, and a target Q network based on the state space, action space, and reward function using a neural network; Among them, the online policy network is used to calculate the first action to be executed in the current state at the current moment; the target policy network is used to calculate the second action to be executed in the next state transferred after executing the first action; the online Q network is used to evaluate the value of executing the first action in the current state at the current moment; the target Q network is used to evaluate the value of executing the second action in the next state at the next moment.
7. The method according to claim 6, wherein The training process of the protection model based on deep reinforcement learning includes: Input the current state corresponding to the ego-vehicle at the current moment into the online policy network, output the first action, and calculate the reward obtained at the current moment; Input the first action into a preset simulation environment and vehicle model to obtain the next state corresponding to the ego-vehicle at the next moment; Take the current state, first action, reward, and next state corresponding to each moment as an experience sample to form experience sample data; An experience pool is established, and the experience sample data is stored in the experience pool. A number of samples are randomly drawn from the experience pool to update the parameters corresponding to the online Q-network and the online policy network; Based on the updated parameters of the online Q-network and the online policy network, the parameters corresponding to the target policy network and the target Q-network are softly updated.
8. A vehicle control device based on deep reinforcement learning, characterized in that, The device includes: A first processing module, configured to construct a state space of deep reinforcement learning based on the self-vehicle driving data, the driver's speed control intention, and the self-vehicle surrounding environment data; A second processing module, configured to construct an action space of deep reinforcement learning based on the torque request upper limit value; A third processing module, configured to construct a reward function of deep reinforcement learning based on safety and efficiency; A fourth processing module, configured to construct and train a protection model based on deep reinforcement learning based on the state space, the action space, and the reward function; A fifth processing module, configured to input the current self-vehicle driving data, the current driver's speed control intention, and the self-vehicle surrounding environment data into the trained protection model to obtain the current torque request upper limit value; A sixth processing module, configured to perform torque control on the self-vehicle based on the magnitude relationship between the current driver torque request and the current torque request upper limit value.
9. A vehicle, characterized in that, The vehicle includes: A memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.
Citation Information
Cited By
Reinforcement learning-based low-attachment road surface four-wheel drive torque distribution method
CN120863365A