An automatic driving car-following speed control method, a computer device and a storage medium

CN116923401BActive Publication Date: 2026-09-08CENT SOUTH UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210320191.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-29
Publication Date
2026-09-08
Estimated Expiration
2042-03-29

AI Technical Summary

Technical Problem

然而,基于运动学的数值跟驰模型很难考虑驾驶员的驾驶风格差异性和驾驶场景的复杂性

Benefits of technology

[0035] The autonomous driving car-following speed control method, computer device, and storage medium provided in this invention include: acquiring current driving information data of a target vehicle, wherein the driving information data includes location information and speed information, and the target vehicle is an autonomous driving vehicle; determining the driving information data of the vehicle in front and behind the target vehicle based on the location information; inputting the current driving information data of the target vehicle, the driving information data of the vehicle in front, and the driving information data of the vehicle behind into a trained deep reinforcement learning model, and outputting the acceleration of the target vehicle; thus, for mixed traffic flow situations of human driving and autonomous driving vehicles in the future vehicle-to-everything (V2X) environment, the autonomous driving vehicle can more comprehensively receive the driving information of the vehicle behind and fully perceive the driving states of the vehicle in front and behind in the mixed traffic flow, thereby improving the stability of road traffic flow and reducing the probability of traffic accidents; the method of this invention is simple in design and easy to calculate; the car-following model of the autonomous driving vehicle based on deep reinforcement learning and considering the vehicle behind can overcome the defects of existing car-following models, improve the car-following safety of autonomous driving vehicles, and effectively improve road traffic safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116923401B_ABST
    Figure CN116923401B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose an automatic driving car-following speed control method, a computer device and a storage medium. Driving information data of a target vehicle is obtained, wherein the driving information data comprises position information and speed information, and the target vehicle is an automatic driving vehicle. Driving information data of front and rear vehicles of the target vehicle is determined based on the position information, so as to construct an automatic driving car-following environment. A state space is determined based on a reinforcement learning framework, the state space being relative distance and relative speed between the automatic driving vehicle and the front and rear vehicles, and an action being acceleration of the automatic driving vehicle. A reward function in the reinforcement learning algorithm is designed to guide the automatic driving vehicle to avoid collision with the front and rear vehicles. The driving information data of the target vehicle, the driving information data of the front vehicle and the driving information data of the rear vehicle are input into a trained deep reinforcement learning model, and acceleration of the target vehicle is output, so as to control car-following speed of the automatic driving vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous driving, and more particularly to an autonomous driving following speed control method, computer device, and storage medium. Background Technology

[0002] As a primary driving behavior, car-following significantly impacts driving safety and traffic flow stability. Car-following theory is one of the foundations of traffic flow theory, and much research in the field of transportation is built upon it. Many traffic simulation tools (such as VISSIM, AIMSUM, and SUMO) and autonomous driving control algorithms have also developed car-following modules. Therefore, car-following models are of great significance for improving safety and efficiency and realizing autonomous driving.

[0003] Considering the two-vehicle mode of following the preceding vehicle is a major research topic in classic vehicle car-following models. Autonomous vehicles only interact with the vehicle in front and react based on its state. Due to the driver's visual limitations, the driver pays more attention to the vehicle in front during car-following. In real-world driving environments, especially when the autonomous vehicle is braking, the following vehicle also needs to be observed. However, in a vehicle-to-everything (V2X) traffic system, vehicles can transmit driving information to each other, allowing the following vehicle to transmit more information to the preceding vehicle, enabling the autonomous vehicle to accurately perceive the state of the following vehicle. Furthermore, before fully realizing autonomous driving, there will be a transition period where mixed traffic flows of autonomous vehicles and human-driven vehicles will emerge. Autonomous vehicles need to maintain safe interaction with both the preceding and following vehicles, considering the states of both when making decisions. Considering the state of the following vehicle also helps the autonomous vehicle control its speed, avoiding rear-end collisions while ensuring the safety of the preceding vehicle. Therefore, considering the following vehicle's car-following mode helps improve the safety and efficiency of road traffic flow.

[0004] Traditional car-following models include kinematic numerical car-following models and machine learning-based car-following models. However, kinematic numerical car-following models struggle to account for differences in driver styles and the complexity of driving scenarios. Furthermore, machine learning, which mimics human driving behavior, has limited ability to make judgments about the specific driving environment. Summary of the Invention

[0005] In view of this, the present invention provides an autonomous driving following speed control method, computer device and storage medium, so that autonomous driving vehicles can better adapt to the complexity of following environment and maintain safe driving between vehicles in front and behind, thereby improving road driving safety.

[0006] To achieve the above objectives, the technical solution of the present invention is implemented as follows:

[0007] In a first aspect, the present invention provides an autonomous driving following speed control method, the method comprising:

[0008] Acquire the current driving information data of the target vehicle, wherein the driving information data includes location information and speed information, and the target vehicle is an autonomous vehicle;

[0009] Based on the location information, determine the driving information data of the vehicles in front of and behind the target vehicle;

[0010] Based on the current driving information data of the target vehicle, the driving information data of the preceding vehicle, and the driving information data of the following vehicle, the trained deep reinforcement learning model is input, and the acceleration of the target vehicle is output.

[0011] The acquisition of the target vehicle's current driving information data includes:

[0012] Vehicle driving information data, including location and speed information, is collected by cameras and roadside detectors on the target area, and the driving information data is transmitted to the target vehicle through a roadside communication unit.

[0013] The step of determining the driving information data of the vehicles preceding and following the target vehicle based on the location information includes:

[0014] The lane position data of the target vehicle is determined based on the location information of the target vehicle;

[0015] The autonomous driving vehicle following the lane position data is determined, and the starting point of the following process is determined.

[0016] Acquire driving information data of the vehicles in front and behind in the current lane at the starting moment to construct a car-following simulation environment for autonomous vehicles.

[0017] The construction of the car-following simulation environment for autonomous vehicles includes:

[0018] In the deep reinforcement learning model, the agent is defined as an autonomous vehicle, and the state space is the relative distance and relative speed [Δv] between the autonomous vehicle and the vehicles in front and behind. ls Δv fs Δy ls Δy fs The action space is the acceleration of the autonomous vehicle, and the state is updated according to kinematic principles, with the formula: v s (t+1)=v s (t)+a(t)*ΔT

[0019]

[0020] Δvls (t+1)=v l (t+1)-v s (t+1); Δv fs (t+1)=v s (t+1)-v f (t+1)

[0021] Δy ls (t+1)=y l (t+1)-y s (t+1); Δy fs (t+1)=y s (t+1)-y f (t+1)

[0022] Where v represents speed, y represents longitudinal position, s represents autonomous vehicle, l represents the preceding vehicle, f represents the following vehicle, and ΔT is the measurement time step.

[0023] The deep reinforcement learning model's network structure includes a value function network, a target value function network, a Q1 network, a Q2 network, and a policy network.

[0024] This also includes:

[0025] The design of the reward function guides the autonomous vehicle to safely follow between the preceding and following vehicles. The reward function is set by decreasing the safety level and the selected target action. The safety level is determined based on the speed relationship, relative distance and minimum braking distance between the autonomous vehicle and the preceding and following vehicles in the three-vehicle mode. The reward is then set for the target vehicle based on the safety level and the desired acceleration and deceleration.

[0026] Before inputting the driving information data based on the current driving information of the target vehicle, the driving information data of the preceding vehicle, and the driving information data of the following vehicle into the trained deep reinforcement learning model, the method further includes:

[0027] The target vehicle and the driving information data of the preceding and following vehicles are input into the initial deep reinforcement learning model, which outputs acceleration and updates the target state.

[0028] The deep reinforcement learning model is iterated individually and alternately based on the target state until the set loss function satisfies the convergence condition, thereby obtaining the trained deep reinforcement learning model.

[0029] The step of inputting the current driving information data of the target vehicle, the driving information data of the preceding vehicle, and the driving information data of the following vehicle into the trained deep reinforcement learning model, and outputting the acceleration of the target vehicle, includes:

[0030] The current target state is determined based on the target vehicle and the driving information data of the vehicles in front and behind.

[0031] Based on the target state, the trained deep reinforcement learning model outputs the acceleration of the target vehicle.

[0032] In a second aspect, the present invention provides a computer device, comprising: a processor and a memory for storing a computer program capable of running on the processor;

[0033] When the processor runs the computer program, it implements the above-described autonomous driving following speed control method.

[0034] Thirdly, the present invention provides a computing storage medium storing a computer program, which is executed by a processor to implement any of the above-described autonomous driving following speed control methods.

[0035] The autonomous driving car-following speed control method, computer device, and storage medium provided in this invention include: acquiring current driving information data of a target vehicle, wherein the driving information data includes location information and speed information, and the target vehicle is an autonomous driving vehicle; determining the driving information data of the vehicle in front and behind the target vehicle based on the location information; inputting the current driving information data of the target vehicle, the driving information data of the vehicle in front, and the driving information data of the vehicle behind into a trained deep reinforcement learning model, and outputting the acceleration of the target vehicle; thus, for mixed traffic flow situations of human driving and autonomous driving vehicles in the future vehicle-to-everything (V2X) environment, the autonomous driving vehicle can more comprehensively receive the driving information of the vehicle behind and fully perceive the driving states of the vehicle in front and behind in the mixed traffic flow, thereby improving the stability of road traffic flow and reducing the probability of traffic accidents; the method of this invention is simple in design and easy to calculate; the car-following model of the autonomous driving vehicle based on deep reinforcement learning and considering the vehicle behind can overcome the defects of existing car-following models, improve the car-following safety of autonomous driving vehicles, and effectively improve road traffic safety. Attached Figure Description

[0036] Figure 1 A flowchart illustrating an autonomous driving following speed control method provided in an embodiment of the present invention;

[0037] Figure 2 A flowchart illustrating another autonomous driving following speed control method provided in an embodiment of the present invention;

[0038] Figure 3 This is a schematic diagram of the network structure of the deep reinforcement learning model provided in an embodiment of the present invention;

[0039] Figure 4 This is a schematic diagram of the structure of an autonomous driving following speed control device provided in an embodiment of the present invention;

[0040] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0041] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terminology and / or as used herein includes any and all combinations of one or more of the associated listed items.

[0042] With the continuous iteration and updates of artificial intelligence, the application of reinforcement learning methods to study control strategies for autonomous driving has become a new hot topic. Compared with traditional supervised and unsupervised machine learning, the biggest characteristic of reinforcement learning is that the agent learns by interacting with the environment and aims to maximize rewards. The driving process of autonomous vehicles is similar to the mechanism of reinforcement learning. During driving, the vehicle needs to interact with its surrounding environment, obtain feedback from the environment, make decisions, and execute commands. Therefore, based on a large amount of real-world driving data, applying deep reinforcement learning methods and considering the following vehicle's car-following approach can provide valuable insights for the control of autonomous vehicles.

[0043] Please see Figure 1 This invention provides an autonomous driving following speed control method, applicable to any vehicle terminal equipped with autonomous driving capabilities. The method can be executed by a computer device, which may be a terminal or a server. The terminal can specifically be a desktop computer, laptop computer, smartphone, personal digital assistant, or tablet computer; the server can be a single server device or a server cluster. The autonomous driving following speed control method includes the following steps:

[0044] Step 101: Obtain the current driving information data of the target vehicle, wherein the driving information data includes location information and speed information, and the target vehicle is an autonomous vehicle;

[0045] Step 102: Determine the driving information data of the vehicle in front of and behind the target vehicle based on the location information;

[0046] Step 103: Input the current driving information data of the target vehicle, the driving information data of the preceding vehicle, and the driving information data of the following vehicle into the trained deep reinforcement learning model, and output the acceleration of the target vehicle.

[0047] Through the above embodiments of the present invention, for mixed traffic flow situations involving human driving and autonomous vehicles in the future Internet of Vehicles environment, autonomous vehicles can more comprehensively receive driving information from following vehicles and fully perceive the driving status of preceding and following vehicles in the mixed traffic flow, thereby improving the stability of road traffic flow and reducing the probability of traffic accidents. The method of the present invention is simple in design and easy to calculate. The car-following model of autonomous vehicles based on deep reinforcement learning and considering following vehicles can overcome the defects of existing car-following models, improve the car-following safety of autonomous vehicles, and effectively improve road traffic safety.

[0048] In one embodiment, obtaining the current driving information data of the target vehicle includes:

[0049] Vehicle driving information data, including location and speed information, is collected by cameras and roadside detectors on the target area, and the driving information data is transmitted to the target vehicle through a roadside communication unit.

[0050] In one embodiment, determining the driving information data of the vehicles preceding and following the target vehicle based on the location information includes:

[0051] The lane position data of the target vehicle is determined based on the location information of the target vehicle;

[0052] The autonomous driving vehicle following the lane position data is determined, and the starting point of the following process is determined.

[0053] Acquire driving information data of the vehicles in front and behind in the current lane at the starting moment to construct a car-following simulation environment for autonomous vehicles.

[0054] In one embodiment, the construction of the car-following simulation environment for autonomous vehicles includes:

[0055] In the deep reinforcement learning model, the agent is defined as an autonomous vehicle, and the state space is the relative distance and relative speed [Δv] between the autonomous vehicle and the vehicles in front and behind. ls Δv fs Δy ls Δy fs The action space is the acceleration of the autonomous vehicle, and the state is updated according to kinematic principles, with the formula: v s (t+1)=v s (t)+a(t)*ΔT

[0056]

[0057] Δv ls (t+1)=v l (t+1)-v s (t+1); Δv fs (t+1)=v s (t+1)-v f (t+1)

[0058] Δy ls (t+1)=y l (t+1)-y s (t+1); Δy fs (t+1)=y s (t+1)-y f (t+1)

[0059] Where v represents speed, y represents longitudinal position, s represents autonomous vehicle, l represents the preceding vehicle, f represents the following vehicle, and ΔT is the measurement time step.

[0060] In one embodiment, the network structure of the deep reinforcement learning model includes a value function network, a target value function network, a Q1 network, a Q2 network, and a policy network.

[0061] In one embodiment, it further includes:

[0062] The design of the reward function guides the autonomous vehicle to safely follow between the preceding and following vehicles. The reward function is set by decreasing the safety level and the selected target action. The safety level is determined based on the speed relationship, relative distance and minimum braking distance between the autonomous vehicle and the preceding and following vehicles in the three-vehicle mode. The reward is then set for the target vehicle based on the safety level and the desired acceleration and deceleration.

[0063] In one embodiment, before inputting the driving information data based on the current driving information of the target vehicle, the driving information data of the preceding vehicle, and the driving information data of the following vehicle into the trained deep reinforcement learning model, the method further includes:

[0064] The target vehicle and the driving information data of the preceding and following vehicles are input into the initial deep reinforcement learning model, which outputs acceleration and updates the target state.

[0065] The deep reinforcement learning model is iterated individually and alternately based on the target state until the set loss function satisfies the convergence condition, thereby obtaining the trained deep reinforcement learning model.

[0066] In one embodiment, the step of inputting the current driving information data of the target vehicle, the driving information data of the preceding vehicle, and the driving information data of the following vehicle into a trained deep reinforcement learning model, and outputting the acceleration of the target vehicle, includes:

[0067] The current target state is determined based on the target vehicle and the driving information data of the vehicles in front and behind.

[0068] Based on the target state, the trained deep reinforcement learning model outputs the acceleration of the target vehicle.

[0069] For example, please refer to Figure 2 This will be illustrated through a specific embodiment.

[0070] Step 201: Use cameras and road detectors on the road in the target area to collect vehicle driving information data, including location and speed information, and transmit the driving information data to the autonomous vehicle through the roadside communication unit;

[0071] Step 202: Based on the vehicle's location information and lane position data, determine the autonomous vehicle following the car, and determine the starting point of the following process. Obtain the driving information data of the vehicle in front and behind in the current lane at the starting point, thereby constructing a car-following simulation environment for the autonomous vehicle.

[0072] Step 203: Based on the above car-following simulation environment, determine that the agent in the deep reinforcement learning SAC algorithm is an autonomous vehicle, and its state is the relative distance and relative speed [Δv] between the autonomous vehicle and the vehicles in front and behind. ls Δv fs Δy ls Δy fs The motion is the acceleration of the autonomous vehicle. v represents velocity, y represents longitudinal position, s represents the autonomous vehicle, l represents the preceding vehicle, f represents the following vehicle, and ΔT is the measurement time step. Therefore, the state can be updated according to kinematic principles, as shown in the formula:

[0073] v s (t+1)=v s (t)+a(t)*ΔT

[0074]

[0075] Δv ls (t+1)=v l (t+1)-v s (t+1); Δv fs (t+1)=v s (t+1)-v f (t+1)

[0076] Δy ls (t+1)=y l (t+1)-y s (t+1); Δy fs (t+1)=y s (t+1)-y f (t+1)

[0077] Step 204: Based on the above states and corresponding actions, design a reward function to guide the autonomous vehicle to select acceleration to avoid collisions with the vehicles in front and behind. The reward is set by decreasing the safety level and the selected action. The safety level is determined based on the speed relationship, relative distance, and minimum braking distance between the autonomous vehicle and the vehicles in front and behind in the three-vehicle mode. The reward is then set based on the safety level and the desired acceleration / deceleration.

[0078] Step 205: Determine the network structure and associated hyperparameters of the two value networks, two Q networks, and one policy network in the deep reinforcement learning SAC algorithm. For the value function network V... ψ (s), where the network parameters are ψ, the input is the state, and the output is a value function; for a Q-network Q θ (s, a), network parameters are θ, inputs are states and actions, outputs are state and action values; for the policy network π φ (a|s), the network parameter is φ, the input is the state, and the output is the action and the corresponding entropy value. Each network has two hidden layers, each with 256 neurons. The network structure of the SAC algorithm is as follows: Figure 3 The corresponding hyperparameters were determined based on existing research.

[0079] Step 206: Based on the car-following simulation environment in Step 202, input the initialized scene state into the SAC algorithm, output acceleration, thereby updating the state and obtaining the reward at the current moment, and recording (s) t a t r t s t+1 The data is stored in the experience replay buffer and sampled from the experience replay buffer during the algorithm learning process based on the minimum batch size.

[0080] Step 207: Update the network parameters in the SAC algorithm until the car-following behavior in the current scene ends. The value function network calculates the gradient. Update the network; the objective value function network is based on... Update the network; the Q network calculates the gradient. Update the network; the policy network calculates the gradient.

[0081] Step 208: The car-following behavior based on the above scenario ends. The car-following scenario is re-initialized, and iterative updates are performed continuously until the SAC algorithm converges.

[0082] This method utilizes traffic flow detection equipment to acquire vehicle driving information in the target area and constructs a car-following environment under a three-vehicle mode to simulate the mixed traffic flow of future autonomous vehicles and human-driven vehicles. It employs the Soft Actor-Critic (SAC) algorithm from deep reinforcement learning as the speed control strategy for autonomous vehicles. The algorithm's state is defined as the relative distance and relative speed between the autonomous vehicle and the vehicles in front and behind, and the action is acceleration. A reward function based on state and corresponding actions is designed to guide the autonomous vehicle in controlling acceleration between the vehicles in front and behind to ensure collision avoidance. Traffic flow data collected in the target area is used as the external input to the model and as the training set data. For the mixed traffic flow of human drivers and autonomous vehicles in the future connected vehicle environment, autonomous vehicles can more comprehensively receive driving information from following vehicles and fully perceive the driving states of vehicles in front and behind in the mixed traffic flow, thereby improving road traffic flow stability and reducing the probability of traffic accidents. The method of this invention is simple in design and easy to calculate; based on deep reinforcement learning and considering the following vehicles, the car-following model of autonomous vehicles can overcome the defects of existing car-following models, improve the car-following safety of autonomous vehicles, and effectively improve road traffic safety.

[0083] This invention also provides an autonomous driving following speed control device, such as... Figure 4 As shown, the device includes:

[0084] The acquisition module 51 is used to acquire the current driving information data of the target vehicle, wherein the driving information data includes location information and speed information, and the target vehicle is an autonomous vehicle;

[0085] The determining module 52 is used to determine the driving information data of the vehicle in front of and behind the target vehicle based on the location information;

[0086] The output module 53 is used to output the acceleration of the target vehicle based on the current driving information data of the target vehicle, the driving information data of the preceding vehicle, and the driving information data of the following vehicle, which are input into the trained deep reinforcement learning model.

[0087] In an optional embodiment, the acquisition module 51 is further configured to:

[0088] Vehicle driving information data, including location and speed information, is collected by cameras and roadside detectors on the target area, and the driving information data is transmitted to the target vehicle through a roadside communication unit.

[0089] In an optional embodiment, the determining module 52 is further configured to:

[0090] The lane position data of the target vehicle is determined based on the location information of the target vehicle;

[0091] The autonomous driving vehicle following the lane position data is determined, and the starting point of the following process is determined.

[0092] Acquire driving information data of the vehicles in front and behind in the current lane at the starting moment to construct a car-following simulation environment for autonomous vehicles.

[0093] In an optional embodiment, the determining module 52 is further configured to:

[0094] In the deep reinforcement learning model, the agent is defined as an autonomous vehicle, and the state space is the relative distance and relative speed [Δv] between the autonomous vehicle and the vehicles in front and behind. ls Δv fs Δy ls Δy fs The action space is the acceleration of the autonomous vehicle, and the state is updated according to kinematic principles, with the formula: v s (t+1)=v s (t)+a(t)*ΔT

[0095]

[0096] Δv ls (t+1)=v l (t+1)-v s (t+1); Δv fs (t+1)=v s (t+1)-v f (t+1)

[0097] Δy ls (t+1)=y l (t+1)-y s (t+1); Δy fs (t+1)=y s (t+1)-y f (t+1)

[0098] Where v represents speed, y represents longitudinal position, s represents autonomous vehicle, l represents the preceding vehicle, f represents the following vehicle, and ΔT is the measurement time step.

[0099] In an optional embodiment, the device further includes a control module for:

[0100] The design of the reward function guides the autonomous vehicle to safely follow between the preceding and following vehicles. The reward function is set by decreasing the safety level and the selected target action. The safety level is determined based on the speed relationship, relative distance and minimum braking distance between the autonomous vehicle and the preceding and following vehicles in the three-vehicle mode. The reward is then set for the target vehicle based on the safety level and the desired acceleration and deceleration.

[0101] In an optional embodiment, the apparatus further includes a training module for:

[0102] The target vehicle and the driving information data of the preceding and following vehicles are input into the initial deep reinforcement learning model, which outputs acceleration and updates the target state.

[0103] The deep reinforcement learning model is iterated individually and alternately based on the target state until the set loss function satisfies the convergence condition, thereby obtaining the trained deep reinforcement learning model.

[0104] In an optional embodiment, the output module 53 is further configured to:

[0105] The current target state is determined based on the target vehicle and the driving information data of the vehicles in front and behind.

[0106] Based on the target state, the trained deep reinforcement learning model outputs the acceleration of the target vehicle.

[0107] It should be noted that the autonomous driving following speed control device provided in the above embodiments is only illustrated by the division of the above-described program modules when implementing the autonomous driving following speed control method. In practical applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the ride-hailing driver safety evaluation device can be divided into different program modules to complete all or part of the processing described above. In addition, the autonomous driving following speed control device provided in the above embodiments and the corresponding autonomous driving following speed control embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0108] This invention provides a computer device, such as... Figure 5 As shown, the computer device includes: a processor 110 and a memory 111 for storing computer programs capable of running on the processor 110; wherein, Figure 5The processor 110 shown in the diagram does not refer to a single processor 110, but rather to the positional relationship of the processor 110 relative to other devices. In practical applications, there can be one or more processors 110; similarly, Figure 5 The memory 111 shown in the diagram has the same meaning, that is, it is only used to refer to the positional relationship of memory 111 relative to other devices. In practical applications, there can be one or more memories 111.

[0109] When the processor 110 is used to run the computer program, it performs the following steps:

[0110] Acquire the current driving information data of the target vehicle, wherein the driving information data includes location information and speed information, and the target vehicle is an autonomous vehicle;

[0111] Based on the location information, determine the driving information data of the vehicles in front of and behind the target vehicle;

[0112] Based on the current driving information data of the target vehicle, the driving information data of the preceding vehicle, and the driving information data of the following vehicle, the trained deep reinforcement learning model is input, and the acceleration of the target vehicle is output.

[0113] In an optional embodiment, the processor 110 is further configured to perform the following steps when running the computer program:

[0114] Vehicle driving information data, including location and speed information, is collected by cameras and roadside detectors on the target area, and the driving information data is transmitted to the target vehicle through a roadside communication unit.

[0115] In an optional embodiment, the processor 110 is further configured to perform the following steps when running the computer program:

[0116] The lane position data of the target vehicle is determined based on the location information of the target vehicle;

[0117] The autonomous driving vehicle following the lane position data is determined, and the starting point of the following process is determined.

[0118] Acquire driving information data of the vehicles in front and behind in the current lane at the starting moment to construct a car-following simulation environment for autonomous vehicles.

[0119] In an optional embodiment, the processor 110 is further configured to perform the following steps when running the computer program:

[0120] In the deep reinforcement learning model, the agent is defined as an autonomous vehicle, and the state space is the relative distance and relative speed [Δv] between the autonomous vehicle and the vehicles in front and behind.ls Δv fs Δy ls Δy fs The action space is the acceleration of the autonomous vehicle, and the state is updated according to kinematic principles, with the formula: v s (t+1)=v s (t)+a(t)*ΔT

[0121]

[0122] Δv ls (t+1)=v l (t+1)-v s (t+1); Δv fs (t+1)=v s (t+1)-v f (t+1)

[0123] Δy ls (t+1)=y l (t+1)-y s (t+1); Δy fs (t+1)=y s (t+1)-y f (t+1)

[0124] Where v represents speed, y represents longitudinal position, s represents autonomous vehicle, l represents the preceding vehicle, f represents the following vehicle, and ΔT is the measurement time step.

[0125] In an optional embodiment, the processor 110 is further configured to perform the following steps when running the computer program:

[0126] The design of the reward function guides the autonomous vehicle to safely follow between the preceding and following vehicles. The reward function is set by decreasing the safety level and the selected target action. The safety level is determined based on the speed relationship, relative distance and minimum braking distance between the autonomous vehicle and the preceding and following vehicles in the three-vehicle mode. The reward is then set for the target vehicle based on the safety level and the desired acceleration and deceleration.

[0127] In an optional embodiment, the processor 110 is further configured to perform the following steps when running the computer program:

[0128] The target vehicle and the driving information data of the preceding and following vehicles are input into the initial deep reinforcement learning model, which outputs acceleration and updates the target state.

[0129] The deep reinforcement learning model is iterated individually and alternately based on the target state until the set loss function satisfies the convergence condition, thereby obtaining the trained deep reinforcement learning model.

[0130] In an optional embodiment, the processor 110 is further configured to perform the following steps when running the computer program:

[0131] The current target state is determined based on the target vehicle and the driving information data of the vehicles in front and behind.

[0132] Based on the target state, the trained deep reinforcement learning model outputs the acceleration of the target vehicle.

[0133] The computer device also includes at least one network interface 112. The various components of the device are coupled together via a bus system 113. It is understood that the bus system 113 is used to implement communication between these components. In addition to a data bus, the bus system 113 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 3 The general designated all buses as Bus System 113.

[0134] The memory 111 can be volatile memory or non-volatile memory, or both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); the magnetic surface memory can be disk storage or magnetic tape storage. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memory 111 described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.

[0135] The memory 111 in this embodiment of the invention is used to store various types of data to support the operation of the device. Examples of such data include: any computer programs used to operate on the device, such as operating systems and applications; contact data; phonebook data; messages; pictures; videos, etc. The operating system includes various system programs, such as the framework layer, core library layer, driver layer, etc., used to implement various basic services and handle hardware-based tasks. Applications may include various applications, such as media players, browsers, etc., used to implement various application services. Here, the program implementing the method of this embodiment of the invention may be included in the application.

[0136] This embodiment also includes a computer storage medium storing a computer program. The computer storage medium can be a magnetic random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc.; it can also be various devices including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc. When the computer program stored in the computer storage medium is run by a processor, it implements the above-described vehicle recognition method. For the specific steps implemented when the computer program is executed by the processor, please refer to [link to relevant documentation]. Figure 1 The description of the illustrated embodiments will not be repeated here.

[0137] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0138] In this document, the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, which includes not only the elements listed but also other elements not expressly listed.

[0139] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for controlling the following speed in autonomous driving, characterized in that, The method includes: Acquire the current driving information data of the target vehicle, wherein the driving information data includes location information and speed information, and the target vehicle is an autonomous vehicle; Based on the location information, determine the driving information data of the vehicles in front of and behind the target vehicle; Based on the current driving information data of the target vehicle, the driving information data of the preceding vehicle, and the driving information data of the following vehicle, a pre-trained deep reinforcement learning model is input, and the acceleration of the target vehicle is output; obtaining the current driving information data of the target vehicle includes: Vehicle driving information data, including location and speed information, is collected by cameras and roadside detectors on the road sections in the target area, and the driving information data is transmitted to the target vehicle through a roadside communication unit. The step of determining the driving information data of the vehicles preceding and following the target vehicle based on the location information includes: The lane position data of the target vehicle is determined based on the location information of the target vehicle; The autonomous driving vehicle following the lane position data is determined, and the starting point of the following process is determined. Acquire driving information data of the vehicles in front and behind in the current lane at the starting time, and construct a car-following simulation environment for autonomous vehicles; the construction of the car-following simulation environment for autonomous vehicles includes: In the deep reinforcement learning model, the agent is defined as an autonomous vehicle, and the state space consists of the relative distance and relative speed between the autonomous vehicle and the vehicles in front and behind. The action space is the acceleration of the autonomous vehicle, and the state is updated according to kinematic principles, using the following formula: in, Represents speed, Represents vertical position. Represents autonomous vehicles. Representing the preceding vehicle, Representing the following car, To measure the time step.

2. The autonomous driving following speed control method according to claim 1, characterized in that, The network structure of the deep reinforcement learning model includes a value function network, an objective value function network, a Q1 network, a Q2 network, and a policy network.

3. The autonomous driving following speed control method according to claim 1, characterized in that, Also includes: The design of the reward function guides the autonomous vehicle to safely follow between the preceding and following vehicles. The reward function is configured by decreasing the safety level and the selected target action. The safety level is determined based on the speed relationship, relative distance and minimum braking distance between the autonomous vehicle and the preceding and following vehicles in the three-vehicle mode. The reward is then set for the target vehicle based on the safety level and the desired acceleration and deceleration.

4. The autonomous driving following speed control method according to claim 3, characterized in that, Before inputting the driving information data based on the current driving information of the target vehicle, the driving information data of the preceding vehicle, and the driving information data of the following vehicle into the trained deep reinforcement learning model, the method further includes: The target vehicle and the driving information data of the preceding and following vehicles are input into the initial deep reinforcement learning model, which outputs acceleration and updates the target state. The deep reinforcement learning model is iterated individually and alternately based on the target state until the set loss function satisfies the convergence condition, thereby obtaining the trained deep reinforcement learning model.

5. The autonomous driving following speed control method according to claim 4, characterized in that, The deep reinforcement learning model, which is trained based on the current driving information data of the target vehicle, the driving information data of the preceding vehicle, and the driving information data of the following vehicle, outputs the acceleration of the target vehicle, including: The current target state is determined based on the target vehicle and the driving information data of the vehicles in front and behind. Based on the target state, the trained deep reinforcement learning model is input, and the acceleration of the target vehicle is output.

6. A computer device, characterized in that, include: Processor and memory used to store computer programs that can run on the processor; When the processor runs the computer program, it implements the autonomous driving following speed control method according to any one of claims 1 to 5.

7. A computer storage medium, characterized in that, The computer storage medium stores a computer program, which is executed by a processor to implement the autonomous driving following speed control method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Automatic driving vehicle lane changing decision control method based on hierarchical reinforcement learning

    CN114013443A