A reinforcement learning collision avoidance control method integrating dynamic evaluation of experience value

Through the modular structure reinforced learning collision avoidance control method, dynamic evaluation and classification of empirical samples, the problems of low sample utilization rate and poor adaptability of sparse reward environments in the existing technology are solved, and efficient acquisition of safe passage strategies and security guarantees of smart vehicles are achieved.

CN119538590BActive Publication Date: 2025-06-06CHANGCHUN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510031632.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-06-06
Estimated Expiration
2045-01-09

Smart Images

  • Figure CN119538590B_ABST
    Figure CN119538590B_ABST
Patent Text Reader

Abstract

A reinforcement learning collision avoidance control method integrating dynamic evaluation of experience value is provided to solve the problems of poor sample utilization rate of current intelligent driving control strategies. The present invention relates to the field of intelligent driving. The present invention comprises a risk assessment module, an experience sample classification module, an experience sample value assessment module and an experience pool allocation module. Among them, the risk assessment module evaluates the risk parameters of each environmental step in real time, the experience sample classification module classifies the experience samples according to the risk parameters and stores them in three experience pools of safety, observation and danger, the experience value assessment module evaluates the value of the experience samples in real time and dynamically, and sorts the experience samples in the three experience pools respectively, the experience pool allocation module evaluates the value of the three experience pools in real time and dynamically, determines the extraction ratio of the three experience pools, transmits them to the intelligent body for experience playback, updates the safe passage strategy, and repeats the above process until the optimal safe passage strategy is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical field:

[0001] The present invention belongs to the field of intelligent driving, and specifically is a reinforcement learning collision avoidance control method integrating dynamic evaluation of experience value. Technical background:

[0002] In daily traffic scenarios, the number of traffic participants is increasing day by day, which undoubtedly increases the complexity of the traffic environment and leads to huge challenges for intelligent driving technology. The development goals of intelligent driving technology are to improve road safety, reduce traffic congestion, improve travel efficiency, and provide drivers with a more comfortable and convenient driving experience.

[0003] Intelligent driving technology is gradually driving a revolutionary change in the transportation sector by integrating advanced sensors, high-precision maps, powerful computing platforms and complex algorithms to provide enhanced road safety, optimized traffic flow management, improved driving efficiency and a more comfortable and convenient travel experience. Although intelligent driving technology has excellent performance, it still has limitations in adapting to driving challenges in changing environments.

[0004] As an intelligent driving technology, reinforcement learning has shown great potential in intelligent driving technology, but it also has some significant disadvantages. First, it usually requires a large number of samples for effective learning, which is particularly prominent when the data acquisition cost is high or the environment is complex. Secondly, reinforcement learning faces challenges in dealing with environments with sparse rewards, and the agent may find it difficult to obtain sufficient feedback from the environment to optimize its strategy. Finally, reinforcement learning models may be difficult to generalize to unseen states or actions, which limits its application in diverse tasks. At present, in order to cope with the above challenges, solutions have been proposed in the field of reinforcement learning. Although this method has advantages, it still has limitations. Patent CN 118430246 A uses the collision time principle to evaluate the risk factor in the experience pool classification to classify the experience pool, which obviously improves the utilization rate of samples and significantly improves the performance of the reinforcement learning algorithm. However, this experience pool classification method also has certain limitations. First, if the entropy decreases too quickly, the agent may give up exploration too early, thus falling into a suboptimal strategy. Second, when using collision time to evaluate the risk factor, more environmental state information is required, which may increase the computational cost of the reinforcement learning model, and the final risk parameter is also a discrete state with low precision, which may cause the action learned by the agent to show jitter. Finally, if only risk experience samples and ordinary experience samples are considered, the intelligent driving strategy learned by the agent may fall into the local optimum, causing potential safety risks for intelligent vehicles. In addition, in the current experience sample extraction, random experience sample extraction is usually used, which helps to break the time correlation between experience samples and enhance the stability of the learning process. However, the random experience sample extraction method does not further screen the importance of experience samples in the experience pool, which may cause important experience samples to be ignored; at the same time, irrelevant or repeated experience samples may be over-sampled, which may make the learning process inefficient and not conducive to the agent learning from the most recent and more relevant experience, thus affecting the timely update and adaptation of the learning strategy. Summary of the invention:

[0005] In view of the shortcomings of the prior art and to solve the problems existing in the above technical background, the present invention provides a reinforcement learning collision avoidance control method integrating dynamic evaluation of experience value. The method adopts a modular structure, makes full use of the global optimization ability of reinforcement learning, and utilizes the joint action between various modules to achieve the acquisition of the optimal safe passage strategy in different scenarios.

[0006] The technical solution adopted by the present invention to solve the technical problem is as follows:

[0007] The present invention is a reinforcement learning collision avoidance control method integrating dynamic evaluation of experience value, the method comprising an environment, a risk evaluation module, an experience sample classification module, an experience sample value evaluation module, an experience pool allocation module and an intelligent agent; wherein the risk evaluation module receives the state of the current environment, and evaluates the risk parameters of each environment step in real time according to the control obstacle function; the experience sample classification module receives the above risk parameters, and classifies the experience samples according to the risk parameters, stores the safe experience samples in the safe experience pool, stores the to-be-observed experience samples in the to-be-observed experience pool, stores the dangerous experience samples in the dangerous experience pool, and outputs the experience pool after sample classification, which is recorded as experience pool A; the experience sample value evaluation module The value of all experience samples in experience pool A is evaluated dynamically in real time, and the experience samples are sorted in the three experience pools of safety, observation and danger according to the value of the experience samples. The sorted experience pool is recorded as experience pool B; the experience pool allocation module evaluates the value of the three experience pools of safety, observation and danger in experience pool B in real time and dynamically, and determines the extraction ratio of the three experience pools of safety, observation and danger according to the value of the experience pool, and extracts a batch of experience samples from the three experience pools of safety, observation and danger according to the extraction ratio; the intelligent agent receives the above batch of experience samples, replays the experience, and learns to update the strategy of safe passage; repeats the above process until the optimal safe passage strategy is obtained;

[0008] The method comprises the following steps:

[0009] Step 1: Reinforcement learning model design:

[0010] Step 1.1, state space design:

[0011] For intelligent driving tasks in the environment, the relative distance between the ego vehicle and the surrounding vehicles in the environment can intuitively reflect the relative motion relationship between the ego vehicle and the surrounding vehicles. Therefore, the state space in reinforcement learning is defined as follows:

[0012]

[0013] I i is the sensor's perception area for whether there are other vehicles on lane i, n is the number of lanes, lo and la are the relative distances between the vehicle and the obstacle in the longitudinal and lateral directions, Δlo and Δla are the corresponding change rates of lo and la, yaw and Δyaw are the vehicle's yaw angle and yaw angle change rate.

[0014] Step 1.2, action space design:

[0015] The action space is a continuous two-dimensional action space, which includes the lateral and longitudinal control quantities of the vehicle. Therefore, the action space in reinforcement learning is defined as follows:

[0016] a=[a 1,a 2 ],U 1 ≤a 1 ≤D 1 ; U 2 ≤a 2 ≤D 2 (2)

[0017] a 1 is the vehicle front wheel steering angle control value; a 2 is the vehicle throttle and brake control amount; U 1 and U 2 They are respectively 1 and a 2 The lower bound of 1 and D 2 They are respectively 1 and a 2 The upper bound of .

[0018] Step 1.3, reward function design:

[0019] The present invention defines the reward function of the intelligent driving task in the collision avoidance scenario as formula (3):

[0020]

[0021] ε is the risk parameter between the vehicle and the obstacle, la id and la hv are the lane boundary position and the lateral position of the vehicle, la center is the current lane center position, r risk is the reward item for vehicle risk, r invasion is the reward term between the vehicle and the lane boundary, r center is the reward between the vehicle and the lane centerline, r exist Reward item for vehicle accident violation.

[0022] Step 2: Construction of risk assessment module:

[0023] The risk assessment module includes an obstacle control function, which combines the state information of the ego vehicle and the obstacle to output a risk parameter ε between the ego vehicle and the obstacle. The obstacle control function is defined as shown in equations (4), (5) and (6).

[0024]

[0025] h(lo)=(lo safe ) 2 -(lo) 2 (5)

[0026] h(la)=(la safe )2 -(la) 2 (6)

[0027] lo and la are the relative distances between the vehicle and the obstacle in the longitudinal and lateral directions, respectively. safe and la safe are the relative safety distances between the vehicle and the obstacle in the longitudinal and lateral directions respectively.

[0028] Step 3: Construction of the experience sample classification module:

[0029] The empirical sample classification module defines the risk parameter threshold parameter ε 1 , ε 2 , when ε≤ε 1 When the experience sample is a safe experience sample, when ε 1 ≤ε≤ε 2 When the experience sample is the experience sample to be observed, when ε 2 ≤ε, the experience samples are dangerous experience samples. The safe, to-be-observed and dangerous samples constitute the safe experience pool, the to-be-observed experience pool and the dangerous experience pool respectively, recorded as experience pool A. In experience pool A, the safe experience samples are [l α ,s,a,r,s_] five-tuple form is stored in the security experience pool, and the experience samples to be observed are in the form of [l 1 ,s,a,r,s_] five-tuple form is stored in the waiting experience pool, and the dangerous experience samples are stored in the form of [t 1 ,t 2 ,s,a,r,s_] six-tuple form is stored in the dangerous experience pool; the capacity of the three experience pools in experience pool A follows M 1 =M 2 =M 3 Relationship, where M 1 is the capacity of the safety experience pool; M 2 is the capacity of the experience pool to be observed; M 3 is the capacity of the dangerous experience pool; l α is the temperature loss of the Soft Actor-Critic algorithm, l 1 is the policy loss of the actor network in the Soft Actor-Critic algorithm, t 1 Critic in the Soft Actor-Critic algorithm 1 The network’s timing difference error, t 2 Critic in the Soft Actor-Critic algorithm 2 The temporal difference error of the network, s is the current state, a is the action, r is the reward, and s_ is the next state.

[0030] Step 4: Construction of experience sample value assessment module:

[0031] The experience sample value evaluation module includes a safety experience value evaluator, an experience value evaluator to be observed, and a dangerous experience value evaluator; the safety experience value evaluator has two evaluation criteria, namely, the temperature loss l of the Soft Actor-Critic algorithm α and reward r; the observed experience value evaluator has two evaluation criteria, namely the policy loss l of the actor network in the Soft Actor-Critic algorithm 1 and reward r; the risk experience value evaluator has three evaluation criteria, namely, the critic in the Soft Actor-Critic algorithm 1 The network's timing difference error t 1 , critic in the Soft Actor-Critic algorithm 2 The network's timing difference error t 2 and reward r; input the value of the evaluation criterion j of all experience samples in the corresponding experience pool into the corresponding experience value evaluator, output the value of all experience samples in the corresponding experience pool, and sort the experience samples in experience pool A in descending order according to the value of the experience samples to obtain experience pool B; in experience pool B, the safety experience samples are sorted in the order of [V i ,l α ,s,a,r,s_] six-tuple form is stored in the security experience pool, and the experience samples to be observed are in the form of [V i ,l 1 ,s,a,r,s_] six-tuple form is stored in the waiting experience pool, and the dangerous experience samples are stored in the form of [V i ,t 1 ,t 2 ,s,a,r,s_] seven-tuple form is stored in the dangerous experience pool; the capacity of the three experience pools in the experience pool B follows M 1 =M 2 =M 3 Relationship, where M 1 is the capacity of the safety experience pool; M 2 is the capacity of the experience pool to be observed; M 3 V is the capacity of the dangerous experience pool; i is the value of all experience samples in the corresponding experience pool, s is the current state, a is the action, r is the reward, s_ is the next state, and the experience value evaluator is defined as formula (7), formula (8), formula (9), formula (10) and formula (11).

[0032]

[0033] P j =p(η j ≤δ ij ) (8)

[0034]

[0035]

[0036]

[0037] η j is the average value of evaluation criterion j, δ ij is the value of the evaluation criterion j of all experience samples in the corresponding experience pool, P j is the probability estimate that the value of evaluation criterion j of all experience samples is greater than or equal to the average value of evaluation criterion j, F j is the impact factor of evaluation criterion j in the experience pool, ω j is the influence weight of evaluation criterion j in the experience pool, V i is the value of all experience samples in the corresponding experience pool.

[0038] Step 5: Construction of experience pool allocation module:

[0039] The experience pool allocation module combines the experience pool B in step 4 and outputs a batch of experience samples with a number of D. To extract a batch of experience samples with a number of D from the experience pool B in step 4, D samples need to be extracted from the three experience pools respectively. 1 , D 2 and D 3 The number of empirical samples, the number of empirical samples D 1 , D 2 and D 3 According to the real-time dynamic adjustment of the experience pool allocation module, the experience pool allocation module is defined as formula (12), formula (13), formula (14) and formula (15),

[0040]

[0041] S k = p(V i ≥β k ) (13)

[0042]

[0043] D k =D·P k (k=1,2,3) (15)

[0044] V i is the value of all experience samples in the corresponding experience pool, β k is the average value of all experience samples in the corresponding experience pool, S kis the probability that the value of all experience samples in the corresponding experience pool is greater than or equal to the average value of all experience samples in the corresponding experience pool, P k is the extraction ratio of the corresponding experience pool, D k is the number of experience samples drawn from the corresponding experience pool.

[0045] Step 6: Update reinforcement learning network parameters:

[0046] Step 6.1. Select reinforcement learning algorithm:

[0047] Different reinforcement learning methods have their own advantages and disadvantages. In order to complete the intelligent driving task, the present invention selects the Soft Actor-Critic algorithm as the reinforcement learning algorithm for intelligent vehicles. The Soft Actor-Critic algorithm can select actions in a continuous action space. Due to the maximum entropy guidance, the strategies learned by the intelligent agent are more random, which helps to improve the generalization of the control strategy.

[0048] Step 6.2: Extraction of experience samples:

[0049] When the network parameters of the reinforcement learning agent are updated, a batch of experience samples with a size of D needs to be extracted from the experience pool, and D are provided from the three experience pools respectively. 1 , D 2 and D 3 The number of empirical samples, D 1 , D 2 and D 3 The definition of D is the same as in step 5 k .

[0050] Step 6.3, experience playback:

[0051] The experience samples from the safe experience pool, the experience samples from the to-be-observed experience pool, and the experience samples from the dangerous experience pool are split and combined into D-dimensional vectors of state s, action a, reward r, and next state s_, which are transmitted to the agent for experience replay. The agent calculates the loss function Loss value of each network for back propagation to update parameters, thereby learning the updated safe passage strategy. The specific process of experience replay is as follows:

[0052] Initialize the network parameters of each neural network in the Soft Actor-Critic algorithm ω 1 ,ω 2 ,ω 1_ ,ω 2_ , is the Actor network parameter, ω 1 Critic 1 Network parameters, ω 2Critic 2 Network parameters, ω 1 _ Target Critic 1 Network parameters, ω 2 _ Target Critic 2 Network parameters. When replaying the experience, the environment information is obtained from the environment to form the state s at time t in step 1.1 t Status t Input to the Actor network, and output the action a in the action space range in step 1.2 through global optimization. Action a is executed in the driving environment, and the reward function in step 1.3 calculates the reward value r at time t t , and get the state s at time t+1 t+1 When the number of experience samples in the three experience pools reaches D, D experience samples are taken from experience pool B in step 4. The experience samples are input to the Critic 1 Network, Critic 2 Network, Target Critic 1 Network and Target Critic 2 The network gets the Q value and target Q value Use equation (16) to calculate the time series difference target,

[0053]

[0054]

[0055]

[0056] γ is the discount factor. By minimizing the loss function in equations (17) and (18), Critic 1 Network and Critic 2 The network is updated.

[0057] Sampling actions using a reparameterized approach

[0058]

[0059]

[0060] The temperature coefficient α is updated using the temperature loss function in equation (19), and the Actor network is updated using the strategy loss function in equation (20).

[0061] α is the entropy regularization coefficient, the target critic 1 Network and Target Critic 2The network parameters are updated using equations (21) and (22),

[0062] ω 1 _ ←τω 1 +(1-τ)ω 1 _ (twenty one)

[0063] ω 2 _ ←τω 2 +(1-τ)ω 2 _ (twenty two)

[0064] τ is a soft update coefficient. The above training process is iterated repeatedly until the algorithm converges. After the algorithm converges, the optimal set of network parameters is selected and loaded into the Actor network to complete the reinforcement learning training of the present invention. Description of the drawings:

[0065] Figure 1 It is a flow chart of the method of the present invention.

[0066] Figure 2 It is an application scenario diagram of the present invention. Specific implementation method:

[0067] The present invention is described in detail below in conjunction with the accompanying drawings.

[0068] The present invention is a reinforcement learning collision avoidance control method integrating dynamic evaluation of experience value, the method includes an environment, a risk evaluation module, an experience sample classification module, an experience sample value evaluation module, an experience pool allocation module and an intelligent agent; wherein the risk evaluation module receives the state of the current environment, and evaluates the risk parameters of each environment step in real time according to the control obstacle function; the experience sample classification module classifies the experience samples according to the risk parameters, and divides the experience samples into safe experience samples, to-be-observed experience samples and dangerous experience samples, and stores them in three experience pools of safety, to-be-observed and dangerous, respectively, which are recorded as experience pool A; the experience sample value evaluation module dynamically evaluates the value of all experience samples in experience pool A in real time, and sorts the experience samples in the three experience pools according to the value of the experience samples, and the sorted experience pools are recorded as experience pool B; the experience pool allocation module dynamically evaluates the value of the three experience pools of safety, to-be-observed and dangerous in experience pool B in real time, and determines the extraction ratio of the three experience pools according to the value of the experience pool, and extracts a batch of experience samples; the intelligent agent receives a batch of experience samples, replays the experience, and learns to update the safe passage strategy; the above process is repeated until the optimal safe passage strategy is obtained. Reference Figure 1 The schematic diagram specifically includes the following steps:

[0069] Step 1: Environment construction:

[0070] The present invention has high versatility and adaptability to different environments, and improves the generalization of the reinforcement learning model. Since the driving safety of intelligent driving vehicles should be ensured first in real life, the main application environment of the present invention is the collision avoidance scenario. Figure 2 As shown, it specifically includes the following three environments:

[0071] Step 1.1: Construction of Scene I:

[0072] All surrounding vehicles in scenario I are stationary vehicles, and stationary vehicles are randomly parked on the 420m road. In order to prevent two parallel cars from blocking the road, each 60m road segment is divided into four sub-segments. The ego vehicle and surrounding vehicles are placed separately in a sub-segment. The location and lane selection of the placed vehicles are randomly initialized by Gaussian sampling method.

[0073] Step 1.2: Construction of Scene II:

[0074] All the surrounding vehicles in Scenario II are moving vehicles at a constant speed. All the settings in this scenario are the same as those in Scenario I in step 1.1. However, in this scenario, all the surrounding vehicles move at a constant speed of 30 m / s.

[0075] Step 1.3, Construction of Scene III:

[0076] All surrounding vehicles in scenario III are moving vehicles that can accelerate, decelerate, and change lanes. All settings in this scenario are the same as scenario II in step 1.2. However, in this scenario, all surrounding vehicles have the possibility of accelerating, decelerating, or changing lanes.

[0077] Step 2: Reinforcement learning model design:

[0078] Step 2.1, state space design:

[0079] The state is an important part of the Markov decision process, which represents the environmental feature information related to the training task. For the intelligent driving task in the environment, the relative distance between the ego vehicle and the surrounding vehicles in the environment can intuitively reflect the relative motion relationship between the ego vehicle and the surrounding vehicles. Therefore, the state space in reinforcement learning is defined as follows:

[0080]

[0081] I i is whether there are other vehicles on lane i within the sensor sensing range (+50m, -5m), n is the number of lanes, lo and la are the relative distances between the vehicle and the obstacle in the longitudinal and lateral directions, Δlo and Δla are the corresponding change rates of lo and la, yaw and Δyaw are the vehicle yaw angle and yaw angle change rate.

[0082] Step 2.2, action space design:

[0083] The action space is a continuous two-dimensional action space, which includes the lateral and longitudinal control quantities of the vehicle. Therefore, the action space in reinforcement learning is defined as follows:

[0084] a=[a 1 ,a 2 ],U 1 ≤a 1 ≤D 1 ; U 2 ≤a 2 ≤D 2 (twenty four)

[0085] a 1 is the vehicle front wheel steering angle control value; a 2 is the vehicle throttle and brake control amount; U 1 and U 2 They are respectively 1 and a 2 The lower bound of 1 and D 2 They are respectively 1 and a 2 The upper bound of .

[0086] Step 2.3, reward function design:

[0087] In order to ensure the safety of intelligent driving vehicles, during the training process of reinforcement learning, the reward value plays an evaluation role in the intelligent agent's selection of control actions, thereby guiding the intelligent agent to learn a reasonable control strategy. According to the environment in step 1, the present invention defines the reward function of the intelligent driving task in the collision avoidance scenario as formula (25),

[0088]

[0089] ε is the risk parameter between the vehicle and the obstacle, la id and la hv are the lane boundary position and the lateral position of the vehicle, la center is the current lane center position, r risk is the reward item for vehicle risk, r invasion is the reward term between the vehicle and the lane boundary, r center is the reward between the vehicle and the lane centerline, r exist Reward item for vehicle accident violation.

[0090] Step 3: Construction of risk assessment module:

[0091] The risk assessment module includes an obstacle control function, which combines the state information of the ego vehicle and the obstacle to output a risk parameter ε between the ego vehicle and the obstacle. The obstacle control function is defined as shown in equations (26), (27) and (28).

[0092]

[0093] h(lo)=(lo safe ) 2 -(lo) 2 (27)

[0094] h(la)=(la safe ) 2 -(la) 2 (28)

[0095] lo and la are the relative distances between the vehicle and the obstacle in the longitudinal and lateral directions, respectively. safe and la safe are the relative safety distances between the vehicle and the obstacle in the longitudinal and lateral directions respectively.

[0096] Step 4: Construction of the experience sample classification module:

[0097] The empirical sample classification module defines the risk parameter threshold parameter ε 1 , ε 2 , where ε 1 =1.5,ε 2 =3; when ε≤ε 1 When the experience sample is a safe experience sample, when ε 1 ≤ε≤ε 2 When the experience sample is the experience sample to be observed, when ε 2 ≤ε, the experience samples are dangerous experience samples. The safe, to-be-observed and dangerous samples constitute the safe experience pool, the to-be-observed experience pool and the dangerous experience pool respectively, recorded as experience pool A. In experience pool A, the safe experience samples are [l α ,s,a,r,s_] five-tuple form is stored in the security experience pool, and the experience samples to be observed are in the form of [l 1 ,s,a,r,s_] five-tuple form is stored in the waiting experience pool, and the dangerous experience samples are stored in the form of [t 1 ,t 2 ,s,a,r,s_] six-tuple form is stored in the dangerous experience pool; the capacity of the three experience pools in experience pool A follows M 1 =M 2 =M 3 Relationship, where M 1 is the capacity of the safety experience pool; M 2 is the capacity of the experience pool to be observed; M 3 is the capacity of the dangerous experience pool; lα is the temperature loss of the Soft Actor-Critic algorithm, l 1 is the policy loss of the actor network in the Soft Actor-Critic algorithm, t 1 Critic in the Soft Actor-Critic algorithm 1 The network’s timing difference error, t 2 Critic in the Soft Actor-Critic algorithm 2 The temporal difference error of the network, s is the current state, a is the action, r is the reward, and s_ is the next state.

[0098] Step 5: Construction of experience sample value assessment module:

[0099] The experience sample value evaluation module includes a safety experience value evaluator, an experience value evaluator to be observed, and a dangerous experience value evaluator; the safety experience value evaluator has two evaluation criteria, namely, the temperature loss l of the Soft Actor-Critic algorithm α and reward r; the observed experience value evaluator has two evaluation criteria, namely the policy loss l of the actor network in the Soft Actor-Critic algorithm 1 and reward r; the risk experience value evaluator has three evaluation criteria, namely, the critic in the Soft Actor-Critic algorithm 1 The network's timing difference error t 1 , critic in the Soft Actor-Critic algorithm 2 The network's timing difference error t 2 and reward r; input the value of the evaluation criterion j of all experience samples in the corresponding experience pool into the corresponding experience value evaluator, output the value of all experience samples in the corresponding experience pool, and sort the experience samples in experience pool A in descending order according to the value of the experience samples to obtain experience pool B; in experience pool B, the safety experience samples are sorted in the order of [V i ,l α ,s,a,r,s_] six-tuple form is stored in the security experience pool, and the experience samples to be observed are in the form of [V i ,l 1 ,s,a,r,s_] six-tuple form is stored in the waiting experience pool, and the dangerous experience samples are stored in the form of [V i ,t 1 ,t 2 ,s,a,r,s_] seven-tuple form is stored in the dangerous experience pool; the capacity of the three experience pools in the experience pool B follows M 1 =M 2 =M 3Relationship, where M 1 is the capacity of the safety experience pool; M 2 is the capacity of the experience pool to be observed; M 3 V is the capacity of the dangerous experience pool; i is the value of all experience samples in the corresponding experience pool, s is the current state, a is the action, r is the reward, and s_ is the next state. The experience value evaluator is defined as formula (29), formula (30), formula (31), formula (32) and formula (33),

[0100]

[0101] P j =p(η j ≤δ ij ) (30)

[0102]

[0103]

[0104]

[0105] η j is the average value of evaluation criterion j, δ ij is the value of the evaluation criterion j of all experience samples in the corresponding experience pool, P j is the probability estimate that the value of evaluation criterion j of all experience samples is greater than or equal to the average value of evaluation criterion j, F j is the impact factor of evaluation criterion j in the experience pool, ω j is the influence weight of evaluation criterion j in the experience pool, V i is the value of all experience samples in the corresponding experience pool.

[0106] Step 6: Construction of experience pool allocation module:

[0107] The experience pool allocation module combines the experience pool B in step 5 and outputs a batch of experience samples with a number of D. To extract a batch of experience samples with a number of D from the experience pool B in step 5, D samples need to be extracted from the three experience pools respectively. 1 , D 2 and D 3 The number of empirical samples, the number of empirical samples D 1 , D 2 and D 3 According to the real-time dynamic adjustment of the experience pool allocation module, the experience pool allocation module is defined as formula (34), formula (35), formula (36) and formula (37),

[0108]

[0109] S k = p(V i ≥β k ) (35)

[0110]

[0111] D k =D·P k (k=1,2,3) (37)

[0112] V i is the value of all experience samples in the corresponding experience pool, β k is the average value of all experience samples in the corresponding experience pool, S k is the probability that the value of all experience samples in the corresponding experience pool is greater than or equal to the average value of all experience samples in the corresponding experience pool, P k is the extraction ratio of the corresponding experience pool, D k is the number of experience samples drawn from the corresponding experience pool.

[0113] Step 7: Update reinforcement learning network parameters:

[0114] Step 7.1. Select reinforcement learning algorithm:

[0115] Different reinforcement learning methods have their own advantages and disadvantages. In order to complete the intelligent driving task in the environment of step 1, the present invention selects the Soft Actor-Critic algorithm as the reinforcement learning algorithm for intelligent vehicles. The Soft Actor-Critic algorithm can select actions in a continuous action space. Due to the guidance of maximum entropy, the strategies learned by the intelligent agent are more random, which helps to improve the generalization of the control strategy.

[0116] Step 7.2: Extraction of experience samples:

[0117] When the network parameters of the reinforcement learning agent are updated, a batch of experience samples with a size of D needs to be extracted from the experience pool, and D are provided from the three experience pools respectively. 1 , D 2 and D 3 The number of empirical samples, D 1 , D 2 and D 3 The definition of D is the same as in step 6 k .

[0118] Step 7.3, experience playback:

[0119] The experience samples from the safe experience pool, the experience samples from the to-be-observed experience pool, and the experience samples from the dangerous experience pool are split and combined into D-dimensional vectors of state s, action a, reward r, and next state s_, which are transmitted to the agent for experience replay. The agent calculates the loss function Loss value of each network for back propagation to update parameters, thereby learning the updated safe passage strategy. The specific process of experience replay is as follows:

[0120] α is set to -2, τ is 0.005, and γ is 0.99 to initialize the network parameters of each neural network in the Soft Actor-Critic algorithm. ω 1 ,ω 2 ,ω 1_ ,ω 2_ , is the Actor network parameter, ω 1 Critic 1 Network parameters, ω 2 Critic 2 Network parameters, ω 1 _ Target Critic 1 Network parameters, ω 2 _ Target Critic 2 Network parameters. When replaying the experience, the environment information obtained from the environment in step 1 constitutes the state s at time t in step 2.1 t Status t Input to the Actor network, and output the action a in the action space range in step 2.2 through global optimization. Action a is executed in the driving environment, and the reward function in step 2.3 calculates the reward value r at time t t , and get the state s at time t+1 t+1 When the number of experience samples in the three experience pools reaches D, D experience samples are taken from experience pool B in step 5. The experience samples are input to the Critic 1 Network, Critic 2 Network, Target Critic 1 Network and Target Critic 2 The network gets the Q value and target Q value Use equation (16) to calculate the time series difference target,

[0121]

[0122]

[0123]

[0124] γ is the discount factor. By minimizing the loss function in equations (39) and (40), Critic 1 Network and Critic 2 The network is updated.

[0125] Sampling actions using a reparameterized approach

[0126]

[0127]

[0128] The temperature coefficient α is updated using the temperature loss function in equation (41), and the Actor network is updated using the strategy loss function in equation (42).

[0129] α is the entropy regularization coefficient, the target critic 1 Network and Target Critic 2 The network parameters are updated using equations (43) and (44),

[0130] ω 1 - ←τω 1 +(1-τ)ω 1 - (43)

[0131] ω 2 - ←τω 2 +(1-τ)ω 2 - (44)

[0132] τ is a soft update coefficient. The above training process is iterated repeatedly until the algorithm converges. After the algorithm converges, the optimal set of network parameters is selected and loaded into the Actor network to complete the reinforcement learning training of the present invention.

[0133] In summary: the present invention proposes a reinforcement learning collision avoidance control method that integrates dynamic evaluation of experience value, which usually does not require a large number of samples for learning, can be learned in situations where data acquisition costs are high or the environment is complex, and can also be learned in environments with sparse rewards, and the agent can obtain sufficient feedback from the environment to optimize its strategy. Finally, the reinforcement learning model can generalize to unseen states or actions, which broadens its application in diverse tasks.

Claims

1. A reinforcement learning collision avoidance control method integrating dynamic evaluation of experience value, characterized by: The method includes an environment, a risk assessment module, an experience sample classification module, an experience sample value assessment module, an experience pool allocation module and an intelligent agent; wherein the risk assessment module receives the state of the current environment, and evaluates the risk parameters of each environment step in real time according to the control obstacle function; the experience sample classification module classifies the experience samples according to the risk parameters, and divides the experience samples into safe experience samples, to-be-observed experience samples and dangerous experience samples, and stores them in three experience pools of safety, to-be-observed and dangerous, respectively, which are recorded as experience pool A; the experience sample value assessment module dynamically evaluates the values ​​of all experience samples in experience pool A in real time, and sorts the experience samples in the three experience pools according to the values ​​of the experience samples, and the sorted experience pools are recorded as experience pool B; the experience pool allocation module dynamically evaluates the values ​​of the three experience pools of safety, to-be-observed and dangerous in experience pool B in real time, and determines the extraction ratio of the three experience pools according to the values ​​of the experience pools, and extracts a batch of experience samples; the intelligent agent receives a batch of experience samples, replays the experience, and learns to update the safe passage strategy; the above process is repeated until the optimal safe passage strategy is obtained; The risk assessment module controls the obstacle function by combining the state information of the vehicle and the obstacle, and outputs the risk parameter ε between the vehicle and the obstacle. The control obstacle function is defined as shown in formula (1), formula (2) and formula (3). h(lo)=(lo safe ) 2 -(it) 2 (2) <h2 style=";text-align:left;direction:ltr">h(la)=(la<h2 style=";text-align:left;direction:ltr"> safe <h2 style=";text-align:left;direction:ltr"> )<h2 style=";text-align:left;direction:ltr"> 2 <h2 style=";text-align:left;direction:ltr"> -(la)<h2 style=";text-align:left;direction:ltr"> 2 <h2 style=";text-align:left;direction:ltr"> (3) Among them, lo and la are the relative distances between the vehicle and the obstacle in the longitudinal and lateral directions, lo safe and la safe are the relative safety distances between the vehicle and the obstacle in the longitudinal and lateral directions respectively; The experience sample classification module defines risk parameter threshold parameters ε1, ε2. When ε≤ε1, the experience sample is a safe experience sample. When ε1≤ε≤ε2, the experience sample is an experience sample to be observed. When ε2≤ε, the experience sample is a dangerous experience sample. The safe, to-be-observed and dangerous samples constitute a safe experience pool, an experience pool to be observed and a dangerous experience pool, respectively, which are recorded as experience pool A. In experience pool A, safe experience samples are [l α ,s,a,r,s_] quintuples are stored in the safe experience pool, the to-be-observed experience samples are stored in the to-be-observed experience pool in the form of [l1,s,a,r,s_] quintuples, and the dangerous experience samples are stored in the dangerous experience pool in the form of [t1,t2,s,a,r,s_] quintuples; l α is the temperature loss of the Soft Actor-Critic algorithm, l1 is the policy loss of the actor network in the Soft Actor-Critic algorithm, t1 is the temporal difference error of the critic1 network in the Soft Actor-Critic algorithm, t2 is the temporal difference error of the critic2 network in the Soft Actor-Critic algorithm, s is the current state, a is the action, r is the reward, and s_ is the state at the next moment; The experience sample value evaluation module includes a safety experience value evaluator, an experience value evaluator to be observed, and a danger experience value evaluator; the safety experience value evaluator has two evaluation criteria, namely, the temperature loss l of the Soft Actor-Critic algorithm α and reward r; the experience value evaluator to be observed has two evaluation criteria, namely, the policy loss l1 and reward r of the actor network in the Soft Actor-Critic algorithm; the dangerous experience value evaluator has three evaluation criteria, namely, the temporal difference error t1 of the critic1 network in the Soft Actor-Critic algorithm, the temporal difference error t2 and reward r of the critic2 network in the Soft Actor-Critic algorithm; the value of the evaluation criterion j of all experience samples in the corresponding experience pool is input into the corresponding experience value evaluator, the value of all experience samples in the corresponding experience pool is output, and the experience samples in the experience pool A are sorted according to the value of the experience samples to obtain the experience pool B; the safe experience samples in the experience pool B are ranked by [V i ,l α ,s,a,r,s_] six-tuple form is stored in the security experience pool, and the experience samples to be observed are in the form of [V i ,l1,s,a,r,s_] six-tuple form is stored in the waiting experience pool, and the dangerous experience samples are stored in the form of [V i ,t1,t2,s,a,r,s_] seven-tuple form is stored in the dangerous experience pool; V i is the value of all experience samples in the corresponding experience pool, s is the current state, a is the action, r is the reward, s_ is the next state, and the experience value evaluator is defined as formula (4), formula (5), formula (6), formula (7) and formula (8), Where η j is the average value of evaluation criterion j, δ ij is the value of the evaluation criterion j of all experience samples in the corresponding experience pool, P j is the probability estimate that the value of evaluation criterion j of all experience samples is greater than or equal to the average value of evaluation criterion j, F j is the impact factor of evaluation criterion j in the experience pool, ω j is the influence weight of evaluation criterion j in the experience pool, V i is the value of all experience samples in the corresponding experience pool; The experience pool allocation module, in combination with the experience pool B, outputs a batch of experience samples with a number of D. To extract a batch of experience samples with a number of D from the experience pool B, it is necessary to provide experience samples with numbers D1, D2, and D3 from the three experience pools respectively. The numbers of experience samples D1, D2, and D3 are dynamically adjusted in real time according to the experience pool allocation module. The experience pool allocation module is defined as formula (9), formula (10), formula (11), and formula (12). S k =p(V i ≥β k ) (10) D k =D·P k (k=1,2,3) (12) Where V i is the value of all experience samples in the corresponding experience pool, β k is the average value of all experience samples in the corresponding experience pool, S k is the probability that the value of all experience samples in the corresponding experience pool is greater than or equal to the average value of all experience samples in the corresponding experience pool, P k is the extraction ratio of the corresponding experience pool, D k is the number of experience samples drawn from the corresponding experience pool; The state space in the reinforcement learning method is defined as in formula (13): In the formula, I i is whether there are other vehicles on lane i within the sensor sensing area, n is the number of lanes, lo and la are the relative distances between the vehicle and the obstacle in the longitudinal and lateral directions, Δlo and Δla are the corresponding change rates of lo and la, yaw and Δyaw are the vehicle yaw angle and yaw angle change rate; The action space in the reinforcement learning method is defined as formula (14): a=[a1,a2],U1≤a1≤D1; U2≤a2≤D2 (14) Wherein, the action space a is a continuous two-dimensional action space, including the lateral and longitudinal control quantities of the vehicle, a1 is the front wheel steering angle control quantity of the vehicle; a2 is the throttle and brake control quantity of the vehicle; U1 and U2 are the lower bounds of a1 and a2 respectively; D1 and D2 are the upper bounds of a1 and a2 respectively; The reinforcement learning is trained by interaction with the environment. For different environments, the reward function in the reinforcement learning method is defined as formula (15): Among them, ε is the risk parameter between the vehicle and the obstacle, la id and la hv are the lane boundary position and the lateral position of the vehicle, la center is the current lane center position, r risk is the reward item for vehicle risk, r invasion is the reward term between the vehicle and the lane boundary, r center is the reward between the vehicle and the lane centerline, r exist Reward item for vehicle accident violation.

Citation Information

Patent Citations

  • Automatic driving lane changing decision control method based on rule fusion reinforcement learning

    CN115257745A

  • Automatic driving method and system based on experience playback constraint strategy optimization

    CN118410856A