Vehicle longitudinal control method and system based on WOA-improved TD3 algorithm

By improving the TD3 algorithm and Actor-Critic network mechanism, combined with the WOA optimization algorithm, the decision-making lag problem of the autonomous driving system in complex scenarios is solved, and the adaptability and safety of the vehicle's longitudinal control are improved.

CN120116966BActive Publication Date: 2025-09-05ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510612183.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-09-05
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

When faced with complex scenarios, the existing autonomous driving systems' decision-making algorithms lack real-time performance, resulting in delayed responses, lack of flexibility and self-optimization capabilities, and affecting safety and efficiency.

Method used

A vehicle longitudinal control method based on the WOA-improved TD3 algorithm is adopted to establish a longitudinal dynamics model. An asymmetric Actor-Critic network mechanism is introduced. The prediction model is trained through the priority experience replay mechanism and the WOA optimization algorithm to optimize the vehicle longitudinal control.

Benefits of technology

It improves the convergence speed and stability of training, has adaptive capabilities, and can achieve safe and effective vehicle control in complex scenarios, reduce traffic congestion, and maintain a safe following distance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120116966B_ABST
    Figure CN120116966B_ABST
Patent Text Reader

Abstract

The present invention relates to a vehicle longitudinal control method and system based on a WOA-improved TD3 algorithm. A longitudinal dynamics model is established, and a driving state strategy for associated vehicles is configured based on the driving condition of the vehicle to be controlled. An asymmetric Actor-Critic network mechanism is introduced to establish a TD3 algorithm network prediction model based on a Markov decision process. The expected acceleration output by the model is iteratively trained using a priority experience replay mechanism and a WOA optimization algorithm. The trained prediction model is used to perform longitudinal control of the vehicle to be controlled. The system includes a controller, a drive system, a transmission system, and a sensor device configured on the vehicle to be controlled. The sensor device obtains the driving state of the associated vehicle of the vehicle to be controlled and transmits it to the controller. The controller uses a method to obtain the expected acceleration, which is converted into a control instruction. The drive system controls the acceleration and braking of the vehicle to be controlled based on the control instruction, and the transmission system cooperates with the drive system to achieve power transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of general control or regulation systems; functional units of such systems; and monitoring or testing devices for such systems or units, and in particular to a vehicle longitudinal control method and system based on a WOA-improved TD3 algorithm. Background Art

[0002] With the advancement of technology, Advanced Driving Assistance Systems (ADAS) have emerged. Using various sensors installed on the vehicle, ADAS senses the surrounding environment, collects data, and identifies, detects, and tracks static and dynamic objects while the car is in motion. Combined with navigation map data, ADAS performs systematic calculations and analysis, thereby proactively avoiding driving risks and effectively improving driving comfort and safety.

[0003] By 2023, the global penetration rate of vehicles equipped with advanced driver assistance systems (ADAS) will reach 45%. Level 2 systems (68%) primarily utilize a combination of PID and MPC algorithms, while Level 3 systems primarily employ a combination of MPC algorithms and reinforcement learning. Because ADAS can provide timely assistance and compensation when drivers make operational errors, preventing casualties and economic losses, improving the reliability and safety of partially automated driving systems in vehicles is of great engineering significance.

[0004] Existing autonomous driving systems lack real-time decision-making algorithms when faced with complex scenarios such as vehicles suddenly cutting in ahead or pedestrians crossing the road. This can lead to delayed responses and a lack of flexibility and self-optimization capabilities, resulting in reduced safety. Current research focuses on how to fully utilize road conditions, avoid traffic jams, and improve efficiency while ensuring safe driving. Summary of the Invention

[0005] The present invention solves the problems existing in the prior art and provides a vehicle longitudinal control method and system based on the WOA-improved TD3 algorithm.

[0006] The technical solution adopted by the present invention is a vehicle longitudinal control method based on the WOA-improved TD3 algorithm. The method establishes a longitudinal dynamics model and configures a driving state strategy for the associated vehicle based on the driving condition of the vehicle to be controlled. It also introduces an asymmetric actor-critic network mechanism and establishes a TD3 algorithm network prediction model based on a Markov decision process. The expected acceleration output is iteratively trained using a prioritized experience replay mechanism and a WOA optimization algorithm.

[0007] The trained prediction model is used to perform longitudinal control on the vehicle to be controlled.

[0008] Preferably, the associated vehicles include the preceding vehicle and the adjacent vehicle of the vehicle to be controlled, and the longitudinal dynamics model definition satisfies:

[0009] ,

[0010] in, is the distance error between the vehicle to be controlled and the preceding vehicle at time k, is the distance between the vehicle to be controlled and the preceding vehicle at time k, is the expected safety distance at time k, is the speed difference between the vehicle to be controlled and the preceding vehicle at time k, is the speed of the vehicle to be controlled at time k, is the speed of the preceding vehicle at time k, is the position of the preceding vehicle, is the position of the vehicle to be controlled, is the acceleration of the vehicle to be controlled at time k, is the calculation time step.

[0011] Preferably, the driving state strategy of the associated vehicle includes a straight-ahead state, a lane-changing state, and a braking state; random disturbances are set in conjunction with the driving state strategy; in the present invention, random disturbances are generally set in the form of noise.

[0012] Preferably, establishing the prediction model comprises the following steps:

[0013] S3.1 Establish the state space and action space of the vehicle to be controlled;

[0014] S3.2 Design the reward and penalty functions for the multi-objective control of the vehicle to be controlled;

[0015] S3.3 introduces an asymmetric actor-critic architecture and establishes a prediction model consisting of one actor network and two critic networks.

[0016] Preferably, the reward and penalty function of the multi-objective control includes a safety reward function, a relative speed penalty, an acceleration smoothness penalty, a collision penalty and a continuous following reward.

[0017] Preferably, the Actor network includes three fully connected hidden layers, and the learning rate of the Actor network is The WOA algorithm is optimized and iterated before delivery;

[0018] The two critic networks each include four fully connected hidden layers, and the learning rate of the critic network is The WOA algorithm is optimized and passed after iteration.

[0019] Preferably, training the prediction model comprises the following steps:

[0020] S4.1 Build a priority experience replay buffer pool to store interaction data, initialize it, and use the optimal vector decoding after initial WOA optimization as the hyperparameter of the prediction model; obtain the initial state S 0 = [ d a c t , d d e s , v , v p ] ;

[0021] S4.2 In the current state, the Actor network Select and execute actions based on the current strategy , obtain the next state and reward and punishment function state; store the experience data into the priority experience replay buffer pool;

[0022] S4.3 Sampling a batch of experience data from the priority experience replay buffer pool, sampling probability Priority in the priority experience replay mechanism Related;

[0023] S4.4 Use the target policy network of the actor network to select the next action and calculate the target Q value based on the critic network;

[0024] S4.5 updates the critic network and calculates the loss of the current critic network. Based on the priority experience replay mechanism, the importance sampling weight of each experience sample is calculated.

[0025] S4.6 Based on the preset delayed update rules for the Actor network, if the update conditions are met, the Actor network is updated once and the loss of the Actor network is calculated. Otherwise, proceed directly to S4.7;

[0026] S4.7 Based on the calculated sampling probability , randomly select a set of experience data from the priority experience replay buffer pool , and update the experience priority according to the TD error;

[0027] S4.8 Update the target critic network;

[0028] S4.9 When the set number of training rounds is reached or the performance meets the requirements, the training ends and the optimal expected acceleration is obtained. Otherwise, the WOA optimization iteration continues and steps S4.2 to S4.8 are repeated.

[0029] Preferably, the sampling probability satisfy ,in, For samples Priority, is the priority strength, idk is the index, representing each experience in the priority replay buffer pool, p idxis the priority indexed by idx;

[0030] Importance sampling weights satisfy ,in, is the importance sampling correction factor, The size of the priority experience replay buffer pool;

[0031] TD error satisfies ,in, is the instant reward obtained by the vehicle to be controlled at the current moment, is the value estimate of the action taken by the vehicle to be controlled at the next moment, is the value estimate of the action taken by the vehicle to be controlled at the current moment, is the discount factor.

[0032] Preferably, the WOA optimization iteration includes the following steps:

[0033] S4.1.1 Determine the hyperparameters to be optimized and the search range;

[0034] S4.1.2 Initialize the WOA algorithm, set the maximum number of iterations, initialize the population and randomly generate a group; initialize the fitness of each whale, using the cumulative reward value after training as the fitness function; initialize the global optimal solution and optimal fitness;

[0035] S4.1.3 For each whale, use the modified TD3 algorithm model with the current hyperparameter configuration; calculate the fitness using the cumulative reward value, and update the global optimal solution if it is greater than the optimal fitness;

[0036] S4.1.4 Update the current optimal solution.

[0037] A vehicle longitudinal control system based on the WOA-improved TD3 algorithm, the system comprising a controller, a drive system, a transmission system, and a sensing device configured on any vehicle to be controlled. The sensing device obtains the driving status of a vehicle associated with the vehicle to be controlled and transmits it to the controller. The controller uses the vehicle longitudinal control method based on the WOA-improved TD3 algorithm to obtain a desired acceleration. The controller converts the desired acceleration into a control instruction. The drive system controls the acceleration and braking of the vehicle to be controlled based on the control instruction. The transmission system cooperates with the drive system to achieve power transmission.

[0038] The present invention relates to a vehicle longitudinal control method and system based on a WOA-improved TD3 algorithm, which establishes a longitudinal dynamics model and configures a driving state strategy for associated vehicles based on the driving condition of a to-be-controlled vehicle; introduces an asymmetric Actor-Critic network mechanism, establishes a TD3 algorithm network prediction model based on a Markov decision process, and iteratively trains the expected acceleration output by the model using a priority experience replay mechanism and a WOA optimization algorithm; longitudinally controls the to-be-controlled vehicle using the trained prediction model; the system includes a controller, a drive system, a transmission system, and a sensor device configured on any to-be-controlled vehicle; the sensor device acquires the driving state of an associated vehicle of the to-be-controlled vehicle and transmits it to the controller; the controller acquires the expected acceleration using a method; the controller converts the expected acceleration into a control instruction; the drive system controls the acceleration and braking of the to-be-controlled vehicle based on the control instruction; and the transmission system cooperates with the drive system to achieve power transmission.

[0039] The beneficial effects of the present invention are:

[0040] (1) The transmission system controls the vehicle based on the desired acceleration control. The vehicle moves at the desired speed, allowing the vehicle to make its own control decisions.

[0041] (2) Through the coordinated application of the WOA optimization algorithm, dual-critic network, priority experience buffer replay mechanism, and curriculum learning mechanism, the convergence speed of training is effectively improved, the sample utilization rate is high, the training stability is strong, and it has adaptive capabilities. Parameter adjustment can be automatically completed without manual intervention, and the end-to-end process is more effectively realized;

[0042] (3) After the experiment, the average distance error was 0.54m and the average collision rate was 0, which made the following distance less different from the expected distance. It can effectively avoid traffic jams and maintain safe following while making full use of the road. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a flow chart of the method of the present invention;

[0044] Figure 2 Schematic diagram of the network framework of the WOA outer loop and the improved TD3 inner loop in the implementation of the prediction model of the present invention;

[0045] Figure 3 3 is a comparison chart of the effects of the embodiment of the present invention and the existing TD3 algorithm. DETAILED DESCRIPTION

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0047] The present invention relates to a vehicle longitudinal control method based on a WOA-improved TD3 algorithm, the method comprising the following steps:

[0048] (1) Establish a longitudinal dynamic model;

[0049] (2) configuring the driving state strategy of the associated vehicle based on the driving condition of the vehicle to be controlled;

[0050] (3) Introducing an asymmetric Actor-Critic network mechanism and establishing a TD3 algorithm network prediction model based on the Markov decision process;

[0051] (4) Iterative training is performed on the expected acceleration output using the priority experience replay mechanism and the WOA optimization algorithm;

[0052] (5) Use the trained prediction model to perform longitudinal control on the vehicle to be controlled.

[0053] The following describes the method in conjunction with the specific implementation process.

[0054] (1) Establish a longitudinal dynamic model;

[0055] The associated vehicles include the preceding vehicle and the adjacent vehicles of the vehicle to be controlled. The longitudinal dynamics model definition satisfies:

[0056] ,

[0057] in, is the distance error between the vehicle to be controlled and the preceding vehicle at time k, is the distance between the vehicle to be controlled and the preceding vehicle at time k, is the expected safety distance at time k, is the speed difference between the vehicle to be controlled and the preceding vehicle at time k, is the speed of the vehicle to be controlled at time k, is the speed of the preceding vehicle at time k, is the position of the preceding vehicle, is the position of the vehicle to be controlled, is the acceleration of the vehicle to be controlled at time k, is the calculation time step.

[0058] (2) configuring the driving state strategy of the associated vehicle based on the driving condition of the vehicle to be controlled;

[0059] The driving state strategy of the associated vehicle includes a straight-ahead state, a lane-changing state, and a braking state; random disturbances are set in conjunction with the driving state strategy, generally noise.

[0060] In the driving state strategy, when the vehicle starts to move, the onboard sensors start working to identify whether there is a vehicle in front. If there is no vehicle in front or the distance between the vehicle in front and the vehicle to be controlled exceeds the sensor detection range, the vehicle enters the cruise mode. At this time, the vehicle's expected speed is the cruising speed set by the driver. If a vehicle is identified in front and the actual distance between the two vehicles is less than the sensor detection range, and the expected safety distance is less than the actual distance between the two vehicles, the vehicle maintains the cruise mode. If the expected safety distance is greater than the actual distance between the two vehicles and the speed of the vehicle in front is less than the cruising speed of the vehicle, the vehicle enters the following mode.

[0061] In this process, the straight-moving state, lane-changing state, and braking state are contents that are easily understood by those skilled in the art.

[0062] (3) Introducing an asymmetric Actor-Critic network mechanism and establishing a TD3 algorithm network prediction model based on the Markov decision process;

[0063] In this embodiment, the expected acceleration of the decision output is defined, the vehicle state space is established using physical quantities related to vehicle motion and dynamics, and the safety and efficiency of the following process are considered. The reward and penalty functions in the decision process are designed respectively. Specifically, establishing the prediction model includes the following steps:

[0064] S3.1 Establish the state space and action space of the vehicle to be controlled;

[0065] Among them, the state space S = [ d a c t , d d e s , v , v p ] , action space A = [ a ] , is the desired acceleration of the vehicle to be controlled.

[0066] S3.2 Design the reward and penalty functions for the multi-objective control of the vehicle to be controlled;

[0067] The reward and penalty function of the multi-objective control includes a safety reward function R1, a relative speed penalty R2, an acceleration smoothness penalty R3, a collision penalty R4, and a continuous following reward R5; the reward and penalty function calculation formula is: ;

[0068] In this embodiment, the safety reward function R1, relative speed penalty R2, acceleration smoothness penalty R3, collision penalty R4, and continuous following reward R5 are set to satisfy:

[0069] ,

[0070] ,

[0071] ,

[0072] ,

[0073] ,

[0074] Where, 、 、 、 、 is a constant, and in this embodiment, the values ​​are 0.2, 0.05, 1, 200, and 1 respectively. is the expected acceleration.

[0075] S3.3 introduces an asymmetric actor-critic architecture and establishes a prediction model consisting of one actor network and two critic networks.

[0076] In this embodiment, the Actor network is used for policy optimization, the Critic network is used for value estimation, and the Adam optimizer is used to independently optimize the Actor and Critic networks.

[0077] The Actor network consists of three fully connected hidden layers. Specifically, each layer contains 256 neurons, uses the ReLU activation function, inputs the current state space S, and outputs the continuous action value a. The tanh activation function is used, and the learning rate of the Actor network is The WOA algorithm is optimized and iterated before delivery;

[0078] The two critic networks each consist of four fully connected hidden layers. Specifically, each layer contains 256 neurons and uses the ReLU activation function. The input is a combination of the state space S and the action space A. The output of each critic network is a scalar Q value. The learning rate of the critic network is The WOA algorithm is optimized and passed after iteration.

[0079] (4) Iterative training is performed on the expected acceleration output using the priority experience replay mechanism and the WOA optimization algorithm;

[0080] In each training round, WOA is used to perform hyperparameter optimization and update the hyperparameters of the TD3 algorithm.

[0081] First, the present invention introduces a priority experience playback mechanism, specifically:

[0082] Calculate the TD error at each iteration, , is the instant reward obtained by the vehicle to be controlled at the current moment, is the value estimate of the action taken by the vehicle to be controlled at the next moment, It is the value estimate of the action taken by the vehicle to be controlled at the current moment. The larger the value, the higher the priority. is the discount factor;

[0083] Calculate sampling probability ,in, For samples Priority, is the priority strength, idk is the index, representing each experience in the priority replay buffer pool, p idx is the priority indexed by idx;

[0084] Defining importance sampling weights ,in, is the importance sampling correction factor, The size of the priority experience replay buffer pool.

[0085] Secondly, the present invention introduces WOA optimization iteration, which includes the following steps:

[0086] S4.1.1 Determine the TD3 hyperparameters to be optimized and the search range, as shown in Table 1 below;

[0087] Table 1 TD3 hyperparameters that need to be optimized and their search ranges

[0088]

[0089] S4.1.2 Initialize the WOA algorithm and set the maximum number of iterations Initialize the population to 50 (define the iteration counter T). And randomly generate a group; initialize the fitness of each whale , It is The total reward value of the iteration round, here is the total iteration round, that is, the cumulative reward value after training is used as the fitness function; initialize the global optimal solution and optimal fitness ;

[0090] S4.1.3 Evaluate the fitness of the current hyperparameter combination:

[0091] For every whale Use current hyperparameters Configure the improved TD3 algorithm model; calculate the fitness using the cumulative reward value, =Current cumulative reward value, if it is greater than the optimal fitness, that is , then update the global optimal solution , ;

[0092] S4.1.4 Update the current optimal solution;

[0093] The distance between the current value and the current optimal solution:

[0094] ,in, is the gap vector between the current value and the optimal solution, is the current optimal hyperparameter, is the position of the current hyperparameter individual;

[0095] Let the coefficient vector introducing randomness be , r ∈ [ 0 , 1 ] is a random vector used to increase the randomness of the search;

[0096]

[0097] Among them, the coefficient vector controlling the search step size , a r ∈ [ 0 , 2 ] , is the convergence constant that decreases linearly with iteration, l ∈ [ − 1 , 1 ] , is a random value, A randomly selected location.

[0098] After obtaining the optimal vector decoding, training the prediction model includes the following steps:

[0099] S4.1 Build a priority experience replay buffer pool to store interaction data, including experience and priority ,in, For status, For action, is the reward value, For the next state;

[0100] Use SumTree data to accelerate priority sampling and update storage interaction data;

[0101] Initialization: the optimal vector decoding after initial optimization of WOA is used as the hyperparameter of the prediction model;

[0102] The learning rate, target network update rate, discount factor, exploration strategy noise standard deviation and other hyperparameter values ​​of the prediction model are calculated and provided by WOA optimization iteration; the TD3 agent and Actor network of the initialization prediction model are and Critic Network 、 ; Initialize the target network of the Actor network and the Critic network 、 、 , and copy it as the current network parameters;

[0103] Get the initial state S 0 = [ d a c t , d d e s , v , v p ] ; Use Gaussian noise as exploration noise ,in, To explore the noise standard deviation, the calculated value comes from the WOA optimization iteration;

[0104] S4.2 In the current state, the Actor network Select and execute actions based on the current strategy , , get the next state and the reward and punishment function status ; Empirical data Stored in the priority experience replay buffer pool, where 、 Represent the state quantity at the current moment and the next moment respectively, 、 Represent the execution action and reward and punishment function at the current moment respectively; "experience data" refers to the data that can be sampled from the priority experience replay buffer pool based on sampling probability;

[0105] S4.3 Sampling a batch of experience data from the priority experience replay buffer pool , calculate the sampling probability according to the priority of each experience, the sampling probability Priority in the priority experience replay mechanism Relevance; Priority reflects the probability of an experience being selected in the replay pool, usually calculated based on the importance or novelty of the experience;

[0106] S4.4 Use the target policy network of the actor network to select the next action and calculate the target Q value based on the critic network;

[0107] Use the target actor network strategy to select the next action , , is the target strategy noise, which has a mean of 0 and a variance of Normal distribution , and is clipped to the range [-c, c], is the noise clipping range, usually [ − 3 , 2 ] , is the discount factor, which is provided by the calculated value after the WOA optimization iteration. Minimum update Actor network strategy, It consists of the current reward and the discounted minimum value, , thereby preventing the Q value (Q') of the Critic network from being too high;

[0108] S4.5 Update the critic network and calculate the loss of the current critic network. Based on the priority experience replay mechanism, calculate the importance sampling weight of each experience sample.

[0109] Use gradient descent to update the parameters of the Critic network 、 And calculate the loss of the current Critic network:

[0110]

[0111]

[0112] Calculate the importance sampling weight for each sampled experience ;

[0113] S4.6 updates the Actor network based on the preset delayed update rules. If the update conditions are met, the Actor network is updated once and the loss of the Actor network is calculated. Otherwise, proceed directly to S4.7.

[0114] The goal of the actor network is to maximize the Q value estimated by the critic network while exploiting uncertainty to optimize exploration. The critic network updates more frequently than the actor network, which can improve the stability of the actor network strategy. Generally speaking, the default rule for delayed updates of the actor network means that after every preset number of critic updates, the actor network is updated once and the loss of the actor network is calculated. Here, the "preset number" is 2.

[0115] Update Actor network parameters using gradient ascent And calculate the loss of the Actor network:

[0116] ;

[0117] in, The sample size used to draw a batch of samples from the priority experience replay buffer pool for calculation and update.

[0118] S4.7 Based on the calculated sampling probability , randomly select a set of experience data from the priority experience replay buffer pool , and update the experience priority according to the TD error, ,in, is a very small constant used for surface priority 0, is the current Q value of the Critic network;

[0119] S4.8 Update the target critic network;

[0120] Use soft update to update the target network parameters:

[0121]

[0122] Where, is the soft update coefficient, which is provided by the calculated value after WOA optimization iteration, usually ;

[0123] S4.9 When the set number of training rounds is reached or the performance meets the requirements, the training ends and the optimal expected acceleration is obtained. Otherwise, the WOA optimization iteration continues and steps S4.2 to S4.8 are repeated.

[0124] Ultimately, the controlled vehicle can autonomously learn how to use the strategy through the accumulated value of the reward and punishment function without human intervention.

[0125] (5) Longitudinal control of the vehicle to be controlled using the trained prediction model

[0126] The present invention also relates to a vehicle longitudinal control system based on the WOA-improved TD3 algorithm. The system includes a controller, a drive system, a transmission system and a sensing device configured on any vehicle to be controlled. The sensing device obtains the driving status of the associated vehicle of the vehicle to be controlled and transmits it to the controller. The controller adopts the vehicle longitudinal control method based on the WOA-improved TD3 algorithm to obtain the expected acceleration. The controller converts the expected acceleration into a control instruction. The drive system controls the acceleration and braking of the vehicle to be controlled based on the control instruction. The transmission system cooperates with the drive system to realize power transmission.

[0127] like Figure 3As shown in the figure, compared with the existing TD3 algorithm, the proposed vehicle longitudinal control method based on the WOA-improved TD3 algorithm significantly accelerates learning and quickly reaches convergence. Through multiple evaluations, the average distance error was 0.54m and the average collision rate was 0. This ensures that the following distance is close to the desired distance, effectively avoiding traffic congestion and fully utilizing the road while maintaining safe following.

[0128] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0129] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0130] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0131] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0132] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0133] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A vehicle longitudinal control method based on the WOA-improved TD3 algorithm, characterized by: The method establishes a longitudinal dynamics model and configures a driving state strategy for the associated vehicle based on the driving condition of the vehicle to be controlled; introduces an asymmetric Actor-Critic network mechanism and establishes a TD3 algorithm network prediction model based on a Markov decision process. Establishing the prediction model includes the following steps: S3.1 Establish the state space and action space of the vehicle to be controlled; S3.2 Design the reward and penalty functions for the multi-objective control of the vehicle to be controlled; S3.3 introduces an asymmetric actor-critic architecture and establishes a prediction model consisting of an actor network and two critic networks; The expected acceleration output is iteratively trained using a priority experience replay mechanism and a WOA optimization algorithm. Training the prediction model includes the following steps: S4.1 Build a priority experience replay buffer pool to store interaction data, initialize it, and use the optimal vector decoding after initial WOA optimization as the hyperparameter of the prediction model; obtain the initial state , where d act is the distance between the vehicle to be controlled and the preceding vehicle, d des is the expected safety distance, v is the speed of the vehicle to be controlled, v p is the speed of the preceding vehicle; S4.2 In the current state, the Actor network Select and execute actions based on the current strategy , obtain the next state and reward and punishment function state; store the experience data into the priority experience replay buffer pool; S4.3 Sampling a batch of experience data from the priority experience replay buffer pool, sampling probability Priority in the priority experience replay mechanism Related; S4.4 Use the target policy network of the actor network to select the next action and calculate the target Q value based on the critic network; S4.5 Update the critic network and calculate the loss of the current critic network. Based on the priority experience replay mechanism, calculate the importance sampling weight of each experience sample. S4.6 Based on the preset delayed update rules of the Actor network, if the update conditions are met, the Actor network is updated once and the loss of the Actor network is calculated. Otherwise, proceed directly to S4.7; S4.7 Based on the calculated sampling probability , randomly select a set of experience data from the priority experience replay buffer pool , and update the experience priority according to the TD error, where S i is the state of the current sample, a i is the execution action of the current sample, R i is the reward and punishment function of the current sample, S i+1 is the state of the next sample; S4.8 Update the target critic network; S4.9 When the set number of training rounds is reached or the performance meets the requirements, the training ends and the optimal expected acceleration is obtained. Otherwise, the WOA optimization iteration continues, and steps S4.2 to S4.8 are repeated. The trained prediction model is used to perform longitudinal control on the vehicle to be controlled.

2. The vehicle longitudinal control method based on the WOA-improved TD3 algorithm according to claim 1, characterized in that: The associated vehicles include the preceding vehicle and the adjacent vehicles of the vehicle to be controlled. The longitudinal dynamics model definition satisfies: , in, is the distance error between the vehicle to be controlled and the preceding vehicle at time k, is the distance between the vehicle to be controlled and the preceding vehicle at time k, is the expected safety distance at time k, is the speed difference between the vehicle to be controlled and the preceding vehicle at time k, is the speed of the vehicle to be controlled at time k, is the speed of the preceding vehicle at time k, is the position of the preceding vehicle, is the position of the vehicle to be controlled, is the acceleration of the vehicle to be controlled at time k, is the calculation time step.

3. The vehicle longitudinal control method based on the WOA-improved TD3 algorithm according to claim 2, characterized in that: The driving state strategy of the associated vehicle includes a straight-ahead state, a lane-changing state, and a braking state; Random disturbances are set in conjunction with the driving state strategy.

4. The vehicle longitudinal control method based on the WOA-improved TD3 algorithm according to claim 1, characterized in that: The reward and penalty function of the multi-objective control includes a safety reward function, a relative speed penalty, an acceleration smoothness penalty, a collision penalty and a continuous following reward.

5. The vehicle longitudinal control method based on the WOA-improved TD3 algorithm according to claim 1, characterized in that: The Actor network includes three fully connected hidden layers, and the learning rate of the Actor network is The WOA algorithm is optimized and iterated before delivery; The two critic networks each include four fully connected hidden layers, and the learning rate of the critic network is The WOA algorithm is optimized and passed after iteration.

6. The vehicle longitudinal control method based on the WOA-improved TD3 algorithm according to claim 1, characterized in that: Sampling probability satisfy ,in, For samples Priority, is the priority strength, idk is the index, representing each experience in the priority replay buffer pool, p idk is the priority indexed by idk; Importance sampling weights satisfy ,in, is the importance sampling correction factor, The size of the priority experience replay buffer pool; TD error satisfies ,in, is the instant reward obtained by the vehicle to be controlled at the current moment, is the value estimate of the action taken by the vehicle to be controlled at the next moment, is the value estimate of the action taken by the vehicle to be controlled at the current moment, is the discount factor.

7. The vehicle longitudinal control method based on the WOA-improved TD3 algorithm according to claim 6, characterized in that: The WOA optimization iteration includes the following steps: S4.1.1 Determine the hyperparameters to be optimized and the search range; S4.1.2 Initialize the WOA algorithm, set the maximum number of iterations, initialize the population and randomly generate a group; initialize the fitness of each whale, using the cumulative reward value after training as the fitness function; initialize the global optimal solution and optimal fitness; S4.1.3 For each whale, use the modified TD3 algorithm model with the current hyperparameter configuration; calculate the fitness using the cumulative reward value, and update the global optimal solution if it is greater than the optimal fitness; S4.1.4 Update the current optimal solution.

8. A vehicle longitudinal control system based on the WOA-improved TD3 algorithm, characterized by: The system includes a controller, a drive system, a transmission system and a sensing device configured on any vehicle to be controlled. The sensing device obtains the driving status of the associated vehicle of the vehicle to be controlled and transmits it to the controller. The controller adopts the vehicle longitudinal control method based on the WOA-improved TD3 algorithm described in one of claims 1 to 7 to obtain the expected acceleration. The controller converts the expected acceleration into a control instruction. The drive system controls the acceleration and braking of the vehicle to be controlled based on the control instruction. The transmission system cooperates with the drive system to realize power transmission.

Citation Information

Patent Citations

  • Intelligent vehicle driving method based on reinforcement learning

    CN117270394A

  • TCN-BiLSTM-WOA-based vehicle bridge crossing safety evaluation method under strong wind

    CN118761335A