Personalized control method for transverse and longitudinal cooperation of intelligent vehicle based on reinforcement learning
By introducing a horizontal and vertical collaborative personalized control method based on reinforcement learning in intelligent vehicle control, combining vehicle dynamics model and driving style differentiation, the problems of flexibility and personalized control in the existing technology are solved, and higher control accuracy and stability are achieved.
Patent Information
- Application Number
- CN202510324309.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-06
AI Technical Summary
The existing intelligent vehicle control methods are difficult to provide sufficient flexibility and real-time in dynamic environments and complex driving scenarios, and lack consideration of driving style, resulting in a lack of personalization and diversity in control effects.
A personalized control method based on reinforcement learning based on intelligent vehicle horizontal and vertical coordination is adopted, combining vehicle dynamics model, dynamic safety boundary constraints and differentiated driving styles, a strategy network and a value network are built, and the optimal control model is obtained through training to achieve vehicle path tracking control.
It improves the safety, stability and diversity of driving styles of intelligent vehicle motion control, enhances control accuracy and accuracy, and meets diverse control needs.
Smart Images

Figure CN120096597A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent vehicle motion control, and specifically relates to a personalized control method for lateral and longitudinal coordination of intelligent vehicles based on reinforcement learning. Background Art
[0002] With the continuous development of autonomous driving technology, improving the control accuracy and stability of intelligent vehicles has become one of the research focuses. Existing autonomous driving control methods mainly include the classical control theory algorithm (PID) commonly used for longitudinal control, the linear quadratic regulator (LQR) commonly used for lateral control, and the model-based predictive control (MPC) method. Although these methods perform well in known environments, their limitations are also obvious. When facing dynamic environments or complex driving scenarios, it is difficult to provide sufficient flexibility and real-time performance. In addition, these traditional control methods lack consideration of driving style and are difficult to meet the personalized driving needs of different drivers, resulting in a lack of personalization and diversity in the control effect of the vehicle.
[0003] Reinforcement learning (RL) does not require any prior data collection for its self-learning optimization model. It usually combines the advantages of deep networks in processing high-dimensional data and has strong generalization ability. It can adapt to complex and changing urban road environments and has been widely used in autonomous driving in recent years. By learning the optimal control strategy based on environmental interaction, it has higher adaptability in different driving environments and scenarios. However, most of the existing intelligent vehicle control methods based on reinforcement learning do not take into account the dynamic factors of the vehicle and the diversity of driving styles, resulting in the need to further improve its effectiveness and stability in practical applications. Summary of the invention
[0004] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides a personalized control method for the lateral and longitudinal coordination of intelligent vehicles based on reinforcement learning, in order to effectively improve the safety, stability and driving style diversity of intelligent vehicle motion control by combining vehicle dynamics models, dynamic safety boundary constraints and differentiated driving styles, thereby improving the precision and accuracy of control and meeting diverse control needs.
[0005] To achieve the above object, the present invention adopts the following technical solution:
[0006] The personalized control method of the intelligent vehicle lateral and longitudinal coordination based on reinforcement learning of the present invention is characterized in that it includes the following steps:
[0007] Step 1: Establish a global coordinate system with the initial position of the vehicle in the scene as the origin, east as the X axis, and north as the Y axis, and obtain vehicle posture information and target path information in the global coordinate system;
[0008] The vehicle body coordinate system is established with the vehicle center of mass as the origin, the vehicle forward direction as the x-axis, the left side of the vehicle as the y-axis, and the vertical ground direction as the z-axis. The vehicle speed information is obtained in the vehicle body coordinate system to construct the state of the intelligent body. ,in, is the horizontal coordinate of the vehicle's center of mass in the global coordinate system, is the ordinate of the vehicle’s center of mass in the global coordinate system, is the heading angle of the vehicle in the global coordinate system, is the desired heading angle, is the longitudinal speed, is the lateral speed, is the target longitudinal speed, is the front wheel turning angle, is the yaw angular velocity, is the desired yaw rate, is the sideslip angle of the center of mass, is the desired sideslip angle of the center of mass, Any point on the vehicle centerline lateral distance to the target path;
[0009] Step 2: Get the vehicle's control parameters and use them to build the agent's actions ,in, is the steering wheel angle, For the The braking torque of each wheel, , is the throttle opening;
[0010] Step 3: According to the vehicle dynamics model, by limiting the yaw rate and the center of mass slip angle , constructing the vehicle dynamics safety boundary ;
[0011] Step 4: Divide the driving styles into conservative, normal, and radical. According to the different driving styles, vehicle dynamics safety boundary constraints B, and target trajectories of tracking control, construct the reward function of any driving style of the agent. ;in, Reward function indicating a conservative driving style, Reward function for normal driving style, The reward function for driving style as radical;
[0012] Step 5: Construct a policy network and a value network, and train the policy network and the value network based on the reward functions of the three driving styles to obtain an optimal control model for processing the state of the vehicle at the current time step, outputting the action of the current time step as a control instruction for vehicle path tracking.
[0013] The personalized control method for lateral and longitudinal coordination of intelligent vehicles based on reinforcement learning described in the present invention is also characterized in that: and Using equation (5) and equation (6) respectively, we can get:
[0014] (5)
[0015] (6)
[0016] In formula (5) and formula (6), represents the vehicle wheelbase, and , , are the distances from the front axle and rear axle to the center of mass, represents the vehicle stability factor, and , , are the cornering stiffness of the front and rear tires, Indicates the mass of the vehicle;
[0017] Furthermore, the step 3 comprises:
[0018] Step 3.1: Use equations (7) and (8) to get the front tire slip angle and rear tire slip angle :
[0019] (7)
[0020] (8)
[0021] Step 3.2: According to the range of tire slip angle [- , ], and use formula (9) to construct the center of mass sideslip angle Safe area:
[0022] (9)
[0023] In formula (9), Indicates the upper limit of the tire slip angle;
[0024] Step 3.3: Use equation (10) to get the yaw rate Safe area:
[0025] (10)
[0026] Step 3.4: A closed envelope is formed by equations (9) and (10) and used as a vehicle dynamics safety constraint , when the vehicle's yaw rate and the center of mass slip angle When it is within the envelope, it means that the vehicle meets the dynamic safety boundary constraints. .
[0027] Furthermore, in step 4, the reward function is constructed using formula (11): :
[0028] (11)
[0029] In formula (11), is the tracking error term, and is obtained by equation (12); is the stability error term and is obtained by formula (13); is the dynamic safety boundary term and is obtained by formula (14); is the control quantity smoothing term, and is obtained by formula (15);
[0030] (12)
[0031] In formula (12), , , , is the adjustment coefficient of each sub-item; is the position deviation, and ; , , is the offset influence factor of each position, and , ; is the heading angle deviation, and ; is the speed deviation, and ;
[0032] (13)
[0033] In formula (13), is the influence term of yaw angular velocity, is the influence term of the sideslip angle at the center of mass;
[0034] (14)
[0035] In formula (14), , are the two reward values constrained by the safety boundary;
[0036] (15)
[0037] In formula (15), , , is the influencing factor of the smoothing term of each control quantity output, is the difference in steering wheel angles between adjacent time steps, is the difference in throttle opening between adjacent time steps, is the interval between adjacent time steps The difference in braking torque of each wheel;
[0038] in, , , Different values are taken when corresponding to different driving styles, that is, the value when the driving style is cautious is smaller than the value when the driving style is normal, which is smaller than the value when the driving style is aggressive;
[0039] , , , , , , The values are different when corresponding to different driving styles, that is, the value when the driving style is cautious is greater than the value when the driving style is normal, which is greater than the value when the driving style is aggressive.
[0040] Further, the step 5 comprises:
[0041] Step 5.1: Randomly initialize the parameters of the policy network and the parameters of the value network ;
[0042] Step 5.2: Set the capacity of the current round of experience playback buffer to , the current step The state of The input is processed in the policy network and the current step is output. Next action , when the vehicle performs an action Then, the reward function of any driving style in step 4 is used to calculate the current step Rewards , and get the next step Down state , and the current step The end sign , so that the experience tuple Recorded as Samples are stored in the current round of experience playback buffer;
[0043] Step 5.3: If the number of samples in the current round of experience replay buffer reaches capacity , then execute step 5.4; otherwise, return to step 5.2 and repeat the sampling operation;
[0044] Step 5.4: Use equation (16) to calculate the advantage function value of any t-th sample in the current round of experience replay buffer: :
[0045] (16)
[0046] In formula (16), is the discount factor, For the current round of value network The estimated value of For the current round of value network An estimate of the value of
[0047] Step 5.5: Use formula (17) to calculate the strategy ratio value of the tth sample in the current round of experience replay buffer: :
[0048] (17)
[0049] In formula (17), The current round strategy exist Next Select The probability of Indicates the previous round strategy exist Next Select The probability of, when the iteration is the first round, the parameters of the previous round of policy network ;
[0050] Step 5.6: Use equation (18) to calculate the target value of the tth sample in the current round of experience replay buffer: :
[0051] (18)
[0052] Step 5.7: Use formula (19) to calculate the current round of policy network pruning objective function ;
[0053] (19)
[0054] In formula (19), is the clipping factor, Indicates that Restricted to Within the interval, Indicates taking a smaller value. It means to find the expectation of all samples in the current round of experience playback buffer;
[0055] Step 5.8: Use formula (20) to calculate the mean square error of the current round value network :
[0056] (20)
[0057] Step 5.9: Use formula (21) to get the parameters of the next round of policy network :
[0058] (twenty one)
[0059] In formula (21), is the learning rate of the policy network, for The gradient of
[0060] Step 5.10: Use formula (22) to get the parameters of the next round of value network :
[0061] (twenty two)
[0062] In formula (22), is the learning rate of the value network, for The gradient of
[0063] Step 5.11: , Assign values to , After that, return to step 5.2 and execute sequentially until the maximum number of iterations is reached, thereby obtaining the optimal parameters of the policy network and the optimal parameters of the value network The optimal path following control model for any driving style.
[0064] The present invention provides an electronic device, comprising a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the generation method, and the processor is configured to execute the program stored in the memory.
[0065] The present invention provides a computer-readable storage medium, on which a computer program is stored, characterized in that the computer program executes the steps of the generation method when executed by a processor.
[0066] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0067] 1. The present invention uses vehicle dynamics simulation software equipped with a high-precision vehicle dynamics model as a training environment, and incorporates the constraints of the dynamic safety boundary into the design of the reward function, ensuring that the reinforcement learning algorithm can guarantee the vehicle's handling stability under various working conditions, thereby improving the accuracy and stability of path tracking control.
[0068] 2. The present invention divides driving styles and constructs a differentiated model based on driving styles, so that the control algorithm reflects differentiated characteristics, achieves personalized control effects, meets diverse driving needs, and improves the user's driving experience; at the same time, the steering wheel angle, throttle opening, and wheel braking torque are used as the output of the reinforcement learning controller, and the high-precision lateral and longitudinal coordinated control effect of the vehicle is achieved by training the algorithm model. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 It is a flow chart of a personalized control method for lateral and longitudinal coordination of an intelligent vehicle based on reinforcement learning of the present invention;
[0070] Figure 2 The dynamic boundary diagram of the personalized control method for lateral and longitudinal coordination of intelligent vehicles based on reinforcement learning of the present invention;
[0071] Figure 3 A network training logic diagram of a personalized control method for horizontal and vertical coordination of intelligent vehicles based on reinforcement learning according to the present invention;
[0072] Figure 4 This is a training path diagram for the personalized control method for lateral and longitudinal coordination of intelligent vehicles based on reinforcement learning of the present invention. DETAILED DESCRIPTION
[0073] In this embodiment, a personalized control method for lateral and longitudinal coordination of intelligent vehicles based on reinforcement learning is applied to the field of intelligent vehicle path tracking control, and a control model that can realize lateral and longitudinal coordinated control and distinguish driving styles can be trained, such as Figure 1 As shown, the following steps are included:
[0074] Step 1: Establish a global coordinate system with the initial position of the vehicle in the scene as the origin, east as the X-axis, and north as the Y-axis, and obtain the vehicle posture information and target path information in the global coordinate system; establish a body coordinate system with the center of mass of the vehicle as the origin, the vehicle's forward direction as the x-axis, the left side of the vehicle as the y-axis, and the vertical ground direction as the z-axis, and obtain the vehicle speed information in the body coordinate system; the design of the intelligent body state first requires task analysis. Taking this embodiment as an example, the main goal is to achieve the task of path tracking control with high precision, and the sub-goals include speed tracking and body stability control; then filter the relevant information. The shorter the time in which the event represented by a certain state information is fed back, the easier it is for the neural network to learn how to process it and establish decision correlation. After comprehensive consideration, the state of the intelligent body is finally selected: ,in, is the horizontal coordinate of the vehicle's center of mass in the global coordinate system, is the ordinate of the vehicle’s center of mass in the global coordinate system, is the heading angle of the vehicle in the global coordinate system, is the desired heading angle, is the longitudinal speed, is the lateral speed, is the target longitudinal speed, is the front wheel turning angle, is the yaw angular velocity, is the desired yaw rate, is the sideslip angle of the center of mass, is the desired sideslip angle of the center of mass, Any point on the vehicle centerline The lateral distance to the target path.
[0075] Among them, the derivation process of the expected yaw rate and the expected sideslip angle of the center of mass is:
[0076] First, the two-degree-of-freedom dynamics model of the vehicle is established using equations (1) and (2):
[0077] (1)
[0078] (2)
[0079] In formula (1) and formula (2), is the total lateral force, is the lateral force on the front wheel, is the rear wheel lateral force; is the total moment about the z-axis.
[0080] Then, using equations (3) and (4), we get the differential equation of vehicle motion:
[0081] (3)
[0082] (4)
[0083] In formula (3) and formula (4), , are the distances from the front axle and rear axle to the center of mass, , are the cornering stiffness of the front and rear tires, Indicates the vehicle mass, represents the lateral acceleration, is the moment of inertia about the z-axis.
[0084] When the vehicle is in a steady state, Substituting into (3) and (4) we can obtain and :
[0085] (5)
[0086] (6)
[0087] Formula (5) and formula (6), represents the vehicle wheelbase, and , represents the vehicle stability factor, and .
[0088] Step 2: Get the vehicle's control parameters and use them to build the agent's actions ,in, is the steering wheel angle, For the The braking torque of each wheel, , is the throttle opening.
[0089] Step 3: According to the vehicle dynamics model, by limiting the yaw rate and the center of mass slip angle , constructing the vehicle dynamics safety boundary ;
[0090] Step 3.1: Using the small angle assumption, use equations (7) and (8) to obtain the front tire slip angle: and rear tire slip angle :
[0091] (7)
[0092] (8)
[0093] Step 3.2: According to the range of tire slip angle [- , ],Will Substituting into equation (8), we can get the center of mass side slip angle of equation (9): Safe area:
[0094] (9)
[0095] In formula (9), Indicates the upper limit of the tire slip angle.
[0096] Step 3.3: According to the moment balance equation And the lateral acceleration expression under the small angle assumption , we can get the rear wheel lateral force as: , when the rear wheel reaches the maximum lateral force, , and the yaw rate is obtained by combining equation (10) Safe area:
[0097] (10)
[0098] Step 3.4: Use equations (9) and (10) to form a closed envelope, such as Figure 2 As shown, AD and BC are the constraints of formula (9), AB and DC are the constraints of formula (10). When the yaw rate of the vehicle is and the center of mass slip angle When it is within the envelope, it means that the vehicle meets the dynamic safety boundary constraints. .
[0099] Step 4: Divide the driving styles into conservative, normal, and radical. According to the different driving styles, vehicle dynamics safety boundary constraints B, and target trajectories of tracking control, construct the reward function of any driving style of the agent. ;in, Reward function indicating a conservative driving style, Reward function for normal driving style, The reward function for driving style as radical is constructed using formula (11):
[0100] (11)
[0101] In formula (11), is the tracking error term, and is obtained by equation (12); is the stability error term and is obtained by formula (13); is the dynamic safety boundary term and is obtained by formula (14); is the control quantity smoothing term, and is obtained by formula (15);
[0102] (12)
[0103] In formula (12), , , , is the adjustment coefficient of each sub-item; is the position deviation, and ; , , is the offset influence factor of each position, and , ; is the heading angle deviation, and ; is the speed deviation, and ;
[0104] (13)
[0105] In formula (13), is the influence term of yaw angular velocity, is the influence term of the sideslip angle at the center of mass;
[0106] (14)
[0107] In formula (14), , are two reward values constrained by the safety boundary; the negative reward outside the boundary is set to be larger than the absolute value of the positive reward inside the boundary. Specifically, The purpose is to avoid the vehicle being in a critical state and performing "greedy" reward-seeking behavior. At the same time, the size of the defined safety margin is the largest for the aggressive type, the second for the normal type, and the smallest for the cautious type.
[0108] (15)
[0109] In formula (15), , , is the influencing factor of the smoothing term of each control quantity output, is the difference in steering wheel angles between adjacent time steps, is the difference in throttle opening between adjacent time steps, is the interval between adjacent time steps The difference in braking torque of the wheels.
[0110] The cautious driving style has the highest requirements for vehicle safety and stability, the highest requirements for control output smoothness, and the lowest requirements for tracking accuracy; the aggressive driving style has the lowest requirements for vehicle safety and stability, the lowest requirements for control output smoothness, and requires accurate tracking of the path; the normal driving style is between the two. Based on this, the design results of each coefficient are as follows:
[0111] in, , , Different values are taken when corresponding to different driving styles, that is, the value when the driving style is cautious is smaller than the value when the driving style is normal, which is smaller than the value when the driving style is aggressive;
[0112] , , , , , , The values are different when corresponding to different driving styles, that is, the value when the driving style is cautious is greater than the value when the driving style is normal, which is greater than the value when the driving style is aggressive.
[0113] Step 5: Construct a policy network and a value network, and train the policy network and the value network based on the reward functions of the three driving styles to obtain an optimal control model for processing the state of the vehicle at the current time step, outputting the action of the current time step as a control instruction for vehicle path tracking.
[0114] The logic diagram of the training network is as follows Figure 3 As shown in the figure, the training environment used is vehicle dynamics simulation software. The training is achieved by linking the simfle file generated by the vehicle dynamics simulation software with the pycarsimlib experimental package under the pytorch framework of deep reinforcement learning. The establishment of the intelligent agent training environment also includes defining vehicle structural parameters and obtaining a variety of road information such as lane changing, lane keeping, and turning for vehicle path tracking control. Figure 4 This is one of the paths used for training; the specific process of training the network is:
[0115] Step 5.1: The policy network and value network consist of an input layer, a hidden layer, and an output layer; the hidden layer activation function is ReLU, and the output layer variable The activation function is Tanh, and the output layer variable and The activation function is Sigmoid; the parameters of the policy network are randomly initialized and the parameters of the value network ;
[0116] Step 5.2: Set the capacity of the current round of experience playback buffer to , through the interaction between the vehicle and the training environment, specifically, the current step The state of The input is processed in the policy network and the current step is output. Next action , when the vehicle performs an action Then, the reward function of any driving style in step 4 is used to calculate the current step Rewards , and get the next step Down state , and the current step The end sign , if the vehicle's current step The lateral distance from the target path exceeds the set value e. This ends the current round and the environment is reset to the initial state for resampling; otherwise , so that the experience tuple Recorded as Samples are stored in the current round of experience playback buffer;
[0117] Step 5.3: If the number of samples in the current round of experience replay buffer reaches capacity , then execute step 5.4; otherwise, return to step 5.2 and repeat the sampling operation;
[0118] Step 5.4: Use equation (16) to calculate the advantage function value of any t-th sample in the current round of experience replay buffer: :
[0119] (16)
[0120] In formula (16), is the discount factor, For the current round of value network The estimated value of For the current round of value network An estimate of the value of
[0121] Step 5.5: Use formula (17) to calculate the strategy ratio value of the tth sample in the current round of experience replay buffer: :
[0122] (17)
[0123] In formula (17), The current round strategy exist Next Select The probability of Indicates the previous round strategy exist Next Select The probability of, when the iteration is the first round, the parameters of the previous round of policy network ;
[0124] Step 5.6: Use equation (18) to calculate the target value of the tth sample in the current round of experience replay buffer: :
[0125] (18)
[0126] Step 5.7: Use formula (19) to calculate the current round of policy network pruning objective function ;
[0127] (19)
[0128] In formula (19), is the clipping factor, Indicates that Restricted to Within the interval, Indicates taking a smaller value. It means to find the expectation of all samples in the current round of experience playback buffer;
[0129] Step 5.8: Use formula (20) to calculate the mean square error of the current round value network :
[0130] (20)
[0131] Step 5.9: Use formula (21) to get the parameters of the next round of policy network :
[0132] (twenty one)
[0133] In formula (21), is the learning rate of the policy network, for ; In the implementation process, in order to reduce the gradient calculation error, the gradient needs to be averaged; Specifically, the n samples in the current round of experience replay buffer are randomly divided into small batches with b samples in each batch and a total of c batches; First, the policy network clipping objective function of each small batch of the current round is calculated using formula (19), and the gradient of the policy network clipping objective function of each small batch is calculated and then the average value of c batches is taken, and the gradient mean of each small batch of the current round is used to replace the gradient of the entire buffer of the current round, and then it is substituted into (21) to calculate the policy network parameters of the next round.
[0134] Step 5.10: Use formula (22) to get the parameters of the next round of value network :
[0135] (twenty two)
[0136] In formula (22), is the learning rate of the value network, for Similarly, the gradient of the mean square error of the value network is also averaged, and the processing method is the same as that of the policy network.
[0137] Step 5.11: , Assign values to , After that, return to step 5.2 and execute sequentially until the maximum number of iterations is reached, thereby obtaining the optimal parameters of the policy network and the optimal parameters of the value network The optimal path following control model for any driving style.
[0138] To improve the generalization ability of the model, fixed road training is used in the early stage of training; in the middle stage of training, the model that has performed well is loaded in a new environment and continued to be trained; in the late stage of training, the model is trained using randomly initialized road information to improve the adaptability of the model.
[0139] In this embodiment, an electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0140] In this embodiment, a computer-readable storage medium stores a computer program on the computer-readable storage medium, and the computer program executes the steps of the above method when executed by a processor.
Claims
1. A personalized control method for lateral and longitudinal coordination of intelligent vehicles based on reinforcement learning, characterized in that: The steps include: Step 1: Establish a global coordinate system with the initial position of the vehicle in the scene as the origin, east as the X axis, and north as the Y axis, and obtain vehicle posture information and target path information in the global coordinate system; The vehicle body coordinate system is established with the vehicle center of mass as the origin, the vehicle forward direction as the x-axis, the left side of the vehicle as the y-axis, and the vertical ground direction as the z-axis. The vehicle speed information is obtained in the vehicle body coordinate system to construct the state of the intelligent body. ,in, is the horizontal coordinate of the vehicle's center of mass in the global coordinate system, is the ordinate of the vehicle’s center of mass in the global coordinate system, is the heading angle of the vehicle in the global coordinate system, is the desired heading angle, is the longitudinal speed, is the lateral speed, is the target longitudinal speed, is the front wheel turning angle, is the yaw angular velocity, is the desired yaw rate, is the sideslip angle of the center of mass, is the desired sideslip angle of the center of mass, Any point on the vehicle centerline lateral distance to the target path; Step 2: Get the vehicle's control parameters and use them to build the agent's actions ,in, is the steering wheel angle, For the The braking torque of each wheel, , is the throttle opening; Step 3: According to the vehicle dynamics model, by limiting the yaw rate and the center of mass slip angle , constructing the vehicle dynamics safety boundary ; Step 4: Divide the driving styles into conservative, normal, and radical. According to the different driving styles, vehicle dynamics safety boundary constraints B, and target trajectories of tracking control, construct the reward function of any driving style of the agent. ;in, Reward function indicating a conservative driving style, Reward function for normal driving style, The reward function for driving style as radical; Step 5: Construct a policy network and a value network, and train the policy network and the value network based on the reward functions of the three driving styles to obtain an optimal control model for processing the state of the vehicle at the current time step, outputting the action of the current time step as a control instruction for vehicle path tracking.
2. According to claim 1, a personalized control method for lateral and longitudinal coordination of intelligent vehicles based on reinforcement learning is characterized by: Step 1 and Using equation (5) and equation (6) respectively, we can get: (5) (6) In formula (5) and formula (6), represents the vehicle wheelbase, and , , are the distances from the front axle and rear axle to the center of mass, represents the vehicle stability factor, and , , are the cornering stiffness of the front and rear tires, Indicates the mass of the vehicle.
3. The personalized control method for lateral and longitudinal coordination of intelligent vehicles based on reinforcement learning according to claim 1 is characterized by: The step 3 comprises: Step 3.1: Use equations (7) and (8) to get the front tire slip angle and rear tire slip angle : (7) (8) Step 3.2: According to the range of tire slip angle [- , ], and use formula (9) to construct the center of mass sideslip angle Safe Area: (9) In formula (9), Indicates the upper limit of the tire slip angle; Step 3.3: Use equation (10) to get the yaw rate Safe Area: (10) Step 3.4: A closed envelope is formed by equations (9) and (10) and used as a vehicle dynamics safety constraint , when the vehicle's yaw rate and the center of mass slip angle When it is within the envelope, it means that the vehicle meets the dynamic safety boundary constraints. .
4. The personalized control method for lateral and longitudinal coordination of intelligent vehicles based on reinforcement learning according to claim 1 is characterized by: In step 4, the reward function is constructed using formula (11): : (11) In formula (11), is the tracking error term, and is obtained by equation (12); is the stability error term and is obtained by equation (13); is the dynamic safety boundary term and is obtained by formula (14); is the control quantity smoothing term, and is obtained by formula (15); (12) In formula (12), , , , is the adjustment coefficient of each sub-item; is the position deviation, and ; , , is the offset influence factor of each position, and , ; is the heading angle deviation, and ; is the speed deviation, and ; (13) In formula (13), is the influence term of yaw angular velocity, is the influence term of the sideslip angle at the center of mass; (14) In formula (14), , are the two reward values constrained by the safety boundary; (15) In formula (15), , , is the influencing factor of the smoothing term of each control quantity output, is the difference in steering wheel angles between adjacent time steps, is the difference in throttle opening between adjacent time steps, is the interval between adjacent time steps The difference in braking torque of each wheel; in, , , Different values are taken when corresponding to different driving styles, that is, the value when the driving style is cautious is smaller than the value when the driving style is normal, which is smaller than the value when the driving style is aggressive; , , , , , , The values are different when corresponding to different driving styles, that is, the value when the driving style is cautious is greater than the value when the driving style is normal, which is greater than the value when the driving style is aggressive.
5. The personalized control method for lateral and longitudinal coordination of intelligent vehicles based on reinforcement learning according to claim 1 is characterized by: The step 5 comprises: Step 5.1: Randomly initialize the parameters of the policy network and the parameters of the value network ; Step 5.2: Set the capacity of the current round experience playback buffer to , the current step The state of The input is processed in the policy network and the current step is output. Next action , when the vehicle performs an action Then, the reward function of any driving style in step 4 is used to calculate the current step Rewards , and get the next step Down state , and the current step The end sign , so that the experience tuple Recorded as Samples are stored in the current round of experience playback buffer; Step 5.3: If the number of samples in the current round of experience replay buffer reaches capacity , then execute step 5.4; otherwise, return to step 5.2 and repeat the sampling operation; Step 5.4: Use equation (16) to calculate the advantage function value of any t-th sample in the current round of experience replay buffer: : (16) In formula (16), is the discount factor, For the current round of value network The estimated value of For the current round of value network An estimate of the value of Step 5.5: Use formula (17) to calculate the strategy ratio value of the tth sample in the current round of experience replay buffer: : (17) In formula (17), The current round strategy exist Next Select The probability of Indicates the previous round strategy exist Next Select The probability of, when the iteration is the first round, the parameters of the previous round of policy network ; Step 5.6: Use equation (18) to calculate the target value of the tth sample in the current round of experience replay buffer: : (18) Step 5.7: Use formula (19) to calculate the current round of policy network pruning objective function ; (19) In formula (19), is the clipping factor, Indicates that Restricted to Within the interval, Indicates taking a smaller value. It means to find the expectation of all samples in the current round of experience playback buffer; Step 5.8: Use formula (20) to calculate the mean square error of the current round value network : (20) Step 5.9: Use formula (21) to get the parameters of the next round of policy network : (21) In formula (21), is the learning rate of the policy network, for The gradient of Step 5.10: Use formula (22) to get the parameters of the next round of value network : (22) In formula (22), is the learning rate of the value network, for The gradient of Step 5.11: , Assign values to , After that, return to step 5.2 and execute sequentially until the maximum number of iterations is reached, thereby obtaining the optimal parameters of the policy network and the optimal parameters of the value network The optimal path following control model for any driving style.
6. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports a processor to execute the generation method described in any one of claims 1 to 5, and the processor is configured to execute the program stored in the memory.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of any generation method described in claims 1-5 are executed.
Citation Information
Cited By
Reinforcement learning-based low-attachment road surface four-wheel drive torque distribution method
CN120863365A
Two-wheeled vehicle automatic driving road condition decision-making method and device based on near-end strategy optimization
CN121671669A
A method and device for autonomous driving of two-wheeled vehicles based on near-end strategy optimization
CN121671669B