A longitudinal decision-making method for autonomous driving vehicles based on model-based reinforcement learning
Through the vertical decision-making method of autonomous driving vehicles based on model reinforcement learning, combined with the vehicle longitudinal and full-vehicle vertical dynamic models, intelligent speed planning is achieved, which solves the problem that electric vehicles are difficult to achieve energy saving, efficiency, comfort and safety in complex traffic environments, especially the comfort and energy consumption problems when following the vehicle on uneven roads.
Patent Information
- Application Number
- CN202410753949.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-12
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-06-12
AI Technical Summary
The prior art is difficult to achieve an energy-saving, efficient, comfortable and safe driving experience of electric vehicles in complex traffic environments, especially the problem of reduced comfort and excessive energy consumption when following the vehicle on uneven roads.
The vertical decision-making method of autonomous driving vehicles based on model reinforcement learning is adopted. By establishing a vehicle longitudinal dynamic model and a full-vehicle vertical dynamic model, combining inter-vehicle communication technology to obtain the driving data of the vehicle in front, determine the minimum safe distance following the vehicle, and update the model through greedy actions to achieve intelligent speed planning.
Effectively balance energy saving and comfort, achieve a more efficient, safe and pleasant driving experience, and solve the problems of reduced driving comfort and excessive energy consumption caused by uneven roads.
Smart Images

Figure CN118770280B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving decision planning, and in particular to a longitudinal decision-making method for an autonomous driving vehicle based on model-based reinforcement learning. Background Art
[0002] In the context of the current global energy crisis and increasingly serious environmental pollution, electric vehicles (EVs) have attracted widespread attention from countries around the world as one of the sustainable transportation solutions. With the continuous advancement of battery technology and the gradual reduction of costs, the market acceptance and penetration of electric vehicles are gradually increasing. However, the widespread promotion of electric vehicles still faces many challenges in terms of driving range, energy efficiency and driving experience. Among them, how to achieve energy-saving, efficient, comfortable and safe driving through intelligent speed planning is a key issue that needs to be solved in the development of electric vehicle technology.
[0003] Traditional speed planning methods often rely on pre-set rules or simple feedback control strategies, which may not achieve optimal performance when faced with complex traffic environments and changing driving conditions. In recent years, with the rapid development of artificial intelligence technology, reinforcement learning, as an advanced machine learning paradigm, has shown great potential in dealing with sequential decision-making problems. Reinforcement learning has achieved remarkable results in many fields by allowing agents to interact with the environment, explore and optimize behavioral strategies autonomously. Applying reinforcement learning to speed planning for electric vehicles is expected to achieve more intelligent and personalized driving decisions, improve the energy efficiency of electric vehicles, and ensure passenger comfort and driving safety.
[0004] In addition, the driving comfort and safety of electric vehicles are crucial to improving user experience and promoting electric vehicles. A comfortable driving experience can reduce passenger fatigue and improve passenger satisfaction, while high safety is the basic prerequisite for ensuring passenger trust and acceptance of electric vehicles. Speed planning based on reinforcement learning not only needs to consider how to save energy and use energy efficiently, but also needs to comprehensively consider passenger comfort and driving safety. By learning and adapting to complex traffic environments, reinforcement learning can provide electric vehicles with a more flexible and adaptable speed planning strategy, thereby achieving the dual goals of comfort and energy saving while ensuring safety.
[0005] Therefore, studying the electric vehicle speed planning method based on reinforcement learning has important theoretical and practical significance for promoting the development of electric vehicle technology and realizing green intelligent transportation. Summary of the invention
[0006] The purpose of the present invention is to provide a longitudinal decision-making method for autonomous driving vehicles based on model-based reinforcement learning. Through intelligent speed planning technology, it effectively balances energy saving and comfort, achieves a more efficient, safe and pleasant driving experience, and solves the problems of reduced following vehicle driving comfort and excessive energy consumption caused by uneven roads, providing a new idea for solving safe, energy-saving, efficient and comfortable autonomous driving tasks; through ecological driving control strategies, it reduces automobile energy consumption and improves ride comfort, which is of great significance for improving the overall travel experience of passengers, ensuring health and safety, improving work efficiency, promoting sustainable transportation development, and enhancing social and economic benefits.
[0007] To achieve the above object, the present invention provides a longitudinal decision-making method for an autonomous driving vehicle based on model-based reinforcement learning, comprising the following steps:
[0008] S1. Establish a longitudinal dynamics model of the vehicle and use an internal resistance model to analyze the energy consumption of the battery during vehicle driving;
[0009] S2. Construct a vertical dynamics model of the entire vehicle, use weighted root mean square acceleration as an objective indicator for evaluating ride comfort, and use the concept of annoyance rate in experimental psychology to correct the evaluation results;
[0010] S3, based on vehicle-to-vehicle communication technology, obtain and analyze the driving data of the preceding vehicle and determine the minimum safe distance for following the vehicle;
[0011] S4. Based on the model-based reinforcement learning method, greedy actions are used to update the vehicle energy consumption and longitudinal dynamics model, and the model is trained repeatedly to achieve the optimal result.
[0012] Preferably, in step S1, a vehicle longitudinal dynamics model is established and the energy consumption of the battery during vehicle driving is analyzed. The specific steps are as follows:
[0013] S11. Establish a vehicle longitudinal dynamics model as follows:
[0014] ;
[0015] in, v is the longitudinal velocity of the vehicle, is the wheel torque, R is the tire radius, For braking force, is the road load, is the total mass of the vehicle including the inertia of rotating components in the powertrain;
[0016] Road load The expression is as follows:
[0017] ;
[0018] in, , and represents the load factor related to the road conditions, M is the total mass of the vehicle, θ is the slope of the road; represents the acceleration due to gravity;
[0019] S12: According to the calculation formula of step S11, the torque of the motor is determined by the transmission ratio and efficiency of the reduction gear. and speed , as shown below:
[0020] ;
[0021] ;
[0022] in, is the final drive efficiency, is the final drive ratio, is the wheel speed;
[0023] S13. Calculate the electric energy consumed by the motor and inverter as follows:
[0024] ;
[0025] in, is the power consumed by the battery, is the total efficiency of the motor and the inverter, which is a function of the motor torque and the motor speed;
[0026] S14, battery state of charge based on internal resistance model SOC The rate of change is as follows:
[0027] ;
[0028] in, is the open circuit voltage, is the internal resistance of the battery, is the battery capacity.
[0029] Preferably, in step S2, a vertical dynamics model of the entire vehicle is constructed, and the specific steps are as follows:
[0030] S21, the vertical dynamics model of the whole vehicle is as follows:
[0031] ;
[0032] in, M is the mass matrix, C Damping matrix,K is the spring matrix; is the acceleration vector, is the velocity vector, Z is the displacement vector; is the system output;
[0033] The state space is as follows:
[0034] ;
[0035] ;
[0036] in, is the stiffness coefficient of the tire; represents the identity matrix; represents an empty matrix; and They correspond to the road surface profiles contacted by the left and right wheels respectively; is the distance between the front and rear axles; is the longitudinal velocity of the vehicle;
[0037] S22. The time domain data is converted into frequency domain data using power spectral density technology, and the weighted root mean square acceleration is used as an objective indicator for evaluating ride comfort. The calculation of this indicator involves assigning a corresponding weighting coefficient to each frequency band, as shown below:
[0038] ;
[0039] in, is the RMS value of the vertical vibration acceleration of the seat of the autonomous driving vehicle, Representative Weighting factors for each 1 / 3 octave band; and Respectively The upper and lower limits of the frequency range of a 1 / 3 octave; is the energy distribution of vibration acceleration in the frequency domain;
[0040] S23. The concept of annoyance rate in experimental psychology is used to correct the evaluation results; the annoyance rate is composed of a random fuzzy evaluation model, a membership function, and a probability distribution, as shown below:
[0041] ;
[0042] ;
[0043] ;
[0044] in, Indicates the annoyance rate; is the membership function, Represents the minimum vibration value that cannot be felt by passengers. Represents the actual vibration acceleration; is the scaling parameter; is the vibration parameter, and is a fixed constant, This is the maximum vibration value that passengers cannot tolerate.
[0045] Preferably, in step S3, the minimum safe distance for following a vehicle is determined as follows:
[0046] ;
[0047] in, The shortest safe distance, is the brake idle travel time, is the linear growth time of braking deceleration, is the initial braking speed, is the braking deceleration.
[0048] Preferably, in step S1, a reinforcement learning method is used to transform the control problem into a control strategy that is a reinforcement learning solution in a discrete form on an infinite horizon with a distance d. The specific steps are as follows:
[0049] S41. Minimize the expected cost function , as shown below:
[0050] ;
[0051] in, In the initial state and control strategies The expected cost of is the discount factor, is the number of steps, is the instantaneous cost of each step k when traveling from d(k) to d(k+1) over a fixed unit distance Δd; represents the control strategy at the k-step state;
[0052] S42. Weighted sum of multiple cost items , as shown below:
[0053] ;
[0054] in, Indicates battery SOC utilization rate Δ SOC The weight coefficient of Indicates the travel time per unit distance The weight coefficient of represents the vehicle vertical comfort constraint cost The weight coefficient of Represents the distance cost to the vehicle in front The weight coefficient of
[0055] S43, the energy consumption and efficiency of step k are as follows:
[0056] ;
[0057] ;
[0058] in, represents the battery SOC utilization rate at step k; The unit distance travel time of k steps; represents the longitudinal velocity of k steps; represents the longitudinal velocity of k+1 steps; Indicates the rate of change of the battery state of charge SOC at time k;
[0059] S44, vertical comfort function , as shown below:
[0060] ;
[0061] in, is the maximum comfortable speed of the vehicle at step k+1; Indicates speed deviation; represents the speed of the vehicle at k+1 steps;
[0062] S45, in order to maintain a safe distance with the vehicle in front to prevent collision, the safety function is constructed , as shown below:
[0063] ;
[0064] in, is the relative distance between the vehicle and the preceding vehicle at k steps, is a constant value of the safe distance between vehicles, α Represents a value close to 0;
[0065] S46, state vector, as shown below:
[0066] ;
[0067] in, is the longitudinal speed, is the height, is the road slope, For prior knowledge, is the observed speed of the preceding vehicle; Indicates the relative distance between the vehicle and the vehicle ahead;
[0068] S47. Based on Q-learning, random data-driven reinforcement learning is used to adaptively optimize the longitudinal control strategy of the vehicle.
[0069] Preferably, step S47 is based on Q-learning, using random data driven reinforcement learning to adaptively optimize the vehicle longitudinal control strategy, and the specific steps are as follows:
[0070] S471, Q-learning optimal control strategy and optimal price , as shown below:
[0071] ;
[0072] ;
[0073] Among them, the Q factor represents the action value function, which is expressed as:
[0074] ;
[0075] S472. In order to determine the value of the q function, the action value function is updated as follows:
[0076] ;
[0077] S473, control input u Simplified to the relative offset of the discretized vehicle speed, as follows:
[0078] ;
[0079] in, is the discretized unit of the car speed.
[0080] Therefore, the present invention adopts the above-mentioned longitudinal decision-making method for autonomous driving vehicles based on model-based reinforcement learning, and effectively balances energy saving and comfort through intelligent speed planning technology to achieve a more efficient, safe and pleasant driving experience. At the same time, it solves the problems of reduced following vehicle driving comfort and excessive energy consumption caused by uneven roads, and provides a new idea for solving safe, energy-saving, efficient and comfortable autonomous driving tasks; through the ecological driving control strategy, it reduces automobile energy consumption and improves ride comfort, which is of great significance for improving the overall travel experience of passengers, ensuring health and safety, improving work efficiency, promoting sustainable transportation development, and enhancing social and economic benefits.
[0081] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 It is a longitudinal decision model structure diagram of a longitudinal decision method for an autonomous driving vehicle based on model-based reinforcement learning according to the present invention;
[0083] Figure 2 It is an overall structural diagram of an electric vehicle of the present invention according to a longitudinal decision-making method of an autonomous driving vehicle based on model-based reinforcement learning;
[0084] Figure 3 It is a schematic diagram of calculating the maximum comfortable speed of a longitudinal decision-making method for an autonomous driving vehicle based on model-based reinforcement learning according to the present invention;
[0085] Figure 4 It is a schematic diagram of the reinforcement learning model training process of the longitudinal decision-making method of an autonomous driving vehicle based on model-based reinforcement learning of the present invention. DETAILED DESCRIPTION
[0086] In order to make the purpose, technical scheme and advantages disclosed in the embodiments of the present invention more clearly understood, the embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the embodiments of the present invention and are not intended to limit the embodiments of the present invention. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0087] like Figure 1 As shown, a longitudinal decision-making method for an autonomous driving vehicle based on model-based reinforcement learning includes the following steps:
[0088] S1. Establish a longitudinal dynamics model of the vehicle and use an internal resistance model to analyze the energy consumption of the battery during vehicle driving;
[0089] S2. Construct a vertical dynamics model of the entire vehicle, use weighted root mean square acceleration as an objective indicator for evaluating ride comfort, and use the concept of annoyance rate in experimental psychology to correct the evaluation results;
[0090] S3, based on vehicle-to-vehicle communication technology, obtain and analyze the driving data of the preceding vehicle and determine the minimum safe distance for following the vehicle;
[0091] S4. Based on the model-based reinforcement learning method, greedy actions are used to update the vehicle energy consumption and longitudinal dynamics model, and the model is trained repeatedly to achieve the optimal result.
[0092] Example
[0093] S1. Establish a longitudinal dynamics model of the vehicle and use an internal resistance model to analyze the energy consumption of the battery during vehicle driving;
[0094] S11. In the process of building the vehicle model, a simplified one-dimensional longitudinal dynamics model is selected. The model is based on the quasi-static assumption. Under this assumption, it is considered that the impact of the short-term dynamic response of the motor and its power system on the overall model can be ignored. Therefore, the longitudinal dynamics are as follows:
[0095] ;
[0096] in, v is the longitudinal velocity of the vehicle, is the wheel torque, R is the tire radius, For braking force, is the total vehicle mass including the inertia of rotating components in the powertrain, is the road load.
[0097] Road load , as shown below:
[0098] ;
[0099] in, , and is the load factor related to road conditions, M is the total mass of the vehicle, θ is the slope of the road; Represents the acceleration due to gravity.
[0100] S12: Calculate the wheel torque according to the calculation formula of step S11, combined with the vehicle speed and expected acceleration .
[0101] In actual vehicles, in order to ensure safety, during deceleration, the torque of regenerative braking will be distributed to the wheel torque and braking force Under the assumptions of this model, unless there is a limit on the wheel torque, it is assumed that all regenerative braking is achieved through the wheel torque during deceleration; in this case, the regenerative braking that exceeds the wheel torque will be converted into braking force, and the torque of the motor is determined by the transmission ratio and efficiency of the reduction gear. and speed , as shown below:
[0102] ;
[0103] ;
[0104] in, For the final drive efficiency, is the final drive ratio, is the wheel speed.
[0105] S13. Calculate the electric energy consumed by the motor and inverter as follows:
[0106] ;
[0107] in, is the power consumed by the battery, is the total efficiency of the motor, including the efficiency of the inverter, which is a function of the motor torque and the motor speed.
[0108] S14. For batteries, based on the internal resistance model, the battery state of charge SOC The rate of change is as follows:
[0109] ;
[0110] in, is the open circuit voltage, is the internal resistance of the battery, is the battery capacity; the overall structure of the vehicle, such as Figure 2 shown.
[0111] S2. Construct a vertical dynamics model of the entire vehicle, use weighted root mean square acceleration as an objective indicator for evaluating ride comfort, and use the concept of annoyance rate in experimental psychology to correct the evaluation results;
[0112] S21. Since the standard quarter car model is too simplified and cannot fully capture the vibration characteristics of the vehicle, a more complex model is used, that is, a full vehicle model including seats. The vertical dynamic model of the full vehicle is as follows:
[0113] ;
[0114] in, M is the mass matrix, C Damping matrix, K is the spring matrix; is the acceleration vector, is the velocity vector, Z is the displacement vector; is the system output.
[0115] The state space is as follows:
[0116] ;
[0117] ;
[0118] in, is the stiffness coefficient of the tire; represents the identity matrix; represents an empty matrix; and They correspond to the road surface profiles contacted by the left and right wheels respectively; l is the distance between the front and rear axles; v is the longitudinal velocity of the vehicle.
[0119] The input data of the full vehicle model is mainly the road profile in the time domain. Although the spatial road profile remains consistent, the time domain data will change due to different driving speeds. In the state space formula, the output is usually manifested as time domain acceleration with irregular fluctuations. In contrast, the acceleration pattern in the frequency domain shows more stable characteristics.
[0120] Therefore, the present invention uses power spectral density technology to convert time domain data into frequency domain data. In frequency domain analysis, the vibration frequency band of 0.5 to 80 Hz has a particularly significant impact on human perception, and there are significant differences in the impact of different frequency bands within this range. In order to more accurately evaluate these differences, the focus is placed on the vibration in a specific frequency band, and the frequency band is further subdivided into 23 parts through a 1 / 3 octave filter. According to the recommendations of ISO 2631-1-1997 standard, weighted root mean square acceleration is used as an objective indicator for evaluating ride comfort. The calculation of this indicator involves assigning corresponding weighting coefficients to each frequency band, as shown below:
[0121] ;
[0122] in, is the RMS value of the vertical vibration acceleration of the seat of the autonomous driving vehicle, Representative Weighting factors for each 1 / 3 octave band; and Respectively The upper and lower limits of the frequency range of a 1 / 3 octave; is the energy distribution of vibration acceleration in the frequency domain.
[0123] S23. Although weighted root mean square acceleration can be used as an objective indicator to evaluate ride comfort, it cannot fully reflect the individual sensitivity differences of passengers. In fact, ride comfort is a subjective experience, which means that even when facing the same vibration, different passengers may have significantly different feelings. In order to more accurately measure the proportion of passengers who cannot tolerate vibration, the concept of annoyance rate in experimental psychology is used to correct the evaluation results. The annoyance rate consists of a random fuzzy evaluation model, a membership function, and a probability distribution, as shown below:
[0124] ;
[0125] ;
[0126] ;
[0127] in, Indicates the annoyance rate; is the membership function, Represents the minimum vibration value that cannot be felt by passengers. Represents the actual vibration acceleration; is the scaling parameter; is the vibration parameter, a and b is a fixed constant, This is the maximum vibration value that passengers cannot tolerate.
[0128] Although vibration perception varies depending on passenger expectations and activities, ISO 2631-1 provides an approximate indication of the response to different vibration amplitudes. and Usually set to 0.125 and 2.3m / s² respectively. a and b The values of are 0.4726 and 0.5478 respectively.
[0129] In the present invention, the annoyance rate of a specific road section is calculated using a conventional road quality assessment method. The road profile of the driving trajectory is evenly divided into several sections, and the corresponding annoyance rate is calculated based on the speed and road conditions of each section. In order to improve the riding experience of passengers, the goal of intelligent speed control is to control the annoyance rate within 25%, thereby ensuring that 75% of passengers feel comfortable. The speed that meets this standard is regarded as an important basis for measuring vertical comfort and is directly applied to the speed control of autonomous vehicles.
[0130] like Figure 3 As shown in the figure, the annoyance rate at different speeds is calculated and recorded at the end of each section. Among them, the circle represents the situation where the annoyance rate is less than 25%, while the other figure represents the annoyance rate above 25%. In order to keep the annoyance rate within 25%, the maximum comfortable speed (MCS) of each section is determined. MCS, as a priori knowledge of vertical comfort, provides an important reference for real-time speed control.
[0131] S3, based on vehicle-to-vehicle communication technology, obtain and analyze the driving data of the preceding vehicle and determine the minimum safe distance for following the vehicle;
[0132] Combined with the real-time speed of the following vehicle and the current road conditions, the braking distance of the vehicle is calculated and used as the shortest safe distance, as shown below:
[0133] ;
[0134] in, The shortest safe distance, is the brake idle travel time, is the linear growth time of braking deceleration, is the initial braking speed, is the braking deceleration;
[0135] S4, based on the model-based reinforcement learning method, the vehicle energy consumption and longitudinal dynamics model are updated using greedy actions, and the model is trained repeatedly to achieve the optimal result;
[0136] For this driving problem, it can be expressed as finding a steady-state control strategy on an infinite horizon from a random perspective. In this form, given the stable distribution of the vehicle driving environment, the optimal control strategy obtained by reinforcement learning RL can be generally applicable to different driving situations. Therefore, the control problem can be expressed as finding a general control strategy as an RL solution in a discrete form on an infinite horizon with a distance of d. The specific steps are as follows:
[0137] S41. Minimize the expected cost function , as shown below:
[0138] ;
[0139] in, In the initial state and control strategies The expected cost of is the discount factor; is the number of steps, is the instantaneous cost of each step k when traveling from d(k) to d(k+1) over a fixed unit distance Δd; Represents the control strategy in the k-step state.
[0140] S42. Weighted sum of multiple cost items , as shown below:
[0141] ;
[0142] in, Indicates battery SOC utilization rate Δ SOC The weight coefficient of Indicates the travel time per unit distance The weight coefficient of represents the vehicle vertical comfort constraint cost The weight coefficient of Represents the distance cost to the vehicle in front The weight coefficient of .
[0143] S43, the energy consumption and efficiency of step k are as follows:
[0144] ;
[0145] ;
[0146] in, represents the battery SOC utilization rate at step k; The unit distance travel time of k steps; represents the longitudinal velocity of k steps; represents the longitudinal velocity of k+1 steps; Represents the rate of change of the battery state of charge SOC at time k.
[0147] Driving speed is a key factor affecting vertical comfort, and MCS provides the vehicle with information about the vertical comfort of the road ahead. To ensure passenger comfort, autonomous vehicles should control their speed within a specific range. ;in, is the maximum comfortable speed of the vehicle at k steps. This will only cause discomfort to a few passengers. When the vehicle speed is in this range, its impact on vertical comfort is acceptable, so the eigenvalue is set to zero.
[0148] However, once the speed exceeds this range, the intelligent system will be punished accordingly. is used as a guide to adjust driving speed. At the same time, the degree of penalty is related to the desired speed deviation It is inversely proportional to the speed, and is intended to ensure that the speed deviation can be controlled within the expected range.
[0149] S44, vertical comfort function , as shown below:
[0150] ;
[0151] in, is the maximum comfortable speed of the vehicle at step k+1; Indicates speed deviation; represents the speed of the vehicle at step k+1.
[0152] S45, in order to maintain a safe distance with the vehicle in front to prevent collision, the safety function is constructed , as shown below:
[0153] ;
[0154] in, is the relative distance between the vehicle and the preceding vehicle at k steps, is a constant value representing the safe distance between vehicles. α To represent a value close to 0, take 0, 1 to avoid the denominator being 0.
[0155] S46, state vector, as shown below:
[0156] ;
[0157] in, is the longitudinal speed, is the height, is the road slope, For prior knowledge, is the observed speed of the preceding vehicle; Indicates the relative distance between the vehicle and the vehicle in front.
[0158] The state vector is composed of the height (altitude) h , road slope θ , prior knowledge Defined driving cycle and vehicle speed , and the observed speed of the preceding vehicle The state corresponding to the relative distance between cars.
[0159] Prior knowledge Samples are collected from the MCS at certain distance intervals to represent future vertical comfort information. In the present invention, it is assumed that the speed of the vehicle ahead and the distance between vehicles (up to 60 meters) can be provided by sensors or vehicle-to-vehicle communication systems. Other information, such as height, slope and vehicle speed, and road conditions can be obtained from the roadside unit and vehicle system.
[0160] S47. In the present invention, the longitudinal control strategy of the vehicle is derived based on the RL algorithm of Q-learning, and the longitudinal control strategy of the vehicle is adaptively optimized by RL driven by random data. The specific steps are as follows:
[0161] S471, Q-learning optimal control strategy and optimal price , as shown below:
[0162] ;
[0163] ;
[0164] Among them, the Q factor represents the action value function, which is expressed as:
[0165] ;
[0166] S472. In order to determine the value of the q function, the action value function is updated as follows:
[0167] ;
[0168] In Q-learning, the control input is determined by the q function, and through interaction with the environment, the q function value is updated after the agent performs an action. In this process, the agent can observe the instantaneous cost and changes in state variables, and then use this information to update the q function value.
[0169] In addition, based on the Q-learning structure, MBRL (model-based reinforcement learning) is combined with vehicle powertrain dynamics and longitudinal dynamics. In MBRL, by building and utilizing the vehicle powertrain battery SOC The consumed data-driven model improves the sampling efficiency of Q-learning, making it more efficient in practical applications.
[0170] S473, control input u Simplified to the relative offset of the discretized vehicle speed, as follows:
[0171] ;
[0172] in, is the discrete unit of the car speed; only when the corresponding motor torque and motor speed Control input can only be applied if allowed .for , , and A given driving cycle data represented by , Q-learning iterations using vehicle simulations.
[0173] Reinforcement learning algorithms such as Figure 4 As shown, in the outer loop, the present invention adopts a greedy strategy to select actions as control input; in the inner loop, the q-function value is updated for all possible actions using an approximate environment model built based on the driving batch data. This reinforcement learning structure means that the powertrain dynamics and longitudinal dynamics are represented by approximate modeling, which are learned through virtual interactions of the agents.
[0174] Changes in driving cycle information, such as arrive , the state variable transformation of the preceding vehicle speed behavior, such as arrive , as experience replay is passed from the actual environment to the approximate environment model. However, this information is stochastic and difficult to generalize directly.
[0175] In order to make more efficient use of data, the stored actual experience is used multiple times for off-policy learning, providing a more efficient sampling method. In this way, the learning process can converge quickly without the time-consuming acquisition of actual experience. Using these stored experiences, different control inputs can be tested and the estimated rewards can be used to and the estimated subsequent state transitions To update the Q function value.
[0176] In order to handle different vehicle speeds and relative distances between vehicles, all allowed actions are executed in a "for loop". ,in Iterate to update the Q function; then at time step k, it is as follows:
[0177] ;
[0178] Among them, controlled by Estimated next speed of the vehicle , as shown below:
[0179] ;
[0180] in, is the next relative distance, as follows:
[0181] ;
[0182] in, Indicates different vehicle speeds. Represents different distances from the vehicle in front; both are discretized uniformly as follows:
[0183] ;
[0184] ;
[0185] in, represents the discretized quantity of vehicle speed, The discretized quantity representing the relative distance. In the quantization process, the nearest neighbor quantization method is used to obtain the discrete values of the q-table.
[0186] In addition, the approximate cost can be estimated by step S21 , through the approximate model Calculate the battery SOC Utilization rate Δ SOC ,and , from the data of the simulated outer loop:
[0187] ;
[0188] in, The status is x Approximate battery SOC Usage Δ SOC , u represents a step control input, is the learning rate of the model.
[0189] It can also be calculated and initialized based on the known powertrain dynamics, so the approximate cost , as shown below:
[0190] ;
[0191] Therefore, the present invention adopts the above-mentioned longitudinal decision-making method for autonomous driving vehicles based on model-based reinforcement learning, and effectively balances energy saving and comfort through intelligent speed planning technology to achieve a more efficient, safe and pleasant driving experience. At the same time, it solves the problems of reduced following vehicle driving comfort and excessive energy consumption caused by uneven roads, and provides a new idea for solving safe, energy-saving, efficient and comfortable autonomous driving tasks; through the ecological driving control strategy, it reduces automobile energy consumption and improves ride comfort, which is of great significance for improving the overall travel experience of passengers, ensuring health and safety, improving work efficiency, promoting sustainable transportation development, and enhancing social and economic benefits.
[0192] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A longitudinal decision-making method for an autonomous driving vehicle based on model-based reinforcement learning, characterized in that: The following steps are involved: S1. Establish a longitudinal dynamics model of the vehicle and use the internal resistance model to analyze the energy consumption of the battery during vehicle driving. The specific steps are as follows: S11. Establish a vehicle longitudinal dynamics model as follows: ; in, is the longitudinal velocity of the vehicle, is the wheel torque, R is the tire radius, For braking force, is the road load, is the total mass of the vehicle including the inertia of rotating components in the powertrain; Road load The expression is as follows: ; in, , and represents the load factor related to the road conditions, M is the total mass of the vehicle, θ is the slope of the road; represents the acceleration due to gravity; S12: According to the calculation formula of step S11, the torque of the motor is determined by the transmission ratio and efficiency of the reduction gear. and speed , as shown below: ; ; in, is the final drive efficiency, is the final drive ratio, is the wheel speed; S13. Calculate the electric energy consumed by the motor and inverter as follows: ; in, is the power consumed by the battery, and the overall efficiency of the inverter, which is a function of motor torque and motor speed; S14, battery state of charge based on internal resistance model SOC The rate of change is as follows: ; in, is the open circuit voltage, is the internal resistance of the battery, is the battery capacity; S2. Construct a vertical dynamics model of the entire vehicle, use weighted root mean square acceleration as an objective indicator for evaluating ride comfort, and use the concept of annoyance rate in experimental psychology to correct the evaluation results; S3, based on vehicle-to-vehicle communication technology, obtain and analyze the driving data of the preceding vehicle and determine the minimum safe distance for following the vehicle; S4. Based on the model-based reinforcement learning method, greedy actions are used to update the vehicle energy consumption and longitudinal dynamics model, and the model is trained repeatedly to achieve the optimal result.
2. The longitudinal decision-making method for an autonomous driving vehicle based on model-based reinforcement learning according to claim 1, characterized in that: In step S2, a vertical dynamics model of the entire vehicle is constructed. The specific steps are as follows: S21, the vertical dynamics model of the whole vehicle is as follows: ; in, M is the mass matrix, C Damping matrix, K is the spring matrix; is the acceleration vector, is the velocity vector, Z is the displacement vector; is the system output; The state space is as follows: ; ; in, is the stiffness coefficient of the tire; represents the identity matrix; represents an empty matrix; and They correspond to the road surface profiles contacted by the left and right wheels respectively; is the distance between the front and rear axles; is the longitudinal velocity of the vehicle; S22. The time domain data is converted into frequency domain data using power spectral density technology, and the weighted root mean square acceleration is used as an objective indicator for evaluating ride comfort. The calculation of this indicator involves assigning a corresponding weighting coefficient to each frequency band, as shown below: ; in, is the RMS value of the vertical vibration acceleration of the seat of the autonomous driving vehicle, Representative Weighting factors for each 1 / 3 octave band; and Respectively The upper and lower limits of the frequency range of a 1 / 3 octave; is the energy distribution of vibration acceleration in the frequency domain; S23. The concept of annoyance rate in experimental psychology is used to correct the evaluation results; the annoyance rate is composed of a random fuzzy evaluation model, a membership function, and a probability distribution, as shown below: ; ; ; in, Indicates the annoyance rate; is the membership function, Represents the minimum vibration value that cannot be felt by passengers. Represents the actual vibration acceleration; is the scaling parameter; is the vibration parameter, and is a fixed constant, This is the maximum vibration value that passengers cannot tolerate.
3. The longitudinal decision-making method for an autonomous driving vehicle based on model-based reinforcement learning according to claim 1, characterized in that: In step S3, the minimum safe distance for following a vehicle is determined as follows: ; in, The shortest safe distance, is the brake idle travel time, is the linear growth time of braking deceleration, is the initial braking speed, is the braking deceleration.
4. The longitudinal decision-making method for an autonomous driving vehicle based on model-based reinforcement learning according to claim 1, characterized in that: In step S4, the control problem is transformed into a control strategy that is a reinforcement learning solution in a discrete form on an infinite horizon with a distance d by using the reinforcement learning method. The specific steps are as follows: S41. Minimize the expected cost function , as shown below: ; in, In the initial state and control strategies The expected cost of is the discount factor, is the number of steps, is the instantaneous cost of each step k when traveling from d(k) to d(k+1) over a fixed unit distance Δd; represents the control strategy at the k-step state; S42. Weighted sum of multiple cost items , as shown below: ; in, Indicates battery SOC utilization rate Δ SOC The weight coefficient of Indicates the travel time per unit distance The weight coefficient of represents the vehicle vertical comfort constraint cost The weight coefficient of Represents the distance cost to the vehicle in front The weight coefficient of S43, the energy consumption and efficiency of step k are as follows: ; ; in, represents the battery SOC utilization rate at step k; The unit distance travel time of k steps; represents the longitudinal velocity of k steps; represents the longitudinal velocity of k+1 steps; Indicates the rate of change of the battery state of charge SOC at time k; S44, vertical comfort function , as shown below: ; in, is the maximum comfortable speed of the vehicle at step k+1; Indicates speed deviation; represents the speed of the vehicle at k+1 steps; S45, in order to maintain a safe distance with the vehicle in front to prevent collision, the safety function is constructed , as shown below: ; in, is the relative distance between the vehicle and the preceding vehicle at k steps, is a constant value of the safe distance between vehicles, α Represents a value close to 0; S46, state vector, as shown below: ; in, is the longitudinal speed, is the height, is the road slope, For prior knowledge, is the observed speed of the preceding vehicle; Indicates the relative distance between the vehicle and the vehicle ahead; S47. Based on Q-learning, random data-driven reinforcement learning is used to adaptively optimize the longitudinal control strategy of the vehicle.
Citation Information
Patent Citations
Longitudinal decision-making system and longitudinal decision-making determination method for automatic driving vehicle
CN111391830A
Automatic driving longitudinal decision control method in vehicle-road cooperation environment
CN112896186A