Self-adaptive cruise control method based on reinforcement learning driver style
By applying reinforcement learning technology in the vehicle adaptive cruise control system and updating controller parameters to match the driver's style, the problem of mismatch between the autonomous driving process and the driver's style is solved, achieving a more natural driving experience and higher control accuracy.
Patent Information
- Application Number
- CN202510276356.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-05-06
AI Technical Summary
The autonomous driving process under the vehicle's adaptive cruise control is difficult to match the driver's driving style.
Using a reinforcement learning-based method, by constructing an agent, obtaining information about driver intervention in driving, updating controller parameters, and optimizing controllers to match the driver's driving style.
It achieves the matching of the adaptive cruise controller with the driver's style while ensuring safety, improving the naturalness and driving experience of autonomous driving.
Smart Images

Figure CN119928891A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of vehicle control technology, and in particular to an adaptive cruise control method based on reinforcement learning of driver style. Background Art
[0002] The adaptive cruise control system senses the environment around the vehicle and automatically adjusts the speed, allowing the controlled vehicle to follow the vehicle in front at the desired distance. It can effectively reduce the driver's workload while ensuring the driver's safety, and improve the level of highway intelligence and operating efficiency. With the development of unmanned driving and intelligent networked vehicles, higher requirements are placed on the control needs and control accuracy of vehicle adaptive cruise control systems.
[0003] Model-based autonomous driving control methods include nonlinear model predictive control, linear quadratic Gaussian model control, backstepping control, sliding mode control, high gain scheduling control, nonlinear neural network adaptive control, etc. Although the control methods based on these models can solve the adaptive cruise control of vehicles to a certain extent and achieve the corresponding control goals, the existing self-driving models of major manufacturers are all fixed parameters, which are difficult to match the driving style of the driver. Summary of the invention
[0004] The technical problem to be solved by the present invention is that the automatic driving process under the vehicle adaptive cruise control is difficult to match the driver's driving style.
[0005] To this end, the present invention provides an adaptive cruise control method based on reinforcement learning of driver style.
[0006] The technical solution adopted by the present invention to solve its technical problem is:
[0007] An adaptive cruise control method based on reinforcement learning driver style, comprising:
[0008] Step 1: Build a vehicle adaptive cruise controller and obtain controller parameters;
[0009] Step 2: Obtain information about the driver's intervention in driving during the vehicle's adaptive cruise control process, construct an intelligent agent corresponding to each controller parameter, and update the controller parameters according to the driver's intervention in driving to optimize the controller;
[0010] Step three, using the optimized controller to perform adaptive cruise control.
[0011] Furthermore, in step 2, the information on the driver's intervention in driving includes: the ratio k_a1 between the actual acceleration of the vehicle and the acceleration calculated by the controller u(t) from the time the driver presses the accelerator pedal to the time the driver releases the accelerator pedal, the ratio k_a2 between the actual deceleration of the vehicle and the acceleration calculated by the controller u(t) from the time the driver presses the brake pedal to the time the driver releases the brake pedal, the number of times the driver intervenes in the automatic driving, the vehicle speed V_self, the relative speed V_front and the relative position D_s of the front vehicle during each intervention.
[0012] Furthermore, when the number of times the driver intervenes in the automatic driving reaches a preset number N_driver, the model starts to update the controller parameters and optimize the controller based on the N_driver intervention data.
[0013] Furthermore, in step 2, the updated controller parameters are input into the original controller and executed in the simulation environment. The reward obtained is fed back to each agent, and the controller parameters are updated again until the difference between the current reward value and the reward value obtained in the previous simulation environment is less than a preset value or the number of iterations reaches a preset value. Then, the controller parameters are stopped from being updated to obtain the final controller.
[0014] Furthermore, in step 2, the agents respectively adopt the confidence interval upper limit strategy (UCB) as the action output strategy of each agent at each time step t, and the specific strategy is expressed as: Among them: A t is the action selected within t time steps; is the reward for action a, is the empirical mean term, representing the historical average reward of action a; is the exploration item, which is inversely proportional to the number of explorations of the action; n(a) is the number of times action a is selected, and c is the exploration parameter, which controls the degree of exploration (c is set to ).
[0015] Furthermore, the reward function includes an acceleration ratio reward function R A And the safety distance reward function R safe , the reward function is r t =R A +R safe .
[0016] Furthermore, the acceleration ratio reward function is: Where α is a positive weight coefficient used to adjust the magnitude of the reward. a is the ratio k_a between the actual acceleration of the vehicle in the previously recorded intervention data and the acceleration calculated by the controller u(t), Indicates the deviation of the ratio from the target value.
[0017] Furthermore, the safety distance reward function is: Among them: κ is a positive weight coefficient used to reward the situation where the safety distance is greater than the threshold; γ is a positive weight coefficient used to punish the situation where the safety distance is less than the set threshold. When the safety distance is greater than the set safety distance, a positive reward is given; when the safety distance is less than the set safety distance, a negative reward is given.
[0018] The beneficial effect of the present invention is that the present application controls the exploration depth of the intelligent agent through the intelligent agent selection parameter strategy with exploration parameters, avoiding unnecessary waste of computing resources. The present application designs the acceleration ratio based on the acceleration value of the driver stepping on the accelerator or brake recorded during intervention and the controller model theory as the reinforcement learning target of the cruise controller to implement the strategy of the driver-style cruise controller. And through the reward function of the acceleration ratio and the safe distance, a more driver-style cruise controller can be implemented under the premise of ensuring safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0020] Figure 1 It is a schematic diagram of the flow chart of controller optimization in the present invention.
[0021] Figure 2 It is a structural schematic diagram of the desired following vehicle distance measurement in the present invention.
[0022] Figure 3 It is an algorithm structure block diagram of the adaptive cruise control method based on neural network approximation in the present invention.
[0023] Figure 4 It is a flow chart of the algorithm for the intelligent agent to update controller parameters in the present invention. DETAILED DESCRIPTION
[0024] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.
[0025] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", "axial", "radial", "circumferential" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.
[0026] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0027] An adaptive cruise control method based on reinforcement learning driver style, comprising the following steps
[0028] Step 1: Build the vehicle controller
[0029] S1.1 Establishing the nonlinear dynamic model of the vehicle
[0030]
[0031] Where: p(t), v(t) are the actual position and speed of the vehicle, is the derivative of p(t), and the derivative of the following vehicle position is the following vehicle speed, is the derivative of v(t), and the derivative of the following vehicle velocity is the following vehicle acceleration; M is the real mass of the following vehicle; F is the real output force of the vehicle; F f =Mgc r Indicates the rolling resistance of the vehicle; Indicates the air resistance of the vehicle; F i =Mgsin(γ) represents the slope resistance; g,c r ,c f ,A f ,ρ a,γ distribution represents the gravitational acceleration, rolling resistance coefficient, air resistance coefficient, vehicle frontal area, air density and road slope angle.
[0032] Since the vehicle's driving resistance is affected by various factors such as vehicle mass, weather changes, and road conditions, c r ,c f ,A f ,ρ a , γ and other parameters have time-varying characteristics; in addition, various resistances are related to the independent components on the vehicle, and there is a part of the vehicle running resistance that cannot be accurately modeled, resulting in a large control error when directly using formula (1) for controller design.
[0033] S1.2 uses the parallel learning characteristics of radial basis neural network to approximate the vehicle's driving resistance, thereby achieving an effective description of the vehicle's nonlinear resistance.
[0034] It should be noted that Lemma 1: The radial basis neural network can approximate unknown nonlinear smooth functions online, and the input of the neural network is: Among them, q represents the dimension of the neural network input; is a vector parameter, R represents a real number set, m represents the number of neurons, is a Gaussian function: In the formula, μ l ,x l is the center and width of the Gaussian function, the radial basis neural network can set the neuron to approximate Q(z) so that:
[0035] Q(z)=W *T Ψ(z)+ε, the approximation error ε can be adjusted to the ideal vector parameter W * Make it as small as possible, that is, is a small positive number, represents the mean of the approximation error, The set of real numbers representing z as a q-dimensional vector.
[0036] Based on the above Lemma 1, the nonlinear driving resistance of the vehicle is approximated, and equation (1) can be set as
[0037]
[0038] Wherein, u(t) = F(t) / M, M is the real mass of the following vehicle; F(t) is the real output force of the vehicle; The nonlinear function in the neural network can approximate the nonlinear driving resistance of the vehicle. W *T To make The smallest "optimal" weight vector, where W *T is the vector parameter W* The transpose of y is the tracking vehicle distance output by the system.
[0039] S1.2, the sensor measures the distance between vehicles and defines the distance tracking error e d , and design the virtual control amount α of the spacing surface error d .
[0040] ACC (Adaptive Cruise Control System) Figure 1 The pilot car in the vehicle adjusts its own speed adaptively to keep the desired following distance between the two vehicles. Its structure is as follows: Figure 1 As shown. The spacing tracking error surface is defined as:
[0041] e d (t) = p(t) - p L (t)-d (3)
[0042] Among them, p L (t) is the actual position of the vehicle in front. The goal of adaptive cruise control is to design a controller so that the system error state reaches a sliding plane described by a polynomial that tracks the speed error and moves on the sliding plane. The control system operates stably, that is, e(t)→0, so that p L (t)-p(t)→d, where d is the desired vehicle spacing.
[0043] In order to ensure that the vehicle can be controlled to automatically follow the preceding vehicle and run safely at the expected vehicle distance d under uncertain driving resistance, the control objectives are set as follows:
[0044] (1) The actual vehicle spacing approaches the expected vehicle spacing, and the spacing error is as small as possible, i.e., lim t→∞ e d =p(t)+dp L (t)→0;
[0045] (2) The following vehicle can respond quickly to changes in the leading vehicle, and the speed error between the two vehicles is as small as possible, that is,
[0046] In order to achieve the above control objectives, the spacing surface error is defined as:
[0047] e d (t) = p(t) + dp L (t) (4)
[0048] Pitch error e d (t) is differentiated, and we can get:
[0049]
[0050] In the above formula, taking v(t) as the virtual control input, we can get:
[0051]
[0052] In the above formula, k d is a constant. In this way, equation (5) can be transformed into
[0053]
[0054] Consider the Lyapunov function
[0055]
[0056] The derivative is
[0057]
[0058] From formula (9), we can see that the spacing surface error is stable. Since v(t) is not a control input, v(t) is called a virtual control variable, which can be expressed as
[0059]
[0060] Combining (5), we can get It can be expressed as
[0061]
[0062] at this time Recorded as
[0063]
[0064] S1.3, using nonlinear filter to convert the virtual control amount α of the spacing surface error d Transformed into the virtual control quantity s of the speed surface, and the speed surface error e in this control system is constructed v .
[0065] In order to avoid the need to use the virtual control variable α when designing the controller d (t) is differentiated, which leads to complex interference terms. Combined with the first-order inertia link in the DSC dynamic surface control design, a nonlinear filter with a sampling time constant τ is introduced to replace the differential calculation. The nonlinear filter satisfies:
[0066]
[0067] Where s is the virtual control amount of the speed surface error, and the speed surface error e is defined on this basis. v And the boundary layer error E is:
[0068] e v(t) = v (t) - s (t) (14)
[0069] E=s-α d (15)
[0070] Velocity surface error e v represents the error between the actual value of the vehicle speed v(t) and the expected value (virtual control quantity s), and the boundary layer error E represents e v The expected value s and the nonlinear filter input α d The error between .
[0071] The specific form is:
[0072]
[0073] in, M d The estimated value of is, and its adaptive law is β is the setting coefficient, M d It mainly represents the bounded interference caused by the pilot vehicle speed, the own vehicle speed, the vehicle spacing strategy, etc. σ(t) is a positive function and satisfies and Where σ1 and σ2 are positive constants. The nonlinear filter designed by equation 16 is introduced by Compensate for the influence caused by the boundary layer error, and then make E→0. Therefore, the control variable α is used to d (t) is converted into a virtual control quantity s.
[0074] Combining equation (2), equation (14), and equation (16) to calculate e v The derivative is
[0075]
[0076] S1.4, based on the previous, the Lyapunov function is redefined and the radial basis neural network function is combined to optimize the parameters to design the adaptive law and the actual controller.
[0077] First of all, it should be noted that Lemma 2: In a neural network, w is the weight coefficient of the neural network node. There exists a constant ε>0 that satisfies:
[0078]
[0079] Lemma 3: For a given parameter w0>0, and There is an inequality:
[0080]
[0081] where κ = e-(κ+1) ,κ=0.2785 is a fixed constant.
[0082] Because of the new quantity e in the previous step (17) v , Therefore, the Lyapunov function is redefined as:
[0083]
[0084] in, is an estimate of W, yes is the estimated value of , ζ1 and ζ2 are constants greater than 0.
[0085] Differentiate equation (14) and substitute equation (13) into it, and we can get:
[0086]
[0087] From Lemma 1, we can see that the basis function of the radial basis neural network is a Gaussian function, and the error function is defined as the mean square error, so is the Euclidean distance deviation, θ is the variance of the point estimate. To approximate Q(z), the smaller the deviation, the smaller the variance. Combined with Lemma 3, we can get the following two inequalities:
[0088]
[0089]
[0090] Substituting equation (20) into equation (19) yields:
[0091]
[0092] The adaptive law of the design parameters is:
[0093]
[0094] Based on equation (21) and combined with equation (22), the actual controller u(t) is designed as:
[0095]
[0096] The entire neural network sliding mode controller algorithm structure diagram is as follows: Figure 2 As shown, Figure 2 Where NF is the nonlinear filter of formula (16), RBFNN is the radial basis neural network, and P is the vehicle drive device.
[0097] Step 5: Learn the driver's driving style and update the vehicle controller
[0098] The threshold number of driver interventions N_driver is set to trigger the reinforcement learning program. This means that each time the assisted driving function is turned on until the vehicle stops, if the driver complains that the following vehicle is too slow or too fast, and the driver steps on the accelerator or brake pedal once or several times, it is considered an intervention. When the threshold number N_driver is reached, the reinforcement learning program is started to learn the driver's individual driving style.
[0099] S2.1 Obtaining driver intervention in autonomous driving
[0100] During each intervention in the autonomous driving process, the ratio k_a1 between the actual acceleration of the car and the acceleration calculated by the controller u(t) should be recorded from the time the driver presses the accelerator pedal to the time the driver releases the accelerator pedal. Also, the ratio k_a2 between the actual deceleration of the car and the acceleration calculated by the controller u(t) should be recorded from the time the driver presses the brake pedal to the time the driver releases the brake pedal. k_a1 and k_a2 can be used as references for the improvement targets of the reinforcement learning of the later self-driving model.
[0101] During each intervention, the speed of the vehicle V_self, the relative speed of the front vehicle V_front and the relative position D_s are recorded. The data of the driver's intervention in the automatic driving is recorded as I_ad. I_ad is 1, which indicates the data of the marked driver's intervention in accelerating and decelerating the vehicle, and 0 indicates the data of the normal operation of the self-driving model. The number of times the driver intervenes in the automatic driving is counted, and the threshold of the number of driver interventions is set as N_driver. When N_driver is reached, the model starts to perform model reinforcement learning based on the N_driver intervention data, and starts to update the controller, so that the original self-driving model can achieve the driver's expected acceleration or deceleration in the N_driver scenario data.
[0102] S2.2 controller parameter update
[0103] The parameters of the controller are k v , k d , ζ1, ζ2, β, so five agents are used to adopt the confidence interval upper limit strategy (UCB) as the action output strategy of each agent at each time step t. The specific strategy is expressed as:
[0104]
[0105] in: is the empirical mean term, representing the historical average reward of action a. In this embodiment, the agent performs 9 actions a: adding parameter values (0.1, 0.01, 0.001, 0.0001), subtracting parameter values (0.1, 0.01, 0.001, 0.0001), and keeping the original value (0); is the exploration term, which is inversely proportional to the number of explorations of the action;
[0106] A t is the action selected within t time steps; is the reward for action a; n(a) is the number of times action a is selected, and c is the exploration parameter that controls the degree of exploration (c is set to ).
[0107] After these five agents choose to add or subtract parameter values, N(increase) (or N(decrease)) increases (i.e. n(a) increases), and the weight in each exploration item will be gradually reduced and updated in real time according to the newly obtained rewards. As t increases with each step, the overall weight of the exploration items of all actions is gradually reduced. Through the above mechanism, each agent can automatically reduce the "exploration" of actions that have been fully explored according to the UCB strategy, while giving priority to exploring actions with high uncertainty, and ultimately maximizing the cumulative reward.
[0108] S2.3 Optimized Controller
[0109] In one update strategy, each of the five agents is only responsible for increasing or decreasing the size of a parameter value, and the five agents output five optimized parameters. The five agents output the updated parameters to the controller u(t) respectively, and the reward obtained by the controller in the simulation environment of the input intervention data is fed back to each agent. During each update, the reward value obtained by the optimized controller in the simulation environment is judged, and the difference between the reward value obtained by the last optimized controller in the simulation environment is judged. When the reward value difference is less than 0.001 or the number of iterations reaches 10,000, the iteration is stopped and the final optimized controller is output. Specifically, the reward function includes acceleration reward and safety distance reward.
[0110] S2.3.1 Acceleration Proportional Reward Function
[0111] When the agent takes action, the acceleration a obtained in the simulation environment of the intervention data t The a obtained by comparing with the original design controller u(t) design The closer the ratio k_a is to the ratio between the actual acceleration of the car in the previously recorded intervention data and the acceleration calculated by the controller u(t), the higher the reward. The reward function is as follows:
[0112]
[0113] Where α is a positive weight coefficient used to adjust the magnitude of the reward. a is the ratio k_a between the actual acceleration of the vehicle in the previously recorded intervention data and the acceleration calculated by the controller u(t), Indicates the deviation of the ratio from the target value.
[0114] S2.3.2 Safety distance reward function
[0115] When the safety distance is greater than the set safety distance, a positive reward is given; when the safety distance is less than the set safety distance, a negative reward is given to ensure safety.
[0116]
[0117] in:
[0118] κ is a positive weight coefficient used to reward situations where the safety distance is greater than the threshold;
[0119] γ is a positive weight coefficient used to penalize the situation where the safety distance is less than the set threshold.
[0120] So the final reward function is r t =R A +R safe , that is, the reward function is:
[0121]
[0122] It should be noted that the reward value r obtained by the nth optimized controller after execution in the simulation environment is n , the reward value r obtained by the optimized controller after execution in the simulation environment n-1 , the difference between the two reward values is recorded as L. According to the relationship between the value of action a used in the n-1th controller parameter update and L, the value of action a in the n+1th controller parameter update is determined to adjust the parameters so that the reward value r obtained after the n+1th controller is executed in the simulation environment is n+1 With r n The difference between them is reduced.
[0123] For example, after the 100th optimization, the actions selected by A1, A2, A3, A4, and A5 (A1-5 represent the agents in the above controller) are [+0.1, -0.01, +0.001, -0.0001, +0.01] respectively. The reward value obtained by the controller after the five parameters are adjusted and updated in the simulation environment is 0.64+0.73, which is fed back to the reward value of the action. According to the empirical mean term of the UCB formula, the sum of the 100 reward values U divided by the number of times 100 is u. It is considered that at the 100th time, the reward of A1 action +0.1 is The reward values of A1, A2, A3, A4, and A5 are all u (for example, u is 1.2474). The reward values of other actions corresponding to these five parameters maintain the reward values of the action when it was last selected, that is, the reward update distribution of (+0.1, +0.01, +0.001, +0.0001, 0, -0.0001, -0.001, -0.01, -0.1) for each action a in A1, A2, A3, A4, and A5 is shown in the following table:
[0124]
[0125] In the above table, when each parameter is updated under its own 9 actions, the action with the largest reward value is selected as the next action by combining the reward value of each parameter action and the exploration item score in the UCB strategy. According to the new distribution obtained by the UCB strategy, it can be seen that some experience mean items are very high, but as the number of times the action is selected increases, the reward of the exploration item is decreasing. The exploration item reward of other actions that were not selected last time is passively high, thereby increasing the probability of becoming the next selected action. Therefore, according to the reward distribution, the next action selected by the five agents for the 101st time is [+0.001, -0.01, +0.0001, +0.0001, +0.01]. The optimized parameters are configured to the controller and executed in the simulation environment to obtain a new reward value. As the program continues to iterate, the optimal parameters will be fixed to a very small interval, and the reward value will become very small. When the difference between the nth reward value and the n+1th reward value is less than 0.001, the iteration stops.
[0126] Step three: output a controller whose parameters have been debugged through reinforcement learning, and use the controller to perform adaptive loops to achieve following and braking cruise control in the same style as the driver.
[0127] Based on the above ideal embodiments of the present invention, the relevant staff can make various changes and modifications without departing from the technical concept of the present invention through the above description. The technical scope of the present invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.
Claims
1. An adaptive cruise control method based on reinforcement learning driver style, characterized in that: include, Step 1: Build a vehicle adaptive cruise controller and obtain controller parameters; Step 2: Obtain information about the driver's intervention in driving during the vehicle's adaptive cruise control process, construct an intelligent agent corresponding to each controller parameter, and update the controller parameters according to the driver's intervention in driving to optimize the controller; Step three, using the optimized controller to perform adaptive cruise control.
2. The adaptive cruise control method based on reinforcement learning driver style according to claim 1, characterized in that: In step 2, the information on the driver's intervention in driving includes: the ratio k_a1 between the actual acceleration of the vehicle and the acceleration calculated by the controller u(t) during the time from the driver pressing the accelerator pedal to the time from the driver releasing the accelerator pedal, the ratio k_a2 between the actual deceleration of the vehicle and the acceleration calculated by the controller u(t) during the time from the driver pressing the brake pedal to the time from the driver releasing the brake pedal, the number of times the driver intervenes in the automatic driving, the vehicle speed V_self, the relative speed V_front and the relative position D_s of the front vehicle during each intervention.
3. The adaptive cruise control method based on reinforcement learning driver style according to claim 1, characterized in that: When the driver intervenes in autonomous driving for a preset number of times, N_driver, the model starts to update the controller parameters and optimize the controller based on the N_driver intervention data.
4. The adaptive cruise control method based on reinforcement learning driver style according to claim 1, characterized in that: In step 2, the updated controller parameters are input into the original controller and executed in the simulation environment. The reward obtained is fed back to each agent, and the controller parameters are updated again until the difference between the current reward value and the reward value obtained in the previous simulation environment is less than the preset value or the number of iterations reaches the preset value. Then, the controller parameters are stopped from being updated to obtain the final controller.
5. The adaptive cruise control method based on reinforcement learning driver style according to claim 4, characterized in that: In step 2, the agents respectively adopt the confidence interval upper limit strategy (UCB) as the action output strategy of each agent at each time step t. The specific strategy is expressed as: Among them: A t is the action selected within t time steps; is the reward for action a, is the empirical mean term, representing the historical average reward of action a; is the exploration item, which is inversely proportional to the number of explorations of the action; n(a) is the number of times action a is selected, and c is the exploration parameter, which controls the degree of exploration (c is set to ).
6. The adaptive cruise control method based on reinforcement learning driver style according to claim 4, characterized in that: The reward function includes an acceleration proportional reward function R A And the safety distance reward function R safe , the reward function is r t =R A +R safe .
7. The adaptive cruise control method based on reinforcement learning driver style according to claim 6, characterized in that: The acceleration proportional reward function is: Where α is a positive weight coefficient used to adjust the magnitude of the reward. a is the ratio between the actual acceleration of the vehicle in the previously recorded intervention data and the acceleration calculated by the controller u(t) Indicates the deviation of the ratio from the target value.
8. The adaptive cruise control method based on reinforcement learning driver style according to claim 6, characterized in that: The safety distance reward function is: Among them: κ is a positive weight coefficient used to reward the situation where the safety distance is greater than the threshold; γ is a positive weight coefficient used to punish the situation where the safety distance is less than the set threshold. When the safety distance is greater than the set safety distance, a positive reward is given; when the safety distance is less than the set safety distance, a negative reward is given.