Reinforcement learning and game theory integrated dynamic cooperative steering control method

By integrating reinforcement learning with non-cooperative game theory, the dynamic allocation of control rights and strategy optimization in the human-machine co-driving system are realized, which solves the problems of control rights allocation and strategy coordination in the human-machine co-driving scenario and improves the safety, comfort and tracking accuracy of the system.

CN120630727BActive Publication Date: 2025-10-10CHANGCHUN UNIV OF TECH

Patent Information

Application Number
CN202511087884.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-10-10
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

Existing technologies make it difficult to achieve dynamic allocation of control rights and strategy coordination in human-machine co-driving scenarios, especially in dynamic traffic environments. Traditional methods lack environmental adaptability and real-time response capabilities, resulting in reduced safety and comfort.

Method used

The method integrates reinforcement learning and non-cooperative game theory. By generating the expected path, reinforcement learning module and non-cooperative game module, the dynamic allocation of control rights between the driver and the controller and the solution of Nash equilibrium are achieved. The control strategy is optimized by combining the reward function and prediction model.

Benefits of technology

It improves the strategic coordination of the human-machine co-driving system in dynamic environments, enhances safety, comfort and tracking accuracy, alleviates human-machine control conflicts, and improves the system's adaptability and real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120630727B_ABST
    Figure CN120630727B_ABST
Patent Text Reader

Abstract

The application relates to a method for dynamic cooperative steering control of man-machine combined vehicles based on reinforcement learning and game theory, which aims to optimize the dynamic weight distribution of the driver and the controller to ensure the safety and path tracking accuracy of the man-machine combined vehicle. The application relates to the field of intelligent driving. The application comprises four modules, namely, a desired path module, a reinforcement learning module, a non-cooperative game module and a man-machine combined vehicle. The desired path module is used for generating a desired path sequence of the driver and the controller and a desired path at the current moment; the reinforcement learning module receives the desired path at the current moment, vehicle dynamic information and optimal steering angles of the driver and the controller, and trains to generate a driver control weight; the non-cooperative game module combines the driver control weight, the desired path sequence of the driver and the controller and the state information of the ego vehicle, and optimizes the optimal steering angles of the driver and the controller through non-cooperative game to realize cooperative steering control of the man-machine combined vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical field:

[0001] The present invention belongs to the field of intelligent driving, and specifically relates to a human-machine dynamic collaborative steering control method integrating reinforcement learning with game theory. Background technology:

[0002] With the rapid development of autonomous driving technology, the level of vehicle intelligence continues to increase, and we are currently entering a transition phase from Level 2 assisted driving to Level 3 conditional autonomous driving. However, as fully autonomous driving has yet to achieve full technical and legal maturity, human drivers will remain an indispensable partner in autonomous driving systems for a long time. Therefore, human-machine co-driving, as a "transitional, long-term coexistence" intelligent driving model, has become a key research direction for intelligent driving systems. Currently, human-machine co-driving vehicles typically involve both a controller system and a human driver in driving decisions. However, differences in intentions and inconsistent control behaviors between the human and the machine lead to problems such as reduced safety and comfort, trajectory deviation, and human-machine control conflict. This is particularly true in dynamic traffic environments, where driving tasks are complex and dynamic, and driving and vehicle states are subject to constant change. Traditional fixed-weight control strategies struggle to coordinate control dominance between the human and the machine, resulting in reduced overall system performance. To address these challenges, achieving dynamic collaborative control between the human and the machine has become a key technology in the design of human-machine co-driving systems.

[0003] Currently, research on the allocation of control rights in human-machine collaborative control primarily focuses on two technical approaches: First, reinforcement learning-based methods learn dynamic control rights allocation strategies through interactive training with human-machine co-driving scenarios. These methods exhibit strong environmental adaptability and autonomous regulation capabilities. However, these methods lack explicit modeling of human-machine interaction, making it difficult to reflect the constraint logic and policy interpretability in the human-machine interaction relationship. Furthermore, they lack policy coordination in situations where multiple control intentions conflict. Second, game theory-based methods construct a game model between the driver and the controller system and utilize Nash equilibrium and other solutions to optimally determine the actions of the driver and controller. These methods offer rigorous theoretical reasoning and interpretability. However, these methods rely on preset static parameters and rule-driven approaches, lack online learning capabilities, and struggle to cope with dynamic environmental changes.

[0004] The patent CN115071758A proposes a human-machine co-driving control right switching method based on reinforcement learning. According to the driver state and environmental information, the control right between the driver and the controller system is dynamically allocated to improve the adaptability and safety of the co-driving system. However, the reinforcement learning model is only used for the decision layer of the control right switching, and the weight allocation process is not cooperatively integrated with the trajectory tracking control strategy, lacking deep participation in the control strategy execution layer. Moreover, this method cannot depict the strategy conflict process between human and machine, limiting the fine coordination ability in complex driving behavior coordination. The patent CN116729417A builds a non-cooperative game model of the driver and the controller system, and uses the driving safety field and the human-machine conflict degree as evaluation indexes to establish the weight allocation strategy of human-machine co-driving control, realizing the cooperative optimization of path tracking and safety. However, the control weight of this method depends on the pre-set evaluation index, lacking adaptive mechanism, and it is difficult to respond to dynamic driving behavior and risk changes in real time. SUMMARY

[0005] In view of the deficiencies of the prior art and the problems existing in the background art, the present application proposes a human-machine dynamic cooperative steering control method based on the fusion of reinforcement learning and game theory. This method fully combines the advantages of reinforcement learning and non-cooperative game, and builds a human-machine cooperative steering control strategy for dynamic traffic scenarios. Among them, reinforcement learning learns and updates the control strategy weight in real time by perceiving the current driving state and vehicle dynamic information, realizes the dynamic allocation of human-machine control right, has high environmental adaptability, and enhances the response ability of the system to risk and intention changes under uncertain conditions. At the same time, non-cooperative game models the strategy conflict between the driver and the controller, and realizes the Nash equilibrium solution of the optimal control strategy in the dynamic game framework, which can effectively depict the inconsistency between human and machine behaviors and improve the decision rationality. Through the fusion design of reinforcement learning and game theory, this method forms a closed-loop linkage between weight dynamic adjustment and control strategy solving, not only improves the strategy coordination of the system in the human-machine co-driving scene, but also realizes the cooperative optimization of multiple target control objectives such as safety, comfort, tracking accuracy and human-machine conflict, significantly enhancing the cooperative control ability of the human-machine co-driving system in dynamic driving environment.

[0006] The technical solution adopted by the present application to solve the technical problems is as follows:

[0007] The application is a dynamic cooperative steering control method of man-machine fusion of reinforcement learning and game theory, which comprises four modules of expected path module, reinforcement learning module, non-cooperative game module and man-machine co-driving vehicle; wherein, the expected path module generates expected path sequence of the driver and the controller and the expected path at the current time; the reinforcement learning module receives the expected path at the current time, vehicle dynamic information, optimal steering angle of the driver and the controller, converts the received information into the reinforcement learning state space through information preprocessing, and generates the driver control weight through reinforcement learning training; the non-cooperative game module combines the driver control weight, the expected path sequence of the driver and the controller and the state information of the ego vehicle, obtains the optimal steering angle of the driver and the controller through three main steps of prediction model establishment, target function design and Nash equilibrium solution, and realizes the cooperative steering control of the man-machine co-driving vehicle.

[0008] The method comprises the following steps:

[0009] Step 1, expected path generation and processing:

[0010] The expected path module generates expected path sequence of the driver and the controller and the expected path at the current time Wherein R dm and are the expected path sequence of the driver and the expected yaw angle sequence of the driver generated by the driver model, R c and are the expected path sequence of the controller and the expected yaw angle sequence of the controller, R dm1 and are the expected path of the driver and the expected yaw angle of the driver at the current time, R c1 and are the expected path of the controller and the expected yaw angle of the controller at the current time.

[0011] Step 2, reinforcement learning module construction and training:

[0012] Step 2.1, state space design:

[0013] The reinforcement learning module first obtains the expected path R t at the current time, vehicle dynamic information and optimal steering angle of the driver and the controller and Wherein, x and y are the longitudinal and lateral positions of the ego vehicle, v x is the longitudinal speed of the ego vehicle, is the yaw angle of the ego vehicle, β is the mass side slip angle of the ego vehicle, a y is the lateral acceleration of the ego vehicle, x f is the longitudinal position coordinate of the front vehicle, v xfis the longitudinal speed of the preceding vehicle;

[0014] The expected path R at the current moment is obtained by the information preprocessing. t , vehicle dynamic information C t , optimal turning angle between driver and controller and Convert to reinforcement learning state space in, is the driver's lateral displacement error, is the driver's yaw angle error, is the lateral displacement error of the controller, is the controller yaw angle error, DTC is the collision distance between the vehicle and the preceding vehicle, TTC is the collision time between the vehicle and the preceding vehicle, W is the driver weight change rate, e δ is the deviation between the driver's optimal steering angle and the controller's optimal steering angle, w d Control weights for the driver.

[0015] Step 2.2, reward function design:

[0016] The present invention designs a reward function based on four aspects: safety, comfort, tracking accuracy, and human-machine error. Safety considers driving risk and driving stability, comfort considers the lateral acceleration of the vehicle and the rate of change of the driver's weight, tracking accuracy considers the path tracking error between the driver and the controller, and human-machine conflict considers the deviation between the driver's optimal turning angle and the controller's optimal turning angle, as shown in equations (1) to (7):

[0017] R total =τ1R risk +τ2R stab +τ3R lat +τ4R rate +τ5R error +τ6R conf , (1)

[0018]

[0019] R stab =2·((-(tanh 2 (k3·β))+0.5), (3)

[0020]

[0021]

[0022]

[0023]

[0024] Among them, Rtotal is the total reward function of reinforcement learning, R risk is the driving risk reward function, R stab is the driving stability reward function, R lat is the lateral acceleration reward function, R rate is the driver weight change rate reward function, R error is the path tracking error reward function, R conf For human-machine conflict reward function, DTC threshold and TTC threshold are the collision distance threshold and collision time threshold between the vehicle and the preceding vehicle, n is the normalization constant, τ i is the weight of each reward function, σ i is the penalty magnitude, k i and l i are the slope adjustment parameter and the threshold change adjustment parameter respectively, where i = {1, 2, ..., 6}, and tanh represents the hyperbolic tangent function.

[0025] Step 2.3, reinforcement learning training:

[0026] The reinforcement learning training uses the DQN algorithm to train the output driver control weight w d DQN achieves efficient learning of the driver's control weights by constructing a state-action value function and combining it with a deep neural network to approximate the high-dimensional state space. Compared with continuous action methods, the DQN algorithm has advantages in training stability, fast convergence, and simple deployment when dealing with discrete action problems. It is suitable for real-time decision-making on control rights in human-machine co-driving.

[0027] Step 3: Establish and solve the non-cooperative game control model:

[0028] Step 3.1: Prediction model establishment:

[0029] The prediction model uses the vehicle kinematic model to establish a driver-controller interaction model, as shown in Equation (8):

[0030]

[0031] Where,

[0032] Among them, ξ is the state variable, δ d and δ c are the driver's steering angle and the controller's steering angle, A1 is the continuous-time state matrix, B1 and B2 are the continuous-time control input matrices of the driver and the controller, C is the output matrix, and z is the control output; v y ,ω,y, are the lateral velocity, yaw rate, lateral position and yaw angle of the vehicle respectively, m is the vehicle mass, l a and l b are the distances from the center of mass to the front and rear axles, C f and C r are the front and rear tire cornering stiffnesses, I z is the yaw moment of inertia, and G is the steering system transmission ratio.

[0033] Step 3.2: Discretize the model in time domain:

[0034] The driver-controller interaction model is discretized to discretize the continuous system into a finite time domain to facilitate the solution of the non-cooperative game, as shown in Equation (9):

[0035]

[0036] Where,

[0037] Among them, A is the discrete time state matrix, B d and B c are the discrete-time control input matrices of the driver and controller, T s is the system sampling time, k is the current moment, and k+1 is the next moment.

[0038] By iterating formula (9), the predicted time domain N can be obtained p The predicted output under is as follows:

[0039]

[0040] Among them, N p is the prediction time domain, N c is the prediction time domain, N c ≤N p , k+j is the jth sampling step from the current k moment forward in discrete time, j={1,…,N p}.

[0041] Arranging formula (10) can obtain the final discretized prediction output equation, as shown in formula (11):

[0042] Z(k)=Ψξ(k)+Θ d U d (k)+Θ c U c (k), (11)

[0043] Where,

[0044] Among them, Z is the control output matrix, Ud and U c are the control input matrices of the driver and the controller, Ψ is the state variable coefficient matrix, Θ d and Θ c are the control input coefficient matrices of the driver and the controller respectively.

[0045] Step 3.3, objective function design:

[0046] After the driver-controller interaction model is established, objective functions for both are constructed to quantify the path tracking capability and control behavior stability. The objective function consists of two parts: the first part is the path tracking error term, which is used to measure the degree of deviation between the desired path and the actual path; the second part is the angle input penalty term, which is used to reflect the smoothness and physical feasibility of the control behavior, as shown in Equation (12):

[0047]

[0048] Where,

[0049] Among them, V d and V c are the objective functions of the driver and the controller, Q d and Q c are the path tracking weights of the driver and the controller, P d and P c are the steering angle weights of the driver and the controller, and are the single-step path tracking weights of the driver and the controller, and are the single-step turning angle weights of the driver and the controller, respectively, where

[0050] r dm and r c are the desired outputs of the driver and the controller, respectively.

[0051] In order to realize the mechanism of dynamic adjustment of control rights in human-machine collaboration, the dynamic control weight generated by the reinforcement learning module is introduced into the objective function of the driver and the controller, and is used as the weight factor of the path tracking error term. The driver control weight w trained by the reinforcement learning module is d That is the driver's path tracking weight Q d The weight factor is adjusted in real time with the vehicle dynamic information in the time domain, so that the driver and the controller have adaptability in the non-cooperative game process, as shown in formula (13):

[0052]

[0053] Among them, Q d and Q c are the path tracking weights of the driver and controller, w d Control weights for the driver.

[0054] Step 3.4, Nash equilibrium solution:

[0055] The Nash equilibrium solution is solved by simultaneously solving the predictive control optimization problem of both parties in the game, as shown in formula (14):

[0056]

[0057] First, define the path tracking error between the driver and the controller as follows:

[0058]

[0059] Among them, e d and e c are the path tracking errors of the driver and the controller, respectively.

[0060] Then, by substituting Equation (15) into Equation (14), we can obtain the objective functions of the driver and controller with the error term added, as shown in Equation (16):

[0061]

[0062] in, and are the structural transformation matrices of the driver and controller objective function path tracking weight matrices, and are the structural transformation matrices of the driver and controller objective function angle weight matrices respectively.

[0063] Finally, the least square method and convex iteration method are used to find the optimal turning angle sequence of the driver and the controller. and The first element of each sequence is selected as the optimal turning angle of the driver and the controller, as shown in formula (17):

[0064]

[0065] in, and are the optimal turning angles for the driver and the controller respectively.

[0066] Step 4: Human-machine collaborative steering control:

[0067] The human-machine co-driving vehicle receives the optimal turning angle from the driver and the controller and Then, the control angle δ of the human-machine co-driving vehicle can be obtained by adding the two together. f , realizing collaborative steering control of the human-machine co-driving vehicle, that is, the vehicle itself.

[0068] The beneficial effects of the present invention are as follows: The present invention relates to a human-machine dynamic collaborative steering control method that integrates reinforcement learning and game theory. By integrating reinforcement learning with non-cooperative game strategy optimization, it achieves dynamic allocation of control rights between the driver and the controller. While ensuring driving safety, it also balances driving comfort and path tracking accuracy. This method not only adapts to control fluctuations caused by driving and vehicle information, but also effectively alleviates human-machine control conflicts, improves the stability and intelligence of collaborative control, and exhibits excellent system adaptability and real-time performance. It has promising application prospects in the field of human-machine shared control of intelligent driving vehicles. Description of the drawings:

[0069] Figure 1 This is a schematic diagram of the human-machine dynamic collaborative steering control that integrates reinforcement learning and game theory in the present invention. Specific implementation method:

[0070] The present invention will be described in detail below with reference to the accompanying drawings.

[0071] The present invention proposes a human-machine dynamic collaborative steering control method that integrates reinforcement learning and game theory. The method forms a closed-loop control system through four modules: an expected path module, a reinforcement learning module, a non-cooperative game module, and a human-machine co-driving vehicle, thereby realizing collaborative steering control of the human-machine co-driving vehicle. The expected path module generates an expected path sequence of the driver and the controller and an expected path at the current moment. The reinforcement learning module receives the expected path at the current moment, vehicle dynamic information, and the optimal turning angle of the driver and the controller, and converts the received information into a reinforcement learning state space through information preprocessing, and then generates the driver control weight through reinforcement learning training. The non-cooperative game module combines the driver control weight, the expected path sequence of the driver and the controller, and the vehicle state information, and obtains the optimal turning angle of the driver and the controller through three main steps: prediction model establishment, objective function design, and Nash equilibrium solution, thereby realizing collaborative steering control of the human-machine co-driving vehicle. Figure 1 The schematic diagram specifically includes the following steps:

[0072] Step 1: System model construction:

[0073] Step 1.1: Vehicle dynamics model construction:

[0074] The present invention focuses on the lateral and yaw motions of the vehicle, directly taking the front wheel angle of the vehicle as the system input for modeling, and the left and right wheels of the vehicle have the same motion state, and a simplified two-degree-of-freedom vehicle dynamics model is obtained. The lateral and yaw motions of the vehicle are as shown in Equation (18):

[0075]

[0076] Where, F yf =C f α f , F yr =C r α r .

[0077] Among them, m is the mass of the vehicle, v x and v y are the longitudinal and lateral velocities of the vehicle, is the yaw angle of the vehicle, ω is the yaw velocity of the vehicle, δ f The control angle of the human-machine co-driving vehicle, I z is the yaw moment of inertia, l a and l b are the distances from the center of mass to the front and rear axles, F yf and F yr are the lateral forces on the front and rear wheels respectively, C f and C r are the front and rear tire cornering stiffnesses, α f and α r are the sideslip angles of the front and rear wheels, is the lateral velocity derivative of the ego vehicle, is the yaw rate derivative of the vehicle, is the yaw angle derivative of the ego vehicle.

[0078] By deducing Equation (18), we can obtain the simplified vehicle dynamics equation, as shown in Equation (19):

[0079]

[0080] Where y is the lateral position of the vehicle, is the lateral position derivative of the ego vehicle.

[0081] Step 1.2: Driver model construction:

[0082] The present invention adopts a dual-point preview driver model to simulate the steering control of the vehicle by a human driver. The control strategy is as follows: during the execution of the path tracking task, based on the current position information, the driver's preliminary steering intention is generated through the far preview point. Then, the ideal position of the vehicle in a short time is predicted with the help of the near preview point. The prediction result is used as feedback information to dynamically correct the preliminary driver's steering intention, so that the lateral deviation between the vehicle and the future expected position approaches zero. By iterating the above process at each moment, path tracking can be achieved, as shown in Equation (20):

[0083]

[0084] Among them, θ n and θ f They are respectively the near point preview sight angle and the far point preview sight angle, Y en and Y en are the lateral deviations of near point and far point preview, L n and L f They are the near point preview distance and the far point preview distance respectively. is the yaw angle of the vehicle.

[0085] The driver obtains the near point preview sight angle θ through preview operation n and the far point preview sight angle θ f , after the visual information is transmitted and processed by the brain delay module, the expected steering wheel angle can be generated, which is the driver model angle, as shown in formula (21):

[0086]

[0087] Among them, δ dm is the driver model angle, K p Preview steering gain for the driver, K f is the driver preview feedback gain, t p is the driver's reaction time, and s is the complex frequency domain variable in Laplace transform.

[0088] The simplified vehicle dynamics equation shown in equation (19) can be modeled and discretized through equations (8) to (10) in the content of the invention to obtain the driver model control output matrix Z dm ; The driver model control output matrix Z dm The lateral position and yaw angle in are extracted separately, and the driver's expected path sequence and driver's expected yaw angle sequence generated by the driver model can be obtained, as shown in formula (22):

[0089]

[0090] Among them, R dm and The driver expected path sequence and driver expected yaw angle sequence are generated for the driver model respectively.

[0091] Take the driver's expected path sequence R generated by the driver model dm and the driver's expected yaw angle sequence generated by the driver model The first term of can be used to obtain the driver's expected path and the driver's expected yaw angle at the current moment, as shown in formula (23):

[0092]

[0093] Among them, R dm1 and are the driver's expected path and driver's expected yaw angle at the current moment respectively.

[0094] Step 2: Generate and process the expected path:

[0095] The desired path module generates and processes the driver and controller desired path sequences and the expected path at the current moment Among them, R c and are the controller expected path sequence and the controller expected yaw angle sequence, R c1 and are the controller expected path and controller expected yaw angle at the current moment respectively.

[0096] Step 3: Construction and training of reinforcement learning modules:

[0097] Step 3.1, state space design:

[0098] The reinforcement learning module first obtains the expected path R at the current moment t , vehicle dynamic information and the optimal turning angle between the driver and the controller and Where x is the longitudinal position of the vehicle, β is the side slip angle of the vehicle's center of mass, and a y is the lateral acceleration of the vehicle, x f is the longitudinal position of the preceding vehicle, v xf is the longitudinal speed of the preceding vehicle.

[0099] Through information preprocessing, the expected path R at the current moment is t , vehicle dynamic information C t , optimal turning angle between driver and controller and Convert to reinforcement learning state space in, is the driver's lateral displacement error, is the driver's yaw angle error, is the lateral displacement error of the controller, is the controller yaw angle error, DTC is the collision distance between the vehicle and the preceding vehicle, TTC is the collision time between the vehicle and the preceding vehicle, W is the driver weight change rate, e δ is the deviation between the driver's optimal steering angle and the controller's optimal steering angle, w d is the driver control weight; the specific calculation process is as shown in equations (24) to (31):

[0100]

[0101]

[0102]

[0103]

[0104] DTC=|xx f |, (28)

[0105]

[0106] W=|w d (k)-w d (k-1)|, (30)

[0107]

[0108] Among them, k is the current moment and k-1 is the previous moment.

[0109] Step 3.2, reward function design:

[0110] The present invention designs a reward function based on four aspects: safety, comfort, tracking accuracy, and human-machine error. Safety considers driving risk and driving stability, comfort considers the rate of change of the vehicle's lateral acceleration and the driver's weight, tracking accuracy considers the path tracking error between the driver and the controller, and human-machine conflict considers the deviation between the driver's optimal turning angle and the controller's optimal turning angle, as shown in equations (1) to (7) in the invention content.

[0111] Step 3.3, reinforcement learning algorithm selection:

[0112] The reinforcement learning training uses the DQN algorithm to train the output driver control weight w d DQN achieves efficient learning of optimal control weights by constructing a state-action value function and combining it with a deep neural network to approximate a high-dimensional state space. Compared to continuous action methods, the DQN algorithm has advantages in training stability, fast convergence, and simple deployment when dealing with discrete action problems. It is suitable for real-time decision-making on control rights in human-machine co-piloting.

[0113] Step 3.4, reinforcement learning module training:

[0114] The state space of the reinforcement learning training of the present invention is 12-dimensional, the action space is 31 discrete values, and the driver control weight w d The value range is 0-1, the simulation time of each training round is 10 seconds, and the sampling period is 0.01 seconds, that is, each round contains a maximum of 1000 time steps. In each time step, the reinforcement learning agent is based on the current state s t Select the current action at , after executing the action, feedback the next state s t+1 With instant rewards t , the formed empirical quadruple (s t ,a t ,r t ,s t+1 ) are stored in the experience pool and subsequently used for network training through the experience replay mechanism.

[0115] In terms of the network structure of the DQN algorithm, the Q function adopts a three-layer feedforward neural network: the input layer accepts a 12-dimensional state space, and after two layers of ReLU activation hidden layers containing 24 neurons each, it outputs 31 discrete actions.

[0116] For updating the DQN algorithm strategy, the temporal difference method is used. The target Q value output by the target network in DQN and the immediate reward obtained at the current moment are used as the temporal difference target. The target Q value output by the DQN target network and the actual Q value output by the main network are calculated by the mean square error to obtain the loss function. The parameters of the main network in DQN are updated by gradient descent of the loss function until the strategy converges. The loss function L(θ) is as follows:

[0117]

[0118] Among them, r t is the immediate reward, γ is the discount factor used to determine the importance of future rewards, The next state s for the DQN target network t+1 The maximum Q value estimate of all actions, Q(s t ,a t ,θ) is the DQN target network in the current state s t Q value estimation, s t is the current state, a t is the current action, s t+1 is the next state, a t+1 For the next action, E represents the mathematical expectation.

[0119] At this point, the reinforcement learning training module has been completed, realizing effective learning of the dynamic weight allocation strategy in human-machine co-driving.

[0120] Step 4: Establish and solve the non-cooperative game control model:

[0121] Step 4.1: Prediction model establishment:

[0122] By reorganizing the simplified vehicle dynamics equation shown in Equation (19), and using the lateral velocity, yaw rate, lateral position, and yaw angle of the vehicle as state variables, and the steering angle of the driver model and the steering angle of the controller as control inputs, a driver-controller interaction model can be established, as shown in Equation (8) in the invention content.

[0123] Step 4.2: Discretize the model in time domain:

[0124] In order to facilitate the solution of the non-cooperative game problem, the interaction model of the driver-controller system is discretized, and the continuous-time system of formula (8) in the invention is converted into a discrete system in a finite time domain, as shown in formula (9) in the invention.

[0125] By iterating the formula (9) in the invention content, the predicted time domain N can be obtained. p The predicted output under is shown in formula (10) in the invention content.

[0126] By reorganizing formula (10) in the content of the invention, the final discretized prediction output equation can be obtained, as shown in formula (11) in the content of the invention.

[0127] Step 4.3, objective function design:

[0128] After the driver-controller interaction model is established, objective functions for both are constructed to quantify the path tracking capability and control behavior stability. The objective function consists of two parts: the first part is the path tracking error term, which is used to measure the degree of deviation between the desired path and the actual trajectory; the second part is the corner input penalty term, which is used to reflect the smoothness and physical feasibility of the control input, as shown in Equation (33):

[0129]

[0130] Among them, V d and V c are the objective functions of the driver and the controller, Q d and Q c are the path tracking weights of the driver and the controller, P d and P c are the steering angle weights of the driver and the controller, z is the control output, r dm and r c are the expected outputs of the driver model and the controller, δ d and δ c are the driver's steering angle and the controller's steering angle, respectively. k+j is the jth sampling step from the current k moment forward in discrete time, j = {1, ..., N p}.

[0131] Formula (33) is simplified to facilitate the subsequent solution of Nash equilibrium, as shown in Formula (12) in the Summary of the Invention.

[0132] In order to realize the mechanism of dynamic adjustment of control rights in human-machine collaborative control, the dynamic control weight generated by the reinforcement learning module is introduced into the objective function of the driver and the controller, and is used as the weight factor of the path tracking error term. The driver control weight w trained by the reinforcement learning module is d That is the driver's path tracking weight Q d ; The weight factor is adjusted in real time with the vehicle dynamic information in the time domain, so that the driver and the controller have adaptability in the non-cooperative game process, as shown in formula (13) in the invention content.

[0133] Step 4.4, Nash equilibrium solution:

[0134] The Nash equilibrium solution is usually solved by simultaneously solving the predictive control optimization problems of both parties in the game, as shown in formula (14) in the invention content.

[0135] First, the path tracking error between the driver and the controller is defined as shown in Equation (15) in the Summary of the Invention.

[0136] Then, formula (15) in the content of the invention is substituted into formula (14) in the content of the invention to obtain the driver and controller objective functions with the error term added, as shown in formula (16) in the content of the invention.

[0137] Then, the least squares solution of equation (16) in the invention summary is obtained by the QR algorithm, as shown in equation (34):

[0138]

[0139] Where, ζ d (k)=[ξ(k)R dm (k)] T ,ζ c (k)=[ξ(k)R c (k)] T ,

[0140]

[0141] in, and are the optimal control inputs of the driver and the controller, Γ d , Γ c , Λ d , Λ c ,ζ d ,ζ c is the transformation matrix used for the quadratic programming solution in the least squares solution.

[0142] From formula (34), we know that the driver’s optimal control input is It depends not only on the state variable ξ and the driver's expected output sequence R dm , also depends on the controller's desired output sequence R c , and vice versa; this hinders the solution of Nash equilibrium to a certain extent. Therefore, a convex iterative method is introduced to solve the Nash equilibrium between the driver and the automation system, as shown in formula (35):

[0143]

[0144] in, and are the optimal control input sequences for the driver and the controller, respectively.

[0145] Finally, take and The first element is the optimal steering angle between the driver and the controller and As shown in formula (17) in the invention content.

[0146] Step 5: Human-machine collaborative steering control:

[0147] The human-machine co-driving vehicle receives the optimal turning angle from the driver and the controller and Then, the control angle δ of the human-machine co-driving vehicle can be obtained by adding the two together. f , to achieve the coordinated steering control of the human-machine co-driving vehicle, i.e., the self-driving vehicle, as shown in formula (36):

[0148]

[0149] In summary: The present invention relates to a human-machine dynamic collaborative steering control method that integrates reinforcement learning and game theory. The method forms a closed-loop control system through four modules: an expected path module, a reinforcement learning module, a non-cooperative game module, and a human-machine co-driving vehicle, thereby realizing collaborative steering control of the human-machine co-driving vehicle. The expected path module generates the expected path sequence of the driver and the controller and the expected path at the current moment. The reinforcement learning module receives the expected path at the current moment, vehicle dynamic information, and the optimal turning angle of the driver and the controller, converts the received information into a reinforcement learning state space through information preprocessing, and then generates the driver control weight through reinforcement learning training. The non-cooperative game module combines the driver control weight, the expected path sequence of the driver and the controller, and the vehicle state information, and obtains the optimal turning angle of the driver and the controller through three main steps: predictive model establishment, objective function design, and Nash equilibrium solution, thereby realizing collaborative steering control of the human-machine co-driving vehicle. This method realizes the dynamic allocation of control rights between the driver and the controller by integrating reinforcement learning and non-cooperative game strategy optimization. Under the premise of ensuring driving safety, it effectively takes into account both handling comfort and path tracking accuracy. This method can not only adapt to the control fluctuation requirements caused by driver status and vehicle dynamic information, but also effectively alleviate human-machine control conflicts, improve the stability and intelligence of collaborative control, and has good system adaptability and real-time performance. It has good application prospects in the field of human-machine shared control of intelligent driving vehicles.

Claims

1. A human-machine dynamic collaborative steering control method integrating reinforcement learning and game theory, characterized by: This method forms a closed-loop control system through four modules: an expected path module, a reinforcement learning module, a non-cooperative game module, and a human-machine co-driving vehicle, to achieve cooperative steering control of the human-machine co-driving vehicle. The expected path module generates the expected path sequence of the driver and the controller, as well as the expected path at the current moment. The reinforcement learning module receives the expected path at the current moment, vehicle dynamic information, and the optimal steering angle of the driver and the controller, converts the received information into a reinforcement learning state space through information preprocessing, and then generates the driver control weight through reinforcement learning training. The non-cooperative game module combines the driver control weight, the expected path sequence of the driver and the controller, and the vehicle state information, and obtains the optimal steering angle of the driver and the controller through three steps: predictive model establishment, objective function design, and Nash equilibrium solution, thereby achieving cooperative steering control of the human-machine co-driving vehicle. The expected path module is used to generate the expected path sequence of the driver and the controller and the expected path at the current moment where R dm and are the driver expected path sequence and driver expected yaw angle sequence generated by the driver model, R c and are the controller expected path sequence and the controller expected yaw angle sequence, R dm1 and are the driver's expected path and the driver's expected yaw angle at the current moment, R c1 and are the controller expected path and controller expected yaw angle at the current moment respectively; The reinforcement learning module first obtains the expected path R at the current moment t , vehicle dynamic information and the optimal turning angle between the driver and the controller and Where x and y are the longitudinal and lateral positions of the vehicle, respectively, and v x is the longitudinal velocity of the vehicle, is the yaw angle of the vehicle, β is the sideslip angle of the center of mass of the vehicle, a y is the lateral acceleration of the vehicle, x f is the longitudinal position of the preceding vehicle, v xf is the longitudinal speed of the preceding vehicle; The expected path R at the current moment is obtained by the information preprocessing. t , vehicle dynamic information C t , optimal turning angle between driver and controller and Convert to reinforcement learning state space in, is the driver's lateral displacement error, is the driver's yaw angle error, is the lateral displacement error of the controller, is the controller yaw angle error, DTC is the collision distance between the vehicle and the preceding vehicle, TTC is the collision time between the vehicle and the preceding vehicle, W is the driver weight change rate, e δ is the deviation between the driver's optimal steering angle and the controller's optimal steering angle, w d Control weights for the driver; The present invention designs the reward function of the reinforcement learning module based on four aspects: safety, comfort, tracking accuracy, and human-machine error. Safety considers driving risk and driving stability, comfort considers the lateral acceleration of the vehicle and the rate of change of the driver's weight, tracking accuracy considers the path tracking error between the driver and the controller, and human-machine conflict considers the deviation between the driver's optimal turning angle and the controller's optimal turning angle, as shown in the following formula: R total =τ1R risk +τ2R stab +τ3R lat +τ4R rate +τ5R error +τ6R conf , R stab =2·((-(tanh 2 (k3·β))+0.5), Among them, R total is the total reward function of reinforcement learning, R risk is the driving risk reward function, R stab is the driving stability reward function, R lat is the lateral acceleration reward function of the vehicle, R rate is the driver weight change rate reward function, R error is the path tracking error reward function, R conf For human-machine conflict reward function, DTC threshold and TTC threshold are the collision distance threshold and collision time threshold between the vehicle and the preceding vehicle, n is the normalization constant, τ i is the weight of each reward function, σ i is the penalty magnitude, k i and l i are the slope adjustment parameter and the threshold change adjustment parameter, respectively, where i = {1, 2, ..., 6}, and tanh represents the hyperbolic tangent function; Reinforcement learning training is performed according to the total reward function of the reinforcement learning, and the driver control weight w is output. d ; The non-cooperative game module receives the driver's control weight w d , driver and controller expected path sequence R s , vehicle status information The driver and the controller play a non-cooperative game; where v y is the lateral velocity of the vehicle, ω is the yaw angular velocity of the vehicle; The first step of the non-cooperative game between the driver and the controller is to establish a prediction model, which is specifically shown in the following formula: Z(k)=Ψξ(k)+Θ d U d (k)+Θ c U c (k), Where Z is the control output matrix, is the state variable, U d and U c are the control input matrices of the driver and the controller, Ψ is the state variable coefficient matrix, Θ d and Θ c are the control input coefficient matrices of the driver and the controller respectively, and k is the current moment; The second step of the non-cooperative game between the driver and the controller is to design the objective function, which is specifically shown in the following formula: Among them, V d and V c are the objective functions of the driver and the controller, Q d and Q c are the path tracking weights of the driver and the controller, P d and P c are the steering angle weights of the driver and the controller respectively; the driver control weight w trained by the reinforcement learning module d That is the driver's path tracking weight Q d , as shown in the following formula: The third step of the non-cooperative game between the driver and the controller is to solve the Nash equilibrium, and use the least squares method and convex iteration method to find the optimal turning angle sequence of the driver and the controller. and The first element of each sequence is selected as the optimal steering angle for the driver and the controller, respectively, and is passed to the reinforcement learning module and the human-machine co-driving vehicle, as shown in the following formula: in, and are the optimal turning angles for the driver and the controller, respectively; The human-machine co-driving vehicle receives the optimal turning angle of the driver and the controller and Then add the two together to get the control angle δ of the human-machine co-driving vehicle f , realizing collaborative steering control of the human-machine co-driving vehicle, that is, the vehicle itself.

Citation Information

Patent Citations

  • Man-machine co-driving transverse and longitudinal combined control method based on non-cooperative game

    CN116729417A

  • Game-based multi-vehicle collaborative lane changing decision and control method

    CN115230706A

  • Personalized man-machine cooperative control method driven by data and mechanism fusion

    CN117585057A

Cited By

  • A Fuzzy Rule Self-Learning Human-Machine Collaborative Control Method

    CN122166150B