A fuzzy rule self-learning human-machine collaborative control method

By combining fuzzy rule self-learning and reinforcement learning, a multi-source driving risk quantification model is constructed to optimize the allocation of human-machine control rights. This solves the problem of unified representation and dynamic adjustment of multi-source risks in human-machine collaborative control, and improves the stability and safety of vehicles in complex driving scenarios.

CN122166150AActive Publication Date: 2026-06-09CHANGCHUN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHANGCHUN UNIV OF TECH
Filing Date
2026-05-11
Publication Date
2026-06-09

Smart Images

  • Figure CN122166150A_ABST
    Figure CN122166150A_ABST
Patent Text Reader

Abstract

The application is based on a fuzzy rule self-learning human-machine collaborative control method, aiming at solving the problems of insufficient fusion of multi-source risk information and dynamic adjustment of human-machine control right and fuzzy rules. The application relates to the technical field of automatic driving vehicles. A desired path module combines vehicle state information to output a driver desired path and a controller desired path; a driving risk prediction module outputs a driver weight, a controller weight and a current fuzzy rule according to the driver desired path and the controller desired path, the vehicle state information and the optimized fuzzy rule; a reinforcement learning module performs online optimization on the fuzzy rule according to a total driving risk, the vehicle state information and the current fuzzy rule, and outputs the optimized fuzzy rule; and a non-cooperative game module integrates the driver weight, the controller weight, the driver desired path and the controller desired path and the vehicle state information, and optimizes a driver optimal steering angle and a controller optimal steering angle through non-cooperative game, so as to realize collaborative steering control of a human-machine co-driving vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of autonomous vehicle technology, specifically a human-machine collaborative control method based on fuzzy rule self-learning. Background Technology

[0002] With the continuous development of autonomous driving technology and the increasing intelligence of vehicles, current autonomous driving systems are gradually evolving from assisted driving to higher levels of intelligent driving. However, for a considerable period, autonomous driving systems will still struggle to completely operate independently of drivers in complex and dynamic traffic environments, especially in scenarios where risks evolve rapidly and human-machine decision-making may differ. The joint participation of the driver and the autonomous driving system in vehicle control will remain an important human-machine co-driving mode. In this mode, how to ensure driving safety while also considering the driver's intentions, reducing human-machine control conflicts, and improving vehicle operational stability and coordination have become critical technical issues that urgently need to be addressed in the field of intelligent driving.

[0003] In complex driving scenarios, vehicle control is influenced not only by the vehicle's own state but also by factors such as driver fatigue, attention levels, and differences in human-machine control intentions. Particularly in human-machine collaborative control, where the driver and controller jointly participate in vehicle manipulation, the allocation of control authority directly affects the vehicle's safety, ride comfort, and handling stability. Therefore, how to uniformly represent multi-source risk information such as collision risk and fatigue risk, and on this basis, achieve dynamic allocation of human-machine control authority and optimize the solution for the human-machine collaborative controller, remains a key technical problem to be solved.

[0004] Among the existing publicly available technologies, the solutions related to this application can be roughly divided into four categories. The first category is human-machine collaboration solutions. For example, patent CN109885040B discloses a vehicle driving control allocation system in human-machine co-driving, which mainly coordinates and allocates human-machine control based on the driver's state and vehicle state. However, it has not yet established a risk recognition mechanism that uses multiple sources of risk, such as collision risk and fatigue risk, as unified inputs and drives the online evolution of fuzzy rules. The second category is game theory and human-machine co-driving coordination solutions. For example, patent CN120630727B discloses a human-machine dynamic collaborative steering control method that integrates reinforcement learning and game theory. It mainly describes the interaction between the driver and the controller through game modeling and uses equilibrium solution to achieve collaborative control. However, it has not yet extended to the fuzzy rule update layer to achieve dynamic adjustment of weights. The third category is fixed fuzzy rule risk identification and behavior planning solutions. For example, patent CN112068542A uses preset fuzzy rules for risk judgment and driving behavior planning. However, these solutions primarily rely on offline-defined fuzzy rules to complete risk reasoning, risk level judgment, or behavior planning output. These fuzzy rules are typically pre-set based on expert experience and cannot be dynamically adjusted according to changes in the scenario. The fourth category is reinforcement learning optimization solutions. For example, patent CN115071758B discloses a reinforcement learning-based method for switching control in human-machine co-driving, primarily utilizing reinforcement learning to enhance the dynamic adaptability of control switching or control strategy optimization. However, it has not yet directly applied reinforcement learning to the fuzzy rule layer, nor has it constructed a unified closed-loop framework around multi-source risk fusion identification, online rule updates, dynamic allocation of human-machine control, and game-theoretic collaborative control. In summary, while existing technologies have explored human-machine collaborative adjustment, game-theoretic coordinated control, fixed fuzzy rule risk judgment, and reinforcement learning optimization, their technical focus is mostly distributed across different functional levels, and a unified framework spanning multi-source risk fusion identification, online self-learning and updating of fuzzy rules, dynamic generation of human-machine control, and non-cooperative game-theoretic collaborative control has not yet been formed. Existing solutions generally lack continuous transmission and closed-loop coupling of state variables, rule parameters and control decision results between the preceding and following stages. Therefore, in complex dynamic driving scenarios, it is still difficult to achieve a collaborative control process that adjusts synchronously with the evolution of risks, changes in human and machine intentions and changes in vehicle state. Summary of the Invention

[0005] To address the aforementioned issues, this invention proposes a fuzzy rule self-learning human-machine collaborative control method. This method applies deep reinforcement learning to risk cognition and the fuzzy rule layer, enabling online adaptive evolution of fuzzy rules. Simultaneously, a multi-source driving risk quantification model is constructed to characterize the degree of risk and its changing trends. Based on this, while considering safety, stability, and the interpretability of control decisions, a human-machine interaction conflict and adaptive control allocation mechanism is introduced, significantly improving the intelligence and safety level of the human-machine collaborative driving system.

[0006] The technical solution adopted by this invention to solve the technical problem is as follows:

[0007] This invention provides a fuzzy rule self-learning human-machine collaborative control method, including a desired path module, a collision risk identification module, a driving state identification module, a driving risk fuzzy identification module, a driving risk prediction module, a reinforcement learning module, a non-cooperative game module, and a human-machine co-driving vehicle. The desired path module outputs the desired path for the driver and controller based on vehicle state information; the driving risk prediction module outputs driver weights, controller weights, and the current fuzzy rule based on the desired path for the driver and controller, vehicle state information, and optimized fuzzy rules; the collision risk identification module receives vehicle state information and outputs the collision risk; the driving state identification module outputs fatigue risk based on driver eye-tracking signals; the driving risk fuzzy identification module outputs the total driving risk based on the collision risk and fatigue risk; the reinforcement learning module outputs optimized fuzzy rules based on the total driving risk, vehicle state information, and the current fuzzy rule; the non-cooperative game module integrates driver weights, controller weights, the desired path for the driver and controller, and vehicle state information. Under the adjustment of driver weights and controller weights, it optimizes the solution for the optimal steering angle for the driver and controller through prediction model construction, objective function design, and Nash equilibrium solution, thereby achieving collaborative steering control of the human-machine co-driving vehicle.

[0008] The method includes the following steps:

[0009] Step 1: Generate the desired path:

[0010] Step 1.1: The input to the desired path module is vehicle status information. The output is the driver's desired path. and controller expected path ;

[0011] in, The average speed of the vehicle. The average acceleration of the vehicle. Variance of vehicle acceleration;

[0012] Step 2: Construct a driving risk prediction module:

[0013] Step 2.1: Integrate decision-making information:

[0014] The driving risk prediction module is input as the driver's desired path. Controller Desired Path Vehicle status information and optimized fuzzy rules The output is the driver weights. Controller weights and current fuzzy rules The driver weight and controller weight are dynamically determined based on the combined deviation of the driver's and controller's desired paths, vehicle status information, and the current fuzzy rules, and the sum of the driver weight and controller weight is 1.

[0015] in, This represents the optimized fuzzy rule. Represents the fuzzy rules of the previous time step. This represents the parameter adjustment amount for the output of the reinforcement learning policy network;

[0016] Step 2.2, LSTM-K-means clustering collision risk prediction:

[0017] The collision risk identification module is used to identify potential collision risks in the driving environment. The input is vehicle status information; the output is the collision risk. .

[0018] In this invention, before performing risk prediction, the collision risk identification module first performs a structured model of the driver's driving behavior characteristics. Specifically, it extracts three types of driving behavior features—average vehicle speed, average vehicle acceleration, and vehicle acceleration variance—from the vehicle's historical operation dataset, transforming the original vehicle state data into a statistically meaningful driving behavior feature vector.

[0019] Building upon this, the K-means clustering method is introduced to perform unsupervised clustering analysis on the driving behavior feature vectors. Based on the clustering results, driving behaviors are divided into several categories, and the driver's driving style is labeled according to the distribution characteristics of each category in the speed and acceleration dimensions, resulting in three driving style labels: conservative, normal, and aggressive. This step achieves adaptive recognition of driving styles without manual labeling.

[0020] Furthermore, the driving style labels are introduced as auxiliary features into the temporal prediction model to construct a Long Short-Term Memory (LSTM) network that integrates driving style information, thereby modeling and predicting the temporal evolution of vehicle states. By combining driving style labels and vehicle state information, the collision risk at future moments is predicted, and the collision risk is output. .

[0021] Through the above steps, the present invention can simultaneously take into account both the driver's short-term maneuvering behavior and long-term driving style characteristics during the risk prediction process, effectively improving the accuracy and stability of collision risk prediction and providing a reliable basis for subsequent human-machine collaborative control weight allocation.

[0022] The collision risk formula is as shown in equation (1):

[0023] (1)

[0024] in, It is the location of the bicycle. It's the position of the car in front. It is a defined safe distance because the closer the two vehicles are, the faster the risk increases. A power function better reflects the high risk at close range. For exponential parameters, For the change in velocity, This refers to the collision time.

[0025] Step 2.3: Fatigue risk prediction based on eye-tracking features and deep learning:

[0026] The driving state recognition module is based on deep learning and uses an eye tracker to collect raw eye movement signals. The raw input feature is the duration of gaze. blinking frequency Eye movement speed Continuous eye movement features were analyzed, and further state information such as driver attention level, fatigue level, and cognitive load was extracted. This state information was represented as follows: Output fatigue risk ;

[0027] in, Attention level As for the degree of fatigue, This is cognitive load.

[0028] Step 2.4, Total Driving Risk:

[0029] The driving risk fuzzy recognition module identifies collision risks based on preset fuzzy rules optimized by a reinforcement learning module. and fatigue risk Perform fuzzification, fuzzy inference, and defuzzification calculations to output the total driving risk. The specific sub-steps are as follows:

[0030] Step 2.4.1: Confirm Input Variables: In this invention, the input variables for the driving risk fuzzy recognition module mainly include collision risk. and fatigue risk The input variables are processed by the driving risk fuzzy recognition module, and the output is the total driving risk. ;

[0031] Step 2.4.2, Fuzzification Processing: In the calculation of total driving risk, each input variable needs to be fuzzified, that is, its continuous value is converted into a corresponding fuzzy level. For collision risk... and fatigue risk Fuzzification is mapped to corresponding fuzzy rules of high, medium, and low, respectively.

[0032] Step 2.4.3, Fuzzy Rule Inference: Once the input variables have been fuzzified, the next step is to fuse multiple fuzzy risk sources through fuzzy rule inference to calculate the total driving risk. Fuzzy rule inference is based on a predefined fuzzy rule base. Each rule describes the relationship between various input variables and determines the corresponding risk output. The fuzzy rules include:

[0033] Rule 1: If the risk of collision is high and the risk of fatigue is high, then the overall risk of driving is high.

[0034] Rule 2: If the risk of collision is high and the risk of fatigue is moderate, then the overall driving risk is moderate.

[0035] Rule 3: If the risk of collision is high and the risk of fatigue is low, then the overall driving risk is low.

[0036] Rule 4: If the risk of collision is moderate and the risk of fatigue is high, then the overall driving risk is high.

[0037] Rule 5: If the risk of collision is moderate and the risk of fatigue is moderate, then the overall risk of driving is moderate.

[0038] Rule 6: If the risk of collision is moderate and the risk of fatigue is low, the overall driving risk is low.

[0039] Rule 7: If the risk of collision is low and the risk of fatigue is high, then the overall risk of driving is high.

[0040] Rule 8: If the risk of collision is low and the risk of fatigue is moderate, then the overall driving risk is moderate.

[0041] Rule 9: If the risk of collision is low and the risk of fatigue is low, then the overall risk of driving is low.

[0042] Step 2.4.4: Obtain the total driving risk: Use fuzzy rules to... and Blurring out the total driving risk .

[0043] Step 3: Construction of reinforcement learning module and optimization of fuzzy rules:

[0044] Step 3.1, State-space design and fuzzy rule adaptive optimization:

[0045] The reinforcement learning module is input with vehicle status information. Total driving risk and current fuzzy rules The output is the optimized fuzzy rules. ;

[0046] in, The average speed of the vehicle. The average acceleration of the vehicle. Variance of vehicle acceleration;

[0047] Step 3.2, Reward Function Design:

[0048] The reinforcement learning reward function is designed based on six aspects: security assurance capability, risk perception capability, risk evolution trend, coordination of control allocation, robustness of fuzzy rules, and conflict of human-machine control rights, as shown in formulas (2)-(8):

[0049] (2)

[0050] (3)

[0051] (4)

[0052] (5)

[0053] (6)

[0054] (7)

[0055] (8)

[0056] in, To reinforce the total reward function of learning, For the security assurance capability reward function, For the risk perception capability reward function, For the risk evolution trend reward function, Assign a coordinating reward function to control. For the robustness reward function of fuzzy rules, Reward function for conflicts of human-machine control , As a weighting coefficient for security assurance capabilities, , This is the weighting coefficient for risk perception capability. The reward coefficient represents the risk evolution trend. This represents the coordination coefficient for the current allocation of control. This represents the expected proportion of control allocated based on total driving risk. These are the weighting coefficients. For the robustness weight coefficients of fuzzy rules, For the fuzzy rules at the current moment, For the fuzzy rules of the previous moment, This represents the penalty coefficient for conflicts of human-machine control. This represents the combined deviation between the driver's and the controller's desired path.

[0057] Step 3.3: Determine the optimized fuzzy rules:

[0058] Based on vehicle status information, total driving risk, and current fuzzy rules, the fuzzy rules are updated online, and the optimized fuzzy rules are output. To support risk prediction and coordinated control.

[0059] Step 4: Establishing and solving the non-cooperative game model:

[0060] The input to the non-cooperative game theory module is the driver's desired path. Controller Desired Path Vehicle status information, driver weight With controller weights The output is the driver's steering angle. and controller steering angle Specifically, it includes the following sub-steps:

[0061] Step 4.1, Prediction Model Construction:

[0062] The predictive model uses the vehicle dynamics model to establish a driver-controller interaction model, as shown in equation (9):

[0063] (9)

[0064] In the formula, , , , , .

[0065] in, For state variables, The derivative of the state variable. and These are the driver's steering angle and the controller's steering angle, respectively. It is a continuous-time state matrix. and These are the continuous-time control input matrices for the driver and the controller, respectively. For the output matrix, To control the output, , , , These are the vehicle's lateral velocity, yaw rate, lateral position, and yaw angle, respectively. For the overall vehicle quality, and These are the distances from the center of gravity to the front and rear axles, respectively. and These are the lateral stiffness of the front and rear tires, respectively. For the moment of inertia of yaw rotation, This represents the longitudinal velocity.

[0066] Step 4.2, Discretization:

[0067] During human-machine co-driving control, the vehicle needs to receive input from both the driver and the controller in real time, and then the human-machine co-driving controller controls the magnitude of both inputs separately through weights. Therefore, the vehicle system model of the human-machine co-driving controller needs to be discretized, as shown in equation (10):

[0068] (10)

[0069] In the formula, .

[0070] in, The discrete-time state matrix, and These are the discrete-time control input matrices for the driver and the controller, respectively. For system sampling time, For the current moment, For the next moment.

[0071] Therefore, by repeatedly iterating equation (10) using the principle of model predictive control, we can obtain... The predicted output of the step is as shown in equation (11):

[0072] (11)

[0073] in, To predict the time domain, To control the time domain, , for discrete time from the current Moving forward every moment Each sampling step, .

[0074] By rearranging equation (11), we can obtain the final discretized prediction output equation, as shown in equation (12):

[0075] (12)

[0076] In the formula, , , ,

[0077] in, To control the output matrix, and These are the control steering angle input matrices for the driver and the controller, respectively. The state matrix, The coefficient matrix of the state variables. and The control input coefficient matrix for the driver and controller;

[0078] Step 4.3, Objective Function Design:

[0079] After the driver-controller interaction model is established, objective functions for both are constructed to quantify path tracking capability and control behavior stability. The objective function consists of two parts: the first part is the path tracking error term, which measures the degree of deviation between the expected path and the actual path; the second part is the corner input penalty term, which reflects the smoothness and physical feasibility of the control behavior, as shown in equation (13):

[0080] (13)

[0081] In the formula, , , , , , , , .

[0082] in, and Let these be the objective functions for the driver and the controller, respectively. and These are the desired paths for the driver and the controller, respectively. and These are the control steering angle input matrices for the driver and the controller, respectively. and These are the path tracking weights for the driver and the controller, respectively. and These are the steering angle weights for the driver and the controller, respectively. and The single-step path tracking weights are for the driver and the controller, respectively. and These are the single-step steering angle weights for the driver and the controller, respectively. .

[0083] Step 4.4, Solving for Nash Equilibrium:

[0084] The Nash equilibrium is obtained by simultaneously solving two optimization problems. The loss function that minimizes the objective functions of the driver and controller is shown in equation (14):

[0085] (14)

[0086] The path tracking error between the driver and the controller is as shown in equation (15):

[0087] (15)

[0088] in, and These are the path tracking errors for the driver and the controller, respectively.

[0089] Then, substituting equation (15) into equation (14), we obtain the objective functions for the driver and controller with the error term added, as shown in equation (16):

[0090] (16)

[0091] in, and These are the structural transformation matrices of the path tracking weight matrices for the objective functions of the driver and controller, respectively. and These are the structural transformation matrices of the steering angle weight matrix of the objective function for the driver and controller, respectively.

[0092] Finally, the optimal control steering angle input matrix for the driver and controller is obtained using the least squares method and the convex iteration method. and The first element of each sequence is selected as the optimal steering angle for the driver and the controller, respectively, as shown in equation (17):

[0093] (17)

[0094] in, To find the driver's optimal steering angle for game optimization, The optimal steering angle of the controller is obtained by solving the game theory problem. and This is the optimal control steering angle input matrix for the driver and controller.

[0095] Step 5, Human-Machine Cooperative Steering Control Strategy:

[0096] The input for the human-machine co-driving vehicle is the driver's optimal steering angle. and the controller's optimal steering angle The output is vehicle status information. The optimal steering angle for the driver and the optimal steering angle for the controller are linearly superimposed to obtain the final control steering angle of the human-machine co-driving vehicle. This enables coordinated steering control of vehicles driven by both humans and machines.

[0097] The technical effects of this invention are as follows: by introducing a fuzzy rule self-learning mechanism, the current fuzzy rules can be optimized and updated online according to the total driving risk, vehicle status information and current control requirements, thereby improving the adaptability and accuracy of weight allocation in the human-machine collaborative control process, effectively coordinating the conflict of control intentions between humans and machines, improving the path tracking performance, lateral control stability and driving safety of vehicles in complex driving scenarios, and improving the collaboration in the human-machine co-driving process. Attached Figure Description

[0098] Figure 1 This is a flowchart of the method of the present invention.

[0099] Figure 2 This is a flowchart of the LSTM-K-means clustering collision risk prediction process of the present invention. Detailed Implementation

[0100] The following is combined Figure 1 and Figure 2 The present invention will be described in detail.

[0101] This invention proposes a fuzzy rule self-learning human-machine collaborative control method and constructs a closed-loop collaborative control framework for complex driving scenarios, including a desired path module, a driving risk prediction module, a collision risk identification module, a driving state identification module, a driving risk fuzzy identification module, a reinforcement learning module, a non-cooperative game module, and a human-machine co-driving vehicle. The system comprises the following modules: The desired path module outputs the desired path for the driver and controller based on vehicle state information; the collision risk identification module receives vehicle state information and outputs the collision risk; the driving state identification module outputs the fatigue risk based on the driver's eye-tracking signals; the driving risk fuzzy identification module outputs the total driving risk based on the collision risk and fatigue risk; the driving risk prediction module outputs the driver weight, controller weight, and current fuzzy rule based on the desired path for the driver and controller, vehicle state information, and optimized fuzzy rules; the reinforcement learning module outputs optimized fuzzy rules based on the total driving risk, vehicle state information, and current fuzzy rules; and the non-cooperative game theory module integrates the driver weight, controller weight, desired path for the driver and controller, and vehicle state information. Under the adjustment of the driver weight and controller weight, it optimizes the solution for the optimal steering angle for the driver and controller through prediction model construction, objective function design, and Nash equilibrium solution. The human-machine co-driving vehicle obtains the final control steering angle based on the optimal steering angle for the driver and controller, thereby achieving cooperative steering control of the human-machine co-driving vehicle. The specific implementation is as follows:

[0102] Step 1: System Model Construction

[0103] Step 1.1, Vehicle dynamics model construction:

[0104] In this invention, the lateral motion of the vehicle is mainly studied, while the longitudinal speed of the vehicle remains constant, and the influence of the vehicle suspension and aerodynamics is ignored. It is assumed that the motion states of the left and right wheels of the vehicle are consistent, and a simplified two-degree-of-freedom vehicle model is selected as the dynamic model of the vehicle, as shown in the following formula (18).

[0105] (18)

[0106] In the formula, .

[0107] in, For the overall vehicle weight; and These are the vehicle's longitudinal and lateral speeds, respectively. Let be the derivative of the vehicle's yaw angle; and These are the lateral forces of the front and rear tires, respectively. This indicates the final control steering angle of the human-machine co-driving vehicle; This is the moment of inertia of yaw rotation; This refers to the yaw rate; and These are the distances from the center of mass to the front and rear axles, respectively. and These are the lateral stiffness of the front and rear wheel tires, respectively. and These are the slip angles of the front and rear wheels, respectively. The derivative of the vehicle's lateral velocity. This is the derivative of the yaw rate.

[0108] The lateral dynamics state equations of the vehicle are expressed as equations (19)-(22):

[0109] (19)

[0110] (20)

[0111] , (twenty one)

[0112] , (twenty two)

[0113] in, This refers to the lateral position of the vehicle. Let be the derivative of the lateral position of the vehicle. The yaw angle of the vehicle. Let be the derivative of the vehicle's yaw angle.

[0114] Step 1.2, Driver Model Construction:

[0115] The driver model is used to describe the driver's decision-making behavior regarding lateral vehicle control under the influence of risk perception and their own state. Treating the driver as a decision-making agent with the ability to follow the desired path, their lateral control objective can be expressed as minimizing the deviation from the desired path.

[0116] Define the driver's desired path as The actual lateral position of the vehicle is The driver path tracking error is As in equation (23):

[0117] , (twenty three)

[0118] Define the steering angle output by the driver. As in equation (24):

[0119] , (twenty four)

[0120] in, The driver control gain is dynamically adjusted according to the total driving risk, as shown in equation (25):

[0121] (25)

[0122] in, The control gain is used as a reference. For collision risk; For fatigue risk; , This represents the risk sensitivity coefficient.

[0123] Step 2, Design based on real vehicle data:

[0124] This invention uses a CR-V as the test vehicle to verify the aforementioned fuzzy rule self-learning human-machine collaborative control method. The disclosed parameters include: vehicle length 4735mm, width 1859mm, height 1682mm, wheelbase 2791mm, front track 1582mm, rear track 1572mm, curb weight 1700kg, maximum engine power 137kW, and maximum torque 320N·m.

[0125] When establishing a two-degree-of-freedom lateral dynamics model based on the test vehicle, the total vehicle mass m = 1700 kg and the wheelbase L = 2.791 m are taken. In this embodiment, the vehicle's center of gravity position is calibrated according to the front and rear axle load distribution, and the distance from the center of gravity to the front axle is taken. =1.17m, distance from center of mass to rear axle =1.62m; Lateral rotational inertia is taken as... =2950kg·m 2 Front wheel lateral stiffness is taken as =8.2×10⁴ N / rad, rear wheel lateral stiffness is taken as =9.0 × 10⁴ N / rad. Wherein, the... , and The calibration values ​​of the embodiment obtained by combining the vehicle parameters of the model and the vehicle dynamics identification are used to complete the parameterization of the vehicle dynamics model shown in equation (18).

[0126] To demonstrate the feasibility of the fuzzy rule self-learning human-machine collaborative control method, the specific implementation process of the method will be described by combining the vehicle state information at a certain discrete moment, the driver and controller's desired path information, the driving risk prediction results, the fuzzy rule self-learning optimization results, and the human-machine collaborative control output. The process can be implemented and verified on the CarSim and Simulink co-simulation platform.

[0127] Step 3: Generate the desired path:

[0128] Step 3.1, the desired path module, the input of which is vehicle status information. The output is the driver's desired path. and controller expected path ;

[0129] in, The average speed of the vehicle. The average acceleration of the vehicle. Variance of vehicle acceleration;

[0130] Step 4: Construction of the driving risk prediction module:

[0131] Step 4.1: Integrate decision-making information:

[0132] The driving risk prediction module takes the driver's desired path as its input. Controller Desired Path Vehicle status information and optimized fuzzy rules The output is the driver weights. Controller weights and current fuzzy rules ;

[0133] in, For the optimized fuzzy rules, Represents the fuzzy rules of the previous time step. This represents the parameter adjustment amount for the output of the reinforcement learning policy network;

[0134] Step 4.1.1: Generation of driver and controller weights:

[0135] After obtaining the driver's desired path, the controller's desired path, vehicle status information, and the current fuzzy rules, the driving risk prediction module first calculates the comprehensive deviation between the driver's desired path and the controller's desired path to characterize the degree of difference between human and machine control intentions. Secondly, it combines the vehicle status information to characterize the vehicle's current operating state. Then, it adjusts the controller's intervention level based on the current fuzzy rules to generate a controller participation coefficient, and further obtains the driver's weight. and controller weights .

[0136] Furthermore, let the controller participation coefficient be... ,in The controller weights and driver weights are as shown in equation (26):

[0137] (26)

[0138] Among them, the greater the comprehensive deviation between the driver's expected path and the controller's expected path, the more unstable the vehicle's operating state, and the higher the level of controller intervention triggered by the current fuzzy rule, the higher the controller participation coefficient. The larger the value, the greater the controller's weight and the smaller the driver's weight; conversely, the smaller the value, the less the controller's participation coefficient. As the weight of the driver decreases, the weight of the controller increases accordingly.

[0139] This enables the dynamic generation of driver and controller weights, and provides weight inputs for the subsequent non-cooperative game module to solve for the optimal steering angles of the driver and controller.

[0140] Step 4.2, LSTM-K-means clustering collision risk prediction:

[0141] like Figure 2 As shown, it first extracts three types of features from the dataset: average vehicle speed, average vehicle acceleration, and vehicle acceleration variance. Then, it constructs a set of driving behavior features based on the above feature vectors. The feature set The data is fed into a K-means clustering model to perform unsupervised classification of driving behavior, dividing it into several categories. Based on the characteristics of the cluster centers, the driving style is labeled, resulting in three driving style labels: conservative, normal, and aggressive. ∈{conservative, normal, radical}, and finally, perform LSTM prediction.

[0142] The collision risk identification module constructs a time-series prediction model based on a long short-term memory network (LSTM), taking the vehicle's historical state sequence and driving style labels as joint inputs, as shown in equation (27):

[0143] (27)

[0144] Indicates length is Historical characteristic sequence, (·) represents the time series prediction function.

[0145] In this invention, the time-series prediction model network adopts a three-layer LSTM structure to perform layer-by-layer modeling and feature abstraction of the time-series characteristics of vehicle state information.

[0146] The first LSTM layer is used to extract low-level temporal features of vehicle state information, focusing on the short-term variation patterns of vehicle average speed, vehicle average acceleration, and vehicle acceleration variance.

[0147] Let the first layer of LSTM be at time... The output hidden state is .

[0148] in, These are the low-level temporal features output by the first LSTM layer;

[0149] The second-layer LSTM, based on the output of the first layer, further models the mid-level time dependency between the driver's operating behavior and collision risk, to reflect the trend of risk accumulation and evolution over time, as shown in equation (28):

[0150] (28)

[0151] in, The low-level temporal features are the output of the second-layer LSTM. This represents the memory cell state of the second-layer LSTM. This represents the memory cell state before the second-layer LSTM update. , , These represent the input gate, forget gate, and output gate, respectively. For the Sigmoid function, Represents element-wise product. It is the hyperbolic tangent activation function;

[0152] The third LSTM layer is used to extract higher-level long-term temporal features, enhancing the model's ability to express complex driving scenarios and changes in collision risk, thereby improving the stability of collision risk prediction.

[0153] Its state update process is as shown in equation (29):

[0154] (29)

[0155] in, This represents the output of the second-layer LSTM at time t, used to characterize mid-term driving behavior features. The low-level temporal features are output from the third-layer LSTM. This refers to the memory cell state of the third-layer LSTM. This represents the memory cell state before the third-layer LSTM update. , , These represent the input gate, forget gate, and output gate, respectively. It is the hyperbolic tangent activation function;

[0156] The output of the three-layer LSTM network serves as the collision risk prediction result, which is then sequentially input into the driving risk fuzzy recognition module for dynamic adjustment of fuzzy rules and human-machine collaborative control decisions.

[0157] By employing a three-layer LSTM structure, this approach balances prediction accuracy with real-time computation requirements, making it suitable for online driving risk identification and human-machine collaborative control scenarios. It enables collision risk prediction to not only reflect the current vehicle state but also characterize differences in driving behavior and the evolution of risk over time, ultimately outputting the moment... Collision risk value .

[0158] The specific method for calculating collision risk is described in Formula (1) in the invention description.

[0159] Step 4.3: Fatigue risk prediction based on eye-tracking features and deep learning:

[0160] The driving state recognition module is used for real-time identification and quantitative assessment of driver fatigue. It is based on a Transformer deep learning model and utilizes an eye tracker to collect raw input features, including fixation duration. blinking frequency Eye movement speed Continuous driver eye movement feature signals are collected, and further state information such as driver attention level, fatigue level, and cognitive load are extracted. This state information is represented as follows: Output fatigue risk .

[0161] in, Attention level As for the degree of fatigue, This is cognitive load.

[0162] In this invention, to further enhance the modeling capability of driving states and risk evolution, a Transformer-based temporal modeling network is introduced into the driving state recognition module. The Transformer network consists of an embedding layer, a position encoding layer, a multi-head self-attention layer, a feedforward neural network layer, and an output layer. Specific implementation sub-steps are as follows:

[0163] Step 4.3.1, Feature Embedding Layer:

[0164] The feature embedding layer maps driver eye movement features to a unified-dimensional feature vector space, enabling efficient representation and fusion of multi-source continuous features. The original input features include continuous eye movement features such as fixation duration, blink frequency, and eye movement speed, at time... The set of eye movement features can be represented as: ;

[0165] in, Indicates the duration of fixation. Indicates blink frequency, This represents eye movement speed. The feature embedding layer maps the original low-dimensional features to a high-dimensional feature space through a linear mapping method. The embedding process is as shown in equation (30):

[0166] (30)

[0167] in, Embed the weight matrix for the features; It is the bias vector; For a moment The corresponding high-dimensional feature embedding representation.

[0168] Step 4.3.2, Location Encoding Layer:

[0169] Since the Transformer itself lacks temporal awareness, this invention introduces a positional encoding layer after feature embedding to inject temporal order information into the features at each time step. The positional encoding employs a learnable positional encoding method, enabling the model to distinguish state features at different time steps.

[0170] The position-encoded input features are shown in equation (31):

[0171] (31)

[0172] in, Indicates time The feature embedding vector; Indicates time The corresponding learnable positional encoding vector; This is a temporal feature representation after fusing time sequence information.

[0173] Stacking the features from all time steps yields the position-encoded input sequence: ,

[0174] in, Indicates the timing length.

[0175] By introducing a learnable positional encoding method, the model can adaptively learn the relative importance between time steps according to different driving scenarios and driving state characteristics, thereby effectively improving the modeling ability of the evolution characteristics of driver fatigue state over time.

[0176] Step 4.3.3, Multi-head Self-Attention Layer:

[0177] A multi-head self-attention layer is used to characterize the global dependencies between features at different time steps. By constructing query, key, and value vectors, adaptive weighting of the correlation between historical moments and the current state is achieved. The calculation process is shown in equation (32):

[0178] (32)

[0179] in, To prevent the value from being too large, It turns scores into weights. To calculate the attention score, To sum the information in a weighted manner, For query, As key, The value is...

[0180] Step 4.3.4, Feedforward Neural Network Layer:

[0181] The feedforward neural network layer is used to perform nonlinear transformation and feature enhancement on the attention output. It is preferred to adopt a two-layer fully connected structure and introduce nonlinear activation functions between each layer, as shown in equation (33):

[0182] (33)

[0183] in, Represents a nonlinear activation function; , This is the weight matrix; , This is a bias term.

[0184] Step 4.3.5, Residual Connectivity and Layer Normalization:

[0185] Residual connections and layer normalization structures are introduced after the multi-head self-attention layer and the feedforward neural network layer to alleviate the gradient vanishing problem during deep network training and improve network training stability and convergence speed.

[0186] The residual connection is used to superimpose the layer input and the layer output, as shown in equation (34):

[0187] (34)

[0188] in, Presentation layer input, This represents the output of a multi-head self-attention layer or a feedforward neural network layer.

[0189] Subsequently, the results after residual connection are subjected to layer normalization, and the calculation formula is as shown in equation (35):

[0190] (35)

[0191] in, and These represent the characteristic mean and variance, respectively. To prevent smooth terms with a denominator of zero, and These are learnable parameters.

[0192] Step 4.3.6, Output Layer:

[0193] The output layer maps the high-dimensional features encoded by the Transformer into a fatigue risk prediction result vector, and the output result serves as the fatigue risk. To the fuzzy recognition module for driving risks.

[0194] The high-dimensional features encoded by the Transformer are shown in equation (36):

[0195] (36)

[0196] in, For a moment The corresponding high-dimensional feature representation; The high-dimensional features represent the feature dimensions; these features comprehensively reflect the driver's fatigue level, attention change trends, and their potential impact on driving safety.

[0197] Transformer model for historical time windows The eye movement features within the driver's eye are weighted and fused to output the driver's fatigue risk prediction value at the current moment, as shown in equation (37):

[0198] (37)

[0199] Step 4.4, Fuzzy Reasoning of Total Driving Risk:

[0200] The driving risk fuzzy recognition module, by fuzzifying the total driving risk, comprehensively characterizes the overall driving risk level of the vehicle at the current moment, serving as an important basis for human-machine collaborative control and control allocation. The total driving risk is obtained by fusing multiple sub-risk factors, including collision risk. and fatigue risk .

[0201] First, the collision risk and fatigue risk are taken as input variables and fuzzified by the corresponding membership functions respectively. The continuous risk values ​​are mapped to fuzzy sets such as "low risk", "medium risk" and "high risk". The rules for each risk factor under different risk levels are shown in the invention content

[0033] -

[0041] .

[0202] In this embodiment, collision risk With fatigue risk For two input variables, the same membership function is preferred for fuzzy description. For any input risk variable x∈[0,1], three fuzzy subsets, "low risk", "medium risk" and "high risk", are set respectively, and their corresponding membership functions are shown in equations (38)-(40):

[0203] Low-risk membership function: (38)

[0204] Medium-risk membership function: (39)

[0205] High-risk membership function: (40)

[0206] The parameter range corresponding to low risk is [0, 0.2, 0.4], the parameter range corresponding to medium risk is [0.2, 0.5, 0.8], and the parameter range corresponding to high risk is [0.6, 0.8, 1].

[0207] In this embodiment, the total driving risk generated by fuzzy generation The output membership function parameters are preferably set as follows: low risk corresponds to the interval [0, 0.25, 0.45], medium risk corresponds to the interval [0.30, 0.50, 0.70], and high risk corresponds to the interval [0.55, 0.75, 1]. Preferably, trapezoidal membership functions are used for low and high risk, and triangular membership functions are used for medium risk. By setting the above parameters, the risk levels can have a moderate overlap at the interval boundaries, thereby ensuring a smooth transition of fuzzy inference results and avoiding abrupt changes in the allocation of control rights.

[0208] Secondly, based on the fuzzy rules optimized by the reinforcement learning module, fuzzy inference is performed on the fuzzified risk variables.

[0209] Subsequently, the membership distribution of the total driving risk under each risk level is obtained through fuzzy reasoning, and it is transformed into specific risk values ​​using a defuzzification method. The preferred defuzzification method is the centroid method, and its calculation formula is shown in equation (41):

[0210] (41)

[0211] in, The membership function represents the total fuzzy risk. For risk-valued variables.

[0212] Final output: Total driving risk The overall driving risk index is input into the subsequent reinforcement learning module to dynamically adjust the control allocation ratio between the driver and the controller, thereby achieving human-machine co-driving vehicle control that balances safety and collaboration.

[0213] Step 5: Construction of reinforcement learning modules and optimization of rules:

[0214] Step 5.1, State-space design and fuzzy rule adaptive optimization:

[0215] The reinforcement learning module takes vehicle status information as input. Total driving risk and current fuzzy rules The state space can be obtained as follows: The output is the optimized fuzzy rules. ;

[0216] Step 5.2, Reward Function Design:

[0217] This invention designs a reinforcement learning reward function based on six aspects: security assurance capability, risk perception capability, risk evolution trend, coordination of control allocation, robustness of fuzzy rules, and conflict of human-machine control rights; the formulas are shown in formulas (2) to (8) in the invention content.

[0218] Step 5.3, Reinforcement Learning Algorithm Selection:

[0219] In driving risk identification and human-machine collaborative control scenarios, the system state is characterized by strong continuity, high dimensionality, and significant uncertainty, and the fuzzy rules are continuously adjustable variables. Therefore, this invention employs a reinforcement learning algorithm based on policy gradients to optimize and update the fuzzy rules. Preferably, the reinforcement learning algorithm uses either the Deep Deterministic Policy Gradient Algorithm (DDPG) or the Proximal Policy Optimization (PPO) algorithm. DDPG is suitable for fine-tuning in a continuous action space, enabling smooth updates of fuzzy rules; PPO improves the stability and convergence of the training process by introducing a policy constraint mechanism. Both algorithms can meet the stability and real-time requirements of online adaptive adjustment of fuzzy rules.

[0220] The reinforcement learning action space is defined as the adjustment amount of the fuzzy rule parameters: ,

[0221] in, To output the action, This represents the parameter adjustment amount of the reinforcement learning policy network output.

[0222] Step 5.4, Reinforce the learning and training process:

[0223] The reinforcement learning training process includes two stages: offline training and online fine-tuning.

[0224] During the offline training phase, the reinforcement learning agent interacts with the human-machine co-driving system in a simulation environment and learns policies through a large amount of driving scenario data.

[0225] Subsequently, risk prediction and collaborative control are completed based on the updated fuzzy rules, and the reward function is calculated according to the system operation results as shown in equation (42):

[0226] (42)

[0227] in, To mitigate collision risk, To mitigate the risk of fatigue, Indicates the degree of conflict between human and machine control. For each weighting coefficient, i = 1, 2, 3, 4.

[0228] During the online operation phase, the reinforcement learning module updates parameters based on the trained strategy and adjusts fuzzy rules to adapt to changes in driver state and environmental uncertainty, thereby ensuring system safety and stability.

[0229] Step 5.5: Feedback on the optimized fuzzy rules:

[0230] After completing the policy update at the current moment, the reinforcement learning module outputs the optimized fuzzy rules, as shown in equation (43):

[0231] (43)

[0232] in, For the fuzzy rules of the previous moment, The parameters of the reinforcement learning strategy network are adjusted based on the environmental state and the reward function.

[0233] The optimized fuzzy rules The feedback mechanism sends data back to the driving risk prediction module, which is used to update the shape of the risk membership function in real time, assign weights based on the optimized fuzzy rules, and set rule trigger thresholds.

[0234] By continuously feeding back the optimized fuzzy rules and incorporating them into the risk identification and collaborative control decisions of the next moment, the problem of insufficient adaptability of traditional fixed fuzzy rules in complex driving scenarios is avoided, thereby improving the accuracy of risk assessment and the safety and stability of human-machine collaborative control.

[0235] Step 6: Establishing and solving the non-cooperative game model:

[0236] Step 6.1, Prediction Model Construction:

[0237] By rearranging the simplified vehicle dynamics equations shown in equations (19)-(22), and taking the lateral velocity, yaw rate, lateral position and yaw angle of the vehicle as state variables, and the turning angle of the driver model and the turning angle of the controller as control inputs, a driver-controller interaction model can be established, as shown in equation (9) in the invention content.

[0238] Step 6.2, Time-domain discretization of the model:

[0239] To facilitate the solution of non-cooperative game problems, the interaction model between the driver and the controller system is discretized, and the continuous-time system in equation (9) of the invention is transformed into a discrete system in the finite time domain, as shown in equation (10) of the invention.

[0240] By iterating over equation (10) in the invention, the prediction time domain can be obtained. The predicted output is shown in Equation (11) in the invention description.

[0241] By organizing equation (11) in the invention content, the final discretized prediction output equation can be obtained, as shown in equation (12) in the invention content.

[0242] Step 6.3, Objective Function Design:

[0243] After the driver-controller interaction model is established, objective functions for both are constructed to quantify path tracking capability and control behavior stability. The objective function consists of two parts: the first part is the path tracking error term, which measures the degree of deviation between the expected path and the actual trajectory; the second part is the corner input penalty term, which reflects the smoothness and physical feasibility of the control input, as shown in equation (44):

[0244] (44)

[0245] in, and Let these be the objective functions for the driver and the controller, respectively. and These are the path tracking weights for the driver and the controller, respectively. and These are the steering angle weights for the driver and the controller, respectively. To control the output, and These are the expected outputs of the driver model and the controller, respectively. and These are the driver's steering angle and the controller's steering angle, respectively. For discrete time from the current Moving forward every moment Each sampling step, .

[0246] The simplified formula (44) is used for subsequent Nash equilibrium solutions; its objective function is shown in the invention content formula (13).

[0247] Step 6.4, Solving for Nash Equilibrium:

[0248] The solution to the Nash equilibrium is usually achieved by simultaneously solving the prediction and control optimization problems of both sides of the game, as shown in equation (14) in the invention.

[0249] First, the path tracking error between the driver and the controller is defined as shown in equation (15) in the invention.

[0250] Then, substitute equation (15) in the invention into equation (14) in the invention to obtain the driver and controller objective functions with added error terms, as shown in equation (16) in the invention.

[0251] Subsequently, the least squares solution of equation (16) in the invention content is obtained by using the QR algorithm, as shown in equation (45):

[0252] (45)

[0253] In the formula, , .

[0254] in, and These are the optimal control steering angle input matrices for the driver and the controller, respectively. , , , , , This is the transformation matrix used for solving the quadratic programming problem in the least squares solution.

[0255] From equation (45), we know that the driver's optimal control steering angle input matrix is... It depends not only on the state matrix and the driver's expected path It also depends on the controller's desired path. Conversely, this also hinders the solution of the Nash equilibrium to some extent. Therefore, the convex iteration method is introduced to solve the Nash equilibrium between the driver and the automated system, as shown in equation (46):

[0256] (46)

[0257] in, and These are the optimal control steering angle input matrices for the driver and the controller, respectively.

[0258] Finally, take and The first element is the optimal steering angle for the driver and controller. and As shown in formula (17) in the invention description.

[0259] Step 7: Human-machine collaborative control:

[0260] The human-machine co-driving vehicle receives the optimal steering angle from the driver and controller. and Then, the two values ​​are superimposed to obtain the final control steering angle of the human-machine co-driving vehicle. This enables coordinated steering control of vehicles driven by both humans and machines, as shown in equation (47):

[0261] (47)

[0262] In summary, this invention provides a fuzzy rule self-learning human-machine collaborative control method. It introduces reinforcement learning into the online optimization of fuzzy rules, utilizes multi-source state information to achieve adaptive evolution of the risk cognition layer, and constructs a unified risk assessment mechanism that integrates driving style, fatigue risk, and collision risk. This method can effectively alleviate human-machine conflict and improve driving safety and comfort, and has good application prospects in the field of human-machine shared control of intelligent driving vehicles.

Claims

1. A human-machine collaborative control method based on fuzzy rule self-learning, characterized in that, The desired path module combines vehicle status information to output the desired path for the driver and controller; the driving risk prediction module combines the desired path for the driver and controller, vehicle status information, and optimized fuzzy rules to output the driver weight, controller weight, and current fuzzy rules. The collision risk identification module receives vehicle status information and outputs the collision risk. The driving state recognition module outputs fatigue risk based on the driver's eye movement characteristic signals; The driving risk fuzzy recognition module outputs the total driving risk based on the collision risk and fatigue risk. The reinforcement learning module combines the total driving risk, vehicle status information, and the current fuzzy rules to output optimized fuzzy rules. The non-cooperative game module integrates driver weights, controller weights, the expected paths of the driver and controller, and vehicle state information. Under the adjustment of driver and controller weights, it optimizes the optimal steering angle of the driver and controller through three steps: prediction model construction, objective function design, and Nash equilibrium solution, thereby realizing the cooperative steering control of the human-machine co-driving vehicle.

2. The fuzzy rule self-learning human-machine collaborative control method according to claim 1, characterized in that, The input to the desired path module is vehicle status information. ; The output is the desired path between the driver and the controller; where, The average speed of the vehicle. The average acceleration of the vehicle. Variance of vehicle acceleration; The driving risk prediction module is input with the driver's and controller's desired path, vehicle status information, and optimized fuzzy rules. The output is the driver weights. Controller weights and current fuzzy rules ;in, For the optimized fuzzy rules, Represents the fuzzy rules of the previous time step. This represents the parameter adjustment amount for the output of the reinforcement learning policy network; The collision risk identification module is used to identify potential collision risks in the driving environment. The input is vehicle status information; the output is the collision risk. The collision risk formula is: , in, It is the location of the bicycle. It's the position of the car in front. It is a defined safe distance because the closer the two vehicles are, the faster the risk increases. A power function better reflects the high risk at close range. For exponential parameters, For the change in velocity, For the collision time, For a moment The risk of collision; The driving state recognition module is based on deep learning and uses an eye tracker to collect raw eye movement signals. The raw input feature is the duration of gaze. blinking frequency Eye movement speed Continuous eye movement features were analyzed, and further state information such as driver attention level, fatigue level, and cognitive load was extracted. This state information was represented as follows: Output fatigue risk ;in, Attention level, As for the degree of fatigue, For cognitive load, For at any time The risk of fatigue; The driving risk fuzzy recognition module identifies collision risks based on preset fuzzy rules optimized by a reinforcement learning module. and fatigue risk Perform fuzzification, fuzzy inference, and defuzzification calculations to output the total driving risk. ;in, For at any time The overall driving risk is as follows.

3. The human-machine collaborative control method for fuzzy rule self-learning according to claim 1, characterized in that, The reinforcement learning module designs a reinforcement learning reward function based on six aspects: safety assurance capability, risk perception capability, risk evolution trend, control allocation coordination, fuzzy rule robustness, and human-machine control conflict. This function is applied based on vehicle state information and total driving risk. and current fuzzy rules The system performs optimization learning and ultimately outputs optimized fuzzy rules. ; The reinforcement learning reward function is shown in the following formula: , , , , , , , in, To reinforce the total reward function of learning, For the security assurance capability reward function, For the risk perception capability reward function, For the risk evolution trend reward function, Assign a coordinating reward function to control. For the robustness reward function of fuzzy rules, Reward function for conflicts of human-machine control , As a weighting coefficient for security assurance capabilities, , This is the weighting coefficient for risk perception capability. The reward coefficient represents the risk evolution trend. This represents the coordination coefficient for the current allocation of control. This represents the expected proportion of control allocated based on total driving risk. These are the weighting coefficients. For the robustness weight coefficients of fuzzy rules, This represents the combined deviation between the driver's and the controller's desired path. This represents the penalty coefficient for conflicts of human-machine control.

4. The fuzzy rule self-learning human-machine collaborative control method according to claim 1, characterized in that, The non-cooperative game theory module, under the adjustment of driver and controller weights, performs game optimization based on the driver's and controller's desired paths, vehicle state information, driver weights, and controller weights, and outputs the optimal steering angle for the driver and controller. Specifically, it includes the following steps: Step 1: Prediction Model Construction A prediction model is constructed by combining a simplified two-degree-of-freedom vehicle model, as shown in the following formula: , in, To control the output matrix, and These are the control steering angle input matrices for the driver and the controller, respectively. The state matrix, The coefficient matrix of the state variables. and The control input coefficient matrix for the driver and controller; Step 2, Objective function design: Based on the driver's desired path, the controller's desired path, and the vehicle state information obtained from the interactive prediction model, an objective function is designed, as shown in the formula: , in, To control the output matrix, and Let these be the objective functions for the driver and the controller, respectively. and These are the desired paths for the driver and the controller, respectively. and These are the control steering angle input matrices for the driver and the controller, respectively. and These are the path tracking weights for the driver and the controller, respectively. and These are the steering angle weights for the driver and the controller, respectively. Step 3: Solving for Nash equilibrium: The optimal steering angle matrix for the driver and controller is obtained using the least squares method and the convex iteration method. and The first element of each sequence is taken as the optimal steering angle for the driver and the controller, respectively. , , in, To find the driver's optimal steering angle for game optimization, The optimal steering angle of the controller is obtained by solving the game theory problem. and This is the optimal steering angle matrix for the driver and controller; The final control steering angle of the human-machine co-driving vehicle is obtained by linearly superimposing the driver's optimal steering angle and the controller's optimal steering angle. This enables coordinated steering control of vehicles driven by both humans and machines.