A Highly Reliable Trajectory Tracking Control Method and System for Four-Wheel Coordinated Motion Units Based on Deep Reinforcement Learning

CN122569520APending Publication Date: 2026-08-14NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

当前轮转向执行器发生局部失效故障、性能衰退或前轮转角传感器测量异常时,实际前轮转向角往往难以准确跟踪期望前轮转向角控制量,从而导致四轮协调运动单元出现轨迹偏离、转向响应滞后及运行稳定性下降等问题,难以满足复杂应用场景下对高精度、高稳定性和高可靠性的使用要求

Benefits of technology

[0116] (1) This invention integrates the front wheel steering angle signal collected by the front wheel steering angle sensor with the front wheel steering angle estimate obtained by back-reaming the state space equation of the motion unit, and uses an adaptive multi-weight combination for weighted calculation. This can improve the estimation accuracy of the actual front wheel steering angle estimate and the real-time fault variable when the front wheel steering actuator has a local fault or the sensor measurement value is distorted, thereby providing a more reliable fault basis for subsequent missing steering angle compensation control and additional yaw torque distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122569520A_ABST
    Figure CN122569520A_ABST
Patent Text Reader

Abstract

This invention discloses a highly reliable trajectory tracking control method and system for a four-wheel coordinated motion unit based on deep reinforcement learning. The method includes: calculating the actual front wheel steering angle estimate and the real-time fault variable estimate by fusing front wheel steering angle sensor signals and back-derived front wheel steering angle estimates from state-space equations; outputting the desired front wheel steering angle control quantity based on lateral error, heading error, and control law parameters adaptively adjusted by an RBF neural network; calculating the additional yaw torque based on the real-time fault variable estimate and the desired front wheel steering angle control quantity, and implementing real-time dynamic fault-tolerant torque distribution for the four wheels based on a quadratic programming algorithm; and introducing a GRU framework and combining it with deep reinforcement learning to dynamically optimize the weight parameters of multiple control links. This invention improves the trajectory tracking accuracy, yaw stability, and continuous operation reliability of the four-wheel coordinated motion unit under conditions of partial failure of the front wheel steering actuator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electromechanical system control technology, specifically to a highly reliable four-wheel coordinated motion unit trajectory tracking control method and system based on deep reinforcement learning. Background Technology

[0002] In the process of a four-wheel coordinated motion unit (FMU) performing trajectory tracking tasks, continuous and stable motion control is typically required under complex road conditions, varying operating conditions, and high-precision control requirements. For a FMU employing independent four-wheel drive and coordinated with front-wheel steering actuators for steering control, trajectory tracking performance depends not only on the calculation accuracy of the desired front wheel steering angle control value but also on the system's comprehensive adjustment capability for lateral error, heading error, center of gravity sideslip state, and yaw attitude changes. Most existing trajectory tracking control methods are based on ideal models and are primarily designed for path following under fault-free conditions. When the front wheel steering actuator experiences partial failure, performance degradation, or abnormal measurement by the front wheel steering angle sensor, the actual front wheel steering angle often fails to accurately track the desired front wheel steering angle control value. This leads to problems such as trajectory deviation, lag in steering response, and decreased operational stability in the FMU, making it difficult to meet the requirements for high precision, high stability, and high reliability in complex application scenarios.

[0003] Furthermore, existing fault-tolerant trajectory tracking control methods often suffer from problems such as insufficient real-time fault variable identification accuracy, untimely additional yaw torque compensation, and difficulty in coordinating weight parameters across multiple control components when handling local faults in the front wheel steering actuator. On the one hand, traditional fault estimation methods often rely on single sensor signals or fixed-weight fusion methods. When front wheel steering angle measurement is distorted or the reliability of the information source changes, it is difficult to accurately obtain the true front wheel steering state and real-time fault variables, thus affecting the subsequent missing steering angle compensation control and additional yaw torque calculation effects. On the other hand, traditional four-wheel torque distribution methods often use fixed rules or static parameter configurations, making it difficult to coordinate and distribute the output torque of the four wheels in real time according to the fault severity and operating status, easily leading to insufficient additional yaw torque compensation and reduced trajectory holding capability. At the same time, there is a clear parameter coupling relationship between fault estimation, front wheel steering missing compensation control, additional yaw torque calculation, and four-wheel dynamic fault-tolerant torque distribution. Relying on manual experience for parameter tuning makes it difficult to simultaneously consider the collaborative optimization of multiple control modules and dynamic adaptability under complex operating conditions. Summary of the Invention

[0004] The technical problem to be solved by this invention is to overcome the shortcomings of the existing technology and propose a highly reliable four-wheel coordinated motion unit trajectory tracking control method that can take into account the local fault estimation of the front wheel steering actuator, the front wheel steering loss compensation control, the additional yaw torque compensation and multi-parameter collaborative optimization.

[0005] To solve the above problems, the present invention adopts the following technical solution:

[0006] First, this invention proposes a highly reliable trajectory tracking control method for a four-wheel coordinated motion unit based on deep reinforcement learning, comprising the following steps:

[0007] S1. Acquire front wheel steering angle sensor signal The estimated value of the front wheel steering angle is obtained by back-deriving the state-space equation of the motion unit. , And adopt adaptive multi-weight combination Perform weighted fusion to calculate the estimated value of the actual front wheel steering angle. and real-time fault variable estimates ;

[0008] S2. Acquire lateral displacement signal and heading displacement signal Construct the comprehensive combined error of front wheel steering angle and the nonlinear calculus convergence function of the multi-weighted comprehensive error of the front wheel steering angle. An RBF neural network is used to approximate the equivalent disturbance caused by the fault online. Combined with the actual estimated front wheel steering angle and real-time fault variable estimates Output the desired front wheel steering angle control value Perform missing corner compensation control;

[0009] S3, Receive real-time fault variable estimates and desired front wheel steering angle control amount A multi-weighted proportional-integral combined state error convergence function is constructed, and the additional yaw moment control quantity of the fused second-order continuous variable structure is calculated, outputting the additional yaw moment. Then, based on the quadratic programming algorithm, Distribute the torque to the four wheels and calculate the real-time dynamic fault-tolerant torque of each wheel. , , , ;

[0010] S4. By introducing the GRU framework and combining it with deep reinforcement learning, the weight parameters involved in the estimation of the front wheel steering angle, the missing steering angle compensation control, the solution of the additional yaw torque, and the dynamic fault-tolerant torque distribution of the four wheels are jointly and dynamically optimized to achieve the overall optimal trajectory tracking control capability.

[0011] Preferably, step S1 specifically includes the following steps:

[0012] S101. Obtain the front wheel angle using the front wheel angle sensor installed on the front wheel of the motion unit. ;

[0013] S102. Using the state-space equations of the motion unit, the two types of front wheel steering angles are derived by reverse deduction. , The state-space equation of the motion unit is expressed as:

[0014] ,

[0015] in, , , , , , ; This refers to lateral error; For heading error; For the first The linear lateral stiffness of each wheel, =1, 2, 3, 4; For overall quality; For partial failure estimation of the front wheel steering actuator; , These are the distances from the center of mass to the front and rear axles, respectively. To bypass Moment of inertia of the shaft; , The longitudinal velocity and lateral velocity of the center of mass are respectively determined; The desired heading angular velocity;

[0016] Two types of front wheel steering angles can be obtained by directly inversely solving the state-space equations of the motion unit. , The estimated value is calculated using the following formula:

[0017] ,

[0018] ,

[0019] S103. Select weighting coefficients and Calculate the partial failure estimation of the front wheel steering actuator The calculation formula is:

[0020] ,

[0021] In the formula, >0, >0 and + =1;

[0022] Real-time fault variable estimates Represented as:

[0023] ,

[0024] In the formula, This is the control law for the front wheel steering angle actuator.

[0025] Preferably, step S2 specifically includes the following steps:

[0026] S201. Acquire the lateral displacement signal from the sensor. and heading displacement signal Calculate the lateral error and heading error ,in, and These are the desired lateral trajectory and the desired heading trajectory; the combined weighting coefficients for the front wheel steering angle error are selected. and Calculate the combined error of the front wheel steering angle. ;

[0027] S202, Based on the comprehensive combined error of the front wheel steering angle Select the front wheel steering angle convergence combination weight coefficient , and The nonlinear calculus convergence function for calculating the multi-weighted comprehensive error of the front wheel steering angle is specifically expressed as:

[0028] ;

[0029] S203, Construct the convergence speed of the comprehensive combined error including the front wheel steering angle. Equivalent disturbance amplitude caused by failure Front wheel steering angle actuator control law :

[0030] ,

[0031] In the formula,

[0032] ,

[0033] In the formula, , The linear lateral stiffness of the front and rear wheels are respectively. For trajectory curvature;

[0034] S204. Adaptive optimization of the control law for the front wheel steering angle actuator is performed by introducing a radial basis function neural network, based on the nonlinear calculus convergence function of the multi-weighted comprehensive error of the front wheel steering angle. As network input, the convergence speed of the combined error of the front wheel steering angle. and As the network output, the nonlinear disturbance caused by a local fault in the steering device is approximated online using the Gaussian radial basis function, and the parameters are adjusted in real time. and ; and The network output is represented as:

[0035] ,

[0036] In the formula, These are the weights of the radial basis function neural network; The Gaussian function for a radial basis function neural network is expressed as follows:

[0037] ,

[0038] In the formula, As the center vector, For width parameters;

[0039] S205, Based on the nonlinear calculus convergence function of the multi-weighted comprehensive error of the front wheel steering angle. The adaptive update criterion for network parameters is constructed based on the trend of their changes, and the weights of the radial basis function neural network are adjusted accordingly. , center vector and width parameter Perform online updates, defining the adaptability metric for the weights as follows:

[0040] ,

[0041] Representing the present The convergence trend, if This indicates that the system state is converging toward the convergence target surface; if This indicates poor convergence; To optimize the objective, network parameters are adjusted. , , Perform online updates, and thus affect and Changes are then fed back to the control law. In this process, a closed-loop adaptive adjustment is formed.

[0042] Preferably, step S3 specifically includes the following calculation steps:

[0043] S301. Obtain the centroid sideslip angle via GPS / GNSS system. And obtain the yaw angle through the gyroscope in the IMU. The state vector that makes up the four-wheel coordinated motion unit is specifically represented as follows:

[0044] ,

[0045] S302. Construct the comprehensive state error of the adaptive weighted four-wheel coordinated motion unit. This characterizes the tracking effect of the four-wheel coordinated motion unit on the yaw moment, specifically expressed as:

[0046] ,

[0047] In the formula, For reference only. , These are weight parameters;

[0048] S303. Calculate the convergence function of the multi-weighted proportional-integral combined state error:

[0049] ,

[0050] In the formula, These are weight parameters;

[0051] S304. Construct an additional yaw moment control law that integrates the second-order continuous variable structure, as the input target for subsequent four-wheel torque distribution, specifically expressed as:

[0052] ,

[0053] Based on the state vector obtained in step S301, the comprehensive state error constructed in step S302, and the convergence function of the multi-weight proportional-integral combined state error obtained in step S303, the additional yaw moment control quantity of the fused second-order continuous variable structure is calculated, and the additional yaw moment ΔM is output. Z ;

[0054] S305. The longitudinal tire force of the wheel is calculated using the wheel-end drive torque, wheel angular velocity, wheel angular acceleration, and wheel radius. The lateral tire force is calculated using an IMU, steering angle sensor, and wheel speed sensor. The vertical force of the wheel is obtained through a wheel load sensor. ,in =1, 2, 3, 4;

[0055] S306. Construct an objective function based on a quadratic programming algorithm for distributing torque across four wheels. Divided into two parts, for and The weighted sum; defining the objective function for optimizing the overall adhesion utilization rate of the four wheel tires in the first part. The function expression is as follows:

[0056] ;

[0057] In the formula, It is the tire-road adhesion coefficient;

[0058] S307, Define the second part of step S306: the objective function for tracking the desired yaw moment and total longitudinal tire force. The function expression is as follows:

[0059] ;

[0060] In the formula, , , , For the expected , ;

[0061] S308. Based on steps S306 and S307, a total objective function based on a quadratic programming algorithm is constructed for distributing the torque of the four wheels, using weighting coefficients. The optimization objectives of overall tire adhesion utilization rate, desired yaw moment, and total longitudinal tire force tracking are weighted and coordinated. Under the constraint of output torque of each wheel, the optimal output torque of the four wheels is solved to achieve additional yaw moment compensation and coordinated torque distribution of the four wheels under local steering system failure conditions in the four-wheel coordinated motion unit. Specifically, this is expressed as follows:

[0062] ;

[0063] In the formula, These are the weighting coefficients.

[0064] Preferably, in step S4, the GRU framework is introduced, and combined with deep reinforcement learning, the weight parameters are jointly and dynamically optimized, specifically including the following steps:

[0065] S401, using real-time fault variable estimates lateral error and heading error Sideslip angle of center of mass A state space for a deep deterministic policy gradient fusion GRU is constructed to characterize the comprehensive operational state under trajectory tracking error, fault state, and attitude change, specifically represented as follows:

[0066] ;

[0067] S402. Construct a deep deterministic strategy gradient action space fused with GRU, wherein the action space includes the weighting coefficients of the integrated combined error, the weighting coefficients of the weighted fusion, and the weighting coefficients of the overall objective function, to jointly adjust key parameters in front wheel steering control, fault estimation, and torque coordination distribution, thereby achieving collaborative optimization of multiple control links; the action space is specifically represented as follows:

[0068] ;

[0069] S403. Define the reward function for the deep deterministic policy gradient of the fused GRU. The reward function consists of three parts. The first part is used to suppress excessive lateral bias. The reward function is expressed as follows:

[0070] ;

[0071] S404. Define the second part of the reward function, specifically as follows:

[0072] ;

[0073] S405. Define the third part, the reward function, as follows:

[0074] ;

[0075] In the formula, and These are the boundary coefficients and the velocity, respectively. and road surface friction coefficient related;

[0076] S406. The lateral deviation reward, heading deviation reward, and stability reward are weighted and combined to construct the total reward function, which is specifically expressed as follows:

[0077] ;

[0078] S407. Construct the action-value function in the deep deterministic policy gradient algorithm. By representing the cumulative benefit that can be obtained by taking corresponding parameter adjustment actions in the current state, evaluate the long-term impact of different parameter combinations on the trajectory tracking control effect, specifically expressed as follows:

[0079] ;

[0080] In the formula, This is the state at this moment. For this moment's action, Reward this moment. For the state at the next moment, for and The expected distribution;

[0081] S408. Construct a target value function based on the reward function and the target network output, aiming to minimize the deviation between the current action value estimate and the target value. Train the Critic network with the following expression:

[0082] ;

[0083] In the formula, As a discount factor, ;

[0084] S409. Construct the Critic network loss function to improve the accuracy of value assessment by minimizing the action value estimation error. Its expression is:

[0085] ;

[0086] S410. Update the Actor network based on the policy gradient output by the Critic network, as shown in the expression:

[0087] ;

[0088] S411. A soft update mechanism is used to update the parameters of the target Actor network and the target Critic network, allowing the target network to slowly track changes in the online network and reduce training oscillations. The calculation process is as follows:

[0089] ,

[0090] In the formula, This is a smoothing factor.

[0091] Preferably, step S4 introduces the GRU framework and combines it with deep reinforcement learning to perform joint dynamic optimization of the weight parameters, and also includes the following steps:

[0092] Current observation status Hidden state from the previous moment A single-layer gated recurrent unit (GRU) in the common input Actor network extracts historical state information through reset and update gates. The reset gate... and the update gate They are represented as follows:

[0093] ,

[0094] ,

[0095] In the formula, It is the Sigmoid activation function. This represents the concatenated vector of the hidden state from the previous time step and the observed state at the current time step.

[0096] According to the reset door Filter historical state information and calculate candidate hidden states. Its expression is:

[0097] ,

[0098] in, The hyperbolic tangent activation function is used. Represents the Hadamard product;

[0099] According to the update gate Hidden state from the previous moment and candidate hidden state Perform fusion to obtain the hidden state at the current moment. Its expression is:

[0100] ,

[0101] Hide the current state With current observation status Feature splicing is performed and fed into subsequent fully connected layers of the Actor network, outputting continuous control actions. Its expression is:

[0102] ,

[0103] In the formula, Represents the policy function of the Actor network;

[0104] The continuous control action Mapped to a multi-control module weight parameter vector:

[0105] ,

[0106] The weights are used for adaptive multi-weight combination in front wheel steering angle estimation, error combination weight in missing steering angle compensation control, and objective function weight in additional yaw moment solution and four-wheel dynamic fault-tolerant torque distribution, respectively, to achieve joint dynamic optimization of the weight parameters of each link.

[0107] Meanwhile, this invention proposes a highly reliable four-wheel coordinated motion unit trajectory tracking control system based on deep reinforcement learning, which consists of four parts: a front wheel steering actuator local failure fault estimation module, a front wheel steering actuator angle fault-tolerant control module, a four-wheel yaw torque coordinated distribution additional control module, and a trajectory tracking deep reinforcement learning parameter optimization module. Specifically, it includes:

[0108] The front wheel steering actuator partial failure estimation module is used to acquire signals from the front wheel steering angle sensor. The estimated front wheel steering angle obtained by back-deriving from the state-space equations of the motion unit is integrated. and Adaptive multi-weight combination is adopted Weighted fusion is performed to calculate the estimated value of the actual front wheel steering angle. and real-time fault variable estimates and the estimated value of the actual front wheel steering angle The output is sent to the front wheel steering actuator angle fault-tolerant control module, which then outputs the estimated value of the real-time fault variable. The output is sent to the four-wheel yaw torque coordination and distribution additional control module;

[0109] The front wheel steering actuator angle tolerance control module is used to acquire lateral displacement signals from the sensors. and heading displacement signal Construct the comprehensive combined error of front wheel steering angle and the nonlinear calculus convergence function of the multi-weighted comprehensive error of the front wheel steering angle. An RBF neural network is used to approximate the equivalent disturbance caused by the fault online. Combined with the actual estimated front wheel steering angle and real-time fault variable estimates The desired front wheel steering angle control value output Perform missing corner compensation control;

[0110] An additional control module for coordinated distribution of yaw moment across four wheels is used to receive real-time estimates of fault variables. and desired front wheel steering angle control amount A multi-weighted proportional-integral combined state error convergence function is constructed, and based on this function, the additional yaw moment control quantity of the fused second-order continuous variable structure is calculated, outputting the additional yaw moment. Then, based on the quadratic programming algorithm, Distributed to four wheels, enabling real-time dynamic fault-tolerant torque for all four wheels. , , , calculate;

[0111] The trajectory tracking deep reinforcement learning parameter optimization module is used to dynamically optimize the weight parameters of each control module by introducing the GRU framework with strong memory of historical state information and combining the online decision-making and global search capabilities of deep reinforcement learning. The optimized parameters are then sent to the front wheel steering actuator local failure fault estimation module, the front wheel steering actuator angle fault-tolerant control module, and the four-wheel yaw torque coordination distribution additional control module to achieve the overall optimal trajectory tracking control capability.

[0112] Meanwhile, the present invention proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps described in the present invention.

[0113] Furthermore, the present invention proposes an electronic device including a processor and a memory, wherein the memory stores a computer program, characterized in that, when the computer program is executed by the processor, it implements the steps of the method described in the present invention.

[0114] Finally, the present invention proposes a computer program product, comprising a computer program, characterized in that, when the computer program is executed by a processor, it implements the steps of the method described in the present invention.

[0115] Compared with the prior art, the present invention adopts the above technical solution and has the following beneficial effects:

[0116] (1) This invention integrates the front wheel steering angle signal collected by the front wheel steering angle sensor with the front wheel steering angle estimate obtained by back-reaming the state space equation of the motion unit, and uses an adaptive multi-weight combination for weighted calculation. This can improve the estimation accuracy of the actual front wheel steering angle estimate and the real-time fault variable when the front wheel steering actuator has a local fault or the sensor measurement value is distorted, thereby providing a more reliable fault basis for subsequent missing steering angle compensation control and additional yaw torque distribution.

[0117] (2) By constructing a nonlinear calculus convergence function for the comprehensive combined error of the front wheel steering angle and the multi-weight comprehensive error of the front wheel steering angle, and combining it with the RBF neural network to approximate the equivalent disturbance caused by the fault online and to adaptively adjust the control law parameters, this invention can more effectively compensate for the missing steering angle caused by the local fault of the front wheel steering actuator, and improve the front wheel steering tracking accuracy, heading holding capability and running stability during the trajectory tracking process.

[0118] (3) This invention constructs a multi-weight proportional integral combined state error convergence function, uses the additional yaw moment control law of the fusion second-order continuous variable structure to output additional yaw moment, and further combines the longitudinal tire force, lateral tire force and vertical force of the wheel to establish a four-wheel torque distribution based on the quadratic programming algorithm. When the steering ability is reduced due to a local fault of the front wheel steering actuator, the yaw moment gap can be compensated in time and the coordinated distribution of the output torque of the four wheels can be realized, thereby improving the trajectory holding ability, yaw stability and continuous operation reliability of the four-wheel coordinated motion unit.

[0119] (4) This invention constructs a state space containing fault state, trajectory error and attitude information, constructs an action space containing multiple key weight parameters, and introduces a GRU framework with historical state information memory in deep reinforcement learning. By combining lateral deviation, heading deviation and stability reward to jointly and dynamically optimize multiple control parameters, it can improve the parameter matching and coordination between multiple control links, thereby improving the overall trajectory tracking control capability of the system under complex working conditions and local fault conditions of the front wheel steering actuator. Attached Figure Description

[0120] Figure 1 This is an architecture diagram of a highly reliable four-wheel coordinated motion unit trajectory tracking control method based on deep reinforcement learning, which is involved in an embodiment of the present invention.

[0121] Figure 2 This is a schematic diagram of the Actor network structure fused with GRU in the trajectory tracking deep reinforcement learning parameter optimization module according to an embodiment of the present invention.

[0122] Figure 3 This is a schematic diagram of the four-wheel coordinated motion unit structure involved in an embodiment of the present invention.

[0123] Figure 4 This is a comparison chart of the training of the trajectory tracking deep reinforcement learning parameter optimization module involved in the embodiments of the present invention. Detailed Implementation

[0124] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.

[0125] Example 1: This example is a specific implementation of a highly reliable four-wheel coordinated motion unit trajectory tracking control method based on deep reinforcement learning. It can run in an industrial environment consisting of a host computer control platform and a four-wheel coordinated motion unit hardware platform. The host computer control platform is used to run the trajectory tracking control algorithm, the deep reinforcement learning parameter optimization algorithm, and the data acquisition and visualization program. The host computer control platform uses a PC with an Intel Core i7 series or higher processor, 16GB or more of memory, and Windows 11 operating system. The deep reinforcement learning algorithm runtime environment includes Python, PyTorch, and MATLAB / Simulink co-simulation environments. The four-wheel coordinated motion unit hardware platform includes a front wheel steering device, four independent drive motors, wheel speed sensors, a front wheel angle sensor, an IMU inertial measurement unit, a GPS / GNSS positioning module, and wheel load sensors. The system includes: a front wheel steering angle sensor for real-time acquisition of front wheel steering angle signals; an IMU (Inertial Measurement Unit) for acquiring yaw rate, attitude information, and vehicle motion status information; a GPS / GNSS positioning module for acquiring the position, velocity, and sideslip status information of the moving units; wheel speed sensors for acquiring the rotational speed of each wheel; and wheel load sensors for acquiring the vertical load information of each wheel. (See also...) Figure 1 This includes the following steps:

[0126] S1. Use the front wheel steering actuator partial failure fault estimation module to collect front wheel steering angle sensor signals. The estimated front wheel steering angle obtained by back-deriving from the state-space equations of the motion unit is integrated. and Adaptive multi-weight combination is adopted Weighted fusion is performed to calculate the estimated value of the actual front wheel steering angle. and real-time fault variable estimates and the estimated value of the actual front wheel steering angle The output is sent to the front wheel steering actuator angle fault-tolerant control module, which then outputs the estimated value of the real-time fault variable. The output is supplied to the four-wheel yaw torque coordination and distribution additional control module.

[0127] S2. Use the front wheel steering actuator angle tolerance control module to collect the lateral displacement signal from the sensor. and heading displacement signal Construct the comprehensive combined error of front wheel steering angle and the nonlinear calculus convergence function of the multi-weighted comprehensive error of the front wheel steering angle. An RBF neural network is used to approximate the equivalent disturbance caused by the fault online. Combined with the actual estimated front wheel steering angle and real-time fault variable estimates The desired front wheel steering angle control value output To implement corner compensation control.

[0128] S3. Use a four-wheel yaw moment coordination distribution additional control module to receive real-time fault variable estimates. and desired front wheel steering angle control amount A multi-weighted proportional-integral combined state error convergence function is constructed, and based on this function, the additional yaw moment control quantity of the fused second-order continuous variable structure is calculated, outputting the additional yaw moment. Then, based on the quadratic programming algorithm, Distributed to four wheels, enabling real-time dynamic fault-tolerant torque for all four wheels. calculate.

[0129] S4. Using a trajectory tracking deep reinforcement learning parameter optimization module, by introducing a GRU framework with strong memory of historical state information, and combining the online decision-making and global search capabilities of deep reinforcement learning, the weight parameters of the front wheel steering actuator angle fault-tolerant control module, the front wheel steering actuator local failure estimation module, and the four-wheel yaw moment coordination and allocation additional control module are respectively optimized. Dynamic optimization is performed, and the optimized parameters are sent to the front wheel steering actuator local failure fault estimation module, the front wheel steering actuator angle fault tolerance control module, and the four-wheel yaw torque coordination distribution additional control module to achieve the overall optimal trajectory tracking control capability.

[0130] Specifically, step S1 includes the following steps:

[0131] S101. Obtain the front wheel angle using a front wheel angle sensor installed on the front wheel of the motion unit. .

[0132] S102 uses the state-space equations of the motion unit to deduce the two types of front wheel steering angles δ. c1 δ c2 The state-space equation of the motion unit can be expressed as:

[0133]

[0134] In the formula, , , , , , .

[0135] like Figure 3As shown, the four-wheel coordinated motion unit adopts a four-wheel independent drive structure and achieves front wheel steering control through a front wheel steering actuator. The figure shows the abstract structure and dimensional parameters of the four-wheel coordinated motion unit. This refers to lateral error; For heading error; For the first The linear lateral stiffness of each wheel, =1, 2, 3, 4; For overall quality; For partial failure estimation of the front wheel steering actuator; , These are the distances from the center of mass to the front and rear axles, respectively. To bypass Moment of inertia of the shaft; , The longitudinal velocity and lateral velocity of the center of mass are respectively determined; The desired heading angular velocity. These dimensional parameters provide a unified geometric and dynamic basis for subsequent fault estimation, missing angle compensation control, and four-wheel dynamic fault-tolerant torque distribution.

[0136] Two types of front wheel steering angles can be obtained by directly inversely solving the state-space equations of the motion unit. Estimated value:

[0137]

[0138]

[0139] S103. Select weighting coefficients and Calculate the partial failure estimation of the front wheel steering actuator The calculation formula is

[0140]

[0141] In the formula, >0, >0 and The trajectory tracking deep reinforcement learning parameter optimization module is used for optimization.

[0142] In addition, real-time fault variable estimates It can be represented as:

[0143]

[0144] In the formula, For the front wheel steering angle actuator control law; fault disturbance Set to extremely small Gaussian noise, health factor The parameters are set as follows:

[0145]

[0146] In the formula, This serves as a baseline value for the degree of failure. Setting it to 0.6 indicates that the effectiveness of the actuator has decreased to 60%. For the range of fluctuation, Setting it to 0.05 means that the health factor fluctuates randomly within a range of 0.05.

[0147] Specifically, the front wheel steering actuator angle tolerance control module acquires lateral displacement signals from the sensors. and heading displacement signal Construct the comprehensive combined error of front wheel steering angle and the nonlinear calculus convergence function of the multi-weighted comprehensive error of the front wheel steering angle. An RBF neural network is used to approximate the equivalent disturbance caused by the fault online. Combined with the actual estimated front wheel steering angle and real-time fault variable estimates The desired front wheel steering angle control value output To perform corner compensation control, the following steps are taken:

[0148] S201. The running road trajectory is set as a circular trajectory, and the trajectory equation is expressed as follows:

[0149]

[0150] Acquire lateral displacement signal from sensor and heading displacement signal lateral error Calculate heading error , and The desired lateral trajectory and the desired heading trajectory are respectively determined by selecting the combined weighting coefficient of the front wheel steering angle error. and Calculate the combined error of the front wheel steering angle. .

[0151] The models for the lateral and heading errors of the four-wheel coordinated motion unit in this embodiment can be expressed as follows:

[0152]

[0153] In the formula, The desired angular velocity of the heading; For trajectory curvature; Let be the radius of curvature of the trajectory.

[0154] From the above formula, we can obtain that

[0155]

[0156] In the formula,

[0157]

[0158] In the formula, , These are the linear lateral stiffnesses of the front and rear wheels, respectively.

[0159] S202, Based on the comprehensive combined error of the front wheel steering angle Select the front wheel steering angle convergence combination weight coefficient , and The nonlinear calculus convergence function for calculating the multi-weighted comprehensive error of the front wheel steering angle is specifically expressed as:

[0160]

[0161] S203, Construct the convergence speed of the comprehensive combined error including the front wheel steering angle. and the equivalent disturbance amplitude caused by failure Front wheel steering angle actuator control law .

[0162]

[0163] S204. Adaptively optimize the control law of the front wheel steering angle actuator in step S203 by introducing a radial basis function neural network and using the nonlinear calculus convergence function of the multi-weight comprehensive error of the front wheel steering angle. As network input, the convergence speed of the combined error of the front wheel steering angle. and As the network output, the nonlinear disturbance caused by a local fault in the steering device is approximated online using the Gaussian radial basis function, and the parameters are adjusted in real time. and This is to improve the compensation capability for missing steering angles and the calculation accuracy of the desired front wheel steering angle. and The network output can be represented as:

[0164]

[0165] In the formula, These are the weights of the radial basis function neural network.

[0166] The Gaussian function for a radial basis function neural network is expressed as follows:

[0167]

[0168] In the formula, As the center vector, This is the width parameter.

[0169] The adaptive update law of radial basis function neural networks is as follows:

[0170]

[0171] In the formula, For inertia ratio, .

[0172] The weight update equation for the above formula is:

[0173]

[0174] In the formula, , , The initial learning rates are set to 0.05, 0.01, and 0.01, respectively.

[0175] Will , , The initial value is set as follows: , , .

[0176] S205. Based on the nonlinear calculus convergence function of the multi-weighted comprehensive error of the front wheel steering angle in step S203. The adaptive update criterion for network parameters is constructed based on the trend of their changes, and the weights of the radial basis function neural network are adjusted accordingly. , center vector and width parameter Perform online updates, defining the adaptability metric for the weights as follows:

[0177]

[0178] Used to evaluate the current The convergence trend. If This indicates that the system state is converging toward the convergence target surface; if This indicates poor convergence, or even deviation from the trend. Therefore, we use... To optimize the objective, network parameters are adjusted. , , Perform an online update. Once the network parameters are updated, and This also changes accordingly, and then feeds back into the preceding control law. This forms a closed-loop adaptive adjustment.

[0179] Specifically, the four-wheel yaw moment coordination and distribution additional control module receives the real-time fault variable estimate F and the desired front wheel steering angle control value. A multi-weighted proportional-integral combined state error convergence function is constructed, and based on this function, the additional yaw moment control quantity of the fused second-order continuous variable structure is calculated, outputting the additional yaw moment. Then, based on the quadratic programming algorithm, Distributed to four wheels, enabling real-time dynamic fault-tolerant torque for all four wheels. Calculation. The specific steps are as follows:

[0180] S301. Obtain the centroid sideslip angle via GPS / GNSS system. And obtain the yaw angle through the gyroscope in the IMU. The state vector that makes up the four-wheel coordinated motion unit is specifically represented as follows:

[0181]

[0182] S302. Construct the comprehensive state error of the adaptive weighted four-wheel coordinated motion unit. This characterizes the tracking effect of the four-wheel coordinated motion unit on the yaw moment, specifically expressed as:

[0183]

[0184] In the formula, For reference only. , These are the weight parameters.

[0185] S303. Calculate the convergence function of the multi-weighted proportional-integral combined state error:

[0186]

[0187] In the formula, These are the weight parameters.

[0188] S304. Construct an additional yaw moment control law that integrates the second-order continuous variable structure, as the input target for subsequent four-wheel torque distribution.

[0189] The control law consists of equivalent control terms With correction item composition:

[0190]

[0191] In the formula, , To control the gain, It is an auxiliary variable.

[0192] The additional yaw moment control law for the integrated second-order continuous variable structure is specifically expressed as follows:

[0193]

[0194] Based on the state vector obtained in step S301, the comprehensive state error constructed in step S302, and the convergence function of the multi-weight proportional-integral combined state error obtained in step S303, the additional yaw moment control quantity of the fused second-order continuous variable structure is calculated, and the additional yaw moment is output. It is used to compensate for the loss of steering ability caused by a partial failure of the front wheel steering actuator.

[0195] S305. The longitudinal tire force of the wheel is calculated using the wheel-end drive torque, wheel angular velocity, wheel angular acceleration, and wheel radius. The lateral tire force is calculated using an IMU, steering angle sensor, and wheel speed sensor. The vertical force of the wheel is obtained through a wheel load sensor. .in =1, 2, 3, 4.

[0196] S306. Construct an objective function based on a quadratic programming algorithm for distributing torque across four wheels. Divided into two parts, for and The weighted sum. The objective function for optimizing the overall adhesion utilization rate of the four wheel tires is defined in the first part. The function expression is as follows:

[0197]

[0198] In the formula, It is the tire-road adhesion coefficient.

[0199] S307, Define the second part of step S406: the objective function for tracking the desired yaw moment and total longitudinal tire force. The function expression is as follows:

[0200]

[0201] In the formula, , , , For the expected , .

[0202] S308. Based on steps S306 and S307, a total objective function based on a quadratic programming algorithm is constructed for distributing the torque of the four wheels, using weighting coefficients. The optimization objectives of overall tire adhesion utilization, desired yaw moment, and total longitudinal tire force tracking are weighted and coordinated. Under the constraint of output torque on each wheel, the optimal output torque of the four wheels is solved to achieve additional yaw moment compensation and coordinated torque distribution among the four wheels under local steering system failure conditions in the four-wheel coordinated motion unit. This is expressed as:

[0203]

[0204] In the formula, These are the weighting coefficients.

[0205] Specifically, the trajectory tracking deep reinforcement learning parameter optimization module introduces the GRU framework, which has a strong memory for historical state information, and combines the online decision-making and global search capabilities of deep reinforcement learning to optimize the weight parameters of the front wheel steering actuator angle fault-tolerant control module, the front wheel steering actuator local failure estimation module, and the four-wheel yaw moment coordination allocation additional control module. Dynamic optimization is performed, and the optimized parameters are sent to the front wheel steering actuator partial failure estimation module, the front wheel steering actuator angle fault-tolerant control module, and the four-wheel yaw torque coordination distribution additional control module to achieve overall optimal trajectory tracking control capability. Specific steps include:

[0206] S401, using real-time fault variable estimation values lateral error and heading error Sideslip angle of center of mass A state space for a deep deterministic policy gradient fusion GRU is constructed, enabling the parameter optimization module to simultaneously characterize the comprehensive operating state of the four-wheel coordinated motion unit under trajectory tracking error, fault state, and attitude change, specifically represented as follows:

[0207]

[0208] S402. Construct the action space of the deep deterministic policy gradient fused with GRU. This action space includes weight coefficients from the front wheel steering actuator angle fault-tolerant control module, the front wheel steering actuator local failure estimation module, and the four-wheel yaw moment coordination distribution supplementary control module. This allows the parameter optimization module to directly adjust key parameters in front wheel steering control, fault estimation, and torque coordination distribution, thereby achieving collaborative optimization among multiple control modules. Because... , The action space is specifically represented as follows:

[0209]

[0210] S403. To effectively train the parameter optimization module, a reward function for the gradient of the deep deterministic policy fused with GRU is defined. The reward function consists of three parts. The first part is used to suppress excessive lateral bias. The reward function is expressed as follows:

[0211]

[0212] S404. To guide the parameter optimization module to reduce the heading angle deviation of the four-wheel coordinated motion unit during trajectory tracking, and to improve the heading following capability and steering consistency under local fault conditions of the steering device, the second part of step S503, the reward function, is defined as follows:

[0213]

[0214] S405. The method for determining the stability region of the centroid side slip phase diagram is as follows:

[0215]

[0216] In the formula, and These are the boundary coefficients and the velocity, respectively. and road surface friction coefficient related.

[0217] To enhance the stability of the four-wheel coordinated motion unit and avoid increased sideslip, abnormal yaw response, or uneven tire adhesion utilization caused by excessive pursuit of trajectory tracking accuracy, the reward function in the third part of step S503 is defined as follows:

[0218]

[0219] S406. A weighted combination of lateral deviation reward, heading deviation reward, and stability reward is constructed to create a total reward function, used to comprehensively evaluate the impact of the current parameter combination on trajectory tracking accuracy, heading maintenance capability, and overall vehicle stability. Specifically, this is expressed as:

[0220]

[0221] S407. Construct the action-value function in the deep deterministic policy gradient algorithm. By representing the cumulative benefit that can be obtained by taking corresponding parameter adjustment actions in the current state, evaluate the long-term impact of different parameter combinations on the trajectory tracking control effect, and provide a valuable basis for policy learning in the parameter optimization module. Specifically, it can be expressed as:

[0222]

[0223] In the formula, This is the state at this moment. For this moment's action, Reward this moment. For the state at the next moment, for and The expected distribution.

[0224] S408. Construct a target value function based on the reward function and the target network output, aiming to minimize the deviation between the current action value estimate and the target value. Train the Critic network to accurately evaluate the impact of parameter adjustment actions on the overall control performance of the system. Its expression is:

[0225]

[0226] In the formula, As a discount factor, .

[0227] S409. Construct the Critic network loss function to improve the accuracy of value assessment by minimizing the action value estimation error. This enables the parameter optimization module to stably learn the parameter coupling relationship between front wheel steering control, fault estimation, and torque coordination distribution. Its expression is:

[0228]

[0229] S410. Update the Actor network based on the policy gradient output by the Critic network, enabling the Actor network to output better continuous control actions. This achieves dynamic optimization of the weight parameters of multiple control modules and improves the trajectory tracking capability of the four-wheel coordinated motion unit under complex working conditions and local fault conditions of the steering device. Its expression is:

[0230]

[0231] S411. A soft update mechanism is used to update the parameters of the target Actor network and the target Critic network, allowing the target network to slowly track changes in the online network, reducing training oscillations, and improving the training stability and convergence reliability of the parameter optimization module. The calculation process is as follows:

[0232]

[0233] In the formula, This is a smoothing factor.

[0234] Specifically, the GRU structure refinement process consists of steps S501-S506:

[0235] S501 Figure 2 As shown, the current observation state Hidden state from the previous moment A single-layer gated recurrent unit (GRU) in the common input Actor network extracts historical state information through reset and update gates. The reset gate... and the update gate They are represented as follows:

[0236]

[0237]

[0238] In the formula, It is the Sigmoid activation function. This represents the concatenated vector of the hidden state from the previous time step and the observed state at the current time step.

[0239] S502 Reset the door according to step 5601 Filter historical state information and calculate candidate hidden states. Its expression is:

[0240]

[0241] in, The hyperbolic tangent activation function is used. This represents the Hadamard product.

[0242] S503 Update the door according to step S501 Hidden state from the previous moment and candidate hidden state Perform fusion to obtain the hidden state at the current moment. Its expression is:

[0243]

[0244] S504 Hide the current time state described in step S503 With current observation status Feature splicing is performed and fed into subsequent fully connected layers of the Actor network, outputting continuous control actions. Its expression is:

[0245]

[0246] in, This represents the policy function of the Actor network.

[0247] S505 The continuous control action described in step S504 is performed. Mapped to a multi-control module weight parameter vector:

[0248]

[0249] The signals are sent to the front wheel steering actuator partial failure estimation module, the front wheel steering actuator angle fault-tolerant control module, and the four-wheel yaw torque coordination and distribution additional control module, respectively, to realize the estimation of the actual front wheel steering angle, the front wheel steering loss compensation control, the solution of the additional yaw torque, and the key parameters in the process of dynamic fault-tolerant torque distribution of the four wheels. Joint dynamic optimization.

[0250] S506 Through the above steps, the GRU is used to extract the historical state change characteristics of the four-wheel coordinated motion unit during the continuous control process, and the continuous control action a is used... t The adjustment amount of the weight parameters of the multiple control modules is used to enable the parameter optimization module to combine the current fault state, trajectory error and historical operating state to collaboratively optimize the weight parameters of the front wheel steering actuator local failure fault estimation module, the front wheel steering actuator angle fault tolerance control module and the four-wheel yaw torque coordination distribution additional control module, thereby improving the overall trajectory tracking control capability under the condition of front wheel steering actuator local failure.

[0251] refer to Figure 4 The trajectory tracking deep reinforcement learning parameter optimization module described in this embodiment of the invention is compared with the ordinary deep reinforcement learning parameter optimization module. The horizontal axis of the curve represents the number of training rounds, and the vertical axis represents the reward value. It can be observed that in the early stage of training, the network explores a large range and takes highly random actions, resulting in large fluctuations in the reward curve, all of which are negative. The agent is still in the exploration stage and it is difficult to obtain high immediate rewards. When the number of training rounds reaches about 400, the reward value gradually stabilizes and shows a good convergence trend. Around round 25, the average reward of ordinary deep reinforcement learning is about -3595.71, while that of GRU-DDPG is about -3002.19, indicating that GRU-DDPG can obtain a higher initial reward in the exploration stage, averaging about 593.52 higher. Before 400 rounds, the minimum reward value of ordinary deep reinforcement learning reached -5235.50, the maximum was 54.47, and the standard deviation was 1240.31, showing significant fluctuations. In contrast, the minimum reward value of GRU-DDPG was -3551.30, the maximum was 90.29, and the standard deviation was 991.45, demonstrating better fluctuation suppression. After 400 training rounds, the reward curve of GRU-DDPG rose rapidly and stabilized, with the average reward increasing to -14.20 and the standard deviation being 39.49, showing a good convergence trend. Ordinary deep reinforcement learning converged more slowly, and the average reward for the last 900 rounds was only -223.93 with a standard deviation of 127.70. Throughout the training process, the average reward of GRU-DDPG was -544.57, better than DDPG's -939.18. It can be seen that the trajectory tracking deep reinforcement learning parameter optimization module described in this embodiment of the invention can achieve higher and more stable rewards.

[0252] Example 2: This example proposes a highly reliable four-wheel coordinated motion unit trajectory tracking control system based on deep reinforcement learning. It consists of four parts: a front wheel steering actuator partial failure estimation module, a front wheel steering actuator angle fault-tolerant control module, a four-wheel yaw torque coordinated distribution additional control module, and a trajectory tracking deep reinforcement learning parameter optimization module. Specifically, it includes:

[0253] The front wheel steering actuator partial failure estimation module is used to acquire signals from the front wheel steering angle sensor. The estimated front wheel steering angle obtained by back-deriving from the state-space equations of the motion unit is integrated. and Adaptive multi-weight combination is adopted Weighted fusion is performed to calculate the estimated value of the actual front wheel steering angle. and real-time fault variable estimates and the estimated value of the actual front wheel steering angle The output is sent to the front wheel steering actuator angle fault-tolerant control module, which then outputs the estimated value of the real-time fault variable. The output is sent to the four-wheel yaw torque coordination and distribution additional control module;

[0254] The front wheel steering actuator angle tolerance control module is used to acquire lateral displacement signals from the sensors. and heading displacement signal Construct the comprehensive combined error of front wheel steering angle and the nonlinear calculus convergence function of the multi-weighted comprehensive error of the front wheel steering angle. An RBF neural network is used to approximate the equivalent disturbance caused by the fault online. Combined with the actual estimated front wheel steering angle and real-time fault variable estimates The desired front wheel steering angle control value output Perform missing corner compensation control;

[0255] An additional control module for coordinated distribution of yaw moment across four wheels is used to receive real-time estimates of fault variables. and desired front wheel steering angle control amount A multi-weighted proportional-integral combined state error convergence function is constructed, and based on this function, the additional yaw moment control quantity of the fused second-order continuous variable structure is calculated, outputting the additional yaw moment. Then, based on the quadratic programming algorithm, Distributed to four wheels, enabling real-time dynamic fault-tolerant torque for all four wheels. , , , calculate;

[0256] The trajectory tracking deep reinforcement learning parameter optimization module is used to dynamically optimize the weight parameters of each control module by introducing the GRU framework with strong memory of historical state information and combining the online decision-making and global search capabilities of deep reinforcement learning. The optimized parameters are then sent to the front wheel steering actuator local failure fault estimation module, the front wheel steering actuator angle fault-tolerant control module, and the four-wheel yaw torque coordination distribution additional control module to achieve the overall optimal trajectory tracking control capability.

[0257] Example 3: This example proposes a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method described in this invention.

[0258] Example 4: This example proposes an electronic device, including a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, it implements the steps of the method described in this invention.

[0259] Example 5: This example proposes a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in this invention.

[0260] It should be noted that the processing flow of embodiments 2-5 corresponds to the specific steps of the method provided in embodiment 1 of the present invention, and has the corresponding functional modules and beneficial effects of the method. Technical details not described in detail in this embodiment can be found in the method provided in embodiment 1 of the present invention.

[0261] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0262] The specific implementation schemes described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific implementation schemes of the present invention and are not intended to limit the scope of the present invention. Any equivalent changes and modifications made by those skilled in the art without departing from the concept and principles of the present invention should fall within the scope of protection of the present invention.

Claims

1. A highly reliable trajectory tracking control method for a four-wheel coordinated motion unit based on deep reinforcement learning, characterized in that, Includes the following steps: S1. Acquire front wheel steering angle sensor signal The estimated value of the front wheel steering angle is obtained by back-deriving the state-space equation of the motion unit. , And adopt adaptive multi-weight combination Perform weighted fusion to calculate the estimated value of the actual front wheel steering angle. and real-time fault variable estimates ; S2. Acquire lateral displacement signal and heading displacement signal Construct the comprehensive combined error of front wheel steering angle and the nonlinear calculus convergence function of the multi-weighted comprehensive error of the front wheel steering angle. An RBF neural network is used to approximate the equivalent disturbance caused by the fault online. Combined with the actual estimated front wheel steering angle and real-time fault variable estimates Output the desired front wheel steering angle control value Perform missing corner compensation control; S3, Receive real-time fault variable estimates and desired front wheel steering angle control amount A multi-weighted proportional-integral combined state error convergence function is constructed, and the additional yaw moment control quantity of the fused second-order continuous variable structure is calculated, outputting the additional yaw moment. Then, based on the quadratic programming algorithm, Distribute the torque to the four wheels and calculate the real-time dynamic fault-tolerant torque of each wheel. , , , ; S4. By introducing the GRU framework and combining it with deep reinforcement learning, the weight parameters involved in the estimation of the front wheel steering angle, the missing steering angle compensation control, the solution of the additional yaw torque, and the dynamic fault-tolerant torque distribution of the four wheels are jointly and dynamically optimized to achieve the overall optimal trajectory tracking control capability.

2. The method according to claim 1, characterized in that, Step S1 specifically includes the following steps: S101. Obtain the front wheel angle using the front wheel angle sensor installed on the front wheel of the motion unit. ; S102. Using the state-space equations of the motion unit, the two types of front wheel steering angles are derived by reverse deduction. , The state-space equation of the motion unit is expressed as: , in, , , , , , ; This refers to lateral error; For heading error; For the first The linear lateral stiffness of each wheel, =1, 2, 3, 4; For overall quality; For partial failure estimation of the front wheel steering actuator; , These are the distances from the center of mass to the front and rear axles, respectively. To bypass Moment of inertia of the shaft; , The longitudinal velocity and lateral velocity of the center of mass are respectively determined; The desired heading angular velocity; Two types of front wheel steering angles can be obtained by directly inversely solving the state-space equations of the motion unit. , The estimated value is calculated using the following formula: , , S103. Select weighting coefficients and Calculate the partial failure estimation of the front wheel steering actuator The calculation formula is: , In the formula, >0, >0 and + =1; Real-time fault variable estimates Represented as: , In the formula, This is the control law for the front wheel steering angle actuator.

3. The method according to claim 1, characterized in that, Step S2 specifically includes the following steps: S201. Acquire the lateral displacement signal from the sensor. and heading displacement signal Calculate the lateral error and heading error ,in, and These are the desired lateral trajectory and the desired heading trajectory; the combined weighting coefficients for the front wheel steering angle error are selected. and Calculate the combined error of the front wheel steering angle. ; S202, Based on the combined error of the front wheel steering angle Select the front wheel steering angle convergence combination weight coefficient , and The nonlinear calculus convergence function for calculating the multi-weighted comprehensive error of the front wheel steering angle is specifically expressed as: ; S203, Construct the convergence speed of the comprehensive combined error including the front wheel steering angle. Equivalent disturbance amplitude caused by failure Front wheel steering angle actuator control law : , In the formula, , In the formula, , The linear lateral stiffness of the front and rear wheels are respectively. For trajectory curvature; S204. Adaptive optimization of the control law for the front wheel steering angle actuator is performed by introducing a radial basis function neural network, based on the nonlinear calculus convergence function of the multi-weighted comprehensive error of the front wheel steering angle. As network input, the convergence speed of the combined error of the front wheel steering angle. and As the network output, the nonlinear disturbance caused by a local fault in the steering device is approximated online using the Gaussian radial basis function, and the parameters are adjusted in real time. and ; and The network output is represented as: , In the formula, These are the weights of the radial basis function neural network; The Gaussian function for a radial basis function neural network is expressed as follows: , In the formula, As the center vector, For width parameters; S205, Based on the nonlinear calculus convergence function of the multi-weighted comprehensive error of the front wheel steering angle. The adaptive update criterion for network parameters is constructed based on the trend of their changes, and the weights of the radial basis function neural network are adjusted accordingly. , center vector and width parameter Perform online updates, defining the adaptability metric for the weights as follows: , Representing the present The convergence trend, if This indicates that the system state is converging toward the convergence target surface; if This indicates poor convergence; To optimize the objective, network parameters are adjusted. , , Perform online updates, and thus affect and Changes are then fed back to the control law. In this process, a closed-loop adaptive adjustment is formed.

4. The method according to claim 1, characterized in that, Step S3 specifically includes the following calculation steps: S301. Obtain the centroid sideslip angle via GPS / GNSS system. And obtain the yaw angle through the gyroscope in the IMU. The state vector that makes up the four-wheel coordinated motion unit is specifically represented as follows: , S302. Construct the comprehensive state error of the adaptive weighted four-wheel coordinated motion unit. This characterizes the tracking effect of the four-wheel coordinated motion unit on the yaw moment, specifically expressed as: , In the formula, For reference only. , These are weight parameters; S303. Calculate the convergence function of the multi-weighted proportional-integral combined state error: , In the formula, These are weight parameters; S304. Construct an additional yaw moment control law that integrates the second-order continuous variable structure, as the input target for subsequent four-wheel torque distribution, specifically expressed as: , Based on the state vector obtained in step S301, the comprehensive state error constructed in step S302, and the convergence function of the multi-weight proportional-integral combined state error obtained in step S303, the additional yaw moment control quantity of the fused second-order continuous variable structure is calculated, and the additional yaw moment ΔM is output. Z ; S305. The longitudinal tire force of the wheel is calculated using the wheel-end drive torque, wheel angular velocity, wheel angular acceleration, and wheel radius. The lateral tire force is calculated using an IMU, steering angle sensor, and wheel speed sensor. The vertical force of the wheel is obtained through a wheel load sensor. ,in =1, 2, 3, 4; S306. Construct an objective function based on a quadratic programming algorithm for distributing torque across four wheels. Divided into two parts, for and The weighted sum; defining the objective function for optimizing the overall adhesion utilization rate of the four wheel tires in the first part. The function expression is as follows: ; In the formula, It is the tire-road adhesion coefficient; S307, Define the second part of step S306: the objective function for tracking the desired yaw moment and total longitudinal tire force. The function expression is as follows: ; In the formula, , , , For the expected , ; S308. Based on steps S306 and S307, a total objective function based on a quadratic programming algorithm is constructed for distributing the torque of the four wheels, using weighting coefficients. The optimization objectives of overall tire adhesion utilization rate, desired yaw moment, and total longitudinal tire force tracking are weighted and coordinated. Under the constraint of output torque of each wheel, the optimal output torque of the four wheels is solved to achieve additional yaw moment compensation and coordinated torque distribution of the four wheels under local steering system failure conditions in the four-wheel coordinated motion unit. Specifically, this is expressed as follows: ; In the formula, These are the weighting coefficients.

5. The method according to claim 1, characterized in that, In step S4, the GRU framework is introduced, and combined with deep reinforcement learning, the weight parameters are jointly and dynamically optimized. This includes the following steps: S401, using real-time fault variable estimates lateral error and heading error Sideslip angle of center of mass A state space for a deep deterministic policy gradient fusion GRU is constructed to characterize the comprehensive operational state under trajectory tracking error, fault state, and attitude change, specifically represented as follows: ; S402. Construct a deep deterministic strategy gradient action space fused with GRU, wherein the action space includes the weighting coefficients of the integrated combined error, the weighting coefficients of the weighted fusion, and the weighting coefficients of the overall objective function, to jointly adjust key parameters in front wheel steering control, fault estimation, and torque coordination distribution, thereby achieving collaborative optimization of multiple control links; the action space is specifically represented as follows: ; S403. Define the reward function for the deep deterministic policy gradient of the fused GRU. The reward function consists of three parts. The first part is used to suppress excessive lateral bias. The reward function is expressed as follows: ; S404. Define the second part of the reward function, specifically as follows: ; S405. Define the third part, the reward function, as follows: ; In the formula, and These are the boundary coefficients and the velocity, respectively. and road surface friction coefficient related; S406. The lateral deviation reward, heading deviation reward, and stability reward are weighted and combined to construct the total reward function, which is specifically expressed as follows: ; S407. Construct the action-value function in the deep deterministic policy gradient algorithm. By representing the cumulative benefit that can be obtained by taking corresponding parameter adjustment actions in the current state, evaluate the long-term impact of different parameter combinations on the trajectory tracking control effect, specifically expressed as follows: ; In the formula, This is the state at this moment. For this moment's action, Reward this moment. For the state at the next moment, for and The expected distribution; S408. Construct a target value function based on the reward function and the target network output, aiming to minimize the deviation between the current action value estimate and the target value. Train the Critic network with the following expression: ; In the formula, As a discount factor, ; S409. Construct the Critic network loss function to improve the accuracy of value assessment by minimizing the action value estimation error. Its expression is: ; S410. Update the Actor network based on the policy gradient output by the Critic network, as shown in the expression: ; S411. A soft update mechanism is used to update the parameters of the target Actor network and the target Critic network, allowing the target network to slowly track changes in the online network and reduce training oscillations. The calculation process is as follows: , In the formula, This is a smoothing factor.

6. The method according to claim 5, characterized in that, Step S4 introduces the GRU framework and combines it with deep reinforcement learning to perform joint dynamic optimization of the weight parameters. It also includes the following steps: Current observation status Hidden state from the previous moment A single-layer gated recurrent unit (GRU) in the common input Actor network extracts historical state information through reset and update gates. The reset gate... and the update gate They are represented as follows: , , In the formula, It is the Sigmoid activation function. This represents the concatenated vector of the hidden state from the previous time step and the observed state at the current time step. According to the reset door Filter historical state information and calculate candidate hidden states. Its expression is: , in, The hyperbolic tangent activation function is used. Represents the Hadamard product; According to the update gate Hidden state from the previous moment and candidate hidden state Perform fusion to obtain the hidden state at the current moment. Its expression is: , Hide the current state With current observation status Feature splicing is performed and fed into subsequent fully connected layers of the Actor network, outputting continuous control actions. Its expression is: , In the formula, Represents the policy function of the Actor network; The continuous control action Mapped to a multi-control module weight parameter vector: , The weights are used for adaptive multi-weight combination in front wheel steering angle estimation, error combination weight in missing steering angle compensation control, and objective function weight in additional yaw moment solution and four-wheel dynamic fault-tolerant torque distribution, respectively, to achieve joint dynamic optimization of the weight parameters of each link.

7. A highly reliable four-wheel coordinated motion unit trajectory tracking control system based on deep reinforcement learning, characterized in that, It consists of four parts: a front wheel steering actuator partial failure estimation module, a front wheel steering actuator angle fault-tolerant control module, a four-wheel yaw torque coordination and distribution supplementary control module, and a trajectory tracking deep reinforcement learning parameter optimization module. Specifically, it includes: The front wheel steering actuator partial failure estimation module is used to acquire signals from the front wheel steering angle sensor. The estimated front wheel steering angle obtained by back-deriving from the state-space equations of the motion unit is integrated. and Adaptive multi-weight combination is adopted Weighted fusion is performed to calculate the estimated value of the actual front wheel steering angle. and real-time fault variable estimates and the estimated value of the actual front wheel steering angle The output is sent to the front wheel steering actuator angle fault-tolerant control module, which then outputs the estimated value of the real-time fault variable. The output is sent to the four-wheel yaw torque coordination and distribution additional control module; The front wheel steering actuator angle tolerance control module is used to acquire lateral displacement signals from the sensors. and heading displacement signal Construct the comprehensive combined error of front wheel steering angle and the nonlinear calculus convergence function of the multi-weighted comprehensive error of the front wheel steering angle. An RBF neural network is used to approximate the equivalent disturbance caused by the fault online. Combined with the actual estimated front wheel steering angle and real-time fault variable estimates The desired front wheel steering angle control value output Perform missing corner compensation control; An additional control module for coordinated distribution of yaw moment across four wheels is used to receive real-time estimates of fault variables. and desired front wheel steering angle control amount A multi-weighted proportional-integral combined state error convergence function is constructed, and based on this function, the additional yaw moment control quantity of the fused second-order continuous variable structure is calculated, outputting the additional yaw moment. Then, based on the quadratic programming algorithm, Distributed to four wheels, enabling real-time dynamic fault-tolerant torque for all four wheels. , , , calculate; The trajectory tracking deep reinforcement learning parameter optimization module is used to dynamically optimize the weight parameters of each control module by introducing the GRU framework with strong memory of historical state information and combining the online decision-making and global search capabilities of deep reinforcement learning. The optimized parameters are then sent to the front wheel steering actuator local failure fault estimation module, the front wheel steering actuator angle fault-tolerant control module, and the four-wheel yaw torque coordination distribution additional control module to achieve the overall optimal trajectory tracking control capability.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

9. An electronic device comprising a processor and a memory, wherein the memory stores a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.