A fuzzy tracking control method based on policy iteration

By employing a strategy-iterative fuzzy tracking control method, utilizing the TS fuzzy model and fuzzy discount coupled algebraic Riccati equations, the dependence of the steer-by-wire system on the precise model is solved, achieving online learning and highly robust trajectory tracking control.

CN122194706BActive Publication Date: 2026-07-21ANHUI UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ANHUI UNIV
Filing Date
2026-05-15
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing steer-by-wire systems rely heavily on complete dynamic information and initial stable control strategies, making it difficult to effectively cope with severe nonlinear characteristics such as friction and backlash, as well as external disturbances. This limits their application, especially in complex environments.

Method used

A fuzzy tracking control method based on strategy iteration is adopted. By establishing a TS fuzzy model, an augmented system is constructed, and the fuzzy discounted coupled algebraic Riccati equation is used in combination with the Hamilton-Bellman equation to iteratively solve the control gain and disturbance gain matrices, thereby realizing the online update of the optimal control strategy.

Benefits of technology

It eliminates the reliance on precise mathematical models, possesses powerful online learning capabilities, and can adaptively update the optimal control strategy, thereby improving the trajectory tracking accuracy and robustness of the steer-by-wire system in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122194706B_ABST
    Figure CN122194706B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of vehicle control system, and particularly relates to a fuzzy tracking control method based on policy iteration, aiming at the nonlinear and disturbance problems existing in the steer-by-wire system, first, the T-S fuzzy model of the system is established and the reference trajectory is introduced to build the augmented system; secondly, the fuzzy discount coupled algebraic Riccati equation containing the value matrix and the disturbance suppression level is constructed based on the Hamilton-Bellman equation; then, the online running data is obtained to iteratively solve the policy evaluation equation, and the control gain matrix and the interference gain matrix corresponding to each fuzzy rule are updated synchronously until convergence, so as to obtain the optimal control strategy and the interference strategy based on the optimal value matrix; finally, the optimal control strategy is applied to the steer-by-wire system. The present application does not need the accurate mathematical model of the system, and realizes the high-precision and strong-robust trajectory tracking control of the steer-by-wire system through complete data driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle control system technology, and in particular to a fuzzy tracking control method based on strategy iteration. Background Technology

[0002] As a key technology for intelligent vehicles and autonomous driving, steer-by-wire systems offer greater freedom in vehicle design by eliminating traditional mechanical connections and using electrical signals to transmit driving commands. This innovative design significantly improves driving safety and comfort while reducing energy consumption. However, the elimination of mechanical connections and the increased reliance on sensors and actuators introduce significant nonlinearities and uncertainties into the system. Actuator saturation, backlash, and friction constitute important sources of nonlinearity, posing a significant challenge to robust and accurate controller design.

[0003] Although various nonlinear control techniques have been proposed to address such problems, existing tracking control methods typically rely on two main components: a feedback term derived from the solution of the coupled algebraic Riccati equation, and a feedforward term obtained by solving an auxiliary differential equation.

[0004] However, these methods often require complete dynamic information of the system (accurate mathematical and physical models) or an initial stable control strategy. Furthermore, such information is difficult to obtain in practical steer-by-wire systems, greatly limiting their practical applicability, especially in uncertain or complex driving environments. To address these issues, we propose a fuzzy tracking control method based on strategy iteration. Summary of the Invention

[0005] In view of this, the purpose of this invention is to propose a fuzzy tracking control method based on strategy iteration, in order to solve the technical problem that the existing trajectory tracking control method of the steer-by-wire system relies heavily on the complete dynamic information of the system (i.e., accurate mathematical model) and the initial stable control strategy, which makes it difficult to be effectively applied in real complex environments when faced with severe nonlinear characteristics such as friction and backlash and external disturbances.

[0006] To achieve the above objectives, this invention provides a fuzzy tracking control method based on strategy iteration, comprising the following steps:

[0007] S1. Establish a TS fuzzy model of the steer-by-wire system, introduce a reference trajectory, and construct an augmented system containing the reference trajectory based on the TS fuzzy model. When the augmented system applies a control strategy, it is recorded as a controlled augmented system.

[0008] S2. Based on the Hamilton-Bellman equations, construct the fuzzy discounted coupled algebraic Riccati equations corresponding to the augmented system. These equations contain a positive definite symmetric matrix used to evaluate the performance of the controlled augmented system. and preset disturbance suppression level ;

[0009] The fuzzy discount coupling algebraic Riccati equation introduces a discount factor greater than or equal to zero. And use it to construct an exponentially decaying weight term. To balance the weight distribution of the controlled augmented system performance on the time axis, and through the control gain matrix With interference gain matrix The coupling achieves a zero-sum game between minimizing control costs and maximizing the impact of disturbances;

[0010] S3. Acquire real-time operating data of the online-collected steer-by-wire system, perform an integral transformation along the augmented system's operating trajectory within a finite time window, and use the real-time operating data to perform an equivalent transformation on the fuzzy discounted coupled algebraic Riccati equation, constructing and iteratively solving the strategy evaluation equation to update the positive definite symmetric value matrix. In each iteration, based on the updated positive definite symmetric value matrix... and perturbation suppression level Calculate and synchronously update the control gain matrix corresponding to each fuzzy rule. and interference gain matrix Repeat the iteration until the positive definite symmetric value matrix of two consecutive iterations is obtained. The norm of the difference is less than the preset minimum value The convergence condition is met, and the optimal control strategy and the optimal disturbance strategy are obtained based on the converged optimal positive definite symmetric value matrix.

[0011] S4. Apply the optimal control strategy to the steer-by-wire system to achieve trajectory tracking control.

[0012] Preferably, in step S1, the TS fuzzy model of the steer-by-wire system is established, which specifically includes the following steps:

[0013] S11. The dynamic equations of the steer-by-wire system are established as follows:

[0014]

[0015] in, For the steering wheel angle, For the angular velocity of the steering wheel, For the steering wheel angular acceleration, For rotational inertia, The viscous damping coefficient is... For nonlinear uncertainties, To control the input, Lumped disturbance terms are treated as disturbance input terms in control law design. The transmission ratio;

[0016] S12, Define the state vector as follows ,in , Define control input The dynamic equations are expressed as a state-space model:

[0017] ;

[0018] The output equation is ;

[0019] in, , , , For state vectors, Let be the derivative of the state vector with respect to time. For the output vector, The local state matrix, For local input matrices, For local perturbation matrix, For local output matrix, superscript Represents the transpose of a vector;

[0020] S13. Using the TS fuzzy modeling method, local linear approximation is performed on the nonlinear uncertainty terms to obtain the overall fuzzy model:

[0021] ;

[0022] ;

[0023] in, To represent the total number of fuzzy rules, This represents a fuzzy rule index. The normalized weight function satisfies , As a given variable vector, , , , The first The local state matrix, local input matrix, local perturbation matrix, and local output matrix corresponding to each fuzzy rule;

[0024] S14, Define the reference trajectory as ;

[0025] in, The reference trajectory state vector, Let be the derivative of the reference trajectory state vector with respect to time. Output vector as reference trajectory. and It is a constant matrix;

[0026] S15. Introducing an extended state vector and extended output vector This yields an augmented system:

[0027]

[0028]

[0029] in, , , , , and These are the transposes of the state vector and the reference trajectory state vector, respectively. and These are the transposes of the output vector and the reference trajectory output vector, respectively. , , , The augmentation system is in the first The local augmented state matrix, local augmented input matrix, local augmented output matrix, and local augmented perturbation matrix under the fuzzy rules.

[0030] Preferably, in step S2, the fuzzy discount coupling algebraic Riccati equation is:

[0031]

[0032] in, The first one to be solved The positive definite symmetric value matrix corresponding to the fuzzy rules. A discount factor greater than or equal to zero This is the preset disturbance suppression level;

[0033] and , , The state penalty weighting matrix is... This is the system trajectory tracking error output matrix. and It is a positive definite weighted matrix. , , , These are the transposes of the corresponding matrices.

[0034] Preferably, step S3 includes the following steps:

[0035] S31. Based on the system operation data collected online, utilize the system operation trajectory within a limited time window... Integral transformations within the range are used to construct state difference data matrices. State-quadratic integral data matrix State-input cross-integral data matrix and state-perturbation cross-integral data matrix :

[0036] ;

[0037] ;

[0038] ;

[0039] ;

[0040] in, , , ..., For continuous sampling times during system operation, The preset sampling time interval, For Kronecker product, This is the integration time variable, used to perform integration calculations on system operating data within a finite time window, and is related to time. Belonging to the same type of time dimension, To expand the state vector's transformation value;

[0041] ;

[0042] It is an exponentially decaying term. For the first The transpose of the control input under the given fuzzy rules. No. The transpose of the corresponding interference input under the fuzzy rule. For integration variables;

[0043] S32. Initialization parameters: Set the preset minimum value For the control gain matrix and interference gain matrix Select it in the first The initial control gain matrix corresponding to the fuzzy rule and interference gain matrix Set the number of iterations =0;

[0044] S33, in the In the next iteration, based on the control gain matrix of the current iteration and interference gain matrix Solve the following data-driven equation using the integral data matrix to obtain the positive definite symmetric value matrix of the current iteration. :

[0045] ;

[0046] in, ;

[0047] ;

[0048] ;

[0049] in, The regression matrix is ​​composed of data collected online. The identity matrix corresponding to the state vector. Denotes the matrix vectorization operator, where ;

[0050] S34. Construct the performance function of the controlled augmented system. Its integral expression is:

[0051]

[0052] Construct the corresponding Hamiltonian function based on the performance function. :

[0053]

[0054] Based on the theory of optimality, and applying the extreme value condition Using the obtained positive definite symmetric value matrix ,renew The control gain matrix of the next iteration and interference gain matrix :

[0055] ;

[0056] in, For the system's performance function, Let Hamiltonian be the system function. The performance function with respect to the extended state vector The partial derivatives, Let be the transformation matrix of the partial derivative functions. and These are the partial derivatives of the Hamiltonian function with respect to the control input and interference input under the current fuzzy rules, respectively.

[0057] S35. Repeat steps S33 to S34 iteratively until the convergence condition is met. Obtain the optimal positive definite symmetric value matrix .

[0058] Preferably, based on the optimal positive definite symmetric value matrix The optimal control strategy and optimal interference strategy The expressions are as follows:

[0059] ;

[0060] .

[0061] Preferably, under the control of the optimal control strategy, the steer-by-wire system satisfies the following performance constraints:

[0062] .

[0063] in, For exponentially decaying weight terms, To expand the state vector The transpose of the matrix, The state penalty weighting matrix is... For optimal control strategy The transpose of the matrix, Optimal interference strategy The transpose of the matrix, For time derivative.

[0064] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described above.

[0065] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0066] A steer-by-wire system includes a steering motor, a feedback motor, an electronic control unit, and a front wheel actuator. The electronic control unit is configured to execute a fuzzy tracking control method based on strategy iteration, wherein:

[0067] The steering motor is used to execute control commands and drive the front wheels to steer;

[0068] The feedback motor is used to collect steering wheel angle and speed signals;

[0069] The electronic control unit is used to run the optimal control strategy;

[0070] The front wheel actuator is used to achieve steering.

[0071] The beneficial effects of this invention are as follows:

[0072] I. This invention utilizes TS fuzzy rules to effectively approximate and handle the complex nonlinear dynamics (such as friction and backlash) within the steer-by-wire system. Subsequently, by employing a data-driven strategy iteration method, it directly uses data generated during online system operation to iteratively solve the algebraic Riccati equations. This mechanism eliminates the dependence on precise mathematical and physical models and avoids the stringent requirements of traditional methods on the initial stable control strategy, enabling the system to possess powerful online learning capabilities and adaptively update the optimal control strategy based on real-time operating data.

[0073] Second, this invention endows the closed-loop system with extremely high robustness by treating unmodeled dynamics and external disturbances as "worst-case interference strategies" for game training, enabling it to have excellent suppression capabilities against disturbances under complex road conditions. At the same time, through rigorous Lyapunov stability analysis, this method provides a strict theoretical guarantee for the global asymptotic stability of the closed-loop system and the monotonic convergence of strategy iteration, greatly improving the accuracy and reliability of trajectory tracking in the steer-by-wire system. Attached Figure Description

[0074] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0075] Figure 1 This is a flowchart of the steps of the present invention;

[0076] Figure 2 This is a flowchart of the strategy iteration process in step S3 of the present invention;

[0077] Figure 3 This is a diagram showing the convergence result of the norm error of the control gain matrix and interference gain matrix during the strategy iteration process of this invention.

[0078] Figure 4 This is a magnified view of the trajectory tracking results and tracking error of the system state output and reference state output of the present invention. Detailed Implementation

[0079] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0080] like Figures 1 to 4 As shown, a fuzzy tracking control method based on policy iteration includes the following steps:

[0081] S1. Establish a TS fuzzy model of the steer-by-wire system, introduce a reference trajectory, and construct an augmented system containing the reference trajectory based on the TS fuzzy model. When the augmented system applies a control strategy, it is recorded as a controlled augmented system.

[0082] This step aims to transform the nonlinear steering-by-wire system into a fuzzy model that is easy to process, and combine it with the target trajectory. Specifically, it includes the following steps:

[0083] S11. The dynamic equations of the steer-by-wire system are established as follows:

[0084]

[0085] in, For the steering wheel angle, For the angular velocity of the steering wheel, For the steering wheel angular acceleration, For rotational inertia, The viscous damping coefficient is... This refers to nonlinear uncertainties (such as the nonlinear components in friction and restoring torque). To control the input, Lumped disturbance terms are treated as disturbance input terms in control law design. The transmission ratio;

[0086] S12. Define the state vector and state-space model:

[0087] Define the state vector as ,in , Define control input The dynamic equations are expressed as a state-space model:

[0088] ;

[0089] The output equation is ;

[0090] in, , , , For state vectors, Let be the derivative of the state vector with respect to time. For the output vector, The local state matrix, For local input matrices, For local perturbation matrix, For local output matrix, superscript Represents the transpose of a vector;

[0091] S13. Using the TS fuzzy modeling method, nonlinear uncertainties are addressed. By performing local linear approximation, the overall fuzzy model is obtained:

[0092] ;

[0093] ;

[0094] in, To represent the total number of fuzzy rules, This represents a fuzzy rule index. The normalized weight function satisfies , As a given variable vector, , , , The first The local state matrix, local input matrix, local perturbation matrix, and local output matrix corresponding to each fuzzy rule;

[0095] S14. Introducing Reference Trajectories and Constructing Augmented Systems

[0096] Define the reference trajectory generator as ;

[0097] in, The reference trajectory state vector, Let be the derivative of the reference trajectory state vector with respect to time. Output vector as reference trajectory. and It is a constant matrix;

[0098] S15. Introducing an extended state vector and extended output vector This yields an augmented system:

[0099]

[0100]

[0101] in, and These are the transposes of the state vector and the reference trajectory state vector, respectively. and These are the transposes of the output vector and the reference trajectory output vector, respectively. , , , The augmentation system is in the first The local augmented state matrix, local augmented input matrix, local augmented output matrix, and local augmented perturbation matrix under the given fuzzy rules are defined as follows: , , , The system trajectory tracking error is defined as... ;

[0102] in, For the first The system trajectory tracking error output matrix under the given fuzzy rules, and satisfying the following conditions: .

[0103] S2. Based on the Hamilton-Bellman equations, construct the fuzzy discounted coupled algebraic Riccati equations corresponding to the augmented system. These equations contain a positive definite symmetric matrix used to evaluate the performance of the controlled augmented system. and preset disturbance suppression level The fuzzy discount coupling algebraic Riccati equation introduces a discount factor greater than or equal to zero. And use it to construct an exponentially decaying weight term. To balance the weight distribution of the controlled augmented system performance on the time axis, and through the control gain matrix With interference gain matrix The coupling achieves a zero-sum game between minimizing control costs and maximizing the impact of disturbances.

[0104] In step S2, by applying the Hamilton-Jacobi-Bellman (HJB) equation, the fuzzy discounted coupled algebraic Riccati equation describing the optimal state of the system is derived as follows:

[0105]

[0106] in, The first one to be solved The positive definite symmetric value matrix corresponding to each fuzzy rule is the kernel matrix of the performance cost function, used to evaluate the long-term cumulative performance cost of the controlled system.

[0107] A discount factor greater than or equal to zero means that an exponentially decaying weight term is introduced into the performance metric. (See step S4 for the specific form) to balance the performance evaluation weights of the system's current and future performance, and to ensure the strict convergence of the infinite time integral;

[0108] The "coupling" in this equation means that it contains both gain terms related to the control input and gain terms related to the lumped disturbance, reflecting a zero-sum game coupling relationship between minimizing control costs and maximizing disturbance effects.

[0109] This is the preset disturbance suppression level;

[0110] and , , The state penalty weighting matrix is... This is the system trajectory tracking error output matrix. and It is a positive definite weighted matrix. , , , These are the transposes of the corresponding matrices.

[0111] S3. Acquire real-time operating data of the online-collected steer-by-wire system, perform an integral transformation along the augmented system's operating trajectory within a finite time window, and use the real-time operating data to perform an equivalent transformation on the fuzzy discounted coupled algebraic Riccati equation, constructing and iteratively solving the strategy evaluation equation to update the positive definite symmetric value matrix. In each iteration, based on the updated positive definite symmetric value matrix... and perturbation suppression level Calculate and synchronously update the control gain matrix corresponding to each fuzzy rule. and interference gain matrix Repeat the iteration until the positive definite symmetric value matrix of two consecutive iterations is obtained. The norm of the difference is less than the preset minimum value The convergence condition is met, and the optimal control strategy and the optimal disturbance strategy are obtained based on the converged optimal positive definite symmetric value matrix.

[0112] Step S3 includes the following steps:

[0113] S31, Based on continuous sampling time of the system , , ..., Real-time operational data collected internally is used to analyze data along the system's trajectory within a finite time window. Integral transformations within the range are used to construct state difference data matrices. State-quadratic integral data matrix State-input cross-integral data matrix and state-perturbation cross-integral data matrix Specifically as follows:

[0114] ;

[0115] ;

[0116] ;

[0117] ;

[0118] in, The preset sampling time interval is explicitly reflected in the matrix. exponential decay term middle, This is the Kronecker product, used to transform the correlation matrix into a vector dimension to facilitate subsequent least squares solutions. As a time integration variable, it is used to perform integration calculations on system operating data within a finite time window, and is related to time. Belonging to the same type of time dimension, To expand the transition value of the state vector, ;

[0119] It is an exponentially decaying term. For the first The transpose of the control input under the given fuzzy rules. No. The transpose of the corresponding interference input under the fuzzy rule. For integration variables;

[0120] S32. Initialization parameters: Set the preset minimum value As a convergence criterion, it is applied to the control gain matrix. and interference gain matrix Select it in the first The initial control gain matrix corresponding to the fuzzy rule and interference gain matrix Set the number of iterations =0;

[0121] S33, in the In the next iteration, based on the control gain matrix of the current iteration and interference gain matrix Using the aforementioned integral data matrix, solve the following data-driven equation (i.e., the policy evaluation equation) to obtain the positive definite symmetric value matrix of the current iteration. :

[0122] ;

[0123] in, ;

[0124] ;

[0125] ;

[0126] in, for Vectorized arrangement of independent elements The regression matrix is ​​composed of data collected online. The identity matrix corresponding to the state vector. Represents the matrix vectorization operator. ;

[0127] S34. Construct the performance function of the controlled augmented system. Its integral expression is:

[0128]

[0129] Construct the corresponding Hamiltonian function based on the performance function. :

[0130]

[0131] Based on the theory of optimality, and applying the extreme value condition Using the obtained positive definite symmetric value matrix ,renew The control gain matrix of the next iteration and interference gain matrix ;

[0132] ;

[0133] in, For the system's performance function, Let Hamiltonian be the system function. The performance function with respect to the extended state vector The partial derivatives, Let be the transformation matrix of the partial derivative functions. and These are the partial derivatives of the Hamiltonian function with respect to the control input and interference input under the current fuzzy rules, respectively.

[0134] S35. Repeat steps S33 to S34 iteratively until the convergence condition is met. Otherwise Continue iterating to obtain the optimal positive definite symmetric value matrix. And simultaneously obtain the optimal positive definite symmetric value matrix. The corresponding theoretical optimal control gain matrix and the theoretically optimal interference gain matrix .

[0135] To prove the convergence of the policy iteration method, a sequence of value functions is constructed and differencing is performed. When convergence is achieved, the following equation is satisfied:

[0136]

[0137] in, and These represent the steer-by-wire system in the first... The general control gain matrix and general disturbance gain matrix corresponding to the fuzzy rule are used as abstract theoretical benchmark variables in the derivation of the strategy iterative optimization and convergence proof in this invention. However, in the specific... In the next iteration, this is specified as the control gain matrix of the current iteration. and interference gain matrix And eventually converges to the theoretically optimal control matrix. and the theoretically optimal interference gain matrix .

[0138] Based on the above equation derivation, it can be concluded that when the iteration approaches infinity... and Therefore, it has been rigorously proven that the strategy iteration method of the present invention has the following convergence result:

[0139] ;

[0140] ;

[0141] ;

[0142] The convergence result shows that, regardless of the initial state of the system, the data-driven strategy iteration method of the present invention can stably find the optimal game saddle point.

[0143] Combination Figure 3 As shown, this invention demonstrates the convergence trajectory of key control parameters during strategy iteration. Figure 3 middle x-axis This represents the number of iterations.

[0144] Numerical stability of parameter gain: The two figures on the left show the norm of the control gain matrix. (Top left) and the norm of the interference gain matrix (Bottom left) The change with the number of iterations. It can be seen that in the initial stage of iteration ( 1) The gain value jumps due to the initial strategy adjustment, but from the second iteration onwards, the gain norm quickly approaches a stable constant, proving that the control law has searched for a stable solution near the game saddle point.

[0145] Quantitative representation of the convergence criterion: The two figures on the right show the norm of the difference in the gain matrix between two adjacent iterations, i.e. (top right) and (Bottom right). These two figures directly correspond to the convergence condition in step S35. Experimental data show that as the number of iterations increases, the difference between adjacent gains... After equaling 2, the value rapidly decays and approaches 0.

[0146] Conclusion Analysis: This rapid convergence trend fully demonstrates that the data-driven strategy iterative algorithm designed in this invention has extremely high computational efficiency. Even without relying on the internal system model, it can achieve accurate convergence of the control strategy within a very short number of iterations (about 2 iterations), thereby meeting the dual requirements of real-time performance and robustness of the steer-by-wire system.

[0147] S4. Apply the optimal control strategy to the steer-by-wire system to achieve trajectory tracking control.

[0148] Based on the optimal positive definite symmetric value matrix The optimal control strategy obtained and optimal interference strategy The expressions are as follows:

[0149] ;

[0150] .

[0151] The optimal control strategy is applied to the steer-by-wire system;

[0152] Under the control of the optimal control strategy, the steer-by-wire system meets the following performance constraints:

[0153] .

[0154] This formula rigorously guarantees, both physically and mathematically, that no matter how severe the nonlinear friction, road impact, or other lumped disturbances the steer-by-wire system encounters... The sum of the global tracking error energy generated by the system and the control energy consumed (left side of the formula) will always be strictly limited to the total disturbance energy (right side of the formula). Within a few times, the system achieved strong robust anti-interference capability.

[0155] Combination Figure 4The figure shows the actual trajectory tracking performance of this optimal control strategy in a steer-by-wire system. The horizontal axis represents time, and the red dashed line represents the reference trajectory output vector. The blue dashed line represents the actual output vector of the steer-by-wire system. In the initial stage of the simulation ( (seconds), due to the system being in a state of severe nonlinear disturbance and a large initial state error, the actual output trajectory deviates from the reference trajectory; however, under the optimal control strategy... After intervention, the blue trajectory quickly converged towards the red target trajectory, and... They perfectly overlapped after a few seconds. Figure 4 The magnified view further illustrates the tracking error signal more intuitively. The attenuation process (shown by the green dashed line) shows that the value converges smoothly to zero within a very short time and no longer produces significant oscillations. This verifies that the control method of the present invention not only has excellent robustness but also extremely high trajectory tracking accuracy.

[0156] To support the execution of the above methods, embodiments of the present invention also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above methods.

[0157] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the various steps of the aforementioned method.

[0158] Furthermore, the present invention also provides a hardware architecture for a steer-by-wire system, including a steering motor, a feedback motor, an electronic control unit, and a front wheel actuator, wherein: the steering motor is used to execute control commands and drive the front wheels to steer; the feedback motor is used to collect steering wheel angle and speed signals; the electronic control unit is used to calculate and run the optimal control strategy online; and the front wheel actuator is used to realize the steering action.

[0159] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of the invention (including the claims) is limited to these examples; within the framework of the invention, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of the different aspects of the invention as described above, which are not provided in the details for the sake of brevity.

[0160] This invention is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A fuzzy tracking control method based on policy iteration, characterized in that, Includes the following steps: S1. Establish a TS fuzzy model of the steer-by-wire system, introduce a reference trajectory, and construct an augmented system containing the reference trajectory based on the TS fuzzy model. When the augmented system applies a control strategy, it is recorded as a controlled augmented system. S2. Based on the Hamilton-Bellman equations, construct the fuzzy discounted coupled algebraic Riccati equations corresponding to the augmented system. These equations contain a positive definite symmetric matrix used to evaluate the performance of the controlled augmented system. and preset disturbance suppression level ; The fuzzy discount coupling algebraic Riccati equation introduces a discount factor greater than or equal to zero. And use it to construct an exponentially decaying weight term. To balance the weight distribution of the controlled augmented system performance on the time axis, and through the control gain matrix With interference gain matrix The coupling achieves a zero-sum game between minimizing control costs and maximizing the impact of disturbances; In step S2, the fuzzy discount coupling algebraic Riccati equation is: in, The first one to be solved The positive definite symmetric value matrix corresponding to the fuzzy rules. A discount factor greater than or equal to zero This is the preset disturbance suppression level; and , , The state penalty weighting matrix is... This is the system trajectory tracking error output matrix. and It is a positive definite weighted matrix. , , , These are the transposes of the corresponding matrices; S3. Acquire real-time operating data of the online-collected steer-by-wire system, perform an integral transformation along the augmented system's operating trajectory within a finite time window, and use the real-time operating data to perform an equivalent transformation on the fuzzy discounted coupled algebraic Riccati equation, constructing and iteratively solving the strategy evaluation equation to update the positive definite symmetric value matrix. In each iteration, based on the updated positive definite symmetric value matrix... and perturbation suppression level Calculate and synchronously update the control gain matrix corresponding to each fuzzy rule. and interference gain matrix Repeat the iteration until the positive definite symmetric value matrix of two consecutive iterations is obtained. The norm of the difference is less than the preset minimum value The convergence condition is met, and the optimal control strategy and the optimal disturbance strategy are obtained based on the converged optimal positive definite symmetric value matrix. S4. Apply the optimal control strategy to the steer-by-wire system to achieve trajectory tracking control.

2. The fuzzy tracking control method based on policy iteration according to claim 1, characterized in that, In step S1, a fuzzy model of the steer-by-wire system (TS) is established, which specifically includes the following steps: S11. The dynamic equations of the steer-by-wire system are established as follows: in, For the steering wheel angle, For the angular velocity of the steering wheel, For the steering wheel angular acceleration, For rotational inertia, The viscous damping coefficient is... For nonlinear uncertainties, To control the input, Lumped disturbance term, The transmission ratio; S12, Define the state vector as follows ,in , Define control input The dynamic equations are expressed as a state-space model: ; The output equation is ; in, , , , For state vectors, Let be the derivative of the state vector with respect to time. For the output vector, The local state matrix, For local input matrices, For local perturbation matrix, For local output matrix, superscript Represents the transpose of a vector; S13. Using the TS fuzzy modeling method, local linear approximation is performed on the nonlinear uncertainty terms to obtain the overall fuzzy model: ; ; in, The total number of fuzzy rules, This represents a fuzzy rule index. The normalized weight function satisfies , Given a vector of presupposition variables, , , , The first The local state matrix, local input matrix, local perturbation matrix, and local output matrix corresponding to each fuzzy rule; S14, Define the reference trajectory as ; in, The reference trajectory state vector, Let be the derivative of the reference trajectory state vector with respect to time. Output vector as reference trajectory. and It is a constant matrix; S15. Introducing an extended state vector and extended output vector This yields an augmented system: in, , , , , and These are the transposes of the state vector and the reference trajectory state vector, respectively. and These are the transposes of the output vector and the reference trajectory output vector, respectively. , , , The augmentation system is in the first The local augmented state matrix, local augmented input matrix, local augmented output matrix, and local augmented perturbation matrix under the fuzzy rules.

3. The fuzzy tracking control method based on policy iteration according to claim 2, characterized in that, Step S3 includes the following steps: S31. Based on the system operation data collected online, utilize the system operation trajectory within a limited time window... Integral transformations within the range are used to construct state difference data matrices. State-quadratic integral data matrix State-input cross-integral data matrix and state-perturbation cross-integral data matrix : ; ; ; ; in, , , ..., For continuous sampling times during system operation, The preset sampling time interval, For Kronecker product, As a time integration variable, it is used to perform integration calculations on system operating data within a finite time window, and is related to time. Belonging to the same type of time dimension, To expand the state vector's transformation value; ; It is an exponentially decaying term. For the first The transpose of the control input under the given fuzzy rules. No. The transpose of the corresponding interference input under the fuzzy rule. For integration variables; S32. Initialization parameters: Set the preset minimum value For the control gain matrix and interference gain matrix Select it in the first The initial control gain matrix corresponding to the fuzzy rule and interference gain matrix Set the number of iterations =0; S33, in the In the next iteration, based on the control gain matrix of the current iteration and interference gain matrix Solve the following data-driven equation using the integral data matrix to obtain the positive definite symmetric value matrix of the current iteration. : ; in, ; ; ; in, The regression matrix is ​​composed of data collected online. The identity matrix corresponding to the state vector. Represents the matrix vectorization operator, where ; S34. Constructing the performance function of the controlled augmented system Its integral expression is: Construct the corresponding Hamiltonian function based on the performance function. : Based on the theory of optimality, and applying the extreme value condition Using the obtained positive definite symmetric value matrix ,renew The control gain matrix of the next iteration and interference gain matrix : ; in, For the system's performance function, Let Hamiltonian be the system function. The performance function with respect to the extended state vector The partial derivatives, Let be the transformation matrix of the partial derivative functions. and These are the partial derivatives of the Hamiltonian function with respect to the control input and interference input under the current fuzzy rules, respectively. S35. Repeat steps S33 to S34 iteratively until the convergence condition is met. Obtain the optimal positive definite symmetric value matrix .

4. The fuzzy tracking control method based on policy iteration according to claim 3, characterized in that, Based on the optimal positive definite symmetric value matrix The optimal control strategy and optimal interference strategy The expressions are as follows: ; 。 5. The fuzzy tracking control method based on policy iteration according to claim 4, characterized in that, Under the control of the optimal control strategy, the steer-by-wire system meets the following performance constraints: ; in, For exponentially decaying weight terms, To expand the state vector The transpose of the matrix, The state penalty weighting matrix is... For optimal control strategy The transpose of the matrix, Optimal interference strategy The transpose of the matrix, For time derivative.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes a computer program, it implements the steps of the method as described in any one of claims 1 to 5.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 5.

8. A steer-by-wire system, comprising a steering motor, a feedback motor, an electronic control unit, and a front wheel actuator, characterized in that, The electronic control unit is configured to perform the method as described in any one of claims 1 to 5, wherein: The steering motor is used to execute control commands and drive the front wheels to steer; The feedback motor is used to collect steering wheel angle and speed signals; The electronic control unit is used to run the optimal control strategy; The front wheel actuator is used to achieve steering.