An accelerated integrated value iteration control method for dual-spin stabilized systems
Through the integrated value iteration method, combined with the relaxation factor and adaptive relaxation function, the problems of poor adaptability and slow convergence of the RTAC system are solved, fast and stable intelligent optimization control is achieved, and the control efficiency and stability of the system are improved.
Patent Information
- Application Number
- CN202310712294.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-16
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2043-06-16
AI Technical Summary
Existing technologies are unable to effectively solve the problems of poor adaptability, slow convergence speed and low control efficiency of RTAC systems. Especially in the optimal control of nonlinear systems, traditional optimization control methods are difficult to achieve fast and stable intelligent optimization.
The integrated value iteration method is adopted, combining traditional value iteration and new value iteration. By introducing relaxation factors and adaptive relaxation functions, the convergence speed of adaptive regulation is designed to ensure the stability of the system and the admissibility of the control strategy, and the adaptive dynamic programming method is used to accelerate the iteration process.
The rapid convergence and stable control of the RTAC system are achieved, the control efficiency of the nonlinear system is improved, and the stability and interference suppression capability of the system under limited interference are ensured.
Smart Images

Figure CN116654295B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of spacecraft. Background Art
[0002] The attitude control system of a spacecraft is a crucial component. This system acquires and maintains the spacecraft's orientation in space and its attitude relative to a reference coordinate system, and is a key technology for achieving standardization of spacecraft platforms. Therefore, designing a suitable attitude control system is crucial for the stable on-orbit operation of a spacecraft and the normal operation of its payloads, and it continues to attract extensive research and discussion. Dual-spin stabilization systems are a common approach for spacecraft attitude control. Their basic principle is to ensure that the spacecraft body rotates at high speed around its principal axis of inertia in space and utilize the gyroscopic effect to maintain its inertial orientation. This method offers high reliability, a simple control system, and significant resistance to large disturbance torques, making it widely used in spacecraft attitude control systems. The rotary / translational actuator (RTAC) experiment with rotational excitation was initially developed as a simplified dual-spin spacecraft model. This system is mathematically and qualitatively equivalent to a dual-spin spacecraft, sharing similar average equations and exhibiting similar dynamic behavior. Therefore, the attitude control of a dual-spin spacecraft can be achieved through simulation and control of the RTAC system. However, the system's rotational and translational motions are coupled, and the system exhibits internal nonlinearities, uncertainties, and interference, resulting in dynamic complexity. This makes direct optimization of the system difficult, and traditional optimization control methods suffer from poor adaptability, slow convergence, and low control efficiency. Therefore, it is necessary to develop advanced intelligent optimization control methods for RTAC systems.
[0003] Intelligent optimization methods are widely used in fields such as control theory, computer science, and computational mathematics. Optimization concepts play a crucial role in advanced control design based on artificial intelligence and are of great significance for the construction of various intelligent systems. However, unlike general linear situations, optimal control of nonlinear systems is often difficult to solve. Reinforcement learning, characterized by agent-environment interaction, is closely related to dynamic programming in intelligent optimization design. In an adaptive evaluation framework, reinforcement learning is combined with approximate structures to approximate complex optimization problems. In recent years, adaptive dynamic programming (ADP) has been widely used to solve complex optimal control problems and has achieved many outstanding results in adaptive optimal control design. Therefore, the present invention implements intelligent control of RTAC systems based on the ADP method. The core task of ADP is to iteratively solve the Hamilton-Jacobi-Bellman (HJB) equations for nonlinear systems. The iterative ADP algorithm primarily consists of value iteration and policy iteration. In the policy iteration algorithm, an admissible control strategy is required to initialize the iterative algorithm to ensure the admissibility of the iterative control strategy generated by the policy iteration. However, in each iteration, the policy iteration uses a successive approximation method to evaluate the policy, which introduces additional computational cost. Value iteration algorithms can be initialized with any positive semidefinite cost function, but the admissibility of the iterative control strategy is unknown, and there is no guarantee that a stable control strategy will be obtained during the iteration process. Currently, few methods can effectively achieve accelerated convergence of iterative ADP methods. Therefore, there is an urgent need to design advanced intelligent optimization controllers that can accelerate the convergence of the cost function while ensuring system stability, while still obtaining the optimal control strategy, improving the optimal control efficiency of nonlinear systems, and enhancing the control performance of ADP methods. Summary of the Invention
[0004] Based on an iterative adaptive evaluation framework, this paper proposes an integrated value iteration method with guaranteed convergence rate to solve the intelligent optimization control problem of nonlinear systems. The paper focuses on the convergence rate of the value iteration method and proposes a novel integrated value iteration scheme. By introducing a relaxation factor and designing an adaptive relaxation function, the convergence rate of the cost function can be adjusted during the iteration process. At the same time, this integrated value iteration scheme does not introduce additional computational cost and ensures system stability.
[0005] The RTAC nonlinear benchmark problem considers a nonlinear fourth-order dynamic system consisting of a translational oscillator and a centrifugally rotating pendulum. Figure 1The translational oscillator shown consists of a cart with mass M connected to a fixed wall by a linear spring with stiffness k. Its motion is limited to one dimension, that is, only in the horizontal plane, so gravity has no effect. The translational position of the cart is q, and the speed of the cart is The acceleration of the car is The pendulum ball installed at the center of the cart can rotate in the horizontal plane. Its mass is m and the angle of rotation is θ. The angular velocity of the pendulum ball is The angular acceleration of the pendulum ball is The moment of inertia of the pendulum's center of mass is I, the distance between the pendulum's center of mass and its rotation point is e, and N represents the control torque applied to the pendulum. For this system, the control objective is to stabilize the oscillator by applying the control torque to the centrifugally rotating pendulum. The designed controller must ensure internal stability and exhibit good interference rejection for certain signals, even with limited control effectiveness.
[0006] Therefore, through mechanism modeling, the model of RTAC system can be obtained as follows:
[0007]
[0008] According to the above RTAC system model, let the translation position q of the car and the speed of the car be The angle θ of the ball's rotation, and the angular velocity of the ball's rotation The four components x1, x2, x3, and x4 of the system state are respectively, then the state of the system is The control torque N applied to the ball is the control input u of the system. In addition, the coupling between translation and rotation is assumed to be The ordinary differential equation satisfied by the known system state is the state equation of the control system. The differential form of the system state x is written as So the state equation of the RTAC system can be obtained as:
[0009]
[0010] Next, we will study the control of the RTAC system, that is, to suppress the horizontal vibration of the system and stabilize the translational position of the car and the rotation angle of the pendulum ball to the system equilibrium point, so that [x1, x2, x3, x4] T =[0,0,0,0] T Therefore, the present invention realizes stable control of the RTAC system based on the integrated value iteration method. The detailed steps of the integrated value iteration intelligent control design are described as follows:
[0011] Step 1: Problem transformation. The problem of achieving oscillator stability in the RTAC system is transformed into an optimal control problem for a nonlinear system. The state equation of the RTAC system is discretized using the Euler method. The discrete time interval is selected as 0.1s. The current moment is set as k. The state components of the discretized system are expressed as x 1k 、x 2k 、x 3k 、x 4k , and the corresponding system state at the next moment is x k+1 , the control strategy of the system is expressed as u k , so the corresponding system state space expression can be obtained as follows:
[0012]
[0013] The system can be viewed as a fourth-order nonlinear non-affine system, namely
[0014]
[0015] Among them, F(·,·) is a continuous system function, x k is the system state vector, represents the set of non-negative integers, i.e. If x0 is the initial state of the system, then x0 is the only equilibrium point of the system when u = 0, that is, F(0,0) = 0, which means that there is a control sequence that can make the system state x0 when k → 0 k →0. Assume that the optimal feedback control strategy of the system is u(x k ), the utility function is U(x k ,u(x k )), select it as a quadratic form, that is, Where Q and R are positive definite matrices whose dimensions match the system state and control. Let the cost function of the system be V(x k ,u(x k )), for the optimal control problem of this system, the goal is to find a suitable feedback control strategy u(x k ) makes the system stable and minimizes the following infinite time cost function:
[0016]
[0017] Where U(0,0)=0 and Here, the system cost function V(x k ,u(x k )) and feedback control strategy u(x k ) is abbreviated as V(x k ) and u k The optimal cost function of the system at the current time k is expressed as The corresponding optimal control strategy is expressed as Then the optimal cost function of the system at the next moment k+1 is V * (F(x k ,u k )), according to the Bellman optimality principle, the HJB equation of the system can be obtained
[0018]
[0019] The corresponding expression of the optimal control strategy is
[0020]
[0021] However, for nonlinear systems, the exact solution of the HJB equation is difficult to obtain, so the ADP method is used to obtain its approximate optimal solution, that is, to obtain the approximate optimal control strategy.
[0022] Step 2: Construct an integrated value iteration control framework. Using the adaptive judgment framework, combining traditional value iteration and new value iteration, the HJB equation of the nonlinear system is approximately solved.
[0023] For traditional value iteration, its cost function and control strategy are expressed as V T (x k ) and u Tk , let the iteration index and iteration termination error be i=1,2,… and δ. The cost function and control strategy in the iterative process are and Then the cost function of the next iteration step is And the cost function at the next moment k+1 is Use any semi-positive function to initialize the initial cost function, and update it through the cost function
[0024]
[0025] and strategy improvement
[0026]
[0027] Alternate iteration until the absolute value of the difference between adjacent cost functions is When , the iterative process stops and the approximately optimal control strategy is obtained.
[0028] For the new value iteration, the system cost function and control strategy under this method are expressed as and The corresponding cost function and control strategy in the iterative process are expressed as and Then the cost function of the next iteration step is And the cost function at the next moment k+1 is Compared with traditional value iteration, the relaxation factor ω>0 is introduced. The corresponding cost function update and strategy improvement process are as follows:
[0029]
[0030] and
[0031]
[0032] Similarly, by iterating (10) and (11) alternately, the absolute value of the difference between adjacent cost functions under the new value iteration method is obtained. When ω=1, the iterative process stops and the approximate optimal control strategy is obtained. In particular, when ω=1, the new value iteration method is equivalent to the traditional value iteration method.
[0033] Step 3: Introduce an adaptive relaxation function to establish an acceleration value iteration scheme.
[0034] According to the method of successive super-relaxation, it can be found that if the relaxation factor satisfies a certain range, the larger the relaxation factor, the faster the cost function sequence converges. Therefore, when ω>1, the proposed new value iteration method has a faster convergence speed than the traditional value iteration. However, the upper limit of the relaxation factor that guarantees the convergence of the cost function is unknown. Therefore, when ω>1, how to ensure the convergence speed of the iterative algorithm is an important issue. Therefore, a practical relaxation factor design method is proposed to ensure the convergence speed. The iterative process is divided into an acceleration stage and a convergence stage, such as Figure 2 shown.
[0035] In the acceleration phase, the relaxation factor is greater than 1, which accelerates the convergence process of the cost function. After the acceleration phase, the relaxation factor is set to 1 to ensure that the iterative cost function converges to the optimal cost function. The monotonicity of the cost function sequence cannot be guaranteed during the iteration process. When the acceleration phase is set too long or the relaxation factor is set too large, the iterative cost function may be greater than the optimal cost function of the system, resulting in oscillation in the acceleration phase and a decrease in the convergence phase. Therefore, the selected relaxation coefficient should not be too large. In addition, a relaxation function ω(i) with respect to the iteration index i is defined, where α>0 and β>1 are both adjustable parameters of the relaxation function and ω(i)∈(1,β). In order to ensure that the function is within this range, the relaxation function ω(i) is set to an exponential function with the natural constant e as the base, and α is set as the variable parameter at the exponential position, and (β-1) is the parameter at the coefficient position, that is, the relaxation function is set as follows:
[0036] ω(i)=(β-1)e -αi +1 (12)
[0037] Based on this, we can get that for any iterative index The function ω(i) is monotonically decreasing and satisfies When β = 1, the relaxation function ω(i) = 1, which means that the new value iteration is transformed into the traditional value iteration. According to the relaxation function in (11), the corresponding cost function can be updated as
[0038]
[0039] The relaxation function is used to gradually reduce the relaxation factor to 1, that is, gradually make β=1, thereby realizing the transition from the new fast value iteration scheme to the traditional value iteration scheme.
[0040] Step 4: Implement intelligent control of the RTAC system.
[0041] Based on the above-mentioned integrated value iteration framework and the introduction of the adaptive relaxation function, an algorithm scheme is established. First, the parameters in the system utility function, cost function, iteration index, relaxation factor size, relaxation function parameters, and iteration termination error are initialized. Then, according to equations (10) and (11) or (13), alternating iterations are performed until the iteration termination error is reached. In this way, the approximate optimal cost function and control strategy of the RTAC system can be obtained, and intelligent optimization control of the system can be achieved.
[0042] The innovation of the present invention is:
[0043] This paper proposes an integrated value iteration scheme with adjustable convergence speed for complex nonlinear RTAC systems by introducing a relaxation factor and an adaptive relaxation function. Combining traditional value iteration with a novel value iteration method, the influence of the relaxation factor on system convergence is studied. To facilitate adjusting the convergence speed of the cost function during the iteration process, a relaxation function is designed to adaptively accelerate the iteration process. This approach achieves rapid system convergence while ensuring that the control strategy derived from integrated value iteration stabilizes the closed-loop system. Ultimately, intelligent optimal control of the nonlinear RTAC system is achieved, improving control efficiency while maintaining effective system control. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 RTAC system schematic
[0045] Figure 2 Schematic diagram of the accelerated integrated value iteration method
[0046] Figure 3 Convergence curve of the norm of the approximate cost function weight vector (ω=1)
[0047] Figure 4 Convergence curve of the norm of the approximate cost function weight vector (ω=2)
[0048] Figure 5 Convergence curve of the approximate cost function weight vector norm (α=0.01,β=4)
[0049] Figure 6 Convergence curve of the weight matrix parameters of the approximate cost function (α=0.01,β=4)
[0050] Figure 7 System state trajectory (α=0.01,β=4)
[0051] Figure 8 System control input trajectory (α=0.01,β=4) DETAILED DESCRIPTION
[0052] This section verifies the effectiveness of the proposed algorithm by conducting specific experiments. The quadratic utility function is selected as That is, Q and R in the system utility function are selected as I4 and 0.05I, where I4 and I are 4×4 and 1×1 unit matrices respectively, and the maximum iteration index is set to 600. The following function approximation structure is selected to approximate the system cost function:
[0053]
[0054] in, is the weight parameter vector of the approximate structure, Represents a set of real numbers, and its initial value is selected as 0. 100 initial states are randomly selected to iterate the approximate cost function and approximate control strategy through equations (10) and (11).
[0055] In order to verify the effectiveness of introducing relaxation factors and adaptive relaxation functions on the rapid convergence of the cost function in the iterative process, the relaxation factors ω are selected as 1 and 2 respectively in the simulation experiment, and the adaptive relaxation function parameters are selected as α = 0.01, β = 4, that is, corresponding to the traditional value iteration, the integrated value iteration method under the introduction of relaxation factors and adaptive relaxation functions can obtain the corresponding system convergence effect diagram as shown in the figure below. Figure 3 、 4 , 5. From the above experimental results, it can be seen that the cost function tends to converge at about 500 iteration steps under the traditional value iteration method, while the cost function tends to converge at about 300 and 200 iteration steps under the introduction of relaxation factor and adaptive relaxation function respectively. Therefore, when the relaxation factor ω>1, the convergence of the system cost function in the iteration process can be accelerated, and the adaptive relaxation function has stronger adaptability and more accurate adjustment range than the directly assigned relaxation factor. At the same time, based on the accelerated integrated value iteration, the convergence diagram of the weight parameters of the approximate cost function of the system is shown in Figure 6As shown in the figure, it can be seen that the cost function in the iterative process gradually approaches the optimal value. Using the approximate optimal control strategy obtained by this method to control the system, the corresponding state trajectory and control trajectory can be obtained as follows: Figure 7 and 8 It can be seen from this that the accelerated integrated value iteration framework proposed in the present invention can not only realize the intelligent optimization control of nonlinear systems and ensure the stability of the system, but also accelerate the convergence process of the system cost function and improve the control efficiency.
Claims
1. An accelerated integrated value iteration control method for a dual-spin stabilized system, characterized by: RTAC nonlinear benchmark problem Consider a nonlinear fourth-order dynamic system consisting of a nonlinear interaction between a translational oscillator and a centrifugally rotating pendulum. The oscillator consists of a cart of mass M connected to a fixed wall by a linear spring of stiffness k. Its motion is limited to one dimension, that is, only in the horizontal plane, so gravity has no effect. The translational position of the cart is q, and the velocity of the cart is The acceleration of the car is The pendulum ball installed at the center of the cart can rotate in the horizontal plane. Its mass is m and the angle of rotation is θ. The angular velocity of the pendulum ball is The angular acceleration of the pendulum ball is The moment of inertia of the pendulum's center of mass is I, the distance between the pendulum's center of mass and its rotation point is e, and N represents the control torque applied to the pendulum. For this system, the control objective is to stabilize the oscillator by applying the control torque to the centrifugally rotating pendulum. The designed controller must ensure internal stability. The model of the RTAC system obtained through mechanism modeling is: According to the above RTAC system model, let the translation position q of the car and the speed of the car be The angle θ of the ball's rotation, and the angular velocity of the ball's rotation The four components x1, x2, x3, and x4 of the system state are respectively, then the state of the system is The control torque N applied to the ball is the control input u of the system; in addition, the coupling between translation and rotation is assumed to be The ordinary differential equation satisfied by the known system state is the state equation of the control system. The differential form of the system state x is written as So the state equation of the RTAC system is: Next, we will study the control of the RTAC system, that is, to suppress the horizontal vibration of the system and stabilize the translational position of the car and the rotation angle of the pendulum ball to the system equilibrium point, so that [x1, x2, x3, x4] T =[0,0,0,0] T The stable control of the RTAC system is achieved based on the integrated value iteration method. The detailed steps of the integrated value iteration intelligent control design are described as follows: Step 1: Problem transformation; The problem of achieving oscillator stability in the RTAC system is transformed into an optimal control problem for a nonlinear system. The state equation of the RTAC system is discretized using the Euler method. The discrete time interval is selected as 0.1s. The current moment is set as k. The state components of the discretized system are expressed as x 1k 、x 2k 、x 3k 、x 4k , and the corresponding system state at the next moment is x k+1 , the control strategy of the system is expressed as u k , so the corresponding system state space expression is as follows: The system is considered as a fourth-order nonlinear non-affine system, namely Among them, F(·,·) is a continuous system function, x k is the system state vector, represents the set of non-negative integers, i.e. If x0 is the initial state of the system, then x0 is the only equilibrium point of the system when u = 0, that is, F(0,0) = 0, which means that there is a control sequence that can make the system state x0 when k → 0 k →0; let the optimal feedback control strategy of the system be u(x k ), the utility function is U(x k ,u(x k )), select it as a quadratic form, that is, Where Q and R are positive definite matrices whose dimensions match the system state and control; let the cost function of the system be V(x k ,u(x k )), for the optimal control problem of this system, the goal is to find a suitable feedback control strategy u(x k ) makes the system stable and minimizes the following infinite time cost function: Where U(0,0)=0 and Here, the system cost function V(x k ,u(x k )) and feedback control strategy u(x k ) is abbreviated as V(x k ) and u k ; The optimal cost function of the system at the current time k is expressed as The corresponding optimal control strategy is expressed as Then the optimal cost function of the system at the next moment k+1 is V * (F(x k ,u k )), according to Bellman optimality principle, the HJB equation of the system is obtained The corresponding expression of the optimal control strategy is However, for nonlinear systems, the exact solution of the HJB equation is difficult to obtain, so the ADP method is used to obtain its approximate optimal solution, that is, to obtain the approximate optimal control strategy; Step 2: Build an integrated value iteration control framework; Its cost function and control strategy are expressed as V T (x k ) and u Tk , let the iteration index and iteration termination error be i=1,2,... and δ; the cost function and control strategy in the iteration process are and Then the cost function of the next iteration step is And the cost function at the next moment k+1 is Use any semi-positive function to initialize the initial cost function, and update it through the cost function and strategy improvement Alternate iteration until the absolute value of the difference between adjacent cost functions is When , the iterative process stops and the approximate optimal control strategy is obtained; For the new value iteration, the system cost function and control strategy under this method are expressed as and The corresponding cost function and control strategy in the iterative process are expressed as and Then the cost function of the next iteration step is And the cost function at the next moment k+1 is The corresponding cost function update and strategy improvement process are as follows: and Similarly, by iterating (10) and (11) alternately, the absolute value of the difference between adjacent cost functions under the new value iteration method is obtained. When ω = 1, the iterative process stops and the approximate optimal control strategy is obtained; among them, when ω = 1, the new value iteration method is equivalent to the traditional value iteration method; Step 3: Introduce an adaptive relaxation function to establish an acceleration value iteration scheme; The iterative process is divided into the acceleration phase and the convergence phase: In the acceleration phase, the relaxation factor is greater than 1, which accelerates the convergence process of the cost function. After the acceleration phase, the relaxation factor is set to 1 to ensure that the iterative cost function converges to the optimal cost function. A relaxation function ω(i) with respect to the iteration index i is defined, where α>0 and β>1 are both adjustable parameters of the relaxation function and ω(i)∈(1,β). In order to ensure that the function is within this range, the relaxation function ω(i) is set to an exponential function with the natural constant e as the base, and α is set as the variable parameter at the exponential position, and (β-1) is set as the parameter at the coefficient position, that is, the relaxation function is set as follows: ω(i)=(β-1)e -αi +1 (12) Based on this, for any iterative index The function ω(i) is monotonically decreasing and satisfies When β = 1, the relaxation function ω(i) = 1, which means that the new value iteration is transformed into the traditional value iteration; according to the relaxation function in (11), the corresponding cost function is updated as The relaxation function is used to gradually reduce the relaxation factor to 1, that is, gradually make β = 1, thereby realizing the transition from the new fast value iteration scheme to the traditional value iteration scheme; Step 4: Implement intelligent control of the RTAC system; The parameters in the system utility function, cost function, iteration index, relaxation factor size, relaxation function parameters, and iteration termination error are initialized, and then alternating iterations are performed according to equations (10) and (11) or (13) until the iteration termination error is reached, thereby obtaining the approximate optimal cost function and control strategy of the RTAC system and realizing intelligent optimization control of the system.
Citation Information
Patent Citations
Feedback control circuit for magnetic suspension and propulsion system
CA957738A
Dissipative structure theory based TORA (Translation oscillators with a rotating actuator) system self-adaption control method
CN104199291A