An Optimal Control Method for Complex-Valued Continuous-Time Nonlinear Systems

By dividing the control model of complex nonlinear systems into real and imaginary models, and constructing a unified complex form cost function, combined with three-layer complex forward neural network, the lack of optimal control algorithm for complex models in the existing technology is solved, and the effective optimal control of complex nonlinear systems and the asymptotic stability of the system are achieved.

CN115840357BActive Publication Date: 2025-06-17四川启睿克科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211318982.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-26
Publication Date
2025-06-17
Estimated Expiration
2042-10-26

AI Technical Summary

Technical Problem

The prior art lacks the optimal control algorithm application for complex models, especially in methods based on adaptive dynamic programming, and it is difficult to effectively control complex nonlinear systems.

Method used

A complex nonlinear system optimal control algorithm based on adaptive dynamic programming method is proposed. By segmenting the control model into real-part model and imaginary model, the cost function of real-part and imaginary parts is established, a unified complex form cost function is constructed, and the optimal control is achieved using a three-layer complex forward neural network.

Benefits of technology

The optimal control of complex nonlinear systems is achieved, and the effectiveness of the control strategy and the asymptotic stability of the system are ensured through segmentation model and construction cost function.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115840357B_ABST
    Figure CN115840357B_ABST
Patent Text Reader

Abstract

The present invention discloses an optimal control method for a complex-valued continuous-time nonlinear system. By splitting the control model into a real part model and an imaginary part model, two cost functions for the real part and the imaginary part are established to construct a complete cost function in complex form. A policy iteration optimal control algorithm is constructed through the cost function, a new sampling convergence analysis method is used to ensure convergence, and the optimal control algorithm is implemented through a three-layer complex-valued feedforward neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of optimal control, and particularly relates to an optimal control method for a complex-valued continuous-time nonlinear system. Background Art

[0002] Optimal control is a popular technical field in recent years. Since most problems in engineering practice have control constraints, modern variational theories for seeking optimal control strategies under constraint conditions have emerged. The most commonly used research methods are dynamic programming and the maximum principle. Many directions have been derived from optimal control, such as approximate control methods that combine dynamic programming and reinforcement learning to form adaptive dynamic programming, predictive control, etc. Complex numbers not only have important theoretical significance in the development of the logarithmic domain. With the development of science and technology, complex numbers have been widely used in system analysis, signal analysis, generalized integrals, quantum mechanics, fluid mechanics, etc. There are many discussions about complex numbers, involving the use of complex numbers, introducing the concept of complex numbers into neural networks, which all involve the research of complex neural networks. There are few algorithms for complex control in the method based on adaptive dynamic programming. Summary of the Invention

[0003] In traditional optimal control of nonlinear systems, it is basically for real-valued optimal control. Especially in the optimal control method based on adaptive dynamic programming, there is a lack of algorithm application for complex models. The present invention provides a new optimal control algorithm for a complex nonlinear system based on the adaptive dynamic programming method. Based on the proposed control model of the nonlinear system, it is divided into a real part model and an imaginary part model. By innovatively establishing the cost functions of the real part and the imaginary part to construct the cost function for the complex system, the policy iteration optimal control algorithm is completed through the constructed function, and a three-layer complex-valued feedforward neural network is constructed to achieve optimal control of the complex nonlinear system.

[0004] The present invention proposes a new control model and a method for constructing a cost function to complete the optimal control algorithm for a complex-valued continuous-time nonlinear system. The data used in the function model are all complex numbers. The traditional control model is divided into a real part model and an imaginary part model. By innovatively establishing the cost functions of the real part and the imaginary part, a unified complex-form cost function is constructed and a new sampling convergence method is used to complete the convergence analysis. The optimal control algorithm is realized by constructing a complex-valued feedforward neural network.

[0005] The present invention achieves the above object through the following technical solutions:

[0006] An optimal control method for a complex-valued continuous-time nonlinear system, comprising the following steps:

[0007] Step 1, dividing the control model of the nonlinear system into a real part model and an imaginary part model;

[0008] Step 2, construct the cost functions of the real and imaginary parts to construct the cost function for the complex number system;

[0009] Step 3, complete the policy iteration optimal control algorithm through the cost function, and construct a three-layer complex forward neural network to achieve the optimal control of the complex nonlinear system.

[0010] For a class of complex continuous-time nonlinear systems, the system dynamics are described as:

[0011]

[0012] where \(z(t)\in\mathbb{R}\) n is a complex state, \(f(\cdot)\in\mathbb{R}\) n is an unknown nonlinear function and satisfies \(f(0) = 0\), \(u(t)\) is the control input, \(C\in\mathbb{R}\) n×1 is the gain matrix.

[0013] Assume that the state \(z(t)\) is separated into the real and imaginary parts, expressed as:

[0014] \(z(t)=x(t)+iy(t),\ (2)\)

[0015] where \(x(t)\) and \(y(t)\) are the real and imaginary parts of \(z(t)\) respectively. \(f(z(t))\) is separated as:

[0016] \(f(z(t))=f\) R (x(t),y(t))+if I (x(t),y(t))\ (3)\)

[0017] According to (1)-(3), rewrite the real and imaginary parts as:

[0018]

[0019]

[0020] Then (4) and (5) are rewritten as:

[0021]

[0022] According to (6), inspired by the idea of adaptive dynamic programming, define the cost function as follows:

[0023]

[0024]

[0025] where \(Q\) and \(R\) are both positive definite matrices with corresponding dimensions.

[0026] The derivative of (7) with respect to \(t\) is:

[0027]

[0028] where is the sum of the partial derivatives of V with respect to x and y respectively.

[0029] A further solution is: According to the definition of the Hamiltonian function:

[0030]

[0031] It is found that the optimal control satisfies the Hamilton-Jacobi-Bellman equation, and the equation is as follows:

[0032]

[0033] Through (10), the following optimal control is established:

[0034]

[0035]

[0036] 2. Sampling Convergence Analysis

[0037] Through sampling, (7) is rewritten in discrete form to perform convergence analysis, and the discrete form is expressed as follows:

[0038]

[0039] where T is the sampling period.

[0040] (13) is also rewritten as:

[0041]

[0042] Taking the derivative of both sides of (14) with respect to t, we have:

[0043]

[0044] Lemma 1: Equation (15) is equivalent to the Hamilton-Jacobi-Bellman function (9).

[0045] In the formula It is proved that the cost function V(x(t), y(t), u R (t), u I (t)) is driven by u R and u I and reaches V(x(∞), y(∞), u R (∞), u I (∞)) = 0 along the corresponding state. It is obtained by integrating (9) in the time range [t, ∞]:

[0046]

[0047] Then we obtain:

[0048]

[0049] After taking the derivative of the above equation with respect to time t, the following equation holds:

[0050]

[0051] This is equivalent to the definition of (15). According to (15), we obtain:

[0052]

[0053] Obviously, for Equation (19) always holds; let T → ∞, then we obtain:

[0054]

[0055] This is equivalent to (9).

[0056] Note that u R and u I are synchronized in time, and they are iterated simultaneously to affect the state at the next moment. Regarding them as a control pair, the control pair u R and u I satisfies the optimal structures (11) and (12), so the optimal control pair is obtained through the following algorithm:

[0057] Starting from the allowed initial control pair and :

[0058] Step 1: Use the control pair and to solve the cost function V k in (9), as follows:

[0059]

[0060] Or equivalently:

[0061]

[0062] Step 2: Update the control pair according to (11) and (12):

[0063]

[0064]

[0065] Then go back to Step 1, and Stop when converging to the optimum.

[0066] Theorem 2. The control pair is obtained through the iterative step k in (23) and (24), and V k > 0, which satisfies the Hamilton-Jacobi-Bellman equation (21). As the cost function value decreases with the iterative steps, such as: V k+1 < V k . Therefore, each iterative step of the cost function is convergent, and the critical value is defined as V * , such as:

[0067] Proof: According to we get:

[0068]

[0069] This shows that:

[0070]

[0071] According to to drive the state to the next step, we have:

[0072] Then the HJB equation is From this, we get:

[0073]

[0074] Next, we calculate to get:

[0075]

[0076]

[0077] Obtained through (23) and (24):

[0078]

[0079] Next, we get:

[0080]

[0081] Since Q and R are symmetric positive definite, there exist orthogonal matrices M and N, such as Q = M T Λ1M, R = N T Λ2N, where represent the eigenvalues of Q and R respectively; (29) is written as:

[0082]

[0083] To simplify the expression, let And Then we can get:

[0084]

[0085]

[0086] We get V from (31) k+1 ≤V k . Because V k is positive definite, V k will converge to V * .

[0087] 3. Stability Analysis

[0088] Theorem 1. Considering the optimal solution V of the system (1) under the Hamilton-Jacobi-Bellman equation * and the optimal controllers (11) and (12), the system is asymptotically stable;

[0089] Proof: According to (7), it is easy to calculate it and obtain the derivative result as follows:

[0090]

[0091] Because V k will converge to V * , the optimal cost function value V * is positive definite. Select it as the Lyapunov function, and according to (32), we get the result:

[0092]

[0093] Based on the above analysis, we have This shows that the system (1) is asymptotically stable.

[0094] 4. Complex Forward Neural Network;

[0095] Currently, most of the data in neural networks are real-valued. Since the system model data is complex-valued, a complex forward network similar to the real-valued forward neural network is used to solve it. A three-layer neural network is designed, and the action-evaluation framework is used to implement the iterative algorithm in this paper. The specific method is as follows.

[0096] Complex Forward Neural Network Design

[0097] First, an evaluation network for approximating the cost function is proposed as follows

[0098]

[0099] where is the weight matrix of the first-layer evaluation network, with the superscript d representing the iteration step, and z c is the input of the evaluation network, denoted as z c ={x + iy, u R + iu I}, and O is the output of the first-layer evaluation network. O can be divided into a real part and an imaginary part:

[0100] O(t)= O R (t)+ iO I (t) (35)

[0101] Then, using the traditional method of "splitting" complex values, the activation function processes the real and imaginary parts of O1 separately, and the output V can be defined as:

[0102]

[0103] where represents the weight matrix of the second-layer evaluation network, and f c (·) represents the activation function of the evaluation network, and is selected as the activation function. Through the continuous function (7), its discretization is performed, and the evaluation network error can be regarded as:

[0104]

[0105] The minimum error function is defined as:

[0106]

[0107] To minimize the error function, the following method is used to update the weights of the evaluation neural network:

[0108]

[0109]

[0110]

[0111]

[0112] where μ is the learning rate.

[0113] Complex action network design

[0114] In this part, a three-layer action network is designed for approximate control as follows:

[0115]

[0116] where is the weight matrix of the first-layer action network, and z aRepresents the input of the action network, with the expression z a = {x + iy};

[0117] O1 is the output of the first-layer action network, and O1 is separated into real and imaginary parts:

[0118] O1(t) = O R1 (t) + iO I1 (t) (44)

[0119]

[0120] where f a (·) represents the activation function of the action network, and the activation function is selected Then U2 is separated into the sum of the following real and imaginary parts:

[0121] O2(t) = O R2 (t) + iO I2 (t) (46) Again, the activation function is used to obtain the output:

[0122]

[0123] The action network error is regarded as:

[0124]

[0125] where represents the expected value, which is set to 0. The minimum error function of the action network is defined as:

[0126]

[0127] To minimize the error function, the following method is used to update the weights of the action network:

[0128]

[0129]

[0130]

[0131]

[0132] The beneficial effects of the present invention are as follows:

[0133] An optimal control method for a complex-valued continuous-time nonlinear system according to the present invention divides a control model into a real part model and an imaginary part model, constructs a complete cost function in complex form by establishing two cost functions for the real part and the imaginary part, constructs a policy iteration optimal control algorithm through the cost function, uses a new sampling convergence analysis method to ensure convergence, and realizes the optimal control algorithm through a three-layer complex forward neural network. BRIEF DESCRIPTION OF THE DRAWINGS

[0134] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0135] Figure 1 It is an optimal control example of the complex-valued signal processing system of the present invention.

[0136] Figure 2 It is the real part evolution trajectory of the error e of the present invention.

[0137] Figure 3 It is the imaginary part evolution trajectory of the error e of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0138] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will describe the technical solutions of the present invention in detail. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope protected by the present invention.

[0139] In any embodiment, as Figure 1 shown, an optimal control method for a complex-valued continuous-time nonlinear system according to the present invention includes:

[0140] Step 1, dividing the control model of the nonlinear system into a real part model and an imaginary part model;

[0141] Step 2, establishing cost functions for the real part and the imaginary part to construct a cost function for the complex system;

[0142] Step 3, completing a policy iteration optimal control algorithm through the cost function, and constructing a three-layer complex forward neural network to realize the optimal control of the complex nonlinear system.

[0143] In a class of complex continuous-time nonlinear systems, the system dynamics are described as:

[0144]

[0145] where \(z(t)\in\mathbb{R}\) n is a complex state, \(f(\cdot)\in\mathbb{R}\) n is an unknown non - linear function and satisfies \(f(0) = 0\), \(u(t)\) is the control input, \(C\in\mathbb{R}\) n×1 is the gain matrix.

[0146] Assume that the state \(z(t)\) is separated into its real and imaginary parts, expressed as follows:

[0147] \(z(t)=x(t)+iy(t)\ (2)\)

[0148] where \(x(t)\) and \(y(t)\) are the real and imaginary parts of \(z(t)\) respectively. \(f(z(t))\) is separated as:

[0149] \(f(z(t)) = f\) R (x(t), y(t))+if I (x(t), y(t))\ (3)\)

[0150] According to (6), inspired by the idea of adaptive dynamic programming, the cost function is defined as follows:

[0151]

[0152] where \(Q\) and \(R\) are both positive - definite matrices with corresponding dimensions.

[0153] Through (10), the following optimal control is established:

[0154]

[0155]

[0156] Since the system - model data is complex - valued, a complex - valued feed - forward network similar to the real - valued feed - forward neural network is used to solve it. A three - layer neural network is designed, and the actor - critic framework is used to implement the iterative algorithm in this paper. The specific method is as follows.

[0157] Complex - valued Feed - forward Neural Network Design

[0158] First, a critic network for approximating the cost function is designed as follows

[0159]

[0160] where is the weight matrix of the first - layer critic network, the superscript \(d\) represents the iteration step, \(z\) c is the input of the critic network, expressed as \(z\) c =\{x + iy,u R +iu I}, where \(O\) is the output of the first-layer evaluation network, and \(O\) can be divided into a real part and an imaginary part:

[0161] \(O(t)=O\) R (t)+iO I (t) (35)

[0162] Then, using the traditional method of "splitting" complex values, the activation function processes the real and imaginary parts of \(O_1\) separately, and the output \(V\) can be defined as:

[0163]

[0164] where represents the weight matrix of the second-layer evaluation network, and \(f\) c (·) represents the activation function of the evaluation network, and is selected as the activation function. Through the continuous function (7), its discretization is carried out, and the evaluation network error can be regarded as:

[0165]

[0166] The minimum error function is defined as:

[0167]

[0168] To minimize the error function, the following method is used to update the weights of the evaluation neural network:

[0169]

[0170]

[0171]

[0172]

[0173] where \(\mu\) is the learning rate.

[0174] Complex action network design

[0175] In this part, a three-layer action network is designed for approximate control as follows:

[0176]

[0177] where is the weight matrix of the first-layer action network, and \(z\) a represents the input of the action network, and the expression is \(z\) a =\(\{x + iy\}\),

[0178] \(O_1\) is the output of the first-layer action network, and \(O_1\) is separated into a real part and an imaginary part:

[0179] O1(t) = O R1 (t) + iO I1 (t) (44)

[0180]

[0181] where f a (·) represents the activation function of the action network, and the activation function is selected Then O2 is separated into the sum of the following real and imaginary parts:

[0182] O2(t) = O R2 (t) + iO I2 (t) (46)

[0183] Once again, the output is obtained using the activation function:

[0184]

[0185] The action network error is regarded as:

[0186]

[0187] where represents the expected value, which is set to 0. The minimum error function of the action network is defined as:

[0188]

[0189] To minimize the error function, the following method is used to update the weights of the action network:

[0190]

[0191]

[0192]

[0193]

[0194] To verify the effectiveness of the model, based on Figure 1 a simulation model is built, and the required parameters are as follows:

[0195] Considering Figure 1 the nonlinear system signal processing model shown, for the complex learning system representing the signal system, its dynamics are given by:

[0196] f1(z(t)) = z(t) + (1 + 0.1i)z 2 (t) + (0.3 + 0.01)z 3 (t)

[0197] A complex unknown system represents the nonlinear model of a signal estimation system, which is represented as follows:

[0198] s(t) = (-0.5 + 0.3i)z(t) + (0.8 - 0.5i)z 2 (t - 1) + (0.2 - 0.1i)z 2 (t - 2)

[0199] f2(z(t)) = s(t) + (0.2 + 0.1i)s 2 (t) + (0.02 + 0.01i)s 3 (t) + v

[0200] where v is additive complex generalized Gaussian distributed noise and z represents a complex-valued signal.

[0201] In this example, the output signal of the nonlinear learning system is used to approximate the output of the nonlinear unknown system to achieve signal tracking. Then the error dynamics of the two systems are represented by system (1):

[0202]

[0203] where f(z(t)) = f2(z(t)) - f1(z(t)), and C is selected to be 2.

[0204] In addition, in (7), the cost function is reprocessed as where both Q and R are set to 1.

[0205] From Figure 2 and Figure 3 it can be seen that a complex error composed of real and imaginary parts will eventually stabilize to 0, which means that the output signal of the complex learning system estimates the output of the unknown system.

[0206] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claimed rights. Additionally, it should be noted that in the above specific implementation manner, the various specific technical features described can be combined in any appropriate manner without contradiction. To avoid unnecessary repetition, the present invention does not separately describe various possible combination manners. Furthermore, any arbitrary combination can be made between different embodiments of the present invention as long as it does not violate the idea of the present invention, and it should equally be regarded as the content disclosed by the present invention.

Claims

1. An optimal control method for a complex-valued continuous-time nonlinear system, characterized in that, Including the following steps: Step 1, splitting the nonlinear system control model into a real part model and an imaginary part model; Step 2, establishing cost functions for the real part and the imaginary part to construct a cost function for the complex system; Step 3, completing the policy iteration optimal control algorithm through the cost function and constructing a three-layer complex forward neural network to achieve optimal control of the complex nonlinear system; The splitting of the nonlinear system control model into a real part model and an imaginary part is specifically as follows: z(t) = x(t) + iy(t), (2) where x(t) and y(t) are the real part and the imaginary part of z(t) respectively; where f(z(t)) is separated into: f(z(t)) = f R (x(t), y(t)) + if I (x(t), y(t)); (3) Establishing cost functions for the real part and the imaginary part to construct a cost function for the complex system, and completing the policy iteration optimal control algorithm through the cost function is specifically as follows: where both Q and R are positive definite matrices with corresponding dimensions; Establishing the following optimal control through the Hamilton-Jacobi-Bellman equation: Completing the policy iteration optimal control algorithm through the cost function, specifically as follows: Updating the control pair according to (11) and (12): Then return to step 1, and stop when convergence reaches the optimum.