Spacecraft non-cooperative game method based on model predictive control deep neural network

By adopting a deep neural network based on model prediction control and a rolling time-domain composite interference estimator in non-cooperative spacecraft games, the impact of multi-source interference on spacecraft control is solved, and the rapid solution and effective control of non-cooperative games of multi-spacecraft are achieved.

CN120029067AActive Publication Date: 2025-05-23TIANJIN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510172978.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-23
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively resist multi-source interference and obtain a game balance solution in non-cooperative games of multi-spacecraft, resulting in unpredictable dynamic behavior.

Method used

A non-cooperative game method of spacecraft based on model prediction control deep neural network is designed to optimize the estimation error of multi-source interference through a rolling time domain composite interference estimator, and use deep neural networks to generate control inputs online to realize the control of multi-spacecraft.

Benefits of technology

This method can quickly solve the non-cooperative game problem of multi-spacecraft, effectively suppress the impact of multi-source interference on spacecraft control, and improve real-time performance and computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029067A_ABST
    Figure CN120029067A_ABST
Patent Text Reader

Abstract

A spacecraft non-cooperative game method based on a model predictive control deep neural network belongs to the technical field of spacecraft control, and comprises the following steps: establishing a multi-spacecraft relative motion model; designing a composite interference estimator with decision variables; obtaining a non-cooperative game equilibrium solution under an offline condition; obtaining a non-cooperative game equilibrium solution under an online condition; according to the invention, through designing the composite interference estimator with the decision variable, accurate estimation of multi-source interference can be realized; a multi-spacecraft non-cooperative game problem is solved offline through model prediction control, and a large amount of data of solutions of the non-cooperative game problem is obtained; a deep neural network is designed, and the neural network is trained by using the obtained data of a large amount of solutions of the non-cooperative game problem, so that the neural network can be applied to the solution of the multi-spacecraft non-cooperative game problem online; according to the method, rapid solving of the multi-spacecraft non-cooperative game problem is achieved, and meanwhile the influence of multi-source interference on multi-spacecraft control is greatly restrained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of spacecraft control technology, and in particular to a spacecraft non-cooperative game method based on a model predictive control deep neural network. Background Art

[0002] In recent years, solving non-cooperative games of spacecraft swarms has attracted increasing attention. In the classical study of non-cooperative games of spacecraft, it usually means that all participants have access to global information, which is almost impossible in most practical situations. Therefore, although the cooperation of multiple spacecraft under a given topology has been widely studied, each controller can only access local / neighborhood information.

[0003] In practice, the orbital motion of spacecraft is inevitably affected by internal disturbances (such as elastic disturbances and actuator noise introduced by antennas and solar flip panels) and external environmental disturbances (such as the non-spherical gravity of the earth, solar radiation pressure and geomagnetic force). Insufficient consideration of these factors may have unpredictable effects on the dynamic behavior of spacecraft groups. In this regard, a new rolling time domain composite disturbance estimator is proposed in this paper. By optimizing the estimation error in a rolling time domain manner, the best disturbance estimation can be obtained.

[0004] In addition to internal and external disturbances, the orbital motion of a spacecraft swarm is usually subject to dynamic constraints (including state and control input constraints). It is well known that model predictive control has advantages in dealing with these inherent dynamic constraints. However, in order to obtain the equilibrium solution of the non-cooperative game, the optimization problem in model predictive control needs to be solved iteratively at each update time, which makes the real-time performance of this method cannot be guaranteed. In order to improve the computational efficiency, the present invention designs a deep neural network based on model predictive control to solve the non-cooperative game problem. Summary of the invention

[0005] The purpose of the present invention is to provide a spacecraft non-cooperative game method based on a model predictive control deep neural network, so as to solve the problem that the prior art is difficult to obtain multi-source interference resistance and game equilibrium solutions in multi-spacecraft non-cooperative games; the present invention utilizes the characteristics of the rolling time domain to optimize the multi-source interference estimation error online, so as to better suppress multi-source interference, and utilizes a deep neural network based on model predictive control to solve the multi-spacecraft non-cooperative game problem, obtain control input, and finally control multiple spacecraft.

[0006] In order to achieve the above object, the present invention adopts the following technical scheme:

[0007] The spacecraft non-cooperative game method based on model predictive control deep neural network is characterized by comprising the following steps:

[0008] Step 1, establish a multi-spacecraft relative motion model: use the linearized CW equation with known system parameters to establish an ideal model of multi-spacecraft relative motion; then characterize the multi-source interference, and add the multi-source interference to the multi-spacecraft relative motion ideal model to obtain the multi-spacecraft relative motion model affected by the multi-source interference; finally, characterize the state constraints and control input constraints of the multi-spacecraft system;

[0009] Step 2, designing a composite interference estimator with decision variables, wherein the decision variables in the composite interference estimator are obtained by means of a rolling time domain;

[0010] Step 3, obtaining the equilibrium solution of the non-cooperative game in the offline state: designing a model predictive controller; designing an iterative optimization problem and solving it to obtain the equilibrium solution of the non-cooperative game in the offline state;

[0011] Step 4, obtain the non-cooperative game equilibrium solution in the online case: design a controller based on a deep neural network, and then use the multi-spacecraft non-cooperative game equilibrium solution in the offline case as the data in the training set to train the neural network variables in the designed deep neural network-based controller, and finally generate the multi-spacecraft non-cooperative game equilibrium solution online by the deep neural network to control multiple spacecraft.

[0012] Furthermore, the step 1 specifically includes the following steps:

[0013] Step 1.1, establish an ideal relative motion model of multiple spacecraft:

[0014] x i (k+1)=A i x i (k)+B i u i (k)

[0015] Among them, x i (k) is the state of the i-th spacecraft, u i (k) is the control input of the i-th spacecraft, A i and B i is a known system matrix;

[0016] Step 1.2, characterize the multi-source interference to the multi-spacecraft system:

[0017] Step 1.2.1: For the unmodelable disturbance d to the multi-spacecraft system i,f (k) is characterized by a bounded norm d i,f (k+1)-d i,f (k) = δ i,f (k) where δ i,f(k) is the interference d that cannot be modeled at the previous and next moments i,f (k) error; is a known positive constant;

[0018] Step 1.2.2: Modelable disturbance d to the multi-spacecraft system i,o (k) is characterized, and its model is:

[0019] w i (k+1)=W i w i (k), d i,o (k) = V i w i (k)

[0020] Among them, w i (k) is the state variable of the modelable interference of the i-th spacecraft; b i,o (k) is the modelable interference; W i and V i is a known matrix;

[0021] Step 1.3, combine the ideal relative motion equation of multiple spacecraft with the multi-source interference to obtain the relative motion equation of multiple spacecraft affected by the multi-source interference:

[0022] x i (k+1)=A i x i (k)+B i u i (k)+B i (d i,f (k)+d i,o (k))

[0023] Among them, spacecraft i is subject to state constraints and control quantity constraints, which can be expressed as:

[0024]

[0025] Among them, b i,x , b i,u ,h i,x and h i,u is a known matrix.

[0026] Furthermore, in step 2, the composite interference estimator with decision variables is established as follows:

[0027]

[0028] in, is d i,f (k) estimated value; is d i,o(k) estimated value; w i (k) is the estimated value; v i,f (k) and v i,w (k) is an intermediate variable; b i,f (k) and b) i,w (k) is the decision variable; L i,f and L i,w is the estimator gain.

[0029] Furthermore, the step 3 specifically includes the following steps:

[0030] Step 3.1, design the model predictive controller:

[0031]

[0032] Among them, u i (k) is the virtual control quantity; c i (k) is the decision variable; K i is the feedback gain.

[0033] Step 3.2, design an iterative optimization problem and solve it to obtain the equilibrium solution of the non-cooperative game in the offline case:

[0034] Step 3.2.1, the iterative optimization problem is:

[0035]

[0036] Constraints

[0037]

[0038] in,

[0039]

[0040] in, and are η i (s|k) and c i,η The predicted value of (s|k), and For state x i Nominal value of (s|k); is the error e i,of The nominal value of (s|k-1); and (s|k) represents the prediction of time s at time k; I is the identity matrix; is the state prediction of the j-th spacecraft; Q i , R i , Q ij and P i is the given weight matrix; N p For the prediction time domain; and The constraints imposed on the optimization problem;

[0041] Step 3.2.2, iteratively solve the optimization problem designed in step 3.2.1.

[0042] Furthermore, the step 4 specifically includes the following steps:

[0043] Step 4.1, design a controller based on deep neural network:

[0044]

[0045] in, is the output of the deep neural network; the output of the nth neuron in the mth layer of the deep neural network is Function f(a)=max{0,a} is the activation function. is the weight pair of the nth neuron in the mth layer

[0046] is the weight value; the output layer uses is the number of layers of its neural network, is the number of its neurons, is the number of neurons in the mth layer; design and adjust the weights The loss function is:

[0047]

[0048] in, is the hth sample output of the designed deep neural network, o i,h (k) is the true value of the hth sample, N p For batches of data; implementation where ρ i >0 is the precision parameter;

[0049] Step 4.2, training the neural network variables in the designed deep neural network-based controller:

[0050] Step 4.2.1, use the state data in the deep neural network-based controller as the input training set of the deep neural network:

[0051]

[0052] in, is the adjacency matrix parameter; is the state x at time k i (k) predicted sequence;

[0053] Step 4.2.2, use the decision variable data in the deep neural network-based controller as the output training set of the deep neural network:

[0054]

[0055] in, Yes middle Approximation of is the sequence solved; Z[0,N p -1] represents from 0 to N p The set of positive integers starting from -1;

[0056] In step 4.3, the output of the deep neural network is used online to obtain the equilibrium solution of the non-cooperative game in the online case and use it as the control input of the multi-spacecraft non-cooperative game.

[0057] Compared with the prior art, the present invention has the following beneficial technical effects:

[0058] The present invention designs a composite interference estimator with decision variables to achieve accurate estimation of multi-source interference; in addition, by utilizing model predictive control to solve multi-spacecraft non-cooperative game problems offline, a large amount of data on solutions to non-cooperative game problems is obtained; finally, by designing a deep neural network and using the obtained large amount of data on solutions to non-cooperative game problems to train the neural network, it can be applied online to solving multi-spacecraft non-cooperative game problems; this method can achieve rapid solution to multi-spacecraft non-cooperative game problems and at the same time well suppress the influence of multi-source interference on multi-spacecraft control. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 A flowchart of the spacecraft non-cooperative game method based on the model predictive control deep neural network provided by the present invention.

[0060] Figure 2-Figure 5 It is a state curve diagram of four spacecraft under coordinated control in an embodiment of the present invention. DETAILED DESCRIPTION

[0061] The present invention is further described in detail below in conjunction with the accompanying drawings:

[0062] See also Figure 1 The present invention provides a non-cooperative game method for spacecraft based on a model predictive control deep neural network, comprising the following steps:

[0063] Step 1, establish a multi-spacecraft relative motion model: use the linearized CW equation with known system parameters to establish an ideal model of multi-spacecraft relative motion; then characterize the multi-source interference, and add the multi-source interference to the multi-spacecraft relative motion ideal model to obtain the multi-spacecraft relative motion model affected by the multi-source interference; finally, characterize the state constraints and control input constraints of the multi-spacecraft system;

[0064] Step 2, designing a composite interference estimator with decision variables, wherein the decision variables in the composite interference estimator are obtained by means of a rolling time domain;

[0065] Step 3, obtaining the equilibrium solution of the non-cooperative game in the offline state: designing a model predictive controller; designing an iterative optimization problem and solving it to obtain the equilibrium solution of the non-cooperative game in the offline state;

[0066] Step 4, obtain the non-cooperative game equilibrium solution in the online case: design a controller based on a deep neural network, and then use the multi-spacecraft non-cooperative game equilibrium solution in the offline case as the data in the training set to train the neural network variables in the designed deep neural network-based controller, and finally generate the multi-spacecraft non-cooperative game equilibrium solution online by the deep neural network to control multiple spacecraft.

[0067] As a specific implementation of this embodiment, step 1 specifically includes the following steps:

[0068] Step 1.1: Establish an ideal relative motion model of multiple spacecraft, as shown in formula (1):

[0069] x i (k+1)=A i x i (k)+B i u i (k) (1)

[0070] Among them, x i (k) is the state of the i-th spacecraft, u i (k) is the control input of the i-th spacecraft, A i and B i is a known system matrix;

[0071] Step 1.2, characterize the multi-source interference to the multi-spacecraft system:

[0072] Step 1.2.1: For the unmodelable disturbance d to the multi-spacecraft system i,f (k) is characterized by a bounded norm d i,f (k+1)-d i,f (k) = δ i,f (k) where δ i,f (k) is the interference d that cannot be modeled at the previous and next moments i,f (k) error; is a known positive constant;

[0073] Step 1.2.2: Modelable disturbance d to the multi-spacecraft system i,o (k) is characterized, and its model is:

[0074] w i (k+1)=W i w i (k), d i,o (k) = V i w i (k)

[0075] Among them, w i (k) is the state variable of the modelable interference of the i-th spacecraft; d i,o (k) is the modelable interference; W i and V i is a known matrix;

[0076] Step 1.3, combine the ideal relative motion equation of multiple spacecraft with the multi-source interference to obtain the relative motion equation of multiple spacecraft affected by the multi-source interference:

[0077] x i (k+1)=A i x i (k)+B i u i (k)+B i (d i,f (k)+d i,o (k)) (2)

[0078] Among them, spacecraft i is subject to state constraints and control quantity constraints, which can be expressed as:

[0079]

[0080] Among them, b i,x , b i,u ,h i,x and h i,u is a known matrix.

[0081] As a specific implementation of this embodiment, in step 2, a composite interference estimator with decision variables is established, formulas (5)-(6):

[0082]

[0083] in, is di,f (k) estimated value; is d i,o (k) estimated value; w i The estimated value of (k); v i,f (k) and v i,w (k) is an intermediate variable; b i,f (k) and b) i,w (k) is the decision variable; L i,f and L i,w is the observer gain.

[0084] As a specific implementation of this embodiment, step 3 specifically includes the following steps:

[0085] Step 3.1, design model predictive controller:

[0086]

[0087] Among them, u i (k) is the virtual control quantity; c i (k) is the decision variable; K i is the feedback gain;

[0088] Step 3.2, design an iterative optimization problem and solve it to obtain the equilibrium solution of the non-cooperative game in the offline case:

[0089] Step 3.2.1, the iterative optimization problem is:

[0090]

[0091] Constraints

[0092]

[0093] in,

[0094]

[0095] in, and is η i (s|k) and c i,η The predicted value of (s|k), and For state x i Nominal value of (s|k); is the error e i,of The nominal value of (s|k-1); and (s|k) represents the prediction of time s at time k; I is the identity matrix; is the state prediction of the j-th spacecraft; Q i , R i , Q ij and P i is the given weight matrix; N p For the prediction time domain; and The constraints imposed on the optimization problem;

[0096] Step 3.2.2, iteratively solve the optimization problem designed in step 3.2.1.

[0097] As a specific implementation of this embodiment, step 4 specifically includes the following steps:

[0098] Step 4.1, design a controller based on deep neural network:

[0099]

[0100] in, is the output of the deep neural network; the neural network structure designed in the present invention is a multi-layer neural network, that is, each layer is composed of multiple neurons; for the i-th spacecraft, let is the number of layers of its neural network, is the number of its neurons, is the number of neurons in the mth layer; once the number of layers and neurons in the neural network is determined, the deep neural network is determined; furthermore, the output of the nth neuron in the mth layer (input layer and hidden layer) of the deep neural network is Among them, the function f(a)=max{0,a} is the activation function, is the weight pair of the nth neuron in the mth layer for is the weight value, which will be adjusted during the neural network training process until the deep neural network can best approximate the functional relationship between the input value and the target output value; the output layer uses a linear output, that is, To adjust the weight Usually, a loss function is designed for a deep neural network. The parameter of the loss function is the mean of the squared difference between the approximate output value of the deep neural network and the target output value; the loss function is:

[0101]

[0102] in, is the hth sample output of the designed multi-layer neural network, o i,h(k) is the true value of the hth sample, N p is the batch of data. By minimizing the value of the loss function, the output value of the deep neural network is close to the target output value, and finally a certain approximation accuracy is achieved. It is believed that the deep neural network can achieve the approximation of the target mapping, that is, to achieve where ρ i >0 is the precision parameter;

[0103] Step 4.2, training the neural network variables in the designed deep neural network-based controller:

[0104] Step 4.2.1, use the state data in the deep neural network-based controller as the input training set of the deep neural network:

[0105]

[0106] The output training set of the deep neural network is:

[0107]

[0108] in, is the adjacency matrix parameter; is the state x at time k i (k) predicted sequence; Yes middle Approximation of is the sequence solved in formula (8); Z[0,N p -1] represents from 0 to N p -1; Deep neural networks can be trained by using formulas (10) and (11);

[0109] In step 4.3, the output of the deep neural network is used online to obtain the equilibrium solution of the non-cooperative game in the online case and use it as the control input of the multi-spacecraft non-cooperative game.

[0110] The present invention will be described below in conjunction with specific embodiments:

[0111] This embodiment uses the above-mentioned spacecraft non-cooperative game method based on model predictive control deep neural network to control multiple spacecraft. It should be noted that in formula (8) of this embodiment, Q i =1, R i =1,Q ij =1,N p =5 is the weight matrix used in the cost function.

[0112] In this embodiment, see Figure 2-Figure 5 , the state curves of four spacecraft are given, among which, (i=1, ..., 4; j=1, ..., 6) represents the jth state of the i-th spacecraft; Figure 2 The six states representing spacecraft 1 all converge to 0 under the designed controller. Figure 3 The six states representing spacecraft 2 all converge to 0 under the designed controller. Figure 4 The six states representing spacecraft 3 all converge to 0 under the designed controller. Figure 5 The final states of the six states representing spacecraft 3 all converge to 0 under the designed controller; in summary, Figure 2-Figure 5 It can be seen that the states of the four spacecraft finally converged to 0.

Claims

1. A non-cooperative game method for spacecraft based on model predictive control deep neural network, characterized in that: The following steps are involved: Step 1, establish a multi-spacecraft relative motion model: use the linearized CW equation with known system parameters to establish an ideal model of multi-spacecraft relative motion; then characterize the multi-source interference, and add the multi-source interference to the multi-spacecraft relative motion ideal model to obtain the multi-spacecraft relative motion model affected by the multi-source interference; finally, characterize the state constraints and control input constraints of the multi-spacecraft system; Step 2, designing a composite interference estimator with decision variables, wherein the decision variables in the composite interference estimator are obtained by means of a rolling time domain; Step 3, obtaining the equilibrium solution of the non-cooperative game in the offline state: designing a model predictive controller; designing an iterative optimization problem and solving it to obtain the equilibrium solution of the non-cooperative game in the offline state; Step 4, obtain the non-cooperative game equilibrium solution in the online case: design a controller based on a deep neural network, and then use the multi-spacecraft non-cooperative game equilibrium solution in the offline case as the data in the training set to train the neural network variables in the designed deep neural network-based controller, and finally generate the multi-spacecraft non-cooperative game equilibrium solution online by the deep neural network to control multiple spacecraft.

2. The spacecraft non-cooperative game method based on model predictive control deep neural network according to claim 1 is characterized in that: The step 1 specifically includes the following steps: Step 1.1, establish an ideal relative motion model of multiple spacecraft: x i (k+1)=A i x i (k)+B i u i (k) Among them, x i (k) is the state of the i-th spacecraft, u i (k) is the control input of the i-th spacecraft, A i and B i is a known system matrix; Step 1.2, characterize the multi-source interference to the multi-spacecraft system: Step 1.2.1: For the unmodelable disturbance d to the multi-spacecraft system i,f (k) is characterized by a bounded norm d i,f (k+1)-d i,f (k) = δ i,f (k) where δ i,f (k) is the interference d that cannot be modeled at the previous and next moments i,f (k) error; is a known positive constant; Step 1.2.2: Modelable disturbance d to the multi-spacecraft system i,o (k) is characterized, and its model is: w i (k+1)=W i w i (k),d i,o (k)=V i w i (k) Among them, w i (k) is the state variable of the modelable interference of the i-th spacecraft; d i,o (k) is the modelable interference; W i and V i is a known matrix; Step 1.3, combine the ideal relative motion equation of multiple spacecraft with the multi-source interference to obtain the relative motion equation of multiple spacecraft affected by the multi-source interference: x i (k+1)=A i x i (k)+B i u i (k)+B i (d i,f (k)+d i,o (k)) Among them, spacecraft i is subject to state constraints and control quantity constraints, which can be expressed as: Among them, b i,x , b i,u ,h i,x and h i,u is a known matrix.

3. The spacecraft non-cooperative game method based on model predictive control deep neural network according to claim 2 is characterized in that: In step 2, the composite interference estimator with decision variables is established as follows: in, is d i,f (k) estimated value; is d i,o (k) estimated value; w i (k) is the estimated value; v i,f (k) and v i,w (k) is an intermediate variable; b i,f (k) and b i,w (k) is the decision variable; L i,f and L i,w is the estimator gain.

4. The spacecraft non-cooperative game method based on model predictive control deep neural network according to claim 3 is characterized in that: The step 3 specifically includes the following steps: Step 3.1, design model predictive controller: Among them, u i (k) is the virtual control quantity; c i (k) is the decision variable; K i is the feedback gain; Step 3.2, design an iterative optimization problem and solve it to obtain the equilibrium solution of the non-cooperative game in the offline case: Step 3.2.1, the iterative optimization problem is: Constraints in, in, and are η i (s|k) and c i,η The predicted value of (s|k), and For state x i Nominal value of (s|k); is the error e i,of The nominal value of (s|k-1); and s|k) represents the prediction of time s at time k; I is the identity matrix; is the state prediction of the j-th spacecraft; Q i , R i , Q ij and P i is the given weight matrix; N p For the prediction time domain; and The constraints imposed on the optimization problem; Step 3.2.2, iteratively solve the optimization problem designed in step 3.2.

1.

5. The spacecraft non-cooperative game method based on model predictive control deep neural network according to claim 4 is characterized in that: The step 4 specifically includes the following steps: Step 4.1, design a controller based on deep neural network: in, is the output of the deep neural network; the output of the nth neuron in the mth layer of the deep neural network is Function f(a)=max{0,a} is the activation function. is the weight pair of the nth neuron in the mth layer is the weight value; the output layer uses is the number of layers of its neural network, is the number of its neurons, is the number of neurons in the mth layer; design and adjust the weights The loss function is: in, is the hth sample output of the designed deep neural network, o i,h (k) is the true value of the hth sample, N p For batches of data; implementation where ρ i >0 is the precision parameter; Step 4.2, training the neural network variables in the designed deep neural network-based controller: Step 4.2.1, use the state data in the deep neural network-based controller as the input training set of the deep neural network: in, is the adjacency matrix parameter; is the state x at time k i (k) predicted sequence; Step 4.2.2, use the decision variable data in the deep neural network-based controller as the output training set of the deep neural network: in, Yes middle Approximation of is the sequence solved; Z[0,N p -1] represents from 0 to N p The set of positive integers starting from -1; In step 4.3, the output of the deep neural network is used online to obtain the equilibrium solution of the non-cooperative game in the online case and use it as the control input of the multi-spacecraft non-cooperative game.

Citation Information

Patent Citations

  • Spacecraft formation discrete distributed non-cooperative game method based on dynamic event triggering

    CN112558471A

  • Four-rotor unmanned aerial vehicle optimal formation control method based on non-cooperative differential game

    CN117193359A

  • Method for controlling relative attitude of spacecrafts having multi-source disturbances and actuator saturation

    GB201910669D0