Spacecraft non-cooperative game method based on model predictive control deep neural network

By using a deep neural network based on model predictive control, the problems of multi-source interference resistance and game equilibrium solutions in non-cooperative game among multiple spacecraft were solved, enabling rapid and precise control of multiple spacecraft.

CN120029067BActive Publication Date: 2025-11-25TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510172978.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-11-25
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively resist multi-source interference and obtain game equilibrium solutions in non-cooperative multi-spacecraft games, and the real-time performance of model predictive control is insufficient.

Method used

A deep neural network based on model predictive control is used to optimize multi-source disturbances online through a rolling time-domain composite disturbance estimator. Combined with offline training of the deep neural network to generate control inputs online, a fast solution for non-cooperative game among multiple spacecraft is achieved.

Benefits of technology

It achieves accurate estimation and rapid solution of multi-source interference, suppresses the impact of multi-source interference on spacecraft control, and improves computational efficiency and control accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029067B_ABST
    Figure CN120029067B_ABST
Patent Text Reader

Abstract

The spacecraft non-cooperative game method based on model predictive control deep neural network belongs to the technical field of spacecraft control and comprises the following steps: a relative motion model of multiple spacecrafts is established; a compound disturbance estimator with decision variables is designed; a non-cooperative game equilibrium solution under an offline condition is obtained; and a non-cooperative game equilibrium solution under an online condition is obtained.The compound disturbance estimator with decision variables can realize accurate estimation of multi-source disturbance; the model predictive control is used to solve the non-cooperative game problem of multiple spacecrafts offline, and a large amount of data of solutions of the non-cooperative game problem is obtained; the deep neural network is designed, and the neural network is trained by using the large amount of data of solutions of the non-cooperative game problem, so that the neural network can be applied to the solution of the non-cooperative game problem of multiple spacecrafts online; the method realizes rapid solution of the non-cooperative game problem of multiple spacecrafts, and greatly suppresses the influence of multi-source disturbance on the control of multiple spacecrafts.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of spacecraft control, and particularly relates to a spacecraft non-cooperative game method based on a model predictive control deep neural network. BACKGROUND

[0002] In recent years, the solution of spacecraft swarm non-cooperative game has been increasingly concerned by people. In the classical research of spacecraft non-cooperative game, it is usually meant that all participants can obtain global information, which is almost impossible in most actual situations. Therefore, although extensive research has been conducted on the coordination of multiple spacecrafts under a given topology, each controller in the research can only access local / neighbor information.

[0003] In practice, the orbital motion of a spacecraft is inevitably affected by internal disturbances (such as elastic disturbances introduced by antennas and solar energy turning plates and actuator noise) and external environmental disturbances (such as the non-spherical gravity of the earth, solar radiation pressure and geomagnetic force). Insufficient consideration of these factors can have unpredictable effects on the dynamic behavior of the spacecraft swarm, and the application proposes a new rolling horizon compound disturbance estimator, which can obtain the best disturbance estimation by optimizing the estimation error in a rolling horizon manner.

[0004] In addition to internal and external disturbances, the orbital motion of a spacecraft swarm is usually limited by dynamic constraints (including state and control input constraints). It is well known that model predictive control has an advantage in handling these inherent dynamic constraints. However, in order to obtain the equilibrium solution of the non-cooperative game, the optimization problem in model predictive control needs to be solved iteratively at each update time, which makes it impossible to guarantee the real-time performance of the method. In order to improve the computational efficiency, the application designs a deep neural network based on model predictive control to solve the non-cooperative game problem. SUMMARY

[0005] The application aims to provide a spacecraft non-cooperative game method based on a model predictive control deep neural network to solve the problem of multi-source disturbance resistance and difficulty in obtaining a game equilibrium solution in the prior art in a multi-spacecraft non-cooperative game. The application optimizes the estimation error of multi-source disturbances online using the characteristics of rolling horizon, thereby better suppressing multi-source disturbances, and uses a deep neural network based on model predictive control to solve the multi-spacecraft non-cooperative game problem, obtains control input, and finally controls the multi-spacecraft.

[0006] To achieve the above-mentioned purpose, the application adopts the following technical solutions:

[0007] The spacecraft non-cooperative game method based on a model predictive control deep neural network comprises the following steps:

[0008] Step 1, establishing a multi-spacecraft relative motion model: a linearized C-W equation with known system parameters is used to establish an ideal model of multi-spacecraft relative motion; then, multi-source disturbances are characterized, and the multi-source disturbances are added to the ideal model of multi-spacecraft relative motion to obtain a multi-spacecraft relative motion model affected by multi-source disturbances; finally, the state constraints and control input constraints of the multi-spacecraft system are characterized;

[0009] Step 2, designing a compound disturbance estimator with decision variables, the decision variables in the compound disturbance estimator being obtained by means of a rolling horizon;

[0010] Step 3, obtaining a non-cooperative game equilibrium solution in an offline case: a model predictive controller is designed; an iterative optimization problem is designed and solved to obtain a non-cooperative game equilibrium solution in an offline case;

[0011] Step 4, obtaining a non-cooperative game equilibrium solution in an online case: a deep neural network-based controller is designed, then the multi-spacecraft non-cooperative game equilibrium solution in an offline case is used as data in a training set to train neural network variables in the deep neural network-based controller, and finally the deep neural network generates a multi-spacecraft non-cooperative game equilibrium solution online to control the multi-spacecraft.

[0012] Further, the step 1 specifically comprises the following steps:

[0013] Step 1.1, establishing an ideal relative motion model of multi-spacecraft:

[0014] x i (k+1)=A i x i (k)+B i u i (k)

[0015] wherein x i (k) is the state of the i-th spacecraft, u i (k) is the control input of the i-th spacecraft, A i and B i are known system matrices;

[0016] Step 1.2, characterizing multi-source disturbances to the multi-spacecraft system:

[0017] Step 1.2.1, characterizing unmodelable disturbances d i,f (k) to the multi-spacecraft system, which is characterized by a bounded norm d i,f (k+1)-d i,f (k)=δ i,f (k), wherein δ i,f(k) represents the unmodelable disturbance d between the two consecutive time points. i,f The error of (k); For known positive constants;

[0018] Step 1.2.2, modelable disturbances d experienced by the multi-spacecraft system i,o (k) is characterized by the following model:

[0019] w i (k+1)=W i w i (k), d i,o (k)=V i w i (k)

[0020] Among them, w i (k) represents the modelable disturbance state variables of the i-th spacecraft; b i,o (k) represents a modelable disturbance; W i and V i The matrix is ​​known.

[0021] Step 1.3: Combine the ideal relative motion equations of multiple spacecraft with multi-source interference to obtain the relative motion equations of multiple spacecraft affected by multi-source interference:

[0022] x i (k+1)=A i x i (k)+B i u i (k)+B i (d i,f (k)+d i,o (k))

[0023] Wherein, spacecraft i is subject to state constraints and control constraints, expressed as:

[0024]

[0025] Among them, b i,x b i,u h i,x and h i,u The matrix is ​​known.

[0026] Furthermore, in step 2, the establishment of the composite disturbance estimator with decision variables is as follows:

[0027]

[0028] in, For d i,f The estimated value of (k); For d i,o(k) is an estimate of the value of is w i (k) is an estimate of the value of v i,f (k) and v i,w (k) is an intermediate variable; b i,f (k) and b i,w (k) is a decision variable; L i,f and L i,w is an estimator gain.

[0029] Further, the step 3 specifically comprises the following steps:

[0030] Step 3.1, design a model predictive controller:

[0031]

[0032] wherein u i (k) is a virtual control variable; c i (k) is a decision variable; K i is a feedback gain.

[0033] Step 3.2, design an iterative optimization problem and solve to obtain a non-cooperative game equilibrium in the offline case:

[0034] Step 3.2.1, the iterative optimization problem is:

[0035]

[0036] The constraint is

[0037]

[0038] wherein,

[0039]

[0040] wherein, and are the predicted values of η i (s|k) and c i,η (s|k), respectively, and is the nominal value of the state x i (s|k); is the nominal value of the error e i,of (s|k-1); and (s|k) represents the prediction of s at time k at time k; I is an identity matrix; is the state prediction of the jth spacecraft; Q i , R i , Q ij and P i are given weight matrices; N p is the prediction horizon; and are constraints on the optimization problem;

[0041] Step 3.2.2, iteratively solve the optimization problem designed in step 3.2.1.

[0042] Further, the step 4 specifically comprises the following steps:

[0043] Step 4.1, design a controller based on deep neural network:

[0044]

[0045] wherein, is the output of the deep neural network; the output of the nth neuron in the mth layer of the deep neural network is the activation function is f(a) = max{0, a}, is the weight pair of the nth neuron in the mth layer,

[0046] is the weight value; the output layer uses is the number of layers of the neural network, is the number of neurons, is the number of neurons in the mth layer; the loss function of the designed adjustment weight is:

[0047]

[0048] wherein, is the output of the hth sample of the designed deep neural network, o i,h (k) is the true value of the hth sample, N p is the batch of data; the implementation wherein p i > 0 is the precision parameter;

[0049] Step 4.2, train the neural network variables in the designed deep neural network-based controller:

[0050] Step 4.2.1, take the state data in the deep neural network-based controller as the input training set of the deep neural network:

[0051]

[0052] wherein, is the adjacency matrix parameter; is the predicted sequence of state x i (k) at time k;

[0053] Step 4.2.2, the decision variable data in the deep neural network-based controller is trained as the output training set of the deep neural network:

[0054]

[0055] wherein, is the approximation of in ; is the solved sequence; Z[0,N p -1] represents a set of positive integers from 0 to N p -1;

[0056] Step 4.3, the non-cooperative game equilibrium solution in the online situation is obtained by using the output of the deep neural network online, and is used as the control input of the multi-spacecraft non-cooperative game.

[0057] Compared with the prior art, the present application has the following beneficial technical effects:

[0058] The present application can realize accurate estimation of multi-source interference by designing a composite interference estimator with decision variables; in addition, a large amount of data of solutions of non-cooperative game problems is obtained by using model predictive control to solve the multi-spacecraft non-cooperative game problem offline; finally, a deep neural network is designed, and the neural network is trained by using the large amount of data of solutions of non-cooperative game problems obtained, so that it can be applied online to the solution of the multi-spacecraft non-cooperative game problem; the method can realize fast solution of the multi-spacecraft non-cooperative game problem, and at the same time, the influence of multi-source interference on the control of the multi-spacecraft is well suppressed. BRIEF DESCRIPTION OF DRAWINGS

[0059] Figure 1 The flowchart of the spacecraft non-cooperative game method based on model predictive control deep neural network provided by the present application.

[0060] Figures 2-5 The state curve diagram under the cooperative control of the four spacecrafts in the embodiment of the present application. DETAILED DESCRIPTION

[0061] The present application will be further described in detail below in combination with the drawings:

[0062] Referring to Figure 1 , the present application provides a spacecraft non-cooperative game method based on model predictive control deep neural network, comprising the following steps:

[0063] Step 1, establishing a multi-spacecraft relative motion model: a linearized C-W equation with known system parameters is used to establish an ideal model of multi-spacecraft relative motion; then the multi-source disturbance is characterized, and the multi-source disturbance is added to the ideal model of multi-spacecraft relative motion to obtain a multi-spacecraft relative motion model affected by multi-source disturbance; finally, the multi-spacecraft system state constraint and control input constraint are characterized;

[0064] Step 2, designing a compound disturbance estimator with decision variables, the decision variables in the compound disturbance estimator being obtained by means of a rolling time domain;

[0065] Step 3, obtaining a non-cooperative game equilibrium solution in an offline case: designing a model predictive controller; designing an iterative optimization problem and solving to obtain a non-cooperative game equilibrium solution in an offline case;

[0066] Step 4, obtaining a non-cooperative game equilibrium solution in an online case: designing a deep neural network-based controller, then training neural network variables in the designed deep neural network-based controller with the multi-spacecraft non-cooperative game equilibrium solution in an offline case as data in a training set, and finally generating a multi-spacecraft non-cooperative game equilibrium solution online by the deep neural network to control the multi-spacecraft.

[0067] As a specific embodiment of the present embodiment, Step 1 specifically comprises the following steps:

[0068] Step 1.1, establishing an ideal relative motion model of multi-spacecraft, specifically as formula (1):

[0069] x i (k+1)=A i x i (k)+B i u i (k) (1)

[0070] Wherein, x i (k) is the state of the i-th spacecraft, u i (k) is the control input of the i-th spacecraft, A i and B i are known system matrices;

[0071] Step 1.2, characterizing the multi-source disturbance received by the multi-spacecraft system:

[0072] Step 1.2.1, characterizing the unmodeled disturbance d i,f (k) received by the multi-spacecraft system, which is characterized by a bounded norm d i,f (k+1)-d i,f (k)=δ i,f (k), where δ i,f (k) is the error of the unmodelable disturbance d i,f (k) at the two time instants; is a known constant;

[0073] Step 1.2.2, the modelable disturbance d i,o (k) received by the multi-spacecraft system is characterized, and its model is:

[0074] w i (k+1) = W i w i (k), d i,o (k) = V i w i (k)

[0075] where w i (k) is the state variable of the modelable disturbance of the ith spacecraft; d i,o (k) is the modelable disturbance; W i and V i are known matrices;

[0076] Step 1.3, the ideal relative motion equation of the multi-spacecraft is combined with the multi-source disturbance to obtain the relative motion equation of the multi-spacecraft affected by the multi-source disturbance:

[0077] x i (k+1) = A i x i (k) + B i u i (k) + B i (d i,f (k) + d i,o (k)) (2)

[0078] where the spacecraft i is limited by state constraints and control quantity constraints, which are expressed as:

[0079]

[0080] where b i,x , b i,u , h i,x and h i,u are known matrices.

[0081] As a specific embodiment of the present embodiment, in Step 2, a compound disturbance estimator with decision variables is established, and formulas (5)-(6) are:

[0082]

[0083] where, is di,f The estimated value of (k); For d i,o The estimated value of (k); For w i The estimated value of (k); v i,f (k) and v i,w (k) is an intermediate variable; b i,f (k) and b i,w (k) is the decision variable; L i,f and L i,w This is the observer gain.

[0084] As a specific implementation of this embodiment, step 3 specifically includes the following steps:

[0085] Step 3.1, Design the model predictive controller:

[0086]

[0087] Among them, u i (k) represents the virtual control variable; c i (k) is the decision variable; K i For feedback gain;

[0088] Step 3.2: Design an iterative optimization problem and solve it to obtain the offline non-cooperative game equilibrium solution:

[0089] Step 3.2.1, the iterative optimization problem is:

[0090]

[0091] Constraints

[0092]

[0093] in,

[0094]

[0095] in, and For η i (s|k) and c i,η The predicted value of (s|k), and For state x i The nominal value of (s|k); For error e i,of The nominal value of (s|k-1); and (s|k) represents the prediction of time s at time k; I is an identity matrix; is the state prediction of the jth spacecraft; Q i , R i , Q ij and P i are given weight matrices; N p is the prediction horizon; and are constraints on the optimization problem;

[0096] Step 3.2.2, iteratively solve the optimization problem designed in step 3.2.1.

[0097] As a specific implementation of the embodiment, step 4 specifically comprises the following steps:

[0098] Step 4.1, design a controller based on a deep neural network:

[0099]

[0100] wherein, is the output of the deep neural network; the neural network structure designed in the present application is a multi-layer neural network, i.e., each layer has multiple neurons; for the ith spacecraft, let be the number of layers of its neural network, be the number of neurons, be the number of neurons in the mth layer; once the number of layers and the number of neurons of the neural network are determined, the deep neural network is determined; further, the output of the nth neuron in the mth layer (input layer and hidden layer) in the deep neural network is wherein, the activation function f(a) = max{0, a}, be the weight pair of the nth neuron in the mth layer be is a weight value, which will be adjusted in the neural network training process until the deep neural network can best approximate the functional relationship between the input value and the target output value; the output layer uses a linear output, i.e., In order to adjust the weight A loss function is usually designed for the deep neural network. The parameter of the loss function is the mean square difference between the approximate output value of the deep neural network and the target output value; the loss function is:

[0101]

[0102] wherein, is the output of the hth sample of the designed multi-layer neural network, o i,h(k) is the true value of the hth sample, N p is the batch of data. By minimizing the value of the loss function, so that the output value of the deep neural network approximates the target output value, a certain approximation accuracy is finally achieved, and it is considered that the deep neural network can realize the approximation of the target mapping, that is, the deep neural network can realize the approximation of the target mapping. where ρ i > 0 is the accuracy parameter;

[0103] Step 4.2, training the neural network variable in the designed deep neural network-based controller:

[0104] Step 4.2.1, taking the state data in the deep neural network-based controller as the input training set of the deep neural network:

[0105]

[0106] The output training set of the deep neural network is:

[0107]

[0108] where, is the adjacency matrix parameter; is the predicted sequence of the state x i (k) at time k; is the approximation of ; is the sequence solved in formula (8); Z[0, N p -1] represents a set of positive integers from 0 to N p -1; the deep neural network can be trained by using formulas (10) and (11);

[0109] Step 4.3, using the output of the deep neural network online to obtain the non-cooperative game equilibrium solution under the online condition, and using it as the control input of the multi-spacecraft non-cooperative game.

[0110] The application will be described below in combination with specific embodiments:

[0111] In this embodiment, the spacecraft non-cooperative game method based on the model predictive control deep neural network is used to control the multi-spacecraft. It should be noted that in formula (8) of this embodiment, Q i = 1, R i = 1, Q ij = 1, N p = 5 are weight matrices used in the cost function.

[0112] In this embodiment, referring to Figures 2-5 , the state curves of four spacecraft are given, wherein,​ (i = 1,..., 4; j = 1,..., 6) represents the jth state of the ith spacecraft; Figure 2 The six states of spacecraft 1 converge to 0 under the designed controller, Figure 3 The six states of spacecraft 2 converge to 0 under the designed controller, Figure 4 The six states of spacecraft 3 converge to 0 under the designed controller, Figure 5 The six states of spacecraft 3 converge to 0 under the designed controller; and Figures 2-5 It can be seen that the states of the four spacecrafts all converge to 0.

Claims

1. A spacecraft non-cooperative game method based on model predictive control deep neural network, characterized in that, The method comprises the following steps: Step 1, establishing a multi-spacecraft relative motion model: a linearized C-W equation with known system parameters is used to establish an ideal multi-spacecraft relative motion model; then, multi-source disturbances are characterized, and the multi-source disturbances are added to the ideal multi-spacecraft relative motion model to obtain a multi-spacecraft relative motion model affected by the multi-source disturbances; finally, the multi-spacecraft system state constraints and control input constraints are characterized; Step 2, designing a compound disturbance estimator with decision variables, wherein the decision variables in the compound disturbance estimator are obtained through a rolling horizon method; The compound disturbance estimator with decision variables is designed as follows: , wherein is an estimate of is an estimate of is an estimate of and are intermediate variables; and are decision variables; and are estimator gains; Step 3, obtaining a non-cooperative game equilibrium solution in an offline case: a model predictive controller is designed; an iterative optimization problem is designed and solved to obtain a non-cooperative game equilibrium solution in the offline case; Step 4, obtaining a non-cooperative game equilibrium solution in an online case: a deep neural network-based controller is designed, then the multi-spacecraft non-cooperative game equilibrium solution in the offline case is used as data in a training set to train neural network variables in the deep neural network-based controller, and finally the deep neural network is used to generate the multi-spacecraft non-cooperative game equilibrium solution online to control the multi-spacecraft.

2. The model predictive control deep neural network based spacecraft non-cooperative game method of claim 1, wherein, The step 1 specifically comprises the following steps: Step 1.1, establishing an ideal multi-spacecraft relative motion model: , wherein, is the state of the th spacecraft, is the control input to the th spacecraft, and is a known system matrix; Step 1.2, characterizing multi-source disturbances suffered by the multi-spacecraft system: Step 1.2.

1. Unmodeled disturbances received by the multi-spacecraft system are characterized by being norm-bounded , where is the error of the unmodeled disturbances at the previous and current time instants ; and is a known positive constant. Step 1.2.

2. Modeling of disturbances to the multi-spacecraft system Characterization is performed, the model of which is: , , wherein, is the state variable of the modelable disturbance for the th spacecraft; is the modelable disturbance; and is a known matrix; Step 1.3, combining the ideal multi-spacecraft relative motion equation with the multi-source disturbances to obtain a multi-spacecraft relative motion equation affected by the multi-source disturbances: , wherein the spacecraft subject to state constraints and control constraints, denoted as: , , wherein , , and are known matrices.

3. The spacecraft non-cooperative game method based on model predictive control deep neural network according to claim 2, wherein, The step 3 specifically comprises the following steps: Step 3.1, designing a model predictive controller: , wherein is a virtual control quantity; is a decision variable; is a feedback gain; Step 3.2, designing an iterative optimization problem and solving to obtain a non-cooperative game equilibrium solution in an offline case: The constraint is , Wherein, , Step 3.2.2, iteratively solving the optimization problem designed in step 3.2.

1. , in, and They are respectively and The predicted value, and ; ; ; For state The nominal value; For error The nominal value; and , ; Representative at Always Predicting the timing; ; ; ; , ; ; ; ; , , ; ; ; ; ; ; It is the identity matrix; For the first State prediction quantity for each spacecraft; , , and Given a weight matrix; For prediction in the time domain; and To optimize the constraints imposed on the problem; The step 4 specifically comprises the following steps:

4. The spacecraft non-cooperative game method based on model predictive control deep neural network according to claim 3, wherein, Step 4.1, designing a deep neural network-based controller: Step 4.2, training neural network variables in the designed deep neural network-based controller: , in, This is the output of a deep neural network; the first... Layer The output of each neuron is ,function For activation function, For the first Layer The weight pairs of each neuron ( ), ; These are weight values; the output layer uses... ; The number of layers in its neural network, The number of its neurons, For the first Layers represent the number of neurons; design and adjust weights. Loss function: , wherein, is the designed deep neural network for the sample output, is the true value for the sample, is a batch of data; implementing wherein, is the accuracy parameter; Step 4.2.1, using state data in the deep neural network-based controller as input training set of the deep neural network: Step 4.2.2, using decision variable data in the deep neural network-based controller as output training set of the deep neural network: , wherein ~ is an adjacency matrix parameter; is a prediction sequence of states at time instant Step 4.3, using the output of the deep neural network online to obtain a non-cooperative game equilibrium solution in an online case, and using the non-cooperative game equilibrium solution as control input of the multi-spacecraft non-cooperative game. , wherein ( ) is an approximation of in is a sequence solved for; represents a set of positive integers from to ​​ ​

Citation Information

Patent Citations

  • Spacecraft formation discrete distributed non-cooperative game method based on dynamic event triggering

    CN112558471A

  • Four-rotor unmanned aerial vehicle optimal formation control method based on non-cooperative differential game

    CN117193359A