A non-cooperative spacecraft active tracking method based on deep reinforcement learning
By employing a Transformer-based deep reinforcement learning algorithm, combined with spacecraft dynamics and orbital dynamics, and using actor and critic networks, the model fusion and robustness issues in active tracking of non-cooperative spacecraft were addressed, achieving accurate and robust tracking and efficient control.
Patent Information
- Application Number
- CN202410966496.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-07-18
AI Technical Summary
Existing non-cooperative spacecraft active tracking algorithms based on deep reinforcement learning cannot effectively integrate spacecraft dynamics models and satellite orbital dynamics, thus failing to improve algorithm robustness and tracking accuracy, and failing to effectively extract time-series related information.
By employing a Transformer-based actor network and critic network, combined with a relative orbital motion dynamics model of the target spacecraft and the chasing spacecraft, and using a deep deterministic policy gradient method for end-to-end learning, we design novel reward and loss functions to extract high-level semantic information and temporal relationships, thereby achieving accurate and robust tracking of non-cooperative spacecraft.
It achieves optimal control of second-order control systems, improves the convergence and tracking performance of active tracking algorithms, avoids the complex model construction and parameter adjustment in traditional control theory, and has good robustness against sensor observation interference.
Smart Images

Figure CN119002255B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a non-cooperative spacecraft active tracking method based on deep reinforcement learning, and belongs to the field of aerospace. BACKGROUND
[0002] Currently, the active tracking problem of spacecraft is mainly solved by using traditional control methods, such as PID control, H∞ control, sliding mode control, etc. However, the traditional control method for solving the active tracking problem of spacecraft often relies on a relatively accurate control object model, and needs to design a stable control law through complex and tedious mathematical derivation and carefully adjust the parameters of the controller to ensure the effectiveness of the control algorithm. In fact, the attitude and orbit motion information of non-cooperative spacecraft is often inaccurate, and there are many uncertain high-intensity external disturbances in the relative motion dynamics model of the pursuit spacecraft and the target spacecraft, which poses a great challenge to the current spacecraft active tracking method. In recent years, the vigorous development of deep reinforcement learning algorithm provides a novel perspective for solving the active tracking problem of non-cooperative spacecraft.
[0003] The deep reinforcement learning algorithm mainly interacts with the environment through the agent, and learns the optimal control strategy according to the environment reward obtained. Its advantage is that it does not need to model the environment, but only needs to ensure the accuracy of the dynamics model of the agent itself. The agent constantly updates the strategy through iterative interaction with the environment, thereby obtaining better strategy parameters. Its advantage is that when facing a nonlinear, non-static, and non-deterministic environment, it can still optimize the control strategy of the agent through a large number of online / offline iterations, thereby obtaining a better control effect. In addition, the deep reinforcement learning algorithm can add different external disturbances during training, so that the agent can learn a more robust action strategy and make corresponding adjustments in real time according to the changes in the environment. SUMMARY
[0004] The purpose of the present application is to solve the defects of the existing non-cooperative spacecraft active tracking algorithm based on deep reinforcement learning, such as the inability to fuse the spacecraft dynamics model and satellite orbit dynamics, the inability to effectively improve the robustness of the algorithm while ensuring tracking accuracy, and the inability to effectively extract time-related information about the target from the training samples, and to propose a non-cooperative spacecraft active tracking method based on deep reinforcement learning.
[0005] The specific process of a non-cooperative spacecraft active tracking method based on deep reinforcement learning is as follows:
[0006] Step 1, based on the basic parameters of the target spacecraft and the pursuit spacecraft, the coordinate system of the target spacecraft and the pursuit spacecraft, the universal gravitation suffered by the target spacecraft and the pursuit spacecraft, the universal gravitation suffered by the target spacecraft is converted into the position and speed of the target spacecraft, and the universal gravitation suffered by the pursuit spacecraft is converted into the position and speed of the pursuit spacecraft;
[0007] Step 2, setting state space and action space, constructing actor network and critic network based on Transformer, and loss function of the actor network and loss function of the critic network;
[0008] Step 3, constructing a reward function, obtaining a trained actor network and critic network;
[0009] Step 4, obtaining the difference between the position and the expected position between the target spacecraft and the pursuit spacecraft, and the speed difference between the target spacecraft and the pursuit spacecraft, inputting the trained actor network, and the trained actor network outputting the action of the current time step.
[0010] The beneficial effects of the application are:
[0011] The application provides an end-to-end deep reinforcement learning-based non-cooperative spacecraft active tracking algorithm fusing a spacecraft dynamics model and a satellite orbit dynamics, which can process time sequence information and improve algorithm robustness.
[0012] The application proposes a deep reinforcement learning-based non-cooperative spacecraft active tracking algorithm, taking spacecraft pose information with estimation noise as input, combining the relative orbit motion dynamics model of the target spacecraft and the pursuit spacecraft, and learning an approximate optimal active tracking strategy in an end-to-end manner through a deep deterministic policy gradient method.
[0013] The application proposes an unsupervised deep reinforcement learning-based active visual tracker, which realizes accurate and robust tracking of non-cooperative spacecraft. The tracker realizes optimal control of a second-order control system, avoids complex model construction and parameter adjustment in traditional control theory, and deep reinforcement learning itself has good robustness to state data with disturbances and can cope with possible sensor observation interference.
[0014] The application designs a novel Transformer-based actor network and critic network, effectively extracts high-level semantic information and time sequence relationship from sequence input states, implicitly represents the coupled motion between the target spacecraft and the pursuit spacecraft, greatly improves the convergence of the active tracking algorithm and optimizes the tracking effect.
[0015] The loss function of the actor network and the critic network is improved, so that the optimal control strategy can be learned better, faster and more stably according to the input sequence state, and the tracking effect of the active tracking algorithm is improved. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 The structure diagram of the deep reinforcement learning network actor network and critic network;
[0017] Figure 2 The algorithm flowchart proposed in the application;
[0018] Figure 3 The final effect diagram of the application;
[0019] Figure 4 The control effect diagram of the PID control algorithm in the model of the application. DETAILED DESCRIPTION
[0020] Embodiment one: the specific process of the non-cooperative spacecraft active tracking method based on deep reinforcement learning is as follows:
[0021] Step 1: based on the basic parameters of the target spacecraft and the pursuit spacecraft, the coordinate systems of the target spacecraft and the pursuit spacecraft, and the universal gravitation suffered by the target spacecraft and the pursuit spacecraft, the universal gravitation suffered by the target spacecraft is converted into the position and speed of the target spacecraft, and the universal gravitation suffered by the pursuit spacecraft is converted into the position and speed of the pursuit spacecraft;
[0022] Step 2: setting the state space and the action space, constructing the actor network and the critic network based on the Transformer, and the loss function of the actor network and the loss function of the critic network;
[0023] And through the loss function, the actor network and the critic network can establish the time sequence relationship of the current input state and the historical input state;
[0024] Step 3: constructing the reward function, obtaining the trained actor network and critic network;
[0025] Step 4: obtaining the distance between the position of the target spacecraft and the pursuit spacecraft and the expected position, and the speed difference between the target spacecraft and the pursuit spacecraft, inputting the trained actor network, and the trained actor network outputting the action of the current time step, which is the control instruction of the force of the pursuit spacecraft.
[0026] Specific implementation method two: the difference between the embodiment and the specific implementation method one is that: in step 1, the gravitational force acting on the target spacecraft is converted into the position and speed of the target spacecraft based on the basic parameters of the target spacecraft and the pursuit spacecraft, the coordinate system of the target spacecraft and the pursuit spacecraft, and the gravitational force acting on the target spacecraft and the pursuit spacecraft.
[0027] The specific process is:
[0028] Step 11, defining the basic parameters of the target spacecraft and the pursuit spacecraft; the specific process is:
[0029] The basic parameters of the target spacecraft are: the mass of the target spacecraft itself, the maximum thrust of the target spacecraft, and the adjustment frequency of the target spacecraft; the self-weight is 100 kg, the maximum thrust is 4N, and the adjustment frequency is 10Hz;
[0030] The basic parameters of the pursuit spacecraft are: the mass of the pursuit spacecraft itself, the maximum thrust of the pursuit spacecraft, and the adjustment frequency of the pursuit spacecraft; the self-weight is 100 kg, the maximum thrust is 4N, and the adjustment frequency is 10Hz;
[0031] Step 12, setting the coordinate system of the target spacecraft and the pursuit spacecraft; the specific process is:
[0032] The coordinate system of the target spacecraft and the pursuit spacecraft uses the global geocentric inertial coordinate system, and only considers the two-dimensional motion of the target spacecraft and the pursuit spacecraft, without considering the position transformation, speed transformation and thrust output on the z-axis, and the geocentric inertial coordinate system origin is the earth center;
[0033] The position information of the pursuit spacecraft and the target spacecraft in the present application is the self-coordinate of the pursuit spacecraft and the target spacecraft in the geocentric inertial coordinate system;
[0034] Step 13, defining the gravitational force acting on the target spacecraft and the pursuit spacecraft; the specific process is:
[0035] The gravitational force acting on the target spacecraft F g,tar And the gravitational force acting on the pursuit spacecraft F g,c As shown in the following formula:
[0036]
[0037] r c =r tar +c,c∈[-150,150]
[0038] Wherein,
[0039] r tar The initial orbit radius of the target spacecraft is rc To pursue the target spacecraft's orbit radius;
[0040] G is the gravitational constant, the value is 6.67 x 10 -11 N·m 2 / kg 2 ;
[0041] M is the mass of the earth, the value is 5.977 x 10 24 kg;
[0042] m tar is the target spacecraft's mass, m c is the pursuit spacecraft's mass;
[0043] c is a random variable to adjust the initial position of the pursuit spacecraft, to increase the robustness of the algorithm;
[0044] Step 14, convert the gravitational force on the target spacecraft into the position and velocity of the target spacecraft;
[0045] Convert the gravitational force on the pursuit spacecraft into the position and velocity of the pursuit spacecraft;
[0046] The specific process is as follows:
[0047] Step 141, convert the gravitational force on the target spacecraft into the position and velocity of the target spacecraft; the specific conversion is as follows:
[0048]
[0049] Where, · represents matrix point multiplication, P t,tar represents the position of the target spacecraft in the geocentric inertial coordinate system at time t, V t,tar represents the velocity of the target spacecraft in the geocentric inertial coordinate system at time t, P t+1,tar represents the position of the target spacecraft in the geocentric inertial coordinate system at time t+1, V t+1,tar represents the velocity of the target spacecraft in the geocentric inertial coordinate system at time t+1; Δt is the adjustment frequency; A is the spacecraft information conversion matrix at time t;
[0050] The target spacecraft is not controlled by the deep reinforcement learning network, so there is no F RL term when calculating the position and velocity change of the target spacecraft;
[0051] represents the unit vector of the target spacecraft at the current position, to decompose the gravitational force into component forces; the specific form is as follows:
[0052]
[0053] Where, r2,tar represents the vector position of the target spacecraft in the earth-centered inertial coordinate system, and r1 represents the vector position of the earth's center in the earth-centered inertial coordinate system;
[0054] Step 142, converting the universal gravitation suffered by the pursuit spacecraft into the position and velocity of the pursuit spacecraft; the specific conversion is shown in the following formula:
[0055]
[0056] Wherein, F RL represents the action output by the actor network (the output of the critic network is only used in the updating process, guiding the updating process of the actor network and the critic network), P t,c represents the position of the pursuit spacecraft in the earth-centered inertial coordinate system at time t, V t,c represents the velocity of the pursuit spacecraft in the earth-centered inertial coordinate system at time t, P t+1,c represents the position of the pursuit spacecraft in the earth-centered inertial coordinate system at time t+1, V t+1,c represents the velocity of the pursuit spacecraft in the earth-centered inertial coordinate system at time t+1; Δt is the adjustment frequency; A is the spacecraft information conversion matrix at time t;
[0057] is the unit vector of the pursuit spacecraft at the current position, used to decompose the universal gravitation into component forces; the specific form is shown in the following formula:
[0058]
[0059] Wherein, r 2,c represents the vector position of the pursuit spacecraft in the earth-centered inertial coordinate system, and r1 represents the vector position of the earth's center in the earth-centered inertial coordinate system;
[0060] The other steps and parameters are the same as in the first embodiment.
[0061] The third embodiment is different from the first or second embodiment in that the specific form of the spacecraft information conversion matrix A at time t is shown in the following formula:
[0062]
[0063] The other steps and parameters are the same as in the first or second embodiment.
[0064] Specific embodiment four: different from one of the specific embodiments one to three is that: the state space and the action space are set in step 2, the actor network and the critic network based on the Transformer are constructed, and the loss function of the actor network and the loss function of the critic network; and the actor network and the critic network can establish the time sequence relationship of the current input state and the historical input state through the loss function; comprising the following steps:
[0065] Step 21, setting a state space;
[0066] The state space contains the difference between the position and the expected position between the target spacecraft and the pursuit spacecraft, and the speed difference between the target spacecraft and the pursuit spacecraft;
[0067] The state space is a continuous state space;
[0068] The state space is a dictionary containing a two-dimensional vector representing the position difference and a two-dimensional vector representing the speed difference, and in the present application, the continuous n training samples are integrated as a training sample input into the deep reinforcement learning network for training, therefore the specific form of the training sample in the present application is [Bx4xn], in the present application, n is defined as 10, B is the batch, which is set to 64 in the present application, therefore the input state should be [s t-9 ,s t-8 ,...,s t ];
[0069] Step 22, setting an action space;
[0070] The action space is defined as a two-dimensional vector, and the range of the two-dimensional vector is specified as -4 to 4;
[0071] The action space is a continuous action space: a = [a x ,a y ], a x ∈[-4,4], a y ∈[-4,4];
[0072] Wherein, a is the action, a x is the action applied to the pursuit spacecraft in the x-axis of the geocentric inertial coordinate system, and a y is the action applied to the pursuit spacecraft in the y-axis of the geocentric inertial coordinate system;
[0073] Step 23, constructing an actor network and a critic network;
[0074] The actor network comprises in sequence: an embedding layer, a position coding layer, a first encoding layer, a second encoding layer, a third encoding layer, a first full connection layer, a second full connection layer, a third full connection layer, a first dimension reduction layer, a fourth full connection layer, a second dimension reduction layer, a fifth full connection layer and an output layer.
[0075] The specific processing procedure of the actor network is as follows:
[0076] The current state data is input into the embedding layer, the position coding layer, the first encoding layer, the second encoding layer, the third encoding layer, the first full connection layer, the second full connection layer, the third full connection layer, the first dimension reduction layer, the fourth full connection layer, the second dimension reduction layer, the fifth full connection layer and the output layer in sequence, and the output layer outputs the action under the current state (the input is the current state, and the output is the action under the current state).
[0077] The critic network comprises in sequence: an embedding layer, a position coding layer, a first encoding layer, a second encoding layer, a third encoding layer, a first full connection layer, a second full connection layer, a third full connection layer, a first dimension reduction layer, a fourth full connection layer, a second dimension reduction layer, a splicing layer, a fifth full connection layer and an output layer.
[0078] The specific processing procedure of the critic network is as follows:
[0079] The current time state data is input into the embedding layer, the position coding layer, the first encoding layer, the second encoding layer, the third encoding layer, the first full connection layer, the second full connection layer, the third full connection layer, the first dimension reduction layer, the fourth full connection layer and the second dimension reduction layer in sequence, the second dimension reduction layer outputs the feature data corresponding to the current time action under the current time state, and the current time action is input into the splicing layer (combined) in sequence after being input into the fifth full connection layer and the output layer, and the output layer outputs the evaluation value data (the critic network input is the current state action pair (s t ,a t ), and the output is the evaluation q value of (s t ,a t ));
[0080] In the application, the input has strong time sequence continuity, so the position coding is used to define the sequence of the input information, so as to better extract the time sequence information in the input.
[0081] The Transformer encoder is used to extract the time sequence features in the continuous input sequence, and the full connection layer and the dimension reduction layer are used to meet the output dimension of the network, and the structure of the actor network and the critic network in the application is as shown in Figure 1
[0082] Wherein the input data is first mapped to a higher dimension using an embedding layer, the dimension of the embedding layer is set to 128 bits in the present application; and the sequence of the input data is defined by a position encoding layer to represent its time sequence; the features of the input information, especially the time sequence information, are extracted by three encoding layers r; then the last dimension of the data dimension is reduced to 1 dimension by three fully connected layers, so as to reduce the dimension of the data by a dimension reduction layer, and the last dimension of the data is further reduced to 1 dimension by a fourth fully connected layer.
[0083] In the actor network, the data is reduced to the output dimension suitable for the environment by a fifth fully connected layer, while in the critic network, the data needs to be combined with the input batch action, and the output is directly given for weighted average processing;
[0084] Step 24, constructing a target actor network and a target critic network;
[0085] The target actor network comprises in sequence: an embedding layer, a position encoding layer, a first encoding layer (Encoder Layer), a second encoding layer, a third encoding layer, a first fully connected layer, a second fully connected layer, a third fully connected layer, a first dimension reduction layer, a fourth fully connected layer, a second dimension reduction layer, a fifth fully connected layer, and an output layer.
[0086] The specific processing process of the target actor network is as follows:
[0087] The next time state s t+1 The data is sequentially input into the embedding layer, the position encoding layer, the first encoding layer (Encoder Layer), the second encoding layer, the third encoding layer, the first fully connected layer, the second fully connected layer, the third fully connected layer, the first dimension reduction layer, the fourth fully connected layer, the second dimension reduction layer, the fifth fully connected layer, and the output layer, and the output layer outputs the next time state s t+1 The next action a t+1 Data;
[0088] The target critic network comprises in sequence: an embedding layer, a position encoding layer, a first encoding layer (Encoder Layer), a second encoding layer, a third encoding layer, a first fully connected layer, a second fully connected layer, a third fully connected layer, a first dimension reduction layer, a fourth fully connected layer, a second dimension reduction layer, a splicing layer, a fifth fully connected layer, and an output layer.
[0089] The specific processing process of the target critic network is as follows:
[0090] The next time state s t+1The data is sequentially input into an embedding layer, a position encoding layer, a first encoding layer (EncoderLayer), a second encoding layer, a third encoding layer, a first full connection layer, a second full connection layer, a third full connection layer, a first dimension reduction layer, a fourth full connection layer, a second dimension reduction layer, and the feature data corresponding to the next moment action of the state of the next moment is input into a fifth full connection layer and an output layer in sequence after being spliced (combined) with the next moment action input of the splicing layer, and the output layer outputs target value data of the state action pair (s t+1 ,a t+1 );
[0091] Step 25, setting the loss function of the actor network; the specific process is:
[0092] The loss function of the actor network in the application first calculates the action under the current state, at this time, the action is an action without noise, and the cumulative expected return of the current state and the action is calculated by the critic network, since the gradient ascent is difficult to achieve, the negative cumulative expected return is taken here, and the weighted average of the obtained batch actor network negative cumulative expectation is taken, and higher weight is given to the step with larger gradient, so as to obtain the loss function of the actor network;
[0093] The expression of the loss function of the actor network is:
[0094]
[0095] Wherein, J μ is the loss function of the actor network, Q is the critic network, mu is the actor network, theta is the network parameter, theta μ is the parameter of the actor network, theta Q is the parameter of the critic network, omega t is the weight, t is the time step, B is the batch quantity, which is set to 64 in the application, s t is the current moment state;
[0096] The acquisition process of the loss function of the actor network is:
[0097] In the network updating process, the Q value obtained by the critic network has a dimension of Bx1, the cumulative expected return is obtained based on the Q value obtained by the critic network, and the obtained cumulative expected return is sorted, that is:
[0098] q qt =-Q(s t ,μ(s t |θ μ )|θ Q )
[0099] q q1 <q q2 <...<q q64
[0100] wherein q qt is the cumulative expected return;
[0101] The input of the critic network is the current state-action pair (s t , a t ), but this a t is not the a t obtained in the experience collection process, but the a t obtained by inputting the state into the actor network, and the output of the critic network is the evaluation value q t of the current state and action;
[0102] (here q q1 only represents the maximum Q value, not the first item in a batch of Q values, the same below) Since the gradient descent method is usually used to update the network parameters in programming, the Q value here is taken as negative, and the smaller the Q value, the higher the weight;
[0103] The Q values are sorted in ascending order of cumulative expected return, and different weight values are assigned:
[0104] ω q1 > ω q2 >... > ω q64
[0105]
[0106] wherein ω qt is the weight value corresponding to the Q value;
[0107] The loss function of the actor network is constructed based on q qt and ω qt :
[0108]
[0109] Step 26, set the loss function of the critic network; the specific process is:
[0110] In the present application, in addition to the actor and critic networks, there are also target actor and target critic networks, which are mainly used to calculate the expected value of the state-action pair, and compare it with the actual evaluation value, and calculate the resulting time difference error by weighted average, and the higher the time difference error will be given a higher weight, and the loss will be solved according to the resulting result, the target actor network has the same structure as the actor network, and the target critic network has the same structure as the critic network;
[0111] The expression of the loss function of the critic network is:
[0112]
[0113] where J Q is the loss function of the critic network; r t is the reward value of performing the current action in the current state (obtained by the reward function calculation in step 33); done is a judgment of whether the current state is an end state; Q' is the target critic network; μ' is the target actor network; γ is a decay function approaching 1, which reduces the influence of the uncertainty of future estimates on the target value; s t+1 is the state at time t+1; a t is the action at the current time; θ μ' is the parameter of the target actor network; θ Q' is the parameter of the target critic network.
[0114] The remaining symbols are the same as in the previous step.
[0115] The target actor network μ'(s t+1 | θ μ' ) calculates the action a t+1 in the current state.
[0116] The loss function of the critic network is obtained as follows:
[0117] First, the critic network is used to calculate the evaluation value of the action pair (s t , a t ) in the current state (the input of the critic network is (s t , a t ), and the output is q t ):
[0118] q t = Q(s t , a t | θ Q )
[0119] where q t is the evaluation value at the current time, Q is the critic network, s t is the state at the current time, a t is the action at the current time, and θ Q is the parameter of the critic network.
[0120] Through the target actor network, the action at the next time is obtained (the input of the target actor network is s t+1 , and the output is a t+1 ):
[0121] a t+1 = μ'(s t+1 | θ μ' )
[0122] where a t+1Let μ' be the target actor network, and s be the action predicted at the next moment. t+1 To determine the state for the next moment (before training begins, the collected data will be used to determine the state of the next moment). t ,a t ,r t ,s t+1 D) Store the data in an experience replay pool, and continuously add to it during subsequent training. Once the pool reaches its limit, remove the first data and add the latest data, and so on. See step 34 for details. θ μ' Parameters for the target actor network;
[0123] Using a target commentator network, state-action pairs (s) are calculated by combining the next action, the next state, the reward obtained in the current step, and whether it is a terminating state. t+1 ,a t+1 Target value:
[0124] y t =r t +γ(1-done)Q'(s t+1 ,a t+1 |θ Q' )
[0125] In the formula, y t For state-action pairs (s) t+1 ,a t+1 The target value, r t For the current state s t Perform action a t The reward value (obtained by the reward function in step 33), Q' is the target commenter network, θ Q' Parameters for the target critic network;
[0126] The input to the target critic network is the current state-action pair (s) t ,a t ), except this a t a was not obtained during the experience collection process. t Instead, a is obtained again through the actor's network input state. t The output of the target critic network is an evaluation value q for the current state and action. t ;
[0127] After calculating the target value and the evaluated value, subtract the target value from the evaluated value:
[0128] L t =y t -q t
[0129] The obtained L tBxl is one-dimensional, and the critic network loss function is finally obtained by the weighted average algorithm as described in step 24:
[0130]
[0131] The other steps and parameters are the same as one of the first to third embodiments.
[0132] Embodiment five: different from one of the first to fourth embodiments is that the reward function is constructed in step 3, and the trained actor network, critic network, target actor network, and target critic network are obtained.
[0133] The active tracking strategy learning based on deep reinforcement learning continuously obtains training samples and adjusts network parameters through interaction with the environment; the active tracking network of the non-cooperative spacecraft is obtained. The input data of this algorithm in the training and application processes are the same, which are the position difference and the expected position difference between the target spacecraft and the pursuit spacecraft, and the speed difference between the target spacecraft and the pursuit spacecraft, and then the output is the control instruction of the force of the pursuit spacecraft, which completes the active tracking of the target spacecraft.
[0134] The steps include:
[0135] Step 31, initializing the initial position of the target spacecraft and the pursuit spacecraft; the specific process is:
[0136] The initial position of the target spacecraft is on an orbit 7500 km away from the Earth's center, that is, the initial orbit radius r of the target spacecraft mentioned in step 13;
[0137] The position of the pursuit spacecraft is on and inside a circle with the target spacecraft as the center and 1 km as the radius;
[0138] Step 32, initializing the initial speed of the target spacecraft and the pursuit spacecraft; the specific process is:
[0139] The initial speed of the target spacecraft is set according to the following formula:
[0140]
[0141] Where, v tar is the initial speed of the target spacecraft under the action of the universal gravitation at the initial position; r tar is the initial orbit radius of the target spacecraft, F g,tar is the universal gravitation received by the target spacecraft, m tar is the weight of the target spacecraft, X tar , Y tar , and Z tar are the position coordinates of the target spacecraft in the Earth's center inertial coordinate system. is the velocity unit vector of the target spacecraft at the current position, and · is the dot product, is the velocity of the target spacecraft, x,tar is the velocity of the target spacecraft in the x-axis of the Earth-centered inertial coordinate system, x,tar is the negative value, v y,tar is the velocity of the target spacecraft in the y-axis of the Earth-centered inertial coordinate system, z,tar is the velocity of the target spacecraft in the z-axis of the Earth-centered inertial coordinate system,
[0142] The initial velocity of the pursuit spacecraft is set according to the following formula:
[0143]
[0144] wherein v c is the initial velocity of the pursuit spacecraft under the action of the universal gravitation at the initial position; r c is the orbital radius of the pursuit spacecraft, F g,c is the universal gravitation to which the pursuit spacecraft is subjected, m c is the weight of the pursuit spacecraft, X c , Y c , Z c is the position coordinate of the pursuit spacecraft in the Earth-centered inertial coordinate system; is the velocity unit vector of the pursuit spacecraft at the current position, and · is the dot product, is the velocity of the pursuit spacecraft, x,c is the velocity of the pursuit spacecraft in the x-axis of the Earth-centered inertial coordinate system, x,c is the negative value, v y,c is the velocity of the pursuit spacecraft in the y-axis of the Earth-centered inertial coordinate system, z,c is the velocity of the pursuit spacecraft in the z-axis of the Earth-centered inertial coordinate system,
[0145] Step 33, defining a reward function;
[0146] Step 34, setting experience replay pool parameters, critic network and actor network hyperparameters;
[0147] Step 35, action selection strategy;
[0148] Step 36, extracting n consecutive training samples (n groups (s t , a t , r t , s t+1 , D) from the experience replay pool as a set of training samples, wherein n is set to 10 in the present application;
[0149] In the present application, when the network parameters are updated, the training data at the current time is the training sample from time t-9 to t, and the training data at the next time will be the data from t-8 to t+1.
[0150] Step 37, predicting the action value of the current time step by using the critic network (the critic network input is the state and action of the current time step, and the critic network output is the predicted action value of the current time step);
[0151] updating the actor network parameters and the target actor network parameters;
[0152] Step 38, predicting the action of the current time step by using the actor network (the actor network input is the state of the current time step, and the actor network output is the predicted action of the current time step), while estimating the action value of the next time step by using the critic network, and finally updating the critic network parameters and the target critic network parameters according to the time difference error (L t = y t - q t , y t = r t + γ(1-done)Q'(s t+1 , a t+1 | θ Q' );
[0153] Step 39, repeating steps 33 to 38 until the trained actor network and critic network are obtained.
[0154] The other steps and parameters are the same as one of the first to fourth embodiments.
[0155] Embodiment six: the difference between this embodiment and the first to fifth embodiments is that the reward function is defined in step 33, and the specific process is as follows:
[0156] The reward function will be defined according to the tracked relative position accuracy. The closer the relative position of the pursuit spacecraft and the target spacecraft to the expected relative position, the smaller the corresponding position reward and punishment will be, and the more tracking steps, the higher the tracking reward will be.
[0157] Let the position coordinates (X c , Y c , Z c ) of the pursuit spacecraft in the geocentric inertial coordinate system be P c ;
[0158] Let the position coordinates (X tar , Y tar , Z tar ) of the target spacecraft in the geocentric inertial coordinate system be P t ;
[0159] The reward function is as follows:
[0160] r = - (|| Pc -P t ||-P e )×f p
[0161]
[0162] wherein, P e is the desired relative position between the pursuit spacecraft and the target spacecraft, f p is the adjustment coefficient of the reward function, which is set to 0.01 in the present application; c is a constant term; and r is the reward value;
[0163] The reward function is mainly calculated at each step during network simulation, and the initialization is to give the agent an initial value for an initial state in the subsequent simulation process. The initialization does not participate in the reward function calculation, and the calculation starts after the first action is performed in the initialization state.
[0164] In the present application, the pursuit spacecraft needs to complete the pursuit and keep the relative distance unchanged at the same time, so the desired distance of the pursuit spacecraft relative to the target spacecraft is set to 300.
[0165] In order to encourage the algorithm to improve the tracking time, a constant reward term c is also added. If the deviation between the relative distance and the desired distance is less than 500, a constant term is added to the reward, otherwise it will not be added. In the present application, the reward constant term is set to 1.
[0166] The other steps and parameters are the same as one of the first to fifth embodiments.
[0167] Embodiment seven: the present embodiment is different from one of the first to sixth embodiments in that the step 34 sets the experience replay pool parameters, the critic network and the actor network hyperparameters; the specific process is:
[0168] The total training number of curtains is 2000, the maximum step length of each curtain is 1000 steps, the action random disturbance conforms to the Gaussian distribution, the initial standard deviation is 0.7, the final standard deviation is 0.1, the mean value is 0, and the reward decay coefficient is 0.95;
[0169] The experience replay pool starts to train after collecting 10000 groups of samples initially, the total capacity is 50000, and the experience replay pool adopts the first-in-first-out strategy. (s t ,a t ,r t ,s t+1 ,D) as a group of samples;
[0170] wherein, D is a marker of whether the present curtain is terminated;
[0171] In the present application, if the deviation between the relative distance between the pursuit spacecraft and the target spacecraft and the expected distance exceeds 500 for 10 consecutive times, it is determined that the present act fails, and the next act will be restarted, if there is a step in which the deviation is less than 500, the counting will be restarted.
[0172] The other steps and parameters are the same as one of embodiments 1 to 6.
[0173] Embodiment 8: The difference between the present embodiment and one of embodiments 1 to 7 is that the action selection strategy in step 35; the specific process is:
[0174] The action selection strategy is divided into two cases;
[0175] When the number of samples in the experience replay pool is less than or equal to 10, the action output by the random action generation network (randomly selected from the action space) is adopted, and the parameters of the random action generation network are shared with the actor network parameters;
[0176] The structure of the random action generation network is the same as that of the actor network, and the parameters are copied from the actor network at the beginning of each act, and after the experience exceeds 10, the random action generation network will not be used until the next act begins;
[0177] When the number of samples in the experience replay pool is greater than 10, the action output by the actor network is adopted, and the input of the actor network is a continuous state sequence;
[0178] In the early stage of retraining and at the beginning of each act, there will be no continuous sequence state with a length of n, so in the present application there is also a random action generation network, which will randomly select an action from the same action space, and the network parameters will be shared with the actor network, and the input dimension will be changed through a fully connected layer to adapt to the actor network parameters;
[0179] In the experience collection phase, when the total amount of data in the experience replay pool is less than n or the total number of steps in each act is less than n, the random action generation network will select an action from the action space according to the initialized network parameters for the pursuit spacecraft to execute. In the training phase, when the total number of steps in the first act is less than n, the random action generation network will still be used to generate random actions, and when the first act ends, the actor network parameters will be synchronized to the random action generation network, so that the random action generation network will generate more reasonable action outputs in the next act. The different dimensions of network parameters will change the input dimension through a fully connected layer.
[0180] The other steps and parameters are the same as one of embodiments 1 to 7.
[0181] Specific implementation nine: different from one of the specific implementations one to eight, in the step 37, the critic network is used to predict the action value of the current time step (the critic network input is the state and action of the current time step, and the critic network output is the predicted action value of the current time step);
[0182] The actor network parameters and the target actor network parameters are updated;
[0183] The specific process is as follows:
[0184] As described in step 24, the critic network is used to evaluate the current state-action pair (s t ,a t ), and the evaluation value is taken as negative, and the evaluation value is compared to obtain a weight with a total of 1 according to the size, and the weight is used as the loss of the actor network (the expression of the loss function of the actor network is: The actor network is updated by gradient descent.
[0185] Based on the loss function of the actor network, the actor network parameters are updated using the backpropagation algorithm to minimize the loss function;
[0186] For the parameter update of the target actor network, a new parameter update strategy will be used, and the specific update method is as follows:
[0187] P μ' =(1-τ)×μ+τ×μ'
[0188] Where P μ' is the network parameter of the target actor network, τ is the parameter ratio coefficient defined by the actor network and the target actor network, μ is the actor network, and μ' is the target actor network;
[0189] In the present application, τ is not a constant value, because in the early stage of training, the total number of steps is small, and the network is extremely unstable, and the target actor network needs to be updated in time to better calculate the target value, therefore, in the present application, τ is defined as follows:
[0190]
[0191] Where T is the total number of steps;
[0192] That is, τ will be very small in the early stage, and the target actor network will be basically the same as the actor network, and as the total number of steps T increases, τ will gradually increase, and the maximum value will not exceed 0.5.
[0193] The other steps and parameters are the same as one of the specific implementations one to eight.
[0194] Specific implementation ten: different from one of the specific implementations one to nine: in the step 38, the actor network is used to predict the action of the current time step (the actor network input is the state of the current time step, and the actor network output is the predicted action of the current time step), and the critic network is used to estimate the action value of the next time step, and finally, the temporal difference error (L t = y t -q t , y t = r t + gamma (1-done) Q'(s t+1 , a t+1 | theta Q' ) ) is used to update the critic network parameters and the target critic network parameters;
[0195] The specific process is as follows:
[0196] Based on the loss function of the critic network, the critic network parameters are updated using the back propagation algorithm to minimize the loss function;
[0197] For the parameter update of the target critic network, a new parameter update strategy will be used, and the specific update method is as follows:
[0198] P Q' = (1-tau') * Q + tau' * Q'
[0199] Where P Q' is the network parameter of the target critic network, tau' is the parameter ratio coefficient of the critic network and the target critic network, Q is the critic network, and Q' is the target critic network;
[0200] In the present application, tau' is not a constant value, because in the early stage of training, the total number of steps is small, and the network is extremely unstable, so the target critic network needs to be updated in time to better calculate the target value, therefore, in the present application, tau' is defined as follows:
[0201]
[0202] Where T is the total number of steps;
[0203] That is, tau' will be very small in the early stage, and the target critic network will be basically the same as the critic network, and as the total number of steps T increases, tau' will gradually increase, and the maximum value will not exceed 0.5;
[0204] As step 26, first change the input dimension through the critic network, reduce the input dimension from B*4*n to B*4, and splice the input with the input batch action (dimension B*2) to become a B*6 input, and reduce it to B*1 dimension through the full connection layer, and calculate the target value through the target actor network and the target critic network, compare it with the evaluation value of the current state action pair (s t ,a t ) through the same weighted average method, calculate the loss value of the critic network, and update the critic network parameters through the back propagation algorithm to minimize the loss function;
[0205] Based on the loss function of the critic network, the critic network parameters are updated using the back propagation algorithm to minimize the loss function;
[0206] The target critic network parameter update is consistent with step 37.
[0207] The other steps and parameters are the same as one of the first nine embodiments.
[0208] The performance of the deep reinforcement learning active tracker is evaluated, including the following steps:
[0209] 1. In the present application, the average episode tracking length will be used as a main performance indicator of the deep reinforcement learning active tracker, and the obtained network model will be applied in the test process. In the test process, the flight trajectory of the target spacecraft will be changed to verify the robustness of the obtained network model;
[0210] The present application is divided into training and testing stages. Every 20 episodes, the training will be paused, and the non-cooperative spacecraft pursuit and relative distance maintaining tasks in different environments will be tested according to the existing network parameters. Different scenarios are used to test whether the obtained strategy network is only effective for a single scenario and lacks robustness.
[0211] The average episode tracking length will be tested for 10 episodes in multiple different scenarios, and the tracking steps of the 10-episode test in multiple scenarios will be averaged to obtain the average episode tracking length. The average episode tracking length will evaluate whether the obtained strategy can effectively keep the target within a certain range. Although it may not be guaranteed to be at the desired distance, it can still be kept within an acceptable distance;
[0212] 2. In the present application, the average episode reward will also be used as part of the evaluation indicator. If the average episode tracking length is high but the average episode reward is low, it means that the obtained network model can only maintain tracking but cannot keep the relative position close to the desired relative position;
[0213] The average episode reward is basically consistent with 1, but all the rewards obtained by each episode under each scene are averaged to verify whether the obtained strategy can guarantee that the expected distance can be approached as much as possible under the premise of guaranteeing that the tracking distance is in an acceptable range, and the higher the average episode reward is, the higher the tracking accuracy is;
[0214] 3, the speed of reaching the expected relative position is also used as part of the evaluation index, compared with the traditional PID control algorithm, if the expected position can be reached in a shorter number of steps, it means that the proposed deep reinforcement learning active tracker can reach the expected relative position with less energy consumption;
[0215] Since it is aimed at spacecraft, it is the most ideal case to consume the minimum resources to reach the specified position in space, therefore, a threshold is set, when the error between the relative distance and the expected distance is less than a certain range, it is considered that the pursuit spacecraft reaches the specified position, after that, the size of the output action is further monitored, the smaller the action is, the smaller the energy used is, this evaluation standard can be used as a supplementary evaluation standard.
[0216] The following examples are used to verify the beneficial effects of the present application:
[0217] Example 1:
[0218] In combination Figure 1 It is illustrated that in this embodiment, the end-to-end non-cooperative spacecraft active tracking algorithm based on robust deep reinforcement learning involved in this embodiment, the collected training samples are preprocessed, the feature extraction and function fitting of the input samples are realized through the Transformer encoder and the full connection layer, the time series motion information of the target spacecraft is estimated, and the approximate optimal tracking strategy is given through the deep deterministic policy gradient algorithm, so as to realize the pursuit of the non-cooperative target spacecraft and keep the relative distance unchanged.
[0219] The present application can also have other various embodiments, those skilled in the art can make various corresponding changes and modifications according to the present application without departing from the spirit and essence of the present application, but these corresponding changes and modifications should all belong to the protection scope of the claims attached to the present application.
Claims
1. A non-cooperative spacecraft active tracking method based on deep reinforcement learning, characterized in that: The method specifically comprises the following steps: Step 1, based on the basic parameters of the target spacecraft and the pursuit spacecraft, the coordinate system of the target spacecraft and the pursuit spacecraft, and the gravitational force suffered by the target spacecraft and the pursuit spacecraft, the gravitational force suffered by the target spacecraft is converted into the position and speed of the target spacecraft, and the gravitational force suffered by the pursuit spacecraft is converted into the position and speed of the pursuit spacecraft; Step 2, setting state space and action space, constructing actor network and critic network based on Transformer, and loss function of the actor network and loss function of the critic network; Step 3, constructing a reward function, obtaining a trained actor network and critic network; Step 4, obtaining the distance between the target spacecraft and the pursuit spacecraft and the expected position, and the speed difference between the target spacecraft and the pursuit spacecraft, inputting the trained actor network, and the trained actor network outputting the action at the current time step; The method specifically comprises the following steps: Step 1, based on the basic parameters of the target spacecraft and the pursuit spacecraft, the coordinate system of the target spacecraft and the pursuit spacecraft, and the gravitational force suffered by the target spacecraft and the pursuit spacecraft, the gravitational force suffered by the target spacecraft is converted into the position and speed of the target spacecraft, and the gravitational force suffered by the pursuit spacecraft is converted into the position and speed of the pursuit spacecraft; The method specifically comprises the following steps: Step 11, defining the basic parameters of the target spacecraft and the pursuit spacecraft; the method specifically comprises the following steps: The basic parameters of the target spacecraft are the mass of the target spacecraft itself, the maximum thrust of the target spacecraft, and the adjustment frequency of the target spacecraft; The basic parameters of the pursuit spacecraft are the mass of the pursuit spacecraft itself, the maximum thrust of the pursuit spacecraft, and the adjustment frequency of the pursuit spacecraft; Step 12, setting the coordinate system of the target spacecraft and the pursuit spacecraft; the method specifically comprises the following steps: The coordinate system of the target spacecraft and the pursuit spacecraft uses the geocentric inertial coordinate system, and only considers the two-dimensional motion of the target spacecraft and the pursuit spacecraft, without considering the position transformation, speed transformation and thrust output on the z-axis, and the origin of the geocentric inertial coordinate system is the earth center; The gravitational force F experienced by the target spacecraft g,tar The gravitational force F experienced by the pursuit spacecraft g,c As shown in the following equation: r c = r tar + c, c e [-150, 150] Step 13, defining the gravitational force suffered by the target spacecraft and the pursuit spacecraft; the method specifically comprises the following steps: r tar r is the initial orbit radius of the target spacecraft c r is the orbit radius of the chaser spacecraft G is the gravitational constant, having a value of 6.67 x 10 -11 N·m 2 / kg 2 ; M is the mass of the earth, with a value of 5.977 x 10 24 kg; m tar m c m Wherein, c is a random variable; Step 14, converting the gravitational force suffered by the target spacecraft into the position and speed of the target spacecraft; Converting the gravitational force suffered by the pursuit spacecraft into the position and speed of the pursuit spacecraft; The method specifically comprises the following steps: wherein, · represents matrix point multiplication, P t,tar represents the position of the target spacecraft in the Earth-Centered Inertial coordinate system at time t, V t,tar represents the velocity of the target spacecraft in the Earth-Centered Inertial coordinate system at time t, P t+1,tar represents the position of the target spacecraft in the Earth-Centered Inertial coordinate system at time t+1, V t+1,tar represents the velocity of the target spacecraft in the Earth-Centered Inertial coordinate system at time t+1; Δt is the adjustment frequency; A is the spacecraft information conversion matrix at time t; a unit vector representing the target spacecraft at the current position, in the form of the following equation: wherein r 2,tar represents the vector position of the target spacecraft in the geocentric inertial coordinate system, and r1 represents the vector position of the geocenter of the Earth in the geocentric inertial coordinate system; Step 141, converting the gravitational force suffered by the target spacecraft into the position and speed of the target spacecraft; the specific conversion is shown in the following formula: where F RL represents the action output by the actor network, P t,c represents the position of the pursuit spacecraft in the geocentric inertial coordinate system at time t, V t,c represents the velocity of the pursuit spacecraft in the geocentric inertial coordinate system at time t, P t+1,c represents the position of the pursuit spacecraft in the geocentric inertial coordinate system at time t+1, V t+1,c represents the velocity of the pursuit spacecraft in the geocentric inertial coordinate system at time t+1; Δt is the adjustment frequency; A is the spacecraft information conversion matrix at time t; is the unit vector in the direction of the spacecraft's current position; in particular, it is given by the following expression: where r 2,c represents the vector position of the pursuit spacecraft in the geocentric inertial coordinate system, and r1is the vector position of the Earth's center in the geocentric inertial coordinate system; Step 142, converting the gravitational force suffered by the pursuit spacecraft into the position and speed of the pursuit spacecraft; the specific conversion is shown in the following formula: The specific form of the spacecraft information transpose matrix A at time t is shown in the following formula: The step 2 of setting state space and action space, constructing actor network and critic network based on Transformer, and loss function of the actor network and loss function of the critic network; comprises the following steps: Step 21, setting state space; The state space includes a difference between a position and an expected position between the target spacecraft and the pursuit spacecraft, and a velocity difference between the target spacecraft and the pursuit spacecraft; The state space is a continuous state space; Step 22, an action space is set; The action space is defined as a two-dimensional vector; Action space is continuous action space: a = [a x ,a y ], a x ∈ [-4, 4], a y ∈ [-4, 4] where a is the action, a x is the action applied to the chaser spacecraft in the x-axis of the geocentric inertial coordinate system, y is the action applied to the chaser spacecraft in the y-axis of the geocentric inertial coordinate system. Step 23, an actor network and a critic network are constructed; The actor network sequentially includes an embedding layer, a position encoding layer, a first encoding layer, a second encoding layer, a third encoding layer, a first full connection layer, a second full connection layer, a third full connection layer, a first dimension reduction layer, a fourth full connection layer, a second dimension reduction layer, a fifth full connection layer, and an output layer; The specific processing procedure of the actor network is as follows: The current state data is sequentially input into the embedding layer, the position encoding layer, the first encoding layer, the second encoding layer, the third encoding layer, the first full connection layer, the second full connection layer, the third full connection layer, the first dimension reduction layer, the fourth full connection layer, the second dimension reduction layer, the fifth full connection layer, and the output layer, and the output layer outputs the action under the current state; The critic network sequentially includes an embedding layer, a position encoding layer, a first encoding layer, a second encoding layer, a third encoding layer, a first full connection layer, a second full connection layer, a third full connection layer, a first dimension reduction layer, a fourth full connection layer, a second dimension reduction layer, a splicing layer, a fifth full connection layer, and an output layer; The specific processing procedure of the critic network is as follows: The state data at the current time is sequentially input into the embedding layer, the position encoding layer, the first encoding layer, the second encoding layer, the third encoding layer, the first full connection layer, the second full connection layer, the third full connection layer, the first dimension reduction layer, the fourth full connection layer, and the second dimension reduction layer, the second dimension reduction layer outputs the feature data, and the current time action corresponding to the state at the current time is input into the splicing layer, and then sequentially input into the fifth full connection layer and the output layer, and the output layer outputs the evaluation value data of the state at the current time and the current time action; Step 24, a target actor network and a target critic network are constructed; The target actor network sequentially includes an embedding layer, a position encoding layer, a first encoding layer, a second encoding layer, a third encoding layer, a first full connection layer, a second full connection layer, a third full connection layer, a first dimension reduction layer, a fourth full connection layer, a second dimension reduction layer, a fifth full connection layer, and an output layer; The specific processing procedure of the target actor network is as follows: s t+1 The data is sequentially input into the embedding layer, the position encoding layer, the first encoding layer, the second encoding layer, the third encoding layer, the first full connection layer, the second full connection layer, the third full connection layer, the first dimension reduction layer, the fourth full connection layer, the second dimension reduction layer, the fifth full connection layer, and the output layer, and the output layer outputs the next time state s t+1 The next action a t+1 Data; The target critic network sequentially includes an embedding layer, a position encoding layer, a first encoding layer, a second encoding layer, a third encoding layer, a first full connection layer, a second full connection layer, a third full connection layer, a first dimension reduction layer, a fourth full connection layer, a second dimension reduction layer, a splicing layer, a fifth full connection layer, and an output layer; The specific processing procedure of the target critic network is as follows: s t+1 The data is sequentially input into the embedding layer, the position encoding layer, the first encoding layer, the second encoding layer, the third encoding layer, the first full connection layer, the second full connection layer, the third full connection layer, the first dimension reduction layer, the fourth full connection layer, the second dimension reduction layer, and the second dimension reduction layer outputs the next moment action corresponding to the feature data of the next moment state and inputs the fifth full connection layer and the output layer in sequence after being input into the splicing layer, and the output layer outputs the target value data of the state action pair (s t+1 ,a t+1 ). Step 25, a loss function of the actor network is set; the specific procedure is as follows: The expression of the loss function of the actor network is as follows: where J μ is the loss function of the actor network, Q is the critic network, μ is the actor network, θ is the network parameter, θ μ is the parameter of the actor network, θ Q is the parameter of the critic network, ω t is the weight, t is the time step, B is the batch size, s t is the current time state; The loss function of the actor network is obtained as follows: The Q value obtained by the critic network has a dimension of Bx1, the cumulative expected return is obtained based on the Q value obtained by the critic network, the obtained cumulative expected return is sorted, that is, q qt = -Q(s t , μ(s t | θ μ )) θ Q ) q q1 <q q2 <...<q q64 where q qt is the cumulative expected return; Different weight values are assigned to the Q values according to the cumulative expected return from small to large: ω q1 > ω q2 >... > ω q64 where ω qt is the weight value corresponding to the Q value; Based on q qt And ω qt Loss function for constructing actor network: Step 26, a loss function of the critic network is set; the specific procedure is as follows: The expression of the loss function of the critic network is as follows: where J Q is the critic network loss function; r t is the reward value of performing the current action in the current state; done is a determination of whether the current state is an end state; Q' is the target critic network; μ' is the target actor network; γ is a decay function; s t+1 is the state at time t+1; a t is the action at the current time; θ μ' is the parameter of the target actor network; θ Q' is the parameter of the target critic network; The obtaining process of the critic network loss function is as follows: First, the critic network computes the evaluation value of the current state-action pair (s t ,a t ): q t = Q(s t , a t | θ Q ) wherein q t is the evaluation value at the current time, Q is the critic network, s t is the state at the current time, a t is the action at the current time, and θ Q is the parameter of the critic network. Through the target actor network, the action at the next time is obtained: a t+1 = μ'(s t+1 |θ μ' ) where a t+1 is the predicted action at the next time step by the target actor network, μ' is the target actor network, s t+1 is the state at the next time step, θ μ' is the parameter of the target actor network; The target value of the state-action pair (s t+1 ,a t+1 ) is calculated by the target critic network, combining the next time action, the next time state, the reward obtained in the current step, and whether it is a terminal state: y t = r t + γ(1 - done)Q'(s t+1 , a t+1 | θ Q' ) where y t is the target value of the state-action pair (s t+1 , a t+1 ), r t is the reward value of performing action a t in current state s t , Q' is the target critic network, and Θ Q' is the parameters of the target critic network. After the target value and the evaluation value are calculated, the target value and the evaluation value are subtracted: L t = y t - q t The resulting L t is B x 1 dimensional, and the loss function for the critic network is finally obtained:
2. The non-cooperative spacecraft active tracking method based on deep reinforcement learning according to claim 1, characterized in that: The reward function is constructed in step 3, and the trained actor network, critic network, target actor network and target critic network are obtained. The method comprises the following steps: Step 31, initializing the initial positions of the target spacecraft and the pursuit spacecraft; the specific process is as follows: The initial position of the target spacecraft is an orbit 7500 km away from the center of the earth; The position of the pursuit spacecraft is on and in a circle with the target spacecraft as the center and 1 km as the radius; Step 32, initializing the initial velocities of the target spacecraft and the pursuit spacecraft; the specific process is as follows: The initial velocity of the target spacecraft is set according to the following formula: Wherein, v tar is the initial speed of the target spacecraft under the universal gravitation at the initial position; r tar is the initial orbit radius of the target spacecraft, F g,tar is the universal gravitation received by the target spacecraft, m tar is the weight of the target spacecraft, X tar ,Y tar ,Z tar is the position coordinate of the target spacecraft in the geocentric inertial coordinate system; is the velocity unit vector of the target spacecraft at the current position, · is the dot product, is the speed of the target spacecraft, v x,tar is the speed of the target spacecraft in the x-axis of the geocentric inertial coordinate system, v x,tar to take negative values, v y,tar is the speed of the target spacecraft in the y-axis of the geocentric inertial coordinate system, v z,tar is the speed of the target spacecraft in the z-axis of the geocentric inertial coordinate system; The initial velocity of the pursuit spacecraft is set according to the following formula: wherein v c is the initial velocity of the pursuit spacecraft at the initial position under the universal gravitation; r c is the orbit radius of the pursuit spacecraft, F g,c is the universal gravitation received by the pursuit spacecraft, m c is the weight of the pursuit spacecraft, X c , Y c , Z c is the position coordinate of the pursuit spacecraft in the geocentric inertial coordinate system; is the velocity unit vector of the pursuit spacecraft at the current position, · is the dot product, is the velocity of the pursuit spacecraft, v x,c is the velocity of the pursuit spacecraft in the x-axis of the geocentric inertial coordinate system, v x,c is negative, v y,c is the velocity of the pursuit spacecraft in the y-axis of the geocentric inertial coordinate system, v z,c is the velocity of the pursuit spacecraft in the z-axis of the geocentric inertial coordinate system; Step 33, defining a reward function; Step 34, setting experience replay pool parameters, critic network and actor network hyperparameters; Step 35, action selection strategy; Step 36, extracting n continuous training samples from the experience replay pool as a group of training samples; Step 37, predicting the action value of the current time step by using the critic network; updating the actor network parameters and the target actor network parameters; Step 38, predicting the action of the current time step by using the actor network, estimating the action value of the next time step by using the critic network, and finally updating the critic network parameters and the target critic network parameters according to the time difference error; Step 39, repeating steps 33 to 38 until the trained actor network and critic network are obtained.
3. The non-cooperative spacecraft active tracking method based on deep reinforcement learning according to claim 2, characterized in that: The reward function is defined in step 33; the specific process is as follows: Let the position coordinates of the chaser spacecraft in the geocentric inertial coordinate system (X c ,Y c ,Z c ) = P c ; Let the position coordinates of the target spacecraft in the geocentric inertial coordinate system (X tar ,Y tar ,Z tar ) = P t ; The reward function is shown in the following formula: r = -(||P c - P t || - P e ) x f p where P e is the desired relative position between the chaser spacecraft and the target spacecraft, f p is the adjustment coefficient of the reward function; c is the constant term; and r is the reward value.
4. The non-cooperative spacecraft active tracking method based on deep reinforcement learning according to claim 3, characterized in that: In step 34, the experience replay pool parameters, critic network and actor network hyperparameters are set; the specific process is as follows: The total number of training scenes is 2000, the maximum step length of each scene is 1000 steps, the action random disturbance conforms to the Gaussian distribution, the initial standard deviation is 0.7, the final standard deviation is 0.1, the mean value is 0, and the reward decay coefficient is 0.95; The experience replay pool starts training after collecting 10000 groups of samples initially, with a total capacity of 50000, and the experience replay pool adopts a first-in, first-out strategy, with a sampling rate of (s t a t r t s t+1 D) as a group of samples; Wherein, D is a marker indicating whether the current scene is terminated.
5. The non-cooperative spacecraft active tracking method based on deep reinforcement learning according to claim 4, characterized in that: The action selection strategy in step 35 is as follows: The action selection strategy is divided into two cases; When the number of samples in the experience replay pool is less than or equal to 10, the action output by the random action generation network is used, and the random action generation network parameters are shared with the actor network parameters; The random action generation network structure is the same as the actor network, and the parameters are copied from the actor network at the beginning of each scene until the experience exceeds 10, and the random action generation network will not be used until the next scene starts; When the number of samples in the experience replay pool is greater than 10, the action output by the actor network is used, and the input of the actor network is a continuous state sequence.
6. The non-cooperative spacecraft active tracking method based on deep reinforcement learning according to claim 5, characterized in that: In step 37, the critic network is used to predict the action value of the current time step; updating the actor network parameters and the target actor network parameters; The specific process is as follows: Based on the loss function of the actor network, the actor network parameters are updated using the back propagation algorithm to minimize the loss function; The parameter updating for the target actor network is specifically as shown in the following formula: P μ' = (1 - τ) x μ + τ x μ' wherein P μ' is the network parameter of the target actor network, τ is the parameter ratio coefficient of the actor network and the target actor network, μ is the actor network, and μ' is the target actor network. The definition of τ is as shown in the following formula: T is the total number of steps.
7. The non-cooperative spacecraft active tracking method based on deep reinforcement learning according to claim 6, characterized in that: In the step 38, the actor network is used to predict the action at the current time step, the critic network is used to estimate the action value at the next time step, and finally the critic network parameters and the target critic network parameters are updated according to the time difference error; The specific process is as follows: Based on the loss function of the critic network, the critic network parameters are updated using the back propagation algorithm to minimize the loss function; The parameter updating for the target critic network is specifically as shown in the following formula: P Q' = (1 - τ') x Q + τ' x Q' wherein P Q' is the network parameter of the target critic network, τ' is the parameter ratio coefficient of the critic network and the target critic network, Q is the critic network, and Q' is the target critic network. The definition of τ' is as shown in the following formula: T is the total number of steps.
Citation Information
Patent Citations
Spacecraft multi-space debris collision avoidance autonomous decision-making method based on near-end strategy optimization
CN116125811A
Multi-spacecraft hunting and pursuit game decision-making method based on group independent learning strategy
CN118052290A