Multi-node intelligent collaborative guidance method for non-cooperative target attachment
Through the coordinated guidance architecture of main thrust plus compensated thrust and the multi-agent reinforcement learning algorithm, the problem of posture instability and thrust conflict during the attachment of non-cooperative targets is solved, and the effects of posture stabilization and fuel consumption are achieved.
Patent Information
- Application Number
- CN202310852435.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-12
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2043-07-12
AI Technical Summary
During the attachment process of non-cooperative targets, the node movements are coupled to each other, making it difficult to meet the constraints of the posture smoothness and the control thrust of multiple nodes may cause conflicts, resulting in increased fuel consumption.
The collaborative guidance architecture of main thrust plus compensated thrust is adopted, combined with the multi-agent reinforcement learning algorithm, the guidance parameters of each node are trained through the multi-agent system, and the compensating thrust is calculated using energy optimal control strategy, and a reward function and neural network loss function are constructed to meet the posture stability and control thrust constraints.
The flexible lander has achieved a smooth attitude during the non-cooperative target attachment process, reducing node thrust conflicts and fuel consumption, and improving the safety and accuracy of the attachment process.
Smart Images

Figure CN116620566B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a collaborative guidance method, in particular to a multi-node intelligent collaborative guidance method for non-cooperative target attachment, and belongs to the technical field of deep space exploration. Background Art
[0002] With the development of aerospace technology, the attachment and detection of non-cooperative space targets such as small celestial bodies has become a research focus. During the attachment process of small celestial bodies, due to the weak gravity of the small celestial bodies and the complex environmental disturbances, traditional rigid landers are at risk of rebound and overturning during landing. The flexible structure of the flexible lander can consume the residual kinetic energy during landing, and its surface configuration increases the contact area during landing, thereby avoiding the rebound and overturning of the lander and improving the reliability of the small celestial body attachment mission. The flexible lander adopts a three-node configuration and is connected by flexible materials. Thrusters and sensors are installed at each node. During the attachment process of small celestial bodies, in order to ensure the acquisition of navigation measurement information, the lander's attitude needs to be kept stable. However, environmental disturbances and navigation observation errors can easily lead to inconsistent states of the three nodes, causing the lander to flip, which places high demands on the guidance technology. Currently, research on single-target attachment guidance methods under complex multi-constraints is relatively mature. While optimal control-based guidance methods can meet the requirements for precise attachment, flexible lander guidance still faces the following difficulties: First, the motions of the flexible lander's nodes are interconnected and coupled, making it difficult to meet attitude stability constraints using a direct single-target guidance method. Second, controlling the thrust direction of multiple nodes can conflict, resulting in unnecessary fuel consumption. To achieve stable attachment for the flexible lander, considering the terminal state constraints, attitude stability constraints, and control thrust constraints, each lander node corresponds to an intelligent agent. Multi-agent reinforcement learning methods are used to train multiple agents for multi-node collaborative guidance. Summary of the Invention
[0003] In response to the problem of non-cooperative target attachment of flexible landers, the main purpose of the present invention is to provide a multi-node intelligent collaborative guidance method for non-cooperative target attachment, which adopts a collaborative guidance architecture of main thrust plus compensation thrust, and uses an energy optimal control strategy to calculate the main thrust to meet the terminal state constraints. In view of the characteristics of the multi-node collaborative attachment problem that conforms to the multi-agent Markov decision process, a multi-agent system is constructed to determine the multi-node compensation thrust guidance parameters. The multi-agent is trained through a multi-agent reinforcement learning algorithm. Each node gives guidance parameters according to the corresponding agent and calculates the compensation thrust. During the training process, a reward function is designed based on the three-node states to meet the attitude stability constraint and improve the attachment safety; a neural network loss function is designed based on the three-node actions to reduce node thrust conflicts and thrust saturation, reduce fuel consumption, and achieve precise attachment of the lander to the target point while maintaining the stability of the lander's attitude.
[0004] The non-cooperative target attachment multi-node intelligent collaborative guidance method disclosed in the present invention includes the following steps:
[0005] Step 1: To solve the problem of coordinated attachment of three nodes of the flexible lander, a coordinated guidance architecture of main thrust plus compensation thrust is adopted. The guidance instructions for each node determined based on the coordinated guidance architecture include two parts: the main thrust ensures that each node can achieve double zero attachment to the target, and the compensation thrust is used to keep the attitude of the lander stable during the attachment process and reduce the node thrust conflict. The main thrust is calculated using the zero control displacement deviation / zero control velocity deviation (ZEM / ZEV) method, and the compensation thrust is calculated using the rolling optimization energy optimal control strategy. The guidance parameters to be optimized in the energy optimal control strategy include K ri , K vi and t c , K ri , K vi and t c Confirmed through subsequent steps 2 and 3.
[0006] The specific implementation method of step one is:
[0007] The flexible lander adopts a three-node configuration, which is connected by flexible material wrapping. The three nodes are distributed in a centrally symmetrical form. In the landing point coordinate system O-XYZ, the position, velocity, attitude and angular velocity of the i-th node are expressed as r i 、v i ,q i 、ω i , the node dynamic equation is as follows:
[0008]
[0009] Among them, m is the node mass, g is the gravitational acceleration of the small celestial body surface, I is the node moment of inertia, and the symbol is Indicates quaternion direct multiplication, T i is the node control thrust, F ei is the flexible force on the node, M ei is the flexible moment acting on the node.
[0010] Construct the following collaborative guidance architecture
[0011] T i =T 0i +T ci (2)
[0012] Among them, T 0i is the main thrust of node i, which is used to control the node to achieve double zero adhesion, T ci is the compensation thrust of node i, which is used to keep the attitude of the lander stable during the attachment process and reduce the thrust conflict between nodes.
[0013] Define the flight time of the attachment process as t f , the current time is t, and the remaining flight time is
[0014] t go =t f -t (3)
[0015] Main thrust is calculated using the ZEM / ZEV method
[0016]
[0017] Among them, r fi and v fi are the target position and velocity of the i-th node respectively. The main thrust obtained can make the attachment process meet the terminal state constraints.
[0018] Define the lander centroid position vector r o and the node relative centroid position vector r oi , the overall centroid velocity vector v of the lander o and the node relative centroid velocity vector v oi , the unit normal vector n of the lander plane
[0019]
[0020] According to the three-node landing target position r f1 、r f2 、r f3 , the target position of the three nodes relative to the centroid can be calculated. During the landing process, by applying compensation thrust to keep the three nodes relative to the target position, the lander attitude can be kept stable without disturbance. The compensation thrust is calculated using the optimal control strategy of rolling optimization, and the rolling optimization time is defined as t c , t c To be smaller than t f , construct the energy optimal control problem to solve the compensating thrust.
[0021]
[0022] Among them, r fo and v fo is the target position and velocity of the lander centroid, K ri With K vi The guidance parameter is used to compensate for the thrust, and its nominal value is [6,6,6] T and [2,2,2] T The resulting compensation thrust satisfies the attitude stability constraint.
[0023] Considering the existence of navigation observation errors and environmental interference during the attachment process, and the need to reduce the fuel loss caused by multi-node thrust conflicts, Kri , K vi and t c Set as the guidance parameter to be optimized.
[0024] Step 2: Considering the characteristics of the multi-node cooperative attachment problem that conforms to the multi-agent Markov decision process, a multi-agent system is constructed to determine the multi-node compensatory thrust guidance parameters. Each lander node corresponds to an agent, and an agent is composed of a set of actor-critic neural networks. Multiple nodes correspond to agents to form a multi-agent system. Each set of actor-critic neural networks takes the lander node state given by the lander dynamics as input and the node compensation thrust guidance parameter K ri , K vi and t c As output, construct a system for determining the guidance parameter K ri , K vi and t c The multiple sets of guidance parameters output by the multi-agent system act together on the lander dynamics model to achieve control of all node states and solve the problem of multi-node motion coupling.
[0025] The specific implementation method of step 2 is:
[0026] In the three-node cooperative attachment problem, considering that the main thrust differences between the nodes are small, the lander's attitude is only related to the flexible force, the compensation thrust, and the attitude at the previous moment. Therefore, the problem can be formulated as a Markov decision process. The Markov decision process consists of a set of interacting objects, namely the agent and the environment. The environment includes the flexible lander dynamics, state space, and action space shown in Equation (1).
[0027] Select state space s and action space a
[0028]
[0029] Among them, r other is the position vector of other nodes, the state space is 22-dimensional, and the action space is 10-dimensional.
[0030] The agent is a neural network whose input is state-space variables provided by the environment and whose output is action-space variables. The agent determines guidance parameters based on node states, thereby calculating compensation thrust. This compensation thrust acts on the environment, changing the state space. In a multi-node collaborative attachment mission, each lander node corresponds to an agent, which is composed of an actor-critic neural network. Multiple nodes correspond to agents, forming a multi-agent system. Both the actor network and the critic network consist of an input layer, an output layer, and three intermediate layers. The actor network has an input layer with 22 neurons, an output layer with 10 neurons, and each intermediate layer with 64 neurons. The activation function is the ReLU function. The actor network has an input layer with 32 neurons, an output layer with 1 neuron, and each intermediate layer with 64 neurons. The activation function is the ReLU function.
[0031] The actor network is the policy network π w , whose input is state and output is action, used to fit the policy function π *
[0032] π w (s,a)→π * (s,a) (8)
[0033] Where w is a neural network parameter. The policy function is a probability density function, that is, the probability of executing action a in state s. Set a suitable reward function r(s,a), and calculate the reward value based on the state and the selected action to evaluate the quality of action a in state s. The critic network input is the state and the action output is the reward value, which is used to fit the value function Q π (s,a), the value function represents the quality of action a in state s when using strategy π, and is used to evaluate the strategy.
[0034] Step 3: Construct a reward function to make the lander meet the attitude stability constraint, and improve the attachment safety through the reward function; construct a neural network loss function to make the lander multi-node meet the control thrust constraint, reduce node thrust conflict and thrust saturation, and reduce fuel consumption. Use the multi-agent reinforcement learning algorithm to train the multi-agent constructed in step 2. In the process of training the multi-agent constructed in step 2, the output node guidance parameter K is obtained through continuous interactive training between the multi-agent system and the flexible lander dynamics model. ri , K vi and t c 's intelligent agent.
[0035] The specific implementation method of step three is:
[0036] Based on the Markov decision process, the multi-agent reinforcement learning algorithm is used to train the multi-agent constructed in step 2. During the training process, the initial state is randomly selected, and the Euler angles θ, θ, and θ of the lander's attitude can be calculated from the positions of the three nodes. ψ, design reward function based on smooth landing requirements
[0037]
[0038] Among them, θ f 、 ψ f is the desired attitude Euler angle, which can be calculated from the relative target position. When the lander attitude deviates from the desired attitude, a penalty value related to the deviation value is generated.
[0039] During the continuous interaction between the multi-agent system and the environment, the time data, state, action and corresponding reward value are stored in the experience pool. After a set number of interactions, the neural network parameters of multiple agents are updated. First, the data in the experience pool is randomly sampled, the number of samples is N, and the network loss function is calculated based on the sampled data.
[0040]
[0041] in, and are the loss functions of the critic network and actor network in the agent corresponding to node i, respectively. The smaller the expected loss function, the better. and is the state and action of the i-th node at time t, U is the set of time data of the collected samples, and γ is the learning rate.
[0042] Since the critic network is used to evaluate the strategy, the strategy function can be affected by designing its loss function, that is, changing the probability density function of action selection. Considering the posture stability requirement of the three-node collaborative attachment task, the posture flip loss term is defined
[0043]
[0044] Where α is the lander inclination angle, or the angle between the lander's normal vector and the target normal vector. This angle is used to evaluate the degree of lander rollover. A smaller α indicates a smaller L1 loss term. When the lander's attitude is unstable, the lander inclination angle α is greater than 0. At this point, the difference in normal compensation thrust at different nodes is large, and the smaller the loss term, the more likely it is that a larger normal compensation thrust is needed to quickly restore the attitude to a stable state.
[0045] Considering that the thrust conflict between nodes will cause unnecessary fuel consumption, the thrust conflict loss term is defined
[0046]
[0047] When the control thrust directions are inconsistent, a loss term is generated to reduce thrust conflict.
[0048] Considering the node thrust amplitude constraint, define the thrust constraint loss term
[0049]
[0050] Among them, T max is the maximum thrust, and a loss term occurs when the node control thrust reaches saturation.
[0051] The loss term designed according to the three-node cooperative attachment requirement is:
[0052] L=k1L1+k2L2+k3L3, k1,k2,k3>0 (14)
[0053] Among them, k1, k2, and k3 are parameter items.
[0054] The critic network loss function is designed based on this
[0055]
[0056] The loss function calculated from the sampled data is used to calculate the gradient of the neural network parameters. This gradient descent strategy is then used to update the network parameters, completing the multi-agent training process through continuous iteration. The input to each node's corresponding agent is the node's state and the positions of other nodes, and the output is the guidance parameters for compensating thrust for that node.
[0057] Step 4: During the landing process, each agent outputs the guidance parameter K of the node's compensation thrust according to the node's motion state and the position information of other nodes. ri , K vi and t c Combined with the energy-optimal control strategy for rolling optimization constructed in step 1, the corresponding node compensation thrust is calculated. The main thrust is used to ensure that the attachment process meets the terminal state constraints. The compensation thrust is used to correct the main thrust based on the main thrust, ensuring that the lander meets the attitude stability constraints, improving the safety of the attachment process. At the same time, it reduces node thrust conflicts, meets the control thrust constraints, reduces fuel consumption during the attachment process, and achieves precise attachment of the lander to the target point while maintaining a stable attitude.
[0058] Beneficial effects:
[0059] 1. The present invention discloses a multi-node intelligent collaborative guidance method for non-cooperative target attachment. This method addresses the problem of three-node collaborative attachment of a flexible lander, employing a collaborative guidance architecture combining main thrust and compensating thrust. Determining guidance instructions for each node based on this collaborative guidance architecture involves two steps: calculating the main thrust using the zero-control displacement deviation / zero-control velocity deviation method to ensure that each node can achieve double-zero attachment to the target; and calculating the compensating thrust using a rolling optimization energy-optimal control strategy to maintain a stable lander attitude during attachment and reduce node thrust conflicts.
[0060] 2. The non-cooperative target attachment multi-node intelligent collaborative guidance method disclosed in this invention is based on the characteristics of the multi-agent Markov decision process for the multi-node collaborative attachment problem. A multi-agent system is constructed to determine the multi-node compensatory thrust guidance parameters. Each lander node corresponds to an agent, and an agent is composed of a set of actor-critic neural networks. Multiple nodes correspond to agents to form a multi-agent system. The agent is used to determine the guidance parameter K of the compensatory thrust. ri , K vi and t c , multiple sets of guidance parameters output by the multi-agent system act together on the lander dynamics to achieve control of all node states and solve the problem of multi-node motion involvement and coupling.
[0061] 3. The non-cooperative target attachment multi-node intelligent collaborative guidance method disclosed in the present invention constructs a reward function for making the lander meet the attitude stability constraint, and improves the attachment safety through the reward function; constructs a neural network loss function for making the lander multi-node meet the control thrust constraint, and quickly restores the lander to a stable attitude through the attitude flip loss term, reduces the three-node thrust conflict through the thrust conflict loss term, and reduces the node thrust saturation through the thrust constraint loss term, thereby reducing fuel consumption. The multi-agent reinforcement learning algorithm is used to train the multi-agent, and the output node guidance parameter K is obtained through continuous interactive training between the multi-agent system and the flexible lander dynamics. ri , K vi and t c 's intelligent agent.
[0062] 4. The non-cooperative target attachment multi-node intelligent collaborative guidance method disclosed in the present invention, during the landing process of the lander, each intelligent agent outputs the guidance parameter K of the node compensation thrust according to the motion state of the node and the position information of other nodes ri , K vi and t c , to ensure the rapidity and real-time performance of the execution process. The guidance parameter K ri , K vi and t cCombined with the energy optimal control strategy of rolling optimization, the corresponding node compensation thrust is calculated. The main thrust is corrected based on the compensation thrust, so that the lander meets the attitude stability constraint and improves the safety of the attachment process. At the same time, the node thrust conflict is reduced, the control thrust constraint is met, the fuel consumption of the attachment process is reduced, and the precise attachment of the lander to the target point is achieved while maintaining the stability of the lander's attitude. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 Flowchart of the multi-node intelligent collaborative guidance method for attaching non-cooperative targets;
[0064] Figure 2 Schematic diagram of a multi-agent system;
[0065] Figure 3 This is the three-node flight trajectory of the flexible lander;
[0066] Figure 4 is the inclination curve of the flexible lander;
[0067] Figure 5 is the main thrust curve of the three nodes of the flexible lander;
[0068] Figure 6 Compensate thrust curve for the three nodes of the flexible lander;
[0069] Figure 7 Control thrust curve for the three nodes of the flexible lander. DETAILED DESCRIPTION
[0070] In order to better illustrate the purpose and advantages of the present invention, the invention is further described below in conjunction with embodiments and corresponding drawings.
[0071] In order to verify the feasibility of the method, a simulation of the flexible lander multi-node intelligent collaborative guidance method is carried out using the small celestial body flexible attachment mission as an example. The flexible lander node mass m = 333 kg, and the node moment of inertia I = [15.51, 0, 0; 0, 15.51, 0; 0, 0, 21.08] kg·m 2 , maximum node thrust T max =25N, gravitational acceleration of small celestial body g=[0,0,-0.001]m / s 2The maximum allowable inclination angle of the lander is 10°. The initial positions of the three nodes are selected as [30.6, 10, 50] m, [29.7, 10.52, 50] m, and [29.7, 9.48, 50] m, respectively. The initial velocities of the three nodes are [-0.1, 0, 0] m / s, [-0.1, 0, 0.05] m / s, and [-0.1, 0, 0.1] m / s, respectively. The target positions of the three nodes are [0.6, 0, 0] m, [-0.3, 0.52, 0] m, and [-0.3, -0.52, 0] m, respectively. The target velocities of the three nodes are all [0, 0, 0] m / s, and the flight time of the lander is 100 s.
[0072] like Figure 1 As shown, the non-cooperative target attachment multi-node intelligent collaborative guidance method disclosed in this embodiment is specifically implemented in the following steps:
[0073] Step 1: To solve the problem of coordinated attachment of three nodes of the flexible lander, a coordinated guidance architecture of main thrust plus compensation thrust is adopted. The guidance instructions for each node determined based on the coordinated guidance architecture include two parts: the main thrust ensures that each node can achieve double zero attachment to the target, and the compensation thrust is used to keep the attitude of the lander stable during the attachment process and reduce the node thrust conflict. The main thrust is calculated using the zero control displacement deviation / zero control velocity deviation (ZEM / ZEV) method, and the compensation thrust is calculated using the rolling optimization energy optimal control strategy. The guidance parameters to be optimized in the energy optimal control strategy include K ri , K vi and t c , K ri , K vi and t c Confirmed through subsequent steps 2 and 3.
[0074] The specific implementation method of step one is:
[0075] The flexible lander adopts a three-node configuration, which is connected by flexible material wrapping. The three nodes are distributed in a centrally symmetrical form. In the landing point coordinate system O-XYZ, the position, velocity, attitude and angular velocity of the i-th node are expressed as r i 、v i ,q i 、ω i , the node dynamic equation is as follows:
[0076]
[0077] Among them, m is the node mass, g is the gravitational acceleration of the small celestial body surface, I is the node moment of inertia, and the symbol is Indicates quaternion direct multiplication, T i is the node control thrust, F ei is the flexible force on the node, M eiis the flexible moment acting on the node.
[0078] Construct the following collaborative guidance architecture
[0079] T i =T 0i +T ci (17)
[0080] Among them, T 0i is the main thrust of node i, which is used to control the node to achieve double zero adhesion, T ci is the compensation thrust of node i, which is used to keep the attitude of the lander stable during the attachment process and reduce the thrust conflict between nodes.
[0081] Define the flight time of the attachment process as t f , the current time is t, and the remaining flight time is
[0082] t go =t f -t (18)
[0083] Main thrust is calculated using the ZEM / ZEV method
[0084]
[0085] Among them, r fi and v fi are the target position and velocity of the i-th node respectively. The main thrust obtained can make the attachment process meet the terminal state constraints.
[0086] Define the lander centroid position vector r o and the node relative centroid position vector r oi , the overall centroid velocity vector v of the lander o and the relative centroid velocity vector v oi , the unit normal vector n of the lander plane
[0087]
[0088] According to the three-node landing target position r f1 、r f2 、r f3 , the target position of the three nodes relative to the centroid can be calculated. During the landing process, by applying compensation thrust to keep the three nodes relative to the target position, the lander attitude can be kept stable without disturbance. The compensation thrust is calculated using the optimal control strategy of rolling optimization, and the rolling optimization time is defined as t c , t c To be smaller than t f , construct the energy optimal control problem to solve the compensating thrust.
[0089]
[0090] Among them, r fo and v fo is the target position and velocity of the lander centroid, K ri With K vi The guidance parameter is used to compensate for the thrust, and its nominal value is [6,6,6] T and [2,2,2] T The resulting compensatory thrust can satisfy the attitude stability constraint.
[0091] Considering the existence of navigation observation errors and environmental interference during the attachment process, and the need to reduce the fuel loss caused by multi-node thrust conflicts, K ri , K vi and t c Set as the guidance parameter to be optimized.
[0092] Step 2: Considering the characteristics of the multi-node cooperative attachment problem that conforms to the multi-agent Markov decision process, a multi-agent system is constructed to determine the multi-node compensatory thrust guidance parameters. Each lander node corresponds to an agent, and an agent is composed of a set of actor-critic neural networks. Multiple nodes correspond to agents to form a multi-agent system. Each set of actor-critic neural networks takes the lander node state given by the lander dynamics as input and the node compensation thrust guidance parameter K ri , K vi and t c As output, construct a system for determining the guidance parameter K ri , K vi and t c The multiple sets of guidance parameters output by the multi-agent system act together on the lander dynamics model to achieve control of all node states and solve the problem of multi-node motion coupling.
[0093] The specific implementation method of step 2 is:
[0094] In the three-node collaborative attachment problem, considering that the main thrust differences between the nodes are small, the lander's attitude is only related to the flexible force, the compensation thrust, and the attitude at the previous moment. Therefore, the problem can be formulated as a Markov decision process. The Markov decision process consists of a set of interacting objects, namely the agent and the environment. The environment includes the flexible lander dynamics, state space, and action space shown in Equation (16).
[0095] Select state space s and action space a
[0096]
[0097] Among them, r otheris the position vector of other nodes, the state space is 22-dimensional, and the action space is 10-dimensional.
[0098] The agent is a neural network whose input is the state space variables provided by the environment and whose output is the action space variables. The agent determines the guidance parameters according to the node state, thereby calculating the compensation thrust, which acts on the environment and changes the state space. In the multi-node collaborative attachment task, each lander node corresponds to an agent, which is composed of an actor-critic neural network. Multiple nodes correspond to agents, such as Figure 2 The multi-agent system shown in Figure 2 is shown in Figure 2. Both the actor network and the critic network consist of an input layer, an output layer, and three intermediate neural networks. The actor network has 22 neurons in the input layer, 10 neurons in the output layer, and 64 neurons in each intermediate layer. The activation function is the ReLU function. The actor network has 32 neurons in the input layer, 1 neuron in the output layer, and 64 neurons in each intermediate layer. The activation function is the ReLU function.
[0099] The actor network is the policy network π w , whose input is state and output is action, used to fit the policy function π *
[0100] π w (s,a)→π * (s,a) (23)
[0101] Where w is a neural network parameter. The policy function is a probability density function, that is, the probability of executing action a in state s. Set a suitable reward function r(s,a), and calculate the reward value based on the state and the selected action to evaluate the quality of action a in state s. The critic network input is the state and the action output is the reward value, which is used to fit the value function Q π (s,a), the value function represents the quality of action a in state s when using strategy π, and is used to evaluate the strategy.
[0102] Step 3: Construct a reward function to make the lander meet the attitude stability constraint, and improve the attachment safety through the reward function; construct a neural network loss function to make the lander multi-node meet the control thrust constraint, reduce node thrust conflict and thrust saturation, and reduce fuel consumption. Use the multi-agent reinforcement learning algorithm to train the multi-agent constructed in step 2. In the process of training the multi-agent constructed in step 2, the output node guidance parameter K is obtained through continuous interactive training between the multi-agent system and the flexible lander dynamics model. ri , K vi and t c 's intelligent agent.
[0103] The specific implementation method of step three is:
[0104] Based on the Markov decision process, the multi-agent reinforcement learning algorithm is used to train the multi-agent constructed in step 2. During the training process, the initial state is randomly selected, and the Euler angles θ, θ, and θ of the lander's attitude can be calculated from the positions of the three nodes. ψ, design reward function based on smooth landing requirements
[0105]
[0106] Among them, θ f 、 ψ f is the desired attitude Euler angle, which is zero in this embodiment. When the lander attitude deviates from the desired attitude, a penalty value related to the deviation value is generated.
[0107] During the continuous interaction between the neural network and the environment, the time data, state, action and corresponding reward value are stored in the experience pool. After every 1000 generations of interaction, the neural network parameters are updated. First, the data in the experience pool is randomly sampled, the number of samples is N, and the network loss function is calculated based on the sampled data
[0108]
[0109] in, and are the loss functions of the critic network and the actor network respectively. It is expected that the smaller the loss function, the better. and is the state and action of the i-th node at time t, U is the set of time data of the collected samples, and γ is the learning rate. The number of samples N = 256 and the learning rate γ = 0.95 are selected.
[0110] Since the critic network is used to evaluate the strategy, the strategy function can be affected by designing its loss function, that is, changing the probability density function of action selection. Considering the posture stability requirement of the three-node collaborative attachment task, the posture flip loss term is defined
[0111]
[0112] Where α is the lander inclination angle, or the angle between the lander's normal vector and the target normal vector. This angle is used to evaluate the degree of lander rollover. A smaller α indicates a smaller L1 loss term. When the lander's attitude is unstable, the lander inclination angle α is greater than 0. At this point, the difference in normal compensation thrust at different nodes is large, and the smaller the loss term, the more likely it is that a larger normal compensation thrust is needed to quickly restore the attitude to a stable state.
[0113] Considering that the thrust conflict between nodes will cause unnecessary fuel consumption, the thrust conflict loss term is defined
[0114]
[0115] When the control thrust directions are inconsistent, a loss term is generated to reduce thrust conflict.
[0116] Considering the node thrust amplitude constraint, define the thrust constraint loss term
[0117]
[0118] Among them, T max is the maximum thrust, and a loss term occurs when the node control thrust reaches saturation.
[0119] The loss term designed according to the three-node cooperative attachment requirement is:
[0120] L=k1L1+k2L2+k3L3, k1,k2,k3>0 (29)
[0121] Among them, k1, k2, and k3 are parameter items, and k1=1, k2=0.01, and k3=0.1 are selected.
[0122] The critic network loss function is designed based on this
[0123]
[0124] The loss function calculated from the sampled data is used to calculate the gradient of the neural network parameters. The network parameters are updated using a gradient descent strategy, completing the agent training process through 100,000 iterations. The input to each node in the agent is the node's state and the positions of other nodes, and the output is the guidance parameter for compensating thrust for that node.
[0125] Step 4: During the landing process, each agent outputs the guidance parameter K of the node's compensation thrust according to the node's motion state and other node position information. ri , K vi and t c Combined with the energy-optimal control strategy for rolling optimization constructed in step 1, the corresponding node compensation thrust is calculated. The main thrust is used to ensure that the attachment process meets the terminal state constraints. The compensation thrust is used to correct the main thrust based on the main thrust, ensuring that the lander meets the attitude stability constraints, improving the safety of the attachment process. At the same time, it reduces node thrust conflicts, meets the control thrust constraints, reduces fuel consumption during the attachment process, and achieves precise attachment of the lander to the target point while maintaining a stable attitude.
[0126] The specific implementation method of step four is:
[0127] Each node calculates the main thrust of the node using the ZEM / ZEV guidance method based on its own position and speed. The intelligent agent corresponding to each node determines the guidance parameter K of the node's compensation thrust based on the node's motion state and other node position information. ri , K vi and t c , and the node compensation thrust is calculated by combining the energy optimal control method of rolling optimization. Considering that the flexible lander dynamic parameters have a random error of ±10%; there is a random error of ±10% when the node obtains other node information; there is a disturbance in the actuator, and the actual thrust is
[0128]
[0129] in, is the thrust disturbance of the i-th node at time t, which is a random number in the range of ±1N.
[0130] Simulation is carried out under the above error and disturbance conditions. Figure 3 This is the flight trajectory of the lander. It can be seen that under the combined effect of the main thrust and the compensation thrust, the lander can accurately reach the target point, and the attachment process satisfies the terminal state constraint. Figure 4 is the lander inclination curve. When the initial velocities of the three nodes are inconsistent, this method can keep the lander's attitude stable and improve the safety of the attachment process. Figure 5 、 6 , 7 are the main thrust curve, compensation thrust curve, and control thrust amplitude curve of the three nodes of the flexible lander, respectively. The thrust conflicts among the three nodes are small, and all meet the thrust amplitude constraints.
[0131] The above specific description further illustrates the purpose, technical solutions and beneficial effects of the invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A non-cooperative target attachment multi-node intelligent collaborative guidance method, characterized by: The following steps are included: Step 1: To address the three-node coordinated attachment problem of the flexible lander, a collaborative guidance architecture combining main thrust and compensation thrust is adopted. The guidance instructions for each node determined based on the collaborative guidance architecture include two parts: main thrust ensures that each node can achieve double-zero attachment to the target, and compensation thrust is used to maintain a stable lander attitude during attachment and reduce node thrust conflicts. The main thrust is calculated using the zero displacement deviation / zero velocity deviation ZEM / ZEV method, and the compensation thrust is calculated using the rolling optimization energy optimal control strategy; the guidance parameters to be optimized in the energy optimal control strategy include K ri , K vi and t c , K ri , K vi and t c , determined through subsequent steps 2 and 3; Step 2: Considering the characteristics of the multi-node cooperative attachment problem that conforms to the multi-agent Markov decision process, a multi-agent system is constructed to determine the multi-node compensatory thrust guidance parameters. Each lander node corresponds to an agent, and an agent is composed of a group of actor-critic neural networks. Multiple nodes correspond to agents to form a multi-agent system. Each group of actor-critic neural networks takes the lander node state given by the lander dynamics as input and the guidance parameter K of the node compensatory thrust as the input. ri , K vi and t c As output, construct a system for determining the guidance parameter K ri , K vi and t c The multiple sets of guidance parameters output by the multi-agent system act together on the lander dynamics model to achieve control of all node states and solve the problem of multi-node motion coupling; Step 3: Construct a reward function to make the lander meet the attitude stability constraint, and improve the attachment safety through the reward function; construct a neural network loss function to make the lander multi-node meet the control thrust constraint, reduce node thrust conflict and thrust saturation, and reduce fuel consumption; use the multi-agent reinforcement learning algorithm to train the multi-agent constructed in step 2; in the process of training the multi-agent constructed in step 2, the output node guidance parameter K is obtained through continuous interactive training between the multi-agent system and the flexible lander dynamics. ri , K vi and t c 's intelligent agent; Step 4: During the landing process, each agent outputs the guidance parameter K of the node's compensation thrust according to the node's motion state and the position information of other nodes. ri , K vi and t c , combined with the energy optimal control strategy of rolling optimization constructed in step 1, the corresponding node compensation thrust is calculated; the main thrust is used to ensure that the attachment process meets the terminal state constraint, and the main thrust is corrected on the basis of the main thrust through the compensation thrust, so that the lander meets the attitude stability constraint, improves the safety of the attachment process, and at the same time reduces the node thrust conflict, meets the control thrust constraint, reduces the fuel consumption of the attachment process, and realizes the precise attachment of the lander to the target point while maintaining the stability of the lander attitude.
2. The non-cooperative target attachment multi-node intelligent collaborative guidance method according to claim 1, characterized in that: The specific implementation method of step one is: The flexible lander adopts a three-node configuration, which is connected by flexible material wrapping. The three nodes are distributed in a centrally symmetrical form. In the landing point coordinate system O-XYZ, the position, velocity, attitude and angular velocity of the i-th node are expressed as r i 、v i ,q i 、ω i , the node dynamic equation is as follows: Among them, m is the node mass, g is the gravitational acceleration of the small celestial body surface, I is the node moment of inertia, and the symbol is Indicates quaternion direct multiplication, T i is the node control thrust, F ei is the flexible force on the node, M ei is the flexible moment acting on the node; Construct the following collaborative guidance architecture T i =T 0i +T ci (2) Among them, T 0i Main thrust, used to control the node to achieve double zero adhesion, T ci To compensate for the thrust, it is used to keep the attitude of the lander stable during the attachment process and reduce the thrust conflict between nodes; Define the flight time of the attachment process as t f , the current time is t, and the remaining flight time is t go =t f -t (3) Main thrust is calculated using the ZEM / ZEV method Among them, r fi and v fi are the target position and velocity of the i-th node respectively; the main thrust obtained from this can make the attachment process meet the terminal state constraint; Define the lander centroid position vector r o and the node relative centroid position vector r oi , the overall centroid velocity vector v of the lander o and the relative centroid velocity vector v oi , the unit normal vector n of the lander plane According to the three-node landing target position r f1 、r f2 、r f3 , the target position of the three nodes relative to the centroid can be calculated. During the landing process, the three nodes are kept in the target position relative to each other by applying compensation thrust, which can ensure the stability of the lander attitude without disturbance. The compensation thrust is calculated using the optimal control strategy of rolling optimization, and the rolling optimization time is defined as t c , t c To be smaller than t f , construct the energy optimal control problem to solve the compensating thrust; Among them, r fo and v fo is the target position and velocity of the lander centroid, K ri With K vi The guidance parameter is used to compensate for the thrust, and its nominal value is [6,6,6] T and [2,2,2] T ;The resulting compensation thrust satisfies the attitude stability constraint; Considering the existence of navigation observation errors and environmental interference during the attachment process, and the need to reduce the fuel loss caused by multi-node thrust conflicts, K ri , K vi and t c Set as the guidance parameter to be optimized.
3. The non-cooperative target attachment multi-node intelligent collaborative guidance method according to claim 2, characterized in that: The specific implementation method of step 2 is: In the three-node cooperative attachment problem, considering that the difference in the main thrust of the nodes is small, the attitude of the lander is only related to the flexible force, the compensation thrust, and the attitude at the previous moment. Therefore, the problem can be constructed as a Markov decision process. The Markov decision process contains a set of interacting objects, namely the agent and the environment. The environment includes the flexible lander dynamics, state space, and action space shown in formula (1). Select state space s and action space a Among them, r other is the position vector of other nodes, the state space is 22-dimensional, and the action space is 10-dimensional; The agent is a neural network whose input is the state space variables provided by the environment and whose output is the action space variables. The agent determines the guidance parameters based on the node state, thereby calculating the compensation thrust, which acts on the environment and changes the state space. In the multi-node collaborative attachment mission, each lander node corresponds to an agent, and the agent is composed of an actor-critic neural network. Multiple nodes correspond to agents to form a multi-agent system. The actor network and the critic network are both composed of an input layer, an output layer, and a three-layer intermediate layer neural network. The actor network input layer contains 22 neurons, the output layer contains 10 neurons, and each intermediate layer contains 64 neurons. The activation function is the ReLU function. The actor network input layer contains 32 neurons, the output layer contains 1 neuron, and each intermediate layer contains 64 neurons. The activation function is the ReLU function. The actor network is the policy network π w , whose input is state and output is action, used to fit the policy function π * p w (s,a)→π * (s,a) (8) Where w is the neural network parameter; the policy function is a probability density function, that is, the probability of executing action a in state s; set a suitable reward function r(s,a), and calculate the reward value based on the state and the selected action to evaluate the quality of action a in state s; the critic network input is the state and action output is the reward value, which is used to fit the value function Q π (s,a), the value function represents the quality of action a in state s when using strategy π, and is used to evaluate the strategy.
4. The non-cooperative target attachment multi-node intelligent collaborative guidance method according to claim 3, characterized in that: The specific implementation method of step three is: Based on the Markov decision process, the multi-agent reinforcement learning algorithm is used to train the multi-agent constructed in step 2. During the training process, the initial state is randomly selected, and the Euler angles θ, θ, and θ of the lander's attitude can be calculated from the positions of the three nodes. ψ, design reward function based on smooth landing requirements Among them, θ f 、 ψ f is the desired attitude Euler angle, calculated relative to the target position; when the lander attitude deviates from the desired attitude, a penalty value related to the deviation value is generated; In the process of continuous interaction between the multi-agent system and the environment, the time data, state, action and corresponding reward value are stored in the experience pool; after a set number of interactions, the neural network parameters of multiple agents are updated; the data in the experience pool are randomly sampled, the number of samples is N, and the network loss function is calculated based on the sampled data in, and are the loss functions of the critic network and actor network in the agent corresponding to node i, respectively. The smaller the expected loss function, the better. and is the state and action of the i-th node at time t, U is the set of time data of the collected samples, and γ is the learning rate; Since the critic network is used to evaluate the strategy, the strategy function is affected by designing its loss function, that is, changing the probability density function of action selection; considering the posture stability requirement of the three-node collaborative attachment task, the posture flip loss term is defined Among them, α is the lander inclination angle, that is, the angle between the lander normal vector and the target normal vector, which is used to evaluate the degree of lander rollover. The smaller α is, the smaller the L1 loss term is. When the lander attitude is unstable, the lander inclination angle α>0. At this time, the difference in normal compensation thrust at different nodes is large, and the loss term is smaller. That is, it is expected that a larger normal compensation thrust will be used to achieve rapid restoration of attitude stability. Considering that the thrust conflict between nodes will cause unnecessary fuel consumption, the thrust conflict loss term is defined When the control thrust directions are inconsistent, a loss term is generated to reduce thrust conflict; Considering the node thrust amplitude constraint, define the thrust constraint loss term Among them, T max is the maximum thrust, and a loss term is generated when the node control thrust reaches saturation; The loss term designed according to the three-node cooperative attachment requirement is: L=k1L1+k2L2+k3L3,k1,k2,k3>0 (14) Among them, k1, k2, and k3 are parameter items; The critic network loss function is designed based on this The loss function is calculated based on the sampled data, and the gradient of the neural network parameters is calculated. The gradient descent strategy is used to update the network parameters, and the training process of the multi-agent is completed through continuous iteration. The input of each node corresponding to the agent is the state of this node and the position of other nodes, and the output is the guidance parameter of the compensation thrust of this node.
Citation Information
Patent Citations
State estimation method for flexible attachment system in weak gravitation environment
CN113022898A
Non-cooperative target flexible attachment multi-node fusion estimation method
CN113408623A