Intelligent control method for asteroid flexible probe based on deep reinforcement learning SAC algorithm
By constructing an intelligent control method for flexible probes using the deep reinforcement learning SAC algorithm, the problem of collision bounce when traditional probes land on asteroid surfaces is solved, achieving high-precision and stable attachment of flexible probes and improving the success rate of asteroid exploration missions.
Patent Information
- Application Number
- CN202310204028.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-06
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2043-03-06
AI Technical Summary
Traditional rigid-body probes are prone to collisions and bounces when landing on asteroid surfaces, leading to mission failure. The complex deformation of flexible probes increases the difficulty of control, and existing intelligent technologies have limited development in the field of asteroid exploration.
A flexible detector structural model is constructed using the deep reinforcement learning SAC algorithm. The weak gravitational field of the asteroid is described by the second-order gravitational potential function. An attitude-orbit coupled dynamic model is established. An actor-critic neural network is used for interactive training. A reward function and a cost function are designed. An intelligent controller is constructed to achieve stable attachment of the flexible detector.
It has achieved high-precision and stable attachment of flexible probes on asteroids, with high control accuracy, strong anti-interference ability and robustness, effectively copes with the complex deformation of flexible materials, and improves the mission success rate.
Smart Images

Figure CN116400589B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to an asteroid flexible probe intelligent control method based on a deep reinforcement learning SAC (Soft Actor-Critic) algorithm and belongs to the field of spacecraft guidance and control. BACKGROUND
[0002] Asteroid exploration is a key research direction in the field of space exploration. The analysis of the composition of asteroids helps scientists study major scientific issues such as the origin of the universe. The development of asteroid probes also helps to promote the progress of key scientific technologies such as cosmic resource development and interstellar travel. Therefore, in recent years, many countries have implemented asteroid exploration missions. Among them, the asteroid sampling and returning missions of Europe, the United States and Japan are representative. The Philae asteroid probe of the European Space Agency planned to land on the 67P comet, but due to equipment failure, multiple bounces occurred during landing, and the landing position deviated from the target position of the task, so the task only achieved partial success. Japan successfully implemented two asteroid sampling and returning missions using the Hayabusa-1 and Hayabusa-2 asteroid probes. The Osiris-Rex probe of the United States also successfully sampled on the Bennu asteroid. Compared with flyby and orbit exploration, landing on the surface of an asteroid has higher scientific value for studying asteroids. However, due to the irregular weak gravitational field and complex topography of asteroids, traditional rigid probes are prone to collision with the surface of the asteroid and escape by bouncing when landing on the asteroid, resulting in the failure of the task. Therefore, in recent years, the academic community has proposed the use of flexible materials on the probe to reduce the energy generated by the collision during landing and achieve stable adhesion. However, compared with rigid materials, flexible materials are prone to flexible deformation, which causes complex random changes in the parameters of the probe, which poses a great challenge to the control of flexible probes. In recent years, with the development of artificial intelligence technology, intelligent control technology for complex stochastic systems has also been widely applied. However, the development of intelligent technology in the field of asteroid exploration is still very limited. Therefore, it is of great significance to study the intelligent control method of flexible probes for asteroid exploration missions. SUMMARY
[0003] The asteroid probe may bounce due to collision with the rugged surface of the asteroid during landing on the asteroid, and the asteroid probe is prone to escape after bouncing due to weak and irregular gravity of the asteroid, so that the mission fails. In view of the above problems, the main purpose of the present application is to provide an asteroid flexible probe intelligent control method based on a deep reinforcement learning SAC algorithm, a flexible probe structure model is established, a simplified model of the flexible probe is constructed, a second-order gravity potential function model is introduced to describe a weak gravity field model of the asteroid, an orbit dynamics equation of a rigid control node in a fixed coordinate system of the asteroid is derived, a position of the rigid control node is calculated according to the orbit dynamics equation, and the overall attitude of the spacecraft is calculated in real time according to the position of the rigid control node, and a dynamics model of the flexible probe coupled with the attitude and the orbit is constructed. An actor-critic neural network based on deep reinforcement learning is used to interact with the attitude-orbit coupled dynamics equation, the command thrust value of the thruster and the thruster deflection angle are output through the neural network, the neural network is trained based on the SAC algorithm to design the reward function and the cost function, and the asteroid flexible probe intelligent controller is constructed, so that the flexible probe can effectively cope with the random changes of the flexible probe structure parameters caused by the complex deformation of the flexible material. And the orbit and attitude motion of the flexible probe are controlled at the same time, on the premise of stabilizing the attitude of the flexible probe in real time, the flexible probe realizes high-precision tracking of the expected attachment orbit of the target task. The present application has the advantages of high control precision, strong anti-interference ability and strong robustness.
[0004] The object of the present application is achieved by the following technical solutions.
[0005] The asteroid flexible probe intelligent control method based on the deep reinforcement learning SAC algorithm disclosed by the present application comprises the following steps:
[0006] Step 1: Establish a flexible probe model, use a linear spring model to describe the flexible force generated by the deformation of the flexible airbag, and construct a simplified model of the flexible probe; introduce a second-order gravity potential function model to describe a weak gravity field model of the asteroid, derive an orbit dynamics equation of a rigid control node in a fixed coordinate system of the asteroid, calculate a position of the rigid control node according to the orbit dynamics equation, and calculate the overall attitude of the spacecraft in real time according to the position of the rigid control node, and construct a dynamics model of the flexible probe coupled with the attitude and the orbit. The flexible probe is composed of a flexible airbag and a rigid control node.
[0007] Step 1.1: Establish a flexible probe model connected by a flexible airbag and multiple rigid control nodes.
[0008] Unlike single-rigid-body probes, flexible probes consist of multiple rigid control nodes connected by flexible airbags. Based on the requirements of actual asteroid exploration missions, n rigid control nodes and n flexible airbags are arranged in a regular n-gon configuration. Each adjacent rigid control node is connected by a flexible airbag. The flexible airbag is a sealed buffer airbag used to absorb energy generated during landing collisions, preventing direct contact between the rigid control nodes and the asteroid surface, which could lead to probe bounce and escape or damage to the onboard payload. It also provides support after landing. Each rigid control node consists of a rigid probe platform, thrusters, sensors, sampling devices, an airbag storage chamber, and other payloads. The rigid probe platform carries the payload; the thrusters provide thrust for attitude-orbit coupling motion control of the probe; the sensors measure the distance and velocity between the rigid control nodes and the asteroid surface; the sampling devices collect samples from the asteroid surface after stable attachment; and the airbag storage chamber houses the uninflated flexible airbag.
[0009] Step 1.2: Use a linear spring model to describe the flexible force generated by the deformation of the flexible airbag, and construct a simplified model of the flexible detector.
[0010] During the attachment process of the flexible detector, the flexible airbag deforms, resulting in a flexible force between the rigid control nodes. This flexible force is equivalent to the elastic force generated by a linear spring, thus simplifying the flexible detector model. The initial configuration of the flexible detector is equivalent to a regular n-gon, with the rigid control nodes connected by linear spring components. The flexible force between the i-th and j-th rigid control nodes is shown in formula (1):
[0011]
[0012] Among them, F ij Indicates the magnitude of the flexible force, l ij Let l and l0 represent the actual relative distance and initial relative distance between the two rigid control nodes, respectively; k represents the equivalent elasticity coefficient; and n represents the number of rigid control nodes. Each rigid control node is simplified to a mass point. This yields a simplified model of the flexible detector.
[0013] Step 1.3: The weak gravitational field model of the asteroid is described by introducing a second-order gravitational potential function model. The orbital dynamics equation of the rigid control node in the asteroid fixed coordinate system is derived. The position of the rigid control node is calculated according to the orbital dynamics equation. The overall attitude of the spacecraft is calculated in real time according to the position of the rigid control node. The dynamics equation of attitude-orbit coupling of the flexible probe is constructed.
[0014] Step 1.3.1: To accurately establish the flexible probe pose-orbit coupling motion dynamics model, the following three coordinate systems are constructed, and the dynamics equations of the asteroid and each rigid control node of the flexible probe in the heliocentric inertial coordinate system are established.
[0015] ① Heliocentric inertial coordinate system OXYZ. The origin of the coordinate system is fixed at the center of the sun, the OX axis points to the vernal equinox direction in the asteroid orbit plane, and the OZ axis is along the angular velocity direction of the asteroid orbit motion, and the OY axis is determined by the right-hand screw rule.
[0016] ② Asteroid-fixed coordinate system oxyz. The origin of the coordinate system is fixed at the center of the asteroid, and the ox, oy, and oz axes coincide with the maximum, intermediate, and minimum inertia principal axes of the asteroid, respectively.
[0017] ③ Body-fixed coordinate system o bi x bi y bi z bi . The origin of the coordinate system is fixed at the center of mass of the i-th rigid control node, the o bi x bi axis points in the radial direction away from the center of mass of the probe in the initial configuration plane, the o bi z bi axis points to the normal direction of the initial configuration plane, and the o bi y bi axis is determined by the right-hand screw rule.
[0018] In the process of deriving the differential equations of motion of each rigid control node in the weak gravitational field of the asteroid, the task time required for the attachment process of the flexible probe is much smaller than the orbital period of the asteroid, so the asteroid's orbital motion is ignored. In the heliocentric inertial system, the orbital dynamics equation of the asteroid is expressed as:
[0019]
[0020] where r1 is the position vector from the center of the sun to the center of the asteroid, is its second-order derivative, and μ represents the gravitational constant of the sun. The orbital motion of the asteroid around the sun is described by equation (2).
[0021] In the heliocentric inertial system, the orbital dynamics equation of each rigid control node of the flexible probe is expressed as:
[0022]
[0023] where r si is the position vector from the center of the sun to the center of the controller, is its second-order derivative, and g aiThe gravitational acceleration is caused by the gravitational pull of the asteroid, a ci The control thrust acceleration is generated by the control thrust action, a uni The unknown acceleration is caused by an unknown disturbance, a ei It is the elastic acceleration generated by the flexible interaction between the rigid control nodes of the flexible probe. Equation (3) describes the orbital motion of each rigid control node in the flexible probe around the sun.
[0024] Step 1.3.2: Introduce the second-order gravitational potential function model, derive the orbital dynamics model of the rigid control node in the asteroid fixed coordinate system, and calculate the position of the rigid control node based on the orbital dynamics equation.
[0025] To describe the irregular weak gravitational field of an asteroid, the control process is modeled using the second-order gravitational potential function of the asteroid, expressed as follows:
[0026]
[0027] Where ψ,θ represents the latitude and longitude angles of a specific location in the gravitational field relative to the asteroid, and R... a μ is the maximum semi-major axis of the asteroid's approximate ellipsoid. a C is the gravitational coefficient of an asteroid. 20 and C 22 denoted as the tuning coefficient of the asteroid ellipsoid, and r is the magnitude of the vector at a specific location in the gravitational field.
[0028] The conversion relationship between position coordinates and latitude and longitude in the fixed coordinate system of asteroids
[0029]
[0030] Substituting equation (5) into equation (4) and taking the partial derivatives with respect to the three-axis position components, we obtain the expression for the asteroid's gravitational acceleration as shown below:
[0031]
[0032] Because the asteroid and the flexible probe are very far from the sun, while the asteroid and the flexible probe are very close, the distance between the rigid control node of the flexible probe and the sun is equivalent to the distance between the asteroid and the sun, i.e.: r si If ≈r1, then by subtracting equation (2) from equation (3), we get:
[0033]
[0034] Define the relative position vector between the asteroid and the i-th rigid control node of the flexible probe in the heliocentric inertial frame as:
[0035] ρ si= r si - r1 (8)
[0036] Taking the second derivative of equation (8) and substituting into equation (7), we get
[0037]
[0038] Since the asteroid-fixed coordinate system rotates with angular velocity ω o = [0 0 ω o ] T , the relative derivative relationship between the fixed coordinate system and the rotating coordinate system is given by equation (9)
[0039]
[0040] where ρ i is the position vector of the i-th rigid control node of the flexible probe in the asteroid-fixed coordinate system. Substituting equation (10) into equation (9), we get
[0041]
[0042] where
[0043] ρ i = [x i y i z i ] T
[0044]
[0045]
[0046]
[0047]
[0048] ω o = [0 0 ω o ] T
[0049] Thus, the orbital dynamics equation of the i-th rigid control node is given by equation (12)
[0050]
[0051] Given the initial state and control input of the rigid control node, we can solve equation (12) by numerical integration to obtain the position of the i-th rigid control node in the asteroid-fixed coordinate system.
[0052] Step 1.3.3: Calculate the whole spacecraft attitude in real time according to the position of the rigid control nodes, and construct the dynamics equation of the flexible probe attitude-orbit coupling.
[0053] The orbit motion of the whole flexible probe is described by the motion of the probe mass center. The position of the probe mass center ρ m is obtained in real time according to the position of each rigid control node:
[0054]
[0055] Where, m i represents the mass of the i-th rigid control node. The velocity of the mass center is obtained by taking the derivative of the mass center position with respect to time: v m = dρ m / dt. Due to the complex flexible deformation of the flexible material during the attachment process, the attitude of the flexible probe cannot be described by the method for describing the attitude of a rigid spacecraft. A plane formed by the connecting line of the mass centers of the three rigid control nodes numbered 1, [n / 3]+1 and 2[n / 3]+1 is defined as the attitude reference plane of the flexible probe; where [N] represents the integer part of N. The out-of-plane attitude angle φ is defined as the angle between the normal vector of the upper surface of the flexible probe pointing to the mass center of the flexible probe and the oz axis of the asteroid-fixed coordinate system, and the upper surface of the flexible probe is the surface facing the space at the initial time of the attachment of the flexible probe. The in-plane attitude angle θ is defined as the angle through which the reference plane rotates about the positive normal vector of the reference plane.
[0056] The two attitude angles θ, φ are calculated in real time by the position of the mass centers of the three rigid control nodes, and the calculation formula is:
[0057]
[0058] Where, ρ1(t0) and ρ m (t0) are the position vectors of the rigid control node 1 and the mass center of the probe at the initial time of the attachment.
[0059] During the attachment process, the flexible probe uses the propellers carried to generate thrust by ejecting working medium to track the expected orbit obtained by planning, while keeping the attitude of the flexible probe stable, so as to realize the stable attachment of the flexible probe. The rigid control node i is configured with two propellers; the thrust direction of one propeller points to the upper surface of the probe, and can swing within a predetermined size of space cone, and the thrust direction of the propeller relative to the reference plane is described by the angle α i between the thrust vector and the normal vector of the reference plane pointing to the upper surface, and the angle β i between the projection of the thrust vector in the reference plane and the vector pointing from the mass center of the probe to the mass center of the i-th rigid control node, and the thrust size is T i1; another thruster direction is fixed, the thrust direction is the same as the normal vector of the reference plane pointing to the lower surface, and the thrust size is T i2 .
[0060] The resultant thrust force of the two thrusters of the rigid control node i is represented in the asteroid-fixed coordinate system as:
[0061]
[0062] According to formula (15), the direction of the thrust force generated by the rigid control node is affected by the overall attitude of the probe, and the thrust force also controls the flexible probe orbit and attitude, so the control of the flexible probe attached to the asteroid is a complex coupling process of attitude and orbit control. The control acceleration generated by the thrust force on the rigid control node is represented as: ci = T ci / m i .
[0063] The attitude-orbit coupling dynamics equation of the flexible probe is established according to the above formula (12) (13) (14) (15).
[0064] Step two: adopt the actor-critic neural network based on deep reinforcement learning to interact with the attitude-orbit coupling dynamics equation, output the command thrust value and thruster deflection angle of the thruster through the neural network, and design the reward function and cost function based on the SAC algorithm to train the neural network, and based on the trained neural network, an intelligent controller for the asteroid flexible probe is constructed.
[0065] Step 2.1: adopt the actor-critic neural network based on deep reinforcement learning to interact with the attitude-orbit coupling dynamics equation, and let the neural network output the command thrust value and thruster deflection angle of the thruster.
[0066] Step 2.1.1: establish the actor-critic neural network based on deep reinforcement learning.
[0067] An intelligent control system based on reinforcement learning includes two parts: agent and environment. The agent is the controller in the intelligent control system, which has the ability of learning and decision-making. The environment is the controlled object in the intelligent control system.
[0068] An agent learns a control policy through continuous interaction with the environment, and finally obtains an optimal control policy to achieve the control goal, which is also called training of the agent. The training process is described by Markov decision process (MDP). MDP is usually composed of a four-tuple, denoted as: <S, A, P, R>, where S represents a set of environment states, A represents a set of agent actions, P is a state transition probability of the environment, and R represents a reward function. At any time t in the training process, the agent perceives the state S t ∈Sof the environment, and performs an action A t ∈A, so that the environment is transferred to the next state S t+1 ∈Saccording to the probability P, and obtains a potential reward R t from the environment.
[0069] The agent has two important functional components, namely a policy Π that outputs a decision action according to a state input, and an evaluation Q that evaluates the current policy according to a state and a reward value.
[0070] The function of the evaluation in the agent is to evaluate the quality of the current policy and update the policy by estimating the long-term cumulative reward obtained by the current policy. The expectation of the long-term cumulative reward is defined as a value function. There are two kinds of value functions, namely a state value function and a state-action value function. The state value function V Π (S) represents the expectation of the long-term cumulative reward obtained by the agent when performing an action according to the policy Π in a state S t of the environment:
[0071] V Π (S) = E[G t |S t =S] (16)
[0072] The state value function is expressed in the form of recursion of the expectation of the sum of the current reward value and the next state value function weighted by a discount factor, i.e., the Bellman expectation equation:
[0073] V Π (S) = E[R t+1 + γV Π (S t+1 )|S t =S] (17)
[0074] The state-action value function Q Π(S, A) is also called Q function, which represents the expected long-term cumulative reward obtained by the agent in a certain state of the environment, after performing a certain action, and then continuing to make decisions according to the policy:
[0075] Q Π (S, A) = E[G t |S t = S, A t = A] (18)
[0076] The Q function can also be expressed as the corresponding Bellman expectation equation:
[0077] Q Π (S, A) = E[R t+1 + γV(S t+1 )|S t = S, A t = A] = R t+1 + γE a~Π [Q(S t+1 , A t+1 )|S t = S, A t = A] (19)
[0078] The basic method of deep reinforcement learning is to use two neural networks named Actor and Critic to replace the policy and evaluation in the agent respectively, which is called Actor-Critic algorithm. Both neural networks are composed of an input layer, one or more hidden layers and an output layer. The input of the input layer of the Actor network is the state of the environment, and the output of the output layer is the policy after fitting through the hidden layer. For deterministic algorithms, the network will directly give the value of the action; for stochastic policy, the Actor will output a probability distribution, and then sample the action according to the probability distribution. The input of the input layer of the Critic network is the state of the environment and the reward value obtained after performing the action, and the output of the output layer is the value of the Q function obtained by fitting through the hidden layer.
[0079] The process of fitting the input through each layer to output the output by the neural network is called forward propagation. For example, a neural network with n layers, where the forward propagation formula of the l-1 layer to the l layer (l = 2…n) is:
[0080] a l = σ(z l ) = σ(W l a l-1 + b l ) (20)
[0081] where a l m×1is the vector of data values for each neuron in the lth layer, m is the number of neurons in the layer, σ(z l ) is the activation function of the neural network, z l = W l a l-1 +b l is the transfer value of the data, is the diagonal matrix of weights for each neuron in the lth layer, b l m×1 is the vector of biases for each neuron in the lth layer. The set of weights and biases [W, b] of the neurons in each layer of the neural network is called the parameter set of the neural network. The parameter set of the Actor network is denoted by φ, and the parameter set of the Critic network is denoted by θ.
[0082] The process of setting the cost function of the neural network, using gradient descent method to iteratively optimize to find the minimum value of the cost function, so as to find the optimal parameter set, so that the fitted data value is as close as possible to the actual data value, is called the backpropagation process. Define the cost function of the neural network as J(W, b, x, y), then the gradient of the cost function with respect to the lth layer of the neural network z l is expressed as:
[0083]
[0084] Then, the process of parameter update of the lth layer of the neural network using backpropagation is expressed as:
[0085]
[0086] In the deep reinforcement learning method, the training of the intelligent controller is actually the training of the Actor and Critic neural networks. A replay pool is constructed to store the training data in real time: D = <S N-L ,A N-L ,P N-L ,R N-L …S N ,A N ,P N ,R N >(L is the length of the replay pool); during the training of the intelligent controller, a proper number of training data is sampled from the replay pool, a reasonable cost function is set according to the principle of reinforcement learning, and the parameters of the Actor and Critic neural networks are iteratively updated using the above backpropagation method to obtain the optimal neural network parameters φ * and θ * .
[0087] Step 2.1.2: Realize the interaction between the neural network and the attitude-orbit coupled dynamics equation, output the commanded thrust value of the thruster and the deflection angle of the thruster through the neural network.
[0088] In the process of training the intelligent controller, in order to improve the training efficiency, the environment in the intelligent control system of the flexible probe adopts the attitude-orbit coupled motion dynamics model based on the simplified mathematical model of the structure of the flexible probe established in step one, and is trained in the simulation environment; the agent refers to the intelligent controller of the flexible probe. The state input of the intelligent controller is the real-time mass center position, velocity and planned position, velocity deviation of the probe, as well as the deviation between the real-time attitude and the expected attitude, which is represented as The action output is the commanded thrust value of each thruster and the deflection angle of the thruster mounted on the upper surface, which is represented as Act = [T 11 ,T 12 ,...T n1 ,T n2 ,α1,...α n ,β1,...β n ] T The intelligent controller needs to interact with the attitude-orbit coupled motion dynamics module of the flexible probe in the simulation environment for training. In the process of training the intelligent controller in each round, the thrusters of each rigid control node generate thrust according to formula (15) after receiving the control command of the intelligent controller; each rigid control node generates motion according to the dynamics law in formula (12) under the action of the thrust; the position, velocity and overall attitude of the probe mass center are calculated in real time according to formula (13) and formula (14) through the position and velocity of each rigid control node; the actual state of the probe is subtracted from the planned expected state to obtain the input of the intelligent controller; the intelligent controller calculates the reward value obtained in this round of training according to the input state error and the reward function set by the program. In the training process of the intelligent controller, the parameters input and output by the intelligent controller need to be standardized to improve the stability of the training process; the change range of each dimension of the input parameter is standardized to [-1, 1], and the change range of each dimension of the output parameter is [0, 1], the relationship between the standardized parameter and the original parameter is represented as:
[0089]
[0090] Wherein, Sta max ,Sta min ,Act max ,Act minThe upper and lower bounds of the input and output are the values. After the intelligent controller is trained in the simulation environment, it is deployed on the on-board computer of the flexible probe, and then applied in the actual asteroid attachment mission. In the actual asteroid attachment mission, the error between the actual speed and position of each rigid control node measured by the sensor and the expected value is used as the input of the intelligent controller, and the intelligent controller outputs the command thrust to each thruster to generate the control thrust.
[0091] Step 2.2: Design reward function and cost function based on SAC algorithm to train neural network and construct intelligent controller of asteroid flexible probe.
[0092] The SAC algorithm is a random policy-based algorithm in deep reinforcement learning. The basic principle is that during the training of the agent, not only the long-term cumulative reward value is maximized, but also the information entropy of the policy is maximized, also known as maximum entropy reinforcement learning. The expression of the long-term cumulative reward of the SAC algorithm is:
[0093]
[0094] where H(Π(·|S t )) is the information entropy of the policy; and is defined as:
[0095]
[0096] Contrary to the thermodynamic phenomenon described by the second law of thermodynamics, the generation of action is a process of information entropy reduction, so the information entropy is negative. By maximizing the information entropy of the policy, the randomness of the policy is increased, so that the probability of each action is as dispersed as possible, rather than concentrated on one action, thereby enhancing the exploration of the policy during training and improving the robustness of control. Substituting equation (24) into equation (19), the soft Bellman equation of the Q function is obtained:
[0097]
[0098] Further, the expression of the state value function is obtained:
[0099]
[0100] As shown in equation (27), the state value function can be expressed as the soft maximum value form of the LogSumExp of the Q function in the action domain, which is the most important feature of the SAC algorithm. Its advantage is that the optimal policy will not fall into the local optimum or a certain optimum of the Q function, but will approach the global optimum or all optimal values of the Q function through a smooth method. Therefore, the SAC-based policy is
[0101]
[0102] where, represents KL divergence, describes the degree of similarity between distribution P and distribution Q, the smaller the KL divergence, the highest degree of similarity of the two parts. By minimizing the KL divergence, the probability distribution π K (·|S t ) best approximation probability distribution so that the Q function has the form of soft maximum value.
[0103] According to the basic principle of SAC algorithm, the cost function of Critic neural network in agent based on SAC algorithm is obtained as:
[0104]
[0105] In the process of training the agent based on SAC algorithm, two Critic networks with parameters θ1, θ2 are used, and the smaller one of the two network output Q function values is taken each time. At the same time, in order to make the value of Q function not update too fast and lead to divergence, the target Critic network with slower update is set, and the network parameters are represented by , and the update method is: on the basis of each Critic network parameter update The cost function of Actor neural network is:
[0106]
[0107] In the training process, the Actor neural network outputs a Gaussian distribution Π φ ~N(μ,δ) representing the policy, whose mean μ t and variance δ t , the agent outputs action by sampling Gaussian distribution, and the specific sampling method is shown in equation (31):
[0108] A t = μ t + rand(0,1) δ t (31)
[0109] Where rand(0,1) represents a random number between 0 and 1.
[0110] In the training process of the agent based on SAC algorithm, the fixed temperature coefficient may cause the training process to be unstable due to the continuous change of the reward value, so the temperature coefficient needs to be adjusted in real time. When the strategy explores the unknown area of the environment, the temperature coefficient should be appropriately increased to stimulate exploration; when the characteristics of the current environment are clear and the strategy is fixed, the temperature coefficient should be lowered to make the training gradually converge. By setting the cost function and solving the optimal value by gradient descent method, the adjustment of the temperature coefficient is realized. The cost function used in the adjustment process of the temperature coefficient is:
[0111]
[0112] The reward function of the intelligent controller is designed as follows:
[0113]
[0114] The reward function mainly includes three parts: state-related reward value, action-related reward value and penalty term. The state-related reward value r adopts a segmented reward method, and ε1=ε2=1e-3 is the accuracy value of the centroid position deviation and attitude deviation. When the centroid position deviation and attitude deviation are within the accuracy range, a negative weight is given to the sum as the reward value to promote the further convergence of the deviation; when the centroid position deviation exceeds the accuracy range but the current time value is less than the previous time value, a smaller negative reward is given, otherwise a larger negative reward is given; the segmented reward method is conducive to the stable convergence of the state deviation. The action reward value is -0.05|T cN | 2 , where T c =T c1 +T c2 +T c3 is the total thrust of all thrusters, and this reward value is conducive to reducing the thrust value of the thruster output to save the consumption of propellant on the star. The meaning of the penalty term B is that when the square sum of the centroid position deviation is greater than the maximum value L of the square sum of the maximum centroid position deviation of the three axes or the relative distance between the two nodes is greater than the limit distance K, a large negative reward is given as a penalty, and the training is exited.
[0115] After the design of the reward function of the intelligent controller is completed, an intelligent controller with two Critic networks and an Actor network is established. The two Critic networks of the intelligent controller have the same structure, each network has a state path and an action path respectively at first, each branch has a plurality of hidden layers, each hidden layer has a plurality of neurons, and the two branches are connected through a common hidden layer with a plurality of neurons and finally output Q values. The Actor network has a common hidden layer with a plurality of neurons, and then branches into two branches, each branch has a hidden layer with a plurality of neurons, and respectively fits the mean and variance of the strategy probability distribution.
[0116] The intelligent controller is used to control the orbit and attitude motion of the flexible system with complex deformation characteristics at the same time, realizes stable adhesion of the flexible probe on the asteroid, and has the advantages of high training efficiency, high robustness and high control precision.
[0117] Step three: through the asteroid flexible probe intelligent controller obtained in step two, the flexible probe can effectively cope with the random changes of the structure parameters of the flexible probe caused by the complex deformation of the flexible material, and realize high-precision tracking of the expected adhesion orbit of the target task.
[0118] The neural network trained based on the SAC algorithm is used to process the input state, i.e. the real-time mass center position, velocity and planned position, velocity deviation of the probe, and the deviation between the real-time attitude and the expected attitude, so that the output action, i.e. the instruction thrust value of each thruster and the deflection angle of the thruster installed on the upper surface, can realize that the flexible probe can effectively cope with the random changes of the structure parameters of the flexible probe caused by the complex deformation of the flexible material; and simultaneously control the orbit and attitude motion of the flexible probe, realize high-precision tracking of the expected adhesion orbit of the target task on the premise of real-time stable attitude of the flexible probe.
[0119] Advantages:
[0120] 1. The asteroid flexible probe intelligent control method based on the deep reinforcement learning SAC algorithm disclosed in the application constructs a regular polygon configuration asteroid flexible probe composed of a plurality of flexible air bags and rigid control nodes, absorbs collision energy through the flexible air bag, and avoids rebound escape during adhesion. Then, a simplified model of the flexible probe is derived, and on this basis, an attitude-orbit coupling motion dynamics model in a weak gravity field of an asteroid is established, which effectively improves the simulation precision and calculation speed in the process of simulating the adhesion of the asteroid flexible probe to the asteroid.
[0121] 2. This invention discloses an intelligent control method for a flexible asteroid probe based on the deep reinforcement learning (SAC) algorithm. Based on the SAC algorithm, it employs an actor-commentator neural network (ADN) that interacts with the attitude-orbit coupled dynamic equations. The neural network outputs the thrust command value and thruster deflection angle. A reward function and cost function are designed based on the SAC algorithm to train the neural network. An intelligent controller for the flexible asteroid probe is then constructed based on the trained neural network. This intelligent controller simultaneously controls the orbital and attitude motions of the flexible system with complex deformation characteristics, achieving stable attachment of the flexible probe to the asteroid. It also offers advantages such as high training efficiency, high robustness, and high control precision. Attached Figure Description
[0122] Figure 1 This is a flowchart of the intelligent control method for asteroid flexible probes using the deep reinforcement learning SAC algorithm of this invention.
[0123] Figure 2 This is a schematic diagram of the multi-node distributed flexible detector model of the present invention;
[0124] Figure 3 This is a schematic diagram showing the configuration of the thruster at each rigid control node of the present invention;
[0125] Figure 4 This is a simplified schematic diagram of the flexible detector of the present invention;
[0126] Figure 5 This is a schematic diagram of the attitude and centroid trajectory of the flexible detector of the present invention;
[0127] Figure 6 This is a schematic diagram of the intelligent control system for the flexible detector of the present invention;
[0128] Figure 7 This is a schematic diagram of the neural network structure of the intelligent controller of the present invention;
[0129] Figure 8 This is a schematic diagram of the desired orbit of the centroid of the flexible detector of the present invention;
[0130] Figure 9 This is the training reward value curve of the intelligent controller of the present invention;
[0131] Figure 10 This refers to the thrust vector of the two thrusters at a single rigid control node in this invention.
[0132] Figure 11 The flexible probe of this invention is attached to the actual orbit of an asteroid;
[0133] Figure 12 This is the centroid position deviation curve of the flexible detector of the present invention;
[0134] Figure 13 centroid velocity deviation curve of the flexible probe of the present application;
[0135] Figure 14 posture angle deviation curve of the flexible probe of the present application. DETAILED DESCRIPTION
[0136] In order to better illustrate the purposes and advantages of the present application, the specific embodiments of the present application will be further described in detail below in combination with the drawings.
[0137] Example 1:
[0138] For the attached target being Itokawa asteroid, the asteroid mass is 3.147x10 10 kg, the reference radius is 300 m, and the rotation angular velocity is 1.4424x10 -4 rad / s. As shown in the figure, the asteroid flexible probe intelligent control method based on the deep reinforcement learning SAC algorithm disclosed in the present embodiment is specifically implemented as follows: Figure 1
[0139] Step 1: Establish a flexible probe structure model, construct a simplified model of the flexible probe, describe the weak gravity field model of the asteroid by introducing a second-order gravity potential function model, derive the orbit dynamics equation of the rigid control node in the asteroid fixed coordinate system, calculate the position of the rigid control node according to the orbit dynamics equation, and calculate the spacecraft overall attitude in real time according to the position of the rigid control node, and construct a dynamics model of the flexible probe attitude-orbit coupling.
[0140] Step 1.1: Establish a flexible probe model connected by multiple rigid control nodes.
[0141] As shown in the figure is a multi-node distributed flexible probe structure model. In the present embodiment, n is taken as 3, i.e. the flexible probe is composed of 3 rigid control nodes and 3 flexible airbags, and the initial configuration is a regular triangle. As shown in the figure is the configuration of the thruster of each rigid control node. Each rigid control node is configured with two thrusters; the direction of one thruster is towards the upper surface of the probe and can swing within a predetermined size of space cone, and the maximum opening angle of the space cone is 60 degrees. The direction of the other thruster is fixed, and is along the normal direction of the reference surface towards the lower surface of the probe. Figure 2 Figure 3 Step 1.2: A linear spring model is used to describe the flexible force generated by the deformation of the flexible airbag, and a simplified model of the flexible probe is constructed.
[0142] Step 1.2: A linear spring model is used to describe the flexible force generated by the deformation of the flexible airbag, and a simplified model of the flexible probe is constructed.
[0143] As shown in the figure is a multi-node distributed flexible probe structure model. In the present embodiment, n is taken as 3, i.e. the flexible probe is composed of 3 rigid control nodes and 3 flexible airbags, and the initial configuration is a regular triangle. As shown in the figure is the configuration of the thruster of each rigid control node. Each rigid control node is configured with two thrusters; the direction of one thruster is towards the upper surface of the probe and can swing within a predetermined size of space cone, and the maximum opening angle of the space cone is 60 degrees. The direction of the other thruster is fixed, and is along the normal direction of the reference surface towards the lower surface of the probe. Figure 4 A simplified model of the flexible probe is shown, in which three rigid control nodes of the flexible probe are equivalent to mass points, and the flexible airbag between each two rigid control nodes is equivalent to a linear spring. The mass of each rigid control node is set to 20 kg, and the mass of the linear spring component is negligible. The equivalent spring coefficient k is set to 0.001, and the initial relative distance Δl between nodes is set to 1 m.
[0144] Step 1.3: A second-order gravitational potential function model is introduced to describe the weak gravitational field model of the asteroid, and the orbit dynamics equation of the rigid control node in the asteroid fixed coordinate system is derived. The position of the rigid control node is calculated according to the orbit dynamics equation, and the overall attitude of the spacecraft is calculated in real time according to the position of the rigid control node. The attitude-orbit coupled dynamics equation of the flexible probe is constructed.
[0145] As shown in Figure 5 The attitude reference plane of the flexible probe, the two corresponding attitude angles, and the center of mass orbit are shown.
[0146] Step 2: An actor-critic neural network based on deep reinforcement learning is used to interact with the attitude-orbit coupled dynamics equation, so that the neural network outputs the command thrust value and the thrust deflection angle of the thruster. The reward function and the cost function are designed based on the SAC algorithm to train the neural network, and the intelligent controller of the asteroid flexible probe is constructed.
[0147] Step 2.1: An actor-critic neural network based on deep reinforcement learning is used to interact with the attitude-orbit coupled dynamics equation, so that the neural network outputs the command thrust value and the thrust deflection angle of the thruster.
[0148] As shown in Figure 6 The intelligent control system of the flexible probe is shown. The working principle of the control system is as follows: the thrusters of the three rigid control nodes generate thrust according to formula (15) after receiving the control instructions of the intelligent controller; the three rigid control nodes move according to the dynamics law in formula (12) under the action of the thrust; the position, velocity, and overall attitude of the center of mass of the flexible probe are calculated in real time according to formula (13) and formula (14) through the position and velocity of each rigid control node; the input of the intelligent controller is obtained by subtracting the actual state of the flexible probe from the planned desired state; the intelligent controller calculates the reward value obtained in this round of training according to the input state error and the reward function set by the program. The allowed maximum and minimum values of the input and output parameters are set as follows: the position deviation of each axis: [0, 10] m, the velocity deviation of each axis: [0, 5] m / s, the attitude angle deviation: [-π, π] rad, the output thrust value T ij of the thruster: [0, 1] N, the deflection angle α i : [0, π / 6] rad, the deflection angle β i : [0, 2π] rad.
[0149] Step 2.2: Design reward function and cost function based on SAC algorithm to train neural network and build intelligent controller for asteroid flexible probe.
[0150] Set L = 3, K = 1.2, ε1 = ε2 = 0.001 in the reward function.
[0151] As shown in Figure 7 , the neural network structure of the intelligent controller. Two Critic networks have the same structure. Each network has a state path and an action path respectively, each branch has 2 hidden layers, each hidden layer has 64 neurons, and the two branches are connected through a common hidden layer with 64 neurons and finally output Q value. The Actor network has a common hidden layer with 64 neurons, then branches out into two branches, each with a hidden layer of 64 neurons, to fit the mean and variance of the policy probability distribution.
[0152] As shown in Figure 8 , the centroid expected orbit of the flexible probe attached to the asteroid, used to calculate the deviation between the actual position and the expected position of the centroid in real time.
[0153] Set the key parameters in the training process of the intelligent controller: Critic network learning rate: 0.001, Actor network learning rate: 0.0001, temperature coefficient learning rate: 0.0001, experience pool size: 1000000, initial temperature coefficient: 0.5, target entropy: -12.
[0154] According to the design parameters given above, the intelligent control system of the flexible probe can be trained to obtain the optimal intelligent controller.
[0155] As shown in Figure 9 , the agent has a total of 580 training. During the first 500 rounds of training, the reward value obtained by the agent has a large range of changes, indicating that the agent is constantly adjusting the strategy, exploring and trying. After 500 rounds, the reward value gradually increases, indicating that the agent's exploration has found a direction that can achieve good control effect, so as to continuously optimize the control effect and get higher scores.
[0156] Take the control strategy of the training round with the highest score in the training process to verify the control effect of intelligent control.
[0157] Step three: through the asteroid flexible probe intelligent controller obtained in step two, the flexible probe can effectively cope with the random changes of the flexible probe structure parameters caused by the complex deformation of the flexible material, and realize the expected attached orbit of high-precision tracking target task.
[0158] As Figure 10 shown is the thrust vector generated by two thrusters at each thrust impulse moment in the single rigid control node attachment process. Through the entire training process, the intelligent controller generates consistent control thrust on each rigid control node, that is, T 11 = T 21 = T 31 , T 12 = T 22 = T 32 , a1 = a2 = a3 and b1 = b2 = b3, so as to maintain the attitude stability of the flexible probe as a whole in real time. During the attachment task, the thrusters on the upper surface of the single rigid control node continuously adjust the size and direction of the thrust impulse, and the thrusters on the lower surface adjust the size of the thrust impulse; the three nodes simultaneously output thrust to make the mass center of the flexible probe track the planned orbit.
[0159] As Figure 11 shown is the actual orbit of the flexible probe attached to the asteroid. The mass center orbit is relatively stable, and there is no large-scale deviation. The orbits of the three rigid control nodes are almost parallel to the mass center orbit, and the attitude of the flexible probe is also always stable.
[0160] As Figures 12 to 14 shown are the mass center position deviation, mass center velocity deviation and attitude angle deviation of the probe, respectively. The three-axis mass center position deviation is always maintained within 0.2m, and the terminal position deviation is within 0.05m, achieving high-precision tracking of the planned orbit. The terminal velocity is within 0.025m / s, and the velocity of the flexible probe when contacting the surface of the asteroid is very low, which will not cause violent collision with the surface of the asteroid, ensuring the stability and safety of the attachment. During the attachment process, the angle between the normal of the probe attitude reference surface and the oz axis of the asteroid-fixed coordinate system is deflected within 3 degrees, and the asteroid probe almost does not rotate around the normal of the reference surface, which proves that the attitude of the flexible probe is always stable during the attachment process, which not only ensures the final safe and stable attachment, but also ensures the stable detection of the asteroid surface by the sensor during the attachment process.
[0161] In summary, the intelligent control method based on the deep reinforcement learning SAC algorithm can effectively control the high-coupling motion of the flexible probe, track the planned orbit in real time with high precision on the basis of maintaining the attitude stability of the flexible probe, and thus realize stable and safe attachment on the surface of the asteroid.
[0162] The above detailed description of the specific description, the purpose, technical scheme and beneficial effects of the application are further described in detail, it should be understood that the above description is only a specific embodiment of the present application, and is not used to limit the protection scope of the present application, any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. A method for intelligent control of an asteroid flexible probe based on a deep reinforcement learning SAC algorithm, characterized in that: The method comprises the following steps, Step one: establishing a flexible probe model, using a linear spring model to describe the flexible force generated by the deformation of the flexible air bag, constructing a simplified model of the flexible probe; by introducing a second-order gravitational potential function model to describe the weak gravity field model of the asteroid, deriving the orbit dynamics equation of the rigid control node in the asteroid fixed coordinate system, calculating the position of the rigid control node according to the orbit dynamics equation, and calculating the overall attitude of the spacecraft in real time according to the position of the rigid control node, and constructing a dynamics model of the attitude-orbit coupling of the flexible probe; the flexible probe model is composed of a flexible air bag and a rigid control node; Step two: using the actor-critic neural network based on deep reinforcement learning to interact with the attitude-orbit coupling dynamics equation, outputting the command thrust value and thrust deflection angle of the thruster through the neural network, and training the neural network based on the SAC algorithm to design the reward function and the cost function, and constructing an asteroid flexible probe intelligent controller based on the trained neural network; Step three: through the asteroid flexible probe intelligent controller obtained in step two, the flexible probe can effectively cope with the random changes of the flexible probe structure parameters caused by the complex deformation of the flexible material, and realize the high-precision tracking of the expected attached orbit of the target task.
2. The intelligent control method for asteroid flexible probe based on deep reinforcement learning SAC algorithm according to claim 1, characterized in that: The implementation method of step one is, Step 1.1: establishing a flexible probe model connected by a flexible air bag and multiple rigid control nodes; Unlike a single rigid body probe, a flexible probe is composed of a flexible air bag connecting multiple rigid control nodes. According to the requirements of the actual asteroid exploration task, n rigid control nodes and n flexible air bags are used to form a probe with a regular n-polygon configuration; each of the two adjacent rigid control nodes is connected by a flexible air bag; the flexible air bag is a closed buffer air bag used to absorb the energy generated by the collision during landing, avoid direct contact and collision between the rigid control node and the asteroid surface, causing the probe to bounce and escape or the payload to be damaged, and play a supporting role after landing; the rigid control node is composed of a rigid probe platform, a thruster, a sensor, a sampling device, an air bag storage cabin and other payloads; the rigid probe platform is used to carry the payload; the thruster is used to provide thrust and control the attitude-orbit coupling motion of the probe; the sensor is used to measure the distance and speed between the rigid control node and the asteroid surface; the sampling device is used to collect samples on the asteroid surface after stable attachment; the air bag storage cabin is used to store the flexible air bag in the un-inflated state; Step 1.2: using a linear spring model to describe the flexible force generated by the deformation of the flexible air bag, constructing a simplified model of the flexible probe model; In the process of the flexible probe attachment, the flexible airbag is deformed, and the flexible force is generated between the rigid control nodes; the flexible force is equivalent to the elastic force generated by the linear spring, that is, the initial configuration of the flexible probe is equivalent to a regular n-polygon, and the rigid control nodes are connected by the linear spring components; the flexible force between the ith rigid control node and the jth rigid control node is shown in formula (1): where F ij represents the size of the flexible force, l ij and l0 respectively represent the actual relative distance and the initial relative distance between the two rigid control nodes, k represents the equivalent elastic coefficient, and n represents the number of rigid control nodes; each rigid control node is simplified as a mass point; thus, a simplified model of the flexible probe is obtained; Step 1.3: The second-order gravitational potential function model is introduced to describe the weak gravitational field model of the asteroid, the orbit dynamics equation of the rigid control node in the asteroid fixed coordinate system is derived, the position of the rigid control node is calculated according to the orbit dynamics equation, the overall attitude of the spacecraft is calculated in real time according to the position of the rigid control node, and the dynamics equation of the attitude-orbit coupling of the flexible probe is constructed; Step 1.3.1: In order to accurately establish the dynamics model of the attitude-orbit coupling motion of the flexible probe, the following three coordinate systems are constructed, and the dynamics equations of the asteroid and each rigid control node of the flexible probe in the heliocentric inertial coordinate system are established; ① The heliocentric inertial coordinate system OXYZ; the origin of the coordinate system is fixed at the center of the sun, the OX axis points to the vernal equinox direction in the plane of the asteroid orbit, and the OZ axis is along the angular velocity direction of the asteroid orbit motion, and the OY axis is determined by the right-hand screw rule; ② The asteroid fixed coordinate system oxyz; the origin of the coordinate system is fixed at the center of the asteroid, and the ox axis, the oy axis and the oz axis are coincided with the maximum, the intermediate and the minimum inertia principal axes of the asteroid respectively; iii. The body-fixed coordinate system o bi x bi y bi z bi ; the origin of the coordinate system is fixed at the center of mass of the i-th rigid control node, o bi x bi axis points in the radial direction away from the overall center of mass of the probe in the initial configuration plane, o bi z bi axis points in the direction of the normal of the initial configuration plane, o bi y bi axis is determined by the right-hand screw rule; In the process of deriving the differential equation of the motion of each rigid control node in the weak gravitational field of the asteroid, the task time required in the process of the flexible probe attachment is very small compared with the orbital period of the asteroid, so the orbital motion of the asteroid is ignored; in the heliocentric inertial system, the orbit dynamics equation of the asteroid is expressed as: where r1is the position vector from the sun's center to the asteroid's center, is its second derivative, and μ represents the sun's gravitational constant; the orbit of the asteroid around the sun is described by equation (2). In the heliocentric inertial system, the orbit dynamics equation of each rigid control node of the flexible probe is expressed as: where r si is the position vector from the Sun's center to the controller's center, is its second derivative, g ai is the gravitational acceleration due to the asteroid's gravitational effect, a ci is the control thrust acceleration due to the control thrust effect, a uni is the unknown acceleration due to unknown disturbances, a ei is the elastic acceleration due to the flexible forces between the rigid control nodes of the flexible probe; the orbital motion of each rigid control node around the Sun in the flexible probe is described by equation (3). Step 1.3.2: The second-order gravitational potential function model is introduced, and the orbit dynamics model of the rigid control node in the asteroid fixed coordinate system is derived, and the position of the rigid control node is calculated according to the orbit dynamics equation; In order to describe the irregular weak gravitational field of the asteroid, the second-order gravitational potential function model of the asteroid is adopted, and the expression is as follows: where ψ, is the latitude angle of any position in the gravitational field relative to the asteroid, R a is the maximum semi-major axis of the asteroid's approximate ellipsoid, μ a is the gravitational coefficient of the asteroid, C 20 and C 22 are the tuning item coefficients of the asteroid's ellipsoid, and r is the modulus of the vector of any position in the gravitational field. The conversion relationship between the position coordinates in the asteroid fixed coordinate system and the latitude and longitude is Substitute formula (5) into formula (4) and take the partial derivative of the three-axis position components to obtain the expression of the gravitational acceleration of the asteroid as follows: Since the distance between the asteroid and the sun is very far and the distance between the asteroid and the flexible probe is very close, the distance between the rigid control node of the flexible probe and the sun is equivalent to the distance between the asteroid and the sun, i.e. r si ≈r1, then the difference between formula (3) and formula (2) is obtained: The relative position vector between the ith rigid control node of the asteroid and the flexible probe in the heliocentric inertial system is defined as: p si = r si - r1 (8) Take the second-order derivative of formula (8) and substitute it into formula (7) to obtain Since the asteroid-fixed coordinate system rotates with angular velocity ω o = [0 0 ω o ] T , equation (9) is further expressed as where ρ i is the position vector of the i-th rigid control node of the flexible probe in the asteroid-fixed coordinate system; substituting equation (10) into equation (9) gives In the formula, p i = [x i y i z i ] T ω o = [0 0 ω o ] T The orbit dynamics equation of the ith rigid control node is shown in formula (12) Given the initial state and control input of the rigid control node, the numerical integral of formula (12) is solved, that is, the position of the ith rigid control node in the asteroid fixed coordinate system is obtained. Step 1.3.3: Real-time calculation of the overall spacecraft attitude according to the rigid control node position, and construction of the dynamics equation of the flexible probe attitude-orbit coupling; The trajectory motion of the flexible probe as a whole is described by the motion of the center of mass of the flexible probe; the position of the center of mass of the flexible probe p m is obtained in real time according to the positions of the respective rigid control nodes: wherein m i represents the mass of the i-th rigid control node; the velocity of the mass center is obtained by derivation of the mass center position with respect to time: v m = dρ m / dt; due to the complex flexible deformation of the flexible material during the attaching process, the attitude of the flexible probe cannot be described by the method for describing the attitude of a rigid spacecraft; a plane formed by the connecting line of the mass centers of the three rigid control nodes numbered 1, [n / 3]+1 and 2[n / 3]+1 is defined as the attitude reference plane of the flexible probe; wherein [N] represents the integer part of N; the out-of-plane attitude angle is defined as the angle φ between the normal vector of the mass center of the flexible probe pointing to the upper surface of the flexible probe and the oz axis of the asteroid-fixed coordinate system; the surface facing the space at the initial time of the attaching of the flexible probe is the upper surface; the in-plane attitude angle is defined as the angle θ through which the reference plane rotates about the positive normal vector of the reference plane; Two attitude angles θ, φ are calculated in real time according to the centroid position of the three rigid control nodes, and the calculation formula is: Where, ρ1(t0) and ρ m (t0) is the position vector of the rigid control node 1 and the centroid of the detector at the initial attachment moment; In the process of attaching, the flexible probe uses the carried thruster to generate thrust by ejecting working substance to track the expected orbit obtained by planning, while keeping the attitude of the flexible probe stable, so as to realize stable attachment of the flexible probe; the rigid control node i is configured with two thrusters; the thrust direction of one thruster points to the upper surface of the probe and can swing within a predetermined size of space cone, and the thrust direction of the other thruster is fixed, and the thrust direction of the other thruster is the same as the normal vector of the reference surface pointing to the lower surface of the probe i , and the thrust size is T i ; the thrust direction of the other thruster is fixed, and the thrust direction of the other thruster is the same as the normal vector of the reference surface pointing to the lower surface of the probe i1 ; the thrust size is T i2 . The resultant force of the thrust generated by the two thrusters of the rigid control node i is represented in the asteroid-fixed coordinate system as: According to formula (15), the direction of the thrust generated by the rigid control node is affected by the overall attitude of the probe, and the thrust also controls the flexible probe orbit and attitude, so the control of the flexible probe attached to the asteroid is a complex coupling process of attitude and orbit control. The control acceleration generated by the thrust on the rigid control node is represented as: a ci = T ci / m i ; The dynamics equation of the flexible probe attitude-orbit coupling is established according to the above equations (12), (13), (14), and (15).
3. The intelligent control method for asteroid flexible probe based on deep reinforcement learning SAC algorithm according to claim 2, characterized in that: The implementation method of Step Two is, Step 2.1: An actor-critic neural network based on deep reinforcement learning is used to interact with the attitude-orbit coupling dynamics equation, and the neural network outputs the command thrust value and thruster deflection angle of the thruster; Step 2.1.1: An actor-critic neural network based on deep reinforcement learning is established; An intelligent control system based on reinforcement learning includes two parts: the agent and the environment. The agent is the controller in the intelligent control system, which has the ability to learn and make decisions. The environment is the controlled object in the intelligent control system. The agent learns the control strategy through continuous interaction with the environment, and obtains the optimal control strategy to achieve the control goal. This process is also called training of the agent. The process of training is described using Markov Decision Process; MDP is usually composed of a four-tuple, denoted as: <S, A, P, R>, where S represents the set of environment states, A represents the set of agent actions, P is the state transition probability of the environment, and R represents the reward function; at any time t in the training process, the agent perceives the state S of the environment t ∈S, and takes an action A t ∈A, so that the environment is transferred to the next state S t+1 ∈S, and gets a potential reward R t from the environment; The agent has two important functional components: the policy Π, which makes decisions based on state inputs and outputs actions, and the evaluation Q, which evaluates the current policy based on state and reward values. The function evaluated in the agent is to evaluate the quality of the current policy and update the policy by estimating the long-term cumulative reward obtained by the current policy; the expectation of the long-term cumulative reward is defined as a value function; the value function has two kinds, state value function and state-action value function; the state value function V Π (S) represents the agent in a certain state S of the environment t =S, the expectation of the long-term cumulative reward obtained when the action is decided according to the policy Pi: V Π (S) = E[G t |S t = S] (16) The state value function is represented as the recursive form of the sum of the expected value of the current reward and the next state value function weighted by the discount factor, which is called the Bellman expectation equation: V Π (S) = E[R t+1 + γV Π (S t+1 )|S t = S] (17) State-Action Value Function Q Π (S,A) is also called the Q-function, which represents the expected long-term cumulative reward that the agent would obtain if it were in a certain state of the environment, had performed a certain action, and then continued to make decisions according to the policy: Q Π (S,A) = E[G t |S t = S, A t = A] (18) The Q function can also be represented as the corresponding Bellman expectation equation: Q Π (S,A) = E[R t+1 + γV(S t+1 )|S t = S, A t = A] = R t+1 + γE a~Π [Q(S t+1 ,A t+1 )|S t = S, A t = A] (19) The basic method of deep reinforcement learning is to use two neural networks named Actor and Critic to replace the policy and evaluation in the agent, respectively, which is called the actor-critic algorithm. Both neural networks consist of an input layer, one or more hidden layers, and an output layer. The input of the Actor network input layer is the state of the environment, and the output of the output layer is the policy after fitting through the hidden layer. For deterministic algorithms, the network will directly give the value of the action. For stochastic strategies, the Actor will output a probability distribution, and then sample the action according to the probability distribution. The input of the Critic network input layer is the state of the environment and the reward value obtained after the action is performed, and the output of the output layer is the value of the Q function obtained through the fitting of the hidden layer. The process of input transmission and fitting output by the neural network is called forward propagation. For an n-layer neural network, the forward propagation formula for the propagation from the (l-1)-th layer to the l-th layer (l = 2…n) is: a l = σ(z l ) = σ(W l a l-1 + b l ) (20) where a l m×1 is a vector of data values for each neuron of the lth layer, where l = 2...n, m is the number of neurons in the layer, σ(z l ) is an activation function of the neural network, z l = W l a l-1 +b l is a transfer value of the data, is a diagonal matrix of weights for each neuron of the lth layer, b l m×1 is a vector of biases for each neuron of the lth layer; a set of weights and biases of the neurons of each layer of the neural network [W, b] is referred to as a parameter set of the neural network; the parameter set of the Actor network is denoted by φ, and the parameter set of the Critic network is denoted by θ; The process of setting a cost function of the neural network, using gradient descent method to iteratively optimize the cost function minimum value to find the optimal parameter set, so that the fitted data value is as close as possible to the actual data value, is called the back propagation process; the cost function of the neural network is defined as J(W,b,x,y), and the gradient of the cost function to the z l of the lth layer of the neural network is represented as: Then, the process of parameter updating of the l-th layer of the neural network using backpropagation is represented as: In the deep reinforcement learning method, the training of the intelligent controller is actually the training of the Actor and Critic neural networks; an experience replay pool is constructed to store the training data in real time: D = <S N-L ,A N-L ,P N-L ,R N-L… S N ,A N ,P N ,R N >, and L is the length of the experience replay pool; in the training process of the intelligent controller, a proper number of training data is sampled from the experience replay pool, a reasonable cost function is set according to the principle of reinforcement learning, the parameters of the Actor and Critic neural networks are continuously iterated and updated by using the above back propagation method, and finally the optimal neural network parameters φ * and θ * are obtained; Step 2.1.2: Implement the interaction between the neural network and the attitude-orbit coupling dynamics equation, and output the command thrust value and thruster deflection angle of the thruster through the neural network; In the process of training the intelligent controller, in order to improve the training efficiency, the environment in the intelligent control system of the flexible probe adopts the posture-orbit coupling motion dynamics model based on the simplified mathematical model of the structure of the flexible probe established in step one, and is trained in the simulation environment; the intelligent agent refers to the intelligent controller of the flexible probe; the input state of the intelligent controller is the real-time mass center position, velocity and planned position, velocity deviation of the probe, as well as the deviation between the real-time attitude and the expected attitude, which is represented as Sta=[Δρ m T Δv m T Δφ Δθ] T ; the output action is the command thrust value of each thruster and the deflection angle of the thruster mounted on the upper surface, which is represented as Act=[T 11 ,T 12 ,...T n1 ,T n2 ,α1,...α n ,β1,...β n ] T ; the intelligent controller needs to interact with the posture-orbit coupling motion dynamics module of the flexible probe in the simulation environment for training; in the process of training the intelligent controller in each round, the thrusters of each rigid control node generate thrust according to formula (15) after receiving the control command of the intelligent controller; under the action of the thrust, each rigid control node generates motion according to the dynamics law in formula (12); the position, velocity of the mass center of the probe and the overall attitude are calculated in real time according to formula (13) and formula (14) through the position and velocity of each rigid control node; the difference between the actual state of the probe and the planned expected state is obtained to obtain the input of the intelligent controller; the intelligent controller calculates the reward value obtained by training according to the input state error and the reward function set by the program; in the training process of the intelligent controller, the parameters of the input and output of the intelligent controller need to be standardized to improve the stability of the training process; the change range of each dimension of the input parameter is standardized to [-1, 1], and the change range of each dimension of the output parameter is [0, 1], and the relationship between the standardized parameter and the original parameter is represented as: wherein, Sta max , Sta min , Act max , Act min are the upper and lower bound values of the input and output; after the intelligent controller completes training in the simulation environment, it is deployed to the on-board computer of the flexible probe, and then applied in the actual asteroid attachment mission; in the actual asteroid attachment mission, the errors between the actual velocities and positions of each rigid control node measured by the sensor and the expected values are taken as the inputs of the intelligent controller, and the intelligent controller outputs the command thrust to each thruster to generate the control thrust. Step 2.2: Design the reward function and cost function based on the SAC algorithm to train the neural network, and construct the intelligent controller of the asteroid flexible probe. SAC algorithm is a stochastic policy-based algorithm in deep reinforcement learning. The basic principle is to maximize the long-term cumulative reward value and the policy information entropy during the training of the agent. The expression of the long-term cumulative reward is: where H(Π(·|S t )) is the information entropy of the policy; defined as: The generation of action is a process of information entropy reduction, which is contrary to the thermodynamic phenomenon described by the second law of thermodynamics. Therefore, the information entropy is negative. By maximizing the information entropy of the policy, the randomness of the policy is increased, and the probability of each action is as dispersed as possible, rather than concentrated on one action, thereby enhancing the exploration of the policy during the training process and improving the robustness of the control. Substituting equation (24) into equation (19), the soft Bellman expectation equation of the Q function is obtained: Further, the expression of the state value function is obtained: As shown in equation (27), the state value function can be expressed as the soft maximum value of the LogSumExp of the Q function in the action domain. This is the most important feature of the SAC algorithm, and its advantage is that the optimal policy does not fall into the local optimum or a certain optimum of the Q function, but approaches the global optimum or all optimal values of the Q function through a smooth method. Therefore, the policy based on SAC is wherein, represents the KL divergence, describes the degree of similarity between the distribution P and the distribution Q, the smaller the KL divergence, the highest degree of similarity of the two parts; by minimizing the KL divergence, the probability distribution π K (·|S t ) best approximation probability distribution so that the Q function has the form of soft maximum value; According to the basic principle of the SAC algorithm, the cost function of the Critic neural network in the agent based on the SAC algorithm is obtained as follows: In the process of training the agent based on SAC algorithm, two Critic networks with parameters θ1, θ2 are adopted, and the smaller one of the Q function values output by the two networks is taken each time. At the same time, in order to prevent the value of Q function from updating too fast and causing divergence, a target Critic network with slower update is set, and the network parameters are represented as , i = 1, 2, and the update mode is: The cost function of the Actor neural network is: During training, the Actor neural network outputs a Gaussian distribution Π φ with mean μ ~ N(μ, δ) t and variance δ t The agent samples an action from the Gaussian distribution, as shown in equation (31): A t = μ t + rand(0,1) δ t (31) where rand(0,1) represents a random number between 0 and 1. During the training process of the agent based on the SAC algorithm, the fixed temperature coefficient may cause instability in the training process due to the continuous change of the reward value. Therefore, the temperature coefficient needs to be adjusted in real time. When the policy explores the unknown area of the environment, the temperature coefficient should be appropriately increased to stimulate exploration. When the characteristics of the current environment are clear and the policy is fixed, the temperature coefficient should be lowered to make the training gradually converge. The adjustment of the temperature coefficient is realized by setting the cost function and solving the optimal value by the gradient descent method. The cost function used in the temperature coefficient adjustment process is as follows: The reward function of the intelligent controller is designed as follows: The reward function is mainly divided into three parts: state-related reward value, action-related reward value and penalty term; the state-related reward value r adopts a segmented reward manner, and e1 = e2 = 1e-3 is the accuracy value of the centroid position deviation and the attitude deviation; when the centroid position deviation and the attitude deviation are within the accuracy range, a negative weight is given to the sum as the reward value to promote the further convergence of the deviation; when the centroid position deviation exceeds the accuracy range but the current time value is less than the previous time value, a small negative reward is given, otherwise a large negative reward is given; the segmented reward manner is conducive to the stable convergence of the state deviation; the action reward value is -0.05|T cN | 2 , wherein T c = T c1 + T c2 + T c3 is the total thrust of all thrusters, and the reward value is conducive to reducing the thrust value of the thruster output so as to save the propellant consumption on the star; the meaning of the penalty term B is that when the square sum of the centroid position deviation is greater than the maximum value L of the square sum of the maximum centroid position deviation of the three axes or the relative distance between the two nodes is greater than the limit distance K, a negative reward is given as a penalty and the training in this round is exited. After designing the reward function of the intelligent controller, an intelligent controller with two Critic networks and one Actor network is established. The two Critic networks of the intelligent controller have the same structure. Each network has a state path and an action path, respectively. Each branch has several hidden layers, and each hidden layer has several neurons. The two branches are connected through a common hidden layer with several neurons, and finally output the Q value. The Actor network has a common hidden layer with several neurons, and then branches into two branches, each with a hidden layer with several neurons, to fit the mean and variance of the policy probability distribution.
4. The intelligent control method for asteroid flexible probe based on deep reinforcement learning SAC algorithm according to claim 3, characterized in that: In step three, The neural network trained based on the SAC algorithm processes the input state, i.e., the real-time mass center position and speed of the probe, the deviation of the planned position and speed, and the deviation between the real-time attitude and the expected attitude, so that the output action, i.e., the instruction thrust value of each thruster and the deflection angle of the upper surface-mounted thruster, can enable the flexible probe to effectively cope with the random changes of the structural parameters of the flexible probe caused by the complex deformation of the flexible material; and the orbit and attitude motion of the flexible probe are simultaneously controlled, so that the flexible probe can realize the expected adhering orbit of the high-precision tracking target task on the premise of stabilizing the attitude of the flexible probe in real time.
Citation Information
Patent Citations
Optimal cooperative control method for attachment of flexible spacecraft to asteroid
CN113325862A
Planetary soft landing control method and system based on reinforcement learning, and storage medium
CN113821057A