Autonomous orbit change decision-making method and device for energy-limited satellite based on deep reinforcement learning
Through the self-game method of deep reinforcement learning, satellites learn effective orbit change strategies in zero samples, solving the problems of weak independent decision-making capabilities and low intelligence, and achieving flexible response and task execution of autonomous orbit change under limited energy.
Patent Information
- Application Number
- CN202211492768.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-25
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-11-25
AI Technical Summary
The existing satellites have weak independent orbit change decision-making capabilities and low intelligence, and cannot dynamically adjust orbit change actions based on energy consumption and residual amount, resulting in the inability to independently perform tasks and quickly respond to emergencies in complex space environments.
The self-game method based on deep reinforcement learning is adopted, and the reinforcement learning algorithm is used to train satellites to learn effective orbital change strategies in zero samples, and the deep reinforcement learning model is used to preprocess situation information and control decision behaviors to achieve independent orbital change.
It has achieved that satellites have independent decision-making capabilities under limited energy, can flexibly respond to emergencies, improve their intelligence, and can learn flexible and changeable autonomous orbit change strategies in zero samples, solving the problems of weak independent decision-making capabilities and low intelligence.
Smart Images

Figure CN115828741B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of satellite orbit change decision-making, and specifically relates to a method and device for autonomous orbit change decision-making of energy-limited satellites based on deep reinforcement learning. Background Art
[0002] Satellites are widely used in communications, meteorology, navigation and positioning, and data relay. With the advent of globalization, satellites play a vital role in safeguarding national security. Reconnaissance satellites, communications satellites, navigation and positioning satellites, meteorological satellites, and mapping satellites, each perform their respective functions and collaborate with each other to ensure national and information security. However, unpowered satellite trajectories are easily predictable, potentially exposing satellite information. Therefore, satellites must be able to flexibly adjust their orbits based on the satellite situation in space. This flexible orbit-changing capability plays a crucial role in ensuring satellite survival and information security. However, research on intelligent, autonomous orbit-changing capabilities tailored to missions and the real-time situation of satellites in space is virtually nonexistent.
[0003] Unlike aircraft and other autonomous vehicles, satellites have limited energy resources, and orbital maneuvers consume significant amounts of energy. Once these resources are depleted, satellites can no longer maneuver or change orbits, remaining reliant on unpowered revolution. Therefore, satellites must be controlled to maneuver and change orbits within their limited energy resources.
[0004] Existing satellite orbit change technology solutions all rely on pre-setting information such as the number of orbit changes, the number of orbit change cycles, and the control amount for each orbit change, allowing the satellite to change orbit according to the preset target. However, these solutions are unable to autonomously make intelligent orbit change decisions. For example, Chinese patent application number CN200710301588.4 discloses a method for autonomous satellite orbit change, while Chinese patent application number CN201110409628.3 discloses a method for optimizing geostationary satellite orbit change strategies. Currently proposed solutions that allow satellites to change orbit according to preset targets are unable to autonomously make intelligent orbit change decisions, nor can they dynamically adjust the value of orbit change actions based on the energy consumption and remaining energy of different actions.
[0005] Obviously, the existing technology has the following shortcomings: (1) Weak autonomous decision-making ability. Current satellites do not have the ability to change orbits autonomously and cannot make autonomous orbit change decisions to perform tasks in complex space environments; (2) Low intelligence level. Current satellite missions are mostly determined manually and lack the ability to dynamically adjust, evolve and predict behavior. They cannot respond quickly to emergencies; (3) Current satellite orbit change methods cannot dynamically adjust the orbit change action according to energy consumption and remaining energy, and cannot complete the orbit change while saving energy as much as possible. Summary of the Invention
[0006] One of the objectives of the present invention is to provide an autonomous orbit change decision-making method for energy-limited satellites based on deep reinforcement learning. By generating data through self-game of reinforcement learning, the satellite is controlled to perform orbit change actions under limited energy, and more diverse and possible autonomous intelligent orbit change strategies are explored.
[0007] To achieve the above object, the technical solution adopted by the present invention is:
[0008] A method for autonomous orbit change decision-making of a satellite with limited energy based on deep reinforcement learning, comprising:
[0009] Step 1: Initialize the satellite simulation operating environment and start a game training;
[0010] Step 2: Determine the target orbit, target position, and target attitude of the orbit change according to the preset mission of the satellite agent to be controlled;
[0011] Step 3: Acquire the situation information of the satellite agent to be controlled in the satellite simulation operation environment;
[0012] Step 4: Preprocess the situation information and input it into the deep reinforcement learning model to obtain the decision-making behavior of the satellite agent to be controlled;
[0013] Step 5: Control the satellite agent to be controlled to change its orbit according to the decision behavior;
[0014] Step 6: Calculate the reward for the current track change, including:
[0015] Calculate the reward for the t-th step as:
[0016]
[0017] In the formula, R(s t ,u t ) is the satellite agent based on the state s at step t t Execute action u t After the reward, is the amount of energy consumed when executing step t, is the remaining energy of the satellite agent after the t-th step is executed, is the distance between the satellite agent and the target position after the tth step is executed, is the relative angle between the running direction of the satellite agent and the target position after the t-th step is executed;
[0018] Then the cumulative reward function of step H is:
[0019]
[0020] Where R H (τ) is the cumulative reward of the satellite agent in the Hth step, and τ is the cumulative state-action sequence of the satellite agent (s0, u0, ..., s H ,u H ), 0≤t≤H;
[0021] Step 7: Store and update the deep reinforcement learning model;
[0022] Step 8: The satellite agent changes its orbit to the target orbit, reaches the target position, and has the target attitude as the basis for determining whether the satellite agent has completed the orbit change. If the satellite agent has not completed the orbit change, the process returns to step 3 and continues. If the satellite agent has completed the orbit change, the current game ends and step 9 is executed. If the orbit change is completed within the time limit, the current game is judged as a win. If the orbit change is not completed within the time limit, the current game is judged as a loss.
[0023] Step 9: Determine whether the satellite agent's orbit change success rate under the deep reinforcement learning model decision reaches a preset threshold. If not, return to step 1 to continue training. Otherwise, terminate the training and output the deep reinforcement learning model as the decision model.
[0024] Step 10: Use the decision model output after training to make autonomous orbit change decisions for the satellite.
[0025] Several optional methods are also provided below, but they are not intended to be additional limitations on the above-mentioned overall solution. They are merely further supplements or optimizations. Under the premise that there are no technical or logical contradictions, each optional method can be combined separately for the above-mentioned overall solution, or multiple optional methods can be combined.
[0026] Preferably, the pre-processing of the situation information includes: screening preset situation information, and normalizing the screened situation information.
[0027] Preferably, the filtered situation information includes: satellite position, satellite attitude, sun position, satellite speed, orbit, relative distance from the target position, and relative angle from the target position.
[0028] Preferably, the controlling the satellite agent to be controlled to change orbit according to the decision-making behavior includes:
[0029] Control the satellite agent to change orbit under chemical thrust or electric thrust according to the decision-making behavior;
[0030] Among them, under chemical thrust, the decision behavior output by the deep reinforcement learning model is the radial, lateral and normal control acceleration of the spacecraft (Δv x ,Δv y ,Δv z), controls the satellite agent to change its orbit according to the radial, lateral and normal control acceleration of the spacecraft;
[0031] Among them, under electric thrust, the decision behavior output by the deep reinforcement learning model is (T, α, β), where T is the thrust size, α is the angle between the projection of the thrust vector in the orbital plane and the perpendicular direction of the spacecraft's geocentric vector, and β is the angle between the thrust vector and the orbital plane. According to (T, α, β), the radial, lateral and normal control accelerations of the spacecraft are calculated to control the satellite agent to change orbit, and the radial control acceleration of the spacecraft f is calculated. r , lateral control acceleration f t and the normal control acceleration f n They are:
[0032]
[0033] Where m is the mass of the satellite agent.
[0034] The present invention provides an autonomous orbit change decision-making method for energy-limited satellites based on deep reinforcement learning. (1) An effective orbit change strategy is learned through a self-game method under zero-sample conditions through reinforcement learning algorithm training; (2) With autonomous decision-making capabilities, the satellite can make effective orbit change strategies in a timely manner according to the current battle situation, thereby being able to flexibly respond to emergencies; (3) It can make intelligent orbit change decisions under limited energy according to the real-time situation of the satellite, guide the satellite to perform tasks, solve the problems of weak autonomous decision-making capabilities, low intelligence level, and energy consumption of the satellite, and can learn flexible and changeable autonomous orbit change strategies through reinforcement learning algorithms under zero-sample conditions.
[0035] The second purpose of the present invention is to provide an autonomous orbit change decision-making device for energy-limited satellites based on deep reinforcement learning. By generating data through self-game of reinforcement learning, the satellite is controlled to perform orbit change actions under limited energy, and more diverse and possible autonomous intelligent orbit change strategies are explored.
[0036] To achieve the above object, the technical solution adopted by the present invention is:
[0037] A device for autonomous orbit change decision-making of a satellite with limited energy resources based on deep reinforcement learning, comprising a processor and a memory storing a plurality of computer instructions. When the computer instructions are executed by the processor, the steps of a method for autonomous orbit change decision-making of a satellite with limited energy resources based on deep reinforcement learning are implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 A schematic diagram of the interaction structure between the satellite simulation operating environment and the satellite intelligent agent of the present invention;
[0039] Figure 2This is a flowchart of the training part of the autonomous orbit change decision method for energy-limited satellites based on deep reinforcement learning of the present invention;
[0040] Figure 3 This is a flowchart of the reasoning application part of the autonomous orbit change decision-making method for energy-limited satellites based on deep reinforcement learning in the present invention. DETAILED DESCRIPTION
[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0043] To address the shortcomings of existing technologies in satellite autonomous orbit change decision-making, this embodiment proposes a satellite autonomous orbit change decision-making method based on deep reinforcement learning. By generating data through self-games using reinforcement learning, the method controls the satellite to perform orbit change actions under limited energy, exploring more diverse and possible autonomous intelligent orbit change strategies.
[0044] The method of this embodiment is divided into two parts: training and reasoning application. It is necessary to first perform training to obtain a decision model for satellite intelligent body orbit change and then perform reasoning application.
[0045] Among them, Figure 1 、 2 As shown, the specific implementation plan of the training part is as follows:
[0046] Step 1: Initialize the satellite simulation operating environment.
[0047] The satellite simulation environment is initialized, and the satellite agent units are initialized according to the number of TLE elements to obtain the initial position and initial attitude of the satellite agent in the set scenario. After initialization, a round of game training begins.
[0048] Step 2: Determine the target orbit, target position, and target attitude of the orbit change according to the preset mission of the satellite intelligent body to be controlled.
[0049] Usually, satellite agents have certain operating goals during operation. In this simulation operation environment, tasks are assigned to satellite agents, which include information such as target orbit, target position, target attitude, and the content of the task to be executed.
[0050] Step 3: Obtain the situation information of the satellite agent to be controlled in the satellite simulation operation environment.
[0051] This example incorporates a competitive mindset into satellite orbit change decisions, generating data through self-play through reinforcement learning to control satellite orbit changes under limited energy constraints. Acquiring situational information requires acquiring real-time situational information about the satellite agents in the game, including their real-time position, attitude, orbit, energy consumption, and remaining energy.
[0052] Step 4: After preprocessing the situation information, input it into the deep reinforcement learning model to obtain the decision-making behavior of the satellite agent to be controlled.
[0053] Preprocessing primarily involves filtering and normalizing situational information that influences satellite strategy selection. This filtered information can be adjusted based on the application, such as satellite position, attitude, sun position, velocity, orbit, relative distance from the target, and angle relative to the target. It can also include information such as the energy consumed and remaining energy during the current orbit change.
[0054] When normalizing the filtered situation information, first count all the collected original situation data obs orh The maximum and minimum values of each dimension, all the maximum and minimum values compose the situation maximum vector obs max and the situation minimum vector obs min , the normalized preprocessing method is as follows:
[0055]
[0056] The global state vector obs after normalization preprocessing nor All data will be between [0,1]. Training with normalized data can make the model converge faster.
[0057] The deep reinforcement learning model in this embodiment can use algorithms such as DDPG and PPO. The state and action of the satellite multi-agent at step t are represented by s t and u t express.
[0058] Step 5: Control the satellite agent to be controlled to change its orbit according to the decision-making behavior.
[0059] This embodiment controls the satellite intelligent body to change orbit under chemical thrust or electric thrust according to the decision-making behavior.
[0060] Among them, under chemical thrust, the decision behavior output by the deep reinforcement learning model is the radial, lateral and normal control acceleration of the spacecraft (Δv x ,Δv y,Δv z ), controls the satellite agent to change its orbit according to the radial, lateral and normal control acceleration of the spacecraft.
[0061] Among them, under electric thrust, the decision behavior output by the deep reinforcement learning model is (T, α, β), where T is the thrust size, α is the angle between the projection of the thrust vector in the orbital plane and the perpendicular direction of the spacecraft's geocentric vector, and β is the angle between the thrust vector and the orbital plane. According to (T, α, β), the radial, lateral and normal control accelerations of the spacecraft are calculated to control the satellite agent to change orbit, and the radial control acceleration of the spacecraft f is calculated. r , lateral control acceleration f t and the normal control acceleration f n They are:
[0062]
[0063] Where m is the mass of the satellite agent.
[0064] In addition, when the satellite intelligent body performs the orbit change action, its situation information will change. At this time, it is necessary to calculate the latest situation information, which corresponds to the situation information required in this embodiment, that is, to calculate the satellite position, satellite attitude, sun position, satellite speed, orbit, relative distance from the target position, and relative angle from the target position.
[0065] The above parameter calculations are all based on existing logic. For example, the calculation of the orbit of the satellite intelligent body can refer to the orbit calculation after the orbit change proposed in the "Study on the Combined Orbit Change Scheme of GEO Satellite Electric Propulsion and Chemical Propulsion" proposed by Tian Baiyi and other authors. The calculation of the amount of energy required for the orbit change action can also refer to the above document. After determining the amount of energy required for the orbit change action, the remaining energy of the satellite intelligent body can be obtained. Since the satellite speed corresponds to the orbit, the satellite speed is also determined when the orbit is determined. The satellite position at the next moment can be calculated based on the satellite position, orbit and execution time of the orbit change action at the previous moment. The satellite attitude is directly measured by the sensor on the position intelligent body. The system presets the sun position under each coordinate system (spherical center coordinate system, satellite coordinate system, etc.), and the corresponding sun position can be obtained under the given time. The relative distance to the target position and the relative angle to the target position can be calculated based on the position relationship, which will not be described in detail in this embodiment.
[0066] Step 6: Calculate the reward for the current track change.
[0067] The rewards in this example are related to the outcome of the game and the target task. To guide the satellite to change its orbit to a specified orbit, the reward needs to include the target position, and the target orbit and target position are set in the reward parameters of the deep reinforcement learning model. In the case of limited energy, the energy consumption and remaining energy of each step need to be added to the reward. In this example, the reward for a single step t is calculated as:
[0068]
[0069] In the formula, R(s t ,u t ) is the satellite agent based on the state s at step t t Execute action u t After the reward, is the amount of energy consumed when executing step t, is the remaining energy of the satellite agent after the t-th step is executed, is the distance between the satellite agent and the target position after the tth step is executed, is the relative angle between the satellite agent's running direction and the target position after the tth step is executed.
[0070] Then the cumulative reward function of step H is:
[0071]
[0072] Where R H (τ) is the cumulative reward of the satellite agent in the Hth step, and τ is the cumulative state-action sequence of the satellite agent (s0, u0, ..., s H ,u H ), 0≤t≤H.
[0073] In the policy-based reinforcement learning method, the objective function is:
[0074]
[0075] In the formula, the variable It generally refers to the parameters involved in the mapping from state to action of the satellite agent. Different parameters represent different strategies. E represents expectation. Indicates that the parameter The following strategy, Indicates that the parameter The strategy under s t Execute action u under the condition t The ultimate goal of reinforcement learning is to find the optimal parameters , so that the objective function reaches its maximum value, which can be described by the formula:
[0076]
[0077] Step 7: Store and update the deep reinforcement learning model. The model update in this embodiment adopts a conventional deep reinforcement learning model update method, such as back propagation to update model parameters.
[0078] Step 8: The satellite intelligent body changes its orbit to the target orbit and reaches the target position and target attitude as the basis for judging whether the satellite intelligent body has completed the orbit change. If the satellite intelligent body has not completed the orbit change (that is, the satellite intelligent body has not changed its orbit to the target orbit and has not reached the target position and target attitude), return to step 3 to continue execution. If the satellite intelligent body completes the orbit change (that is, the satellite intelligent body changes its orbit to the target orbit and reaches the target position and target attitude), the current game ends and step 9 is executed. If the orbit change is completed within the limited time, the current game is judged as a victory. If the orbit change is not completed within the limited time, the current game is judged as a failure.
[0079] Step 9: Determine whether the satellite agent's trajectory change success rate under the deep reinforcement learning model reaches a preset threshold. If not, return to step 1 to continue training. Otherwise, terminate training and output the deep reinforcement learning model as the decision model. The success rate can be the ratio of the number of games won to the total number of games played.
[0080] In addition, if Figure 3 As shown, the technical solution of the reasoning part of this embodiment is as follows, which mainly uses the decision model output after the training to make autonomous orbit change decisions for the satellite:
[0081] (1) Initialize the satellite simulation operating environment, initialize the satellite intelligent unit according to the number of TLE roots, and start the deduction.
[0082] (2) Determine the target orbit, target position, and target attitude of the orbit change based on the preset mission of the satellite intelligent body to be controlled.
[0083] (3) Obtain global satellite agent situation information in real time.
[0084] (4) Screen and preprocess the information of the obtained global situation.
[0085] (5) Input the filtered and processed situation data into the satellite orbit change decision model to obtain the current orbit change decision behavior.
[0086] (6) Change the orbit according to the orbit change decision action to obtain new satellite position, satellite attitude, orbit, energy consumption, energy remaining and other situation information.
[0087] (7) Update the new situation information to the front end for display.
[0088] (8) Repeat steps 3-7 until the end of a round.
[0089] In another embodiment, the present application also provides an autonomous orbit change decision-making device for energy-limited satellites based on deep reinforcement learning, comprising a processor and a memory storing a plurality of computer instructions. When the computer instructions are executed by the processor, the steps of the autonomous orbit change decision-making method for energy-limited satellites based on deep reinforcement learning are implemented.
[0090] Regarding the specific limitations of the autonomous orbit change decision-making device for energy-limited satellites based on deep reinforcement learning, please refer to the limitations of the autonomous orbit change decision-making method for energy-limited satellites based on deep reinforcement learning above, which will not be repeated here.
[0091] The memory and processor are electrically connected, directly or indirectly, to enable data transmission or interaction. For example, these components may be electrically connected to each other via one or more communication buses or signal lines. The memory stores a computer program executable on the processor. The processor executes the computer program stored in the memory to implement the deep reinforcement learning-based autonomous orbit change decision-making method for energy-limited satellites according to an embodiment of the present invention.
[0092] The memory may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. The memory is used to store a program, and the processor executes the program after receiving an execution instruction.
[0093] The processor may be an integrated circuit chip with data processing capabilities. The processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor.
[0094] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0095] The above-described embodiments merely illustrate several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that a person skilled in the art would be able to make numerous modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A method for autonomous orbit change decision-making of energy-limited satellites based on deep reinforcement learning, characterized by: The autonomous orbit change decision-making method for energy-limited satellites based on deep reinforcement learning includes: Step 1: Initialize the satellite simulation operating environment and start a game training; Step 2: Determine the target orbit, target position, and target attitude of the orbit change according to the preset mission of the satellite agent to be controlled; Step 3: Acquire the situation information of the satellite agent to be controlled in the satellite simulation operation environment; Step 4: Preprocess the situation information and input it into the deep reinforcement learning model to obtain the decision-making behavior of the satellite agent to be controlled; Step 5: Control the satellite agent to be controlled to change its orbit according to the decision behavior; Step 6: Calculate the reward for the current track change, including: Calculate the reward for the t-th step as: In the formula, R(s t ,u t ) is the satellite agent based on the state s at step t t Execute action u t After the reward, is the amount of energy consumed when executing step t, is the remaining energy of the satellite agent after the t-th step is executed, is the distance between the satellite agent and the target position after the tth step is executed, is the relative angle between the running direction of the satellite agent and the target position after the t-th step is executed; Then the cumulative reward function of step H is: Where R H (τ) is the cumulative reward of the satellite agent in the Hth step, and τ is the cumulative state-action sequence of the satellite agent (s0, u0, ..., s H ,u H ), 0≤t≤H; Step 7: Store and update the deep reinforcement learning model; Step 8: The satellite agent changes its orbit to the target orbit, reaches the target position, and has the target attitude as the basis for determining whether the satellite agent has completed the orbit change. If the satellite agent has not completed the orbit change, the process returns to step 3 and continues. If the satellite agent has completed the orbit change, the current game ends and step 9 is executed. If the orbit change is completed within the time limit, the current game is judged as a win. If the orbit change is not completed within the time limit, the current game is judged as a loss. Step 9: Determine whether the satellite agent's orbit change success rate under the deep reinforcement learning model decision reaches a preset threshold. If not, return to step 1 to continue training. Otherwise, terminate the training and output the deep reinforcement learning model as the decision model. Step 10: Use the decision model output after training to make autonomous orbit change decisions for the satellite.
2. The method for autonomous orbit change decision-making of energy-limited satellites based on deep reinforcement learning as claimed in claim 1, characterized in that: The pre-processing of the situation information includes: screening preset situation information, and normalizing the screened situation information.
3. The autonomous orbit change decision method for energy-limited satellites based on deep reinforcement learning as claimed in claim 2, characterized in that: The filtered situation information includes: satellite position, satellite attitude, sun position, satellite speed, orbit, relative distance from the target position, and relative angle from the target position.
4. The autonomous orbit change decision method for energy-limited satellites based on deep reinforcement learning as claimed in claim 1, characterized in that: The controlling of the satellite agent to be controlled to change orbit according to the decision-making behavior includes: Control the satellite agent to change orbit under chemical thrust or electric thrust according to the decision-making behavior; Among them, under chemical thrust, the decision behavior output by the deep reinforcement learning model is the radial, lateral and normal control acceleration of the spacecraft (Δv x ,Δv y ,Δv z ), controls the satellite agent to change its orbit according to the radial, lateral and normal control acceleration of the spacecraft; Among them, under electric thrust, the decision behavior output by the deep reinforcement learning model is (T, α, β), where T is the thrust size, α is the angle between the projection of the thrust vector in the orbital plane and the perpendicular direction of the spacecraft's geocentric vector, and β is the angle between the thrust vector and the orbital plane. According to (T, α, β), the radial, lateral and normal control accelerations of the spacecraft are calculated to control the satellite agent to change orbit, and the radial control acceleration of the spacecraft f is calculated. r , lateral control acceleration f t and the normal control acceleration f n They are: Where m is the mass of the satellite agent.
5. A deep reinforcement learning-based autonomous orbit change decision-making device for energy-limited satellites, comprising a processor and a memory storing a plurality of computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Satellitic self-determination orbital transfer method
CN101219713A
Method for optimizing orbital transfer strategy of geostationary orbit satellite
CN102424116A
Orbit maneuver optimization method for multiple satellites and multiple reconnaissance targets
CN113408063A
Spacecraft artificial intelligence model training method and system
CN115293033A