A method and system for training an artificial intelligence model of a spacecraft
By developing the SpaceSimGYM platform and introducing DDPG and MADDPG algorithms, the problems of insufficient multi-agent confrontation and unreasonable reward design of spacecraft simulation platforms in the existing technology are solved, and convenient simulation of customized scenarios and agent behaviors are realized, improving the efficiency and learning effect of spacecraft orbit interception.
Patent Information
- Application Number
- CN202210882083.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-26
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-07-26
AI Technical Summary
The existing technology is difficult to effectively deal with complex space combat environments, lacks a spacecraft simulation platform that supports multi-agent confrontation and cooperation, and the unreasonable reward design of reinforcement learning algorithms in spacecraft orbital interception missions leads to learning difficulties.
Develop the SpaceSimGYM platform, combines DDPG and MADDPG algorithms, realizes multi-agent confrontation decision-making, supports custom scenarios and agent behavior, and designs a guided reward mechanism to solve the problem of track interception.
It provides user-friendly custom combat scenarios, reduces algorithm development costs, supports multi-type mission simulation, and builds spacecraft confrontation scenarios through visualization platforms, achieving convenient Initial and Reset scenario switching, improving the efficiency of spacecraft orbit interception and the learning effect of agents.
Smart Images

Figure CN115293033B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of aerospace technology, and specifically, to a method and system for training an artificial intelligence model of a spacecraft. Background Art
[0002] As space gradually becomes an increasingly crowded and highly competitive domain, in order to adapt to the increasingly complex space operations, it is necessary to effectively counter the increasingly complex space offensive means;
[0003] Based on the existing self-developed spacecraft simulation platform SpaceSim platform, the present invention has developed a general-purpose platform SpaceSimGYM for the research and application of intelligent technologies for spacecraft offense and defense confrontation and gaming. The SpaceSimGYM reinforcement learning platform is developed based on SpaceSim. On the basis of SpaceSim, the establishment of an instruction-driven component architecture is completed, and relevant interface functions for supporting SpaceSimGYM to control the simulation are established in SpaceSim. Finally, through the integrated single-agent reinforcement learning algorithm DDPG and multi-agent reinforcement learning algorithm MADDPG, the design of a simulation system architecture supporting machine learning is realized. Summary of the Invention
[0004] The present invention proposes a method and system for training an artificial intelligence model of a spacecraft; by using the DDPG algorithm integrated in the SpaceSimGYM platform, an optimal interception strategy for space non-cooperative targets based on reinforcement learning under continuous thrust is realized.
[0005] The present invention is realized through the following technical solutions:
[0006] A system for training an artificial intelligence model of a spacecraft:
[0007] The training system includes a visual human-machine interface, a deduction environment module, an adversarial scheduling module, a multi-agent adversarial decision-making process module, and a combat scenario;
[0008] The combat scenario provides satellite data for the deduction environment module, and the deduction environment module transmits the original observation data and scenario data to the adversarial scheduling module through a scheduling interface;
[0009] The adversarial scheduling module receives the rule information of the adversarial rule library, and transmits the original observation data to the observation and reward module in the multi-agent adversarial decision-making process module, and transmits the scenario data to the multi-agent adversarial algorithm sub-module in the multi-agent adversarial decision-making process module;
[0010] The multi-agent adversarial algorithm sub-module in the multi-agent adversarial decision-making process module transmits action information to the adversarial scheduling module, and the adversarial scheduling module then transmits the action information back to the simulation environment module through the call interface, and finally displays it on the visual human-machine interface.
[0011] Furthermore, the visual human-machine interface is used for satellite-related settings, scenario-related settings, ground station-related settings, and the scenario JSON file created by scheduling and opposing the scheduling module.
[0012] Furthermore, the simulation environment module is used for satellite payload control and calculation, satellite orbit transfer calculation, and satellite orbital attitude control.
[0013] Furthermore, the adversarial scheduling module includes a Reset function, an Init function, and a Step function;
[0014] The Reset function is used for scenario recovery initialization;
[0015] The Init function is used for scenario file modification and reading;
[0016] The Step function is used for sending actions and instructions, obtaining the scenario environment, and recursive derivation of the scenario.
[0017] Furthermore, the training system supports users to customize the number and types of satellites called, can customize the functions, attributes, and observation capabilities of satellites, and supports the customization of the agent behavior reward rules;
[0018] The simulation environment module supports the use of the Python language and supports the integrated call of common deep learning frameworks such as TensorFlow and PyTorch.
[0019] A machine learning orbit interception method for a spacecraft artificial intelligence model training system:
[0020] The method specifically includes the following steps:
[0021] Step 1, on a near-Earth circular orbit with a fixed confrontation time, perform the observation settings of the red visible light reconnaissance satellite;
[0022] Step 2, set the actions of the red visible light reconnaissance satellite;
[0023] Step 3, design a reward return function to complete the orbit target interception;
[0024] Step 4, real-time status update: After each satellite executes an action, update the status in the scenario according to orbital dynamics, etc.;
[0025] S' = S + a;
[0026] Step 5, termination condition: When all the blue satellites are intercepted, the mission ends.
[0027] Further, in Step 1,
[0028] Suppose there are a total of n red red agents and a total of n blue blue agents. The self - state S t of the red visible - light reconnaissance satellite;
[0029] Step 1.1, the 8 - dimensional physical state S physic of the satellite at the current moment includes: satellite blood volume Health RED ; satellite remaining fuel Fuel RED , total refueled fuel of the satellite Fuel total ; absolute position (p x , p y , p z ) in the VVLH orbital system; absolute velocity (v x , v y , v z ) in the VVLH orbital system; then
[0030] S physic =(Health RED , Fuel RED , p x , p y , p z , v x , v y , v z )
[0031] Step 1.2, determine the 6n - dimensional combat state combat state: During the combat, the positions of both sides are accurate and transparent. Then, the red - side observation items include the relative positions and velocities of all n blue enemy satellites relative to the red visible - light reconnaissance satellite. blue If an enemy satellite has been destroyed, set the observation distance of this satellite to be extremely far to encourage the agent not to track this target. Therefore:
[0032]
[0033]
[0034] Among them: S i =(Dist i , p RELi,x , p RELi,y , p RELi,z , v RELi,x , v RELi,y , v RELi,z ) )
[0035] Step 1.3, determine the physical state other state of the own satellites in 8*(n red -1) dimensions: the physical states of all other own satellites:
[0036]
[0037] Where: S i =(Health REDi , Fuel REDi , p xi , p yi , p zi , v xi , v yi , v zi );
[0038] Step 1.4, the communication state communication state sent by other own satellites; temporarily None in the interception scenario.
[0039] Furthermore, in Step 2,
[0040] The red visible light reconnaissance satellite continuously advances throughout the entire combat duration until the target is intercepted;
[0041] Therefore, the control quantity of the satellite is only the propulsion direction; according to the orbital maneuver theory, completing the approach mission cannot reduce the relative velocity, so the action space is restricted, and the range of the propulsion direction angle is set as δ p ∈[-90°, 90°];
[0042] The propulsion direction angle is a discrete value, with each interval being 3 degrees, and there are a total of A = 60 optional propulsion direction angles; the center of the propulsion direction is the negative velocity direction of the satellite in the relative coordinate system;
[0043] a t =(δ p ).
[0044] Furthermore, in Step 3,
[0045] For the spacecraft orbital game problem, in the pursuit-evasion game scenario, the state of the defender is relative rest, periodic / aperiodic orbiting motion formed according to the C-W equation, or motion with maneuvers; the strategy of the attacker is given by the output of the neural network;
[0046] Set a guiding reward related to the relative distance, and give a reward according to the remaining fuel value at the moment of task completion; the guiding reward uses the difference between the historical value and the current value of the distance between the agent and the target as the step-by-step distance approach reward value within a round to guide the agent to learn;
[0047] Define the step-by-step reward value R for approaching the distance p as follows:
[0048] Suppose in the current step t, the historical distance from the agent to the target point is Dist t , and in the previous step t-1, the historical distance from the agent to the target point is Dist t-1 ;
[0049] Then, rewards or punishments are imposed according to whether the relative historical distance advances or retreats, where the punishment is 1.3 times the advance reward, driving the agent to approach the target as much as possible.
[0050] An electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0051] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the steps of the method described in any one of the above are implemented.
[0052] Advantages of the present invention
[0053] 1. The present invention provides a user-friendly spacecraft reinforcement learning platform that facilitates customizing combat scenarios, reducing the algorithm development cost.
[0054] 2. The present invention supports simulating the confrontation and cooperation processes between multiple agents and can support multiple types of tasks.
[0055] 3. While supporting the use of a visualization platform to construct a spacecraft confrontation scenario, the present invention achieves the convenience of Initial and Reset scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is the composition of the SpaceSimGYM module of the present invention;
[0057] Figure 2 It is the core architecture of SpaceSimGYM;
[0058] Figure 3 It is the relationship between the interface function and the interface operation of SpaceSimGYM;
[0059] Figure 4 It is the code structure of SpaceSimGYM;
[0060] Figure 5 It is the DDPG algorithm structure. DETAILED DESCRIPTION OF THE INVENTION
[0061] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0062] A spacecraft artificial intelligence model training system:
[0063] The training system includes a visual human-machine interface, a deduction environment module Spacesim, an adversarial scheduling module based on an adversarial rule library, a multi-agent adversarial decision-making process module (multi-agent adversarial scenario library Scenarios), and a combat scenario Scene;
[0064] The combat scenario Scene provides satellite data for the deduction environment module Spacesim, and the deduction environment module Spacesim transmits the original observation data and scenario data to the adversarial scheduling module through a scheduling interface;
[0065] The adversarial scheduling module receives the rule information of the adversarial rule library, transmits the original observation data to the observation and reward module in the multi-agent adversarial decision-making process module, and transmits the scenario data to the multi-agent adversarial algorithm sub-module in the multi-agent adversarial decision-making process module;
[0066] The multi-agent adversarial algorithm sub-module in the multi-agent adversarial decision-making process module transmits the action information to the adversarial scheduling module, and the adversarial scheduling module then transmits the action information back to the deduction environment module through a call interface, and finally displays it on the visual human-machine interface.
[0067] The training system uses SpaceSimGYM as the core architecture and combines it with the multi-agent adversarial algorithm sub-module of the multi-agent adversarial decision-making process module to implement multi-agent algorithm training and confrontation, as Figure 2 shown.
[0068] The visual human-machine interface is used for satellite-related settings, scenario-related settings, ground station-related settings, and creating a scenario JSON file by scheduling the adversarial scheduling module.
[0069] The deduction environment module is used for satellite payload control and calculation, satellite orbit transfer calculation, and satellite orbital attitude control.
[0070] The adversarial scheduling module includes a Reset function, an Init function, and a Step function;
[0071] The Reset function is used for scenario recovery initialization;
[0072] The Init function is used for modifying and reading scenario files;
[0073] The Step function is used for sending actions and instructions, obtaining the scenario environment, and recursively processing scenarios.
[0074] The training system supports users to customize the number of satellites called (number of agents) and the types of satellites called (types of agents), can customize the functions, attributes, and observation capabilities of satellites, and supports customizing the behavior reward rules of agents; etc.
[0075] The simulation environment module supports using the Python language and supports the integrated call of common deep learning frameworks such as TensorFlow and PyTorch.
[0076] A machine learning-based orbital interception method for a spacecraft artificial intelligence model training system:
[0077] Based on a machine learning-based orbital interception method, on a near-Earth circular orbit with a fixed confrontation time, according to the Markov decision process model of spacecraft orbital interception under continuous thrust;
[0078] In the case of random evasion by the blue reconnaissance satellite, use the reinforcement learning DDPG algorithm to solve the propulsion sequence decision problem of the red interception satellite under continuous propulsion.
[0079] The method specifically includes the following steps:
[0080] Step 1, on a near-Earth circular orbit with a fixed confrontation time, perform the observation setting of the red visible light reconnaissance satellite;
[0081] Step 2, set the actions of the red visible light reconnaissance satellite;
[0082] Step 3, design a reward return function to complete the orbital target interception;
[0083] Step 4, real-time state update: After each satellite executes an action, update the state in the scenario according to orbital dynamics, etc.;
[0084] S' = S + a
[0085] Step 5, termination condition: When all blue satellites are intercepted, the mission ends.
[0086] In Step 1,
[0087] Suppose there are a total of n red red agents and a total of n blue blue agents, and the self-state S t ;
[0088] Step 1.1, the 8-dimensional physical state S of itself at the current moment physic includes: the satellite's blood volume Health RED ; the remaining fuel Fuel of the satellite RED , the total refueled amount of fuel of the satellite total ; the absolute position (p x , p y , p z ) in the VVLH orbital system; the absolute velocity (v x , v y , v z ) in the VVLH orbital system; then there is
[0089] S physic = (Health RED , Fuel RED , p x , p y , p z , v x , v y , v z )
[0090] Step 1.2, determine the n blue *6-dimensional combat state: During the combat process, the positions of both sides are accurate and transparent. Then, the red side's observation items include a total of n blue enemy satellites, the relative positions and velocities relative to the red side's visible light reconnaissance satellite;
[0091] If an enemy satellite has been destroyed, set the observation distance of this satellite to be extremely far to encourage the agent not to track this target. Therefore:
[0092]
[0093] Among them: S i = (Dist i , p RELi,x , p RELi,y , p RELi,z , v RELi,x , v RELi,y , v RELi,z )
[0094] Step 1.3, determine the 8*(n red - 1)-dimensional physical state other state of its own satellites: the physical states of all other own satellites:
[0095]
[0096] Among them: S i = (Health REDi , FuelREDi , p xi , p yi , p zi , v xi , v yi , v zi );
[0097] Step 1.4, the communication state sent by other friendly satellites; temporarily None in the interception scenario.
[0098] In Step 2,
[0099] To complete the orbital interception target as soon as possible, the red visible light reconnaissance satellite continuously advances throughout the entire combat duration (2400 s) until the target is intercepted;
[0100] Therefore, the control quantity of the satellite is only the propulsion direction; according to the orbital maneuver theory, to complete the approach mission, the relative velocity cannot be reduced, so the action space is restricted, and the range of the propulsion direction angle is set as δ p ∈[-90°, 90°];
[0101] The propulsion direction angle is a discrete value, with each interval being 3 degrees, and there are a total of A = 60 optional propulsion direction angles; the center of the propulsion direction is the negative velocity direction of the satellite in the relative coordinate system;
[0102] a t =(δ p ).
[0103] In Step 3,
[0104] For the spacecraft orbital game problem, in the pursuit-evasion game scenario, the state of the defender can be relative rest, periodic / aperiodic orbiting motion formed according to the C-W equation, or motion with maneuvers; the strategy of the attacker is given by the output of the neural network;
[0105] Due to the uniqueness of the spacecraft pursuit-evasion problem, if the spacecraft only obtains a clear reward value at the end of the round, this delayed reward will make reinforcement learning extremely difficult;
[0106] Set a guiding reward related to the relative distance; to make the return function dense, use a combination of the guiding reward and the sparse reward after the task is successful, and the guiding reward must be related to the continuous variable - the relative distance.
[0107] Give rewards according to the remaining fuel value at the moment when the task is completed; for continuous thrust, the fuel consumption is proportional to the total duration of the thrust action. However, if the fuel consumption is directly placed in the guiding reward, it will interfere with the agent's learning. To reflect the fuel consumption situation during interception, rewards are only given according to the remaining fuel value at the moment when the task is completed, encouraging the agent to complete the interception as soon as possible.
[0108] The guiding reward uses the difference between the historical value and the current value of the distance between the agent and the target as the step-by-step reward value for distance advancement within a round step to guide the agent's learning.
[0109] Define the step-by-step reward value R for distance advancement p as follows:
[0110] Suppose in the current step t, the historical distance from the agent to the target point is Dist t , and in the previous step t - 1, the historical distance from the agent to the target point is Dist t-1 ;
[0111] Then, rewards or punishments are imposed according to whether the relative historical distance advances or retreats, where the punishment for retreat is 1.3 times the reward for advancement, driving the agent to approach the target as much as possible.
[0112] During the round, the guiding reward value is described as follows:
[0113] if Dist t <Dist t-1 :
[0114] R p =(Dist t-1 -Dist t )*5
[0115] else:
[0116] R p =(Dist t-1 -Dist t )*8
[0117] The sparse reward value after the task is successfully completed is described as follows:
[0118] if Dist t <5.0
[0119] R f =F + Fuel red *20
[0120] where F is the fixed winning reward value, which is 80.
[0121] Simulation platform structure design:
[0122] 1) The SpaceSimGYM reinforcement learning platform is developed based on the SpaceSim software. The SpaceSim software can support the whole-process simulation and analysis of space missions and provides relatively complete tool functions related to space calculations.
[0123] 2) SpaceSimGYM is a research and training platform for multi-agent collaborative cooperation and game confrontation algorithms in the space background. Based on several preset combat scenarios and engagement rules, this platform supports users to customize the number of satellites called (the number of agents) and the types of satellites called (the types of agents). It can customize the functions, attributes, and observation capabilities of satellites and support the customization of agent behavior reward rules, etc.
[0124] 3) The SpaceSimGYM platform provides a good foundation for using deep reinforcement learning algorithms to solve the problem of multi-agent distributed confrontation in the space background. It specifically opens the RL-API interface for multi-agent deep reinforcement learning. The environment supports the implementation of algorithms using the Python language and supports the integrated call of common deep learning frameworks such as TensorFlow and PyTorch. The module composition is as Figure 1 shown.
[0125] Figure 3 In it, SpaceSimGYM simulates interface operations for scenario settings, simulation control, and interface instruction control through the initialall, CommandAdd, and StepAllGet interface functions respectively.
[0126] Figure 4 In it, the SpaceSimGYM program consists of modules such as a multi-agent training environment (Train_maeos), a generation trainer (trainers), a combat scenario library (Scenes), a call interface (core), an adversarial rule library (scenario), and an adversarial scheduler (environment).
[0127] Figure 5 In it, the Deep Deterministic Policy Gradient (DDPG) reinforcement learning algorithm uses a multi-hidden-layer feedforward neural network in deep learning to fit the mapping relationship from states to actions.
[0128] An electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0129] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the steps of the method described in any one of the above are implemented.
[0130] The memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory of the method described in the present invention is intended to include but not limited to these and any other suitable types of memories.
[0131] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a high-density digital video disc (DVD)), or a semiconductor medium (such as a solid state disc (SSD)), etc.
[0132] In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware processor, or executed and completed by a combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0133] It should be noted that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0134] The above has introduced in detail a method and system for training an artificial intelligence model of a spacecraft, and expounded on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A spacecraft artificial intelligence model training system, characterized in that: The training system uses SpaceSimGYM as the core architecture, combines with the multi-agent adversarial algorithm sub-module of the multi-agent adversarial decision-making process module to achieve multi-agent algorithm training and confrontation. The SpaceSimGYM reinforcement learning platform is developed based on SpaceSim. On the basis of SpaceSim, the establishment of the instruction-driven component architecture is completed, and relevant interface functions for supporting SpaceSimGYM to control the simulation are established in SpaceSim. Finally, through the integrated single-agent reinforcement learning algorithm DDPG and multi-agent reinforcement learning algorithm MADDPG, the design of the simulation system architecture supporting machine learning is realized; The training system includes a visual human-machine interface, a deduction environment module, an adversarial scheduling module, a multi-agent adversarial decision-making process module, and a combat scenario; The combat scenario provides satellite data for the deduction environment module SpaceSim, and the deduction environment module SpaceSim transmits the original observation data and scenario data to the adversarial scheduling module through the scheduling interface; The adversarial scheduling module receives the rule information of the adversarial rule library, transmits the original observation data to the observation and reward module in the multi-agent adversarial decision-making process module, and transmits the scenario data to the multi-agent adversarial algorithm sub-module in the multi-agent adversarial decision-making process module; The multi-agent adversarial algorithm sub-module in the multi-agent adversarial decision-making process module transmits the action information to the adversarial scheduling module, and the adversarial scheduling module then transmits the action information back to the deduction environment module through the call interface, and finally displays it on the visual human-machine interface; SpaceSimGYM simulates the interface operations of scene setting, simulation control, and interface instruction control through the initialall, CommandAdd, and StepAllGet interface functions respectively.
2. The system according to claim 1, wherein: The visual human-machine interface is used for satellite-related settings, scene-related settings, ground station-related settings, and creating a scene JSON file by scheduling the adversarial scheduling module.
3. The system according to claim 2, wherein: The deduction environment module is used for satellite payload control and calculation, satellite orbit transfer calculation, and satellite orbit attitude control.
4. The system according to claim 3, wherein: The adversarial scheduling module includes a Reset function, an Init function, and a Step function; The Reset function is used for scene recovery initialization; The Init function is used for modifying and reading the scene file; The Step function is used for sending actions and instructions, obtaining the scene environment, and recursively advancing the scene.
5. The system according to claim 4, wherein: The training system supports users to customize the number and types of satellites called, can customize the functions, attributes, and observation capabilities of satellites, and supports customizing the reward rules for agent behavior; The deduction environment module supports the use of the Python language and supports the integrated call of common deep learning frameworks such as TensorFlow and PyTorch.
6. A machine learning orbital interception method for a spacecraft artificial intelligence model training system according to any one of claims 1 to 5, characterized in that: The method specifically includes the following steps: Step 1, on a near-earth circular orbit with a fixed confrontation time, set up the observation of the red visible light reconnaissance satellite; Step 2, set the actions of the red visible light reconnaissance satellite; Step 3, design a reward return function to complete the orbital target interception; Step 4, real-time state update: after each satellite executes an action, update the state in the scenario according to orbital dynamics; S' = S + a; Step 5, termination condition: when all blue satellites are intercepted, the mission ends.
7. The method according to claim 6, characterized in that: In Step 1, Suppose there are a total of n Red agents red and a total of n Blue agents blue The own state S of the Red visible light reconnaissance satellite t ; Step 1.1, the 8-dimensional physical state S of itself at the current moment physic including: satellite blood volume Health RED ; remaining fuel of the satellite Fuel RED , total refueling amount of the satellite Fuel total ; absolute position in the VVLH orbital system (p x , p y , p z ); absolute velocity in the VVLH orbital system (v x , v y , v z ); then there is S physic = (Health RED , Fuel RED , p x , p y , p z , v x , v y , v z ) Step 1.2, determine n blue *The combat state of 6 dimensions: During the combat process, the positions of both sides are accurate and transparent. Then, among the observation items of the red side, there are a total of n blue enemy satellites, and the relative positions and speeds relative to the visible light reconnaissance satellite of the red side itself If the enemy satellite has been destroyed, set the observation distance of this satellite to be extremely far, so as to encourage the agent not to track this target. Therefore: Where: S i =(Dist i , p RELi,x , p RELi,y , p RELi,z , v RELi,x , v RELi,y , v RELi,z ) Step 1.3, determine the physical state other state of the own satellites in 8*(n red -1)-dimension: the physical states of all other own satellites: Where: S i = (Health REDi , Fuel REDi , p xi , p yi , p zi , v xi , v yi , v zi ); Step 1.4, the communication state sent by other friendly satellites; it is temporarily None in the interception scenario.
8. The method according to claim 7, wherein: In Step 2, The red visible light reconnaissance satellite continuously advances throughout the combat duration until the target is intercepted; Therefore, the control quantity of the satellite is only the propulsion direction; according to the orbital maneuver theory, the relative velocity cannot be reduced to complete the approach mission, so the action space is restricted, and the range of the propulsion direction angle is set to δ p ∈[-90°, 90°]; The propulsion direction angle is a discrete value, with an interval of every 3 degrees, and there are a total of A = 60 optional propulsion direction angles; the center of the propulsion direction is the negative velocity direction of the satellite in the relative coordinate system; a t = (δ p ).
9. The method according to claim 8, wherein: In Step 3, For the spacecraft orbital game problem, in the pursuit-evasion game scenario, the state of the defense side is relative rest, periodic / aperiodic orbiting motion formed according to the C-W equation, or motion with maneuvers; the strategy of the attacking side is given by the output of the neural network; Set a guiding reward related to the relative distance, and give a reward according to the remaining fuel value at the moment of task completion; the guiding reward uses the difference between the historical value and the current value of the distance between the agent and the target as the step-by-step distance approach reward value within a round to guide the agent to learn; Define the step-by-step reward value R for approaching the distance p as follows: Assume that in the current step t, the historical distance of the agent to the target point is Dist t , and in the previous step t-1, the historical distance of the agent to the target point is Dist t-1 ; Then impose a reward or penalty according to whether the relative historical distance advances or retreats, where the penalty for retreat is 1.3 times the reward for advance, driving the agent to approach the target as much as possible.
10. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 6 to 8.
Citation Information
Patent Citations
Intelligent game confrontation platform
CN112295229A
Behavior imitation training method for air intelligent game
CN113221444A