Spacecraft Morphology Control Method and System Based on Evolutionary Algorithms and Reinforcement Learning
By employing a spacecraft morphology control method based on evolutionary algorithms and reinforcement learning, the problem of reliance on manual operation for modular spacecraft reconfiguration has been solved. This enables autonomous optimization and control of spacecraft morphology, improves environmental adaptability and mission agility, and shortens the research and development cycle and costs.
Patent Information
- Application Number
- CN202510942197.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-07-09
AI Technical Summary
Existing modular spacecraft reconfiguration methods rely on manual operation by astronauts or assistance from orbital robots, which suffer from limited reconfiguration opportunities, poor autonomy, low efficiency, and lack of flexibility. Furthermore, traditional spacecraft structural designs are fixed and difficult to adapt to complex space environments and diverse mission requirements.
A spacecraft morphology control method based on evolutionary algorithms and reinforcement learning is adopted. Through outer-loop morphology evolution and inner-loop learning training, the morphology of modular spacecraft is autonomously generated. By utilizing the inner and outer loop algorithm architecture of deep evolutionary reinforcement learning, combined with intelligent agents, sensors and actuators, the autonomous optimization and control of spacecraft morphology is achieved.
It has enabled the autonomous generation of the form of modular spacecraft, improved environmental adaptability and mission agility, shortened the research and development cycle and development costs, and enhanced the service life and mission adaptability of spacecraft.
Smart Images

Figure CN120440311B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of spacecraft design optimization and autonomous control, and in particular to a spacecraft morphology control method and system based on evolutionary algorithms and reinforcement learning. Background Technology
[0002] Traditional spacecraft employ highly customized design methods, resulting in relatively fixed topologies and mission adaptability. Functions must be pre-designed according to specific mission requirements. During on-orbit operation, their structure, function, and operating mode remain largely unchanged, leading to high development costs, poor flexibility, and long development cycles. Modular spacecraft can effectively solve these problems.
[0003] By defining existing single units or subsystems as standardized modules and designing controllable connection mechanisms, modular spacecraft can achieve flexible combination and form changes of different modules. By leveraging the advantages of low R&D costs of standardized modules, economical and efficient mass production can be achieved. Furthermore, by upgrading and replacing modules, the service life of spacecraft can be significantly extended, the design of the entire satellite system can be simplified, the production and development cycle can be shortened, and the spacecraft can be better adapted to the needs of complex space environments and diverse missions.
[0004] However, existing modular spacecraft still face limitations in reconfiguration methods during actual operation. Their reconfiguration primarily relies on manual operation by astronauts or assisted operation by orbital robots, which suffers from limited reconfiguration opportunities, poor autonomy, low efficiency, and insufficient flexibility.
[0005] Meanwhile, traditional spacecraft have fixed structural designs that require pre-designed functions based on mission requirements. Existing modular spacecraft reconfiguration methods mainly rely on manual operation by astronauts or assisted operation by orbital robots, making it difficult to achieve true on-orbit autonomous reconfiguration capabilities during operation. Summary of the Invention
[0006] To address the problems existing in the prior art, the present invention aims to provide a spacecraft morphology control method based on evolutionary algorithms and reinforcement learning. This method achieves autonomous morphological generation of modular spacecraft through continuous alternation between outer-loop morphological evolution and inner-loop learning training. Another objective of the present invention is to provide a spacecraft morphology control system based on evolutionary algorithms and reinforcement learning that implements the above method.
[0007] To achieve the above objectives, this invention provides a spacecraft morphology control method based on evolutionary algorithms and reinforcement learning, specifically as follows:
[0008] S1. Establish an initial population, in which individuals are spacecraft of different forms with different functional modules and different combinations of the number of modules;
[0009] S2. Perform inner-loop initialization learning training on all individuals in the population and calculate the fitness value of each individual;
[0010] S3. Select individuals with high fitness values to form an elite population;
[0011] S4. Using the genetic and mutation operations in the genetic algorithm, perform uniform crossover and single-point mutation on the elite population to generate elite offspring;
[0012] S5. Inner Loop Reinforcement Learning: Train the elite offspring obtained in step S4;
[0013] S6. Morphological Assessment: Analyze the spacecraft's mission performance in the current mission scenario to obtain a comprehensive morphological assessment result. Use the assessment result as the fitness value of the spacecraft's morphology and use it as the evaluation basis for external circulation evolution.
[0014] S7. Forming the optimal individual: After completing the inner loop reinforcement learning training, the elite offspring are added to the original population. The individuals in the population are sorted according to their fitness values, and the individuals with the lowest fitness are eliminated until the number of individuals in the population is the same as the initial population size. When the outer loop reaches the set number of generations, the optimal individual is output.
[0015] Furthermore, steps S1, S2, S3, S4, and S7 constitute the morphological evolution process of the outer loop, while steps S5 and S6 constitute the reinforcement learning process of the inner loop.
[0016] Furthermore, the inner loop includes an agent, a sensor, a controller, and an actuator; the agent is the object of study, that is, an individual in the population, and the sensor, controller, and actuator realize the agent's perception and control behavior.
[0017] Furthermore, the reward values obtained during the inner loop learning and training process will be used to calculate the fitness values of individuals in the population and output to the outer loop.
[0018] Furthermore, the external circulation system completes the generational evolution of the population based on the fitness values of individual individuals, thereby achieving the evolution of the spacecraft's morphology.
[0019] Furthermore, in step S1, the definition is... As the initial population set in the genetic algorithm, the initial population contains a total of Individual.
[0020] Furthermore, in step S3, random selection is made from the initial population. Individuals, conducting a tournament, totaling Group tournaments are held, and the winners of each group tournament, i.e., the individuals with the highest fitness values, form the elite population as parents. The elite population consists of... It consists of individuals.
[0021] Furthermore, in step S4, based on the elite genetic algorithm, by using the crossover and mutation operator in the population iteration, the elite population is obtained through uniform crossover and single-point mutation. Offspring of elites.
[0022] Furthermore, in step S5, the inner-loop reinforcement learning stage employs a nested PPO algorithm. The elite offspring will receive learning and training.
[0023] This invention relates to a spacecraft morphology control system based on evolutionary algorithms and reinforcement learning. This system is used to implement the spacecraft morphology control method based on evolutionary algorithms and reinforcement learning described above.
[0024] This invention fully integrates the characteristics of the space environment and mission requirements, and is based on an inner and outer loop algorithm architecture of deep evolutionary reinforcement learning. Through the continuous alternation of outer loop morphological evolution and inner loop learning training, it realizes the autonomous generation of the morphology of modular spacecraft. Attached Figure Description
[0025] Figure 1 This is a framework diagram of the present invention;
[0026] Figure 2 This is a flowchart of the invention;
[0027] Figure 3 This is a flowchart of the inner-loop reinforcement learning algorithm;
[0028] Figure 4 It is the attitude control reward function curve;
[0029] Figure 5 This is the structure diagram of the outer-loop evolutionary algorithm;
[0030] Figure 6 This is a graph showing the change in fitness values of the outer ring population. Detailed Implementation
[0031] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0032] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0033] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0034] The following combination Figures 1-6 Specific embodiments of the present invention will be described in detail below. It should be understood that the specific embodiments described herein are for illustrative and explanatory purposes only and are not intended to limit the present invention.
[0035] This invention relates to a spacecraft morphology control method and system based on evolutionary algorithms and reinforcement learning. It utilizes an inner and outer loop algorithm framework based on deep evolutionary reinforcement learning to achieve modular spacecraft morphology optimization for different mission scenarios. The structural framework diagram of this invention is shown below. Figure 1 As shown.
[0036] Among them, the definition As the initial population set in the genetic algorithm , , ∈ , , , The individuals in the population represent different forms of spacecraft, namely spacecraft with different functional modules and different combinations of module numbers. The initial population consisted of... Individual.
[0037] It should be clarified that the form of a spacecraft is defined as its configuration, that is, the combination of different functional modules and different numbers of modules constitutes a spacecraft.
[0038] The external circulation is the process by which a population undergoes intergenerational evolution.
[0039] First, an initial population is established. Then, all individuals in the population undergo inner-loop initialization training, and the fitness value of each individual is calculated. Here, the fitness value is the result of the fitness function calculation and is a quantitative indicator of individual quality, determined by the morphological evaluation rule proposed later. Subsequently, individuals with higher fitness values are selected to form an elite population. Then, using the genetic and mutation operations in the genetic algorithm, uniform crossover and single-point mutation are performed on the elite population to generate elite offspring.
[0040] The internal cycle is the process by which individuals within a population undergo lifelong evolution:
[0041] The inner loop contains agents, sensors, controllers, and actuators. Agents are the objects of study, i.e., individuals within a population, while sensors, controllers, and actuators enable the agents to perceive and control their behavior.
[0042] First, under the given spacecraft morphology in the outer loop, a training environment is set up. Through continuous interaction between the population of individuals and the mission scenario, and based on the established morphology evaluation rules, the optimal controller, output configuration, and attitude control strategy are trained. The current generation's optimal configuration is formed based on the reward values obtained from mission completion. The reward values obtained during the inner loop learning and training process are used to calculate the fitness values of the population of individuals and output to the outer loop. The outer loop then completes the generational evolution of the population based on the fitness values of the individuals, thereby achieving the evolution of the spacecraft morphology.
[0043] This invention relates to a spacecraft morphology control method based on evolutionary algorithms and reinforcement learning. The specific steps are as follows: Figure 2 As shown:
[0044] The process involves: ① Generating an initial population; ② Inner-loop initialization and training; ③ Obtaining an elite population; ④ Generating elite offspring; ⑤ Inner-loop reinforcement learning; ⑥ Morphological evaluation; and ⑦ Forming the optimal individual. Here, ①, ③, ④, and ⑦ are outer-loop morphological evolution algorithms, while ②, ⑤, and ⑥ are inner-loop reinforcement learning algorithms.
[0045] ① Generate the initial population:
[0046] In the initial stage, a modular spacecraft is first modeled based on different functional modules and different numbers of modules. Vectors are defined. The following shows the functional modules inside the spacecraft and their number.
[0047] (1);
[0048] in, This indicates the number of functional modules in a modular spacecraft, specifically... Indicates the number of attitude and orbit control modules, Indicates the number of propulsion modules, Indicates the number of energy modules, Indicates the number of control modules, Indicates the number of computing modules, Indicates the number of communication modules, Indicates the number of optical sensing modules, Indicates the number of radar sensing modules, Indicates the number of electronic positioning modules, This indicates the number of electromagnetic interference modules. Based on the functional modules inside the spacecraft, their number, and the connection constraints between them, individuals in the population are randomly generated to establish the initial population of the modular spacecraft.
[0049] ②Inner loop initialization learning and training:
[0050] Individuals from the initial population are sent to the inner loop to interact with the task scenario for initial learning and training, obtaining the fitness value of each individual in the population. During the inner loop training, a multi-process approach is used to allocate a separate process for the learning and training of each individual, thereby achieving parallel computation and improving computational efficiency. The individual fitness function is expressed as:
[0051] (2);
[0052] in, Indicates the number of typical task scenarios. Indicates the spacecraft in the mission scenario The reward value obtained through inner-loop reinforcement learning Representing the task scenario The upper limit of reward value, This represents the standard weight of the reward function, used to unify the magnitude of reward and fitness values.
[0053] The inner-loop reinforcement learning phase employs nested PPO algorithms to train individuals across the entire initial population. The PPO algorithm utilizes an Actor-Critic architecture to implement spacecraft configuration and attitude / orbit control strategies. The Actor model uses a neural network. Fit the control strategy function of the spacecraft, where These are the parameters of the Actor network to be optimized. The Actor network takes into account the environmental situation information of the task scenario. Output spacecraft's configuration and maneuvering strategies The Critic model also uses a neural network to fit an evaluation function. ,in These are network parameters. Output The PPO algorithm requires sampling a large amount of data to form an experience buffer for training the neural network and updating the policy. However, a single Actor network cannot train and update its parameters while performing the sampling task. Therefore, a sampling network is designed. Specifically designed for empirical sampling, in which The sampling network parameters are used, while the Actor network is only responsible for continuously updating the parameters and strategies using empirical data.
[0054] The PPO algorithm uses the expected value of the dynamic weight advantage function as its objective cost function, as shown below:
[0055] (3);
[0056] in, , These represent the current time step in the task process. The task scenario environment state and the actions taken by the Actor network; Indicates the current parameter Under the strategy followed, the mission scenario is affected by the spacecraft performing actions. From state Transition to the next state The probability of This indicates that in the parameter The probability of [the outcome]. The objective of the optimization function is to maximize the advantage function. In probability weights Expectations The advantage function can be expressed as:
[0057] (4);
[0058] in, Indicates the spacecraft at time step Environmental rewards obtained through reinforcement learning reward functions This represents the reward discount factor. Indicates from state The expected cumulative discount reward to be obtained, Indicates from state The expected cumulative discount reward to be obtained.
[0059] The reinforcement learning training framework with nested dual PPO algorithms described in this invention is as follows: Figure 3 As shown. vectors Reinforcement learning with input inner loop configuration. Vectors are processed sequentially. The functional modules corresponding to each element in ( Individual posture control module, One propulsion module... (Electromagnetic interference modules), based on module connection constraints and the installation order (attitude and orbit control, propulsion, etc.), randomly selects locations for each module sequentially to form an initial configuration. This method uses a single attitude and orbit control module as the initial unit, iteratively solving the connection space that satisfies the connection constraints of the current configuration, and randomly selecting locations to add new modules to update the configuration, until all functional modules have been traversed. This method yields an initial configuration with connectivity, rationality, and randomness.
[0060] The configuration strategy output by the Actor1 network in the configuration PPO algorithm is a sequence of actions, which is then used to transform the initial configuration. The action sequence specifies the flipping direction of each module. Each action is processed sequentially: the motion space of the corresponding module is calculated; if the flip is feasible (within the motion space) and the configuration after the flip satisfies connectivity and constraints, the flip is performed and the configuration is updated. This process is iterated based on the new configuration until the sequence is processed, resulting in the training configuration.
[0061] The training configuration is input into the attitude and orbit control reinforcement learning process, and attitude and orbit control policies for typical task scenarios are obtained through reinforcement learning. The training configuration completes the state transition by executing the attitude and orbit control policy output by the Actor2 network, and obtains the new attitude and orbit control reinforcement learning environment state and corresponding reward value. The reward function for attitude and orbit control reinforcement learning is described below. It consists of action reward, boundary penalty, success reward, and failure penalty. The Critic2 network evaluates the Actor2 control strategy based on the attitude and orbit control reinforcement learning environment state and reward value, and optimizes the Actor2 network parameters accordingly to improve attitude and orbit control performance. The attitude and orbit control reward function curve of this invention is shown in the figure below. Figure 4 As shown. Simultaneously, the attitude control effect of the controller is evaluated using performance evaluation rules, and the evaluation result is output to the configuration reinforcement learning.
[0062] Subsequently, the Critic1 network evaluates the configurational reconfiguration policy of Actor1 using the current state and evaluation information, and completes the parameter tuning of the Actor1 network. Configurational reinforcement learning evaluates Actor1's policy through a configurational reward function (numerical form). The expression is as follows:
[0063] ;
[0064] in, These represent the spacecraft's mission completion rate reward, mission completion time reward, mission cost penalty, platform cost penalty, and cumulative reward for reinforcement learning during the interaction of typical mission scenarios within the inner loop, respectively. The calculation method is shown in "Morphological Evaluation⑥".
[0065] ③ Obtain an elite population:
[0066] Randomly select from the initial population Individuals, conducting a tournament, totaling Group Championship. Among them... Tournament size refers to the number of individuals randomly selected for comparison in each competition. The winners of each tournament, i.e., the individuals with the highest fitness values, form the elite population as their parents. The elite population consists of... It consists of individuals.
[0067] ④ Producing elite offspring:
[0068] Elite populations are obtained through uniform crossover and single-point mutation. Elite Offspring. Based on the elite genetic algorithm, crossover and mutation operators are used in population iterations to improve population diversity, expand the algorithm's optimization space, and allow the spacecraft morphology to evolve fully. The crossover operator uses a uniform crossover method. The population is traversed, and when an individual triggers the crossover operator through the crossover probability, that individual becomes the parent, and another individual is found as the mother for uniform crossover. The chromosomes of each individual are traversed, i.e., each gene position in its morphological variable matrix is traversed. If a gene position triggers uniform crossover through the uniform crossover probability, the genes at the corresponding gene positions in the parent and mother are exchanged, i.e., the number of functional modules corresponding to the individual is adjusted. During this process, it is checked whether the module type corresponding to the gene position meets the specific module connection constraints in the model constraints. If a specific module connection exists, the genes of the specific connection modules at the exchanged gene positions in the parents are changed to ensure that the offspring always satisfy the module constraints. If uniform crossover is not triggered, the gene position remains unchanged, and the process continues to the next operation position until all gene positions in the morphological variable matrix are completely traversed, completing one uniform crossover.
[0069] The mutation operator employs a single-point mutation method. The population is traversed, and when an individual triggers the mutation operator through mutation probability, the chromosome of that individual is selected for operation; that is, a gene locus is randomly selected from the individual's morphological variable matrix for single-point mutation. During single-point mutation, the number of genes at that gene locus, i.e., the corresponding functional modules, is adjusted to ensure random mutation and reconstruction within the allowed range of module number constraints. Simultaneously, it is checked whether the module type corresponding to that gene locus conforms to the specific module connection constraints in the model constraints. If specific module connections exist, the number of modules at the specific connection module gene locus corresponding to the mutated gene locus is synchronously changed to ensure that the offspring always satisfy the module constraints. This process completes one morphological variable mutation operation. The complete outer loop evolutionary algorithm structure is as follows: Figure 5 As shown.
[0070] ⑤ Inner-loop reinforcement learning:
[0071] The inner-loop reinforcement learning phase calls the same nested PPO algorithm as in step ② for initialization and training. Unlike step ②, this step does not require inputting the entire population into the inner-loop reinforcement learning; instead, it only needs the data obtained in step ④ of the outer loop. Elite offspring are trained and their fitness values are obtained through a fitness function.
[0072] ⑥ Morphological assessment:
[0073] This paper analyzes the mission performance of spacecraft in the current mission scenario and proposes a comprehensive morphological evaluation method. The reward function value obtained during training in the current mission scenario and the mission execution status serve as the basis for the comprehensive morphological evaluation, which includes a component for mission completion rate. Task completion time Task rewards Costs and consequences Platform costs Spacecraft morphology evaluation rules, including indicators. Mission completion rate. Assess the spacecraft mission's progress and completion time. The number of training steps for attitude and orbit control reinforcement learning is evaluated, and its expression is as follows:
[0074] ;
[0075] in, The number of time steps for the attitude and orbit control task execution. This marks the completion of the attitude and orbit control task. The reinforcement learning and training task is now complete. =1, otherwise =0. Task reward This refers to the numerical value of the reward function for attitude and trajectory control reinforcement learning training; cost and consequences. The energy consumption used to evaluate the execution of a spacecraft mission is expressed as follows:
[0076] ;
[0077] in, For the first The control force output by the secondary spacecraft This refers to the number of execution steps for the attitude and orbit control task. Platform cost. The total number of functional modules required for a spacecraft to complete its mission is expressed as follows:
[0078] ;
[0079] in, This represents the total number of all functional modules.
[0080] Configuration reward function As a fitness value for spacecraft morphology, it serves as an evaluation criterion for external circulation evolution to support morphological evolution.
[0081] ⑦ Optimal Individual:
[0082] After completing the inner-loop reinforcement learning training, elite offspring with high fitness values are added to the original population containing their parents, expanding the population size. Then, individuals within the population are sorted according to their fitness values, and the individuals with the lowest fitness are eliminated until the population size matches the initial population size. When the outer loop reaches a certain number of generations, the optimal individual is output. The fitness value change process of the outer-loop population is as follows: Figure 6 .
[0083] This invention addresses the problems of traditional fixed-structure spacecraft, such as long development cycles, high development costs, fixed functions, and difficulty in enabling microsatellites to flexibly cope with complex external environments. It proposes a modular modeling approach for spacecraft. Subsequently, a deep evolutionary reinforcement learning framework with inner and outer loops is employed. The outer loop uses an elite genetic algorithm to topologically reorganize functional modules to achieve morphological evolution of the spacecraft. The inner loop uses the expected value of a dynamic weighted advantage function as the objective cost function and employs reinforcement learning to learn spacecraft configuration and attitude / orbit control strategies. Data generated during the inner loop's training process is used to calculate the fitness values of individuals in the population and output to the outer loop. The outer loop then uses the fitness values of individuals in the population to achieve morphological evolution of the spacecraft until the population converges to the optimal individual.
[0084] This invention achieves morphological optimization and control of modular spacecraft through continuous alternation of morphological evolution and training, thereby improving the spacecraft's environmental adaptability, rapid response, and mission agility.
[0085] Any process or method described in the flowcharts of this invention or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, which can be implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device. The computer-readable medium can be any medium containing a program for storage, communication, propagation, or transmission for use by the execution system, apparatus, or device, including read-only memory, magnetic disks, or optical disks.
[0086] In the description of this specification, references to terms such as "embodiment," "example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, those skilled in the art can combine or combine the different embodiments or examples described in this specification and the features therein without causing contradiction.
[0087] While embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions, and alterations to the above embodiments within the scope of the present invention.
Claims
1. A spacecraft morphology control method based on evolutionary algorithms and reinforcement learning, characterized in that, Specifically: S1. Establish an initial population, in which individuals are spacecraft of different forms with different functional modules and different combinations of the number of modules; S2. Perform inner-loop initialization learning training on all individuals in the population and calculate the fitness value of each individual; S3. Select individuals based on fitness values to form an elite population; S4. Using the genetic and mutation operations in the genetic algorithm, perform uniform crossover and single-point mutation on the elite population to generate elite offspring; S5. Inner Loop Reinforcement Learning: Train the elite offspring obtained in step S4; S6. Morphological Assessment: Analyze the spacecraft's mission performance in the current mission scenario to obtain a comprehensive morphological assessment result. Use the assessment result as the fitness value of the spacecraft's morphology and use it as the evaluation basis for external circulation evolution. S7. Forming the optimal individual: After completing the inner loop reinforcement learning training, the elite offspring are added to the original population. Individuals in the population are sorted according to their fitness values, and individuals with the lowest fitness are eliminated until the number of individuals in the population is the same as the initial population. When the outer loop reaches the set number of generations, the best individual is output. In step S2, during the inner-loop training process, a multi-process approach is used to allocate a separate process for the learning and training of each individual, thereby achieving parallel computing and improving computational efficiency; the individual fitness function is expressed as: ; in, Indicates the number of typical task scenarios. Indicates the spacecraft in the mission scenario The reward value obtained through inner-loop reinforcement learning Representing the task scenario The upper limit of reward value, This represents the standard weight of the reward function, used to standardize the magnitude of reward and fitness values; The inner-loop reinforcement learning phase uses a nested PPO algorithm to train individuals in the entire initial population. The PPO algorithm uses an Actor-Critic architecture to implement spacecraft configuration and attitude control strategies. The PPO algorithm uses the expected value of the dynamic weight advantage function as its objective cost function. ; in, , These represent the current time step in the task process. The task scenario environment state and the actions taken by the Actor network; Indicates the current parameter Under the strategy followed, the mission scenario is affected by the spacecraft performing actions. From state Transition to the next state The probability of This indicates that in the parameter The probability of the following; the Actor model uses a neural network. Fit the control strategy function of the spacecraft, where These are the Actor network parameters to be optimized; the Actor network uses environmental situation information from the task scenario. Output spacecraft's configuration and maneuvering strategies ; These are the sampling network parameters.
2. The spacecraft morphology control method based on evolutionary algorithm and reinforcement learning according to claim 1, characterized in that, Steps S1, S2, S3, S4, and S7 constitute the morphological evolution process of the outer loop, while steps S5 and S6 constitute the reinforcement learning process of the inner loop.
3. The spacecraft morphology control method based on evolutionary algorithms and reinforcement learning according to claim 2, characterized in that, The inner loop contains agents, sensors, controllers, and actuators; the agent is the object of study, that is, an individual in the population, and the sensors, controllers, and actuators realize the agent's perception and control behavior.
4. The spacecraft morphology control method based on evolutionary algorithms and reinforcement learning according to claim 2, characterized in that, The reward values obtained during the inner loop learning and training process will be used to calculate the fitness values of individuals in the population and output to the outer loop.
5. The spacecraft morphology control method based on evolutionary algorithm and reinforcement learning according to claim 4, characterized in that, The external circulation system, based on the fitness values of individual individuals in the population, completes the generational evolution of the population and achieves the evolution of the spacecraft's morphology.
6. The spacecraft morphology control method based on evolutionary algorithm and reinforcement learning according to claim 1, characterized in that, In step S1, define As the initial population set in the genetic algorithm, the initial population contains a total of Individual.
7. The spacecraft morphology control method based on evolutionary algorithm and reinforcement learning according to claim 1, characterized in that, In step S3, randomly select from the initial population Individuals, conducting a tournament, totaling Group tournaments are held, and the winners of each group tournament, i.e., the individuals with the highest fitness values, form the elite population as parents. The elite population consists of... It consists of individuals.
8. The spacecraft morphology control method based on evolutionary algorithm and reinforcement learning according to claim 1, characterized in that, In step S4, based on the elite genetic algorithm, the elite population is obtained through uniform crossover and single-point mutation by using crossover and mutation operators in population iteration. Offspring of elites.
9. The spacecraft morphology control method based on evolutionary algorithm and reinforcement learning according to claim 1, characterized in that, In step S5, the inner-loop reinforcement learning stage employs a nested PPO algorithm. The elite offspring will receive learning and training.
10. A spacecraft morphology control system based on evolutionary algorithms and reinforcement learning, characterized in that, The system is used to implement the spacecraft morphology control method based on evolutionary algorithms and reinforcement learning as described in any one of claims 1-9.
Citation Information
Patent Citations
Task constraint-based spacecraft attitude control system on-orbit reconstruction method
CN107608208A
Layout planning method based on reinforcement learning and genetic algorithm
CN115758981A