A fast constellation configuration design method based on parallel reinforcement learning and genetic algorithm

By combining parallel reinforcement learning and genetic algorithm methods, the problem of easily falling into local optimal solutions and low computational efficiency in constellation configuration design is solved, and an efficient constellation configuration design is achieved, and the optimal solution that meets multi-objective optimization is found.

CN119647293BActive Publication Date: 2025-05-06NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510166858.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-06
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

The existing intelligent optimization methods are prone to fall into local optimal solutions in the constellation configuration design, and the relationship between solution space and configuration parameters is not clear enough, and the calculation efficiency is low.

Method used

A fast constellation configuration design method based on parallel reinforcement learning and genetic algorithm is adopted. By combining the parallel reinforcement learning module and the genetic algorithm module, a multi-objective optimization model is built, parallel computing is used to improve iteration efficiency, and the optimal constellation configuration parameters are found.

Benefits of technology

It greatly improves the iterative process of reinforcement learning, can effectively process a large amount of emergency observation target data, improves the efficiency of constellation configuration design, and finds a set of constellation configuration parameters with the largest Reward, the smallest Loss and meets the constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119647293B_ABST
    Figure CN119647293B_ABST
Patent Text Reader

Abstract

The fast constellation configuration design method based on parallel reinforcement learning and genetic algorithm provided by the present invention comprises the following steps: randomly generating observation targets, constructing observation target data sets, and giving initial constellation configuration parameters at the same time; under constraint conditions, constructing a multi-objective optimization model for constellation configuration parameters, the multi-objective optimization model comprising a parallel reinforcement learning module and a genetic algorithm module; the parallel reinforcement learning module performs training interaction according to the input observation target data sets and constellation configuration parameters, and outputs batch candidate solutions to the genetic algorithm module; the genetic algorithm module evaluates and optimizes the received candidate solutions, obtains the optimal constellation configuration parameters and outputs them. The present invention can maximize the efficiency of reinforcement learning under limited graphics card resources, thereby improving the efficiency of configuration optimization for large-scale satellite constellations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of satellite intelligence technology, and in particular to a fast constellation configuration design method based on parallel reinforcement learning and genetic algorithm. Background Art

[0002] Building a large-scale, low-cost, fast-response low-orbit giant constellation to provide all-day, all-weather, high-performance, and timely ground information services is the forefront and hot spot of the current aerospace development, and represents the development direction of satellite system construction. Building an efficient satellite constellation requires meeting the goal of timely coverage of emergency situations and providing rapid response. Rapid constellation configuration optimization requires comprehensive weighing of multiple objective factors.

[0003] The design and optimization of satellite constellations is a basic and crucial link in the construction of new low-orbit satellite constellation systems. The goal of satellite constellation design and optimization is mainly to determine the constellation satellite configuration and its orbital parameters. Ideally, the optimal design is found in all possible solution set spaces, but because there are many parameters, even under multiple constraints, the solution set space is still very large. Therefore, in current research, multiple data dimensions and parameter calculations lead to a certain amount of time spent on satellite configuration design calculations. Therefore, the core of constellation configuration design research is how to use limited computing resources to quickly solve the optimal solution in an infinite solution set space. Therefore, it is of great significance to study the rapid optimization design of constellation configurations.

[0004] From the evolution history of constellation configuration design methods, they mainly include traditional methods and intelligent optimization methods. Traditional methods mainly include geometric analysis methods and simulation comparison methods. Such methods are simple to calculate and easy to implement, but the design efficiency is low and the system indicators are less considered. Intelligent optimization methods mainly include genetic algorithms, particle swarm algorithms, simulated annealing algorithms, etc. Although such methods consider the optimal overall performance under multi-objective and multi-constraint conditions, they have defects such as easy to fall into local optimal solutions and insufficient search accuracy, and the analytical relationship between the solution space obtained by the algorithm and the configuration parameters is not clear enough.

[0005] Therefore, a technical solution is urgently needed to solve the problems of existing intelligent optimization methods in constellation configuration design, such as easy to fall into local optimal solutions, fuzzy relationship between solution space and configuration parameters, large solution set space and low computational efficiency. Summary of the invention

[0006] In view of the defects of the prior art, the present invention provides a fast constellation configuration design method based on parallel reinforcement learning and genetic algorithm.

[0007] In order to achieve the above technical objectives, the specific technical solutions adopted by the present invention are as follows:

[0008] The present invention provides a fast constellation configuration design method based on parallel reinforcement learning and genetic algorithm, comprising the following steps:

[0009] S1. Randomly generate observation targets, construct observation target data sets, and give initial constellation configuration parameters;

[0010] S2. Under constraints, construct a multi-objective optimization model for constellation configuration parameters, wherein the multi-objective optimization model includes a parallel reinforcement learning module and a genetic algorithm module;

[0011] S3, the parallel reinforcement learning module performs training interaction according to the input observation target data set and constellation configuration parameters, and outputs batch candidate solutions to the genetic algorithm module;

[0012] S4. The genetic algorithm module evaluates and optimizes the received candidate solutions, obtains the optimal constellation configuration parameters and outputs them.

[0013] Further, in S2, the constraints include observation task requirement constraints, on-board resource storage constraints, on-board energy constraints, and observation time window constraints;

[0014] The observation task requirement constraints satisfy:

[0015] ;

[0016] ;

[0017] in, Imaging payload types for mission requirements; The type of payload carried on the satellite; is the load resolution; Imaging resolution for mission requirements;

[0018] The on-board resource storage constraints satisfy:

[0019] ;

[0020] in, For satellites; for the goal; Storage space for satellites; For satellite To target The time when the imaging observations were made; is the storage consumption rate;

[0021] The on-board energy constraints satisfy:

[0022] ;

[0023] in, Indicates the target Subtasks In satellite The kth time window of is observed; Indicates the target The duration of the visible time window; Indicates satellite Observe the power consumption per unit time; Indicates satellite Total power consumption;

[0024] The observation time window constraints satisfy:

[0025] ;

[0026] in, For the goal In satellite All possible observation windows within the imaging observation period; is the observation window; and They represent the time when the observation starts and ends respectively; and They represent the start time and end time of the observation window respectively.

[0027] Furthermore, in S1, the constellation configuration parameters include: the number of orbits, the number of satellites on the orbital plane, the orbital altitude, the orbital intersection angle, the right ascension of the ascending node, and the true anomaly angle.

[0028] Further, in S3, the parallel reinforcement learning module is trained using an Actor-Critic structure;

[0029] The Actor-Critic structure includes an Actor structure and a Critic structure;

[0030] The Actor structure is used to generate a probability for determining the next action and select an optimal action strategy based on the observed state input; the Actor structure includes a parallel Transformer module and a softmax module;

[0031] The Critic structure is used to evaluate the value estimate of a given state and the reward of a specific state based on the observed state input; the Critic structure includes a parallel Transformer module.

[0032] Furthermore, the parallel Transformer module includes a position encoding block, a residual encoding block, a Transformer encoding block and a decoding block.

[0033] Furthermore, in the parallel Transformer module, the encoding calculation of the position encoding block calls the computing power of GPU-0; the encoding calculation of the residual encoding block calls the computing power of GPU-1; and the encoding calculation of the Transformer encoding block calls the computing power of GPU-2.

[0034] Further, the S3 comprises the following steps:

[0035] S31, reading observation target data set and constellation configuration parameters;

[0036] S32, creating an observation state, constructing a four-dimensional dynamic state space and a three-dimensional static state space;

[0037] S33, using the four-dimensional dynamic state space and the three-dimensional static state space as the observation space of the parallel reinforcement learning module, and using the given number of observation targets as the action space of the parallel reinforcement learning module;

[0038] S34, synchronously input the current observation space and the current action space into the Actor structure and the Critic structure;

[0039] S35, the actions and rewards generated by the Actor structure and the Critic structure interact with the environment to obtain a new observation space and action space, and input them into the improved optimization experience replay pool and then return to S34;

[0040] S36. Iterate S34 and S35 a predetermined number of times and output batch candidate solutions to the genetic algorithm module.

[0041] Furthermore, the candidate solution includes a mission completion reward Reward and a satellite repetition time window Loss.

[0042] Furthermore, the improvements of the improved optimized experience replay pool include:

[0043] Added the current size property to track the current amount of experience stored in the experience replay pool;

[0044] Optimized the push method to ensure that old historical data is not overwritten and new data is only added when the experience replay pool is not full;

[0045] Optimize the compute returns method to ensure that only the return value of the currently stored experience is calculated;

[0046] Added the sample sampling method to randomly extract batches of data from the experience replay pool for training.

[0047] Further, the S4 comprises the following steps:

[0048] S41, calling the NSGA-II algorithm to optimize the received candidate solutions;

[0049] S42, constructing a Minimize function based on the optimized candidate solution and solving it, and inputting the optimized candidate solution into the genetic algorithm module again;

[0050] S43, iterating S41 and S42 for a preset number of times;

[0051] S44. Minimize the corresponding constellation configuration parameters using the Minimize function to obtain the optimal constellation configuration parameters and output them.

[0052] Compared with the prior art, the present invention has the following beneficial technical effects:

[0053] The fast constellation configuration design method based on parallel reinforcement learning and genetic algorithm provided by the present invention introduces parallel computing, calls multiple GPU graphics cards to build a parallel interactive algorithm framework, updates network parameters in real time and obtains a batch of candidate solutions, which greatly improves the iterative process of reinforcement learning, can effectively process a large amount of emergency observation target data, and the present invention needs to optimize the strategy by constantly interacting with the environment, and this process is highly complex and high-dimensional. By calling multiple GPUs for parallel computing, the efficiency of this process can be improved. The candidate solution is input into the genetic algorithm module, the genetic algorithm evaluates and optimizes the candidate solution, and processes it with the Minimize function, so as to find a set of constellation configuration parameters with the maximum reward, the minimum loss and satisfying the constraints.

[0054] In the large-scale constellation configuration design experiment of the present invention, the optimization design efficiency of the large-scale constellation orbit construction of 1,000 satellites was accelerated by more than 75% compared with the original. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.

[0056] Figure 1 A flow chart of a fast constellation configuration design method based on parallel reinforcement learning and genetic algorithm provided by an embodiment;

[0057] Figure 2 A schematic diagram of a multi-objective optimization model provided by an embodiment;

[0058] Figure 3 A schematic diagram of a parallel reinforcement learning module provided by an embodiment;

[0059] Figure 4 A flowchart of the parallel reinforcement learning module provided by an embodiment;

[0060] Figure 5 A schematic diagram of the structure of a parallel Transformer module provided in an embodiment;

[0061] Figure 6 A schematic diagram of an improved optimized experience replay pool provided by an embodiment;

[0062] Figure 7 A flowchart of the genetic algorithm module provided by one embodiment. DETAILED DESCRIPTION

[0063] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0064] Reference Figure 1 , an embodiment provides a fast constellation configuration design method based on parallel reinforcement learning and genetic algorithm, comprising the following steps:

[0065] S1. Randomly generate observation targets, construct observation target data sets, and give initial constellation configuration parameters;

[0066] S2. Under constraints, construct a multi-objective optimization model for constellation configuration parameters, wherein the multi-objective optimization model includes a parallel reinforcement learning module and a genetic algorithm module;

[0067] S3, the parallel reinforcement learning module performs training interaction according to the input observation target data set and constellation configuration parameters, and outputs batch candidate solutions to the genetic algorithm module;

[0068] S4. The genetic algorithm module evaluates and optimizes the received candidate solutions, obtains the optimal constellation configuration parameters and outputs them.

[0069] The constellation configuration parameters include: the number of orbits, the number of satellites on the orbital plane, the orbital altitude, the orbital intersection angle, the right ascension of the ascending node and the true anomaly angle.

[0070] Reference Figure 2 The multi-objective optimization model includes a parallel reinforcement learning module and a genetic algorithm module, and the parallel reinforcement learning module and the genetic algorithm module are in a double-layer parallel cycle alternating relationship.

[0071] In one embodiment, the observation target dataset is a large-scale dynamic observation target dataset, which is used to meet the training requirements of parallel reinforcement learning. The specific construction process is as follows:

[0072] 1,500 observation targets are randomly generated in different geographical areas (including polar regions), including 800 dynamic ad hoc targets and 700 static fixed targets. Each target has longitude, latitude, altitude and task priority definitions. The dynamic ad hoc targets contain a series of longitudes and latitudes that can show their movement trajectory.

[0073] The given initial constellation configuration parameters are: the initial number of orbits is defined as 40; the number of satellites on each orbital plane is 25; the optimization interval of the orbital altitude is [200,700], in kilometers, and the initial orbital altitude is defined as 286 kilometers; the optimization interval of the orbital intersection angle is [20,160], in degrees, and the initial orbital intersection angle is defined as 22 degrees; the optimization interval of the ascending node right ascension is [60,100], and the initial ascending node right ascension value is defined as 62; the optimization interval of the true anomaly angle is [60,100], and the initial true anomaly angle value is defined as 62.

[0074] In the constellation configuration optimization design process, the method provided by the present invention comprehensively considers the observation task demand constraints, on-board resource storage constraints, on-board energy constraints, and observation time window constraints; and reasonably arranges the constellation observation tasks under these limited constraints.

[0075] The observation task requirement constraints satisfy:

[0076] ;

[0077] ;

[0078] in, Imaging payload types for mission requirements; The type of payload carried on the satellite; is the load resolution; Imaging resolution for mission requirements;

[0079] Since each satellite in the constellation has limited onboard storage resources, the satellite has a fixed number of power-on and power-off times. To target When performing imaging observations, the satellite Storage space , that is, the on-board resource storage constraints satisfy:

[0080] ;

[0081] in, For satellites; for the goal; Storage space for satellites; For satellite To target The time when the imaging observations were made; is the storage consumption rate;

[0082] In order to ensure the normal operation of the satellite, the energy on the satellite must have a corresponding applicable range; the energy constraints on the satellite satisfy:

[0083] ;

[0084] in, Indicates the target Subtasks In satellite The kth time window of is observed; Indicates the target The duration of the visible time window; Indicates satellite Observe the power consumption per unit time; Indicates satellite Total power consumption;

[0085] The observation time window constraints satisfy:

[0086] ;

[0087] in, For the goal In satellite All possible observation windows within the imaging observation period; is the observation window; and They represent the time when the observation starts and ends respectively; and They represent the start time and end time of the observation window respectively.

[0088] At the same time, it is necessary to consider the mission benefits, mission timeliness and satellite load on the mission to maximize the mission benefits, maximize the mission timeliness and balance the satellite load on the mission.

[0089] The mission maximizes the benefits of:

[0090] The observation target defines the task priority attribute, which binds the user-specified target observation priority to the target observation benefit. The higher the priority, the greater the benefit of the target being observed and imaging data obtained. By calculating the priority and importance, the task strategy is adjusted and the imaging benefit is increased. The calculation method is as follows:

[0091] ;

[0092] in, The energy required to observe the current task; The power consumption per unit time of satellite observation; The available storage capacity for the satellite.

[0093] The timeliness of the task is maximized:

[0094] In the Earth observation constellation, there are multiple satellites that observe the same target multiple times. However, in a fixed assumption period, different satellites observe it at different times, which may differ from each other by minutes or hours. Therefore, the timeliness of task completion is very important. It is necessary to ensure that the task is completed as early as possible within the assumption period. The timeliness of the task is defined as:

[0095] ;

[0096] in, Indicates the task In satellite No. Whether the time window is observed, 1 if observed, 0 if not observed; , For the current task The start and end observation times.

[0097] The satellite load balances the mission:

[0098] Each satellite in the Earth Observation constellation has its own fixed storage and energy. Tasks are assigned to each satellite as evenly as possible. The load balancing is defined as:

[0099] ;

[0100] in For satellite Observation time:

[0101] ;

[0102] The average observation time of the satellite is:

[0103] .

[0104] Reference Figure 3 ,In S3, the parallel reinforcement learning module is trained using an Actor-Critic structure;

[0105] The Actor-Critic structure includes an Actor structure and a Critic structure;

[0106] The Actor structure is used to generate a probability for determining the next action and select an optimal action strategy based on the observed state input; the Actor structure includes a parallel Transformer module and a softmax module;

[0107] The Critic structure is used to evaluate the value estimate of a given state and the reward of a specific state based on the observed state input; the Critic structure includes a parallel Transformer module.

[0108] Reference Figure 5 , the parallel Transformer module includes a position encoding block, a residual encoding block, a Transformer encoding block and a decoding block; the working steps of the parallel Transformer module include:

[0109] A1. Embedded satellite observation space, which includes four-dimensional dynamic state space and three-dimensional static state space;

[0110] A2, Linear coding performs linear state coding on the input observation space obs to achieve linear transformation between the input tensor and the weight matrix. This part of tensor calculation uses GPU-0 computing power for calculation.

[0111] A3. Position encoding block adds position information to the embedding layer. Position encoding helps to understand the position relationship of sequence elements. This part of encoding calculation uses GPU-0 computing power to perform;

[0112] A4. The input of the residual coding block comes from the position coding block. The residual coding block includes a normalization layer, a multi-head self-attention layer, a residual connection, and a discard operation. Among them, the normalization layer is used to normalize the input; the multi-head self-attention layer extracts the relevant features of the input sequence through the multi-head attention mechanism and helps the model to further learn the representation, thereby calculating the attention weights of different positions in the sequence; the residual connection and discard operation add the output of the self-attention layer to the input sequence and apply the discard operation. The residual connection helps to alleviate the gradient vanishing problem, making it easier for the model to learn deep representations; the discard operation is used to randomly discard the output before the residual connection to further improve the generalization ability of the model; the parallel encoding calculation of the residual coding block calls the GPU-1 computing power for execution;

[0113] A5. After receiving the sequence information from the residual coding block and some position coding blocks, the Transformer coding block performs a linear transformation including a feedforward neural network to achieve the final coding output. This part of the coding calculation uses the computing power of GPU-2.

[0114] A6. The decoder makes decisions based on the output encoding. A 5-layer perceptron MLP network is used in the decoding process. The output of the last layer of the decoder is mapped to the probability distribution of the target through a linear layer.

[0115] Reference Figure 4 , said S3 comprises the following steps:

[0116] S31, reading the observation target data set and constellation configuration parameters, wherein the observation target data set is read in csv format, and the constellation configuration parameters (the number of orbits, the number of satellites on the orbital plane, the orbital altitude, the orbital intersection angle, the right ascension of the ascending node, and the true anomaly angle) are read in config configuration file format;

[0117] S32, create observation state, construct four-dimensional dynamic state space and three-dimensional static state space; the observation target data set input in S31 is , each Both contain tuples ,in Represents static input, , and Respectively represent tasks The observable start time, end time and duration of the observation task. represents the weight of the task; For input Dynamic input element at the moment, Indicates the task The state (unselectable due to constraint violation, selected, or not selected), Indicates the time when the task starts to be planned. Indicates the remaining storage space of the satellite currently planning the mission; the above forms a four-dimensional dynamic state space obs[0] and a three-dimensional static state space obs[1];

[0118] S33, using the four-dimensional dynamic state space obs[0] and the three-dimensional static state space obs[1] as the observation space obs of the parallel reinforcement learning module, and using the given number of observation targets as the action space of the parallel reinforcement learning module;

[0119] S34. Synchronously input the current observation space and the current action space into the Actor structure and the Critic structure; the Actor structure includes a parallel Transformer module and a softmax module, and the parallel Transformer module includes a position encoding block, a residual encoding block, a Transformer encoding block and a decoding block, wherein the position encoding block and the residual encoding block are respectively calculated by parallel construction of synchronous encoding on GPU-0 and GPU-1, and then input into the Transformer encoding block and the softmax module to generate the probability for the next action, thereby selecting the optimal action strategy Output; The Critic structure includes parallel Transformer modules, which evaluate the value estimate of a given observation space and thus evaluate the reward of a specific observation space. Output;

[0120] S35, the actions and rewards generated by the Actor structure and the Critic structure interact with the environment to obtain a new observation space and action space, and input them into the improved optimization experience replay pool and then return to S34;

[0121] S36. Iterate S34 and S35 a predetermined number of times and output batch candidate solutions to the genetic algorithm module.

[0122] The actions and rewards generated by the Actor structure and the Critic structure are input into the improved experience replay pool, and the current strategy is continuously trained through historical data to improve sample utilization.

[0123] In one embodiment, the predetermined number of times is 1000 times, and the parallel reinforcement learning module updates the network parameters in real time during the training process, thereby outputting a batch of candidate solutions to the genetic algorithm module; the candidate solutions include the task completion reward Reward and the satellite repetition time window Loss.

[0124] Reference Figure 6 The improvements of the improved optimized experience replay pool include:

[0125] Added the current size property to track the current amount of experience stored in the experience replay pool;

[0126] Optimize the push method to ensure that old historical data is not overwritten and new data is added only when the experience replay pool is not full (if historical data is not used, the previous data must be discarded for each iteration and new data must be generated to update the network parameters, which will result in high resource consumption and waste for complex problems; and a considerable part of historical data is positive for updating new network parameters, which can save training time);

[0127] Optimize the compute returns method to ensure that only the return value of the currently stored experience is calculated;

[0128] Added the sample sampling method to randomly extract batches of data from the experience replay pool for training.

[0129] Reference Figure 7 , said S4 comprises the following steps:

[0130] S41, calling the NSGA-II algorithm to optimize the received candidate solutions;

[0131] S42, constructing a Minimize function based on the optimized candidate solution and solving it, and inputting the optimized candidate solution into the genetic algorithm module again;

[0132] S43, iterating S41 and S42 for a preset number of times;

[0133] S44. Minimize the corresponding constellation configuration parameters using the Minimize function to obtain the optimal constellation configuration parameters and output them.

[0134] The optimal constellation configuration parameters finally output include the number of orbits, the number of satellites on the orbital plane, orbital altitude, orbital intersection angle, right ascension of the ascending node, and true anomaly angle.

[0135] In one embodiment, the preset number of times in S43 is 1000 times.

[0136] The present invention can maximize the operating efficiency of reinforcement learning under limited graphics card resources, thereby improving the efficiency of configuration optimization for large-scale satellite constellations.

[0137] Matters not covered by the present invention are known technologies.

[0138] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0139] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

[0140] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A fast constellation configuration design method based on parallel reinforcement learning and genetic algorithm, characterized in that: The following steps are involved: S1. Randomly generate observation targets, construct observation target data sets, and give initial constellation configuration parameters; S2. Under constraints, construct a multi-objective optimization model for constellation configuration parameters, wherein the multi-objective optimization model includes a parallel reinforcement learning module and a genetic algorithm module; S3, the parallel reinforcement learning module performs training interaction according to the input observation target data set and constellation configuration parameters, and outputs batch candidate solutions to the genetic algorithm module; The parallel reinforcement learning module adopts the Actor-Critic structure for training; The Actor-Critic structure includes an Actor structure and a Critic structure; The Actor structure is used to generate a probability for determining the next action and select an optimal action strategy based on the observed state input; the Actor structure includes a parallel Transformer module and a softmax module; The Critic structure is used to evaluate the value estimate of a given state and the reward of a specific state based on the observed state input; the Critic structure includes a parallel Transformer module; S4. The genetic algorithm module evaluates and optimizes the received candidate solutions, obtains the optimal constellation configuration parameters and outputs them.

2. The fast constellation configuration design method based on parallel reinforcement learning and genetic algorithm as claimed in claim 1, characterized in that: In S2, the constraints include observation mission requirement constraints, on-board resource storage constraints, on-board energy constraints, and observation time window constraints; The observation task requirement constraints satisfy: in, Imaging payload types for mission requirements; The type of payload carried on the satellite; is the load resolution; Imaging resolution for mission requirements; The on-board resource storage constraints satisfy: in, For satellites; for the goal; Storage space for satellites; For satellite To target The time when the imaging observations were made; is the storage consumption rate; The on-board energy constraints satisfy: in, Indicates the target Subtasks In satellite The kth time window of is observed; Indicates the target The duration of the visible time window; Indicates satellite Observe the power consumption per unit time; Indicates satellite Total power consumption; The observation time window constraints satisfy: in, For the goal In satellite All possible observation windows within the imaging observation period; is the observation window; and They represent the time when the observation starts and ends respectively; and They represent the start time and end time of the observation window respectively.

3. The fast constellation configuration design method based on parallel reinforcement learning and genetic algorithm as claimed in claim 1, characterized in that: In S1, the constellation configuration parameters include: the number of orbits, the number of satellites on the orbital plane, the orbital altitude, the orbital intersection angle, the right ascension of the ascending node and the true anomaly angle.

4. The fast constellation configuration design method based on parallel reinforcement learning and genetic algorithm as claimed in claim 1, characterized in that: The parallel Transformer module includes a position encoding block, a residual encoding block, a Transformer encoding block and a decoding block.

5. The rapid constellation configuration design method based on parallel reinforcement learning and genetic algorithm as claimed in claim 4, characterized in that: In the parallel Transformer module, the encoding calculation of the position encoding block calls the computing power of GPU-0; the encoding calculation of the residual encoding block calls the computing power of GPU-1; and the encoding calculation of the Transformer encoding block calls the computing power of GPU-2.

6. The fast constellation configuration design method based on parallel reinforcement learning and genetic algorithm as claimed in claim 1, characterized in that: The S3 comprises the following steps: S31, reading observation target data set and constellation configuration parameters; S32, creating an observation state, constructing a four-dimensional dynamic state space and a three-dimensional static state space; S33, using the four-dimensional dynamic state space and the three-dimensional static state space as the observation space of the parallel reinforcement learning module, and using the given number of observation targets as the action space of the parallel reinforcement learning module; S34, synchronously input the current observation space and the current action space into the Actor structure and the Critic structure; S35, the actions and rewards generated by the Actor structure and the Critic structure interact with the environment to obtain a new observation space and action space, and input them into the improved optimization experience replay pool and then return to S34; S36. Iterate S34 and S35 a predetermined number of times and output batch candidate solutions to the genetic algorithm module.

7. The fast constellation configuration design method based on parallel reinforcement learning and genetic algorithm according to claim 6, characterized in that: The candidate solution includes the mission completion reward Reward and the satellite repetition time window Loss.

8. The fast constellation configuration design method based on parallel reinforcement learning and genetic algorithm as claimed in claim 6, characterized in that: The improvements to the improved optimized experience replay pool include: Added the current size property to track the current amount of experience stored in the experience replay pool; Optimized the push method to ensure that old historical data is not overwritten and new data is only added when the experience replay pool is not full; Optimize the compute returns method to ensure that only the return value of the currently stored experience is calculated; Added the sample sampling method to randomly extract batches of data from the experience replay pool for training.

9. The method for rapid constellation configuration design based on parallel reinforcement learning and genetic algorithm as claimed in claim 1, characterized in that: The S4 comprises the following steps: S41, calling the NSGA-II algorithm to optimize the received candidate solutions; S42, constructing a Minimize function based on the optimized candidate solution and solving it, and inputting the optimized candidate solution into the genetic algorithm module again; S43, iterating S41 and S42 for a preset number of times; S44. Minimize the corresponding constellation configuration parameters using the Minimize function to obtain the optimal constellation configuration parameters and output them.

Citation Information

Patent Citations

  • Satellite space target collaborative observation distributed planning method based on multi-agent reinforcement learning

    CN116187160A

  • Large-scale satellite-ground networking optimization problem modeling method and hybrid learning type multi-parallel method

    CN118118082A