Micro-nano calibration satellite constellation cooperative control method based on GA and DRL coupling optimization

By using the GA and DRL coupled optimization method, a micro-nano satellite constellation mission model was constructed, realizing constellation-level autonomous collaborative control and efficient calibration mission execution. This solved the problems of poor real-time performance and low resource utilization efficiency in traditional methods, and improved the on-orbit calibration capability of micro-nano satellite constellations.

CN121966680APending Publication Date: 2026-05-01SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHWEST JIAOTONG UNIV
Filing Date
2026-02-05
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional microsatellite constellation calibration methods rely on centralized ground scheduling, which has poor real-time performance and lacks global optimization and autonomous coordination capabilities. This results in low calibration frequency, low resource utilization efficiency, and weak maneuverability and limited energy of microsatellites, making it difficult to respond quickly in complex environments.

Method used

A coupled optimization method combining genetic algorithm (GA) and deep learning (DRL) is used to construct a micro-nano satellite constellation mission model, perform global optimization and autonomous collaborative control, and achieve efficient constellation-level calibration mission execution through information sharing and distributed collaborative strategies via inter-satellite links.

Benefits of technology

This has improved the on-orbit calibration capability of micro and nano satellite constellations, enabling them to perform calibration tasks with high timeliness and stability in complex and dynamic environments, reducing reliance on ground control and improving resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121966680A_ABST
    Figure CN121966680A_ABST
Patent Text Reader

Abstract

The invention discloses a micro-nano calibration satellite constellation cooperative control method based on GA and DRL coupling optimization, and relates to the field of satellite networking and constellation optimization design. The method comprises the following steps: constructing a micro-nano satellite constellation task model based on constellation information data; performing global optimization on the constellation parameter vector according to the fitness function in a full constellation parameter space by adopting a genetic algorithm based on the micro-nano satellite constellation task model to obtain an initial constellation configuration set; determining a strategy parameter set according to the initial constellation configuration set by adopting a deep learning method; and carrying out in-orbit deployment according to the strategy parameter set, carrying out state estimation on each satellite, carrying out dynamic updating on the satellites based on the satellite execution control quantity, and completing task allocation and maneuvering negotiation through inter-satellite links by adopting a satellite distributed cooperative strategy so as to realize constellation-level execution control. According to the invention, constellation-level optimal configuration and autonomous cooperative control can be realized, and the in-orbit calibration capability of the micro-nano satellite constellation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of satellite networking and constellation optimization design, and in particular to a micro-nano calibration satellite constellation cooperative control method based on GA and DRL coupling optimization. Background Technology

[0002] With the development of active microwave remote sensing technology, the demand for space calibration has increased significantly. Calibration frequency, timeliness, and geometric configuration stability all have a crucial impact on sensor accuracy. Traditional calibration methods rely on centralized ground scheduling, resulting in poor real-time performance. Furthermore, microsatellite constellations lack global optimization and autonomous collaborative capabilities, leading to low calibration frequency and inefficient resource utilization. In recent years, microsatellites have gradually demonstrated their advantages in space calibration tasks due to their low cost, flexible launch, and strong networking capabilities. However, microsatellites are constrained by factors such as weak orbital maneuverability, limited energy, and complex operating environments, significantly increasing the difficulty of overall constellation calibration planning, resource coordination, and orbit control. Related methods often rely on unified ground scheduling, failing to achieve rapid response to dynamic environments and making it difficult to guarantee constellation calibration efficiency and stability. Therefore, an intelligent method capable of constellation-level optimized configuration and autonomous collaborative control is urgently needed to improve the on-orbit calibration capabilities of microsatellite constellations. Summary of the Invention

[0003] The purpose of this application is to provide a cooperative control method for micro-nano calibration satellite constellations based on GA and DRL coupling optimization, which can realize constellation-level optimized configuration and autonomous cooperative control, and improve the on-orbit calibration capability of micro-nano satellite constellations.

[0004] To achieve the above objectives, this application provides the following solution: This application provides a cooperative control method for micro / nano calibration satellite constellations based on GA and DRL coupling optimization, including: Acquire constellation information data; the constellation information data includes: a predefined satellite constellation, a full constellation parameter space, and an orbital six-root parameter vector established for each satellite in the satellite constellation; A micro-nano satellite constellation mission model is constructed based on the constellation information data. The micro-nano satellite constellation mission model includes an objective function and constraints. The objective function is determined based on the satellite orbital dynamics equations and, under effective calibration events, the number of constellation calibrations and the total speed increment of maneuvers. The effective calibration event is obtained after determining that the satellites within the constellation are operating normally when the satellite visibility conditions meet preset conditions and the continuous time exceeds a set threshold. A genetic algorithm is used based on the micro / nano satellite constellation mission model to globally optimize the constellation parameter vectors according to the fitness function within the entire constellation parameter space, thereby obtaining an initial constellation configuration set. The constellation parameter vectors are determined based on the orbital six-root parameter vectors. The fitness function is determined based on the objective function and the penalty function. The penalty function is determined based on the penalty coefficient under the condition of violating constraints. The strategy parameter set is determined using deep learning methods based on the initial constellation configuration set; Based on the strategy parameter set, the satellites are deployed in orbit and their states are estimated. Based on the satellite execution control variables, the satellites are dynamically updated. Through inter-satellite links, a satellite distributed cooperative strategy is adopted to complete task allocation and maneuver negotiation in order to achieve constellation-level execution control.

[0005] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a cooperative control method for micro / nano calibration satellite constellations based on GA and DRL coupled optimization. A micro / nano satellite constellation mission model is constructed based on constellation information data. A genetic algorithm is used to globally optimize the constellation parameter vectors within the entire constellation parameter space according to a fitness function, achieving integrated global optimization of constellation configuration and mission planning. Deep learning methods are employed to determine the strategy parameter set based on the initial constellation configuration set. On-orbit deployment is performed based on the strategy parameter set, and the state of each satellite is estimated. Furthermore, satellite dynamics are updated based on the satellite's execution control variables. Mission allocation and maneuver negotiation are completed through inter-satellite links using a distributed cooperative strategy, achieving constellation-level execution control. In other words, on-orbit deployment based on deep reinforcement learning enables satellites to autonomously generate calibration schedules and orbital maneuvers based on real-time observation conditions, energy status, and mission requirements, without relying on continuous ground control. This allows for decision-making continuity even during link interruptions, external disturbances, or unexpected missions, achieving highly timely and stable calibration mission execution. Furthermore, by sharing status through inter-satellite links and combining distributed collaborative strategies, constellation-level optimized configuration and autonomous collaborative control can be achieved, enhancing the on-orbit calibration capability of micro-nano satellite constellations. Attached Figure Description

[0006] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0007] Figure 1 This is a flowchart of a cooperative control method for micro / nano calibration satellite constellations based on GA and DRL coupling optimization. Detailed Implementation

[0008] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0009] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0010] In one exemplary embodiment, such as Figure 1 As shown, a cooperative control method for micro / nano calibration satellite constellations based on GA and DRL coupling optimization is provided, including: Step 100: Obtain constellation information data. Constellation information data includes: a predefined satellite constellation, the entire constellation parameter space, and the orbital six-root parameter vector established for each satellite in the constellation.

[0011] Step 200: Construct a micro / nano satellite constellation mission model based on constellation information data. The micro / nano satellite constellation mission model includes an objective function and constraints. The objective function is determined based on the satellite orbital dynamics equations, under effective calibration events, and is determined after judging whether each satellite in the constellation is operating normally. Effective calibration events are determined when satellite visibility conditions meet preset conditions and the continuous time exceeds a set threshold.

[0012] The mathematical expression for the objective function is: .

[0013] .

[0014] .

[0015] .

[0016] in, The objective function is... Number of times the constellation is calibrated; For constellation parameter vectors; The total energy consumed during the calibration task; This represents the increment of the total speed of the maneuver; , , All are weighted coefficients; The target revenue; For serial numbers; Number of satellites; Number the decision steps within the task's time window; For the first The orbital or attitude maneuvers of each micro-nano calibration satellite within the mission time window are numbered; The first in the constellations Micro-nano calibration satellites in the The speed increment required for the next maneuver; For the first The first micro-nano calibration satellite within the mission time window Load power consumption at each discrete decision moment or simulation time step; The first one within the task time window A discrete decision moment or simulation time step; For the first Micro-nano calibration satellites in The power generated by the solar cell array; For time step.

[0017] The constraints include: orbital mechanics constraints, inter-satellite visibility constraints, and energy balance constraints.

[0018] The orbital mechanical constraints specifically include: .

[0019] .

[0020] .

[0021] The inter-satellite visibility constraints include the inter-satellite geometric visibility determination function and the visibility indicator; The interstellar geometric visibility determination function is denoted as: Used to determine the first micro-nano calibration satellite and the first Position vector of a micro-nano calibration satellite in space , Whether an interstellar connection is obscured by Earth is considered to have interstellar visibility if the following conditions are met: ; The visibility indicator is denoted as ,when At the time of its establishment, ;otherwise ; The energy balance constraints specifically include: .

[0022] .

[0023] in, For the first The position vector of a micro-nano calibration satellite in space; The position vector of the calibrated target or reference load in space; The maximum spatial distance allowed for calibration tasks; Angle of incidence; It is a continuous-time variable; This refers to the maximum allowable attitude deflection angle or maximum field of view angle when the satellite is performing a calibration mission. For the first The velocity vector of a micro-nano calibration satellite in space; The velocity vector in space for the calibration target or reference load; The maximum speed allowed for a satellite during its orbital operation or relative motion; The average radius of the Earth; For the first Micro-nano calibration satellites in The state of charge of the battery at that time; The minimum state of battery charge allowed for the safe operation of the satellite; This represents the maximum permissible state of charge for the satellite's batteries. The first one within the task time window A discrete decision moment or simulation time step; For serial numbers; In constellation configuration scheme The cumulative velocity increment required for a micro-nano calibration satellite to perform all orbital or attitude maneuvers within the mission time window; For the first The maximum velocity increment that a micro-nano calibration satellite is allowed to consume throughout its entire mission cycle; It is a constellation configuration scheme, consisting of the set of orbital parameters of each satellite in the constellation.

[0024] Step 300: Using a genetic algorithm based on the micro / nano satellite constellation mission model, the constellation parameter vector is globally optimized within the entire constellation parameter space according to the fitness function to obtain the initial constellation configuration set. The constellation parameter vector is determined based on the orbital root number parameter vector; the fitness function is determined based on the objective function and the penalty function; the penalty function is determined based on the penalty coefficient under the condition of violating constraints.

[0025] Specifically, a genetic algorithm based on a micro / nano satellite constellation mission model is used to globally optimize the constellation parameter vectors within the entire constellation parameter space according to a fitness function, resulting in an initial constellation configuration set, which includes: An initial population is randomly generated within the entire constellation parameter space; the initial population consists of multiple individuals; each individual is a satellite within the constellation; each individual contains a set of orbital root numbers established by N satellites within the satellite constellation.

[0026] Based on the micro-nano satellite constellation mission model and the initial population, the first... The objective function corresponding to each individual is assigned, and constraints and penalties are applied. Selection, crossover, and mutation operations are performed on each individual. If each individual satisfies all constraints, the fitness of the original fitness function is retained. If there is a violation of the constraints, a pre-set deduction is applied according to the degree of violation.

[0027] For each of the individuals, selection, crossover, and mutation operations are performed, specifically including: The fitness of all individuals is normalized, and the normalized fitness is used as the selection probability. A tournament-style selection method is employed to calculate the probability of an individual being selected based on the selection probability. The expression corresponding to the probability of an individual being selected is: .

[0028] A new parent set is generated based on the probability of an individual being selected. Parent individuals in the parent set are crossovered according to the crossover probability to generate new individuals. Random perturbations are applied to some orbital parameters of the new individuals according to the mutation probability. To prevent parameters from going out of bounds, the mutated parameters need to be pruned back to the feasible range. The random perturbations include: .

[0029] The feasible range includes: .

[0030] An adaptive mutation strategy is adopted, based on algebraic... Gradually reduce the intensity of variation: .

[0031] in, The probability of an individual being selected; For the probability of selection; Population size; For the first In the first generation of genetic algorithms, the first generation... Normalized fitness of each gene locus; For the first Generation Zhong, No. The individual in the first The probability of variation at each gene locus; For the first The individual in the first Random perturbation at each gene locus; For the first The minimum allowed value for a gene locus; For the first The maximum allowed value or probability limit for a gene locus; For the first Generation Zhong, No. The intensity of variation at each gene locus; For the first The initial variation intensity of each gene locus; It is the attenuation factor; This represents the maximum number of iterations.

[0032] After each iteration, the optimal fitness is determined, the relative improvement rate between successive generations is calculated, and convergence is checked; the expression for the optimal fitness is: .

[0033] in, For optimal fitness; For the first In the first generation of genetic algorithms, the first generation... Individual constellation configuration The corrected fitness function value; This is the index of the individual in the genetic algorithm population.

[0034] The convergence conditions for convergence determination include: .

[0035] .

[0036] in, The maximum number of iterations; For the first The relative improvement rate between successive generations corresponding to each generation; The threshold for determining convergence; For the first The optimal fitness corresponding to each generation; For the first The optimal fitness corresponding to each generation.

[0037] If the result is convergence, output the optimal or suboptimal constellation parameter vectors under the current iteration, and determine the initial constellation configuration set.

[0038] Step 400: Use deep learning methods to determine the policy parameter set based on the initial constellation configuration set.

[0039] Specifically, the strategy parameter set is determined using deep learning methods based on the initial constellation configuration set, including: The initial state set of the agent is determined by mapping based on the initial constellation configuration set; each state in the initial state set of the agent includes satellite orbital elements, energy state, and mission visibility matrix.

[0040] A single-agent global strategy and a multi-agent interactive approach are employed. The overall training objective is to maximize the expected reward value of all initial states in the agent's initial state set, thereby determining the trajectory set. The expression corresponding to the overall training objective is: .

[0041] The trajectory set is injected into the experience replay pool, and behavior cloning is used for policy training and updating. Furthermore, a projection operator is used for action correction, and a Lagrange penalty term is used for joint optimization to obtain the policy parameter set. The mathematical formula for the joint optimization is as follows: .

[0042] in, The parameters for the deep reinforcement learning strategy are: The expected return function at that time; To perform expectation calculations on the initial state distribution; To reinforce the initial state of the learning process; To reinforce the initial state distribution set of the learning environment; In strategy The following is the expectation calculation performed on the trajectory distribution; To reinforce the state-action-reward sequence trajectory in learning; The accumulated reward value starting from the initial state; Let be the constraint reinforcement learning loss function containing Lagrange multipliers; For the parameter vector of the deep reinforcement learning policy network; These are Lagrange multiplier vectors; For the constraint number index; For the first Lagrange multipliers corresponding to each constraint; This refers to the expectation operation performed on the set of trajectories. For the first The positive part of a constraint function.

[0043] Step 500: Deploy satellites in orbit according to the strategy parameter set and estimate the state of each satellite. Update the satellite dynamics based on the satellite execution control variables. Complete task allocation and maneuver negotiation through inter-satellite links using a satellite distributed cooperative strategy to achieve constellation-level execution control.

[0044] In one embodiment, on-orbit deployment is performed based on a strategy parameter set, and state estimation is performed for each satellite. Furthermore, satellite dynamics are updated based on satellite execution control variables. Task allocation and maneuver negotiation are completed via inter-satellite links using a distributed satellite cooperative strategy to achieve constellation-level execution control. Specifically, this includes: Quantization and compression are performed based on the strategy parameter set to meet onboard resource constraints, and offline verification is conducted. Furthermore, state estimation is performed using filters based on the observations of each satellite to ensure the execution of the online deployment strategy. The expression corresponding to the onboard resource constraints is: .

[0045] After on-orbit deployment, the satellite's dynamics are updated based on its execution control variables. This updates the satellite's propulsion status in an offline manner and includes updates to resource consumption and event statistics. Task allocation and maneuver negotiation are completed via inter-satellite links using a distributed satellite cooperative strategy to achieve constellation-level execution control. In the event of a major constraint violation, a safety backoff is performed based on a safety baseline strategy. Major constraint violations include collisions and energy levels falling below a safety threshold. The expressions corresponding to the satellite execution control variables are: .

[0046] in, This represents the total number of parameters in the policy model after training. For the strategy parameter set; This is the upper limit of the parameter; The time for a single inference of the policy model on a spaceborne platform; For single inference delay; This represents the maximum allowed inference time for the spaceborne system. Storage resources required for the strategy model; Available memory; For satellite control; For the first A satellite at any time Control input; For control input Operators for performing constrained projections; A generic variable used to control input; The set of actions to satisfy the current hard constraints.

[0047] The satellite distributed collaborative strategy specifically includes: Each benchmarking task Released within the constellation, the calibration mission is undertaken by the satellite with the highest competitive value; satellite competitive value The calculation formula is: .

[0048] If the calibration task requires multi-satellite collaboration, a distributed consensus solution should be adopted to share the burden: .

[0049] in, for The expected increase in rewards after undertaking a task; for Consumed in time; This is the cost weighting coefficient; For constellations Time for the first The scheduling decision vector for each task; It is an inter-satellite distributed consistency operator.

[0050] As an optional implementation method, the micro-nano calibration satellite constellation cooperative control method based on GA and DRL coupling optimization also includes: performing on-orbit fine-tuning of the satellite and periodic performance evaluation.

[0051] The on-orbit fine-tuning process specifically includes: Collect local small batch trajectories Calculate the local gradient: .

[0052] Update with a conservative step size: .

[0053] In-orbit updates must meet the following requirements: .

[0054] in, For the first The gradient vector calculated by each agent; For about The gradient operator; For the first Loss function calculated for each satellite sampling batch; For the first Local strategy parameters obtained by the satellite after in-orbit fine-tuning; For the strategy parameter set; The learning rate is fine-tuned online in orbit; For numerical stability parameters; The performance improvement of the new strategy compared to the baseline strategy on the validation set; The minimum performance improvement threshold required for policy updates.

[0055] This application relates to space-based calibration and micro / nano satellite constellation optimization, belonging to the field of remote sensing. It optimizes the micro / nano calibration satellite constellation by using a coupled algorithm of genetic algorithm and deep reinforcement learning to optimize the micro / nano calibration satellite group, thereby significantly increasing the calibration frequency of the micro / nano satellite group for active microwave sensors.

[0056] First, a genetic algorithm is used to globally search the six orbital root numbers of the satellites within the constellation. The fitness is calculated by weighting the total number of calibration attempts within the mission time and subtracting the energy consumed. Through continuous iteration, poorly performing constellation schemes are eliminated, while high-performing schemes are combined and mutated, ultimately evolving into several optimal initial constellation configurations. The constellation obtained by the genetic algorithm (GA) is used as the initial state, i.e., the initial constellation configuration set. Subsequently, each satellite within the selected constellation is treated as an agent, continuously interacting with the complex dynamic environment and learning the optimal strategy based on the rewards obtained, further improving calibration efficiency. A trained deep reinforcement learning (DRL) model is embedded into the constellation to achieve autonomous and intelligent calibration mission planning and orbit control.

[0057] The method mentioned in this application can solve the following technical problems: 1. The initial constellation configuration lacks global optimization capabilities. Most existing calibration constellations rely on empirical design or local optimization algorithms based on simplified models; the real-time performance and adaptability of centralized ground scheduling are insufficient, and traditional calibration tasks mostly depend on ground planning and command issuance.

[0058] 2. Microsatellites and nanosatellites have limited capabilities and lack coordination mechanisms within their constellations. Microsatellites and nanosatellites have weak maneuverability and tight energy budgets, while current technologies lack effective methods for coordinated control and task allocation; they also lack on-orbit intelligent decision-making capabilities and sustainable optimization mechanisms. Existing methods typically involve planning on the ground, with the satellite only performing fixed tasks.

[0059] By constructing a global search constellation configuration optimization framework, combining it with an intelligent control model that has online reasoning and autonomous planning capabilities, and designing inter-satellite collaboration and feedback learning mechanisms, the micro-nano satellite constellation can achieve high-frequency, high-efficiency, low-energy consumption, and robust on-orbit calibration mission execution capabilities in complex dynamic environments.

[0060] S1: Construct a micro-nano satellite constellation mission model, define the six root parameter space of satellite orbits and establish the mission objective function, and determine the constraints, including orbital mechanics constraints, inter-satellite visibility constraints, and energy balance constraints.

[0061] S2: The genetic algorithm performs global optimization, randomly generates multiple constellation configurations, calculates the fitness of each satellite in the constellation, performs crossover and mutation on each selected individual, and finally performs convergence judgment.

[0062] S3: Deep reinforcement learning local optimization models the problem as a multi-agent problem, treating each satellite in the constellation as an agent, defining its environment, state space, action space, and setting a reward function, and learning the best strategy through continuous optimization.

[0063] S4: The GA-DRL coupled optimization process uses several optimal initial constellation configurations output from S1 as the starting point for DRL training. The initial configurations provided by GA reduce the exploration space of DRL, accelerate convergence, avoid local optima, and improve learning efficiency and convergence speed. The trained DRL model autonomously plans the constellation, including orbit control, task allocation, and calibration scheduling. It outputs an optimized constellation operation strategy, achieving efficient calibration task planning.

[0064] S5: Constellation Cooperative Control and Mission Execution. The trained DRL strategy is embedded into the constellation management and control system. During operation, satellites autonomously execute calibration mission scheduling and orbital maneuver control using on-orbit computing power. Information sharing and coordination are conducted within the constellation via inter-satellite links, and the actual execution results are fed back. Online fine-tuning or retraining is performed as needed to maintain long-term adaptability and robustness.

[0065] Step S1, which comprehensively considers satellite parameter constraints, satellite orbit constraints, inter-satellite visibility constraints, and energy balance constraints, involves the following steps to construct a micro / nano satellite constellation mission model: First, the satellite constellation and parameter space are defined. Then, the mission modeling of the micro / nano satellite constellation is formalized, providing standardized inputs and constraints for subsequent GA global search and DRL online optimization. The core output is: parameter space. Fitness function Constraint Sets And the discretization evaluation process.

[0066] Establish a six-element orbital parameter vector for each satellite in the constellation: .

[0067] in, For the semi-major axis, For eccentricity, For the track inclination angle, Right ascension of the ascending node, Angular distance from perigee Assume the initial mean anterior angle; and construct the full constellation parameter space: .

[0068] in, This is an operator used to map random variables to a uniform distribution; This represents the maximum allowable value for the semi-major axis of the track. This represents the maximum permissible value for the orbital eccentricity. This represents the minimum allowable value for the track inclination angle; This represents the maximum permissible value for the track inclination angle; For the number of satellites, Using the sequence number, form a constellation parameter vector. .

[0069] An orbital dynamics model is constructed, and appropriate evaluation criteria are established to define a valid calibration, as well as an energy balance model and inter-satellite visibility and constraint model to ensure the normal operation of satellites within the constellation. Finally, energy consumption and maneuver limitations are defined. The orbital dynamics model is as follows: Based on the Earth's gravitational constant With the perturbation term, construct the satellite orbital dynamics equations: .

[0070] .

[0071] in, The factors include Earth's J2 perturbation, air resistance, and gravitational disturbances caused by a third body. To control acceleration; within a given time step The internal propagation is obtained using numerical integration. , .

[0072] for The first derivative; for Time of the first The velocity vector of a micro-nano calibration satellite in space; for Time of the first The position vector of a micro-nano calibration satellite in space; for The first derivative.

[0073] For ground targets or main sensor locations Determine satellite visibility conditions: Inter-satellite visibility constraints include inter-satellite geometric visibility determination functions and visibility indicators.

[0074] The interstellar geometric visibility determination function is denoted as: Used to determine the first micro-nano calibration satellite and the first Position vector of a micro-nano calibration satellite in space , Whether an interstellar connection is obscured by Earth is considered to have interstellar visibility if the following conditions are met: .

[0075] Visibility indicator quantity is denoted as ,when At the time of its establishment, ;otherwise .

[0076] And calculate the angle of incidence. With relative velocity When satisfied .

[0077] .

[0078] .

[0079] And the continuous time exceeds When this occurs, it is considered a valid calibration event. The minimum time interval allowed during reinforcement learning or system operation.

[0080] Accumulated number of constellation calibrations: .

[0081] Define each satellite in time The state of charge (SOC) changes as follows: .

[0082] .

[0083] .

[0084] in, For the first Micro-nano calibration satellites in The state of charge of the battery at that time; For the first Micro-nano calibration satellites in The state of charge of the battery at that time; For time step; For the first The rated energy storage capacity of the satellite battery; For the first Micro-nano calibration satellite at time Power generated by solar cell arrays; For the first Energy conversion efficiency of a satellite's solar cell array; For the first The effective light-receiving area of ​​the satellite's solar cell array; To reinforce the initial state distribution set of the learning environment; For the first A satellite at any time The angle of incidence of the sun; For the first A satellite at any time Total load power consumption; Power consumption for basic maintenance of satellite platforms; For the first A satellite at any time Inter-satellite power consumption; For the first The satellite calibration payload at time Power consumption; For the first The power consumption of a satellite during orbital or attitude maneuvers.

[0085] The constraints are: .

[0086] Total speed increment of computer: .

[0087] And impose constraints: .

[0088] Construct the objective function, define constraints and penalty functions, and perform fitness calculations to predefine subsequent inputs: Under the premise of ensuring the feasibility of the constraints, the objective function of the task is defined as follows: .

[0089] .

[0090] All are weighted coefficients, representing the benefits of calibration times, energy consumption costs, and maneuver costs, respectively.

[0091] Let the constraint set be: .

[0092] constellation configuration At any moment The first one satisfied Inequality constraint functions; constellation configuration At any moment The first one satisfied Equality constraint functions.

[0093] These include orbital boundary constraints, energy constraints, maneuver constraints, and inter-satellite distance constraints; if a solution violates a constraint, a penalty function is applied. .

[0094] in, This is the penalty coefficient. Within the time interval... Inner step length The propagation trajectory is analyzed and various indicators are calculated to obtain the number of calibrations, energy consumption, and maneuvering, and the fitness value is obtained by combining these results. This serves as the input for subsequent optimization using genetic algorithms and deep reinforcement learning. This refers to the index of the constraint conditions. The genetic algorithm, based on the above model, processes the constellation parameter vector... Perform a global search based on the fitness function. To optimize the objective, several optimal constellation configurations are obtained through selection, crossover, and mutation operations.

[0095] Based on the task model, a genetic algorithm is used to process the constellation parameter vector. Global optimization includes the following steps: First, initialize the population, including every satellite in the constellation; then calculate the fitness of each individual satellite: In the already constructed parameter space Within, an initial population is randomly generated. : .

[0096] in, For population size, . The serial number is 1- .

[0097] Each individual Including constellations The set of six orbital numbers for the satellites: .

[0098] in, The first in the initial population The first individual One orbital parameter; For the first Among the individuals, the first The semi-major axis of the satellite's orbit; For the first Among the individuals, the first The semi-major axis of the satellite's orbit; For the first Among the individuals, the first The orbital inclination of the satellite; For the first Among the individuals, the first The right ascension of the ascending node of a satellite; For the first Among the individuals, the first The right ascension of the ascending node of a satellite; For the first Among the individuals, the first The perigee argument of a satellite.

[0099] Parameters can be generated randomly using the following methods: .

[0100] in, This is an operator used to map random variables to a uniform distribution; This is an operator used to map random variables to a uniform distribution; This represents the maximum allowable value for the semi-major axis of the track. This represents the maximum permissible value for the orbital eccentricity. This represents the minimum allowable value for the track inclination angle; This represents the maximum permissible value for the track inclination angle.

[0101] For each individual (No. Using the task model, calculate the objective function: .

[0102] And impose constraints and penalties: .

[0103] If an individual satisfies all constraints (orbital boundaries, energy balance, inter-satellite distance, maneuver limitations, etc.), its original fitness is retained. If a violation is found, points will be deducted according to the severity of the violation.

[0104] The fitness of all individuals is normalized as follows: .

[0105] Normalized As the probability of selection. For the first The middle generation The corrected fitness function value for each individual; For the first The first one selected in the generation for comparison or crossover The modified fitness of each individual. For the first The middle generation The corrected fitness function value for each individual.

[0106] For each individual calculated, selection, crossover, and mutation operations are performed to ensure that better individuals are obtained.

[0107] The Tournament Selection method is used, based on normalized fitness. Calculate the probability of an individual being selected. : .

[0108] For the first The middle generation Normalized fitness of an individual.

[0109] Generate the parent set in the new population : .

[0110] in, It is acceptable . This represents the number of individuals in each generation of the population in the genetic algorithm. All are sets of parents in the new population Individuals within.

[0111] Randomly selected from the population Individuals form a subset Choose one Individuals enter the parent set.

[0112] Based on the crossover probability of parent individuals Perform a crossover operation to generate new individuals: for any parent... , Perform arithmetic crossover: .

[0113] in . The first offspring generated by the crossover operation; These are the weights or mixing coefficients for the crossover operation in the genetic algorithm; This is the second offspring generated by the crossover operation.

[0114] Group crossover is used to exchange satellite-level data: that is, several sets of satellite IDs are randomly selected. The six roots of satellites within the set are interchanged between their parent generations: .

[0115] in, This is the second offspring generated by the crossover operation; For the first Individuals in Generation The Gene parameters. For the first Individuals in Generation The Gene parameters With the probability of mutation Apply random perturbations to some orbital parameters of the new generation of individuals: .

[0116] in, or . For the first The perturbation step size of each gene locus during mutation operations; For the first The perturbation step size of each gene locus during mutation operations.

[0117] To prevent parameters from going out of bounds, the mutated parameters need to be pruned back to the feasible range: .

[0118] If an adaptive mutation strategy is adopted, then according to algebra... Gradually reduce the intensity of variation: .

[0119] in, As the attenuation factor, This represents the maximum number of iterations. For the first The middle generation The intensity of variation at each gene locus; For the first The initial variation intensity of each gene locus.

[0120] After each iteration, the output individuals are evaluated to determine if they have converged. The best or second-best individuals are combined to obtain a preliminary constellation configuration set, as well as a connection point reserved for subsequent deep learning. The main steps are as follows: Record the optimal fitness after each generation iteration. : .

[0121] And calculate the relative improvement rate between successive generations: .

[0122] If any of the following convergence conditions are met: .

[0123] Then the algorithm is considered convergent.

[0124] Output several sets of optimal or suboptimal constellation parameter vectors. : .

[0125] The preliminary set of globally optimal constellation configurations is obtained: .

[0126] This serves as the initial solution set for the next part of deep reinforcement learning optimization.

[0127] Output constellation parameter set As the initial state set of the agent: . All are constellation parameter sets The constituent elements of.

[0128] Each state Including satellite orbital elements and energy status Task visibility matrix, etc.

[0129] in, Disturbance or noise figure used for adaptive adjustment; The threshold for determining convergence.

[0130] This state set will be input into the deep reinforcement learning training module. Using the global optimization result of GA as the initial structure, the efficiency of dynamic calibration will be further improved through local decision-making of DRL, forming a coupled optimization process.

[0131] With several optimal initial constellation configuration sets Using the DRL model as input, the constellation is autonomously planned, enabling it to acquire autonomous and intelligent calibration mission planning and orbit and attitude coordinated control capabilities. For the first The optimal or suboptimal constellation configuration solution.

[0132] First, initialize the policy by mapping the GA output to the initial state distribution of the DRL, define a multi-agent Markov decision process, and define a reward function that keeps the immediate reward consistent with the long-term goal. The steps are as follows: For each GA output individual The initial position set is obtained by mapping the six orbital numbers of the first part. With the initial velocity set : .

[0133] And set the initial power level. Initial task queue and inter-satellite link matrix .

[0134] Construct the initial state set: .

[0135] Set global state (Including all satellites) (etc.), local observation For satellite Available information (own orbit, energy, neighbor links, target visibility), joint actions .

[0136] System transitions follow a probabilistic transition function: .

[0137] in, It is determined jointly by orbital dynamics, energy equations, and communication delay models.

[0138] Each agent The strategy is to parameterize the probability distribution. .

[0139] Or concentrated strategy It depends on the training paradigm. The joint strategy is... .

[0140] Define the training objective and provide specific algorithm training details, including single-agent global policy and multi-agent communicative methods, ensuring usability in different scenarios. The steps are as follows: The overall training objective is to maximize the expected return over all initial states: .

[0141] in, Represents the trajectory.

[0142] Policy gradient form: .

[0143] Among them, the dominant function . For about The gradient operator; for The strategy function; In strategy Below, the state-action value function; In strategy Below, the state value function. This represents the parameter vector of the deep reinforcement learning policy network.

[0144] Multi-Agent Proximal Policy Optimization (MAPPO) intensive training: Define probability ratio : .

[0145] In order to be in Under the strategy, at any time The instant reward received; In order to be in Under the strategy, at any time The instant reward received.

[0146] Proximal Policy Optimization (PPO) pruning target (sample expectation) : .

[0147] in, For the sake of advantage estimation, The shearing parameter is 0.1–0.3. For time step The expected value operation of the random variable.

[0148] Adding the value function and entropy regularization term to the overall loss : .

[0149] in, Parameterization of value functions As weight, Entropy.

[0150] MADDPG / Centralized Commentator (Suitable for continuous actions and distributed execution) For each intelligent agent Establish a centralized commentator The training objective is to minimize the mean square Bell residual. : .

[0151] Where the target value : .

[0152] From target strategy produce. For time step Expectation operation of random variables; To reinforce the next state in learning; To reinforce the next action in learning. For the experience replay pool The expectation operation is performed on the quadruple of state, action, reward and next state obtained from the sampling. To reinforce the discount factor in learning; For the first Each agent uses target parameters The state-action value function in the next state Next action The estimated value below.

[0153] Agent policy updates employ deterministic policy gradient (DPG): . For the reason about the Agent policy parameters The gradient of the expected return function; For the distribution of empirical data Expectation calculation of the state sample; For the reason about the Gradient operator for the action variables of an agent. For the reason about the Agent policy parameters The gradient operator; For the first An intelligent agent observes Below, by parameters A defined strategy function.

[0154] Stabilization (soft update) is performed using the target network. ).

[0155] The process of experience collection, replay, and GA trajectory guidance, where GA trajectories serve as expert samples, significantly reduces the inefficient sampling caused by initial random exploration, thereby improving DRL convergence speed and policy quality; and constraint safety is implemented to ensure policy safety and feasibility during training and deployment phases, satisfying constraints such as energy, collision, and upper bounds on maneuver, as follows: Generate a heuristic trajectory library using GA output configuration. For each Run the baseline strategy (rule-based scheduling and maneuver strategy) within the simulation environment and collect trajectory sets: .

[0156] Will Inject into the experience replay pool (For offline pre-training and online warm-start), and pre-training of the policy using behavior cloning (BC): . This is the behavioral cloning loss function.

[0157] Datasets generated by genetic algorithms The expected computation performed on the state-action sample pair; For parameters The strategy in the state Action mapping under [the specified action].

[0158] Pre-training steps: . The policy parameter solution that minimizes the behavior cloning loss function.

[0159] Then proceed to RL update.

[0160] Hybrid sampling (GA policy probability) can be used in the early stages of training. To balance exploration and utilization.

[0161] Introduce soft-constraint penalties into the training loss or immediate reward (consistent with the steps above), and introduce action shielding during the action output phase: if The output action violates hard constraints (such as...) Exceeding limits, resulting in a collision distance of less than [a certain value]. If so, then use the projection operator. Correction action: . For the first A satellite at any time Executable actions after safety constraint projection; For the action Operators for secure projection; In the state The feasible region of actions that satisfy all constraints; Let be the square norm of the action deviation.

[0162] During training, the Lagrangian method can also be used to transform constraints into Lagrangian penalties and perform joint optimization: . Let be the objective function for constrained reinforcement learning containing Lagrange multipliers; The parameters for the deep reinforcement learning strategy are: The expected return function at that time; For the first Lagrange multipliers corresponding to each constraint condition; This is the expectation operation on the set of trajectories; For the first The positive part of a constraint function.

[0163] in, ,right Perform dual ascents to approximate a policy that satisfies the constraints.

[0164] After training, the strategy will be implemented. Export as a spaceborne runnable model (quantization / pruning / compilation), and save as a model package. Export must meet on-board computing / storage / latency constraints: number of parameters. Inference delay .

[0165] Deployment interface: Before deployment, use high-fidelity simulation to... For offline validation, if the performance threshold is met (e.g., on the validation set)... If the improvement rate is greater than or equal to the specified percentage and there are no constraint violations, then the model will be uploaded to the satellite's onboard computing unit (in orbit), and an online monitoring module will be set up to record the operational data. .

[0166] Online fine-tuning: Under the premise of ensuring safety shielding, the satellite can collect logs in orbit and periodically transmit compressed experience back to the ground, perform offline and federated fine-tuning, and then issue new models; or perform on-device fine-tuning within a limited range (using small steps and conservative update rules) when there is sufficient on-board computing power.

[0167] Based on the strategy parameter set As input, model deployment and constraint acceptance ensure that the model size, real-time performance, and capabilities meet the requirements of onboard operation. Then, computing power is used to autonomously execute calibration task scheduling and orbital maneuver control. Within the constellation, information sharing and collaboration are achieved through inter-satellite links to ensure optimal mission completion. Feedback on actual execution results allows for fine-tuning or retraining as needed, maintaining the stability and adaptability of the entire system.

[0168] The steps for deploying the model, planning its scale to avoid failure during in-orbit deployment, and performing state estimation for each satellite to ensure correct policy execution are as follows: The policy parameter set obtained during training (or model package) Quantize and compress the data into a spaceborne executable format to meet spaceborne resource constraints. .

[0169] in, This is the upper limit of the parameter. For single inference delay, Available memory.

[0170] Perform offline verification before deployment: Verify the strategy in a high-fidelity simulation environment on the initial distribution. The performance meets the threshold: . This represents the average number of calibrations within a given time window. This represents the average number of calibrations within a given time window. The average velocity increment consumed within a given time window; Evaluation thresholds for runtime performance or model deployment.

[0171] Otherwise, return to the ground for further training or fine-tuning.

[0172] At any moment Each satellite Acquire observations The state estimate is obtained through an estimator (filter): .

[0173] in, Include and local visibility information . For the first The current available battery energy value of each satellite; For the first The satellite from the initial time to the time Historical observation sequence; For the first From the initial moment to The historical control input sequence executed at any given time.

[0174] The estimation error satisfies: . For the first A satellite at any time State estimates; For the first A satellite at any time State estimates; This is the maximum allowable threshold for state estimation error.

[0175] To ensure the safety of decision-making.

[0176] Each satellite Based on local observations / estimations (May include neighbor information) Perform policy reasoning to obtain action distribution or actions: . To determine the optimal policy after training in the observation The probability distribution of actions under the following conditions; In order to observe The action that maximizes the probability of the strategy.

[0177] in, It can contain several sub-actions, for example: . For the first A satellite at any time Attitude control input; For the proposed maneuver; This is the priority vector for the current task.

[0178] These respectively represent attitude control commands and maneuvers ( Commands and local task scheduling / priority allocation.

[0179] Map the policy output to executable control commands: . For the first A satellite at any time Comprehensive control input; To make the action The allocation operator is mapped to satellite-executable control variables.

[0180] And calculate instantaneous or cumulative values ​​for maneuver commands. : . For the first A satellite at any time The speed increment required to perform the maneuver; For the first The rate of increase in the cumulative consumption of each satellite; For the first A satellite at any time The speed increment required to perform the maneuver.

[0181] Perform a safety projection before sending the action: .in, To meet the current hard constraints (SOC, A set of actions (including upper limit, collision distance, etc.).

[0182] After the above steps, the satellite undergoes a dynamic update to reflect the actual state of the propulsion system in the discrete propulsion system and update resource consumption and event statistics. Through inter-satellite links, the satellites collaboratively complete task allocation and maneuver negotiation to achieve constellation-level optimal execution. Online monitoring, logging, and data transmission ensure that the ground can obtain operational data for long-term optimization and safety auditing. In case of emergencies, a safety rollback procedure is in place as follows: Satellite execution control quantity The dynamics are updated according to the equations in Part 1: .

[0183] in, Including natural perturbation and controlled acceleration (and correspond). For the first The position vector of the satellite at the next moment; For the first The position vector of each satellite at the current moment; For the first The velocity vector of the satellite at the next moment; For the first The velocity vector of the satellite at the current moment; The time between the status update and the walk is long.

[0184] Energy update (based on discrete energy model): . For the first A satellite at any time The state of charge of the battery; For the first A satellite at any time Solar power generation capacity; For the first A satellite at any time The load power consumption; For the first The rated energy storage capacity of the satellite battery.

[0185] in, Corresponding motor power consumption included Update calibration instructions Event Count .

[0186] Inter-satellite communication model: If satellites and At any moment Satisfying inter-satellite visibility and link conditions (see Part 1): . For a moment No. With the Visibility indicators between satellites; For the first The position vector of each satellite; For the first The position vector of each satellite; The maximum distance allowed for inter-satellite communication or cooperative operation.

[0187] The set of exchangeable messages .

[0188] The exchanged content may include local estimates Current task priority vector Proposed motor and log summary .

[0189] Distributed coordination strategy: Each benchmarking task Released within the constellation, satellite i calculates competitive value (marginal revenue): .

[0190] in, This refers to the expected increase in rewards after undertaking the task. For consumption (motorization and energy). This is the cost weighting coefficient.

[0191] Auction decision: The task will be undertaken by the entity with the largest... The satellite is the carrier.

[0192] If the mission requires multi-satellite collaboration, a distributed consensus (average / maximization) approach is used to solve for collaborative burden sharing: . For constellations at all times For the The scheduling decision vector for each task; It is a distributed consensus decision operator based on competitive value sets.

[0193] The standard average / maximum or ADMM distributed optimization steps can be used. Collaborative decision-making satisfies the set of constraints. (SOC, Collision safety, etc.

[0194] Each satellite records operational log entries. : . This is a flag variable used to identify whether a constraint violation has occurred.

[0195] Compress logs Periodically (or when the link is available) back to the ground station or relay satellite for ground-based retraining or federated updates.

[0196] If a major constraint violation occurs (collision risk, SOC below the safety threshold), an emergency response procedure will be triggered (e.g., rollback to the safety baseline policy). ).

[0197] Subsequent on-orbit fine-tuning of the satellite, combined with on-device incremental fine-tuning and centralized ground retraining, balances real-time adaptability and policy safety to ensure the long-term stability of the entire system. Performance is also evaluated and feedback is provided to form a closed-loop control system. The operational results guide algorithm and parameter adjustments, maintaining long-term optimization. The steps are as follows: Local fine-tuning (only if on-board computing power allows and security policy permits): Collect local small batch trajectories Calculate the local gradient: . For the first The gradient vector calculated by a satellite or intelligent agent; For parameters The gradient operator; For the first Satellite sampling batches The loss function.

[0198] Update with a conservative step size: . For the first Local strategy parameters of the satellite after in-orbit fine-tuning; These are the final policy parameters after training is complete and validation is passed. The learning rate is fine-tuned online in orbit; This is the direction for normalized gradient updates.

[0199] Before updating, perform a short-term verification in the sandbox / simulation environment. If it is safe and the performance is improved, then replace the on-board model (or perform a soft merge).

[0200] Intensive retraining (ground-led): Onboard periodic backhaul of compressed gradients or model differences ; Ground aggregation: . These are the updated strategy parameters.

[0201] And on The update will be released after larger-scale verification.

[0202] Safety policy: On-orbit updates must meet the following requirements: . This represents the performance improvement of the new strategy on the validation set compared to the baseline strategy. The minimum performance improvement threshold required for policy updates.

[0203] Periodic evaluation indicators (using a sliding window): .

[0204] in, This is the weight vector.

[0205] If any metric shows a significant decrease or the violation rate increases, the rollback / retraining process is triggered.

[0206] Multiple redundancy: Baseline rule strategy can be selected at any time. Alternative : . The strategy adopted during actual on-orbit operation; Used for evaluating strategies Health check function for runtime status; This indicates that the strategy's operational status meets both security and performance requirements.

[0207] Periodically run evaluation missions on the ground or between satellites: calculate the expected value of the objective function over a fixed time window. . For constellations at all times The runtime performance evaluation function.

[0208] like If the expected result is not achieved (e.g., below the baseline of GA output or training expectation), one or more of the following actions will be initiated: Retrain the model on the ground and distribute the new model; adjust the weights or penalty coefficients. or Modify collaboration rules (auction / consensus parameters); expand data collection or add simulation scenarios.

[0209] This application has the following beneficial effects: 1. Achieve integrated global optimization of constellation configuration and mission planning. This application uses a genetic algorithm to perform a global search on the six orbital roots of the constellation to obtain the optimal or near-optimal initial configuration that meets the requirements of calibration coverage, energy constraints, and visibility. This effectively avoids the problems of traditional configuration design relying on experience and easily getting trapped in local optima, thus improving the overall mission reachability of the constellation.

[0210] 2. Significantly enhances on-orbit autonomous decision-making capabilities and real-time performance. This application deploys deep reinforcement learning strategies onto the satellite's on-orbit computing platform, enabling the satellite to autonomously generate calibration scheduling and orbital maneuver commands based on real-time observation conditions, energy status, and mission requirements, without relying on continuous ground control. This method can maintain decision-making continuity in the event of link interruptions, external disturbances, or unexpected missions, achieving highly timely and stable calibration task execution.

[0211] 3. Achieve constellation-level coordinated control and optimal resource utilization. Through inter-satellite links for state sharing and combined with a distributed cooperation mechanism, the constellation can autonomously complete task allocation, observation window coordination, and maneuver coordination planning in orbit, reducing redundant maneuvers and resource conflicts, and improving energy and propellant utilization efficiency. Combined with an online fine-tuning mechanism, the constellation can maintain efficient and robust calibration capabilities during long-term operation, significantly improving overall mission effectiveness.

[0212] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0213] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A cooperative control method for micro / nano calibration satellite constellations based on GA and DRL coupling optimization, characterized in that, include: Obtain constellation information data; The constellation information data includes: a predefined satellite constellation, a full constellation parameter space, and an orbital six-root parameter vector established for each satellite in the constellation; A micro-nano satellite constellation mission model is constructed based on the constellation information data. The micro-nano satellite constellation mission model includes an objective function and constraints. The objective function is determined based on the satellite orbital dynamics equations and, under effective calibration events, the number of constellation calibrations and the total speed increment of maneuvers. The effective calibration event is obtained after determining that the satellites within the constellation are operating normally when the satellite visibility conditions meet preset conditions and the continuous time exceeds a set threshold. A genetic algorithm is used based on the micro / nano satellite constellation mission model to globally optimize the constellation parameter vectors according to the fitness function within the entire constellation parameter space, thereby obtaining an initial constellation configuration set. The constellation parameter vectors are determined based on the orbital six-root parameter vectors. The fitness function is determined based on the objective function and the penalty function. The penalty function is determined based on the penalty coefficient under the condition of violating constraints. The strategy parameter set is determined using deep learning methods based on the initial constellation configuration set; Based on the strategy parameter set, the satellites are deployed in orbit and their states are estimated. Based on the satellite execution control variables, the satellites are dynamically updated. Through inter-satellite links, a satellite distributed cooperative strategy is adopted to complete task allocation and maneuver negotiation in order to achieve constellation-level execution control.

2. The micro / nano calibration satellite constellation cooperative control method based on GA and DRL coupling optimization according to claim 1, characterized in that, The mathematical expression corresponding to the objective function is: ; ; ; ; in, The objective function is... Number of times the constellation is calibrated; For constellation parameter vectors; The total energy consumed during the calibration task; This represents the increment of the total speed of the maneuver; , , All are weighted coefficients; The target revenue; For serial numbers; Number of satellites; Number the decision steps within the task's time window; For the first The orbital or attitude maneuvers of each micro-nano calibration satellite within the mission time window are numbered; The first in the constellations Micro-nano calibration satellites in the The speed increment required for the next maneuver; For the first The first micro-nano calibration satellite within the mission time window Load power consumption at each discrete decision moment or simulation time step; The first one within the task time window A discrete decision moment or simulation time step; For the first Micro-nano calibration satellites in The power generated by the solar cell array; For time step.

3. The micro / nano calibration satellite constellation cooperative control method based on GA and DRL coupling optimization according to claim 1, characterized in that, The constraints include: orbital mechanics constraints, inter-satellite visibility constraints, and energy balance constraints; The orbital mechanical constraints specifically include: ; ; ; The inter-satellite visibility constraints include the inter-satellite geometric visibility determination function and the visibility indicator; The interstellar geometric visibility determination function is denoted as: Used to determine the first micro-nano calibration satellite and the first The position vector of a micro-nano calibration satellite in space , Whether an interstellar connection is obscured by Earth is considered to have interstellar visibility if the following conditions are met: ; The visibility indicator is denoted as ,when At the time of its establishment, ;otherwise ; The energy balance constraints specifically include: ; ; in, For the first The position vector of a micro-nano calibration satellite in space; The position vector of the calibrated target payload or reference target in space; The maximum spatial distance allowed for calibration tasks; Angle of incidence; It is a continuous-time variable; This refers to the maximum allowable attitude deflection angle or maximum field of view angle when the satellite is performing a calibration mission. For the first The velocity vector of a micro-nano calibration satellite in space; The velocity vector in space for the calibration target or reference load; The maximum speed allowed for a satellite during its orbital operation or relative motion; The average radius of the Earth; For the first Micro-nano calibration satellites in The state of charge of the battery at that time; The minimum state of battery charge allowed for the safe operation of the satellite; This represents the maximum permissible state of charge for the satellite's batteries. The first one within the task time window A discrete decision moment or simulation time step; For serial numbers; In the constellation configuration scheme, the first The cumulative velocity increment required for a micro-nano calibration satellite to perform all orbital or attitude maneuvers within the mission time window; For the first The maximum velocity increment that a micro-nano calibration satellite is allowed to consume throughout its entire mission cycle; It is a constellation configuration scheme, consisting of the set of orbital parameters of each satellite in the constellation.

4. The micro / nano calibration satellite constellation cooperative control method based on GA and DRL coupling optimization according to claim 1, characterized in that, Based on the aforementioned micro / nano satellite constellation mission model, a genetic algorithm is employed to globally optimize the constellation parameter vectors within the entire constellation parameter space according to a fitness function, resulting in an initial constellation configuration set, specifically including: An initial population is randomly generated within the entire constellation parameter space; the initial population includes multiple individuals; each individual is a satellite within the constellation; each individual contains a set of orbital root numbers established by N satellites within the satellite constellation; Based on the aforementioned micro / nano satellite constellation mission model and the initial population, the first... The objective function corresponding to each individual is assigned, and constraints and penalties are applied. Then, selection, crossover, and mutation operations are performed on each individual. If each individual satisfies all constraints, the fitness of the original fitness function is retained. If there is a violation of the constraints, a preset deduction is applied according to the degree of violation. After each iteration, the optimal fitness is determined, the relative improvement rate between successive generations is calculated, and convergence is determined; the expression corresponding to the optimal fitness is: ; in, For optimal fitness; For the first In the first generation of genetic algorithms, the first generation... Individual constellation configuration The corrected fitness function value; This is the index of the individual in the genetic algorithm population; If the result is convergence, output the optimal or suboptimal constellation parameter vectors under the current iteration, and determine the initial constellation configuration set.

5. The micro / nano calibration satellite constellation cooperative control method based on GA and DRL coupling optimization according to claim 4, characterized in that, The convergence conditions for convergence determination include: ; ; in, The maximum number of iterations; For the first The relative improvement rate between successive generations corresponding to each generation; The threshold for determining convergence; For the first The optimal fitness corresponding to each generation; For the first The optimal fitness corresponding to each generation.

6. The micro / nano calibration satellite constellation cooperative control method based on GA and DRL coupling optimization according to claim 4, characterized in that, For each of the individuals, selection, crossover, and mutation operations are performed, specifically including: The fitness of all individuals is normalized, and the normalized fitness is used as the selection probability. Using a tournament method, the probability of an individual being selected is calculated based on the selection probability; the expression corresponding to the probability of an individual being selected is: ; A new set of parents is generated in the population based on the probability of the individual being selected. The parent individuals in the parent set are crossovered according to the crossover probability to generate new individuals. A random perturbation is then applied to some orbital parameters of the new individuals according to the mutation probability. To prevent parameters from going out of bounds, the mutated parameters need to be pruned back to the feasible range. The random perturbation includes: ; Feasible range includes: ; An adaptive mutation strategy is adopted, based on algebraic... Gradually reduce the intensity of variation: ; in, The probability of an individual being selected; For the probability of selection; Population size; For the first In the first generation of genetic algorithms, the first generation... Normalized fitness or intermediate evaluation value of each gene locus; For the first Generation Zhong, No. The individual in the first The probability of variation at each gene locus; For the first The individual in the first Random perturbation at each gene locus; For the first The minimum allowed value or probability lower limit for a gene locus; For the first The maximum allowed value or probability limit for a gene locus; For the first Generation Zhong, No. The intensity of variation at each gene locus; For the first The initial variation intensity of each gene locus; It is the attenuation factor; The maximum number of iterations; For serial numbers; After applying a random perturbation, the first Generation Zhong, No. The individual in the first The probability of variation at each gene locus.

7. The micro / nano calibration satellite constellation cooperative control method based on GA and DRL coupling optimization according to claim 1, characterized in that, The strategy parameter set is determined using deep learning methods based on the initial constellation configuration set, specifically including: The initial state set of the agent is determined by mapping the initial constellation configuration set; each state in the initial state set of the agent includes satellite orbital elements, energy state, and mission visibility matrix. A single-agent global strategy and a multi-agent interactive method are employed. The overall training objective is to maximize the expected reward value of all initial states in the initial state set of the agents, thereby determining the trajectory set. The expression corresponding to the overall training objective is: ; The trajectory set is injected into the experience replay pool, and policy training and updating are performed using behavior cloning. Action correction is then performed using a projection operator, and joint optimization is conducted using a Lagrange penalty term to obtain the policy parameter set. The mathematical formula for the joint optimization is as follows: ; in, The parameters for the deep reinforcement learning strategy are: The expected return function at that time; To perform expectation calculations on the initial state distribution; To reinforce the initial state of the learning process; To reinforce the initial state distribution set of the learning environment; In strategy The following is the expectation calculation performed on the trajectory distribution; To reinforce the state-action-reward sequence trajectory in learning; The accumulated reward value starting from the initial state; Let be the constraint reinforcement learning loss function containing Lagrange multipliers; For the parameter vector of the deep reinforcement learning policy network; These are Lagrange multiplier vectors; For the constraint number index; For the first Lagrange multipliers corresponding to each constraint; This refers to the expectation operation performed on the set of trajectories. For the first The positive part of a constraint function.

8. The micro / nano calibration satellite constellation cooperative control method based on GA and DRL coupling optimization according to claim 1, characterized in that, Based on the aforementioned strategy parameter set, on-orbit deployment is performed, and state estimation is conducted for each satellite. Furthermore, satellite dynamics are updated based on satellite execution control variables. Task allocation and maneuver negotiation are completed via inter-satellite links using a distributed satellite cooperative strategy to achieve constellation-level execution control. Specifically, this includes: The strategy parameter set is quantized and compressed to meet onboard resource constraints, and offline verification is performed. Furthermore, state estimation is conducted using filters based on the observations of each satellite to ensure the execution of the online deployment strategy. The expression corresponding to the onboard resource constraints is: ; After on-orbit deployment, the satellite's dynamics are updated based on satellite execution control variables to maintain its offline propulsion status and update resource consumption and event statistics. Task allocation and maneuver negotiation are completed via inter-satellite links using a distributed satellite cooperative strategy to achieve constellation-level execution control. In the event of a major constraint violation, a safety backoff is performed based on a safety baseline strategy. Major constraint violations include collisions and energy levels falling below a safety threshold. The expressions corresponding to the satellite execution control variables are: ; in, This represents the total number of parameters in the policy model after training. For the strategy parameter set; This is the upper limit of the parameter; The time for a single inference of the policy model on a spaceborne platform; For single inference delay; This represents the maximum inference time allowed for the spaceborne system. Storage resources required for the strategy model; Available memory; For satellite control; For the first A satellite at any time Control input; For control input Operators for performing constrained projections; A generic variable used to control input; The set of actions that satisfy the current hard constraints; For a moment The system state estimate.

9. The micro / nano calibration satellite constellation cooperative control method based on GA and DRL coupling optimization according to claim 8, characterized in that, The satellite distributed collaborative strategy specifically includes: Each benchmarking task Released within the constellation, the calibration mission is undertaken by the satellite with the highest competitive value; satellite competitive value The calculation formula is: ; If the calibration task requires multi-satellite collaboration, a distributed consensus solution should be adopted to share the burden: ; in, for The expected increase in rewards after undertaking a task; for Consumed in time; This is the cost weighting coefficient; For constellations Time for the first The scheduling decision vector for each task; It is an inter-satellite distributed consensus operator; For serial numbers.

10. The micro / nano calibration satellite constellation cooperative control method based on GA and DRL coupling optimization according to claim 1, characterized in that, Also includes: Perform on-orbit fine-tuning and periodic performance evaluation of the satellite; The on-orbit fine-tuning process specifically includes: Collect local small batch trajectories Calculate the local gradient: ; Update with a conservative step size: ; In-orbit updates must meet the following requirements: ; in, For the first The gradient vector calculated by each agent; For about The gradient operator; For the first Loss function calculated for each satellite sampling batch; For the first Local strategy parameters obtained by the satellite after in-orbit fine-tuning; For the strategy parameter set; The learning rate is fine-tuned online in orbit; For numerical stability parameters; The performance improvement of the new strategy compared to the baseline strategy on the validation set; The minimum performance improvement threshold required for policy updates.

Citation Information

Cited By

  • An inter-satellite communication window calculation method for remote sensing satellite task

    CN122178992A

  • A method for calculating inter-satellite communication windows for remote sensing satellite missions

    CN122178992B