Planning method for agile observation satellite

By building the SatData benchmark suite and the Sat-Former scheduling model, combined with the Transformer architecture and internal constraint modules, the feasibility and accuracy issues of agile Earth observation satellite constellation scheduling in complex environments were solved, and efficient satellite mission planning was achieved.

CN120746211AActive Publication Date: 2025-10-03HOHAI UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511203624.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-10-03
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Existing agile Earth observation satellite constellation scheduling methods have difficulty in effectively handling complex constraints in large-scale and dynamic environments. Existing methods often simplify physical limitations, resulting in performance degradation.

Method used

Build the SatData benchmark suite and Sat-Former scheduling model, combine the Transformer architecture and internal constraint module, optimize satellite task allocation through reinforcement learning, and ensure accurate simulation and evaluation of constraint conditions.

Benefits of technology

It improves the feasibility and fidelity of satellite mission scheduling, reduces model complexity, and improves scheduling efficiency and accuracy in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120746211A_ABST
    Figure CN120746211A_ABST
Patent Text Reader

Abstract

The invention discloses a planning method for an agile observation satellite, and the method comprises the following steps: obtaining a satellite data set, and constructing a satellite scene model; generating and marking satellite task scheduling data based on the satellite scene model; obtaining a data set with a constellation scheduling label; constructing a scheduling model based on a Transform architecture, and embedding an internal constraint module, a satellite attribute matrix and a task attribute matrix for satellite task allocation; training and testing the scheduling model by adopting the data set, and obtaining a final satellite task scheduling model through reinforcement learning fine tuning; and during target task planning of the agile observation satellites, scheduling the corresponding agile observation satellites through the final satellite task scheduling model. The invention relates to the technical field of satellite communication, and the method comprises the steps: carrying out the satellite task distribution through constructing a benchmark test suite and a scheduling model; the objective of optimizing large-scale AEOS constellation scheduling task planning in a complex environment is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of satellite communication technology, and more particularly to a planning method for an agile observation satellite. Background Art

[0002] Agile Earth Observation Satellites (AEOS) have become a transformative technology in remote sensing, enabling rapid and flexible monitoring of the Earth's surface. By collaboratively planning in a constellation, multiple AEOS satellites can significantly increase revisit frequency and expand coverage, surpassing the capabilities of a single satellite.

[0003] Agile Earth observation satellite constellations offer unprecedented flexibility for monitoring the Earth's surface, but their scheduling remains challenging in large-scale scenarios, dynamic environments where missions can be launched and terminated at any time, and under strict constraints such as power and attitude. Existing approaches often oversimplify these complexities, limiting their performance in the real world.

[0004] Existing methods can be roughly divided into optimization-based methods and neural network-based methods.

[0005] 1. Optimization-based methods.

[0006] Early research relied on exact solvers to optimize satellite allocation. Some methods employed a constrained programming framework for agile satellite scheduling, while others used sequential convex programming to accelerate target acquisition. While these methods guaranteed optimality, their computational cost rose sharply with increasing problem size. Subsequent heuristic algorithms aimed to improve adaptability, including methods that balanced performance and runtime by switching perceptual task allocations, and methods that employed randomized hill climbing strategies for timely allocation optimization. Other approaches included ant colony optimization, evolutionary algorithms, and genetic algorithms. While these methods offered faster runtimes, their performance degraded in large-scale or dynamic scenarios.

[0007] 2. Methods based on neural networks.

[0008] The powerful fitting capabilities of neural networks have driven breakthroughs in constellation scheduling. New neural network-based methods have emerged, formulating the problem as a Markov decision process (MDP) and employing reinforcement learning for scheduling. Pointer networks provide a sequence-to-sequence (Seq2Seq) formulation for optimal allocation of multiple tasks. Methods employing GNNs and deep reinforcement learning to solve Earth observation satellite planning have also emerged, achieving highly competitive performance. Multi-agent RL has also been combined with polynomial-time greedy solvers to balance allocation quality and speed. Despite promising research results, these approaches often oversimplify key physical limitations.

[0009] Therefore, proposing a benchmark and model that combines physical constraints to perform constellation scheduling of agile observation satellites is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention

[0010] In view of this, the present invention provides a planning method for agile observation satellites. By constructing the SatData benchmark test suite and the Sat-Former scheduling model, it effectively addresses the limitations of existing satellite constellation scheduling methods and achieves the goal of optimizing the scheduling task planning of large-scale AEOS constellations in complex environments.

[0011] In order to achieve the above object, the present invention adopts the following technical solutions:

[0012] The present invention provides a planning method for an agile observation satellite, comprising the following steps:

[0013] S1. Obtain satellite data sets and build satellite scenario models based on satellite system models, satellite mission definitions, satellite action space control, and multiple constraints.

[0014] S2. Generate and label satellite mission scheduling data based on the satellite scenario model to obtain a dataset SatData with constellation scheduling labels;

[0015] S3: Build a scheduling model Sat-Former based on the Transformer architecture and embed internal constraint modules, satellite attribute matrix, and task attribute matrix for satellite task allocation.

[0016] S4. Using the dataset SatData with constellation scheduling annotations, train and test the scheduling model Sat-Former, and obtain the final satellite mission scheduling model through reinforcement learning fine-tuning;

[0017] S5. When planning the target mission of the agile observation satellite, the corresponding agile observation satellite is scheduled using the final satellite mission scheduling model.

[0018] Furthermore, the process of generating the satellite dataset includes:

[0019] Perform MRP empirical formula calculation, completion rate check and manual quality inspection for each satellite mission, and decide whether to accept or regenerate the MRP based on the inspection results;

[0020] Satellites that pass the inspection are regarded as qualified satellite assets and form a satellite dataset.

[0021] Furthermore, the satellite scenario model is constructed based on the satellite system model, satellite mission definition, satellite action space control and multiple constraints, specifically including:

[0022] Build satellite system models based on orbital dynamics, attitude control, power systems, and sensor payloads;

[0023] Satellite mission definition based on release time, expiration time, observation duration, and ground target coordinates;

[0024] High-level control is used to dispatch low-level control to control satellite movement space; wherein, the high-level control includes task allocation commands; the low-level control includes power switch commands and attitude pointing instructions;

[0025] Construct dynamic constraints, energy constraints, field of view constraints, continuity constraints and time window constraints to limit the execution of satellite missions; obtain a satellite scenario model including problem definition and scenario.

[0026] Furthermore, the step S2 specifically includes:

[0027] In the satellite scenario model, a greedy algorithm is used to perform initial allocation and obtain a preliminary allocation result;

[0028] Simulating the preliminary allocation results and retaining the observation trajectory of successful simulation;

[0029] Through manual quality review, high-quality scheduling trajectories are marked, and the dataset SatData with constellation scheduling annotations is obtained.

[0030] Furthermore, in step S3, the satellite attribute matrix includes a static satellite matrix and a dynamic satellite matrix; it is expressed as follows:

[0031]

[0032] Where L represents the satellite attribute matrix; L S represents the static satellite matrix; L d represents the dynamic satellite matrix; N L Indicates the number of satellites; d L Represents the satellite feature dimension;

[0033] The task attribute matrix includes a static task matrix and a dynamic task matrix; it is expressed as follows:

[0034]

[0035] Among them, R represents the task attribute matrix, R S represents the static task matrix; R d represents the dynamic task matrix; N R Indicates the number of tasks; d R Represents the task attribute dimension.

[0036] Furthermore, in step S3, the internal constraint module is used to predict the feasibility of the satellite performing the mission; the execution process includes:

[0037] S31. For each satellite-mission pair , internal constraint module Predict the feasibility of satellite i performing mission j , expressed as follows:

[0038]

[0039]

[0040]

[0041]

[0042] in, L represents the predicted logarithm of the feasibility of satellite i completing mission j; i represents the matrix of satellite i, R j represents the matrix of task j; represents the set of real numbers;

[0043] S32. Define approximate labels , to handle tasks that require collaboration among multiple satellites;

[0044] S33. Define the feasibility loss function and introduce time supervision to further guide the internal constraint module internalized constraints;

[0045] The feasibility loss function is expressed as follows:

[0046]

[0047] in, Represents feasibility loss.

[0048] Furthermore, in step S4, the training and testing of the scheduling model Sat-Former using the dataset SatData with constellation scheduling annotations specifically includes:

[0049] S41, using an embedding module to convert the sum into vector form and embed a sinusoidal time step; obtaining a satellite feature vector and a mission feature vector;

[0050] S42, using the Transformer decoder to decode the satellite feature vector and the mission feature vector to obtain the satellite feature h L and task characteristics h R ;

[0051] S43. The matching degree between satellite and mission is quantified by assigning a score matrix M. The feasibility prediction of the embedded constraint module is added to guide the planning. The formula is expressed as:

[0052]

[0053]

[0054] in, represents the predicted logarithm of the feasibility of satellite i to complete mission j; represents a trainable vector, represents the Hadamard product of element-wise multiplication, and F represents the The feasibility matrix obtained after 1 filling;

[0055] S44. Define a model loss function, evaluate the current task allocation result, and optimize the scheduling model Sat-Former; the model loss function is the loss function of the model allocation target:

[0056]

[0057] in, represents the model loss, and m represents the task truth value;

[0058] S45. During the testing phase, when quantizing the assignment score matrix M, infeasible satellite-mission pairs are excluded;

[0059] The formula is:

[0060]

[0061] in, represents a satellite-mission pair; represents the indicator function, represents a predefined feasibility threshold; Represents the jth value in the i-th row of the assignment score matrix.

[0062] Furthermore, in step S4, the final satellite mission scheduling model is obtained by fine-tuning through reinforcement learning, which specifically includes:

[0063] In the supervised pre-training stage, the Sat-Former is initialized with random weights and trained based on the labeled tracks in the SatData to obtain a total loss function;

[0064] The total loss function is expressed as follows:

[0065]

[0066] in, represents the total loss function; and denote the weights of and respectively;

[0067] During the reinforcement learning fine-tuning phase, a comprehensive evaluation is performed based on the corresponding total loss function; and the feasibility threshold is exceeded. The trajectory is added to the SatData for reinforcement learning of the Sat-Former until the model converges to obtain the final satellite mission scheduling model.

[0068] Furthermore, in step S4, the reinforcement learning specifically includes:

[0069] Initialize the value function network to evaluate the current state. Use Sat-Former as the policy function in reinforcement learning, train Sat-Former time-step by time-step, and perform probability sampling using the assignment score matrix output by Sat-Former.

[0070] Reinforcement learning is guided by a reward function, and a reward is calculated based on the current state; the time difference error is calculated based on the calculated reward and the next state; and the Sat-Former and value function network are updated based on the time difference error.

[0071] It can be seen from the above technical solutions that, compared with the prior art, the present invention provides a planning method for agile observation satellites, which has the following beneficial effects:

[0072] This paper constructs a large-scale, high-precision benchmark test suite, SatData, which accurately simulates physical characteristics such as satellite orbital dynamics and attitude control. It incorporates real-world satellite data for testing, ensuring realistic constraints and evaluation metrics. It also provides expert-annotated ground-truth scheduling annotations, addressing the lack of universal evaluation standards in existing benchmarks. This allows for fair comparison of models and ensures comprehensiveness and openness.

[0073] This paper builds a Sat-Former model based on the Transformer architecture, embedding an internal constraint module to explicitly model the physical and operational limits of satellites. This module guides scheduling decisions by predicting feasibility probabilities, improving the feasibility and fidelity of generated solutions.

[0074] The Sat-Former model divides the action space into high-level task assignments and low-level control instructions. The scheduling model focuses on task selection and timing, and the platform automatically converts these into low-level instructions, reducing model complexity. In dynamic data processing, the static properties and dynamic states of satellites and missions are integrated, and sinusoidal time embedding is introduced to enhance the model's ability to capture temporal features.

[0075] The internal constraint module uses binary cross-entropy loss (BCE) supervision to optimize feasibility prediction. During the satellite-mission matching phase, the cross-attention mechanism of the Transformer decoder achieves a deep fusion of constraint and feature matching, improving allocation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0077] Figure 1 A flow chart of a planning method for an agile observation satellite provided by an embodiment of the present invention.

[0078] Figure 2 A schematic diagram of generating a satellite dataset provided by an embodiment of the present invention.

[0079] Figure 3 This is a schematic diagram of generating the dataset SatData provided in an embodiment of the present invention.

[0080] Figure 4 This is a structural diagram of the Sat-Former scheduling model provided by an embodiment of the present invention.

[0081] Figure 5 This is a flowchart of the reinforcement learning fine-tuning provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0082] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0083] The embodiment of the present invention discloses a planning method for agile observation satellites, referring to Figure 1 As shown, the following steps are included:

[0084] S1. Obtain satellite data sets and build satellite scenario models based on satellite system models, satellite mission definitions, satellite action space control, and multiple constraints.

[0085] S2. Generate and label satellite mission scheduling data based on the satellite scenario model to obtain a dataset SatData with constellation scheduling labels;

[0086] S3: Build a scheduling model Sat-Former based on the Transformer architecture and embed internal constraint modules, satellite attribute matrix, and task attribute matrix for satellite task allocation.

[0087] S4. Using the dataset SatData with constellation scheduling annotations, train and test the scheduling model Sat-Former, and obtain the final satellite mission scheduling model by fine-tuning the model through reinforcement learning;

[0088] S5. When planning the target mission of the agile observation satellite, the corresponding agile observation satellite is scheduled using the final satellite mission scheduling model.

[0089] This example first constructs a large-scale, high-precision, and fully open benchmark suite, SatData, which includes 3,907 satellite datasets and 16,410 scenarios. Each scenario contains 1-50 satellites and 50-300 imaging tasks, covering 3,600 time steps. The scenarios are generated using the Basilisk engine high-fidelity simulation platform, accurately simulating physical characteristics such as satellite orbital dynamics and attitude control. Real satellite data is introduced for testing to ensure that the constraints and evaluation indicators are close to reality. The evaluation indicators cover six major dimensions, including task completion rate, turnaround time, and power consumption. Expert-labeled true value scheduling annotations are provided, and the data is publicly accessible, addressing the problem of the lack of common evaluation standards in existing benchmarks and supporting fair comparison of models.

[0090] Secondly, a Sat-Former scheduling model was constructed. Based on the Transformer architecture, an internal constraint module was embedded to explicitly model the physical and operational limits of satellites, guide scheduling decisions, and improve the feasibility and fidelity of generated solutions. A two-stage learning approach combining supervised pre-training and simulation exploration enabled the model to adapt to dynamic environments, discover optimal scheduling strategies, and enhance generalization capabilities. The action space was divided into high-level task assignments and low-level control instructions, reducing model complexity and improving the model's ability to capture temporal features. Dual supervision using binary cross-entropy loss and mean squared error loss enabled a deep fusion of constraint and feature matching, improving scheduling efficiency.

[0091] The embodiments of this invention significantly outperform existing technologies, such as REDA and EOSSP-RCS, in terms of mission completion rate, power consumption, and comprehensive scores. This has promoted standardization and toolchain development. The SatData of this invention will become the first large-scale benchmark for real-world AEOS constellation scheduling. The open-source code and data of Sat-Former facilitate technical replication and methodological innovation.

[0092] The specific steps of this embodiment are described in detail below.

[0093] First, build the benchmark suite SatData, which includes:

[0094] S1, obtains satellite data sets and builds satellite scenario models based on satellite system models, satellite mission definitions, satellite action space control, and multiple constraints;

[0095] This embodiment refers to Figure 2 The figure shows the process of a satellite mission from initial assessment to final confirmation, including calculating the MRP using an empirical formula, performing completion rate checks and manual quality inspections, and deciding whether to accept or regenerate the MRP based on the inspection results. Ultimately, satellites that pass all inspections are considered qualified satellite assets. This generates the satellite datasets in this embodiment, including MRP formula calculations and multiple inspections to ensure stable attitude control for each dataset.

[0096] The MRP empirical formula of this embodiment is:

[0097]

[0098]

[0099]

[0100]

[0101] in, Represents the proportionality constant of the MRP formula; Indicates the integral constant of the MRP formula; represents the differential constant of the MRP formula; Indicates the upper limit of points; represents the maximum value of the inertia tensor projection; represents the maximum angular momentum that all momentum wheels can provide; α, β, γ, δ∼U(0,1) represent random variables with uniform distribution of 0-1.

[0102] In this embodiment, a specific satellite and mission are taken as an example.

[0103] Satellite information is:

[0104] Mass: 210.05 kg

[0105] Moment of inertia: 107.87 kg·m²

[0106] Angular momentum: 50.00 kg·m²·s⁻¹

[0107] The mission points are:

[0108] Longitude: 0.000

[0109] Latitude: 5.239

[0110] Completion progress: 15 / 67

[0111] This embodiment performs manual quality inspection after the completion rate inspection;

[0112] This example evaluates the scheduler using six metrics, including task completion, timeliness, and energy efficiency. The completion rate (CR) measures the proportion of completed tasks to all tasks. The partial completion rate (PCR) assesses the ratio of maximum progress to the total required duration. The weighted completion rate (WCR) is a weighted version of the CR that takes task duration into account. The turnaround time (TAT) calculates the average time taken to complete a task, reflecting the efficiency of the scheduler. The power consumption (PC) quantifies the total energy consumed by the satellite sensor during imaging. Finally, the composite score (CS) aggregates these metrics into a single performance indicator:

[0113]

[0114] In this example, a value greater than 70% indicates a pass; for example, a CR of 5.00% and a PCR of 16.23% indicate a failure. For example, a CR of 90.00% and a PCR of 93.28% indicate a pass. For questionable cases, the MRP is regenerated. Satellites that ultimately pass the inspection are classified as satellite assets and form the satellite dataset for this example.

[0115] This example builds a benchmark suite, SatData, starting from defining the problem setting of AEOS constellation scheduling and describing the ground truth scheduling annotations for generating satellite datasets.

[0116] About the process of problem setting, that is, the process of building the satellite scene model.

[0117] This embodiment constructs a satellite scenario model based on a satellite system model, satellite mission definition, satellite action space control, and multiple constraints.

[0118] The satellite system model consists of four core subsystems: orbital dynamics, attitude control, power system, and sensor payload. It captures the essential physical properties that ensure mission feasibility. These include parameters such as the satellite's low-Earth orbit, orbital elements, mass properties, and moments of inertia, as well as control gains in attitude control and actuator limitations for each satellite. Low control gains result in slow attitude adjustments, while overloaded actuators can destabilize the satellite, increasing the risk of mission failure.

[0119] This embodiment collects satellite features into a static satellite matrix In which N L represents the number of satellites, Represents the static feature dimension of the satellite.

[0120] Imaging tasks arrive dynamically, and each task is defined by a release time, an expiration time, a required observation duration, and ground target coordinates. In this embodiment, these task descriptors form a static task matrix ,have tasks and Task attributes.

[0121] Regarding satellite action space control, this embodiment divides actions into two levels of abstraction, separating high-level scheduling from low-level control. The low-level action space includes power on / off commands and attitude pointing instructions, which are directly assigned to the Basilisk engine to simulate battery cycling, sensor activation, and MRP-based attitude maneuvers. While this provides maximum control flexibility, it introduces excessive complexity to the scheduling model. In contrast, the high-level action space of this embodiment consists of task allocation commands. The scheduler outputs the allocation vector , where each . Indicates that the satellite i sensor is off, and any Then instruct satellite i to activate sensors and redirect to the servicing mission The platform automatically translates these high-level assignments into low-level commands, allowing the scheduling model to focus purely on task selection and timing.

[0122] Real-world AEOS constellation scheduling is subject to multiple constraints. In this implementation, five constraints are enforced: dynamics, energy, field of view (FOV), continuity, and time window. Any advanced allocation that violates these constraints is rejected by the simulator, and only successful observations are recorded for later benchmarking.

[0123] S2, based on the satellite scenario model constructed in S1, generates and annotates satellite mission scheduling data, and obtains the dataset SatData with constellation scheduling annotations;

[0124] This embodiment refers to Figure 3 As shown in Figure 1, different mission scenarios are generated using satellite assets. These scenarios involve the assignment of tasks to multiple satellites at different time points. A greedy algorithm is used to assign initial tasks and generate a mission plan. Figure 3 The three-dimensional time-space grid displays the mission allocation: rows represent satellites, columns represent time points, and each square represents the mission status of a satellite at that point in time. The initial allocation is iteratively filtered based on completion rates to generate satellite trajectories. The resulting SatData benchmark library, annotated with constellation scheduling, contains 16,000 predefined satellite trajectories for evaluating and optimizing mission planning.

[0125] After the benchmark suite SatData is built, the scheduling model Sat-Former needs to be built, which includes:

[0126] S3, builds a scheduling model Sat-Former based on the Transformer architecture and embeds internal constraint modules, satellite attribute matrix, and task attribute matrix for satellite task allocation;

[0127] The Transformer architecture in this embodiment only includes a decoder. Figure 4 As shown in the figure, the static and dynamic data of satellites and missions are first concatenated and embedded. The decoder focuses on the satellite embedding and the mission embedding with cross attention. Then, the internal constraint module predicts the feasibility probability to guide the action selection.

[0128] Figure 4 Satellite static data of medium satellites These include satellite mass, satellite center of mass position, satellite orbit eccentricity, semi-major axis, inclination, right ascension of the ascending node, argument of perihelion, satellite solar panel properties, satellite imaging equipment properties, and satellite momentum wheel properties. Satellite solar panel properties include the orientation of the solar panel relative to the satellite, the area of ​​the solar panel, and the photovoltaic conversion efficiency of the solar panel; satellite imaging equipment properties include payload type (which may be multiple, such as optical or infrared payloads) and half-field of view; and satellite momentum wheel properties include static properties (maximum momentum, motor efficiency, and orientation).

[0129] Satellite dynamic data This includes the satellite's true anomaly, battery properties, payload power, current momentum wheel speed, and the satellite's current attitude. Battery properties include battery capacity and current battery charge percentage; the satellite's current attitude is expressed in the form of the Modified Rodriguez Parameter (MRP).

[0130] Task static data of the task Including the task release time, task deadline, continuous observation time, target latitude and longitude, imaging equipment payload type required for the observation task, task completion status and task mask. Regarding the task completion status, if the task has been completed, it is 1, otherwise it is 0; and the task mask is used to cover up illegal tasks. Task dynamic data Includes time steps and task progress.

[0131] Among them, regarding dynamic data, it is known that each scene in SatData is represented by a static satellite matrix L S and the static task matrix R SBy definition, they capture time-independent properties. Dynamic properties, such as mission progress and satellite attitude, are not contained in these static matrices. Enabling the scheduling model to infer dynamic state from static properties and past decisions would significantly increase complexity without significant benefit. Instead, this embodiment queries the simulator at each time step to retrieve the current dynamic satellite and mission properties. The complete input matrix is ​​formed by concatenating the static and dynamic components:

[0132]

[0133]

[0134]

[0135]

[0136] in, is a dynamic satellite matrix, is the dynamic task matrix; represents the satellite matrix dimension, represents the static parameter dimension of the satellite matrix, Represents the dimension of satellite matrix dynamic parameters; represents the task matrix dimension, represents the dimension of the task static parameter matrix, Represents the dimension of the task dynamic parameter matrix.

[0137] In this embodiment, the constellation attributes at the current time step are as shown in Table 1:

[0138] Table 1 Constellation attribute display table

[0139]

[0140] The attributes of the tasks to be completed are shown in Table 2:

[0141] Table 2 Task attribute display table

[0142]

[0143] The internal constraint module is used to predict the feasibility of the satellite to perform the mission. In this embodiment, for each satellite-mission pair (i, j), the internal constraint module Predict the feasibility of satellite i performing mission j , expressed as follows:

[0144]

[0145]

[0146]

[0147]

[0148] in, L represents the predicted logarithm of the feasibility of satellite i completing mission j; i represents the matrix of satellite i, R j represents the matrix of task j; represents the set of real numbers;

[0149] This embodiment also defines an approximate tag , handle tasks that require collaboration among multiple satellites; define a feasibility loss function and introduce time supervision to further guide the internal constraint module internalized constraints.

[0150] The feasibility loss function is expressed as follows:

[0151]

[0152] This supervision strategy enables learning the feasibility of satellite task assignments, effectively capturing the constraints present in the AEOS-Bench scenario.

[0153] After the model is built, according to S4, the scheduling model Sat-Former is trained and tested, and fine-tuned through reinforcement learning to obtain the final satellite mission scheduling model.

[0154] To match satellites and missions, the Sat-Former scheduling model of this embodiment uses a decoder architecture that jointly processes satellite and mission embeddings and is guided by an internal constraint module. The processing is shown below.

[0155] First, S41, this embodiment projects S and T into the embedding space and appends a sinusoidal time step embedding , expressed as follows:

[0156]

[0157]

[0158] in, represents the satellite feature vector, represents the task feature vector; and Represents an embedding module that looks up categorical data (e.g., sensor patterns) in an embedding matrix, while continuous categorical data (e.g., quality, progress) uses linear projection.

[0159] S42, use the Transformer decoder to decode the satellite feature vector and the mission feature vector to obtain the satellite feature h L and task characteristics h R ;

[0160] S43, quantify the matching degree between satellite and mission by assigning score matrix M, and add feasibility prediction of embedded constraint module To guide planning, the formula is:

[0161]

[0162]

[0163] in, represents a trainable vector, represents the Hadamard product of element-wise multiplication, and F represents the The feasibility matrix obtained after 1 filling is hour, ,otherwise Keep with equal;

[0164] S44. Define the model loss function, evaluate the current task allocation results, and optimize the scheduling model Sat-Former; the model loss function is the loss function of the model allocation target:

[0165]

[0166] Among them, m is the true value of the task;

[0167] S45. During the testing phase, when performing the quantization of the assignment score matrix, infeasible satellite-mission pairs are excluded;

[0168] The formula is:

[0169]

[0170] in, represents a satellite-mission pair; represents the indicator function, represents a predefined feasibility threshold; This embodiment tightly integrates learned constraints with feature matching to achieve efficient satellite task allocation.

[0171] This embodiment constructs a reinforcement learning process. The overall learning process is as follows: Figure 5 As shown in the supervised pre-training stage, Sat-Former is initialized with random weights and trained based on the labeled tracks in SatData to obtain the total loss function , expressed as follows:

[0172]

[0173] in, represents the total loss function; and Respectively and The weight of .

[0174] During the reinforcement learning fine-tuning phase, a comprehensive evaluation is performed based on the corresponding total loss function; and the feasibility threshold is exceeded. The trajectory is added to the SatData for reinforcement learning of the Sat-Former until the model converges to obtain the final satellite mission scheduling model. Reinforcement learning specifically includes:

[0175] S46. Initialize the value function network before reinforcement learning training begins ,in, It is a learnable vector used to evaluate the value of the current state and guide Sat-Former to generate a better scheduling solution.

[0176] S47. Using Sat-Former as a policy function in reinforcement learning , represents the probability of selecting action a in state L, R, where Represents all learnable vectors in Sat-Former, trains Sat-Former time-step by time-step, and outputs the assignment score matrix through Sat-Former Probabilistic sampling is performed and the formula is expressed as:

[0177]

[0178]

[0179] in, represents the satellite-task pair sampled at the current time step, i.e., the action in reinforcement learning, P represents the sampling probability, Represents the jth value in the i-th row of the assignment score matrix.

[0180] S48. Guide reinforcement learning through the reward function r, and calculate the reward based on the current state, using the formula:

[0181]

[0182] in, Indicates the number of satellites that are not currently assigned tasks. Indicates the number of tasks currently being observed, Indicates the number of tasks that have completed observations.

[0183] S49. Use the simulation platform to build a state transfer function , which can receive the current state and the sampled satellite-task pair and return the next new state and reward, expressed as:

[0184]

[0185]

[0186] Among them, L t and R t represent the satellite attribute matrix and mission attribute matrix at the tth time step respectively, represents the mission corresponding to the i-th satellite sampled at the t-th time step, represents all satellite-mission pairs sampled at the tth time step.

[0187] S410, reward obtained according to calculation And the next state calculation time difference error , expressed as follows:

[0188]

[0189] in, is the discount factor.

[0190] S411. Update the Sat-Former and value function network according to the time difference error, which is expressed as follows:

[0191]

[0192]

[0193] in, and represents the learning rate hyperparameter.

[0194] Use the pre-trained Sat-Former for inference to generate trajectories. Each generated trajectory is passed through the total loss The score defined in the calculation formula is used to evaluate. Then, this embodiment collects the performance exceeding the predetermined threshold. These high-quality trajectories are then added back to the SatData training set, and Sat-Former is fine-tuned using the new dataset. This cycle is repeated until convergence. In this way, Sat-Former continuously refines its strategies, discovers new strategies beyond the original annotations, and adapts to simulation-driven exploration and an increasingly diverse range of scenarios.

[0195] The present invention performs satellite scheduling through a unified framework that integrates a standardized benchmark suite and a new scheduling model. The benchmark suite SatData contains 3907 fine-tuned satellite datasets and 16410 scenarios. Each scenario has 1 to 50 satellites and 50 to 300 imaging tasks. These scenarios are generated through a high-fidelity simulation platform to ensure realistic satellite behavior, such as orbital dynamics and resource constraints. A true scheduling annotation is provided for each scenario. Based on this benchmark, the present invention introduces Sat-Former, a Transformer-based scheduling model that includes an attention mechanism for perceiving constraints. The internal constraint module explicitly simulates the physical and operational limits of each satellite. Through simulation-based reinforcement learning, the present invention enables Sat-Former to adapt to different scenarios and provides a robust solution for AEOS constellation scheduling.

[0196] Our benchmark suite, SatData, has four key characteristics: 1) Large scale. SatData includes 16,410 scenarios, each consisting of 1 to 50 satellites, 50 to 300 imaging tasks, and 3,600 time steps. 2) Real-world. All scenarios are generated and evaluated on our simulation platform to ensure accurate physical behavior of the satellites. The test set uses real satellite data from public sources, enabling evaluation on real data. 3) Comprehensiveness. SatData evaluates six metrics, including task completion rate, turnaround time, and power consumption. 4) Public data. Each scenario is annotated with ground-truth tasks through a rigorous process. All benchmark data and annotations are publicly accessible.

[0197] Our Sat-Former is a Transformer-based scheduler designed specifically for the AEOS constellation. At its core is a dedicated internal constraints module that explicitly models the physical and operational limits of each satellite, including sensor field of view, battery state, and attitude control time. This module guides scheduling by predicting feasibility probabilities.

[0198] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0199] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A planning method for an agile observation satellite, characterized in that: The following steps are involved: S1. Obtain satellite data sets and build satellite scenario models based on satellite system models, satellite mission definitions, satellite action space control, and multiple constraints. S2. Generate and label satellite mission scheduling data based on the satellite scene model; Get the dataset SatData with constellation scheduling annotations; S3: Build a scheduling model Sat-Former based on the Transformer architecture and embed internal constraint modules, satellite attribute matrix, and task attribute matrix for satellite task allocation. S4. Using the dataset SatData with constellation scheduling annotations, train and test the scheduling model Sat-Former, and obtain the final satellite mission scheduling model through reinforcement learning fine-tuning; S5. When planning the target mission of the agile observation satellite, the corresponding agile observation satellite is scheduled using the final satellite mission scheduling model.

2. A planning method for an agile observation satellite according to claim 1, characterized in that: In step S1, the process of generating the satellite dataset includes: Perform MRP empirical formula calculation, completion rate check and manual quality inspection for each satellite mission, and decide whether to accept or regenerate the MRP based on the inspection results; Satellites that pass the inspection are regarded as qualified satellite assets and form a satellite dataset.

3. The planning method for an agile observation satellite according to claim 1, wherein: In step S1, the satellite scenario model is constructed based on the satellite system model, satellite mission definition, satellite action space control and multiple constraints, specifically including: Build satellite system models based on orbital dynamics, attitude control, power systems, and sensor payloads; Satellite mission definition based on release time, expiration time, observation duration, and ground target coordinates; High-level control is used to dispatch low-level control to control satellite movement space; wherein, the high-level control includes task allocation commands; the low-level control includes power switch commands and attitude pointing instructions; Construct dynamic constraints, energy constraints, field of view constraints, continuity constraints and time window constraints to limit the execution of satellite missions; obtain a satellite scenario model including problem definition and scenario.

4. A planning method for an agile observation satellite according to claim 3, characterized in that: The step S2 specifically includes: In the satellite scenario model, a greedy algorithm is used to perform initial allocation and obtain a preliminary allocation result; Simulating the preliminary allocation results and retaining the observation trajectory of successful simulation; Through manual quality review, high-quality scheduling trajectories are marked, and the dataset SatData with constellation scheduling annotations is obtained.

5. The planning method for an agile observation satellite according to claim 1, wherein: In step S3, the satellite attribute matrix includes a static satellite matrix and a dynamic satellite matrix; it is expressed as follows: ; Where L represents the satellite attribute matrix; L S represents the static satellite matrix; L d represents the dynamic satellite matrix; N L Indicates the number of satellites; d L Represents the satellite feature dimension; The task attribute matrix includes a static task matrix and a dynamic task matrix; it is expressed as follows: ; Among them, R represents the task attribute matrix, R S represents the static task matrix; R d represents the dynamic task matrix; N R Indicates the number of tasks; d R Represents the task attribute dimension.

6. A planning method for an agile observation satellite according to claim 5, characterized in that: In step S3, the internal constraint module is used to predict the feasibility of the satellite to perform the mission; the execution process includes: S31. For each satellite-mission pair , internal constraint module Predict the feasibility of satellite i performing mission j , expressed as follows: ; ; ; ; in, L represents the predicted logarithm of the feasibility of satellite i completing mission j; i represents the matrix of satellite i, R j represents the matrix of task j; represents the set of real numbers; S32. Define approximate labels , to handle tasks that require collaboration among multiple satellites; S33. Define the feasibility loss function and introduce time supervision to further guide the internal constraint module internalized constraints; The feasibility loss function is expressed as follows: ; in, Represents feasibility loss.

7. A planning method for an agile observation satellite according to claim 6, characterized in that: In step S4, the data set SatData with constellation scheduling annotations is used to train and test the scheduling model Sat-Former, which specifically includes: S41, using an embedding module to convert the sum into vector form and embed a sinusoidal time step; obtaining a satellite feature vector and a mission feature vector; S42, using the Transformer decoder to decode the satellite feature vector and the mission feature vector to obtain the satellite feature h L and task characteristics h R ; S43. The matching degree between satellite and mission is quantified by assigning a score matrix M. The feasibility prediction of the embedded constraint module is added to guide the planning. The formula is expressed as: ; ; in, represents the predicted logarithm of the feasibility of satellite i to complete mission j; represents a trainable vector, represents the Hadamard product of element-wise multiplication, and F represents the The feasibility matrix obtained after 1 filling; S44. Define a model loss function, evaluate the current task allocation result, and optimize the scheduling model Sat-Former; the model loss function is the loss function of the model allocation target: ; in, represents the model loss, and m represents the task truth value; S45. During the testing phase, when quantizing the assignment score matrix M, infeasible satellite-mission pairs are excluded; The formula is: ; in, represents a satellite-mission pair; represents the indicator function, represents a predefined feasibility threshold; Represents the jth value in the i-th row of the assignment score matrix.

8. A planning method for an agile observation satellite according to claim 7, characterized in that: In step S4, the final satellite mission scheduling model is obtained through reinforcement learning fine-tuning, which specifically includes: In the supervised pre-training stage, the Sat-Former is initialized with random weights and trained based on the labeled tracks in the SatData to obtain a total loss function; The total loss function is expressed as follows: ; in, represents the total loss function; and denote the weights of and respectively; During the reinforcement learning fine-tuning phase, a comprehensive evaluation is performed based on the corresponding total loss function; and the feasibility threshold is exceeded. The trajectory is added to the SatData for reinforcement learning of the Sat-Former until the model converges to obtain the final satellite mission scheduling model.

9. A planning method for an agile observation satellite according to claim 8, characterized in that: In step S4, the reinforcement learning specifically includes: Initialize the value function network to evaluate the current state. Use Sat-Former as the policy function in reinforcement learning, train Sat-Former time-step by time-step, and perform probability sampling using the assignment score matrix output by Sat-Former. Reinforcement learning is guided by a reward function, and a reward is calculated based on the current state; the time difference error is calculated based on the calculated reward and the next state; and the Sat-Former and value function network are updated based on the time difference error.

Citation Information

Patent Citations

  • Agile satellite task parallel scheduling method based on data driving

    CN109960544A

  • Agile imaging satellite task planning method based on independent pointer network

    CN113051815A

  • Star group collaborative task planning method based on mixed expert experience playback

    CN117068393A

  • Multi-agent-based large-scale satellite collaborative observation task planning method

    CN117114317A

  • Rapid constellation configuration design method based on parallel reinforcement learning and genetic algorithm

    CN119647293A