Variable-scale satellite group orbit planning method based on graph near-end optimization algorithm

By employing a dynamic graph model based on graph near-end optimization algorithm and decentralized multi-agent reinforcement learning, the orbit planning problem of multi-satellite systems under dynamic node changes and communication-constrained environments is solved, achieving high-precision, real-time collaborative control of satellite constellations, which is suitable for the safe and efficient operation of large-scale satellite constellations.

CN121032282AActive Publication Date: 2025-11-28NORTHWESTERN POLYTECHNICAL UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511544042.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2025-11-28
Estimated Expiration
2045-10-28

AI Technical Summary

Technical Problem

Traditional centralized orbit planning methods are difficult to meet the requirements of multi-satellite systems in dynamic environments where nodes are added or removed at any time and links are unstable. Furthermore, existing reinforcement learning methods rely on a central node, which leads to compromised convergence when communication is limited.

Method used

A variable-scale satellite constellation orbit planning method based on graph near-end optimization algorithm is adopted. By constructing a dynamic graph model and a decentralized multi-agent reinforcement learning framework, each satellite independently completes decision-making and parameter updates. High-precision orbit simulation and policy optimization are performed using GNN-RNN and RK8 methods to achieve distributed training and execution.

Benefits of technology

It enables efficient and stable coordination of satellite constellations in environments with dynamically changing node sizes and limited communication, improves the accuracy and efficiency of orbit planning, reduces the system's dependence on the central node, and adapts to the engineering requirements of large-scale, long-life space missions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121032282A_ABST
    Figure CN121032282A_ABST
Patent Text Reader

Abstract

The invention discloses a variable-scale satellite group orbit planning method based on a graph near-end optimization algorithm. Real-time and efficient control under the conditions that the number of nodes dynamically changes and communication links are limited is achieved. Firstly, the orbit and communication relation of a satellite group on time is abstracted into a dynamic graph; then, a reinforcement learning structure based on superposition of the graph neural network and the recurrent neural network is constructed; a strategy and value function is asynchronously iterated on a satellite by adopting a reinforcement learning near-end strategy optimization algorithm, all calculation and parameter updating are completed at a satellite end, and a central master control satellite is not needed. After local algorithm updating is completed, each satellite exchanges parameter differences with the neighbor satellites meeting the reliability condition, and weighting is carried out according to the satellite link quality and physical distance mixed weight. The method can be widely applied to application scenes needing constellation-level collaboration such as earth imaging, global communication and navigation, and a complete technical system is provided for safe, efficient and autonomous operation of a large-scale satellite group.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of space orbit planning, and application of multi-agent reinforcement learning technology in orbit control and intelligent decision of a star cluster. BACKGROUND

[0002] With the maturity of batch manufacturing and commercial launch services, a constellation or formation composed of dozens to hundreds of satellites has become an important technical route for obtaining high temporal and spatial resolution remote sensing data and providing global continuous communication. Compared with single-satellite systems, such multi-satellite systems exhibit significant complexity in orbit scheduling and cooperative control. In addition, the rapid evolution of the number of satellites and the formation shape across the orbital plane makes it difficult for traditional orbit planning based on fixed-scale variables or centralized solvers to meet the engineering needs of real-time re-planning as the configuration changes.

[0003] The introduction of distributed and self-learning paradigms in multi-satellite orbit planning breaks through the bottleneck of traditional centralized optimization in large-scale applications. Deep reinforcement learning is considered a powerful means to improve the autonomy of satellite clusters because it can achieve end-to-end policy search under unknown disturbances, multi-source nonlinear constraints, and incomplete observation conditions. However, existing reinforcement learning research still follows the "centralized training-distributed execution" architecture: all trajectory data need to be aggregated to the ground or the main star server for unified gradient calculation, and then the updated network parameters are distributed to each node. This mode not only fails to meet real-time requirements in situations where inter-satellite links are limited and backhaul windows are short, but also implicitly requires the satellite scale to remain consistent during the training period and deployment, making it difficult to handle flexible scenarios where nodes are added or deleted at any time. Some research attempts to embed a graph neural network into a policy model to utilize time-varying topology information, but its parameter updates usually rely on centralized aggregation or parameter server frameworks, still unable to break away from global synchronization constraints, resulting in compromised convergence when links are intermittent or some nodes fail. In contrast, the fully decentralized graph reinforcement learning system proposed in the application avoids dependence on a central node from the underlying communication mechanism to the policy iteration process, with each satellite independently completing gradient calculation and exchanging compressed neural network parameter information only with a limited neighborhood, efficiently achieving continuous learning and robust cooperation under conditions of dynamic expansion of node scale, unstable links, and frequent reconstruction of topology.

[0004] In summary, there is an urgent need for a new orbit planning method that can adapt to the environment of independent operation with node addition and deletion, while maintaining high-order dynamic accuracy and computational affordability. The scalable orbit planning system for star cluster proposed in the application not only significantly improves the accuracy and efficiency of multi-satellite cooperative maneuvering, but also maintains the stability of the formation topology and the continuity of the task in a fully decentralized communication-limited environment. The system naturally adapts to changes in the size of the constellation, providing a solid theoretical foundation and engineering feasibility for future large-scale, long-life, and high-complexity space missions. SUMMARY

[0005] To solve the above problems, the application provides a variable scale satellite group orbit planning method based on a graph proximal optimization algorithm, which can be widely applied to application scenarios such as earth imaging, global communication, navigation and other constellation-level cooperation, and provides a complete technical system for safe, efficient and autonomous operation of large-scale satellite groups.

[0006] To solve the above problems, the application adopts the technical scheme of: A variable scale satellite group orbit planning method based on a graph proximal optimization algorithm, comprising the following steps: S1, establishing a dynamic graph model of satellite group orbit planning: at time , first describe the time-varying topological relationship of the satellite group with a dynamic graph sequence ; in the formula , the current in-orbit satellite set is given, the number of satellites increases or decreases in real time with the task; , the edge set satisfying the distance or communication constraint is represented; the position, velocity and remaining fuel of each satellite are packaged into the node feature matrix , the relative distance, relative velocity and link quality between two satellites are packaged into the edge feature tensor . The orbit motion and communication constraint are uniformly coded into a set of data formats with clear structure and automatic updating over time, and subsequent artificial intelligence algorithms only need to process this dynamic network graph to perform global collaborative planning.

[0007] Each satellite carries complete state information for decision-making in the graph, and the th satellite introduces a node feature vector , wherein , position, velocity and acceleration, respectively, describe the satellite motion state; is the remaining fuel, reflecting the maneuvering resource; , indicating the task state, so that the strategy can distinguish the satellite role and the fault condition. This coding compresses heterogeneous information into a fixed length, which is convenient for neural networks to directly process, and ensures that the network dimension does not need to be modified when the number of nodes changes through weight sharing. In the network model, an edge feature vector , the subscript indicates the two satellites connected by the edge, and the time superscript indicates the current time; is the Euclidean distance between the two satellites; the second term describes the difference between the velocity vectors of the two satellites, i.e. the relative velocity, which can reflect the potential convergence / separation trend; the third term This is the communication availability coefficient, where 1 indicates that the link is fully available at this moment, and 0 indicates that the link is interrupted. During the training process, reinforcement learning algorithms will automatically learn satellite pairs that are closer in distance, have lower relative speeds, and have more stable links, and these pairs should be given higher interaction weights in task allocation or formation control.

[0008] To explicitly adjust adjacency weights during message passing, a weighted adjacency matrix is ​​constructed.

[0009] in The formula controls the distance decay rate. It maps physical distance to a continuous coefficient between 0 and 1, naturally weakening long-distance information during graph convolution. This preserves necessary global coupling while suppressing noise propagation. Simultaneously, the exponential form ensures gradient smoothness, facilitating end-to-end training. Overall, these four formulas collaboratively define a variable-scale, real-time updated, and fully decentralized dynamic graph model for satellite swarms, providing a unified and scalable data foundation and information flow mechanism for subsequent GNN-RNN-based execution neural network-evaluation neural network reinforcement learning strategies.

[0010] S2. Fully decentralized multi-agent reinforcement learning framework and execution neural network-evaluation neural network structure design: Within a decentralized intelligent control framework, each satellite can make independent decisions based solely on its own detection data and limited neighborhood communication. Therefore, at any given moment... The Satellite structure local observation vector , in For the characteristics of the node itself, It is a set of neighbors that meet the communication / distance threshold. Based on the corresponding edge features, it is ensured that any decision depends only on "local star + neighbor" information; thus, a local strategy is derived.

[0011] In this model, the symbol Indicates the satellite's current position. The maneuvering command to be executed, i.e., acceleration. (Symbol) This is the internal memory vector stored in the previous time step of the recurrent network. It is used to retain historical information about orbital evolution to support current decisions. The subscripts omit satellite numbers, and the superscripts... This indicates that it originated from the previous time point. (Symbol) This represents the trainable parameters shared by the entire neural network (including graph neural networks and recurrent neural networks); since all satellites use the same set... They can perform reasoning and updates independently and in parallel on their respective onboard computers without relying on any central control node, thus achieving true decentralized collaboration.

[0012] The input data at the current moment is fed into the neural network, executing the neural network (Actor) - judging neural network (Critic). The first layer uses... A graph neural network with layer-wise message passing; for the first layer Hidden vectors have

[0013] Among them, time superscript Indicate the time corresponding to this calculation; subscript The satellite being updated represents all neighboring satellites that meet the communication and distance thresholds; It is the satellite's original feature vector (containing position, velocity, fuel, etc.). This node's input vector is specifically processed. It specifically processes the input vectors of neighboring nodes, and both are in the [number]th [node] of the entire graph. Sharing within layers makes the number of model parameters independent of the number of satellites; The attention weights are obtained through softmax normalization; higher values ​​indicate more neighbors. With the target satellite The closer the distance, the better the link, or the more relevant the state; This is a non-linear activation function used to enhance the network's expressive power. This process is iterated over... After the layers are completed, each satellite integrates the spatial dependency information of its neighborhood in a single inference process.

[0014] go through After layered graph neural networks, nodes Satellite at time Obtain spatial feature vectors This vector is compared with the memory vector stored in the gated loop unit at the previous time step. The data is fed into a single-layer GRU, where it is fused with current spatial information and historical orbital trajectories through gating operations, resulting in a new memory vector.

[0015] Because GRU automatically performs selective forgetting and updating of information, The dimension remains fixed, thus preserving short-term dynamic features while avoiding feature dimension expansion as the sequence lengthens, ensuring the stability and computational controllability of long-term sequence training.

[0016] In the execution neural network (Actor), the memory vector from the previous time step... The mean of the action distribution is generated through two fully connected channels. Covariance Then from the normal distribution The original thrust vector was obtained by sampling. and with differentiable Saturation function compressed to engine limit ,Right now This ensures that the gradient is propagable while also guaranteeing that the output thrust command always falls within the physically permissible range.

[0017] In the Critic network, a GNN-RNN encoder, from which spatial-temporal information has already been extracted, is used, and then the memory vector at its output is... (subscript) Satellite number, superscript The current state value is obtained by following it with a linear regression layer containing only one neuron.

[0018] in These are trainable weights specific to the Critic, which, without changing the current local policy, change from time point... The initial value is the expected value obtained by accumulating future rewards using a discount factor. The Critic output is a scalar state value used to measure the quality of an action and effectively reduce gradient variance, thereby accelerating convergence.

[0019] Both strategy and value are jointly updated using a proximal optimization loss with truncated importance weights; for each time-series empirical fragment, the advantage is first calculated.

[0020] in For discount rate, This is the target network under the old parameters; In order to update the parameters of the execution network A loss function called CLIP is used, which restricts the policy gradient during optimization, thereby maintaining the stability of updates. Specifically, the CLIP loss is calculated as follows:

[0021] CLIP loss is a loss function used to limit the magnitude of policy updates.

[0022] It is an importance-weighted ratio used to measure the current strategy. Compared to the old strategy The differences between them. It is a satellite In time Actions taken at all times (direction, magnitude, etc. of the thrust). It is the advantage function, which represents the improvement in value of taking a certain action compared to the baseline value of the current strategy. It calculates the superiority or inferiority of the current strategy relative to existing strategies. It is a limiting operation that will affect the importance ratio. Limited to Within the range. It is a pruning threshold, usually set to a small constant (e.g., 0.1), to prevent large policy updates and maintain the stability of the training process.

[0023] In addition, the Critic network updates its parameters by minimizing the mean squared error loss. The formula is:

[0024] in, It is to evaluate the network's performance on satellites At any moment status The state value prediction represents the estimated future returns of a satellite in that state. It is the actual discount reward, obtained through environmental feedback (rewards), representing the discount from the current moment. Accumulated returns on future discounts.

[0025] Each satellite utilizes a replay buffer containing only information about itself and its neighbors to compute gradients in parallel and asynchronously synchronize weights. This preserves decentralized execution while maintaining learning consistency through parameter broadcasting when bandwidth allows. The entire design linearly connects three types of information: "spatial interaction, time dependence, and policy optimization," forming a GNN+RNN execution neural network-evaluation neural network algorithm that requires no central node, distributes both training and execution, and is robust to the number of satellites and topology.

[0026] S3. High-precision iterative simulation of satellite orbital state based on the RK8 method: To ensure high accuracy in satellite orbital state evolution during reinforcement learning, this invention models the orbital dynamics of each satellite using a set of first-order ordinary differential equations. The set of equations is shown below:

[0027] in, It is the satellite's state vector, which contains its three-dimensional position. and speed Both are three-dimensional vectors; It is the thrust command output by the execution network (Actor), representing the direction and magnitude of the thrust applied by the satellite in three-dimensional space. Function This study integrates all influencing factors affecting the satellite, including Earth's principal gravity, J2 non-uniformity, atmospheric drag, and orbital perturbations such as solar radiation pressure. The state-acceleration relationship is directly solved by a numerical integrator. To ensure the accuracy and efficiency of numerical calculations, this invention employs the Runge-Kutta 8th-order method (RK8) for iterative simulation of the orbital state. At each environmental step... In the process, the 13-stage, 8th-order method of RK8 is used for updating. First, the slopes of the 13 stages are calculated. Its expression is:

[0028] in, It is the current adaptive step size. These are the coefficients used for updating; their specific values ​​are determined by the RK8 algorithm. These are constant coefficients in the RK8 method, used to adjust the weight of each slope; The acceleration is obtained by executing the network. Next, the current state is calculated based on these slopes:

[0029] in, and These are coefficients that satisfy the consistency of 8th-order algebras, ensuring the accuracy of numerical integration.

[0030] Orbital state calculated using the RK8 method As the next observation in the reinforcement learning environment, it is stored in the local replay buffer along with the corresponding reward. Because this method can effectively control errors while maintaining high-order accuracy, it ensures that the policy gradient more accurately reflects the real satellite orbit dynamics in subsequent reinforcement learning processes.

[0031] S4. Evaluate the decentralized fusion of neural network parameters and the policy generalization mechanism: During the decentralized training phase, each satellite Each maintains a set of evaluation neural network parameter vectors that are only relevant to itself. and according to the gradient obtained from local sampling First, perform a pure local stochastic gradient descent update to obtain intermediate results.

[0032] in To ensure local learning rate, each node can independently improve its value estimation even if it temporarily loses connection; subsequently, while ensuring communication link reliability... Exceeding the threshold and relative distance Not greater than neighborhood Inside, satellite With all Exchange their respective And the decentralized integration is completed in the form of a weighted average:

[0033] Among them, weight Determined by both link reliability and physical proximity, To balance the influence of both, The distance attenuation scale is used; this formula ensures that the weights are non-negative and sum to 1, thus embedding the classical consensus average into communication-constrained, dynamically changing space networks, and allowing neighbors with "good links and short distances" to contribute more to value estimation. Since all calculations rely solely on local neighborhood information, the entire fusion process can be executed asynchronously on each node, without requiring a global clock or centralized parameter server, and naturally supports topology changes: when new satellites are added… For access, as long as communication can be established with several neighbors, the initial weight can be quickly obtained using the same weight formula.

[0034] in It is the most recent synchronization time of the neighbor; when a satellite fails and exits, its parameters are automatically removed from the normalized denominator of its neighbors, without affecting the continuity of value estimation for the remaining nodes. This decentralized fusion mechanism significantly reduces the sensitivity of the evaluation neural network to single-node observation bias without increasing additional bandwidth, improves the generalization ability of the value function under different formation sizes and task scenarios, and ensures that the policy gradient still originates from a completely local information flow, which is in line with the patent design goal of distributed training-distributed execution.

[0035] The beneficial effects of this invention are as follows: 1. Completely decentralized, significantly improving reliability and real-time performance: Adopting a fully distributed architecture, each satellite independently completes decision-making and parameter updates, eliminating dependence on central nodes and the risk of single point of failure; on-board local computing avoids long-link communication delays, and can still achieve millisecond-level response in link-constrained scenarios, greatly improving the system's anti-interference capability and mission continuity.

[0036] 2. Naturally adaptable to dynamic changes in scale, with extremely high engineering flexibility: Through sparse index dynamic graph model and shared parameter design, it supports seamless expansion and contraction of satellite numbers (adding / retiring satellites only requires updating the index); newly added satellites can be quickly initialized through neighborhood parameter fusion, and failed satellites automatically leave the system without interrupting the mission or retraining the model, significantly reducing the cost of large-scale constellation deployment and maintenance.

[0037] 3. Excellent in both accuracy and efficiency, adaptable to complex aerospace scenarios: GNN-RNN fusion extracts inter-satellite space interaction and orbital history features, and the PPO algorithm ensures the accuracy of strategy decision-making; RK8 high-order numerical integration ensures that the orbit simulation is highly consistent with the real environment; the local parameter interaction mechanism optimizes bandwidth overhead, and can still achieve high-precision collaboration (such as position error ≤50m) and efficient mission execution in complex scenarios such as multi-source disturbances, heterogeneous nodes, and link discontinuity. Attached Figure Description

[0038] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0039] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0040] Reference Figure 1 A variable-scale satellite constellation orbit planning method based on graph near-end optimization algorithm includes the following steps: Step 1: Establish a dynamic graphical model for constellation orbit planning This implementation first considers the variable-size satellite constellation in time... The abstraction is a set of dynamic graphs that update automatically over time. This is used to uniformly describe orbital motion and communication topology. The core definitions are as follows:

[0041] subscript in formula Indicates the current time; in the formula Give the current set of satellites in orbit, and its number. The number of tasks can be increased or decreased in real time. Represents the set of edges that satisfy distance or communication constraints; the position, velocity, and remaining fuel of each satellite are encapsulated in the node feature matrix. The relative distance, relative velocity, and link quality between the two satellites are encapsulated into the edge feature tensor. .

[0042] Node feature encoding: In order for each satellite to carry the complete local information required for decision-making, the first feature encoding is defined. Node vectors of satellites

[0043] in These are position, velocity, and acceleration, which describe the satellite's motion state. It represents remaining fuel, reflecting mobile resources; This represents the mission status, allowing the strategy to distinguish between satellite roles and fault conditions. All components are linearly normalized and mapped to the [-1,1] interval to ensure uniformity of the dimensions of different physical quantities and facilitate gradient propagation.

[0044] Edge feature encoding: arbitrarily satisfying Satellite pairs construct edge vectors

[0045] Subscript Indicates the two satellites connected by the edge, with time superscript. Indicates the current moment; It is the Euclidean distance between the two satellites; the second term The difference in the velocity vectors of the two stars, i.e., their relative velocities, reflects potential convergence / separation trends; the third term... This is the communication availability coefficient, where 1 indicates that the link is fully available at this moment, and 0 indicates that the link is interrupted. This three-dimensional encoding encapsulates geometry, dynamics, and communication quality in a unified way, making it easier for the graph neural network to automatically learn during training that neighboring stars with "close distance, low relative speed, and good link" should be given higher interaction weights, without the need for manual parameter tuning.

[0046] Adjacency weight adjustment: To explicitly reduce long-distance noise during message passing, a weighted adjacency matrix is ​​defined.

[0047] in Let be the hyperparameter for distance decay. The exponential mapping compresses physical distances to a continuous interval (0,1], ensuring that the contribution of distant nodes to information aggregation decreases exponentially during graph convolution; simultaneously, the first derivative of the function is continuous, which is beneficial for end-to-end gradient optimization. Stored in a sparse format, computation is performed only on non-zero edges.

[0048] Dynamic graph update mechanism: Each control cycle is executed in three steps: ① The higher-order numerical integrator (see step three) outputs the new graphs for all satellites. ② Updated based on the latest distance and link assessment results and corresponding ③ If a satellite enters or leaves the network, simply insert or delete it. Corresponding index and adjustment The rows and columns can be reconstructed without rebuilding the entire graph. This design ensures that the model remains intact even when the number of nodes changes drastically. The overhead of adding or deleting levels.

[0049] Through the above four sets of formulas and update steps, the dynamic graph model constructed in this embodiment can accurately reflect the spatiotemporal coupling relationship of the satellite constellation within a millisecond-level refresh cycle, and provide a scalable and semantically consistent input tensor for the subsequent GNN-RNN execution-evaluation network.

[0050] Step 2: Design of a fully decentralized multi-agent reinforcement learning framework and execution-judgment network structure First, at every moment within, no. Each satellite constructs a local observation vector based on its own sensor readings and neighborhood information collected via available links.

[0051] in Node features, including inertial position ,speed The actual acceleration in the previous step Remaining fuel With task status codes ;gather Based on satisfying the "distance threshold" And link availability Composed of neighboring satellites; edge features These represent the Euclidean distance, velocity difference modulus, and link availability factor, respectively. This design ensures that any decision relies solely on locally observable data, satisfying decentralized execution constraints.

[0052] Then, Send to the first floor Graph Neural Networks (GNNs) with heavy message passing. For the first... Hidden Vectors (subscript) Indicates satellite node, layer number (Depth of graph neural network) execution

[0053] in .matrix It only applies to the features of this node. It only applies to the features of neighboring nodes, and both are shared across all nodes in the graph; Use LeakyReLU; attention weights

[0054] From a trainable scoring function Output and in the neighborhood Node satellite orientation normalization to ensure This allows nearby satellites with good links or similar states to have a greater impact during information aggregation. Spatial embedding is obtained after layer transfer .

[0055] To introduce historical dynamics dependence, Memory of the previous moment Input single-level gated recurrent unit (GRU):

[0056] Memory vector With a fixed dimension, independent of sequence length, GRU automatically retains or forgets information through reset and update gates, capturing short-term maneuvering trends while avoiding gradient vanishing.

[0057] The execution network (Actor) uses two fully connected mappings. , Will Converted into motion distribution parameters; then the original thrust vector was sampled from the Gaussian distribution. And limited to the engine upper limit through a differential saturation conversion. :

[0058] in These are globally shared parameters for the Actor; bold. This indicates the final acceleration command.

[0059] The Critic network shares the GNN-GRU encoder with the Actor network, but its tail is connected to a single-neuron linear layer that outputs the state value.

[0060] in To be independent The evaluation weights. This value function is defined as, with the current policy unchanged, from time [time]. The expected cumulative discount return is used as a low variance benchmark.

[0061] Strategy optimization: Employs a proximal optimization loss with truncated importance weights. For each length... Advantages of empirical trajectory calculation

[0062] Again

[0063] Update Actor, where It is an importance-weighted ratio used to measure the current strategy. Compared to the old strategy The differences between them. It is a satellite In time Actions taken at all times (direction, magnitude, etc. of the thrust). It is the advantage function, which represents the improvement in value of taking a certain action compared to the baseline value of the current strategy. It calculates the superiority or inferiority of the current strategy relative to existing strategies. It is a limiting operation that will affect the importance ratio. Limited to Within the range. It is a pruning threshold, usually set to a small constant (e.g., 0.1), to prevent large policy updates and maintain the stability of the training process.

[0064] Critic synchronization minimize

[0065] This represents a practically discounted return. Each satellite calculates its gradient and averages parameters asynchronously using only its own data and neighbor data, eliminating the need for a central server. The algorithm maintains consistent convergence rate and final performance without requiring parameter retuning, validating the framework's inherent robustness to changes in the number of nodes and topology.

[0066] Step 3: High-precision iterative simulation of satellite orbital state based on RK8 method This implementation describes the orbital motion and external thrust of each satellite as a unified system of first-order ordinary differential equations.

[0067] in, It is the satellite's state vector, which contains its three-dimensional position. and speed Both are three-dimensional vectors; It is the thrust command output by the execution network (Actor), representing the direction and magnitude of the thrust applied by the satellite in three-dimensional space. Function This study integrates all influencing factors affecting the satellite, including Earth's principal gravity, J2 non-uniformity, atmospheric drag, and orbital perturbations such as solar radiation pressure. The state-acceleration relationship is directly solved by a numerical integrator. To ensure the accuracy and efficiency of numerical calculations, this invention employs the Runge-Kutta 8th-order method (RK8) for iterative simulation of the orbital state. At each environmental step... In the process, the 13-stage, 8th-order method of RK8 is used for updating. First, the slopes of the 13 stages are calculated. Its expression is:

[0068] in, It is the current adaptive step size. These are the coefficients used for updating; their specific values ​​are determined by the RK8 algorithm. These are constant coefficients in the RK8 method, used to adjust the weight of each slope; The acceleration is obtained by executing the network. Next, the current state is calculated based on these slopes:

[0069] in, and These are coefficients that satisfy the consistency of 8th-order algebras, ensuring the accuracy of numerical integration.

[0070] An outer control cycle Composed of several different sizes Cascaded composition: The algorithm accumulates in an internal loop. When the accumulated time first exceeds Back to exactly using linear interpolation The state, and the state As the next observation of the intelligent agent by the environment.

[0071] Step 4: Evaluate the decentralized fusion of neural network parameters and the policy generalization mechanism. During the decentralized training phase, each satellite Each maintains a set of evaluation neural network parameter vectors that are only relevant to itself. and according to the gradient obtained from local sampling First, perform a pure local stochastic gradient descent update to obtain intermediate results.

[0072] in To ensure local learning rate, each node can independently improve its value estimation even if it temporarily loses connection; subsequently, while ensuring communication link reliability... Exceeding the threshold and relative distance Not greater than neighborhood Inside, satellite With all Exchange their respective And the decentralized integration is completed in the form of a weighted average:

[0073] Among them, weight Determined by both link reliability and physical proximity, To balance the influence of both, The distance attenuation scale is used; this formula ensures that the weights are non-negative and sum to 1, thus embedding the classical consensus average into communication-constrained, dynamically changing space networks, and allowing neighbors with "good links and short distances" to contribute more to value estimation. Since all calculations rely solely on local neighborhood information, the entire fusion process can be executed asynchronously on each node, without requiring a global clock or centralized parameter server, and naturally supports topology changes: when new satellites are added… For access, as long as communication can be established with several neighbors, the initial weight can be quickly obtained using the same weight formula.

[0074] in It is the most recent synchronization time of the neighbor; when a satellite fails and exits, its parameters are automatically removed from the normalized denominator of its neighbors, without affecting the continuity of value estimation for the remaining nodes. This decentralized fusion mechanism significantly reduces the sensitivity of the evaluation neural network to single-node observation bias without increasing additional bandwidth, improves the generalization ability of the value function under different formation sizes and task scenarios, and ensures that the policy gradient still originates from a completely local information flow, which is in line with the patent design goal of distributed training-distributed execution.

[0075] Step 4: Evaluate the decentralized fusion of network parameters and the policy generalization mechanism. In this step, each satellite Each maintains its own independent local criterion network parameter vector. and in discrete training weeks The value function is made fully decentralized through a two-stage process of "local update → neighborhood fusion". First, the satellite uses the gradient obtained from sampling in this cycle. Perform a single pure local stochastic gradient descent operation, with the learning rate denoted as . :

[0076] The satellite can still independently and promptly correct its value estimate based on its latest trajectory, avoiding error accumulation. Then, it enters the parameter fusion stage: the satellite only connects to systems that meet the "link reliability" requirement. "and distance" "Neighbors exchange intermediate parameters" Forming a neighborhood set

[0077] To integrate link quality and physical proximity, an unnormalized kernel function is introduced.

[0078] in Controlling the weights of the two factors, the exponential term makes the influence of distant nodes follow the trend. Exponential decay. For Normalization yields consistency weights

[0079] Thus, weighted average fusion is completed locally on the satellite.

[0080] The above formula ensures that the weights are non-negative and sum to 1, and that parameter updates remain within the neighborhood convex hull, suppressing noise from abnormal nodes and not relying on a global clock. Each satellite can asynchronously trigger fusion after detecting that all neighborhood information has arrived, exhibiting inherent topology flexibility. When a new satellite joins, as long as a link meeting the threshold is established, initial parameters can be quickly obtained using the same rules.

[0081] in As the first batch of communicable neighbors, if a satellite fails and leaves the network, its parameters are automatically removed from the denominator of the neighbor normalization, and the remaining nodes seamlessly continue iterating. Because and All are derived from the edge feature tensor of the dynamic graph. This means that distributed consistency of value functions is achieved on a dynamic, weighted, and asynchronous spatial network.

[0082] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A variable-scale satellite constellation orbit planning method based on graph near-end optimization algorithm, characterized in that, Includes the following steps: S1. Establish a dynamic graph model for satellite constellation orbit planning: Abstract the satellite constellation into a dynamic graph structure that evolves continuously over time. Each satellite is mapped to a node in the graph. Whether an edge is established between two satellites is determined by the relative spatial distance and the quality of the communication link. S2. Design a fully decentralized multi-agent reinforcement learning framework: Each satellite is an independent agent that makes decisions based solely on its own observations and neighborhood communication results, without the need for coordination from ground or central nodes. A combination of graph neural networks and recurrent neural networks is used to extract the historical dependence of space interaction patterns and orbital dynamics. The execution of the neural network branch maps the hidden state to a series of thrust commands, including direction and magnitude, and then executes them. The evaluation neural network branches assess the value of the same hidden state, providing a low-variance benchmark for policy gradients. S3. High-precision iterative simulation of satellite orbital state based on RK8 method: The eighth-order Runge-Kutta algorithm is used as the core numerical integrator. During the simulation, the controlled thrust, the non-spherical gravity of the Earth, the second-order oblateness perturbation, and the atmospheric drag are comprehensively considered as the combined acceleration input to the dynamic model to calculate the instantaneous state derivative. Subsequently, multiple slope samples are evaluated in one step, and the position and velocity at the next moment are predicted by high-order combination. The high-order integral error is strictly controlled, thus achieving high-precision orbit state prediction. S4. Implement a decentralized fusion and policy generalization mechanism for evaluating neural network parameters: After completing local gradient descent, each satellite sends updated value network parameters or gradient fragments to neighboring satellites and simultaneously receives data from them. The nodes calculate weights based on link reliability and relative distance, and perform a weighted average of multiple parameters to generate a new local evaluation neural network.

2. The variable-scale satellite constellation orbit planning method based on graph near-end optimization algorithm according to claim 1, characterized in that, In S1, once an edge is established, the real-time distance, relative speed, and link availability metrics are encapsulated as edge features. As the satellite moves along its orbit, the attributes of nodes and edges, as well as the entire topology, are automatically updated to ensure that the model continuously reflects real spatial relationships.

3. The variable-scale satellite constellation orbit planning method based on graph near-end optimization algorithm according to claim 1, characterized in that, In S1, the graph is stored using a sparse index. Adding or removing nodes only requires inserting or deleting index entries, without having to rebuild the entire graph. Therefore, it can naturally adapt to the batch launch or decommissioning of satellites during a mission.

4. The variable-scale satellite constellation orbit planning method based on graph near-end optimization algorithm according to claim 1, characterized in that, In S1, in order to balance global scale and local accuracy, the model supports hierarchical adjacency, which weakens long-distance weak coupling relationships to the point that they are negligible.

5. A variable-scale satellite constellation orbit planning method based on graph near-end optimization algorithm according to claim 1, characterized in that, In S1, the dynamic graph interface is decoupled from the reinforcement learning environment, allowing any external orbit simulator to inject state into the model through standardized feature tensors. This enables algorithm developers to quickly integrate new dynamic or constraint models without altering the core framework.

6. The variable-scale satellite constellation orbit planning method based on graph near-end optimization algorithm according to claim 1, characterized in that, In S2, all parameters are iteratively updated using a proximal policy optimization algorithm, the loss is calculated locally, and gradients can be shared within a limited neighborhood, thus maintaining training stability without introducing global dependencies.

7. A variable-scale satellite constellation orbit planning method based on graph near-end optimization algorithm according to claim 1, characterized in that, In S2, the framework naturally supports dynamic changes in the number of satellites and topology: new nodes can be inferred by loading shared weights, and the offline status of old nodes will not disrupt the training of other nodes.

8. A variable-scale satellite constellation orbit planning method based on graph near-end optimization algorithm according to claim 1, characterized in that, In S2, to reduce the burden, the network layers adopt parameter sharing, and batch inference can be executed in parallel on the same hardware; gradient synchronization only transmits weight differences, reducing communication traffic.

9. A variable-scale satellite constellation orbit planning method based on graph near-end optimization algorithm according to claim 1, characterized in that, In S4, the local consensus mechanism ensures that the value function remains consistent across the board when the topology changes, while avoiding the bandwidth pressure caused by large-scale broadcasting.

Citation Information

Patent Citations

  • Satellite network topology generation method based on deep reinforcement learning

    CN116319355A

  • Low earth orbit satellite constellation network edge calculation multi-stage unloading method based on reinforcement learning

    CN116634498A

  • Satellite computing power network task scheduling processing method and device

    CN117978238A

  • Star group orbit pursuit decision-making method based on multi-near-end reinforcement learning

    CN119962403A

  • Internet constellation dynamic routing optimization method and system

    CN120658306A