Heterogeneous satellite cluster autonomous measurement and control resource allocation method, system and equipment

Through a two-stage decomposition strategy combining genetic algorithm with the Improved-MADDPG framework, the dynamic allocation of measurement and control resources in heterogeneous satellite clusters is solved, and the balance of task completion rate, resource utilization rate and cost-effectiveness ratio is achieved, and operation efficiency and adaptability are improved.

CN120509668APending Publication Date: 2025-08-19NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510641900.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The prior art is difficult to achieve the optimal dynamic allocation of measurement and control resources in heterogeneous satellite clusters, which makes it difficult to balance the task completion rate, resource utilization rate and cost-effectiveness ratio, and traditional algorithms have prominent problems in adaptability and computing complexity in dynamic environments.

Method used

Genetic algorithms are used to combine the improved Improved-MADDPG deep reinforcement learning framework, and resource allocation is performed through a two-stage decomposition strategy. The first stage is to perform three-dimensional matching of the task-satellite-visible time window through the genetic algorithm, and the second stage is to use the improved MADDPG framework to make distributed decisions to achieve optimal dynamic configuration.

Benefits of technology

It significantly improves the task completion rate, resource utilization rate and cost-effectiveness ratio, improves the operation efficiency and convergence speed, and solves the problems of calculation bottlenecks and insufficient adaptability of traditional methods in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120509668A_ABST
    Figure CN120509668A_ABST
Patent Text Reader

Abstract

The invention relates to a heterogeneous satellite cluster autonomous measurement and control resource allocation method, system and device. The problem of multi-target balance of task integrity rate, resource utilization rate and cost-efficiency ratio is effectively solved through a two-stage decomposition strategy. In the first stage, a genetic algorithm is adopted to complete satellite-task-time window three-dimensional matching, and optimal distribution of limited visible windows is achieved. In the second stage, a distributed decision-making mechanism is implemented based on an improved Improved-MADDPG framework, optimal dynamic configuration of measurement and control resources is achieved through global information sharing and local strategy iteration, and key calculation indexes such as an average reward value are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of satellite resource allocation, and relates to a method, system and equipment for allocating autonomous measurement and control resources of a heterogeneous satellite cluster. Background Art

[0002] In recent years, the innovative application of heuristic algorithms and the integration of multiple methods have become a research hotspot for satellite mission planning. In the field of satellite cluster collaborative scheduling, researchers have proposed a two-stage optimization strategy based on an improved genetic algorithm. This strategy introduces population perturbations and an elite elimination mechanism during the encoding phase to enhance global search capabilities. Subsequently, a dynamic resource conflict resolution model is established to address satellite resource competition under multiple constraints. This approach reduces computational complexity while achieving the coordinated optimization of task completion rate and resource utilization. With the further development of metaheuristic algorithms, researchers have innovatively constructed a metaheuristic-exact algorithm integrated model (EHE-DCF) within a divide-and-conquer framework. This two-stage optimization strategy addresses the task allocation challenge of large-scale satellite clusters, effectively balancing solution efficiency and solution optimality. In the area of dynamic autonomous decision-making, researchers have proposed a rolling scheduling framework with a hybrid trigger mechanism based on heuristic algorithms to meet the real-time response requirements of agile satellites. By establishing a dynamic task priority assessment model and a resource utilization prediction mechanism, this framework significantly improves the response speed of emergency tasks while ensuring the efficiency of routine tasks. To address the technical difficulties of dynamic measurement and control of heterogeneous satellites, a dynamic collaborative scheduling model based on discrete particle swarm optimization (DPSO) has been proposed. By constructing a multi-dimensional fitness function, it provides a new paradigm for solving the problems of dynamic satellite-ground matching and optimal resource allocation. However, heuristic algorithms also have significant limitations. For example, they may lack adaptability in dynamic heterogeneous environments, are prone to falling into local optimality, and have difficulty effectively handling high-dimensional multi-objective collaborative optimization problems.

[0003] To address these challenges, hybrid intelligent optimization algorithms are emerging as a breakthrough. This paradigm, combining the advantages of data-driven and knowledge-guided approaches, integrates multiple algorithms to complement each other in areas such as model interpretability, dynamic multi-objective trade-offs, and heterogeneous resource collaboration. For example, the proposed Data-Driven Improved Genetic Algorithm (DDIGA) combines artificial neural networks with genetic algorithms. Using artificial neural networks to pre-train the initial population of the genetic algorithm, it also incorporates frequent patterns and a competitive adaptive local optimization strategy, achieving promising results in dynamic and complex scenarios. Others have proposed a hybrid framework that combines deep reinforcement learning with heuristic rules. Deep reinforcement learning (DRL) is used to model the global dynamic decision-making for multi-antenna task allocation, while a heuristic algorithm is used to rapidly generate local optimal solutions for single-antenna task timing. Cross-layer iterative feedback between the two forms a closed-loop decision optimization, enabling efficient resolution of timing conflicts. This demonstrates the effectiveness of hybrid intelligent algorithm fusion in TT&C collaborative scheduling. However, these traditional technologies still struggle to achieve optimal dynamic allocation of autonomous TT&C resources for heterogeneous satellite clusters. Summary of the Invention

[0004] In response to the problems existing in the above-mentioned traditional technologies, the present invention proposes a method for allocating autonomous measurement and control resources of a heterogeneous satellite cluster, a system for allocating autonomous measurement and control resources of a heterogeneous satellite cluster, and a computer device, which can achieve the optimal dynamic configuration of autonomous measurement and control resources of a heterogeneous satellite cluster.

[0005] To achieve the above objectives, the embodiments of the present invention adopt the following technical solutions: On the one hand, a method for allocating autonomous measurement and control resources of a heterogeneous satellite cluster is provided, comprising the steps of: Obtain resource parameters of a heterogeneous satellite cluster; resource parameters include the number of satellites, tasks to be performed, satellite type, and visibility time window; A mixed integer programming model based on resource parameters is called based on the scheduling requirements of autonomous measurement and control resources of heterogeneous satellite clusters; The mixed integer programming model is optimized for the first phase of tasks, satellites, and visible time windows using a genetic algorithm to generate an initial task allocation plan that matches tasks with satellites. Sort the initial task allocation plan according to the order in which the tasks arrive, and generate a complete initial task execution plan for each satellite; The Improved-MADDPG deep reinforcement learning framework based on genetic algorithm is used to guide target optimization with the complete initial mission execution plan as prior knowledge to obtain the optimal task allocation plan for heterogeneous satellite clusters. In the Improved-MADDPG deep reinforcement learning framework, satellites are regarded as intelligent agents and the task execution process is modeled as a Markov decision process. The optimal task allocation plan is the one that achieves the optimal task yield, completion rate and energy consumption under the measurement and control resource constraints of the heterogeneous satellite cluster.

[0006] On the other hand, a heterogeneous satellite cluster autonomous measurement and control resource allocation system is also provided, including: The resource acquisition module is used to obtain the resource parameters of the heterogeneous satellite cluster; the resource parameters include the number of satellites, the tasks to be performed, the satellite type and the visible time window; The model calling module is used to call the mixed integer programming model built based on the resource scheduling requirements of the autonomous measurement and control resources of the heterogeneous satellite cluster according to the resource parameters; The first allocation module is used to jointly optimize the tasks, satellites and visible time windows in the first stage of the mixed integer programming model through a genetic algorithm to generate an initial task allocation plan that matches tasks with satellites; The program sorting module is used to sort the initial task allocation plan according to the order of task arrival and generate the complete task execution initial plan for each satellite; The second allocation module is used to utilize the Improved-MADDPG deep reinforcement learning framework based on genetic algorithm to guide target optimization with the complete initial mission execution plan as prior knowledge to obtain the optimal task allocation plan for the heterogeneous satellite cluster. Among them, within the Improved-MADDPG deep reinforcement learning framework, the satellite is regarded as an intelligent agent and the task execution process is modeled as a Markov decision process. The optimal task allocation plan is the measurement and control resource allocation plan that achieves the optimal task yield, completion rate and energy consumption under the measurement and control resource constraints of the heterogeneous satellite cluster.

[0007] On the other hand, a computer device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned method for allocating autonomous measurement and control resources of a heterogeneous satellite cluster are implemented.

[0008] One of the above technical solutions has the following advantages and beneficial effects: The aforementioned autonomous TT&C resource allocation method, system, and equipment for heterogeneous satellite clusters effectively balances mission completion rate, resource utilization, and cost-effectiveness through a two-stage decomposition strategy. In the first stage, a genetic algorithm (GA) is used to achieve three-dimensional matching of satellite, mission, and visible time window, achieving optimal allocation within the limited visible time window. In the second stage, a distributed decision-making mechanism based on the improved MADDPG framework is implemented. Through global information sharing and local policy iteration, optimal dynamic configuration of TT&C resources is achieved, significantly improving key computational metrics such as operational efficiency, convergence speed, and average reward value. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0010] Figure 1 1 is a flow chart of a method for allocating autonomous measurement and control resources for a heterogeneous satellite cluster in one embodiment; Figure 2 An overlapping diagram of earth observation time windows in one embodiment; Figure 3 Schematic diagram of a hybrid optimization framework based on a genetic algorithm and improved multi-agent deep deterministic policy gradient (GA-Improved-MADDPG) in one embodiment; Figure 4 A schematic diagram of an improved architecture of the MADDPG algorithm in one embodiment; Figure 5 A schematic diagram of the basic framework of a multi-head attention mechanism in one embodiment; Figure 6 A schematic diagram of an independently designed research and experimental verification scenario in an embodiment; Figure 7 This is a diagram of task rewards in one embodiment; Figure 8 An optimal task allocation diagram for a heterogeneous satellite cluster with a task scale of 50 in one embodiment; Figure 9 An optimal task allocation diagram for a heterogeneous satellite cluster with a task scale of 100 in one embodiment; Figure 10 An optimal task allocation diagram for a heterogeneous satellite cluster with a task scale of 200 in one embodiment; Figure 11An optimal task allocation diagram for a heterogeneous satellite cluster with a task scale of 500 in one embodiment; Figure 12 A comparison chart of average rewards between the improved algorithm and the first-come, first-served (FCFS) algorithm in one embodiment; Figure 13 Schematic diagram of the module framework of a heterogeneous satellite cluster autonomous measurement and control resource allocation system in one embodiment. DETAILED DESCRIPTION

[0011] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and Examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the description of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0012] It should be noted that the reference to "embodiment" in this document means that the specific features, structures or characteristics described in conjunction with the embodiment may be included in at least one embodiment of the present invention. The presentation of this phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It will be understood by those skilled in the art that the embodiments described herein may be combined with other embodiments. The term "and / or" used in the specification of the present invention and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.

[0013] The following describes the implementation of the present invention in detail with reference to the accompanying drawings in the embodiments of the present invention.

[0014] In recent years, with the rapid development of "Internet + Satellite" technology, low-orbit satellite constellation projects, represented by Starlink, have sparked a new revolution in the global space industry. These projects have not only achieved breakthroughs in global communications coverage through large-scale satellite deployments, but have also driven the rapid iteration of space technology with significant commercial value. However, as user needs become increasingly diverse and complex, core contradictions facing satellite systems have become increasingly prominent. This core contradiction manifests itself in a structural imbalance between the large-scale deployment of satellite constellations and the dynamic adaptability of measurement and control resources. Existing research suffers from three cognitive flaws: First, inappropriate dimension selection, excessive focus on macro-variances such as satellite type and orbital parameters, while ignoring micro-parameter characteristics such as payload performance baselines, AI model iterations, and onboard algorithm baselines; second, ambiguous modeling methods, simplifying heterogeneous satellites into homogenized entities, resulting in significant errors in identifying capability gradients; third, rigid scheduling mechanisms, relying on manual rules and static scheduling of individual satellites, are subject to drawbacks such as frequent manual intervention, high mission latency, and insufficient resource utilization. These systemic flaws at these three levels collectively lead to a vicious cycle of "heterogeneous homogeneous processing, extensive resource scheduling, and exponentially declining efficiency," severely hindering the constellation system's evolutionary goal of achieving autonomous collaborative reconfiguration. Therefore, improving the autonomous planning capabilities of heterogeneous satellite clusters to resolve this supply-demand contradiction is urgent.

[0015] The autonomous TT&C resource allocation problem for heterogeneous satellite clusters involves dynamically changing space mission scenarios. A cluster system composed of satellites of various types and configurations must leverage the autonomous decision-making capabilities of onboard intelligent agents to achieve the optimal dynamic match between TT&C resources and user mission requirements, while satisfying constraints. Satellite clusters offer distinct advantages over the fragmented scheduling of individual satellites. Satellite clusters possess three key characteristics: personalization, socialization, and evolutionary capabilities. Personalization emphasizes the differences in satellite types and configuration capabilities. Even for the same mission type, satellites exhibit heterogeneity, and homogenous processing often leads to resource waste. Socialization emphasizes collaboration between satellites, achieving efficient coordination through distributed execution and centralized decision-making. Evolutionary capabilities indicate that the number of satellites will continue to increase in the future, requiring the system to be adaptable and scalable. Autonomous TT&C resource allocation for heterogeneous satellite clusters consists of two main phases: time window allocation and mission TT&C. The visibility time window (VTW) refers to the time period during which a satellite is visible to a ground station while orbiting Earth. During this timeframe, ground stations must be able to observe satellites and receive data or transmit commands, a key constraint in satellite tracking and control. When allocating autonomous tracking and control resources within a heterogeneous satellite cluster, the key lies in leveraging the heterogeneity of satellites, ensuring a rational division of labor and cooperation, and allocating limited tracking and control resources appropriately without violating constraints to maximize mission benefits. This requires not only that the system recognize and exploit the heterogeneity between satellites, but also that, guided by collaborative principles, satellites autonomously execute decisions and commands through a combination of distributed execution and centralized decision-making to successfully complete tracking and control missions.

[0016] Satellite tracking and control resource scheduling is a typical NP-hard combinatorial optimization problem. Its complexity stems from the coupling of multiple dynamic and static constraints: it is necessary to coordinate basic constraints such as mission time window conflicts and tracking and control station capacity limitations, and to cope with dynamic challenges such as inter-satellite communication delays, payload compatibility, and insufficient autonomous decision-making capabilities derived from satellite heterogeneity and clustering. It is also necessary to reconcile the inherent contradiction between the efficiency of a single satellite and the global benefits of the cluster. This multi-level constraint system makes it difficult to directly migrate the traditional single-satellite scheduling model to cluster tracking and control scenarios. To address this complex problem, researchers have gradually overcome technical bottlenecks through innovative mathematical programming algorithms. In the field of collaborative scheduling of heterogeneous satellites, a hybrid algorithm based on branch-and-bound and heuristic pruning was proposed, constructing a multi-agent collaborative model to achieve efficient mission planning for heterogeneous satellites. This collaborative mechanism, combining distributed decision-making with global optimization, significantly improved overall system performance. To address the surge in communication and decision-making costs caused by the expansion of satellite clusters, network flow theory was innovatively introduced into the field of autonomous decision-making, proposing a minimum-cost flow dynamic optimization method based on a core network. This method achieves low-latency autonomous scheduling through real-time task supply and demand matching and a local update mechanism. Furthermore, to address the prominent contradiction between individual and group interests, a two-stage dynamic allocation framework was designed. Based on the minimum conflict set strategy, a multi-objective optimization model was established to achieve a dynamic balance between minimizing scheduling perturbations and maximizing observation benefits during the task insertion phase. Although current research trends indicate that the deep integration of precise algorithms and heuristic strategies, as well as the parallel innovation of multi-agent collaboration and network flow architectures, are providing systematic solutions to the "dynamic multi-constraint optimization" dilemma of satellite cluster measurement and control, significant shortcomings remain. For example, due to the inherently high computational complexity of the problem, the computational time of traditional mathematical programming algorithms (such as branch-and-bound and integer programming) increases exponentially, making it difficult to meet the real-time requirements of dynamic scenarios. In heterogeneous satellite clusters, the introduction of collaborative constraints further exacerbates the problem of model dimensionality explosion. Secondly, faced with the conflict between model accuracy and the cost of simplification, mathematical programming methods often approximate some complex constraints, making overly idealized model assumptions and easily causing deviations between the model and the actual system. These shortcomings essentially reflect the inherent bottlenecks of mathematical programming methods in handling "dynamic, distributed, and multi-objective" coupled optimization problems, and urgently need to be integrated with new paradigms such as reinforcement learning and metaheuristics to break through theoretical boundaries.

[0017] Artificial intelligence-based optimization algorithms offer a new path to overcoming these bottlenecks. In the area of autonomous decision-making, a distributed scheduling method based on multi-agent deep reinforcement learning (MADRL) has been proposed. This method, using a multi-agent deep deterministic policy gradient (MADDPG) network, enables satellites to share decision-making policies, rather than specific states or decision data, significantly reducing communication overhead and improving response speed. Compared to traditional contract network protocols (CNPs), this method reduces communication load while optimizing satellite resource utilization, offering enhanced real-time and robustness. Building on this foundation, a framework for autonomous decision-making based on deep neural networks (DNNs) has been proposed. Its core idea is to replace manual rule design with data-driven approaches, leveraging the autonomous learning capabilities of neural networks to directly link task requirements with resource allocation strategies, achieving end-to-end autonomous optimization. This algorithm meets the real-time dynamic scheduling requirements of large-scale satellite constellations while reducing reliance on manual rules and demonstrating strong adaptability. It provides an efficient and reliable technical path for autonomous collaborative decision-making in complex scenarios. To address the challenges of integrating heterogeneous measurement and control systems and optimizing resources, the problem has been modeled as a Markov decision process (MDP). The DQN algorithm, which combines deep neural networks with Q-learning, significantly optimizes resource utilization efficiency and alleviates resource conflicts in the coordinated scheduling of state-owned and commercial TT&C resources. Currently, with the exponential growth of low-Earth orbit constellations, a proposed deep reinforcement learning distributed scheduling algorithm has demonstrated significant advantages. This algorithm constructs a competitive-cooperative model and, by dynamically sensing the operational status of satellites, generates a TT&C link allocation scheme that satisfies spatiotemporal constraints. This improves response efficiency while adaptively adjusting resource allocation weights, providing a new technical paradigm for TT&C management of large-scale constellations. While traditional supervised learning algorithms have demonstrated application value in specific areas, their adaptability to the dynamic operational environment of satellites remains significantly limited. They suffer from three major drawbacks: First, model training relies on massive amounts of historical data. The unpredictability of real-time on-orbit status can lead to a distributional shift between training and measured data, which in turn risks reduced model decision reliability. Second, the algorithm's high computational overhead conflicts with limited onboard computing power, which can lead to mission delays. Third, the black-box nature of deep networks impairs decision interpretability, posing a safety hazard for space missions requiring precise fault tracing.

[0018] In one embodiment, Figure 1 As shown, a method for allocating autonomous measurement and control resources of a heterogeneous satellite cluster is provided, which may include the following steps S10 to S18: S10, obtaining resource parameters of the heterogeneous satellite cluster; the resource parameters include the number of satellites, tasks to be executed, satellite types, and visible time windows; S12, calling a mixed integer programming model based on the resource scheduling requirements of the autonomous measurement and control resources of the heterogeneous satellite cluster according to the resource parameters; S14, performing a first-stage joint optimization of tasks, satellites, and visible time windows on the mixed integer programming model using a genetic algorithm to generate an initial task allocation plan that matches tasks with satellites; S16, sorting the initial task allocation plan according to the order of task arrival, and generating a complete task execution initial plan for each satellite; S18 uses the Improved-MADDPG deep reinforcement learning framework improved by the genetic algorithm, and uses the complete initial mission execution plan as prior knowledge to guide target optimization and obtain the optimal task allocation plan for the heterogeneous satellite cluster. Among them, within the Improved-MADDPG deep reinforcement learning framework, the satellite is regarded as an intelligent agent and the task execution process is modeled as a Markov decision process. The optimal task allocation plan is the measurement and control resource allocation plan that achieves the optimal task yield, completion rate and energy consumption under the measurement and control resource constraints of the heterogeneous satellite cluster.

[0019] It can be understood that the Heterogeneous Satellite Cluster Measurement & Control Resource Allocation Problem (HSCMC-RAP) addressed in this specification is described as a cluster system consisting of satellite nodes with homogeneous platforms but heterogeneous capabilities. In dynamic mission scenarios, for the application requirements of single-point observation targets (which can be completed with a single measurement and control), it is necessary to rely on the autonomous collaborative decision-making capabilities of satellites to complete the dynamic adaptation and optimal configuration of measurement and control resources under the conditions of meeting multi-dimensional constraints, so as to achieve efficient response to user tasks. The system description is as follows: Figure 2 Therefore, the core challenge of this problem lies in building a target optimization mechanism for single-shot measurement and control scenarios. That is, under the conditions of satisfying multi-dimensional constraints such as inter-satellite communication topology, energy supply threshold, and payload performance boundary, limited time window resources are allocated to achieve distributed collaborative decision-making and achieve the optimal balance between mission success rate, resource allocation efficiency, and energy consumption economy.

[0020] The problem is modeled as a two-stage dynamic resource allocation problem. TT&C resource allocation is decomposed into two phases: time window allocation and mission TT&C. These phases address the dynamic matching of visible time windows (VTWs) and the optimal allocation of TT&C resources, respectively. A hierarchical optimization strategy enhances the system's adaptability to emergent missions and high-latency scenarios, improving the satellite's autonomous decision-making capabilities. Furthermore, in the modeling of heterogeneous satellite clusters, the capability differences between satellites of the same type are carefully considered. Quantitative analysis of satellite capability gradients effectively avoids the decision-making blind spot of "homogenizing diverse scenarios."

[0021] A hybrid measurement and control framework was also constructed, combining the collaborative optimization of a genetic algorithm and an improved multi-agent deep deterministic policy gradient (MADDPG) algorithm. This framework employs a two-stage optimization mechanism: in the first stage, the global search capability of the genetic algorithm is used to pre-optimize the initial policy space, generating a high-quality initial policy distribution for mission-satellite matching. In the second stage, multi-agent collaborative optimization is carried out based on the improved MADDPG algorithm, achieving optimization objectives such as mission yield, completion rate, and energy consumption while satisfying measurement and control resource constraints. This collaborative optimization mechanism effectively overcomes the inefficient exploration of traditional reinforcement learning in very large-scale state spaces, achieving the complementary advantages of data-driven and knowledge-guided approaches.

[0022] Specifically, a systematic improvement scheme was proposed to address the limitations of the MADDPG algorithm, such as slow convergence and insufficient training stability. At the training optimization level, an efficient and stable training framework was constructed by integrating the Fixup parameter initialization method with the Adax adaptive optimizer. The Fixup initialization strategy abandons traditional batch normalization and adopts a layered weight scaling mechanism to mitigate model gradient vanishing or convergence issues. The Adax optimizer dynamically adjusts the parameter update step size to effectively improve computational efficiency while maintaining training accuracy. Combining the advantages of attention mechanisms and residual learning, an Attention-Enhanced Residual Block (AERB) was introduced. This AERB implements dynamic feature filtering through a multi-head self-attention mechanism, enabling the network to focus on key information domains of the input data. Incorporating cross-layer skip connections, a deep feature propagation path was constructed, improving model representation capabilities while effectively mitigating the vanishing gradient problem in deep networks through a gradient direct connection mechanism. The MADDPG algorithm is designed to handle continuous state and action spaces, whereas the tracking and control scheduling problem for heterogeneous satellite clusters is a complex integer programming problem with discrete decision and solution spaces. To meet practical application needs and address the problem of continuous space discretization, this implementation, based on the preprocessing principles of probability extreme value theory and utilizing the Gumbel-Softmax parameterization technique, superimposes Gumbel noise on the discrete action logits output by the policy network and performs temperature annealing. This achieves a continuous relaxation representation of discrete decisions, effectively addressing the action space mismatch problem of the basic MADDPG algorithm in discrete combinatorial optimization problems. Furthermore, a moving average algorithm is employed to smooth data fluctuations and improve the stability of convergence results.

[0023] At the system architecture level, a hybrid "distributed execution-centralized learning" architecture is proposed. Satellites use an onboard agent to achieve autonomous decision-making (distributed execution), and then optimize global returns through a centralized policy network. Simultaneously, the problem is modeled as a sequential decision-making problem using a Markov process, effectively overcoming the drawback of the traditional MADDPG algorithm, which is prone to falling into local optimal solutions during policy search.

[0024] This method for allocating autonomous TT&C resources for a heterogeneous satellite cluster effectively balances mission completion rate, resource utilization, and cost-effectiveness through a two-stage decomposition strategy. In the first stage, a genetic algorithm (GA) is used to construct a fitness function centered on cost-effectiveness, achieving a three-dimensional matching of satellite, mission, and visible time window, thereby optimizing the allocation of limited visible windows. In the second stage, a distributed decision-making mechanism is implemented based on the improved MADDPG framework. Through global information sharing and local policy iteration, optimal dynamic configuration of TT&C resources is achieved. Compared to the baseline MADDPG algorithm, the improved algorithm demonstrates significant advantages in key metrics such as operational efficiency, convergence speed, and average reward value.

[0025] First, in terms of mathematical models, during the autonomous measurement and control resource allocation process of heterogeneous satellite clusters, their working environment changes continuously over time, and the modeling process is extremely complex. In order to transform this practical problem into a mathematical model for easy analysis, the following assumptions are made: (1) The orbital dynamics environment exhibits piecewise stationary characteristics, and its dynamic changes can be modeled by discrete time series; (2) The state of the intersatellite link obeys a Markov random process, and its state transition probability matrix can be obtained a priori; (3) Each satellite node has the ability to perceive its own state and can obtain the working status of its own payload and the remaining energy information in real time; (4) The task is a single-point task, that is, each task must maintain the continuity of execution and cannot be reassigned, and the execution process is not allowed to be interrupted; (5) Satellite node resources have limited and non-renewable characteristics, and the storage capacity and energy reserves are dynamically consumed as the task is executed and cannot be replenished midway; (6) The failure rate between satellites and measurement and control equipment is ignored; (7) Satellite capabilities are heterogeneous, which is manifested in the differences in the task execution cost matrix and is also subject to the constraints of single-task processing (each satellite can only process a single task per unit time period); (8) The satellite task switching process has a resource reconfiguration phase with a fixed time cost, and its duration is independent of the task type.

[0026] This hypothesis system characterizes the link dynamics through Markov processes and combines discrete-time modeling to handle the time-varying nature of the environment. While ensuring the solvability of the model, it effectively retains the main characteristics of the satellite resource allocation problem. Among them, assumptions (3)-(5) constitute the perception basis for autonomous decision-making, and assumptions (6)-(8) define the rigid constraints of the system, together constructing a theoretical analysis framework with engineering guidance value.

[0027] Based on the above assumptions, the mathematical model is given below. First, the necessary parameters are explained: Satellite, with S express, ,in, M Represents the total number of satellites on mission; ,in, s i It consists of a sextet; Indicates the satellite identifier, Represents a collection of heterogeneous capability satellites, whose satellite type belongs to one of the three heterogeneous capability satellites A, B and C. ABC belong to satellites of the same type but with different capabilities. The satellite capabilities are assumed to decrease in the order of C, B and A. For example, when performing the same mission The memory consumed will be few; represents a set of visible time windows; N Represents the total number of tasks that need to be executed in a cycle, Indicates the Satellites for Task No. k Visible time window; Represents the set of all visible time windows of a task; Indicates whether the satellite is available in the current period, that is, whether the satellite is performing other tasks in the current period. ; Indicates the memory of the satellite; Indicates the remaining satellite observation time.

[0028] Task, use T express, ,in N Represents the total number of tasks that need to be executed in a cycle. ,in It consists of a sextet; An identifier representing the task; Indicates the location where the task occurs, consisting of the latitude and longitude where the task occurs; Represents the time attributes of a task, which consists of the start, end, and duration of the task; Represents the benefits of the task; Indicates the satellite storage consumed by the mission, indicating the amount of power consumed by A, B, and C type satellites when observing respectively; Indicates which satellite is performing the mission.

[0029] Therefore, to meet the needs of autonomous TT&C resource scheduling for heterogeneous satellite clusters, this specification constructs a mixed integer programming model that integrates an autonomous collaborative optimization mechanism. This model achieves mathematical abstraction of the spatiotemporal configuration of satellite TT&C resources through the design of multi-dimensional decision variables and constraint systems. Its core architecture can be described as follows: The decision variables are , Indicates the Satellites perform j The effective time of a task, that is, the effective time of task execution, The time when the task starts to execute effectively, is the end time of effective execution of the task. The complete optimization mathematical model (i.e., mixed integer programming model) is: (1) (2) (3) (4) (5) (6) (7) Among them, formula (1) is the target optimization function, which is a hybrid optimization model that pursues the dual requirements of maximizing benefits and minimizing costs, while taking into account the task completion rate. In order to facilitate the solution, the weight factor is introduced to perform a linear weighted combination of the two. Represents the cost-effectiveness of the task, thus taking into account multiple constraints, formula (2) It is a specific calculation method of the weight factor. By using the adaptive adjustment mechanism based on the Sigmoid function to smooth the weight function, it can effectively avoid the influence of parameter mutation on system stability. Furthermore, by introducing the threshold parameter c By standardizing the Sigmoid function, we can precisely adjust the sensitivity of the weight to the target state change, thereby establishing a nonlinear mapping relationship between task completion rate and cost consumption, and achieving dynamic balance control in the collaborative optimization process. is a decision variable, representing Whether the execution is successful Formula (4) means that each task can only be executed once at most. Formula (5) means that each satellite can only execute one task at most at the same time. Formula (6) means that for the same satellite, the switching time between two adjacent tasks must be greater than or equal to , is a constant. Formula (7) shows that the time it takes for a satellite to perform a certain task is greater than the effective time of the task and the conversion time of the equipment. express.

[0030] Genetic algorithms, a classic metaheuristic optimization method, demonstrate significant advantages in solving NP-hard combinatorial optimization problems, such as satellite tracking and control resource allocation, thanks to their swarm intelligence search mechanism and global optimization capabilities. This paper innovatively combines genetic algorithms with the Improved-MADDPG deep reinforcement learning framework to construct an efficient hybrid optimization paradigm. This approach uses genetic algorithms to pre-explore the problem space and generate high-quality initial solution distributions, providing prior knowledge guidance for subsequent reinforcement learning agents. This effectively reduces the dimensionality of the policy search and accelerates algorithm convergence.

[0031] The modeling process is as follows: To address the problem of autonomous measurement and control resource allocation for heterogeneous satellite clusters, this manual proposes the GA-IMADDP algorithm, which divides the task allocation process into two parts. The first part is the allocation of tasks to satellites, which is solved by the GA algorithm; the other part is whether the task is measurement and control, that is, how to use limited measurement and control resources to maximize the mission benefits while minimizing costs and balancing the task completion rate.

[0032] The first step is to build a GA algorithm. The GA algorithm is used to implement task allocation preprocessing, which is mainly divided into several parts: population initialization, encoding method, fitness function, selection operation, crossover operation and mutation operation.

[0033] In one embodiment, population initialization is a basic step of the genetic algorithm, and its quality directly affects the convergence efficiency of the algorithm. This embodiment uses the global random generation method to construct the initial population to ensure the diversity and coverage of the population. Specifically, for each task , first from its available satellite set Uniformly randomly select satellite numbers from , then the visible time window of the selected satellite is set The time window index is randomly selected in [ ]. This process is repeated until an initial population of the desired number of individuals is generated. This dual randomization mechanism ensures spatial coverage of the initial solution while implicitly satisfying the hard constraints on satellite selection and time window.

[0034] In this embodiment, the encoding method uses a two-dimensional encoding strategy based on integers to construct a solution space representation. N The problem of allocating observation tasks is that the chromosome structure of each individual is N For the first A task, whose genetic encoding form is defined as an ordered pair , among which , ; , ; , ; i Indicates the selected satellite number. Indicates the Satellites for Task No. k visible time windows. Therefore, the complete chromosome is represented as , its essence is to realize the joint encoding of task-resource-visible time window through two-dimensional discrete space mapping.

[0035] In this embodiment, the fitness function of the genetic algorithm takes minimizing memory consumption as the core goal and considers the problem of heterogeneous satellite capabilities. For the same set of measurement and control tasks, the cost paid by the satellites is , the three satellite capabilities belong to heterogeneous capabilities. Therefore, the specific form of the fitness function is: (8) in, Indicates the memory consumption of the mission on the satellite, is the time window conflict penalty. Expressed as a penalty coefficient, it controls the time window conflict penalty term, affecting its contribution to the objective function. Two core constraints must be met when determining feasibility: the mission observation duration must be within the visible time window of the satellite and the ground station. Second, the mission time windows of the same satellite must not overlap. The infinite penalty term effectively eliminates illegal solutions and guides the search towards optimization within the feasible region.

[0036] In this embodiment, the selection operation uses a tournament selection mechanism to balance selection pressure and population diversity. The specific process is as follows: appropriate individuals are randomly selected from the population to form a tournament group, which is then sorted in ascending order of fitness (prioritizing individuals with lower memory usage). The top two best individuals are selected as parents for genetic operations. This process is repeated until the required mating pool size is met. This strategy retains high-quality individuals while maintaining population diversity through random sampling, avoiding premature convergence.

[0037] In this embodiment, the crossover operation promotes the recombination of gene segments by designing a single-point crossover operator. Specifically, a crossover point is randomly generated. , the parent chromosome and Exchange the gene segments after point d to generate two offspring individuals: (9) (10) The crossover operation explores new combination patterns of task assignment sequences by exchanging gene segments.

[0038] In this embodiment, the mutation operation enhances the local search capability by introducing a directed mutation strategy. , randomly select a new satellite and reselect the satisfied window index from its corresponding time window set. The mutation process adopts a constraint-aware mechanism to ensure that the newly generated solution is always within the feasible region, thereby improving the search efficiency and preventing the algorithm from converging to a local optimal solution too early during the search process.

[0039] Therefore, the specific operation process of the GA algorithm is as follows Figure 3 As shown in the figure, the genetic algorithm proposed in the first phase constructs a feasible initial task allocation scheme through the joint optimization of mission, satellite, and visible time window. This scheme can provide prior knowledge guidance for the reinforcement learning algorithm based on improved multi-agent deep deterministic policy gradient (Improved-MADDPG) in the second phase, effectively alleviating the dimensionality curse problem faced by reinforcement learning in the exploration of large action spaces, improving the real-time allocation and autonomous decision-making capabilities of satellites, and providing a foundation for reinforcement learning modeling.

[0040] Then comes the modeling of the Improved-MADDPG algorithm. The initial satellite-task assignment scheme obtained from the GA algorithm is sorted in the order of task arrival, and all task plans that need to be executed on each satellite can be obtained. Then the Improved-MADDPG algorithm is used to maximize the mission benefits under limited measurement and control resources, while taking into account the mission completion rate and cost consumption. The MADDPG algorithm is a typical multi-agent reinforcement learning. With its centralized decision-making and distributed execution advantages, it is widely used in satellite measurement and control resource allocation problems. The problem of autonomous measurement and control resource allocation for heterogeneous satellite clusters is a typical combinatorial optimization problem. Its process can be modeled as a Markov decision process. The measurement and control resources that the satellite can choose are only related to the current state and have nothing to do with its past schedulable state. This process can be established as a four-tuple , representing state, action, state transition probability and reward function respectively.

[0041] For the initial satellite-task assignment scheme obtained using a genetic algorithm (GA), the matching results are sorted according to the task arrival sequence through a time-series scheduling strategy, thereby generating a complete initial mission execution plan for each satellite. On this basis, this embodiment proposes an Improved-MADDPG algorithm, which aims to achieve target optimization under the condition of limited measurement and control resources: while maximizing task benefits, it also takes into account both task completion rate improvement and measurement and control cost control. As a typical paradigm of multi-agent reinforcement learning (MARL), MADDPG has demonstrated significant application value in the field of satellite measurement and control resource allocation due to its "centralized training-distributed execution" architectural advantages. The autonomous measurement and control resource allocation problem for heterogeneous satellite clusters studied in this paper has the following characteristics: First, it is an NP-hard combinatorial optimization problem; second, its dynamic decision-making process satisfies the Markov property, that is, the satellite's measurement and control resource selection depends only on the current state and is independent of the historical state. To this end, this paper regards the satellite as an intelligent agent and models the task execution process as a Markov decision process (MDP), formally defined as a four-tuple < S , A ,P,R>, where: S Represents the state space, characterizing the system environment and satellite status; A is the action space, describing whether measurement and control are performed; P is the state transition probability function, describing the law of environmental evolution; R is the reward function, which is used to quantitatively evaluate the comprehensive benefits of the resource allocation plan.

[0042] The state space is defined as the agent's abstract representation of the system environment and satellite operating status, which contains multi-dimensional normalized feature vectors, specifically divided into satellite state and mission state. The state of each satellite agent is composed of a three-dimensional observation vector. ,in Represents the ratio of the remaining time of each satellite to the maximum observation time, that is, the normalized remaining time; is the normalization of the remaining storage, Indicates whether the satellite is performing other tasks in the current period. The state of each task is composed of a four-dimensional observation vector. , which represent which satellite executes each mission, mission duration, consumed storage and reward of the mission respectively.

[0043] The action space represents the agent's resource allocation decision for the task, which is designed as a binary discrete space , corresponding to the "reject" and "accept" tasks respectively. , the satellite will consume time and storage to perform the task and update its time window status; if the action , then the resource state remains unchanged and the action selection must satisfy the Markov property.

[0044] State transition function. State transition function Describes the process of environmental evolution, whose determinism is dominated by the task execution logic and resource consumption rules. Specifically, if you accept the task , the satellite will and Deductions and , and set the task time interval Add to the satellite time window queue, and then perform subsequent task allocation conflict detection. The resource status and time window remain unchanged, and the task feature vector is reset to the next task to be assigned. If the task is rejected , the resource status and time window remain unchanged, and the task feature vector is reset to the attributes of the next task to be assigned.

[0045] The reward function is used to quantify the overall benefits of the resource allocation plan and is a specific representation of the objective function. The reward function, also known as the single-step reward or immediate reward, is an important basis for guiding the learning direction of the algorithm. The reward function is as follows: (11) The reward function is composed of the current state The action taken by the agent The feedback formed with the environment, namely the product of the state transition probability and the task benefit, guides the improvement direction of the task.

[0046] Therefore, a hybrid optimization framework based on a genetic algorithm and an improved multi-agent deep deterministic policy gradient (GA-Improved-MADDPG) is proposed to solve the dynamic allocation of autonomous measurement and control resources for heterogeneous satellite clusters. This study employs a phased optimization framework: in the first phase, a global search strategy using a genetic algorithm pre-screens the matching relationships between missions, heterogeneous satellite capabilities, and time windows under multi-dimensional constraints, establishing a feasible solution space. In the second phase, an optimization model with improved mechanisms is constructed based on the MADDPG multi-agent reinforcement learning framework. Specifically, a series of algorithmic improvements enhance the stability and convergence speed of reinforcement learning, such as the use of a fix-up initialization method to strengthen network stability and the design of a parameter update mechanism based on the Adax optimizer to improve the training efficiency of residual neural networks based on the attention mechanism. This hybrid algorithm combines the global optimization properties of a genetic algorithm with the dynamic decision-making capabilities of deep reinforcement learning to form a hierarchical hybrid intelligent optimization architecture, providing a systematic solution to the constellation resource allocation problem.

[0047] Figure 3The figure shows the hybrid optimization framework of GA-Improved-MADDPG. The system workflow is divided into two stages: first, a set of user mission requirements within a unit period is collected. Combined with schedulable measurement and control resources, a set of time window resources subject to multidimensional constraints is generated through spatiotemporal constraint analysis. A genetic algorithm, acting as a preprocessing module, constructs a fitness function aimed at minimizing the overall cost. It then performs a global search and iterates the elite retention strategy within the solution space, outputting a matching relationship between mission, satellites with heterogeneous capabilities, and time window, which serves as the initialization scheme for the Improved-MADDPG algorithm.

[0048] During the reinforcement learning optimization phase, the classic MADDPG architecture was improved, employing a centralized training and distributed execution (CTDE) architecture to construct a heterogeneous agent collaborative decision-making model, in which each satellite is equipped with an independent actor network for autonomous decision-making. First, to adapt to the discrete decision-making characteristics of measurement and control resource allocation, a Gumbel-Softmax-based reparameterization method was utilized to probabilistically sample and discretize the continuous action space, effectively addressing the compatibility issue between the policy network output and the discrete action space. Second, global state information was shared through a critic network, and an exponential moving average technique was introduced to smooth the immediate reward signal, suppressing the interference of environmental feedback noise on policy updates and significantly improving the algorithm's convergence stability. At the model architecture level, a residual neural network structure based on an attention enhancement mechanism was introduced, combined with a fix-up initialization method and the Adax optimizer to improve the network's convergence speed.

[0049] Among them, the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm is a reinforcement learning algorithm for multi-agent systems proposed in 2017. This algorithm is based on the Actor-Critic (AC) framework, and its core design concept lies in the synergy of centralized training and decentralized execution. Specifically, each agent independently maintains an actor network with autonomous decision-making capabilities. A centralized critic network is introduced during the training phase, which integrates global state information and the joint action space to evaluate the state-action value function. This architectural design effectively addresses the environmental non-stationarity problem inherent in multi-agent systems. The global value function guides the policy optimization direction of each agent, maintaining decision-making autonomy while ensuring system-wide synergy. Its specific implementation mechanism can be summarized as follows: Each agent, guided by the global Q-value provided by the critic network, concurrently updates the actor network parameters using a policy gradient method. Simultaneously, the critic network optimizes its parameters using a temporal difference method. Ultimately, through alternating iterations, policy optimization is achieved progressively, as follows: Policy function. It is defined as the mapping from state to behavior, using The essence is that in a given state s Behavior set in the case The probability distribution function of is expressed as follows: (12) in, is the probability distribution function. This algorithm determines the action strategy. For deterministic strategies, this function degenerates into a mapping: (13) A policy is uniquely determined by actions.

[0050] As a core element of reinforcement learning theory, the expected cumulative return function is designed to overcome the limitations of the immediate reward function. The immediate reward function only represents the agent’s state. s Execute an action a Transfer to state s This method is inherently similar to the local optimal selection mechanism of the greedy algorithm, which can easily lead to policy optimization falling into a local convergence state. In view of this, the reinforcement learning framework introduces a time dimension expansion mechanism, by defining the discounted cumulative return: (14) in, is the current agent, , Represents the expectation of the joint distribution of state and action; is a discount factor that weighs the importance of current rewards against future rewards. The closer it is to 0, the more emphasis is placed on recent rewards, and vice versa.

[0051] Value function. The cumulative expected return function can be mathematically formalized as a state value function and a state-action value function: (15) (16) Where (15) is the state value function, (16) is the state-action value function. (Bellman) formula for temporal recursive modeling: (17) (18) in Indicates that from Departure arrive The state transition probability.

[0052] Strategy update. This mechanism is the core module of the algorithm, and its implementation process follows the distributed strategy optimization paradigm. The algorithm builds a strategy actor network for each satellite agent, and the design parameters are , , M is the number of satellites. The strategy is updated using the stochastic gradient descent method: (19) is the standard expected discounted return, is the logarithmic gradient of the strategy, reflecting the sensitivity of action selection to parameters. D Experience Replay Buffer is an experience replay buffer that stores trajectory data when it reaches the preset capacity threshold. , from which importance sampling is performed to obtain batch data: . is the state-action value function, which is implemented by minimizing the temporal difference error: (20) in, is the target network parameter, It is an experience sampling trajectory, and a soft update mechanism is used to maintain training stability. It effectively solves the coupling oscillation problem between policy evaluation and policy improvement. This method updates the parameters of the target action network and the target policy network as follows: (twenty one) (twenty two) Sampling soft update method, is a hyperparameter, usually taken as ( , such as 0.001). This mechanism ensures a smooth transition of target value estimation by weighted averaging of online network and target network parameters. It is generally a constant. By using the soft update method, the target network parameters can be changed less, thus making the algorithm converge stably. Experimental studies have shown that when When the value approaches 0, the target network is approximately frozen, which can be regarded as a special form of delayed update; When it approaches 1, it degenerates into real-time synchronous update of the online network.

[0053] To address the limitations of the basic MADDPG algorithm, this study proposes an improved multi-agent reinforcement learning framework, Improved-MADDPG. Its core improvement strategies include the following aspects: Taking into account the particularity of multi-satellite measurement and control scenarios, and considering that under the constraints of limited measurement and control resources, each satellite acts as an independent decision-making entity, and its resource selection strategy will form a mutually influential relationship.

[0054] At the model architecture level, a residual neural network structure enhanced by the attention mechanism is proposed, and the Fix-up parameter initialization method and Adax adaptive optimizer are used to accelerate the model convergence process.

[0055] In view of the discrete nature of the action space in measurement and control resource allocation tasks, a discrete-to-continuous hybrid action space mapping mechanism is constructed using the Gumbel-Softmax reparameterization technique. The discretization of the policy network output is achieved through probabilistic sampling, effectively resolving the compatibility issues of traditional continuous action strategies in discrete decision-making scenarios.

[0056] The following are some of the innovations mentioned above: (1) Optimizing the network architecture and training mechanism of the MADDPG algorithm, the algorithm systematically improves the MADDPG algorithm from three dimensions: neural network architecture design, parameter initialization method, and optimizer selection. Figure 4 As shown in the figure, the Improved-MADDPG deep reinforcement learning framework adopts a residual neural network structure enhanced by the attention mechanism, and accelerates the model convergence through the Fix-up parameter initialization method and the Adax adaptive optimizer.

[0057] Specifically, the residual network architecture is designed based on attention enhancement. The traditional MADDPG algorithm uses a two-layer fully connected neural network as the basic architecture of the policy network. This design has two significant defects: first, the parameter space of the fully connected layer grows squared with the input dimension, and the computational efficiency is significantly reduced when processing high-dimensional state space, which is not suitable for the spatial state of the heterogeneous satellite measurement and control cluster; second, the fixed-structure fully connected layer is difficult to effectively capture the dynamic interaction relationship between multiple agents, especially in partially observable environments, which is prone to policy degradation problems. To this end, this paper proposes the introduction of an attention-enhanced residual network (ARN), whose basic framework is as follows: Figure 5As shown in Figure 2, its advantages are reflected in the following two aspects. First, a dynamic feature extraction module is constructed through a multi-head attention mechanism, enabling each agent to adaptively adjust its attention weight to the observation information of other agents based on the current environment state. In partially observable collaborative task scenarios, this helps improve the efficiency of capturing key state features. Second, a deep residual structure is adopted to replace the traditional fully connected layer, and skip connections are used to effectively alleviate the vanishing gradient problem in deep network training.

[0058] Application of the Fix-up Initialization Method. Neural network parameter initialization has a significant impact on the model's convergence speed and generalization performance. Addressing the problem that traditional initialization methods (such as random initialization and Xavier initialization) are prone to causing gradient anomalies in deep networks, the Fix-up initialization technique proposed in 2019 pioneered a new solution. Through a unique parameter scaling mechanism, this technique successfully achieved two major breakthroughs without relying on batch normalization: First, it ensures that the output variance of each residual module remains stable during the initialization phase, effectively avoiding gradient explosion and vanishing gradient phenomena; second, by dynamically adjusting the weight scale of the residual branch, the convergence efficiency of deep networks is significantly improved. This characteristic, which is naturally compatible with the residual network architecture, allows Fix-up to maintain the advantage of network depth while providing a theoretical basis for eliminating the traditional initialization method's strong reliance on batch normalization.

[0059] Adaptive improvements to the Adax optimizer. Addressing the common non-stationary optimization challenges in deep reinforcement learning, the Adax (Adaptive Gradient Descent with Exponential Long Term Memory) optimizer offers a systematic solution to the limitations of the mainstream Adam optimizer used in most neural networks today by innovatively incorporating a second-order moment estimation reconstruction mechanism. Proposed in 2020, this algorithm's core breakthrough lies in the introduction of an inverse temporal weighting strategy with an exponential decay factor. This strategy uses the exponential moving average (EMA) of the accumulated historical gradients to establish a second-order momentum estimation model with long-term memory. This design effectively overcomes the momentum estimation bias of the Adam optimizer when the gradient exhibits exponential decay, mathematically avoiding the risk of non-convergence while maintaining the fast responsiveness of the adaptive learning rate. It is suitable for solving typical multi-objective scenarios, such as optimizing coordinated tracking and control resources for satellite clusters. This algorithm addresses the heterogeneity of gradient distributions across different nodes due to hardware performance differences and diverse mission types. Adaptive multi-objective compatibility is achieved through a unified gradient accumulation mechanism, eliminating the need for hyperparameter tuning for individual nodes. This "one-to-many" optimization feature reduces the complexity of parameter adjustment in multi-target collaborative scenarios from Dimensionality reduction to , significantly improving the deployment efficiency of large-scale distributed systems.

[0060] In one embodiment, a discrete-continuous hybrid action space mapping mechanism is constructed using the Gumbel-Softmax function in the Improved-MADDPG deep reinforcement learning framework.

[0061] Specifically, regarding the feature design improvements for the algorithm's input and output: Addressing the inherent limitation of the original MADDPG algorithm, which is that discrete action spaces are non-differentiable, the improved algorithm utilizes the Gumbel-Softmax function to effectively address the adaptability challenges of the continuous policy gradient algorithm in discrete action spaces. This method uses a reparameterization technique to inject Gumbel noise into the continuous actions output by the policy network and leverages the characteristics of extreme value distributions to construct a differentiable discrete action sampling process. To balance exploration efficiency and policy stability, a hybrid action generation strategy is designed, embedding the ε-greedy stochastic exploration mechanism into the argmax-based deterministic action selection process to achieve a hard mapping of discrete actions. To address the gradient breakage problem caused by discrete operations, a gradient pass-through estimator is used to maintain the continuity of parameter updates during backpropagation, allowing the non-differentiable operations of discrete actions to be encapsulated as differentiable computational graph nodes.

[0062] In one embodiment, the immediate reward is dynamically smoothed by constructing an exponential moving average function in the Improved-MADDPG deep reinforcement learning framework.

[0063] As can be understood, the improved algorithm addresses the convergence oscillation problem caused by high variance in reward signals in multi-agent allocation by proposing a reward correction mechanism based on time-series decay. By constructing an exponential moving average (EMA) function to dynamically smooth immediate rewards, it gives higher weight to recent rewards. This suppresses long-term fluctuations while maintaining sensitivity to recent state changes. This makes it suitable for real-time streaming reward processing, eliminating the need to wait for a fixed window period to complete and adapting to dynamic changes in the environment.

[0064] In some embodiments, a series of simulation experiments are set up to verify the effectiveness of the Improved-MADDPG algorithm. The experimental environment is a 13th Gen Intel(R) Core(TM) i9-13900HX 2.20 GHz, and the operating environment is Python 3.11.7. First, multiple experimental scenarios are designed and corresponding parameter values are set in the Improved-MADDPG algorithm. Then, the Improved-MADDPG algorithm is compared with the basic MADDPG algorithm in terms of profit margin, running time, and convergence. Finally, the convergence of the algorithm and its time efficiency are analyzed and discussed in detail.

[0065] Due to the lack of standard test cases for the Heterogeneous Satellite Cluster Measurement and Control Resource Allocation Problem (HSCMC-RAP), this study independently designed experimental verification scenarios, such as Figure 6 As shown. The experimental design uses real satellite orbital parameters to build a simulation environment. The satellite data source is the publicly released Starlink satellite TLE data. The satellite dynamic characteristics are obtained by analyzing its orbital parameters (TLE). In order to ensure the orbital heterogeneity of the satellite cluster, a sampling method is used to select 10 satellites operating in different orbital planes from the Starlink constellation, approximately 15 to 16 orbits, to establish an experimental satellite group numbered 1-10. The time is set from 3 Dec 2024 04:00:00.000 to 4 Dec 2024 04:00:00.000. This method based on the discretization of orbital parameters can effectively avoid the problem of convergence of orbital characteristics of adjacent satellites, thereby ensuring the differences in key parameters such as orbital altitude, inclination, and ascending node right ascension of the experimental samples, providing a statistically significant test environment for verifying the effectiveness of the resource allocation algorithm. The specific selected satellite numbers and TLE parameters are shown in Table 1: Table 1

[0066] Secondly, the basic attribute parameters of all experimental tasks strictly follow the following standardized settings: (1) Geographical location. It is determined by both longitude and latitude coordinates, where latitude is randomly generated uniformly in the range of 3° to 53° north latitude, and longitude is uniformly distributed in the range of 74° to 133° east longitude.

[0067] (2) Arrival time. The arrival time of the task adopts the 24-hour time system and is discretely modeled by converting it to second-level accuracy. The specific value is obtained by uniform random sampling in the interval [0, 86400] seconds.

[0068] (3) Duration. The observation requirement is set as a single observation mode of a single point target, and the observation duration follows a discrete uniform distribution of 3 to 6 seconds.

[0069] (4) Observed returns. Observed returns are generated using the Monte Carlo method and are uniformly randomly assigned in the integer range of 1-10 to quantify the differences in returns between different tasks.

[0070] (5) Consumption and storage. In view of the heterogeneous characteristics of satellites, memory consumption parameters are modeled differently according to the processing capabilities of the satellite platforms: Type A satellites with weak processing capabilities have the largest memory consumption (5-10 units), followed by Type B satellites with medium processing capabilities (3-7 units), while Type C satellites with high performance achieve the lowest memory usage (2-5 units) through optimization algorithms. This parameter setting effectively reflects the heterogeneous capability characteristic of "the stronger the processing capability of the satellite, the less resource consumption during mission execution". The memory consumption ranges of the three types of satellites are randomly generated through discrete uniform distribution, and strictly meet the inclusion relationship that the consumption range of Type A satellites completely covers Type B, and Type B covers Type C.

[0071] Finally, the time window is generated and set. The generation of the time window is based on the collaborative calculation of the orbital dynamics model and the visibility analysis algorithm. The specific process is as follows: (1) The high-precision orbit predictor is called through the Connect module of the STK software, and the spatial geometric relationship between the satellite and the ground station is calculated using the J4 perturbation model. (2) The actual visible range of the ground station is corrected using terrain elevation data, and the ground station altitude threshold is set to 50 meters to meet actual deployment conditions.

[0072] (3) Traverse the spatiotemporal relationship between the satellite constellation and the mission geographic coordinates, and use the numerical integration method to solve the set of visible time intervals that meet the elevation angle constraint conditions.

[0073] (4) The calculation results are discretized and stored in the form of UTC time periods. Each access time window is represented as a continuous sequence of time-state tuples. The final output format is a two-dimensional array structure of [start time, end time].

[0074] This generation mechanism fully considers the multi-physical field coupling effects such as the Earth's rotation, orbital attenuation and terrain shielding, providing precise time-constrained input parameters for subsequent scheduling optimization.

[0075] Based on the above basic settings, we then set the hyperparameters needed in the Improved-MADDPG algorithm. The specific configuration strategy is shown in Table 2.

[0076] Table 2

[0077] This manual uses a controlled variable experimental design to systematically evaluate the performance of Improved-MADDPG and a baseline algorithm. Based on a task scheduling scenario within a reinforcement learning framework, four experimental groups with different task sizes (50, 100, 200, and 500 task units) were constructed to maximize task benefits.

[0078] As the core evaluation criteria, the focus is on the dynamic response characteristics of the average cumulative reward value as the task scale increases and the corresponding algorithm running time. The average reward is as follows: Figure 7 shown.

[0079] from Figure 7 The data clearly shows that as the complexity of the task scenarios increases, from simple to complex, and the task size increases from 50 to 500, the proposed algorithm shows higher average returns in all task scenarios compared to the basic algorithm. This is likely due to the baseline exploration strategy being insufficient, making it difficult to fully capture and effectively utilize the interaction and collaboration information between agents. In contrast, the average return of the improved algorithm significantly exceeds that of the basic algorithm.

[0080] Algorithm convergence analysis is an important part of distributed satellite mission scheduling algorithm optimization. The analysis is conducted from the perspective of convergence, mainly analyzing the algorithm performance from the individual convergence dimension.

[0081] Individual convergence analysis shows that, as shown in Tables 3 and 4, all agents in both algorithms exhibit highly consistent rewards across different task sizes. This is due to the homogeneous reward function design, which ensures balanced task distribution, allowing all agents (satellites) to consistently achieve their goals and maintain positive returns. However, Improved-MADDPG achieves significant performance improvements through its distributed policy optimization architecture. It initially allows agents to explore differently, adopting distributed execution, followed by centralized decision-making to integrate global information and ultimately arrive at the optimal allocation. The improved algorithm achieves an average individual reward of 19.77 after 50 training epochs, a 296% improvement over the baseline algorithm (4.99). With increasing training epochs, the improved algorithm's advantage continues to expand, converging to 50.01 after 500 epochs, a 12.5% improvement over the baseline algorithm (44.45). Furthermore, the reward values for all agents in the improved algorithm are completely consistent (with zero variance), demonstrating that it maintains exceptional stability while improving performance. This result shows that the improved algorithm fully taps the potential of the strategy while preserving the reliability of task allocation by balancing "individual exploration freedom" and "global collaborative constraints", providing a more efficient solution for the collaborative optimization of multi-agent systems.

[0082] Table 3

[0083] Table 4

[0084] Figures 8 to 11The optimal task allocation scheme for the algorithm at different task scales is demonstrated in [1]. Experimental data shows that when the task scale is gradually expanded from 50 to 500, the system completes 42, 86, 174, and 418 tasks, respectively, with corresponding task completion rates of 84.0%, 86.0%, 87.0%, and 83.6%. Across four scale tests, the algorithm consistently maintains a high completion rate exceeding 83%. This series of experimental results demonstrates that the algorithm not only exhibits excellent scalability but also maintains high stability in task processing performance across different scales.

[0085] Figure 12 Comparative experiments deeply analyze the performance differences between the improved algorithm and the classic measurement and control method, namely the first-come, first-served (FCFS) algorithm. As shown in the figure, when the task scale is below 200, the FCFS algorithm demonstrates a short-term advantage, with its task completion rate (82.3%) 2.2 percentage points higher than the Improved-MADDPG algorithm (80.1%). However, when the scale increases to 500, the FCFS algorithm experiences significant performance degradation, with the completion rate plummeting to 68.5%, and the number of completed tasks remaining at only 342. In contrast, the Improved-MADDPG algorithm maintains a high completion rate of 83.6% at the same scale, successfully processing 418 tasks, a 22.2% increase in task throughput compared to FCFS. Experiments demonstrate that the improved algorithm encounters systemic bottlenecks when handling large-scale tasks. This performance difference demonstrates the improved algorithm's robustness to scalability in complex scenarios. Its innovative dynamic observation mechanism effectively overcomes the efficiency degradation of traditional methods in large-scale task scheduling.

[0086] The proposed Heterogeneous Satellite Cluster Autonomous Tracking and Control Resource Allocation Method (HSCMC-RAP) based on the improved MADDPG algorithm effectively addresses the multi-objective trade-off between mission completion rate, resource utilization, and cost-effectiveness through a two-stage decomposition strategy. In the first stage, a genetic algorithm (GA) is used to construct a fitness function centered on cost-effectiveness, achieving three-dimensional matching of satellites, missions, and time windows, and optimizing allocation within the limited visibility window. In the second stage, a distributed decision-making mechanism is implemented based on the improved MADDPG framework. Through global information sharing and local policy iteration, the optimal dynamic configuration of tracking and control resources is achieved. Comparative experiments demonstrate that the improved algorithm exhibits significant advantages over the baseline MADDPG algorithm.

[0087] It should be understood that although the above process Figure 1 The steps in the flowchart are shown in the order indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Figure 1At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0088] In one embodiment, Figure 13 A heterogeneous satellite cluster autonomous measurement and control resource allocation system 100 is provided, comprising a resource acquisition module 11, a model call module 13, a first allocation module 15, a solution sorting module 17, and a second allocation module 19. The resource acquisition module 11 is used to obtain resource parameters for the heterogeneous satellite cluster; the resource parameters include the number of satellites, tasks to be executed, satellite type, and visibility time window. The model call module 13 is used to call a mixed integer programming model constructed based on the autonomous measurement and control resource scheduling requirements of the heterogeneous satellite cluster according to the resource parameters. The first allocation module 15 is used to perform a first-stage joint optimization of tasks, satellites, and visibility time windows on the mixed integer programming model using a genetic algorithm to generate an initial task allocation solution that matches tasks with satellites. The solution sorting module 17 is used to sort the initial task allocation solutions according to the order in which the tasks arrive, generating a complete initial task execution plan for each satellite. The second allocation module 19 is used to utilize the Improved-MADDPG deep reinforcement learning framework improved based on the genetic algorithm, and use the complete mission execution initial plan as prior knowledge to guide target optimization to obtain the optimal task allocation plan for the heterogeneous satellite cluster; wherein, within the Improved-MADDPG deep reinforcement learning framework, the satellite is regarded as an intelligent agent and the task execution process is modeled as a Markov decision process. The optimal task allocation plan is a measurement and control resource allocation plan that achieves the optimal task yield, completion rate and energy consumption under the measurement and control resource constraints of the heterogeneous satellite cluster.

[0089] The aforementioned autonomous TT&C resource allocation system 100 for heterogeneous satellite clusters effectively addresses the multi-objective balancing problem of mission completion rate, resource utilization, and cost-effectiveness through a two-stage decomposition strategy. In the first stage, a genetic algorithm (GA) is used to achieve three-dimensional matching of satellites, missions, and time windows, achieving optimal allocation within the limited visibility window. In the second stage, a distributed decision-making mechanism based on the improved MADDPG framework is implemented. Through global information sharing and local policy iteration, optimal dynamic allocation of TT&C resources is achieved, significantly improving the key computational metric of reward value.

[0090] In one embodiment, the first-stage joint optimization of tasks, satellites, and time windows of a mixed integer programming model using a genetic algorithm includes: constructing an initial population using a global random generation method; constructing a solution space representation using an integer-based two-dimensional encoding strategy; performing calculations using a fitness function with the core goal of minimizing the memory consumption of the task on the satellite; performing selection operations using a tournament selection mechanism; performing crossover operations using a designed single-point crossover operator; and performing mutation operations by introducing a directed mutation strategy and adopting a constraint-aware mechanism.

[0091] In one embodiment, a residual neural network structure enhanced by an attention mechanism is adopted in the Improved-MADDPG deep reinforcement learning framework, and the model convergence is accelerated through the Fix-up parameter initialization method and the Adax adaptive optimizer.

[0092] For the specific limitations of a heterogeneous satellite cluster autonomous measurement and control resource allocation system 100, please refer to the corresponding limitations of a heterogeneous satellite cluster autonomous measurement and control resource allocation method above, which will not be repeated here.

[0093] In one embodiment, a computer device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in any embodiment of the above-mentioned method for allocating autonomous measurement and control resources of a heterogeneous satellite cluster are implemented.

[0094] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media used in the various embodiments provided herein may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus DRAM (RDRAM), and DDR DRAM.

[0095] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0096] The above embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that a person skilled in the art would be able to make numerous modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of the present invention. Therefore, the scope of the present invention shall be determined by the appended claims.

Claims

1. A method for allocating autonomous measurement and control resources for a heterogeneous satellite cluster, characterized in that: Including steps: Obtain resource parameters of a heterogeneous satellite cluster; the resource parameters include the number of satellites, tasks to be performed, satellite types, and visible time windows; Invoking a mixed integer programming model based on the resource scheduling requirements of the heterogeneous satellite cluster autonomous measurement and control resources according to the resource parameters; The mixed integer programming model is optimized for tasks, satellites, and visible time windows in the first phase by using a genetic algorithm to generate an initial task allocation plan that matches tasks with satellites; Sorting the initial task allocation plan according to the order in which the tasks arrive, and generating a complete initial task execution plan for each satellite; An Improved-MADDPG deep reinforcement learning framework based on a genetic algorithm is used to guide target optimization using the complete mission execution initial plan as prior knowledge to obtain an optimal task allocation scheme for a heterogeneous satellite cluster. Within the Improved-MADDPG deep reinforcement learning framework, satellites are considered as intelligent agents and the task execution process is modeled as a Markov decision process. The optimal task allocation scheme is a measurement and control resource allocation scheme that achieves the optimal task yield, completion rate, and energy consumption while satisfying the measurement and control resource constraints of the heterogeneous satellite cluster.

2. The method for allocating autonomous measurement and control resources for a heterogeneous satellite cluster according to claim 1, wherein: The process of performing the first phase of joint optimization of the mission, satellites and time windows of the mixed integer programming model using a genetic algorithm includes: The initial population is constructed using the global random generation method; The solution space representation is constructed using an integer-based two-dimensional encoding strategy; The calculation is performed using a fitness function with the core goal of minimizing the memory consumption of the mission on the satellite; Use tournament selection mechanism for selection operation; Perform crossover operation through the designed single-point crossover operator; The mutation operation is performed by introducing a directed mutation strategy and adopting a constraint-aware mechanism.

3. The method for allocating autonomous measurement and control resources for a heterogeneous satellite cluster according to claim 1 or 2, characterized in that: In the Improved-MADDPG deep reinforcement learning framework, a residual neural network structure enhanced by the attention mechanism is adopted, and the model convergence is accelerated through the Fix-up parameter initialization method and the Adax adaptive optimizer.

4. The method for allocating autonomous measurement and control resources for a heterogeneous satellite cluster according to claim 3, wherein: In the Improved-MADDPG deep reinforcement learning framework, the Gumbel-Softmax function is used to construct a discrete-continuous hybrid action space mapping mechanism.

5. The method for allocating autonomous measurement and control resources for a heterogeneous satellite cluster according to claim 3, wherein: In the Improved-MADDPG deep reinforcement learning framework, the immediate reward is dynamically smoothed by constructing an exponential moving average function.

6. A heterogeneous satellite cluster autonomous measurement and control resource allocation system, characterized in that: include: Resource acquisition module, used to obtain resource parameters of heterogeneous satellite clusters; The resource parameters include the number of satellites, the mission to be performed, the type of satellite and the visible time window; A model calling module is used to call a mixed integer programming model built based on the autonomous measurement and control resource scheduling requirements of a heterogeneous satellite cluster according to the resource parameters; A first allocation module is configured to perform a first-stage joint optimization of tasks, satellites, and visible time windows on the mixed integer programming model using a genetic algorithm to generate an initial task allocation plan that matches tasks with satellites; A scheme sorting module is used to sort the initial task allocation schemes according to the order in which the tasks arrive, and generate a complete initial task execution plan for each satellite; The second allocation module is used to utilize the Improved-MADDPG deep reinforcement learning framework improved based on the genetic algorithm, and use the complete mission execution initial plan as prior knowledge to guide target optimization to obtain the optimal task allocation plan for the heterogeneous satellite cluster; wherein, within the Improved-MADDPG deep reinforcement learning framework, the satellites are regarded as intelligent agents and the task execution process is modeled as a Markov decision process. The optimal task allocation plan is a measurement and control resource allocation plan that achieves the optimal task yield, completion rate, and energy consumption under the measurement and control resource constraints of the heterogeneous satellite cluster.

7. The autonomous measurement and control resource allocation system for heterogeneous satellite clusters according to claim 6, characterized in that: The process of performing the first phase of joint optimization of the mission, satellites and time windows of the mixed integer programming model using a genetic algorithm includes: The initial population is constructed using the global random generation method; An integer-based two-dimensional encoding strategy is used to construct the solution space representation; The calculation is performed using a fitness function with the core goal of minimizing the memory consumption of the mission on the satellite; Use tournament selection mechanism for selection operation; Perform crossover operation through the designed single-point crossover operator; The mutation operation is performed by introducing a directed mutation strategy and adopting a constraint-aware mechanism.

8. The autonomous measurement and control resource allocation system for heterogeneous satellite clusters according to claim 6 or 7, characterized in that: In the Improved-MADDPG deep reinforcement learning framework, a residual neural network structure enhanced by the attention mechanism is adopted, and the model convergence is accelerated through the Fix-up parameter initialization method and the Adax adaptive optimizer.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method for allocating autonomous measurement and control resources of a heterogeneous satellite cluster as described in any one of claims 1 to 5 are implemented.

Citation Information

Cited By

  • Satellite image data downlink scheduling method, system and device

    CN120849069A

  • Satellite image data downlink scheduling method, system and device

    CN120849069B