A Multi-Star Autonomous Cooperative Scheduling Method Based on Distributed Multi-Agent Reinforcement Learning
The multi-satellite autonomous collaborative scheduling algorithm based on distributed multi-agent reinforcement learning solves the problems of flexibility and timeliness in imaging satellite mission planning, and realizes rapid response and efficient communication in satellite autonomous decision-making and multi-satellite collaborative scheduling.
Patent Information
- Application Number
- CN202411486528.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-23
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-10-23
AI Technical Summary
Existing imaging satellite mission planning modes are carried out on a daily basis, which lacks flexibility and timeliness, fails to fully utilize the satellite's observation capabilities, makes it difficult to respond quickly to dynamically changing environments, and the centralized architecture significantly reduces system performance during communication failures.
A multi-satellite autonomous collaborative scheduling algorithm based on distributed multi-agent reinforcement learning is adopted. Through centralized training and distributed execution framework, a separate decision network is built for each satellite. Distributed autonomous decision-making is carried out by utilizing local communication information, which reduces inter-satellite communication consumption and improves system resilience and response speed.
It achieves good mission planning results with low communication consumption, has rapid response capability and high system resilience, and is suitable for satellite autonomous decision-making and multi-satellite collaborative scheduling.
Smart Images

Figure CN119623910B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of multi-satellite autonomous cooperative scheduling using distributed multi-agent reinforcement learning, enabling satellites to make distributed autonomous decisions based on local communication information. Background Technology
[0002] With the rapid advancement of satellite technology, future satellites will possess advanced functions such as onboard automatic processing, onboard autonomous planning, inter-satellite link communication, and data broadcasting and distribution. High efficiency, autonomy, and intelligence are becoming important trends in the future development of satellite mission planning systems. However, current imaging satellite mission planning typically operates on a daily cycle, employing a "ground-based decision-making, onboard execution" model. This model lacks flexibility and timeliness. Because the satellite-ground link cannot be constantly communicated, the ground cannot grasp the real-time resource status of the satellite, thus failing to fully utilize the satellite's observation capabilities and hindering rapid responses to dynamically changing environments. This traditional model is increasingly unable to meet the rapid response requirements of future missions, urgently necessitating the exploration of new imaging satellite mission planning models to more effectively address the complex challenges facing future space development.
[0003] The design of onboard algorithms typically relies on a multi-satellite collaborative architecture. Current research on collaborative architectures can be categorized into two types: centralized and distributed. Centralized architectures usually depend on a central node that can obtain information from all child nodes and perform comprehensive planning accordingly. Its advantages lie in its simple operation and accurate solutions, but it places high demands on the central node's computing power and the inter-satellite communication environment. If the central node fails, the overall system performance will be significantly reduced. To cope with more complex communication environments and improve system resilience, distributed scheduling methods have gained popularity. In a distributed architecture, each child node undertakes a certain amount of computational work, rather than simply acting as an executor of instructions. That is, each node in the system possesses a degree of autonomy, and the nodes continuously negotiate to obtain a consistent planning result, such as the Contract Network Protocol (CNP) and the Consensus-based Bundle Algorithm (CBBA). However, such algorithms are often accompanied by high communication overhead. Considering the risks of inter-satellite link failure, external interference, communication delay, and data loss, when the communication quality is poor, the on-board system needs the algorithm to rely less on inter-satellite communication. Multi-agent deep reinforcement learning (MADRL) has shown great potential in this scenario.
[0004] MADRL allows multiple agents to learn interactively within an environment, with each agent maximizing its cumulative reward through interactions with the environment and other agents. Centralized training & decentralized execution (CTDE) is a commonly used algorithmic framework in MADRL. It equips each agent with a decision network and introduces a central control network during the training phase. With the aid of global state information, all agent networks are integrated for joint training, thus avoiding the problems of environmental non-stationarity and excessively large state-action spaces, ensuring the training effectiveness of each agent's decision network. During the execution phase, each agent only needs to make independent decisions based on its own locally observed information, achieving rapid task response. Lowe et al. proposed the multi-agent deep deterministic policy gradient (MADDPG) algorithm, which is an extension of the classic deep deterministic policy gradient (DDPG) algorithm in a multi-agent environment. Foerster et al. proposed the Counterfactual Multi-Agent Policy Gradients (COMA) algorithm, which aims to solve the credit allocation problem in multi-agent environments. It estimates the contribution of a single agent's action to the overall task by comparing its expected reward with the counterfactual baseline while keeping the actions of other agents constant. Sunehag et al. proposed Value Decomposition Networks (VDN), a method that decomposes the joint action-value function of multiple agents into individual action-value functions for each agent. Its core idea is to maximize the joint action-value function while maximizing the action-value function of each agent. Subsequent research has focused on improving this algorithm, such as the QMIX and QTRAN algorithms, effectively enhancing its versatility and solution performance.
[0005] Currently, research on applying MADRL to imaging satellite mission planning is still in its early stages. Wang Haijiao proposed a distributed online satellite scheduling algorithm based on MADDGP to address this problem. Its application context is defined as observation requests being submitted to the system one by one, and each satellite within the system needing to respond quickly to the received tasks based on its own acquired observation information. In similar fields, Prasad and Dusparic applied the MADRL model to energy allocation problems, and Yun et al. applied the MADRL algorithm to autonomous UAV collaboration in a distributed environment. The MADRL method still holds significant research potential in the field of satellite mission planning.
[0006] To meet the future development needs of satellites for high timeliness, autonomy, and intelligence, and to fully leverage the potential brought about by technological changes such as satellite autonomous decision-making and multi-satellite collaboration, this invention proposes a multi-satellite autonomous collaborative scheduling algorithm based on distributed multi-agent reinforcement learning. This algorithm effectively overcomes the limitations of current on-board computing power and unstable inter-satellite communication links, and supports multi-satellite distributed autonomous decision-making in local communication environments. Summary of the Invention
[0007] The purpose of this invention is to provide a multi-satellite autonomous cooperative scheduling algorithm based on distributed multi-agent reinforcement learning, supporting distributed autonomous decision-making by each satellite based on local communication information. This invention achieves good mission planning results while maintaining low communication consumption, and possesses rapid response capabilities and good system resilience.
[0008] To achieve the above objectives, the present invention adopts the following technical solution:
[0009] In a first aspect, the present invention provides a multi-satellite autonomous cooperative scheduling algorithm based on distributed multi-agent reinforcement learning, specifically including the following steps:
[0010] S1. Based on target attribute information, satellite attribute information, and various constraints, mission planning is carried out. Mission planning is used to solve how to effectively allocate and schedule multiple satellite resources and formulate satellite observation plans to maximize the completion of user-submitted tasks.
[0011] S2. Construct a partially observable Markov decision process model adapted to the multi-satellite autonomous collaborative planning mode based on the satellite mission planning.
[0012] S3. Considering the limited autonomous planning capabilities of satellites, the distributed multi-agent reinforcement learning algorithm QTRAN is used to construct a separate decision network for each satellite;
[0013] S4. It adopts a centralized training and distributed execution framework for the application mode of ground training and on-board execution.
[0014] S5. Centralized training phase: Centralized training is conducted on all satellites, and collaborative training is carried out by uniting all satellite networks.
[0015] S6, Distributed Execution Phase: Each satellite makes independent decisions based on its local network and interacts with only a subset of satellites.
[0016] Furthermore, the task planning in step S1 includes an objective function and constraints; specifically, the objective function is:
[0017]
[0018] Constraints:
[0019]
[0020] Formula (1-1) is the objective function, representing maximizing the total task reward. Where x... ijk As a decision variable, when x ijk When = 1, it means that task t will be... j Arrange for satellites to be used i The k-th visible time window is executed, at which time ts j =ws ijk ,te j =we ijk Conversely, x ijk =0; Time window benefit r ijk The definition is as follows:
[0021]
[0022] This approach incorporates a mechanism where rewards decay over time, taking into account both the importance of the task and the need for timely response. β is the decay factor. The attenuation coefficient is given by formula (1-2), which represents the storage constraint. Within the planning period, the storage occupied by the satellite in performing the mission must not exceed the satellite's maximum storage capacity. Formula (1-3) represents the mission execution uniqueness constraint, indicating that each mission can be executed at most once. Formula (1-4) represents the attitude transition time constraint, indicating that the interval between two adjacent missions must be greater than their attitude transition time. Formula (1-5) represents the mission continuous observation time constraint, indicating that the actual observation time of the mission must meet the mission imaging duration requirement. Formula (1-6) defines the range of values for the decision variables.
[0023] Furthermore, step S2 specifically includes the following steps:
[0024] Multiple satellites can be viewed as a fully cooperative multi-agent system. This multi-agent system is modeled as a decentralized partially observable Markov decision process (Dec-POMDP), represented by tuples. in, Represents the real environment in which the intelligent agent exists; Let {1,...,N} be the set of agents; the action chosen by the i-th agent is represented as... The actions chosen by all agents constitute a joint action u = [u 1 ,u 2 ,...,u N ]; Z represents the local observation of the i-th agent; Z represents the observation function. P(s′|s,u) represents the state transition probability; R represents the reward function. In a fully cooperative environment, all agents share the same reward function, denoted as r=R(s,u); γ is the discount factor, used to balance immediate rewards and future rewards.
[0025] Furthermore, step S2 also includes the following steps:
[0026] In a multi-satellite autonomous collaborative mission planning system, after receiving a mission request, the satellites within the system make decisions in a distributed manner, generating their own mission planning schemes. This approach eliminates the need for information sharing across the entire system; each satellite makes independent decisions based solely on its own available information, thus enabling rapid response to requests. The satellite set is represented as SAT = {sat1,sat2,...,sat...} N}, containing a total of N satellites, sat i Let represent the i-th satellite; the set of tasks to be planned is represented as T = {t1, t2, ..., t}. |T|} contains |T| tasks, t j This represents the j-th task in the task set.
[0027] Furthermore, step S3 specifically includes the following steps: the multi-agent reinforcement learning algorithm VDN based on the value function decomposition idea adopts the CTDE framework, and each agent deploys a decision network Q. i It can learn and construct its own action value function. The idea of value function decomposition is to construct a hybrid network to fit the joint action value function Q during the centralized training phase. jt Through training, the action value function Q of each agent is made more efficient. i With the joint action value function Q jt The following relationship must be satisfied:
[0028]
[0029] Formula (2-1) is called the individual-to-collective maximum condition, where Let represent the action-observation history of the agents. When this condition is met, the optimal action chosen by each agent based on its own decision network is equivalent to the joint optimal action of the entire system, thus ensuring the overall system optimality even when each agent makes independent decisions; let Represents a vector consisting of the Q-value functions of all agents; when [Q i ] To Q jt When the IGM condition is satisfied, it is called [Q] i ] is Q jt To construct value decomposition relations that satisfy the IGM, VDN proposes an additive decomposition method:
[0030]
[0031] This constraint enables value decomposition that satisfies the IGM condition, but it also imposes structural constraints on the problem. Based on this, the QMIX algorithm extends this additive decomposition by proposing a monotonic decomposition method, as shown in Equation (2-3), which can fit a more complex relationship between the joint Q-value and the Q-values of each agent.
[0032]
[0033] The QTRAN algorithm proposes a decomposition method for this, which decomposes the original joint action value function Q... jt Transform into a new, more easily decomposable function Q′ jt By ensuring that the joint optimal action of the two is the same, the IGM condition is satisfied;
[0034] definition Let represent the optimal action of the i-th agent, and let represent the set of optimal actions of all agents. The QTRAN algorithm provides [Q i A sufficient condition for satisfying IGM:
[0035]
[0036] in
[0037]
[0038] The QTRAN algorithm directly uses the transformation function Q′ jt The definition is as follows:
[0039]
[0040] The construction method defined by formula (2-6) is similar to the VDN algorithm, and obviously [Q i Satisfying Q′ jt The IGM condition, and because Therefore, [Q] satisfies formula (2-4). i This can be considered as Q′ jt Value decomposition;
[0041] Formula (2-4) will serve as the basis for training the QTRAN algorithm, which means [Q i ] and Q′ jt The additive decomposition relationship between them will be characterized by formula (2-4); during the algorithm training process, three types of functions are involved: the Q-value function of each agent [Q i ], Joint action value function Q jt and function V jtDue to the transformation function Q′ jt It can be directly from [Q] i The symbol ] indicates that no further definition is needed. To more clearly demonstrate Q... jt and its transformation function Q′ jt The relationship between the equations (2-3) and (2-6) leads to equation (2-7):
[0042]
[0043] The above analysis shows that [Qi] and Q′ jt The relationship will be through Q jt and V jt To fit the data; during the training process, the QTRAN algorithm additionally introduces the function V. jt V jt It can be considered a correction term used to correct Q. jt and Q′ jt The differences between them can be used to characterize the complex relationships in multi-agent systems.
[0044] Furthermore, step S4 specifically includes: under the CTDE framework, each satellite will deploy a decision network. During the training phase, an additional hybrid network will be constructed to assist learning. During the execution phase, each satellite only needs to make decisions independently based on local observation information.
[0045] Furthermore, step S5 specifically includes: in order to improve the generalization ability of the algorithm in different scenarios, the training process adopts an algorithm training mechanism oriented towards random initial scenarios: several training scenarios are generated, and a scenario is randomly selected for training in each iteration; in the predetermined scenario, each satellite makes its own decision in a distributed manner to obtain its own mission planning scheme; the training data generated by this decision-making process will be put into the experience pool for subsequent algorithm training.
[0046] Furthermore, step S6 specifically includes: for a homogeneous multi-satellite system, during distributed execution, each imaging satellite can be considered to have the same decision-making process. During the execution phase, the satellites... i The pre-trained model Q was pre-deployed. i After receiving the task set, the satellite sat i Based on model Q i Make decisions independently for the entire task set.
[0047] This application addresses the future development needs of satellites for autonomy, intelligence, and rapid response. Considering the current limitations of satellite computing power, simple inter-satellite coordination mechanisms, and unstable communication links, it innovatively proposes a multi-satellite autonomous cooperative scheduling algorithm based on distributed multi-agent reinforcement learning. First, a partially observable Markov decision process model adapted to multi-satellite autonomous cooperative planning is constructed, with each component of the model specifically designed based on actual onboard conditions. Given the limited autonomous planning capabilities of satellites, this application introduces a low-complexity distributed multi-agent reinforcement learning algorithm—QTRAN. This algorithm constructs a separate decision network for each satellite and employs a centralized training and distributed execution framework, suitable for both ground training and onboard execution. In the centralized training phase, collaborative training is conducted by uniting all satellite networks to cultivate the global optimization awareness of each satellite, ensuring the overall mission planning effect. In the distributed execution phase, each satellite makes independent decisions based on its local network, requiring only interaction with a subset of satellites. Finally, this application constructs multi-satellite mission planning scenarios of different scales to verify the algorithm's effectiveness. Simulation results show that the algorithm can effectively reduce inter-satellite communication consumption while ensuring planning effectiveness, and has faster response speed and higher system resilience. Attached Figure Description
[0048] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0049] Figure 1 This is a schematic diagram of the local observation space in which multi-source information is fused in this invention;
[0050] Figure 2 This is the single-star network structure in this invention;
[0051] Figure 3 This is the hybrid network structure in this invention;
[0052] Figure 4 The algorithm execution process at decision step t in this invention
[0053] Figure 5 Analysis of the QTRAN algorithm training and evaluation process. Detailed Implementation
[0054] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0055] Imaging satellites must consider various factors when completing imaging missions. When a user submits an imaging request, mission planning needs to be performed based on target attribute information, satellite attribute information, and various constraints. Mission planning primarily addresses how to effectively allocate and schedule resources from multiple satellites, and formulate satellite observation plans to maximize the completion of the user-submitted task. The effectiveness of mission planning directly impacts the performance of the imaging satellite Earth observation system and is a crucial step in ensuring the system's efficient operation.
[0056] Because imaging satellite mission planning involves many complex constraints, this application makes the following assumptions and simplifications to highlight the research focus:
[0057] (1) All satellites have sufficient power.
[0058] (2) The image download process is not considered.
[0059] (3) All targets are point targets and can be observed by a single satellite within a single time window.
[0060] Table 1 shows the relevant symbols and explanations for imaging satellite mission planning issues.
[0061] Table 1. Symbols and Explanations for Imaging Satellite Mission Planning Issues
[0062]
[0063] Task planning includes objective functions and constraints; specifically,
[0064] Objective function:
[0065]
[0066] Constraints:
[0067]
[0068] Formula (1-1) is the objective function, representing maximizing the total task reward. Where x... ijk As a decision variable, when x ijk When = 1, it means that task t will be... j Arrange for satellites to be used iThe k-th visible time window is executed, at which time ts j =ws ijk ,te j =we ijk Conversely, x ijk =0. Time window benefit r ijk The definition is as follows:
[0069]
[0070] This function introduces a mechanism where the reward decays over time, comprehensively considering both the importance of the task and the need for timely response. β is called the decay factor. The attenuation coefficient is given by formula (1-2). Formula (1-2) represents the storage constraint, which states that the storage occupied by the satellite for performing tasks must not exceed the satellite's maximum storage capacity within the planning period. Formula (1-3) represents the task execution uniqueness constraint, which states that each task can be executed at most once. Formula (1-4) represents the attitude transition time constraint, which states that the interval between two adjacent tasks must be greater than the attitude transition time. Formula (1-5) represents the task continuous observation time constraint, which states that the actual observation time of the task must meet the task imaging duration requirement. Formula (1-6) defines the range of values for the decision variables.
[0071] In multi-satellite autonomous collaborative scheduling, multiple satellites can be viewed as a fully cooperative multi-agent system, sharing the same environment and influencing each other. Since agents often cannot observe the entire environment, multi-agent systems are typically modeled as decentralized partially observable Markov decision processes (Dec-POMDPs), represented as tuples. in, Represents the real environment in which the intelligent agent exists; Let {1,...,N} be the set of agents; the action chosen by the i-th agent is represented as... The actions chosen by all agents constitute a joint action u = [u 1 ,u 2 ,...,u N ]; Z represents the local observation of the i-th agent; Z represents the observation function. P(s′|s,u) represents the state transition probability; R represents the reward function. In a fully cooperative environment, all agents share the same reward function, denoted as r=R(s,u); γ is the discount factor, used to balance immediate rewards and future rewards.
[0072] In 2017, the DeepMind team proposed a multi-agent reinforcement learning algorithm based on the value function decomposition idea—VDN. This algorithm achieved good results in solving multi-agent problems, and subsequently, many improved algorithms were derived based on this idea, becoming an important research direction in the field of multi-agent reinforcement learning. The VDN algorithm uses the CTDE framework, where each agent deploys a decision network Q. i It can learn and construct its own action value function. The idea of value function decomposition is to construct a hybrid network to fit the joint action value function Q during the centralized training phase. jt Through training, the action value function Q of each agent is made more efficient. i With the joint action value function Q jt The following relationship must be satisfied:
[0073]
[0074] Formula (2-1) is called the individual-global-max (IGM) condition, where Let represent the agent's action-observation history. When this condition is met, the optimal action chosen by each agent based on its own decision network is equivalent to the joint optimal action of the entire system, thus ensuring the overall system optimality even when each agent makes independent decisions. Let This represents a vector consisting of the Q-value functions of all agents. When [Q...] i ] To Q jt When the IGM condition is satisfied, it is called [Q] i ] is Q jt Value decomposition. To construct value decomposition relations that satisfy the IGM, VDN proposes an additive decomposition method:
[0075]
[0076] This constraint enables value decomposition that satisfies the IGM condition, but it also imposes structural constraints on the problem. Based on this, the QMIX algorithm extends this additive decomposition by proposing a monotonic decomposition method, as shown in Equation (2-3), which can fit a more complex relationship between the joint Q-value and the Q-values of each agent.
[0077]
[0078] The QTRAN algorithm proposes a more general decomposition method for this, which decomposes the original joint action value function Q... jt Transform into a new, more easily decomposable function Q′ jt The IGM condition is satisfied by ensuring that the joint optimal action of the two is the same.
[0079] definition Let represent the optimal action of the i-th agent, and let represent the set of optimal actions of all agents. The QTRAN algorithm provides [Q i A sufficient condition for satisfying IGM:
[0080]
[0081] in
[0082]
[0083] The QTRAN algorithm directly uses the transformation function Q′ jt The definition is as follows:
[0084]
[0085] The construction method defined by formula (2-6) is similar to the VDN algorithm, and obviously [Q i Satisfying Q′ jt The IGM condition, and because Therefore, [Q] satisfies formula (2-4). i This can be considered as Q′ jt Value decomposition.
[0086] Formula (2-4) will serve as the basis for training the QTRAN algorithm, which means [Q i ] and Q′ jt The additive decomposition relationship between them will be characterized by formula (2-4). Therefore, three types of functions are involved in the algorithm training process: the Q-value function of each agent [Q i ], Joint action value function Q jt and function V jt Due to the transformation function Q′ jt It can be directly from [Q] i The symbol ] indicates that no further definition is needed. To more clearly demonstrate Q... jt and its transformation function Q′ jt The relationship between the equations (2-3) and (2-6) leads to equation (2-7):
[0087]
[0088] Based on the above analysis, [Q] i ] and Q′ jt The relationship will be through Q jt and V jt To fit. Therefore, although the transformation function Q′ jt Similar in form to the VDN algorithm, but the QTRAN algorithm introduces an additional function V during training. jt Vjt It can be considered a correction term used to correct Q. jt and Q′ jt The differences between them can be used to characterize more complex relationships in multi-agent systems.
[0089] In multi-satellite autonomous collaborative scheduling scenarios, communication constraints limit each satellite to communication only with neighboring satellites. Information from more distant satellites requires intermediary satellites. This approach is not only time-consuming but also prone to signal interference and data loss, reducing the robustness of mission planning. Therefore, the algorithm aims to minimize communication overhead while maintaining mission planning effectiveness. The CTDE framework used in multi-agent reinforcement learning algorithms demonstrates its applicability in this context. Within the CTDE framework, each satellite possesses an independent decision-making network. The centralized training phase can be conducted on the ground, integrating all satellite networks using pre-acquired onboard data for joint training. After training, the decision-making network is deployed to the corresponding satellites. In the distributed execution phase, each satellite independently plans based on its decision model, requiring only communication with a subset of satellites to make optimal decisions, thus significantly reducing inter-satellite communication overhead. Therefore, the aforementioned multi-agent reinforcement learning algorithm based on value decomposition is effectively applicable to multi-satellite autonomous collaborative scheduling scenarios.
[0090] In imaging satellite mission planning, uniqueness constraints are set to avoid resource waste. This means that from the perspective of a single satellite, the optimal action to ensure the overall mission planning effect is not necessarily to select and execute the mission, which makes the relationship between the joint Q-value and the Q-values of individual satellites more complex. Therefore, this application chooses the QTRAN algorithm, which has a more general value decomposition method, to solve the multi-satellite autonomous mission planning problem.
[0091] In a multi-satellite autonomous collaborative mission planning system, upon receiving a mission request, the satellites within the system make decisions in a distributed manner, generating their own mission planning schemes. This approach eliminates the need for information sharing across the entire system; each satellite makes independent decisions based solely on its own available information, thus enabling rapid response to requests. The satellite set is represented as SAT = {sat1,sat2,...,sat...} N} contains a total of N satellites, sat i Let represent the i-th satellite; the set of tasks to be planned is represented as T = {t1, t2, ..., t}. |T|} contains |T| tasks, t j This represents the j-th task in the task set.
[0092] Since the design of multi-satellite autonomous mission planning algorithms is inseparable from the characteristics of the satellites themselves, this application defines the characteristics of the research object in detail, as follows:
[0093] (1) All satellites are isomorphic. This indicates that the satellites have similar imaging capabilities, follow common mission planning rules, and can be solved using the same decision-making model.
[0094] (2) Inter-satellite communication is limited. Considering the many problems such as signal interference, data loss, energy limitations, and the time cost incurred by communication, the algorithm design should minimize the consumption of inter-satellite communication.
[0095] From the perspective of a single satellite, the following characteristics are defined:
[0096] (1) Satellites have multiple channels to obtain mission planning information, such as ground transmission, local observation, and inter-satellite communication.
[0097] (2) Due to communication limitations, satellites can only obtain part of the system information through inter-satellite communication.
[0098] (3) Throughout the mission planning process, the satellite's behavior can be summarized as: receiving the mission set—collecting information—mission planning—execution of the plan. During this process, the satellite's behavior is completely independent of other satellites. Inter-satellite interaction is only involved in the information collection phase, but this behavior is initiated autonomously by the satellite and is solely for the purpose of acquiring information. Inter-satellite collaboration refers to strategy-level collaboration, and this strategy network is pre-trained and deployed on each satellite.
[0099] For a fully cooperative multi-agent system, at a certain decision step, after all satellites perform actions on the environment, the environment will return a joint action reward. To train the satellites to achieve tacit cooperation under distributed decision-making, the model is set such that within a predetermined decision step, all satellites make decisions for the same task. This ensures that all satellites face the same decision-making object during training, thus making the information used for decision-making more similar. Under this setting, after receiving the task set, each satellite first sorts the tasks according to a predetermined rule to obtain a unified decision order. This application sets the sorting rule to sort according to task priority from high to low, and the sorted task decision sequence is represented as T′={t′1,t′2,...,t′ |T|}. The satellite's SAT i For task t′ j The action space for making decisions is represented as follows:
[0100]
[0101] The decision space size for each decision step is "number of satellites + 1". This indicates that no satellite is selected. Indicates satellite SAT i Task t′ j Assign it to the satellite with the corresponding index.
[0102] Action selection must strictly satisfy all constraints in the mathematical model in Section 1.1.3. Considering the constraint coupling relationship between tasks, a constraint single-step checking mechanism is designed to screen the action selection space. Specifically:
[0103] (1) For onboard storage constraints, a forward storage check mechanism was designed. Specifically, for the current decision task t′... j Assuming it occupies storage space sm′ j Satellite i Before making a decision, the saved status information of each satellite will be pre-screened. k Remaining storage Does it meet the requirements? Then the action Not selectable.
[0104] (2) Regarding time window constraints, when two time windows overlap or do not meet the transition time requirements, they are said to conflict. A single-step time window check mechanism is designed to address this, specifically: satellite SAT... i Before the initial decision, a 0-1 variable is set to indicate the task t′. j Does a time window satisfy the constraints exist on each satellite? in, Indicates satellite SAT i Assuming task t′ j In satellite SAT k There is a time window, action Satisfy the time window constraint. Conversely, if This indicates an action. Not optional. Each time a decision is made, the flag variable for the remaining tasks must be updated.
[0105] (3) In particular, multiple satellites make decisions based on local observation information, which inevitably leads to the repeated execution of tasks, making it difficult to satisfy the uniqueness constraint (Equation (1-3)). Since this constraint is intended to avoid wasting satellite resources and is not due to the limitation of satellite execution capabilities, it is defined as a soft constraint and used as the learning target of the algorithm to minimize the repeated execution of tasks.
[0106] Under the multi-agent reinforcement learning algorithm, each satellite independently plans its mission, aiming to formulate its own mission execution plan. Therefore, although each satellite can generate a plan for the entire system during this decision-making process, it ultimately only retains the plan relevant to itself. i The set of tasks to be performed on this satellite after the decision can be represented as:
[0107]
[0108] Since mission planning is system-wide, and each satellite in the system makes independent decisions in a distributed manner, the more comprehensive the satellite's global information, the better, to ensure the overall mission planning effectiveness. Satellites acquire information through multiple channels, including ground-based uploads, local satellite sensing, and inter-satellite interactions. While ground-based systems can predict the state of the entire onboard system, errors are inevitable. Onboard information is highly real-time, but communication limitations prevent the timely acquisition of all information. Based on the above analysis, the model designs a local observation space that integrates multi-source information, using satellite SAT... i For example, its local observation space can be represented as:
[0109] O i ={O i1 O i2 ,...,O iN} 2-10
[0110] satellites i It maintains information on all satellites within the system, the difference being that the information sources for different satellites vary. Figure 1 Using four satellites as an example, this illustrates the information sources within a satellite's local observation space. Assume these four satellites are arranged in a ring within the same orbital plane, and each satellite can establish communication links with its two neighboring satellites (e.g., satellite sat1's neighboring satellites are sat2 and sat4). While uploading mission sets to the satellites, the ground also uploads mission planning information (predictive data) for all satellites within the corresponding orbital plane. Upon receiving this information, the satellites first update their own mission planning information based on local sensing. Subsequently, through inter-satellite communication, the satellites can further update the mission planning information of surrounding satellites. However, due to limitations in inter-satellite communication capabilities, satellites cannot obtain information from all satellites within the entire orbital plane in a timely manner. Therefore, the information maintained by each satellite will ultimately include information from the ground, its own satellite, and neighboring satellites, as described in [see section 1]. Figure 1 (c)
[0111] Will make decision steps j satellitesat i Maintenance of satellite SAT k The information is represented as:
[0112]
[0113] Indicates satellite SAT k Execute task t′ j The potential gains (see formula (1-7)); Indicates satellite SAT k For task t′ jThe degree of conflict within the visible time window further subdivides this feature into two categories: in This indicates the number of time windows that conflict with this time window, simply referred to as the conflict number. The total revenue corresponding to time windows that conflict with this time window is referred to as the conflict value. Indicates satellite SAT k Available storage at decision step j; Indicates satellite SAT k The number of tasks already assigned at decision step j. It is a static feature that remains unchanged throughout the decision-making process; It is a dynamic feature and needs to be updated after each decision.
[0114] The global state space s∈S is represented in the same way as the local observation space maintained by a single star. The difference is that the information represented by the global state space comes entirely from the on-board environment, that is, it represents the real-time information of the entire system.
[0115] In the QTRAN algorithm design, after all agents have selected actions, the environment provides joint action rewards, which are recorded at decision step j, where all satellites are task t′. j After the decision is made, the joint action reward from environmental feedback is r. j The corresponding reward function for the joint action is expressed as:
[0116]
[0117] As shown in formula (2-9), when calculating the reward for joint actions, it is only necessary to consider whether each satellite chooses to perform the current task t′. j Assign it to itself, thereby setting the decision variable. Used to represent satellite sat i Do you want to select task t′ to execute? j Specifically, it is expressed as:
[0118]
[0119] Formula (2-12) means: when multiple satellites select observation task t′ j In this case, the joint action reward is the ratio of the maximum benefit gained by each observation satellite to the total number of observation satellites. This is because, in a problem scenario where a task only needs to be observed once, even if the same task is observed multiple times, only one benefit can be obtained. Simultaneously, to avoid resource waste, a penalty is imposed on repeated observation behavior by dividing by the total number of observation satellites, thereby guiding multi-satellite decision-making behavior. Based on this, the cumulative reward obtained in a scenario can be expressed as:
[0120]
[0121] Within the CTDE framework, each satellite will deploy a decision network. During the training phase, an additional hybrid network will be constructed to assist learning. During the execution phase, each satellite will make decisions independently based solely on local observation information. This application presents the specific structures and update methods of the decision networks and hybrid networks for each satellite, and elaborates on the centralized training and distributed execution processes in detail, considering the specific context of satellite mission planning.
[0122] Since each satellite makes decisions under partially observable conditions, a deep recurrent Q-learning (DRQN) network suitable for partially observable Markov decision processes is configured on each satellite. The specific network structure is shown in [link to network structure]. Figure 2 MLP stands for MLP, a basic neural network structure, where the hidden state... It is a key component of the DRQN algorithm, providing the agent with a memory mechanism. By updating at each decision step, it retains past observation information, enabling the agent to better utilize historical information to make decisions. Gating mechanisms in networks, such as LSTM or GRU, are often used to facilitate learning on longer time scales.
[0123] Each satellite sat i It has a separate DRQN network that receives individual observations at each decision step. Actions at the previous time step Output the Q of all actions i (τ i The Q-values are calculated, and during the training phase, an ε-greedy strategy is used to select actions and output the Q-values of the corresponding actions.
[0124] The QTRAN algorithm needs to construct, as follows: Figure 3 Three independent neural networks to fit Q′ jt and [Q i Value decomposition relationship between ].
[0125] Q-network of each satellite: All satellite network outputs Q i The sum of the values constitutes Q′ jt , Figure 3 (c) illustrates this process.
[0126] United Q Networks: Used to calculate the combined action value Q of all satellites jt This network sums the hidden layer features of all individual networks. The hidden state parameters h of each satellite are shared in this way. t Simultaneously, with the coordinated actions of all satellites... t and global state s t As input, the process is as follows Figure 3 As shown in (a).
[0127] State-value networks: Calculate the scalar state value, used to correct Q. jt and Q′ jt The deviation. This network also shares the hidden layer feature parameters of individual networks. As shown in formula (2-5), the calculation of the state value is independent of the action selection of each satellite. Therefore, apart from the hidden layer state parameters, the network only uses the global state s. t As input, Figure 3 (b) demonstrates this process.
[0128] The QTRAN algorithm uses the CTDE framework, which differs between the training and execution phases. The centralized training and distributed execution phases will be described in detail separately.
[0129] The centralized training process is described in Algorithm 2.1. To improve the algorithm's generalization ability in different scenarios, a training mechanism oriented towards random initial scenarios is adopted: several training scenarios are generated, and a scenario is randomly selected for training in each iteration. Under the predetermined scenario, each satellite makes independent decisions in a distributed manner to obtain its own mission planning scheme. The training data generated by this decision-making process will be placed in the experience pool for subsequent algorithm training.
[0130]
[0131]
[0132] The specific decision-making steps for each satellite are as follows: Upon receiving a task set, it acquires and updates local observation information, which can come from various channels such as ground-based uploads, local satellite sensing, and inter-satellite interactions. Then, it determines the task decision-making order according to predetermined rules and makes decisions for each task one by one. Lines 7-11 of Algorithm 2.1 illustrate this sequential decision-making process. First, it performs constraint checks on the current decision task to determine the set of possible actions. Then, based on this, it selects actions according to an ε-greedy strategy and updates the state. When all tasks have been decided, the local satellite's task planning scheme is obtained. Furthermore, during the centralized training phase, it is necessary to additionally save global state information. This global information is independent of the decision-making process of each satellite and is only used as part of the input to the hybrid network during the training phase to assist model training.
[0133] Because the training process of deep reinforcement learning algorithms requires substantial computational resources, which are difficult to support with the limited computing power on satellites, this centralized training process will be conducted on the ground, simulating real-world operational scenarios by pre-collecting satellite-ground data. Once training is complete, the model will be deployed to each satellite to support distributed decision-making within the onboard system.
[0134] The distributed execution process is described in Algorithm 2.2. For homogeneous multi-satellite systems, during distributed execution, each imaging satellite can be considered to have the same decision-making process; therefore, Algorithm 2.2 only illustrates this process from the perspective of a single satellite. During the execution phase, the satellite's... i The pre-trained model Q was pre-deployed. i After receiving the task set, the satellite sat i Based on model Q i It independently makes decisions for the entire task set. Except for selecting the action with the highest Q-value at each decision step, the rest of the decision-making process is exactly the same as the single-star decision-making process shown in Algorithm 2.1. This design allows the model to be directly applied to distributed execution scenarios after training.
[0135]
[0136] Regarding the on-board distributed autonomous scheduling process, the following points need to be emphasized:
[0137] (1) This process is pre-planned. When the satellite receives the mission set T, it begins to make decisions and formulates a plan for the satellite in the future for a certain period of time. Then it will carry out the observation mission according to the plan.
[0138] (2) Since it is pre-planned, after selecting the satellite to execute the current task in line 7 of Algorithm 2.2, it will not immediately order the corresponding satellite to execute the task, but will assume that the task can be executed smoothly and obtain the set reward.
[0139] (3) Considering communication constraints, it is stipulated that the satellite will only obtain information through inter-satellite communication when it receives the task set, and no further communication activities are involved in the decision-making process in lines 5-8 of Algorithm 2.2. The state transition is designed to be updated based on existing data under the assumption that the action can be executed smoothly.
[0140] (4) Since the information of each satellite is real-time, this decision-making model can ensure that the satellite mission planning scheme is not affected by ground prediction data and strictly meets the actual constraints.
[0141] The training and execution processes in a predetermined scenario are described in [link to documentation]. Figure 4 .
[0142] In the training and evaluation process of the model under 12 satellites and 120 mission scenarios, for simplicity, the scenario configuration will be represented as "number of satellites - number of missions", such as 12-120. There are two indicators for evaluating the performance of the QTRAN algorithm: the cumulative reward during the training phase, represented by formula (2-14), and the overall mission planning objective, represented by formula (1-1). Formula (2-14) can be used to observe the training effect of the algorithm, while formula (1-1) is the final evaluation indicator of interest in this application. To distinguish between the two, in the following text, formula (2-14) will be simply referred to as "total reward," and formula (1-1) will be simply referred to as "total benefit."
[0143] During training, the algorithm records the total reward and total profit obtained in each iteration. Figure 5 (a) shows the curves of total reward and total profit as the training process progresses. Since the model interacts with random scenarios, the curves were smoothed to more clearly observe their overall trend. The smoothness is defined as ξ, representing the average of the data from each ξ generations. In this application's experiments, ξ = 200. It can be seen that the curves converge to a relatively stable state after a period of increase. The total reward and total profit show similar upward trends, indicating consistency between the design of the reward function and the final total profit. For the evaluation process, every 200 generations of training, the current model is used to perform an overall test on the evaluation examples. The curve smoothness is set to ξ = 50, which is equivalent to calculating the average of the entire evaluation set. To more clearly observe the changes in the evaluation results as the training process progresses, Figure 5 (b) When displaying the evaluation curve trend, the horizontal axis is enlarged proportionally to match the number of iterations of the training curve. It can be seen that the evaluation curve rises rapidly in the early stages, then the rate of increase slows down and is accompanied by significant fluctuations. This may be due to random decisions during the exploration process, and the curve eventually reaches a relatively stable state. Unless otherwise specified, subsequent experiments will use the above processing method to demonstrate both the training and evaluation processes.
[0144] Figure 5 (c) and Figure 5 (d) illustrates the changes in the number of repeatedly claimed tasks during training and evaluation. In multi-satellite distributed decision-making based on local observations, repeated task claiming is unavoidable. However, with limited resources, repeated claiming leads to resource waste. To ensure planning effectiveness, the algorithm aims for a task to be claimed only once. Figure 5 (c) and Figure 5 As can be seen in (d), the number of tasks claimed only once showed the most significant change, with a clear upward trend; a small number of tasks were claimed twice, and their corresponding curves showed a slight decrease as the training process progressed; a very small number of tasks were claimed three times, and after training, the number almost reached zero.
[0145] This application addresses the future development needs of satellites for autonomy, intelligence, and rapid response. Considering the current limitations of satellite computing power, simple inter-satellite coordination mechanisms, and unstable communication links, it innovatively proposes a multi-satellite autonomous cooperative scheduling algorithm based on distributed multi-agent reinforcement learning. First, a partially observable Markov decision process model adapted to multi-satellite autonomous cooperative planning is constructed, with each component of the model specifically designed based on actual onboard conditions. Given the limited autonomous planning capabilities of satellites, this application introduces a low-complexity distributed multi-agent reinforcement learning algorithm—QTRAN. This algorithm constructs a separate decision network for each satellite and employs a centralized training and distributed execution framework, suitable for both ground training and onboard execution. In the centralized training phase, collaborative training is conducted by uniting all satellite networks to cultivate the global optimization awareness of each satellite, ensuring the overall mission planning effect. In the distributed execution phase, each satellite makes independent decisions based on its local network, requiring only interaction with a subset of satellites. Finally, this application constructs multi-satellite mission planning scenarios of different scales to verify the algorithm's effectiveness. Simulation results show that the algorithm can effectively reduce inter-satellite communication consumption while ensuring planning effectiveness, and has faster response speed and higher system resilience.
[0146] In another embodiment of the present invention, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to at least one processor; wherein the memory stores instructions executable by at least one processor, the instructions being executed by at least one processor to enable at least one processor to perform the above-described method; for details not disclosed herein, please refer to the method section of the above-described embodiments of the present invention.
[0147] In another embodiment of the present invention, a computer-readable storage medium is provided, wherein computer instructions are stored on the medium, the computer instructions being used to cause the computer to perform the above-described method. For specific technical details not disclosed, please refer to the method section of the above-described embodiments of the present invention.
[0148] The medium in this invention can be any combination of one or more computer-readable media. The medium can be a computer-readable signal medium or a computer-readable storage medium. The medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of the medium (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, the medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0149] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A multi-satellite autonomous cooperative scheduling method based on distributed multi-agent reinforcement learning, specifically including the following steps: S1. Based on target attribute information, satellite attribute information, and various constraints, mission planning is carried out. Mission planning is used to effectively allocate and schedule resources of multiple satellites and formulate satellite observation plans to maximize the completion of user-submitted tasks. S2. Construct a partially observable Markov decision process model adapted to the multi-satellite autonomous collaborative planning mode based on satellite mission planning; Step S2 specifically includes the following steps: Multiple satellites can be viewed as a fully cooperative multi-agent system. This multi-agent system is modeled as a decentralized partially observable Markov decision process (Dec-POMDP), represented by tuples. ;in, Represents the real environment in which the intelligent agent exists; Represents a set of intelligent agents ;No. The action chosen by an agent is represented as: The actions chosen by all agents constitute a joint action. ; Indicates the first Local observations of individual agents; Represents the observation function, ; Indicates the state transition probability; The reward function is represented as follows: In a fully cooperative environment, all agents share the same reward function. ; This is a discount factor used to balance immediate rewards and future rewards; S3. Considering the limited autonomous planning capabilities of satellites, the distributed multi-agent reinforcement learning algorithm QTRAN is used to construct a separate decision network for each satellite; Step S3 specifically includes the following steps: The multi-agent reinforcement learning algorithm VDN based on the value function decomposition idea adopts the CTDE framework, and each agent deploys a decision network. It can learn and construct its own action value function. The idea of value function decomposition is to construct a hybrid network to fit the joint action value function during the centralized training phase. Through training, the action value function of each agent is made more efficient. Value function of joint action The following relationship must be satisfied: 2-1 Formula ( -1) is called the individual-to-collective maximum condition, where Let represent the action-observation history of the agents. When this condition is met, the optimal action chosen by each agent based on its own decision network is equivalent to the joint optimal action of the entire system, thus ensuring the overall system optimality even when each agent makes independent decisions; let , indicating that it is composed of all intelligent agents A vector composed of value functions; when right When the IGM condition is met, it is called yes To construct value decomposition relations that satisfy the IGM, VDN proposes an additive decomposition method: 2-2 This constraint enables value decomposition that satisfies the IGM condition, but it also imposes structural constraints on the problem. Based on this, the QMIX algorithm extends this additive decomposition by proposing a monotonic decomposition method, as shown in equation ( ). -3), thus enabling the fitting of joint Values and each intelligent agent More complex relationships between values; 2-3 The QTRAN algorithm proposes a decomposition method for this, which decomposes the original joint action value function. Transform into a new, more easily decomposable function By ensuring that the joint optimal action of the two is the same, the IGM condition is satisfied; definition , indicating the first The optimal actions of each agent, and the set of optimal actions of all agents, are represented as: The QTRAN algorithm provides... A sufficient condition for satisfying IGM: 2-4 in 2-5 The QTRAN algorithm directly uses the transformation function The definition is as follows: 2-6 formula( -6) Satisfy the demand The IGM condition, and because Therefore, the formula ( -4) It can be regarded as Value decomposition; formula( -4) will be used as the training basis for the QTRAN algorithm, which means and The additive decomposition relationship between them will be expressed by the formula ( -4) is used to represent; during the algorithm training process, three types of functions are involved: the Q-value function of each agent. Joint action value function and functions Due to the transformation function Can be directly from It is stated that, in order to display more clearly and its transformation function The relationship between them, the joint formula ( -3) and the formula ( -6) yields the formula ( -7): 2-7 and The relationship will be through and To fit; during the training process, the QTRAN algorithm introduces an additional function. , It can be considered a correction item, used for correction. and The differences between them can be used to characterize the complex relationships in multi-agent systems; S4. It adopts a centralized training and distributed execution framework for the application mode of ground training and on-board execution. S5. Centralized training phase: Centralized training is conducted on all satellites, and collaborative training is carried out by uniting all satellite networks. S6, Distributed Execution Phase: Each satellite makes independent decisions based on its local network and interacts with only a subset of satellites.
2. The method according to claim 1, characterized in that, The task planning in step S1 includes an objective function and constraints; specifically, the objective function is: 1-1 Constraints: , 1-2 , 1-3 , , , 1-4 , , 1-5 , 1-6 formula( -1) is the objective function, representing maximizing the total reward of the task; where, As a decision variable, when At that time, it indicates that the task will be completed. Arranged for satellite The Execute within a visible time window, at which time... , ,on the contrary Time window benefits The definition is as follows: 1-7 Among them, satellite collection , For the number of satellites, use Indicates a certain satellite, For satellite identification, For satellite Maximum storage capacity within the planning period; To carry out the mission Required satellite storage; For the task Required observation duration; and Tasks The start and end times of the actual observation; For the task Priority; task set , For the number of tasks, Indicates a certain task, For mission identification, satellite For the task Visible time window set , Indicates satellite For the task The number of visible time windows: within one orbital period, a satellite can have at most one visible time window for the same mission; Indicates satellite For the task The One visible time window; and These represent the visible time windows. The start and end times, Indicates time window Observation mission The benefits; Indicates satellite In the time window Execute adjacent tasks , The required attitude transition time; similarly, the following can be obtained. and The object being represented; a mechanism is introduced where the reward decays over time, comprehensively considering both the importance of the task and the need for timely response. As the attenuation factor, For the attenuation coefficient, the formula is ( -2) represents a storage constraint: within the planning period, the storage occupied by the satellite for its missions must not exceed the satellite's maximum storage capacity; formula ( -3) is a unique constraint for task execution, indicating that each task will be executed at most once; formula ( -4) represents the attitude transition time constraint, indicating that the interval between two adjacent tasks performed by the satellite must be greater than its attitude transition time; Formula ( -5) represents the mission's continuous observation time constraint, indicating that the actual observation time of the mission must meet the mission's imaging duration requirement; Formula ( -6) Defines the range of values for the decision variables.
3. The method according to claim 1, characterized in that, Step S2 further includes the following steps: In a multi-satellite autonomous collaborative mission planning system, upon receiving a mission request, the satellites within the system make decisions in a distributed manner, generating their own mission planning schemes. This approach eliminates the need for information sharing across the entire system; each satellite makes independent decisions based solely on its own available information, thus enabling rapid response to requests. Here, the satellite set is represented as... , contains One satellite, Indicates the first One satellite; the set of tasks to be planned is represented as , contains One task, Indicates the first task in the task set One task.
4. The method according to claim 1, characterized in that, Step S4 specifically includes: Under the CTDE framework, each satellite will deploy a decision network. During the training phase, an additional hybrid network will be constructed to assist learning. During the execution phase, each satellite only needs to make decisions independently based on local observation information.
5. The method according to claim 1, characterized in that, Step S5 specifically includes: In order to improve the generalization ability of the algorithm in different scenarios, the training process adopts an algorithm training mechanism oriented to random initial scenarios: several training scenarios are generated, and a scenario is randomly selected for training in each iteration; in the predetermined scenario, each satellite makes its own decision in a distributed manner to obtain its own mission planning scheme; the training data generated by this decision-making process will be put into the experience pool for subsequent algorithm training.
6. The method according to claim 1, characterized in that, Step S6 specifically includes: For a homogeneous multi-satellite system, during distributed execution, each imaging satellite can be considered to have the same decision-making process. During the execution phase, the satellites... The pre-trained model was deployed. After receiving the task set, the satellite Based on the model Independently make decisions for the entire task set.
Citation Information
Patent Citations
Satellite observation distributed online planning method based on multi-agent reinforcement learning
CN113128828A
Multi-agent-based large-scale satellite collaborative observation task planning method
CN117114317A