Spacecraft multi-agent distributed autonomous decision evolution method and system
Patent Information
- Application Number
- CN202611016180.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-07-09
AI Technical Summary
[0004]在环境动态性方面,航天器轨道持续受地球引力、日月摄动等影响,导致轨道高度、倾角、偏心率等参数实时漂移,改变智能体间相对位置与通信条件,同时空间目标的运动状态具有随机性,其轨道参数难以长期精准预测,增加决策实时性要求;在环境不确定性方面,大气阻力随太阳活动与轨道高度波动,空间辐射与太阳风暴等突发事件难以提前预测,且电磁干扰影响范围随环境变化而调整,可能导致通信或感知失效;在扰动类型方面,既存在可通过控制补偿的可控扰动,也存在难以完全补偿的不可控扰动,需通过快速策略调整降低影响;在通信约束方面,航天器多智能体之间的通信链路受空间环境影响显著,存在通信延迟、链路中断、传输误差等问题;上述环境特征叠加作用,使航天器多智能体分布式自治决策面临高动态、高不确定与通信受限条件下难题,对现有决策方法的适应性与鲁棒性提出了更高要求
通过构建规则驱动与LSTM强化学习融合的分布式自组织决策网络,在保证常规操作高可靠性的同时增强对复杂动态扰动环境的自适应能力,提升航天器在轨运行过程中的自主决策水平;
Smart Images

Figure CN122596260B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of spacecraft intelligent control, distributed autonomous decision-making, and multi-agent collaboration. Specifically, it relates to a spacecraft multi-agent distributed autonomous decision-making evolution method, and more specifically, it relates to a spacecraft multi-agent distributed autonomous decision-making evolution method and system under complex dynamic space environment conditions such as orbital disturbance uncertainty, strong space environment disturbance, communication link intermittency / delay, limited on-board computing power and energy, and possible agent failure / failure. Background Technology
[0002] With the rapid iteration of aerospace technology and the continuous advancement of complex aerospace missions such as deep space exploration, on-orbit servicing, and cluster networking, spacecraft systems are gradually transforming from independent single-satellite operation to multi-agent collaborative operation. In this process, spacecraft multi-agents need to complete collaborative decision-making and mission execution through local information interaction in complex and dynamic space environments, placing higher demands on the system's autonomous decision-making capabilities and collaborative control levels.
[0003] The complex and dynamic space environment in which spacecraft multi-agent systems operate has typical characteristics such as high vacuum, strong radiation, uncertain orbital disturbances, and susceptibility to communication link interference. These characteristics are fundamentally different from the operating environment of ground-based multi-agent systems, and they directly determine the complexity and difficulty of distributed autonomous decision-making in spacecraft multi-agent systems.
[0004] In terms of environmental dynamics, spacecraft orbits are continuously affected by Earth's gravity and lunar perturbations, causing parameters such as orbital altitude, inclination, and eccentricity to drift in real time, altering the relative positions and communication conditions between agents. Simultaneously, the motion of space targets is random, making long-term accurate prediction of their orbital parameters difficult, increasing the real-time requirements for decision-making. Regarding environmental uncertainty, atmospheric drag fluctuates with solar activity and orbital altitude; sudden events such as space radiation and solar storms are difficult to predict in advance; and the range of electromagnetic interference changes with environmental variations, potentially leading to communication or sensing failures. In terms of disturbance types, there are both controllable disturbances that can be compensated for and uncontrollable disturbances that are difficult to fully compensate for, requiring rapid strategy adjustments to reduce their impact. Regarding communication constraints, communication links between multiple agents on spacecraft are significantly affected by the space environment, resulting in communication delays, link interruptions, and transmission errors. The combined effect of these environmental characteristics presents challenges to distributed autonomous decision-making by multiple agents on spacecraft under highly dynamic, highly uncertain, and communication-constrained conditions, placing higher demands on the adaptability and robustness of existing decision-making methods.
[0005] Existing multi-agent decision-making methods for spacecraft mostly rely on centralized scheduling or a single decision-making mechanism, which has the following shortcomings: First, most decision-making models adopt a single rule-driven or single reinforcement learning architecture, making it difficult to balance the high reliability of routine operations with adaptability to dynamic environments. The autonomous decision-making capability of a single satellite is limited, and it cannot effectively cope with the dynamic uncertainties of complex space environments. Second, the cluster consensus building mechanism is imperfect and does not fully consider the needs of multi-objective optimization. In extreme scenarios such as communication delays and agent failures, it is difficult to achieve efficient collaborative consensus among multiple agents, affecting mission execution efficiency. Third, the decision-making model lacks an effective real-time iterative evolution mechanism, with insufficient generalization and online update capabilities. It cannot adjust decision-making strategies according to changes in the space environment and mission requirements, making it difficult to adapt to the operational requirements of long-term and complex space missions.
[0006] Therefore, how to provide a distributed autonomous decision-making method for complex and dynamic space environments, which can enhance dynamic adaptability while ensuring high reliability of routine operations, and achieve efficient collaborative consensus and online strategy evolution in scenarios with limited communication and agent failures, so as to support the long-term stable and autonomous operation of spacecraft multi-agent systems, is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0007] In view of the above problems, the present invention aims to provide a spacecraft multi-agent distributed autonomous decision-making evolution method and system that overcomes or at least partially solves the above problems.
[0008] To achieve the above objectives, the present invention adopts the following technical solution:
[0009] A spacecraft multi-agent distributed autonomous decision-making evolution method includes: S1. Acquire spacecraft dynamics status, environmental disturbance information, resource status and communication link status and perform feature analysis, establish orbital dynamics model, space disturbance model and communication link model, and output environmental model parameters and constrained characterization; S2. Construct a unified decision input state vector, input a rule-driven-LSTM fusion distributed self-organizing decision network, fuse rule-driven and LSTM reinforcement learning to generate multi-source candidate decisions, and correct decisions through distributed conflict avoidance and negotiation mechanisms, outputting the local decision strategies of each spacecraft agent in the current decision cycle; S3. A cluster consensus building mechanism based on multi-agent reinforcement learning is established by clarifying the core objective and key decision variables of cluster consensus, generating local decisions, and introducing a multi-objective optimization and differentiated reward function system for consensus negotiation. The consensus verification outputs a global collaborative decision-making strategy and a set of corresponding key decision variables that meet the consistency threshold constraint. S4. Through a real-time iterative evolution mechanism of the decision model based on a centralized training and distributed execution CTDE architecture, combined with experience sharing, knowledge distillation and parameter synchronization technologies, the decision network and cluster consensus building mechanism are centrally trained, distributed and synchronized, and deployed in a lightweight manner on satellite.
[0010] Preferably, the space disturbance model includes the lunar perturbation acceleration model, the atmospheric drag acceleration model, and the space electromagnetic interference model; the communication link model includes the communication delay model and the communication link interruption model; The orbital dynamics model outputs orbital state variables and relative motion parameters; the space disturbance model outputs various disturbance intensity estimates; the communication link model outputs communication delay parameters, including signal transmission delay and signal processing delay; and the communication link interruption model obtains the link interruption probability.
[0011] Preferred, the first A spacecraft intelligent agent at any time The decision input state vector is:
[0012] in, For the first A spacecraft intelligent agent at any time The position vector in the geocentric inertial coordinate system or orbital coordinate system. For velocity vector, Relative motion For relative velocity, This represents the intensity characteristics of environmental disturbances at the current moment. Due to signal propagation delay, This represents the probability of link interruption. For mission requirements, This represents the remaining resources; No. The local decision-making strategy output of each agent is:
[0013] in, These are task-level actions, including execution, postponement, handover, and reordering of instructions. For orbit control layer actions, including orbit adjustment direction and amplitude, maneuver window selection, and attitude maneuver trigger commands; For resource allocation actions, For communication and coordination actions.
[0014] Preferably, step S2 includes the following: S21. Taking track dynamics, disturbance intensity index and link quality index as input, the system uses a production rule base to make deterministic judgments on safety constraints, track maintenance, task priority and resource limitations, and outputs candidate rule actions that meet the engineering safety boundary and conventional operating procedures. S22. Based on historical state sequence Input the LSTM-DQN network, use LSTM to perform time-series encoding of orbit drift trend, disturbance evolution trend and link delay fluctuation, and combine DQN output to adaptively learn candidate actions in dynamic and uncertain scenarios; S23. Receive rule-based candidate actions and adaptive candidate actions, and determine the fusion weight based on the comprehensive disturbance strength, link constraint index and task urgency. Use the action consistency mapping function to map the two types of actions to a unified action vector space to form a fused action. S24. Based on the fusion action, combined with the short-term prediction output of the orbital dynamics model and the output of the communication link model, local intention interaction, conflict detection and negotiation correction are performed within the neighborhood topology, and the final local decision-making strategy is output.
[0015] Preferably, the specific content of step S21 is as follows: In the state determination phase, the output results of the troop trajectory dynamics model, spatial disturbance model and communication link model are converted into rule triggering condition variables. At the same time, task and resource condition variables are constructed by combining the task system and platform state information to form the set of condition variables required for rule triggering, and further mapped into discretized condition labels through threshold determination. During the rule matching phase, the condition variables are logically matched based on the production rule base; When multiple rules are triggered simultaneously and their actions conflict, the priority conflict resolution phase begins. A unified rule scoring mechanism is established to rank the candidate rules, and the final rule candidate action is selected based on the scoring results. The LSTM-DQN network in step S22 includes an input layer, an LSTM encoding layer, a feature mapping layer, and a Q-value output layer; The input layer receives a sequence of historical states. The LSTM encoding layer performs temporal encoding on the historical states and outputs the hidden states. It is used to characterize orbital change trends, disturbance evolution trends, and link state evolution trends; the feature mapping layer will... Mapped to decision feature vector ; The Q-value output layer outputs a set of discrete actions. The value of each action It then performs a value assessment and finally outputs adaptive learning candidate actions; The specific content of step S23 is as follows: Based on the model output in step S1, construct the scenario discrimination input, including the comprehensive disturbance strength index, link constraint index, and task urgency index; Using the scene discrimination input as fuzzy input, three levels of fuzzy sets (low, medium, and high) are constructed respectively. The system state of the current decision cycle is represented in a fuzzy manner through the fuzzy membership function. Inference calculations are performed based on a pre-designed fuzzy rule base to determine the fusion weight of rule candidate actions and adaptive candidate actions in the current cycle; After the fuzzy rule reasoning is completed, the rule weights and learning weights are obtained by defuzzification calculation; The rule-based candidate actions and adaptive candidate actions are mapped to a unified action vector space, and the fused action is output by combining the rule weights and the learned weights. The specific content of step S24 is as follows: The agent encapsulates the key quantities required for fusion actions and conflict determination into neighborhood messages, and sends neighborhood decision intentions based on a communication sending strategy constrained by the communication link model. After receiving the decision intentions of the neighborhood, the neighboring agents perform task conflict detection, track conflict detection, and resource conflict detection. If a conflict is detected, a negotiation adjustment is performed in the neighborhood based on the local negotiation evaluation function to maximize the benefits of high-priority tasks and reduce the overall cost while satisfying safety constraints. At the same time, the LSTM module is used to predict the conflict development trend and generate conflict avoidance strategies in advance.
[0016] Preferably, the specific content of step S3 is as follows: S31. Define the core objectives and key decision variables of the cluster consensus. The core objectives include maximizing task completion rate, maximizing resource utilization, and minimizing decision delay. The key decision variables include task allocation scheme, track adjustment strategy, and resource allocation ratio. At the same time, initialize the parameters, reward function weights, and consensus threshold of the multi-agent reinforcement learning model, and set the communication range and interaction frequency of the agents. S32. Obtain the local decision-making strategies generated by each intelligent agent, and at the same time collect its own state, task requirements and local environment information. Send the local decision-making strategies and key states to neighboring intelligent agents through distributed communication to achieve local information sharing. S33. Each agent receives local decision information from neighboring agents, adjusts decision variables and conducts negotiations by combining its own strategy with the global objective through a multi-agent reinforcement learning model; introduces a delay compensation mechanism to predict the trend of decision changes of neighboring agents; introduces a fault detection and replacement mechanism to identify faulty agents and assign tasks to neighboring healthy agents. S34. Based on the preset consensus threshold, verify the consistency of each agent on key decision variables. If the consistency error is less than or equal to the consensus threshold, it is considered that a cluster consensus has been reached and a global collaborative decision strategy is output. If no consensus is reached, return to the consensus negotiation stage and continue to adjust the decision variables until a consensus is reached or the maximum number of negotiations is reached, and a suboptimal decision strategy is output.
[0017] Preferably, the specific content of step S33 is as follows: S331. After each agent receives the neighborhood message set, it first performs time alignment of the neighbor state and decision variable according to the communication delay in the message to obtain the compensated state and compensated decision variable. S332. A multi-agent deep deterministic policy gradient algorithm driven by delay compensation and consensus is adopted. Delay condition variables are explicitly introduced to construct a delay condition value network, and the joint state after delay compensation is used for evaluation. After the policy network outputs the update amount of the decision variable, local update is performed. S333. After completing the local update, each agent and its neighboring agents perform weighted consensus fusion, based on the weights determined by communication latency and link reliability, to obtain the negotiated decision variables; S334. Design a delay adaptive adjustment mechanism to dynamically adjust the consensus negotiation frequency and the step size of the decision variable adjustment based on the average delay of the neighborhood. At the same time, use the LSTM network to predict the decision changes of neighboring agents, correct the policy output update, and adjust its own decision policy in advance. S335. Based on the state feedback information of each agent, a threshold detection method is used to identify faulty agents and isolate them; at the same time, based on the task load and decision-making ability of neighboring agents, a replacement agent is determined through a comprehensive evaluation function. S336. After determining the replacement agent, the policy parameters of the faulty node are migrated to the replacement agent in a soft update manner through policy inheritance and fast recovery mechanism. The policy distillation loss function is introduced to constrain the policy output of the replacement node to be consistent with the policy of the original faulty node. S337. Transform multiple objectives into a single objective using a weighted summation method, and design a differentiated reward function system that combines the global reward based on the overall performance of the cluster with the local reward based on individual performance.
[0018] Preferably, the specific content of step S4 is as follows: S41. Gain experience in local decision-making, consensus-building and negotiation, and the execution process; S42. Introduce an experience screening and priority management mechanism for global experience storage, and divide the experience data into three data sets according to function: local decision-making experience pool, consensus negotiation experience pool, and evaluation feedback buffer. S43. After acquiring global experience data, the decision model is periodically trained and optimized. The training process includes at least local learning model training, consensus negotiation model training, and fusion module parameter calibration. After the model training is completed, the updated model parameters are distributed to the local models of each agent through the parameter synchronization mechanism. S44. The local learning model and consensus negotiation model obtained through centralized training are compressed into a lightweight model suitable for satellite operation through knowledge distillation technology; S45. Each model is updated through a real-time iterative update mechanism that combines active and passive triggering.
[0019] A spacecraft multi-agent distributed autonomous decision-making evolution system, based on the aforementioned spacecraft multi-agent distributed autonomous decision-making evolution method, includes: a space environment feature analysis and modeling module, a rule-driven-LSTM fusion distributed self-organizing decision network, a cluster consensus mechanism based on multi-agent reinforcement learning, and a real-time iterative evolution mechanism for decision models based on CTDE architecture. The space environment feature analysis and modeling module is configured to acquire spacecraft dynamics status, environmental disturbance information, resource status and communication link status and perform feature analysis, establish orbital dynamics model, space disturbance model and communication link model, and output environmental model parameters and constrained characterization. The rule-driven and LSTM-fused distributed self-organizing decision network is configured to construct a unified decision input state vector. It generates multi-source candidate decisions by fusing rule-driven and LSTM reinforcement learning, and corrects the decisions through a distributed conflict avoidance and negotiation mechanism, outputting the local decision strategies of each spacecraft agent in the current decision cycle. The cluster consensus mechanism based on multi-agent reinforcement learning is configured to generate local decisions by clarifying the core objectives and key decision variables of the cluster consensus, and to introduce a multi-objective optimization and differentiated reward function system for consensus negotiation. The consensus verification outputs a global collaborative decision strategy and a set of corresponding key decision variables that meet the consistency threshold constraint. The real-time iterative evolution mechanism of the decision model based on the CTDE architecture is configured to perform centralized training, distributed execution of the decision model based on the CTDE architecture, combined with experience sharing, knowledge distillation and parameter synchronization technologies, to centrally train, distribute and synchronize the decision network and cluster consensus building mechanism and deploy it in a lightweight manner on the satellite.
[0020] Preferably, the rule-driven-LSTM fusion distributed self-organizing decision network includes: a rule-driven module, an LSTM reinforcement learning module, a decision fusion module, and a distributed self-organizing and conflict avoidance module; The cluster consensus mechanism based on multi-agent reinforcement learning includes: a consensus goal and variable definition module, a local decision generation module, a consensus negotiation module, and a consensus verification module; The real-time iterative evolution mechanism of the decision model based on the CTDE architecture includes: a centralized training module, a distributed execution module, an experience storage module, a knowledge distillation module, and an evaluation and feedback module.
[0021] As can be seen from the above technical solution, compared with the prior art, the present invention discloses a spacecraft multi-agent distributed autonomous decision-making evolution method and system, which has the following beneficial effects: By constructing a distributed self-organizing decision network that integrates rule-driven and LSTM reinforcement learning, the adaptive capability to complex dynamic disturbance environments is enhanced while ensuring high reliability of routine operations, thereby improving the autonomous decision-making level of spacecraft during on-orbit operation. By introducing a consensus building mechanism based on multi-agent reinforcement learning, and combining it with multi-objective optimization and delay compensation strategies, multi-agents can still achieve consensus negotiation on key decision variables under limited conditions such as communication delay, link interruption and individual failure, thereby improving the efficiency and stability of cluster collaborative decision-making. Based on the centralized training and distributed execution CTDE architecture, a real-time iterative update mechanism for the model is built. Through experience sharing, incremental training and knowledge distillation, the decision model is continuously optimized and updated online, thereby enhancing the model's generalization ability and long-term operational stability. Construct a distributed autonomous decision-making closed-loop system adapted to complex space environments and on-board resource constraints. Through an integrated design of environment modeling, individual decision-making, cluster consensus, and model evolution, a complete distributed autonomous decision-making closed-loop structure is formed without relying on a central node, thereby improving the overall coordination, robustness, and engineering feasibility of the system in complex and dynamic space environments. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0023] Figure 1 This is a schematic diagram of a spacecraft multi-agent distributed autonomous decision-making evolution method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the distributed self-organizing decision network structure provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the cluster consensus building mechanism provided in this embodiment of the invention; Figure 4 This is a schematic diagram of the real-time iterative evolution mechanism of the decision model provided in this embodiment of the invention; Figure 5 This is a schematic diagram of the co-simulation platform provided in the embodiments of the present invention; Figure 6 This is a schematic diagram comparing simulation results provided in the embodiments of the present invention. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] This invention discloses a spacecraft multi-agent distributed autonomous decision-making evolution method, such as... Figure 1 ,include: S1. Acquire spacecraft dynamics status, environmental disturbance information, resource status and communication link status and perform feature analysis, establish orbital dynamics model, space disturbance model and communication link model, and output environmental model parameters and constrained characterization; S2. Construct a unified decision input state vector, input a rule-driven-LSTM fusion distributed self-organizing decision network, fuse rule-driven and LSTM reinforcement learning to generate multi-source candidate decisions, and correct them through a distributed conflict avoidance and negotiation mechanism, outputting the local decision strategy of each spacecraft agent in the current decision cycle; S3. A cluster consensus building mechanism based on multi-agent reinforcement learning is established by clarifying the core objective and key decision variables of cluster consensus, generating local decisions, and introducing a multi-objective optimization and differentiated reward function system for consensus negotiation. The consensus verification outputs a global collaborative decision-making strategy and a set of corresponding key decision variables that meet the consistency threshold constraint. S4. Through a real-time iterative evolution mechanism of the decision model based on a centralized training and distributed execution CTDE architecture, combined with experience sharing, knowledge distillation and parameter synchronization technologies, the decision network and cluster consensus building mechanism are centrally trained, distributed and synchronized, and deployed in a lightweight manner on satellite.
[0026] To further implement the above technical solutions, the space disturbance model includes the lunar perturbation acceleration model, the atmospheric drag acceleration model, and the space electromagnetic interference model; the communication link model includes the communication delay model and the communication link interruption model. The orbital dynamics model outputs orbital state variables and relative motion parameters; the space disturbance model outputs various disturbance intensity estimates; the communication link model outputs communication delay parameters, including signal transmission delay and signal processing delay; and the communication link interruption model obtains the link interruption probability.
[0027] In this embodiment, the orbital dynamics model is used to describe the orbital motion of the spacecraft intelligent agent and quantify the dynamic drift characteristics of orbital parameters. Based on the circular orbit assumption, and considering the three main perturbation factors—Earth's gravity, lunar and solar perturbations, and atmospheric drag—the spacecraft orbital dynamics model is constructed as follows:
[0028]
[0029] in, Let be the position vector of the spacecraft relative to the Earth's center of mass. For the spacecraft's velocity vector, The gravitational constant of Earth, The acceleration vector is the environmental disturbance. The control acceleration vector of the spacecraft propulsion system; Space disturbance models are used to quantify the intensity and impact of various environmental disturbances. For the three types of uncontrollable disturbances that have the greatest impact on spacecraft decisions—lunar and solar perturbations, atmospheric drag, and space electromagnetic interference—corresponding disturbance models are constructed to achieve a quantitative description of the disturbance impact. The solar-moon perturbation acceleration model is expressed as:
[0030] in, The gravitational constant of the Sun / Moon. This is the position vector of the Sun / Moon relative to the Earth's center of mass; The atmospheric drag acceleration model is constructed piecewise based on orbital altitude and is expressed as follows:
[0031] in, This refers to atmospheric density (which varies with orbital altitude). This is the velocity vector of the spacecraft relative to the atmosphere. The drag coefficient, For the windward area of the spacecraft; The space electromagnetic interference model adopts a random interference model, assuming that the impact of electromagnetic interference on communication signals follows a Gaussian distribution and that the interference intensity fluctuates randomly over time. Its interference model can be expressed as:
[0032] in, The original communication signal, It is Gaussian white noise. It is used to quantify the impact of electromagnetic interference on the information interaction of intelligent agents after the communication signal is interfered with, and to provide support for fault-tolerant decision-making of communication links; By combining the constraints of space communication, a local communication link model between multiple agents in a spacecraft is constructed to quantify the impact of communication delay, link interruption and transmission error on distributed decision-making. The communication delay model adopts a dynamic delay model, which considers the impact of changes in orbital position on communication delay. Due to signal transmission delay Signal processing delay It consists of two parts, represented as:
[0033]
[0034] in, The straight-line distance between the two agents is denoted as . The speed of light; The communication link interruption model uses a binomial distribution model, assuming that the working state of the communication link is divided into two types: normal and interrupted, and the probability of link interruption is... The link adjusts dynamically based on changes in environmental electromagnetic interference intensity and orbital position. When the electromagnetic interference intensity exceeds a threshold or the distance between the two agents exceeds the maximum communication distance, the link interruption probability increases. Significant improvements have been made. This model can simulate the random interruption characteristics of communication links, providing support for the fault-tolerant design of decision models.
[0035] In order to further implement the above technical solution, the first A spacecraft intelligent agent at any time The decision input state vector is:
[0036] in, For the first A spacecraft intelligent agent at any time The position vector in the geocentric inertial coordinate system or orbital coordinate system. For velocity vector, Relative motion , For relative velocity, It is used to characterize the relative geometric relationships and potential conflict risks of spacecraft in a cluster formation; Based on the current environmental disturbance intensity characteristics, calculate the current environmental disturbance acceleration according to the spatial disturbance model in step S1. (Including lunar and solar perturbations and atmospheric drag), and extracting perturbation intensity characteristics through perturbation amplitude or comprehensive perturbation index to reflect the degree of impact of environmental uncertainty on decision-making; To account for signal propagation delay, based on the communication link model established in step S1, the distance between spacecraft is considered. Based on the calculation of signal propagation delay, The link interruptibility probability is obtained by combining the link interruption probability model. This allows for the construction of communication state characteristics; For mission requirements, To provide resource margin, task-related information and resource status information (task requirements, resource margin) are integrated with track status, relative motion status, disturbance intensity, and communication link status to form the input state vector of the decision network. No. The local decision-making strategy output of each agent is:
[0037] in, These are task-level actions, including execution, postponement, handover, and reordering of instructions. For orbit control layer actions, including orbit adjustment direction and amplitude, maneuver window selection, and attitude maneuver trigger commands; This includes resource allocation actions such as power allocation ratio, computing resource allocation ratio, and load time slot adjustment. For communication coordination actions, such as sending / buffering / retransmitting, neighbor selection, and negotiating frequency settings.
[0038] To further implement the above technical solutions, such as Figure 2 The specific content of step S2 includes: S21. In terms of orbital dynamics ( , , ), disturbance intensity index and link quality indicators ( , Taking ) as input, the system performs deterministic determinations on safety constraints, track maintenance, task priority, and resource limitations through a production rule base, and outputs candidate rule actions that meet the engineering safety boundaries and routine operating procedures.
[0039] in, Used to ensure the interpretability, verifiability, and consistency of decisions made under both normal operating conditions and safety risk conditions; S22. Based on historical state sequence Input to an LSTM-DQN network, where each All of them include the orbital dynamics output, disturbance model output and communication link model output of step S1. LSTM is used to perform time-series encoding of orbital drift trend, disturbance evolution trend and link delay fluctuation. Combined with DQN output, candidate actions are adaptively learned in dynamic and uncertain scenarios.
[0040] in, It is used to cope with non-stationary operating conditions such as task changes, enhanced disturbances, and intermittent link interruptions, and improves the robustness of local decision-making and real-time adaptive capabilities. S23. Receive rule candidate actions With adaptive candidate actions And based on the overall strength of the disturbance Link constraint indicators and task urgency Determine the fusion weights , and Using an action-consistent mapping function The two types of actions are mapped to a unified action vector space to form a fused action.
[0041]
[0042] in, It reflects a trade-off between rule reliability and learning adaptability, and ensures that safety-related action components prioritize meeting engineering boundaries through weight constraints; S24. Using fusion movements Based on this, combined with the short-time prediction output of the orbital dynamics model and communication link model output , In the neighborhood topology Within the system, local intent interaction, conflict detection, and negotiation correction are performed to output the final local decision-making strategy. Intent interaction refers to sending decision intent messages. It includes key fields such as task time window, predicted track occupancy, and priority; conflict detection is based on minimum relative distance prediction. Track conflicts are determined by identifying task conflicts based on the intersection of task time windows and resource conflicts based on resource budgets. Negotiation and correction are then implemented to prioritize high-priority tasks and reduce overall costs while meeting safety constraints. Make local adjustments and output the final local decision strategy. .
[0043] To further implement the above technical solution, the specific content of step S21 is as follows: During the state determination phase, the outputs of the trajectory dynamics model, spatial disturbance model, and communication link model are converted into rule-triggered condition variables. Simultaneously, task and resource condition variables are constructed by combining task system and platform state information, forming the set of condition variables required for rule triggering. Furthermore, it is mapped to discretized condition labels through threshold determination. The discretized condition labels include: communication latency level. With disturbance level The specific threshold is given by engineering calibration or task planning parameters; Specifically, the spacecraft position and velocity information output by the orbital dynamics model is used to calculate orbital safety indicators such as relative distance, relative velocity, and orbital drift rate, and compared with preset safety thresholds to form orbital safety condition variables; the atmospheric drag disturbance, lunar perturbation intensity, and space electromagnetic interference intensity output by the space disturbance model are further discretized into disturbance level condition variables to describe the uncertainty level of the current space environment; the parameters such as communication delay, link interruption probability, and bit error rate output by the communication link model are mapped into link quality condition variables to characterize the stability of distributed cooperative communication. During the rule matching phase, the condition variables are logically matched based on the production rule base; The rules are uniformly represented as:
[0044] in, It is a Boolean conditional expression. Output the rule-based action; The rule base is constructed based on the spacecraft cluster operation scenario and includes three categories: safety constraint rules, mission execution rules, and communication coordination rules. Safety constraint rules have the highest priority and are used to ensure the safety of spacecraft operation. When the relative distance is less than the safety threshold and the relative velocity points close together, evasive maneuvers are triggered and unnecessary mission execution is restricted. When the orbital drift rate exceeds the threshold, orbit holding or orbit correction operations are prioritized. When the propellant balance or battery power is lower than the safety threshold, high-energy maneuvers are restricted. Mission execution rules are mainly used to ensure mission completion efficiency. When the mission deadline is approaching and the communication link quality is good, local missions are prioritized and the communication reporting frequency is increased. When the mission has transfer conditions and the neighboring agent has higher execution capabilities, mission handover procedures are generated. Communication coordination rules are used to adapt to changes in link status. When the communication latency is high and the probability of link interruption is high, a low-frequency negotiation and buffered transmission mode is switched. When the bit error rate increases, a retransmission mechanism is initiated or a simplified communication message structure is adopted to reduce the communication burden and improve coordination stability. When multiple rules are triggered simultaneously and their actions conflict, the priority conflict resolution phase begins. A unified rule scoring mechanism is established to rank the candidate rules, and the final rule candidate action is selected based on the scoring results. The rule scoring function is defined as follows:
[0045] in, This indicates the rule priority level; security-related rules have higher priority. This indicates an estimate of task utility gain, which can be obtained from the improvement in task completion or the time efficiency gain. This indicates the resource cost required to perform the action, such as propellant consumption, energy consumption, and communication overhead; parameters , , The weighting coefficient is used to balance the relationship between security, task efficiency, and resource consumption. When multiple rules are triggered, the system first filters them according to a three-level priority principle: security first, task timeliness first, and resource efficiency first. Within the same level of rules, they are sorted according to the scoring function, and finally the rule action with the highest score is selected as the rule candidate output.
[0046] In this embodiment, to ensure that the rule base can adapt to environmental changes and task adjustments during long-term task operation, a dynamic update mechanism for the rule base is designed. A constrained update strategy is adopted, which only allows adjustments to the threshold parameters, action amplitude parameters, and task priority mapping relationships in the rules, while the core structure of the safety constraint rules remains unchanged, so as to ensure the engineering verifiability and safety reliability during system operation. The LSTM-DQN network in step S22 includes an input layer, an LSTM encoding layer, a feature mapping layer, and a Q-value output layer; The input layer receive length is Historical state sequence Each of them Each includes orbital state, perturbation characteristics, link quality, mission state, and resource state; the LSTM encoding layer performs temporal encoding on historical states and outputs hidden states. It is used to characterize orbital change trends, disturbance evolution trends, and link state evolution trends; the feature mapping layer will... Mapped to decision feature vector ; The Q-value output layer outputs a set of discrete actions. The value of each action It then performs a value assessment and finally outputs adaptive learning candidate actions; The set of discrete actions is predefined according to the multi-agent mission execution scenario of the spacecraft. Among them, mission-related actions include executing tasks, delaying tasks, handing over tasks, and reordering tasks; orbit-related actions include maintaining orbit, radial fine-tuning, tangential fine-tuning, and delayed maneuvers; resource-related actions include increasing navigation computing resources, reducing payload ratio, and switching energy-saving modes; and communication-related actions include high-frequency negotiation, low-frequency negotiation, buffered transmission, and priority retransmission. The output adaptive learning candidate actions are:
[0047] The local reward function for adaptive learning is:
[0048] in, This represents the task reward, calculated based on task completion rate, task priority, and task completion time. The safety benefit item is given a positive reward when the agent maintains a safe relative distance and avoids high-risk maneuvers, and a negative reward when the agent violates safety constraints. This represents a positive reward for communication adaptation benefits, which is given when strategies and actions that adapt to communication constraints are taken in environments with high latency or high link loss probability. The resource efficiency benefit term is given a positive reward when the decision action reduces energy consumption, propellant consumption, or balances the computational load while ensuring the completion of the task. The security constraint information, link state information, and disturbance environment information in the reward function are all derived from the output of the space environment model established in step S1, thereby ensuring that the reinforcement learning strategy is always constrained by the engineering environment model during the optimization process, so that the learning results can be consistent with the actual space environment conditions. In the process of policy training and updating, an experience reuse mechanism is also introduced to improve learning efficiency and reduce resource consumption caused by repeated trial and error among different agents. Each agent generates an experience entry after performing an action:
[0049] in, This is the current state. In order to perform the action, As a reward value, For the state at the next moment, Experience tags are used to identify the environmental conditions and task type of the experience sample, including at least a disturbance level tag, a link quality tag, and a task type tag. When communication link conditions permit, agents share experience entries with neighboring agents to expand sample sources and improve learning efficiency. Specifically, the system only performs experience sharing when the communication link model output satisfies the communication latency being less than the sharing threshold and the link interruption probability being lower than the sharing threshold, to avoid increasing communication burden under unstable link conditions. Simultaneously, the system prioritizes sharing high-reward experiences and experience samples obtained in rare scenarios such as strong disturbances or high communication latency, enabling neighboring agents to quickly acquire adaptability to extreme environments. Experience sharing is typically accomplished through neighborhood broadcasting or point-to-point transmission, without requiring global synchronization among all agents, thus ensuring effective experience reuse and policy improvement even under distributed communication conditions. The specific content of step S23 is as follows: Based on the model output in step S1, construct the scenario discrimination input, including the comprehensive disturbance strength index, link constraint index, and task urgency index;
[0050] in, This represents the comprehensive intensity index of the disturbance, which is obtained by weighting the intensity of lunar and solar perturbations, atmospheric drag, and electromagnetic interference. The link constraint index is obtained by weighting the outputs of the communication delay model and the communication link interruption model. This indicator represents the urgency of a task and is calculated based on task priority. Determine input volume based on scenario , , For fuzzy input, three levels of fuzzy sets (low, medium, and high) are constructed respectively, and the system state of the current decision cycle is represented in a fuzzy manner through fuzzy membership functions; Inference calculations are performed based on a pre-designed fuzzy rule base to determine the fusion weight of rule candidate actions and adaptive candidate actions in the current cycle; The fuzzy rule base is constructed by combining the relationship between spatial environment stability, communication conditions, and task requirements. When environmental disturbances are small, communication links are stable, and task urgency is at a low to medium level, the system tends to adopt a rule-driven decision-making mode to ensure the stability and interpretability of the decisions. In this case, the rule weights are adjusted accordingly. Improve, learning weight Correspondingly, the system uncertainty increases when environmental disturbances intensify or link constraints significantly increase. The learning model then demonstrates greater adaptability in complex dynamic environments, thus fuzzy inference is used to increase the learning weights. This enhances the role of data-driven decision-making; when the task is urgent and there are potential security risks, a safety lower limit is set for the rule weights during the inference process to ensure the system's operational security. To avoid learning strategies from making unstable decisions in extreme situations; After fuzzy rule reasoning is completed, the rule weights are obtained by defuzzification calculation. With learning weights ; The rule weights and learning weights satisfy the following constraints:
[0051] To ensure the normalization of the weights for the two types of decisions; in practical applications, the weight range can be further limited according to different operating scenarios: in normal scenarios with stable environments and good communication conditions, the system is dominated by rule-based decision-making, and typically takes the weight range of 100%. Learning weights In dynamic and uncertain scenarios with significant disturbances or unstable communication, increasing the learning decision weights typically involves taking... , In mixed scenarios where both environmental and task conditions are of moderate complexity, the two types of decisions work together, with their weights generally remaining within a certain range. Within the scope; through fuzzy reasoning mechanism, the system can adaptively adjust the contribution ratio of rule-based decision-making and learning-based decision-making according to environmental changes, thereby improving the decision-making flexibility and adaptability in complex dynamic environments while ensuring safety and reliability. Candidate actions for rules With adaptive candidate actions Perform action unification mapping to map the two to a unified action vector space, and output the fused action by combining rule weights and learned weights;
[0052]
[0053] in, This is an action mapping function used to convert task / orbit / resource / communication actions into a unified local decision vector representation; when and When there is a conflict in security constraints, the security constraint action component in the rule module is retained first, and then the non-security components are weighted and fused. The specific content of step S24 is as follows: The intelligent agent will integrate actions The key quantities required for conflict determination are encapsulated as neighborhood messages. And based on the communication transmission strategy constrained by the communication link model, it sends the neighborhood decision intent;
[0054] in, For task execution time window / slot prediction, This refers to the short-term orbital occupancy interval or relative position prediction interval calculated from the orbital dynamics model. Prioritize tasks; , These represent the latency and link reliability output by the communication link model, respectively. The communication transmission strategy is to buffer and retransmit when the probability of link interruption is high; After receiving the decision intentions of the neighborhood, the neighboring agents perform task conflict detection, track conflict detection, and resource conflict detection. Task conflict detection determines a conflict when two agents propose actions to execute on the same task object within the same time window, and the tasks cannot be executed concurrently. The determination function is as follows:
[0055] Orbital conflict detection utilizes the relative motion prediction output from the orbital dynamics model in step S1, and calculates the minimum relative distance within a short-time prediction window. ;when Timely determination of orbital conflict;
[0056] Resource conflict detection determines a resource conflict when the overlap between an ontology's planned action and a neighboring cooperative action causes power / computing / communication resources to exceed the allocation limit; the comprehensive conflict flag is defined as follows:
[0057] If a conflict is detected, a negotiation adjustment is performed in the neighborhood based on the local negotiation evaluation function to maximize the benefits of high-priority tasks and reduce the overall cost while satisfying safety constraints; at the same time, the LSTM module is used to predict the conflict development trend and generate conflict avoidance strategies in advance. The local negotiation evaluation function is:
[0058] in, For task rewards, tasks with high priority and high timeliness yield higher rewards; To ensure safety, the higher value should be used when meeting track safety distance requirements and avoiding high-risk maneuvers. The costs are adjusted, including maneuver costs, energy consumption, communication overhead, and mission delay penalties; The negotiation adjustment rules include: prioritizing security constraints; prioritizing high-priority tasks when security constraints are met; delaying, transferring, or downgrading the execution of low-priority tasks; and adopting a conservative strategy to reduce the orbit adjustment range and the number of negotiation rounds if the link constraints are severe, i.e., high latency / high interruption probability. The collision avoidance method is as follows: LSTM timing coding capability is reused to perform short-time predictions of neighborhood relative motion trends, disturbance change trends, and link quality change trends, outputting a collision risk prediction value.
[0059] when In this case, predictive avoidance actions are triggered before the actual occurrence of the conflict. These predictive avoidance actions include at least: adjusting the task time slot in advance, switching the communication negotiation mode in advance, and performing minor orbital corrections in advance.
[0060] To further implement the above technical solutions, such as Figure 3 The specific content of step S3 is as follows: S31. Define the core objectives and key decision variables of the cluster consensus. The core objectives include maximizing task completion rate, maximizing resource utilization, and minimizing decision delay. The key decision variables include task allocation scheme, track adjustment strategy, and resource allocation ratio. At the same time, initialize the parameters, reward function weights, and consensus threshold of the multi-agent reinforcement learning model, and set the communication range and interaction frequency of the agents. In this embodiment, the following parameters are further specified: Define the key decision variable vector as follows:
[0061] in, This represents a vector representing the proportion of tasks allocated. This represents the vector of orbital adjustment parameters. Represents a vector of resource allocation proportions; The consensus error is:
[0062] in, For weighted average decision variables, For agent weights; consensus threshold is The maximum number of negotiations is The communication cycle is ; Communication delay threshold is The fault judgment threshold is Predicting residual threshold ; S32. Obtain the local decision-making strategies generated by each intelligent agent, and at the same time collect its own state, task requirements and local environment information. Send the local decision-making strategies and key states to neighboring intelligent agents through distributed communication to achieve local information sharing. No. An intelligent agent at time The message sent is:
[0063] in, This is a local observation state. For local actions, For local decision variables, This is an estimate of the communication delay. For link reliability; S33. Each agent receives local decision information from neighboring agents, adjusts decision variables and conducts negotiations by combining its own strategy with the global objective through a multi-agent reinforcement learning model; introduces a delay compensation mechanism to predict the trend of decision changes of neighboring agents; introduces a fault detection and replacement mechanism to identify faulty agents and assign tasks to neighboring healthy agents. S34. Based on the preset consensus threshold, verify the consistency of each agent on key decision variables. If the consistency error is less than or equal to the consensus threshold, it is considered that a cluster consensus has been reached and a global collaborative decision strategy is output. If no consensus is reached, return to the consensus negotiation stage and continue to adjust the decision variables until a consensus is reached or the maximum number of negotiations is reached, and a suboptimal decision strategy is output. Consistency determination uses the defined error When satisfied Output global collaborative decision-making strategies in real time; When the number of negotiations reaches Still not satisfied When the time comes, output the suboptimal decision strategy. The suboptimal decision strategy is defined as the set of decision variables that maximizes the global evaluation function across all negotiation rounds.
[0064] in, It is calculated by weighting the task completion rate, resource utilization rate, and decision delay.
[0065] To further implement the above technical solution, the specific content of step S33 is as follows: S331. Each agent receives a set of neighborhood messages. Then, first, based on the communication delay in the message... By aligning the neighbor states with the decision variables over time, the compensated states are obtained. Compensation decision variables ;
[0066] in, As a state predictor, LSTM is used to predict the output neighbor states of the network at time t. The estimated value; As a decision variable extrapolator, linear extrapolation based on the differences between two historical time points is used to calculate the compensated decision variables; when the predicted residuals satisfy... If this happens, the reliability of the neighbor's message is set to 0, and it will not participate in the consensus weighting in this round of negotiation; S332. The delay-compensated consensus-driven multi-agent deep deterministic policy gradient algorithm DC²-MADDPG is adopted to explicitly introduce delay condition variables to construct a delay condition value network and use the joint state after delay compensation for evaluation. No. The state space of an agent is defined as follows:
[0067] in, For the intensity of environmental disturbance, For communication delay, This is a fault status flag (0 for normal, 1 for fault).
[0068] No. The action space of an agent is defined as:
[0069] These correspond to adjustments in task allocation ratio, track parameters, and resource allocation ratio, respectively. Constructing a delayed conditional value network:
[0070] in, , , .
[0071] After the policy network outputs the update values of the decision variables, it first performs a local update:
[0072] S333. After completing the local update, each agent and its neighboring agents perform weighted consensus fusion, combining the weights determined by communication latency and link reliability, to obtain the negotiated decision variables. ;
[0073] Among them, weight satisfy And is determined by both communication latency and link reliability:
[0074] Through the above mechanism, agents with lower communication latency and higher link reliability have a higher weight in consensus fusion; S334. Design a delay adaptive adjustment mechanism to dynamically adjust the consensus negotiation frequency and the step size of the decision variable adjustment based on the average delay of the neighborhood. At the same time, use the LSTM network to predict the decision changes of neighboring agents, correct the policy output update, and adjust its own decision policy in advance. The average neighborhood delay is:
[0075] when At that time, the negotiation frequency is set to Step size coefficient is ;when At that time, the negotiation frequency is set to And satisfy The step size coefficient is And satisfy And correct the policy output update amount to:
[0076] in, or ; S335. Based on the state feedback information of each agent, a threshold detection method is used to identify faulty agents and isolate them; at the same time, based on the task load and decision-making ability of neighboring agents, a replacement agent is determined through a comprehensive evaluation function. The method for identifying faulty agents is as follows: The node's operational status is comprehensively judged using information such as the agent's periodically reported heartbeat information, state observation data, and consensus variables. When any fault criterion is met, the agent is determined to be potentially faulty. First, if an agent's heartbeat information is not received by neighboring nodes within a continuous time window, and the loss duration exceeds a preset threshold... If this occurs, it can be considered that the node's communication or system operation is malfunctioning; secondly, the system predicts the agent's state based on the dynamic model and records the observed state. With predicted state When comparing the residuals of the two, the following conditions are met:
[0077] This indicates that the agent's operating state has significantly deviated from normal dynamic behavior; furthermore, during the consensus coordination process, if the consensus variable of a certain node deviates from the neighborhood consensus result for a long period of time, and satisfies...
[0078] If the duration exceeds the preset time window, it indicates that the node may be making abnormal decisions; when any of the above conditions are met, the system will trigger the fault detection flag.
[0079] The fault isolation method is as follows: In order to prevent the faulty node from continuing to affect the consensus update and information propagation of the cluster, the agent is removed from the current consensus topology, and its adjacency weight and consensus weight are reset to zero, so that it no longer participates in neighborhood information interaction and decision update. Through this isolation mechanism, the propagation path of abnormal information in the cluster can be effectively blocked, ensuring that the consensus calculation of the remaining healthy nodes can still proceed stably. The method for determining alternative agents is as follows: First, construct a neighborhood candidate set for the faulty node. Then, a suitable healthy agent is selected from this set as an alternative execution node. The selection of the alternative node comprehensively considers factors such as the agent's capability matching degree, communication link reliability, and current task load. The optimal alternative node is determined through a comprehensive evaluation function.
[0080] in, This indicates the degree of matching between the agent's capability vector and the task requirements. This represents the link reliability metric between the node and its neighbors. Indicates the current task load level. The weighting coefficient is used to select intelligent agents with high capability matching, good communication reliability and low task load as replacement nodes, thereby ensuring the execution efficiency and collaborative stability after task takeover. S336. After determining the alternative agent, the policy parameters of the faulty node are migrated to the alternative agent in a soft update manner through policy inheritance and fast recovery mechanism;
[0081] in, This is the migration ratio coefficient; A policy distillation loss function is introduced to constrain the policy output of the replacement node to remain consistent with the policy of the original faulty node;
[0082] Through policy transfer and distillation mechanisms, alternative agents can recover their original collaborative behavior patterns in a short period of time, and continue to participate in the neighborhood consensus update and decision fusion process while undertaking new tasks, thereby achieving stable operation and rapid recovery of multi-agent systems in the event of node failure. S337. Transform multiple objectives into a single objective using a weighted summation method, and design a differentiated reward function system that combines the global reward of the overall cluster performance with the local reward based on individual performance; No. An intelligent agent at time The reward is:
[0083] in, Calculated based on cluster task completion rate, resource utilization rate, and overall decision latency; The calculation is based on the individual task completion quality, energy consumption, and orbital maneuver costs; the third item is a consistency constraint penalty.
[0084] To further implement the above technical solutions, such as Figure 4 The specific content of step S4 is as follows: S41. Gain experience in local decision-making, consensus-building and negotiation, and the execution process; Local decision-making experience is used to train the LSTM-DQN learning module in step S2, defined as follows:
[0085] in, The historical state sequence constructed from the model output of step S1 includes orbit, disturbance, link, task, and resource; Output actions for each module in step S2; The local decision reward value is calculated from task benefits, security constraints, disturbance penalties, link penalties, etc. The execution result label should include at least the following: task execution success / failure, conflict occurrence flag, resource consumption, and execution latency. The consensus negotiation experience is used to train the MARL-CCCM consensus negotiation model in step S3, and is defined as follows:
[0086] in, The joint observations after delay compensation in step S3 are obtained by time alignment from the orbit / link model output in step S1. This is the set of actions negotiated by each agent, including adjustments to decision variables; The set of communication delays is derived from the communication link model in step S1; Key decision variables before and after negotiation; This is a consistency error; The consensus negotiation reward consists of global performance, local performance, and consistency convergence term. The fault handling information should include at least the fault detection flag, isolation action, alternative selection result, and strategy inheritance / distillation loss value. The experience gained during the execution process, i.e., the evaluation feedback, is used to determine whether to trigger a model update, and is defined as:
[0087] in, As an indicator of decision reliability; The generalization adaptation rate is the metric. This can be used as an indicator of consensus achievement rate or average number of negotiation rounds. As an indicator of decision-making and negotiation delay; Indicators for energy consumption / propellant / communication overhead; For fault tolerance and recovery performance indicators in fault scenarios; S42. Introduce an experience screening and priority management mechanism for global experience storage, and divide the experience data into three data sets according to function: local decision-making experience pool, consensus negotiation experience pool, and evaluation feedback buffer. The experience screening and priority management mechanism is as follows: First, uploaded experiences are validated, and samples with missing status fields, abnormal timestamps, or mismatched actions and statuses are removed. Second, an anomaly detection mechanism filters out abnormal samples caused by communication errors or data corruption. After validation and anomaly screening, the remaining samples are scored based on reward value, task completion rate, and consensus convergence speed, and higher-value experiences are assigned higher priority. The experience sampling weight is represented as follows:
[0088] in, The reward is calculated by comprehensively considering factors such as absolute value of reward, TD error, severity of consensus failure, or difficulty of fault recovery. This priority sampling mechanism allows the model training to utilize representative key experiences more frequently, thereby improving training convergence efficiency. Local decision-making experience pool This is used to store the local decision-making experience generated in step S2 and serves as a data source for training and fine-tuning the LSTM-DQN learning model; consensus negotiation experience pool. Used to store the negotiation interaction data from step S3 and for training the delayed conditional MARL consensus model; evaluation feedback buffer This is used to record system performance evaluation indicators, providing a basis for triggering subsequent model updates; considering the limited on-board communication bandwidth, experience uploading adopts a tiered strategy: the agent first caches the experience locally, and when the communication link reliability meets the requirements... And the delay satisfies Upload only when necessary, and prioritize uploading high-value data such as high-reward experience, rare disturbance scenario experience, fault replacement and recovery experience, and consensus failure samples. S43. After acquiring global experience data, the decision model is periodically trained and optimized. The training process includes at least local learning model training, consensus negotiation model training, and fusion module parameter calibration. After the model training is completed, the updated model parameters are distributed to the local models of each agent through the parameter synchronization mechanism. During the training of the local learning model, the training object is the parameters of the LSTM-DQN reinforcement learning module, including the parameters of the LSTM encoding layer and the DQN value network. The training data comes from the local decision experience stored in the local decision experience pool. By minimizing the temporal decision value estimation error, the model can still obtain high decision gains under perturbation changes and link constraints. After training, the updated local learning model parameters are obtained. ; During the consensus negotiation model training process, the training object is the parameters of the step-delay conditional multi-agent reinforcement learning model. When implemented using DC²-MADDPG, this corresponds to the Actor and Critic network parameters. The training data comes from negotiation experience in the consensus negotiation experience pool. The training objective is to improve the consensus achievement rate, reduce the number of negotiation rounds, and decrease the consistency error under conditions of communication latency and node failure, thereby optimizing the overall task performance. After training, updated consensus negotiation model parameters are generated. ; The parameters of the fusion module are calibrated. By utilizing local decision-making experience and evaluation feedback data, the parameters of the fuzzy membership function, weight threshold, and safety lower bound are optimized to make the collaboration between rule-based decision-making and learning decision-making more stable and efficient. After training, an updated set of fusion parameters is obtained. ; After model training is completed, the system distributes the updated model parameters to the local models of each agent through a parameter synchronization mechanism. Parameter synchronization supports two modes: full synchronization and incremental synchronization. When the model structure remains unchanged but the performance is significantly improved, the complete parameter package is sent to each agent. If only some layer parameters or threshold parameters change, only the differential parameters are synchronized to reduce the communication load. At the same time, a unique version number is assigned to each synchronization and the effective timestamp and applicable scenario label are recorded. When the performance of the new version model degrades in local evaluation, the system restores to the previous stable version through a rollback mechanism. To adapt to different task requirements, each agent can make limited fine-tuning to the local model, such as adjusting the parameters of the end layer of the learning module or the threshold parameters of the fusion module, but must not modify the security master rule structure in the rule-driven module. S44. The local learning model and consensus negotiation model obtained through centralized training are compressed into a lightweight model suitable for satellite operation through knowledge distillation technology; For local learning models, the high-capacity LSTM-DQN teacher model is compressed into a lightweight student model through distillation to generate the learning actions required for real-time onboard decision-making. For the consensus negotiation model, the centralized training negotiation policy network is distilled into a lightweight negotiation network, which is used to generate initial values for negotiation actions or to perform rapid negotiation evaluation. The distillation training data comes from the state sequences, delay-compensated observations, and action output samples generated during steps S1–S3. The distillation objective function is expressed as follows:
[0089] in, Indicates losses related to the task outcome. This represents the difference loss between the policy outputs of the teacher model and the student model. To constrain the distillation model to maintain consistency with the original model in terms of consensus performance; to further reduce computational overhead, the system combines lightweight techniques such as parameter pruning, low-bit quantization, layer structure compression, and time window length reduction, so that the model can be adapted to the computing power of the on-board embedded processor. S45. Each model is updated through a real-time iterative update mechanism that combines active and passive triggering.
[0090] Active triggering is determined by the centralized training module based on global evaluation metrics. When the system performance is below a preset threshold, the model update process is automatically initiated. Passive triggering is initiated by each agent based on its local operating status. When an agent experiences consecutive decision failures, significant shifts in environment distribution, prolonged failure to reach consensus, or insufficient performance recovery after fault replacement, the agent will send a model update request to the centralized training module. The update request includes the current model version number, abnormal scene labels, key environmental statistics, and recent decision and negotiation segments. The centralized training module uses this information to determine whether a global model update or a local model update for a specific scenario is needed. After triggering, the source of the problem is first determined. When the problem is mainly manifested as insufficient local decision-making adaptability, the LSTM-DQN learning module is updated first. When the problem is manifested as a decrease in negotiation convergence speed or an increase in consensus failure rate, the consensus negotiation model is updated first. When the conflict between rule-based decision and learning decision increases significantly, the parameters of the fusion module are recalibrated. If the problem involves anomalies in the security boundary, only the rule threshold parameters are allowed to be adjusted, and they need to be verified on the ground before being updated. Through the triggering and judgment mechanism, it can be ensured that the model evolution process has a clear goal and a controllable range, thereby achieving continuous and stable operation of the system in complex space environments.
[0091] In another embodiment, a MATLAB / Simulink and Cesium co-simulation platform is built to verify the engineering feasibility and performance improvement of the method of the present invention, such as... Figure 5As shown, the platform integrates a space environment modeling module, a spacecraft multi-agent modeling module, a decision-making technology simulation module, and a performance evaluation module. It also performs comparative verification and visualization under scenarios such as orbital drift, disturbance enhancement, communication delay / interruption, and agent failure.
[0092] The joint simulation platform uses MATLAB / Simulink as the main environment for dynamics and algorithm simulation, and Cesium digital earth platform as the environment for orbit / state visualization and index monitoring, forming a closed-loop simulation link of dynamics calculation, strategy reasoning, communication interaction, visualization rendering and index acquisition.
[0093] The space environment modeling module calls the orbital dynamics model, space disturbance model, and communication link model constructed in step S1 to quantify the impact of features such as orbital drift, lunar perturbations, atmospheric drag, electromagnetic interference, and communication delay / interruption on decision-making, and outputs constrained representations and link status to the decision-making module; the spacecraft multi-agent modeling module constructs the perception, decision-making, execution, and communication functions of multiple spacecraft agents, supporting single agent status viewing and overall cluster situation monitoring; the decision technology simulation module deploys the RD-LSTM-DSDN, MARL-CCCM, and CTDE-DMRE mechanisms of this invention respectively, and sets up comparison methods for comparative experiments; the performance evaluation module collects and statistically calculates data on single-satellite decision-making, cluster consensus, model evolution, and visualization performance, and outputs index curves or statistical results.
[0094] Combining the characteristics of complex dynamic space environment with the mission requirements of spacecraft multi-agent cluster, three simulation scenarios are designed to verify the key mechanisms of steps S2 to S4 respectively, and quantitative evaluation is carried out through a unified indicator system.
[0095] Scenario 1: A mixed scenario of routine operation and dynamic perturbation, used to verify the single-satellite autonomous decision-making performance of the RD-LSTM-DSDN network; the scenario includes routine operations such as routine orbit maintenance and routine task execution, while adding environmental perturbations such as lunar perturbation, atmospheric drag, and electromagnetic interference to simulate a single-satellite autonomous decision-making scenario in a complex dynamic space environment. Evaluation indicators include single-satellite autonomous decision-making response time, routine operation reliability, and dynamic environment adaptability. Scenario 2: A scenario of communication delay and agent failure, used to verify the cluster consensus building performance of the MARL-CCCM mechanism. The scenario sets communication latency (200~1000ms) and single agent failure (randomly selected agent fails) to simulate a cluster collaborative decision-making scenario under complex communication constraints and agent failures. Evaluation indicators include cluster consensus achievement delay, task completion rate, resource utilization rate, and fault tolerance rate. Scenario 3: Long-term dynamic operation scenario, used to verify the iterative evolution performance of the decision model of the CTDE-DMRE mechanism. The scenario runs continuously for 10,000 time steps, including dynamic changes such as changes in environmental disturbance intensity, task type changes, and adjustments in the number of agents, to simulate the long-term operation scenario of a multi-agent cluster in a spacecraft. Evaluation indicators include online update time of the decision model, model generalization adaptation rate, and long-term operational reliability.
[0096] Comparative experiments were conducted for the three scenarios described above, and various evaluation indicators were collected and compared with the comparison methods. The following verification conclusions were obtained: Simulation results analysis for scenario 1: The single-satellite autonomous decision-making response time of the RD-LSTM-DSDN network is ≤500ms, which is 42.3% shorter than the traditional method; the reliability of routine operations is ≥99.8%, which is 3.1 percentage points higher than the traditional method; the dynamic environment adaptability is ≥92.5%, which is 15.7 percentage points higher than the traditional method; combined with Cesium visualization monitoring data, the orbit rendering accuracy of this invention on the Cesium platform is stable at 8.2m, the visualization frame rate is maintained at 35fps, and the real-time rendering latency of spatial disturbances is ≤42ms, all of which are better than the traditional method's orbit rendering accuracy of 15.3m, visualization frame rate of 22fps, and rendering latency of 78ms. Experimental results show that the RD-LSTM-DSDN network can effectively balance high reliability of routine operations and adaptability to dynamic environments, significantly improve the autonomous decision-making capability of a single satellite, and has good compatibility with the Cesium platform, which can accurately visualize the single-satellite decision execution effect.
[0097] Scenario 2 Simulation Result Analysis: In scenarios with communication latency ≤ 500ms, the MARL-CCCM mechanism achieves cluster consensus latency ≤ 2s, a 51.2% reduction compared to traditional methods; the task completion rate under single agent failure is ≥ 95.3%, outperforming traditional methods by 8.4 percentage points; resource utilization is ≥ 86.7%, an improvement of 7.9 percentage points compared to traditional methods; and fault tolerance is ≥ 98.0%, effectively addressing agent failure issues. Cesium visualization key indicator comparison: The multi-agent state synchronization accuracy of this invention is ≤ 85ms, the fault agent identification visualization latency is ≤ 120ms, and the cluster track collaborative rendering consistency is ≥ 98%, while the corresponding indicators for traditional methods are 160ms, 250ms, and 83%, respectively. Experimental results show that the MARL-CCCM mechanism can quickly achieve cluster consensus under communication latency and agent failure scenarios, improving collaborative decision-making efficiency and fault tolerance. The Cesium platform can clearly visualize the cluster consensus process, fault identification, and collaborative execution effects, with superior synchronization accuracy and rendering consistency.
[0098] Scenario 3 Simulation Result Analysis: The online update time of the decision model using the CTDE-DMRE mechanism is ≤28.5s, a 62.7% speedup compared to traditional methods; the model generalization and adaptation rate is ≥91.2%, an improvement of 12.3 percentage points compared to traditional methods; the long-term operational reliability is ≥97.8%, an improvement of 5.6 percentage points compared to traditional methods; Comparison of key Cesium long-term operational indicators: Within 10,000 time steps, the visualization frame rate of this invention remains stable at 32~38fps, the orbit rendering accuracy fluctuation is ≤±1.5m, and there is no significant stuttering. The stuttering time is ≤100ms / 1000 steps. Traditional methods have large frame rate fluctuations (18~25fps), orbit rendering accuracy fluctuations ≤±4.2m, and stuttering time ≥350ms / 1000 steps. Experimental results show that the CTDE-DMRE mechanism can realize real-time iterative updates of the decision model, improve the model's generalization ability and long-term operational stability. The Cesium platform can stably support the visualization simulation of long-term dynamic operation scenarios. The rendering accuracy and frame rate stability meet the simulation requirements of spacecraft multi-agent simulation and are superior to traditional visualization simulation methods.
[0099] In another embodiment, traditional centralized decision-making methods (Method 1), single reinforcement learning distributed decision-making methods (Method 2), and consensus protocol-based distributed decision-making methods (Method 3) in the field of multi-agent distributed decision-making in spacecraft are selected and compared with the method of the present invention (Method 4) in a comprehensive comparative experiment. The comparison dimensions cover four categories: decision performance, collaborative performance, model performance, and Cesium visualization performance, to ensure the comprehensiveness and objectivity of the comparison. The comparison data are all from simulation results under the same simulation scenario and experimental parameters. The specific comparison is as follows: Figure 6 "—" indicates that the method has no corresponding function.
[0100] A comprehensive comparison of the four core indicators shows that this invention outperforms mainstream methods in the field of multi-agent distributed decision-making in spacecraft in terms of decision performance, collaborative performance, model performance, and Cesium visualization performance. It effectively solves the core problems of existing methods, such as strong dependence on central nodes, insufficient decision reliability, low collaborative efficiency, poor model generalization, lack of online update capability, and poor visualization effect. This fully demonstrates the innovation and superiority of the technical solution presented in this paper, and is more adaptable to the engineering requirements of multi-agent distributed autonomous decision-making in spacecraft under complex dynamic space environments.
[0101] A spacecraft multi-agent distributed autonomous decision-making evolution system, based on a spacecraft multi-agent distributed autonomous decision-making evolution method, includes: a space environment feature analysis and modeling module, a rule-driven-LSTM fusion distributed self-organizing decision network, a cluster consensus mechanism based on multi-agent reinforcement learning, and a real-time iterative evolution mechanism for decision models based on CTDE architecture; The space environment feature analysis and modeling module is configured to acquire spacecraft dynamics status, environmental disturbance information, resource status and communication link status and perform feature analysis, establish orbital dynamics model, space disturbance model and communication link model, and output environmental model parameters and constrained characterization. The rule-driven and LSTM-fused distributed self-organizing decision network is configured to construct a unified decision input state vector. It generates multi-source candidate decisions by fusing rule-driven and LSTM reinforcement learning, and corrects them through a distributed conflict avoidance and negotiation mechanism, outputting the local decision strategy of each spacecraft agent in the current decision cycle. The cluster consensus mechanism based on multi-agent reinforcement learning is configured to generate local decisions by clarifying the core objectives and key decision variables of the cluster consensus, and to introduce a multi-objective optimization and differentiated reward function system for consensus negotiation. The consensus verification outputs a global collaborative decision strategy and a set of corresponding key decision variables that meet the consistency threshold constraint. The real-time iterative evolution mechanism of the decision model based on the CTDE architecture is configured to perform centralized training, distributed execution of the decision model based on the CTDE architecture, combined with experience sharing, knowledge distillation and parameter synchronization technologies, to centrally train, distribute and synchronize the decision network and cluster consensus building mechanism and deploy it in a lightweight manner on the satellite.
[0102] To further implement the above technical solutions, the rule-driven-LSTM fusion distributed self-organizing decision network includes: a rule-driven module, an LSTM reinforcement learning module, a decision fusion module, and a distributed self-organizing and conflict avoidance module. The cluster consensus mechanism based on multi-agent reinforcement learning includes: a consensus goal and variable definition module, a local decision generation module, a consensus negotiation module, and a consensus verification module; The real-time iterative evolution mechanism of the decision model based on the CTDE architecture includes: a centralized training module, a distributed execution module, an experience storage module, a knowledge distillation module, and an evaluation and feedback module.
[0103] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0104] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multi-agent distributed autonomous decision-making evolution method for spacecraft, characterized in that, include: S1. Acquire spacecraft dynamics status, environmental disturbance information, resource status and communication link status and perform feature analysis, establish orbital dynamics model, space disturbance model and communication link model, and output environmental model parameters and constrained characterization; S2. Construct a unified decision input state vector, input a rule-driven-LSTM fusion distributed self-organizing decision network, fuse rule-driven and LSTM reinforcement learning to generate multi-source candidate decisions, and correct them through a distributed conflict avoidance and negotiation mechanism, outputting the local decision strategy of each spacecraft agent in the current decision cycle; S3. A cluster consensus building mechanism based on multi-agent reinforcement learning is established by clarifying the core objective and key decision variables of cluster consensus, generating local decisions, and introducing a multi-objective optimization and differentiated reward function system for consensus negotiation. The consensus verification outputs a global collaborative decision-making strategy and a set of corresponding key decision variables that meet the consistency threshold constraint. S4. Through the real-time iterative evolution mechanism of the decision model based on the centralized training and distributed execution CTDE architecture, combined with experience sharing, knowledge distillation and parameter synchronization technologies, the decision network and cluster consensus building mechanism are centrally trained, distributed and synchronized and deployed in a lightweight manner on the satellite. The specific content of step S2 includes: S21. Taking track dynamics, disturbance intensity index and link quality index as input, the system uses a production rule base to make deterministic judgments on safety constraints, track maintenance, task priority and resource limitations, and outputs candidate rule actions that meet the engineering safety boundary and conventional operating procedures. S22. Based on historical state sequence Input the LSTM-DQN network, use LSTM to perform time-series encoding of orbit drift trend, disturbance evolution trend and link delay fluctuation, and combine DQN output to adaptively learn candidate actions in dynamic and uncertain scenarios; S23. Receive rule-based candidate actions and adaptive candidate actions, and determine the fusion weight based on the comprehensive disturbance strength, link constraint index and task urgency. Use the action consistency mapping function to map the two types of actions to a unified action vector space to form a fused action. S24. Based on the fusion action, combined with the short-term prediction output of the orbital dynamics model and the output of the communication link model, local intention interaction, conflict detection and negotiation correction are performed within the neighborhood topology, and the final local decision-making strategy is output. The specific content of step S3 is as follows: S31. Define the core objectives and key decision variables of the cluster consensus. The core objectives include maximizing task completion rate, maximizing resource utilization, and minimizing decision delay. The key decision variables include task allocation scheme, track adjustment strategy, and resource allocation ratio. At the same time, initialize the parameters, reward function weights, and consensus threshold of the multi-agent reinforcement learning model, and set the communication range and interaction frequency of the agents. S32. Obtain the local decision-making strategies generated by each intelligent agent, and at the same time collect its own state, task requirements and local environment information. Send the local decision-making strategies and key states to neighboring intelligent agents through distributed communication to achieve local information sharing. S33. Each agent receives local decision information from neighboring agents, adjusts decision variables and conducts negotiations by combining its own strategy with the global objective through a multi-agent reinforcement learning model; introduces a delay compensation mechanism to predict the trend of decision changes of neighboring agents; introduces a fault detection and replacement mechanism to identify faulty agents and assign tasks to neighboring healthy agents. S34. Based on the preset consensus threshold, verify the consistency of each agent on key decision variables. If the consistency error is less than or equal to the consensus threshold, it is considered that a cluster consensus has been reached and a global collaborative decision strategy is output. If no consensus is reached, return to the consensus negotiation stage and continue to adjust the decision variables until a consensus is reached or the maximum number of negotiations is reached, and a suboptimal decision strategy is output.
2. The spacecraft multi-agent distributed autonomous decision-making evolution method as described in claim 1, characterized in that, Space disturbance models include lunar perturbation acceleration models, atmospheric drag acceleration models, and space electromagnetic interference models; communication link models include communication delay models and communication link interruption models. The orbital dynamics model outputs orbital state variables and relative motion parameters; the space disturbance model outputs various disturbance intensity estimates; the communication link model outputs communication delay parameters, including signal transmission delay and signal processing delay; and the communication link interruption model obtains the link interruption probability.
3. The spacecraft multi-agent distributed autonomous decision-making evolution method as described in claim 1, characterized in that, No. A spacecraft intelligent agent at any time The decision input state vector is: in, For the first A spacecraft intelligent agent at any time The position vector in the geocentric inertial coordinate system or orbital coordinate system. For velocity vector, Relative motion For relative velocity, This represents the intensity characteristics of environmental disturbances at the current moment. Due to signal propagation delay, This represents the probability of link interruption. For mission requirements, This represents the remaining resources; No. The local decision-making strategy output of each agent is: in, These are task-level actions, including execution, postponement, handover, and reordering of instructions. For orbit control layer actions, including orbit adjustment direction and amplitude, maneuver window selection, and attitude maneuver trigger commands; For resource allocation actions, For communication and coordination actions.
4. The spacecraft multi-agent distributed autonomous decision-making evolution method as described in claim 1, characterized in that, The specific content of step S21 is as follows: In the state determination phase, the output results of the troop trajectory dynamics model, spatial disturbance model and communication link model are converted into rule triggering condition variables. At the same time, task and resource condition variables are constructed by combining the task system and platform state information to form the set of condition variables required for rule triggering, and further mapped into discretized condition labels through threshold determination. During the rule matching phase, the condition variables are logically matched based on the production rule base; When multiple rules are triggered simultaneously and their actions conflict, the priority conflict resolution phase begins. A unified rule scoring mechanism is established to rank the candidate rules, and the final rule candidate action is selected based on the scoring results. The LSTM-DQN network in step S22 includes an input layer, an LSTM encoding layer, a feature mapping layer, and a Q-value output layer; The input layer receives a sequence of historical states. ; The LSTM encoding layer performs temporal encoding on the historical states and outputs the hidden states. It is used to characterize orbital change trends, disturbance evolution trends, and link state evolution trends; the feature mapping layer will... Mapped to decision feature vector ; The Q-value output layer outputs a set of discrete actions. The value of each action It then performs a value assessment and finally outputs adaptive learning candidate actions; The specific content of step S23 is as follows: Based on the model output in step S1, construct the scenario discrimination input, including the comprehensive disturbance strength index, link constraint index, and task urgency index; Using the scene discrimination input as fuzzy input, three levels of fuzzy sets (low, medium, and high) are constructed respectively. The system state of the current decision cycle is represented in a fuzzy manner through the fuzzy membership function. Inference calculations are performed based on a pre-designed fuzzy rule base to determine the fusion weight of rule candidate actions and adaptive candidate actions in the current cycle; After the fuzzy rule reasoning is completed, the rule weights and learning weights are obtained by defuzzification calculation; The rule-based candidate actions and adaptive candidate actions are mapped to a unified action vector space, and the fused action is output by combining the rule weights and the learned weights. The specific content of step S24 is as follows: The agent encapsulates the key quantities required for fusion actions and conflict determination into neighborhood messages, and sends neighborhood decision intentions based on a communication sending strategy constrained by the communication link model. After receiving the decision intentions of the neighborhood, the neighboring agents perform task conflict detection, track conflict detection, and resource conflict detection. If a conflict is detected, a negotiation adjustment is performed in the neighborhood based on the local negotiation evaluation function to maximize the benefits of high-priority tasks and reduce the overall cost while satisfying safety constraints. At the same time, the LSTM module is used to predict the conflict development trend and generate conflict avoidance strategies in advance.
5. The spacecraft multi-agent distributed autonomous decision-making evolution method as described in claim 1, characterized in that, The specific content of step S33 is as follows: S331. After each agent receives the neighborhood message set, it first performs time alignment of the neighbor state and decision variable according to the communication delay in the message to obtain the compensated state and compensated decision variable. S332. A multi-agent deep deterministic policy gradient algorithm driven by delay compensation and consensus is adopted. Delay condition variables are explicitly introduced to construct a delay condition value network, and the joint state after delay compensation is used for evaluation. After the policy network outputs the update amount of the decision variable, local update is performed. S333. After completing the local update, each agent and its neighboring agents perform weighted consensus fusion, based on the weights determined by communication latency and link reliability, to obtain the negotiated decision variables; S334. Design a delay adaptive adjustment mechanism to dynamically adjust the consensus negotiation frequency and the step size of the decision variable adjustment based on the average delay of the neighborhood. At the same time, use the LSTM network to predict the decision changes of neighboring agents, correct the policy output update, and adjust its own decision policy in advance. S335. Based on the state feedback information of each agent, a threshold detection method is used to identify faulty agents and isolate them; at the same time, based on the task load and decision-making ability of neighboring agents, a replacement agent is determined through a comprehensive evaluation function. S336. After determining the replacement agent, the policy parameters of the faulty node are migrated to the replacement agent in a soft update manner through policy inheritance and fast recovery mechanism. The policy distillation loss function is introduced to constrain the policy output of the replacement node to be consistent with the policy of the original faulty node. S337. Transform multiple objectives into a single objective using a weighted summation method, and design a differentiated reward function system that combines the global reward based on the overall performance of the cluster with the local reward based on individual performance.
6. The spacecraft multi-agent distributed autonomous decision-making evolution method as described in claim 1, characterized in that, The specific content of step S4 is as follows: S41. Gain experience in local decision-making, consensus-building and negotiation, and the execution process; S42. Introduce an experience screening and priority management mechanism for global experience storage, and divide the experience data into three data sets according to function: local decision-making experience pool, consensus negotiation experience pool, and evaluation feedback buffer. S43. After acquiring global experience data, the decision model is periodically trained and optimized. The training process includes at least local learning model training, consensus negotiation model training, and fusion module parameter calibration. After the model training is completed, the updated model parameters are distributed to the local models of each agent through the parameter synchronization mechanism. S44. The local learning model and consensus negotiation model obtained through centralized training are compressed into a lightweight model suitable for satellite operation through knowledge distillation technology; S45. Each model is updated through a real-time iterative update mechanism that combines active and passive triggering.
7. A spacecraft multi-agent distributed autonomous decision-making evolution system, characterized in that, A spacecraft multi-agent distributed autonomous decision-making evolution method based on any one of claims 1-6 includes: a space environment feature analysis and modeling module, a rule-driven-LSTM fusion distributed self-organizing decision network, a cluster consensus mechanism based on multi-agent reinforcement learning, and a real-time iterative evolution mechanism for decision models based on CTDE architecture. The space environment feature analysis and modeling module is configured to acquire spacecraft dynamics status, environmental disturbance information, resource status and communication link status and perform feature analysis, establish orbital dynamics model, space disturbance model and communication link model, and output environmental model parameters and constrained characterization. The rule-driven and LSTM-fused distributed self-organizing decision network is configured to construct a unified decision input state vector. It generates multi-source candidate decisions by fusing rule-driven and LSTM reinforcement learning, and corrects the decisions through a distributed conflict avoidance and negotiation mechanism, outputting the local decision strategies of each spacecraft agent in the current decision cycle. The cluster consensus mechanism based on multi-agent reinforcement learning is configured to generate local decisions by clarifying the core objectives and key decision variables of the cluster consensus, and to introduce a multi-objective optimization and differentiated reward function system for consensus negotiation. The consensus verification outputs a global collaborative decision strategy and a set of corresponding key decision variables that meet the consistency threshold constraint. The real-time iterative evolution mechanism of the decision model based on the CTDE architecture is configured to perform centralized training, distributed execution of the decision model based on the CTDE architecture, combined with experience sharing, knowledge distillation and parameter synchronization technologies, to centrally train, distribute and synchronize the decision network and cluster consensus building mechanism and deploy it in a lightweight manner on the satellite.
8. A spacecraft multi-agent distributed autonomous decision-making evolution system as described in claim 7, characterized in that, The rule-driven-LSTM fusion distributed self-organizing decision network includes: a rule-driven module, an LSTM reinforcement learning module, a decision fusion module, and a distributed self-organizing and conflict avoidance module. The cluster consensus mechanism based on multi-agent reinforcement learning includes: a consensus goal and variable definition module, a local decision generation module, a consensus negotiation module, and a consensus verification module; The real-time iterative evolution mechanism of the decision model based on the CTDE architecture includes: a centralized training module, a distributed execution module, an experience storage module, a knowledge distillation module, and an evaluation and feedback module.
Citation Information
Patent Citations
Spacecraft variable-quantity space debris collision avoidance autonomous decision-making method based on safety reinforcement learning
CN117972901A
Heterogeneous spacecraft cluster target distribution method based on multi-agent reinforcement learning algorithm
CN121503221A