Low earth orbit satellite network routing method and system based on multi-agent reinforcement learning and heterogeneous on-board computing

By employing a routing method based on multi-agent reinforcement learning and heterogeneous onboard computing, the poor adaptability and decision-making lag issues of LEO satellite networks under dynamic topology and link fluctuations are addressed. This approach achieves global optimization and distributed collaboration, thereby improving network performance and robustness.

CN122496071APending Publication Date: 2026-07-31INNOVATION ACAD FOR MICROSATELLITES OF CAS +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INNOVATION ACAD FOR MICROSATELLITES OF CAS
Filing Date
2026-03-04
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing LEO satellite network routing technologies suffer from poor adaptability, decision lag, difficulty in achieving global optimization, and challenges in deploying complex intelligent algorithms on satellites when dealing with highly dynamic topologies, fluctuating link quality, complex traffic patterns, and distributed collaboration.

Method used

A routing method employing multi-agent reinforcement learning and heterogeneous spaceborne computing is proposed. Through centralized training on the ground, the multi-agent reinforcement learning algorithm enables each satellite node to learn a cooperative routing strategy, and a lightweight model is deployed on a heterogeneous spaceborne computing system to achieve real-time, autonomous, and cooperative routing decisions.

Benefits of technology

It enhances the adaptability of LEO satellite networks under highly dynamic topologies and fluctuating links, reduces decision lag, achieves global load balancing, improves network throughput and carrying capacity, reduces network construction and operation costs, and enhances system survivability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122496071A_ABST
    Figure CN122496071A_ABST
Patent Text Reader

Abstract

A routing method and system for low Earth orbit satellite networks based on multi-agent reinforcement learning and heterogeneous spaceborne computing is presented. This method defines each satellite node as an agent and performs offline centralized collaborative training on the ground using a multi-agent near-end policy optimization algorithm. A centralized commentator network is used to introduce global state information to guide each agent in learning a collaborative routing strategy. After training, the executor policy model is deployed on a heterogeneous spaceborne computing system on the satellite nodes. This system includes a routing processing board and a high-performance computing board connected via a high-speed interface, responsible for high-speed packet forwarding and real-time model inference, respectively. During on-orbit operation, each satellite node autonomously makes decisions and forwards data packets based on local perception information. This invention enables real-time spaceborne execution of complex intelligent algorithms, effectively reducing end-to-end latency and packet loss rate, enhancing network adaptability and robustness, and can be widely applied to next-generation non-terrestrial communication networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of satellite communication technology, specifically relating to a routing method and system for low Earth orbit (LEO) satellite networks based on multi-agent reinforcement learning and heterogeneous onboard computing. Background Technology

[0002] Low Earth Orbit (LEO) satellite networks, with their significant advantages such as low transmission latency, low path loss, and strong global coverage, have become a key infrastructure for building next-generation ubiquitous global communication networks (such as 6G non-terrestrial networks), and are widely used in many fields such as broadband internet access, IoT data backhaul, remote sensing observation, and emergency communications. However, the unique characteristics of LEO satellite networks also bring unprecedented challenges to network routing design.

[0003] First, the network topology exhibits rapid and dynamic changes. LEO satellites operate at extremely high speeds in orbit, leading to frequent switching of visibility relationships between satellites and the connection status of inter-satellite links (ISLs). Traditional routing protocols, such as shortest path first algorithms based on static topology calculations or periodic updates (e.g., OSPF and its variants), rely on the core assumption that the network state remains stable over relatively long timescales. These protocols struggle to keep pace with the rapid changes in LEO network topology. Their routing update signaling overhead is enormous and they cannot respond in real time, resulting in lagging route selection, an inability to effectively adapt to the dynamic environment, and a sharp decline in network performance.

[0004] Secondly, link quality and network traffic are highly volatile. The communication quality of inter-satellite links is easily affected by various factors such as satellite attitude adjustments, space weather changes, and Doppler shift, leading to drastic fluctuations in parameters such as link latency, bit error rate, and available bandwidth within a short period. Simultaneously, due to the uneven distribution of users and the uneven distribution of service demands, network traffic can form hotspots in both time and space, causing localized congestion. However, most existing routing schemes, whether based on distance vectors or link states, often rely on a single, static metric (such as hop count) for decision-making, lacking the ability to dynamically perceive real-time link quality, node queue length, and network congestion. This can cause data flows to be continuously directed to already congested nodes or links, resulting in severe queuing delays and packet loss, and low network resource utilization.

[0005] Furthermore, there is a lack of distributed intelligent decision-making and collaboration mechanisms. Given the massive scale of the LEO constellation (potentially containing hundreds or thousands of satellites), relying entirely on centralized ground control presents challenges such as high control latency, limited satellite-to-ground link bandwidth, and the risk of single points of failure. Therefore, distributed routing is essential. However, simple distributed strategies, such as opportunistic routing or greedy forwarding based on local information, often only pursue local optima, lacking a global perspective and collaboration between agents. This can lead to conflicting decisions among multiple satellites, such as simultaneously directing traffic to the same area, exacerbating rather than alleviating network congestion, and failing to achieve global resource optimization and load balancing.

[0006] Finally, the deployment of complex intelligent algorithms on spacecraft faces bottlenecks. To address the aforementioned dynamic and collaborative challenges, academia has begun exploring the introduction of artificial intelligence, particularly deep reinforcement learning (DRL), to design adaptive routing strategies. DRL algorithms, through interactive learning with the environment, are expected to learn optimal decisions in complex environments. However, these algorithms are typically computationally intensive, requiring powerful computing capabilities for their inference processes. Traditional spacecraft computers, based on single-core or low-power CPUs, have limited processing power and cannot meet the stringent requirements of real-time inference for deep neural network models. Although some research has attempted to use single-agent DRL, the lack of effective information sharing and collaborative training among agents makes it difficult to solve global problems such as traffic balancing and congestion control at the network level, thus limiting its application in complex LEO network environments.

[0007] In summary, existing LEO satellite network routing technologies have significant shortcomings in handling highly dynamic topologies, fluctuating link quality, complex traffic patterns, and achieving efficient distributed collaboration. Furthermore, limitations imposed by the computing power bottleneck of onboard computing platforms hinder the practical application of advanced intelligent algorithms in orbit. Therefore, a novel routing method and system that integrates advanced intelligent decision-making algorithms with high-performance onboard computing capabilities is urgently needed to fully unleash the communication potential of LEO satellite networks. Summary of the Invention

[0008] This invention aims to overcome the shortcomings of existing technologies and solve the problems of poor adaptability, decision lag, difficulty in achieving global optimization, and difficulty in deploying complex intelligent algorithms on satellites in traditional LEO satellite network routing schemes when dealing with highly dynamic topologies, link fluctuations, and distributed cooperation.

[0009] To achieve the above objectives, this invention provides a low Earth orbit satellite network routing method and system based on multi-agent reinforcement learning and heterogeneous spaceborne computing.

[0010] This invention provides a routing method for low Earth orbit (LEO) satellite networks based on multi-agent reinforcement learning and heterogeneous onboard computing. The core of this method is to treat each satellite node in the LEO satellite network as an independently decision-making agent. Through centralized training using a multi-agent reinforcement learning algorithm, each agent learns a cooperative routing strategy. The trained strategy model is then deployed on the heterogeneous onboard computing systems of the satellite nodes, achieving real-time, autonomous, and cooperative routing decisions in orbit. The method specifically includes the following steps: Step S1: Multi-agent system modeling and environment definition.

[0011] Each satellite node in the LEO satellite network is defined as an independent agent. Each agent has a local observation space, an action space, and a global reward function. The local observation space includes, but is not limited to, the link quality parameters between the satellite node and its neighbors (such as real-time latency, packet loss rate, and signal-to-noise ratio), the queue length of each outgoing port of the node, neighbor congestion status information obtained from neighbor beacons, and the Quality of Service (QoS) requirement field in the data packets to be forwarded. The action space is defined as the set of all possible next-hop forwarding neighbors for the satellite node. The global reward function is designed to guide the agent to learn cooperative optimization objectives, such as minimizing the global average end-to-end latency, minimizing the global packet loss rate, or maximizing the total network throughput.

[0012] Step S2: Centralized offline training based on Multi-Agent Proximal Policy Optimization (MAPPO). A centralized training environment is constructed on a powerful ground-based computing cluster. This environment simulates the dynamic topology, link changes, and traffic patterns of the LEO satellite network. During training, each agent k has an actor policy network π. θk ( ak | ok ), used to base observations on local observations. k Output selection action a k The probability distribution of (i.e., the next hop node). Simultaneously, a centralized critic network is constructed, such as V... φ (s) or Q φ (s, a), where s represents the global state (or the joint observations of all agents), and a represents the joint action of all agents. During training, all agents interact with the environment according to the current policy, and transfer their experience trajectories (including each agent's observations o) to the environment. k Action a k The global reward (r) and the next-time global state (s') are stored in the experience replay pool. By sampling experience, the value of the joint state-action sequence is evaluated using a centralized critic network, and the advantage function A is calculated. k,tAnd update the executor network π for each agent accordingly. θk The goal is to maximize the cumulative expected reward that takes into account the global impact. In this way, the policy network of each agent achieves implicit cooperation during the learning process, and ultimately learns a routing policy that can jointly optimize the global performance of the network.

[0013] Step S3: Lightweighting and deployment of the trained policy model.

[0014] The agent policy network π obtained after offline centralized training converges is used to calculate the agent policy network for each agent. θk Extraction is performed. Optionally, to adapt to the resource limitations of the spaceborne platform, the model can be lightweighted, such as by weight quantization, pruning, or knowledge distillation. The processed executor policy network model file is pre-burned or loaded into the spaceborne intelligent processing unit of the corresponding LEO satellite node.

[0015] Step S4: On-orbit distributed real-time inference and routing decision.

[0016] During the LEO satellite's on-orbit operation phase, the onboard intelligent processing unit on each satellite node operates independently. For each data packet arriving at this node, the following sub-steps are executed: Step S41: Local State Awareness and Observation Vector Construction. The node, through the FPGA and CPU on its routing processing board, senses and collects local network state information in real time, including queue lengths at each outgoing port, real-time quality of neighboring links, and parsed packet header information, and integrates this information into a local observation vector conforming to the executor's network input format. k,t .

[0017] Step S42: Intelligent decision-making task distribution. The CPU on the routing processing board will distribute the constructed observation vector o k,t It is sent to the high-performance computing board via a high-speed inter-board interface.

[0018] Step S43: Actor Network Inference. The GPU or SoC on the high-performance computing board receives the observation vector o. k,t And input it into the preloaded executor policy network π θk The forward inference calculation is performed to output the probability distribution or Q-value of all possible actions (i.e., candidate next-hop nodes).

[0019] Step S44: Optimal action selection. Based on the network output, select the action a* with the highest probability or the largest Q-value. k,t The node selected as the optimal next-hop forwarding node for the data packet is then determined. This decision is transmitted back to the CPU of the routing processing board via the high-speed inter-board interface.

[0020] Step S45: High-speed packet forwarding. The CPU of the routing processing board parses the decision result and instructs its FPGA to forward the data packets from the specified output port at high speed and line speed according to the decision.

[0021] Another aspect of this invention provides a low Earth orbit satellite network routing system based on multi-agent reinforcement learning and heterogeneous spaceborne computing. This system is deployed on each satellite node in the LEO satellite network, serving as the hardware foundation for implementing the aforementioned method. The system includes: Routing processing board: As the data plane, it is responsible for receiving, initially processing, and high-speed forwarding of data packets. It includes at least: Field Programmable Gate Array (FPGA): Used to implement line-speed packet parsing, protocol processing, queue management, and high-speed switching and forwarding based on routing decisions.

[0022] Central Processing Unit (CPU): Coupled with the FPGA, it is responsible for collecting low-level state information from the FPGA, constructing observation vectors, controlling communication with the high-performance computing board, and parsing and issuing routing decision instructions to the FPGA.

[0023] High-performance computing board: Serving as the intelligent decision plane, it is responsible for executing inference for complex multi-agent reinforcement learning algorithm models. It connects to the routing processing board via a high-speed inter-board interface and includes at least: Graphics Processing Unit (GPU) and / or System-on-Chip (SoC): Used to load and run pre-trained executor policy network models, perform high-speed parallel computation based on received observation vectors, and generate optimal routing decisions.

[0024] Power supply unit: Connected to the routing processing board and the high-performance computing board, it provides a stable and reliable power supply to both.

[0025] Compared with the prior art, the beneficial effects of the present invention 1) This invention introduces a "centralized training, distributed execution" paradigm of multi-agent reinforcement learning (especially the MAPPO algorithm), enabling each satellite node (agent) to learn routing strategies not only based on local real-time observations but also incorporating global collaborative wisdom trained through a centralized commentator network. This gives the network superior adaptability when facing the inherently high dynamic topology, fluctuating link quality, and bursty traffic of the LEO constellation.

[0026] 2) This invention employs a fully distributed online execution architecture. After training, each satellite node makes independent decisions relying solely on its local observations and onboard computing resources, without depending on any central control node or ground station. This decentralized architecture eliminates the risk of single points of failure. Even if some nodes fail or the satellite-to-ground link is interrupted, the remaining nodes can continue to work collaboratively based on the learned cooperation strategy, maintaining basic network communication services and greatly enhancing the survivability and robustness of the entire constellation system.

[0027] When a new satellite node joins the constellation, it can quickly integrate into the existing network and work in collaboration with other nodes simply by loading an actor policy model that is adapted to its orbital position and obtained through offline training. This demonstrates excellent scalability without requiring reconfiguration of the entire network or interruption of existing services.

[0028] 3) By placing the entire decision-making closed loop on the satellite, this invention frees the LEO satellite network from excessive reliance on ground control, achieving true "on-orbit intelligence." The satellite can autonomously and rapidly make optimal route adjustments based on real-time network status perception, improving response speed from minutes to milliseconds, significantly enhancing the network's ability to respond to emergencies.

[0029] 4) By achieving global load balancing through multi-agent collaborative learning, this invention can more efficiently utilize valuable on-board resources (such as link bandwidth and queue buffers). Data flows no longer blindly rush to a few "shortest" paths, but are intelligently distributed across all available network resources, thereby improving the overall network throughput and carrying capacity. This delays or avoids the need for expensive hardware expansion or launching more satellites to cope with local congestion, effectively reducing network construction and operation costs. Attached Figure Description

[0030] Figure 1 This is a schematic diagram comparing the traditional ground processing mode with the on-board processing mode of this invention.

[0031] Figure 2 This is a structural block diagram of a spaceborne intelligent processing unit provided in one embodiment of the present invention.

[0032] Figure 3 This is a schematic diagram of the principle framework of the Multi-Agent Proximity Policy Optimization (MAPPO) algorithm used in one embodiment of the present invention.

[0033] Figure 4 This is a schematic diagram illustrating the deployment of a multi-agent reinforcement learning-based LEO satellite network routing system in a constellation, as provided in one embodiment of the present invention.

[0034] Figure 5The diagram shows a performance comparison between the proposed solution and existing technical solutions. Specifically: (a) is a bar chart comparing the average network packet loss rate using the proposed solution (MAPPO algorithm), the single-agent PPO algorithm, and the traditional shortest path algorithm at different packet injection rates (pps); (b) is a bar chart comparing the average end-to-end network delay using the proposed solution (MAPPO algorithm), the single-agent PPO algorithm, and the traditional shortest path algorithm at different packet injection rates (pps). Detailed Implementation

[0035] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0036] The core concept of this invention lies in modeling the dynamically changing routing decision problem in LEO satellite networks as a multi-agent collaborative learning problem. Through centralized, offline training on the ground that considers global information, each satellite node (agent) learns a lightweight policy network capable of making collaborative optimization decisions based on local observations. Subsequently, this policy network is deployed on a specially designed onboard intelligent processing unit with CPU+GPU+FPGA / SoC heterogeneous computing capabilities, thereby achieving real-time, autonomous, and intelligent routing for satellites in orbit.

[0037] Several embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0038] Example 1: A LEO satellite network routing method based on MAPPO and heterogeneous onboard computing This embodiment provides a low Earth orbit satellite network routing method based on multi-agent reinforcement learning and heterogeneous spaceborne computing. Its overall process can be divided into two stages: offline training and online execution.

[0039] Step S1: Multi-agent system modeling and environment definition.

[0040] Each satellite node in the LEO satellite network is defined as an independent agent. Each agent has a local observation space, an action space, and a global reward function. The local observation space includes, but is not limited to, the link quality parameters between the satellite node and its neighbors (such as real-time latency, packet loss rate, and signal-to-noise ratio), the queue length of each outgoing port of the node, neighbor congestion status information obtained from neighbor beacons, and the Quality of Service (QoS) requirement field in the data packets to be forwarded. The action space is defined as the set of all possible next-hop forwarding neighbors for the satellite node. The global reward function is designed to guide the agent to learn cooperative optimization objectives, such as minimizing the global average end-to-end latency, minimizing the global packet loss rate, or maximizing the total network throughput.

[0041] Step S2: Centralized offline training based on Multi-Agent Proximal Policy Optimization (MAPPO).

[0042] A centralized training environment is built on a powerful ground-based computing cluster. This environment simulates the dynamic topology, link changes, and traffic patterns of the LEO satellite network. During training, each agent k has an actor policy network π. θk (a k,t |o k,t ), used to base observations on local observations. k Output selection action a k The probability distribution of (i.e., the next hop node). Simultaneously, a centralized critic network is constructed, such as V... φ (s) or Q φ (s, a), where s represents the global state (or the joint observations of all agents), and a represents the joint action of all agents. During training, all agents interact with the environment according to the current policy, and transfer their experience trajectories (including each agent's observations o) to the environment. k Action a k The global reward (r) and the next-time global state (s') are stored in the experience replay pool. By sampling experience, the value of the joint state-action sequence is evaluated using a centralized critic network, and the advantage function A is calculated. k,t And update the executor network π for each agent accordingly. θk The goal is to maximize the cumulative expected reward that takes into account the global impact. In this way, the policy network of each agent achieves implicit cooperation during the learning process, and ultimately learns a routing policy that can jointly optimize the global performance of the network.

[0043] Step S3: Lightweighting and deployment of the trained policy model, such as... Figure 3 As shown.

[0044] The agent policy network π obtained after offline centralized training converges is used to calculate the agent policy network for each agent. θk Extraction is performed. Optionally, to adapt to the resource limitations of the spaceborne platform, the model can be lightweighted, such as by weight quantization, pruning, or knowledge distillation. The processed executor policy network model file is pre-burned or loaded into the spaceborne intelligent processing unit of the corresponding LEO satellite node.

[0045] Step S4: On-orbit distributed real-time inference and routing decision.

[0046] During the LEO satellite's on-orbit operation phase, the onboard intelligent processing unit on each satellite node operates independently. For each data packet arriving at this node, the following sub-steps are executed: Step S41: Local State Awareness and Observation Vector Construction. The node, through the FPGA and CPU on its routing processing board, senses and collects local network state information in real time, including queue lengths at each outgoing port, real-time quality of neighboring links, and parsed packet header information, and integrates this information into a local observation vector conforming to the executor's network input format. k,t .

[0047] Step S42: Intelligent decision-making task distribution. The CPU on the routing processing board will distribute the constructed observation vector o k,t It is sent to the high-performance computing board via a high-speed inter-board interface.

[0048] Step S43: Actor Network Inference. The GPU or SoC on the high-performance computing board receives the observation vector o. k,t And input it into the preloaded executor policy network π θk The forward inference calculation is performed to output the probability distribution or Q-value of all possible actions (i.e., candidate next-hop nodes).

[0049] Step S44: Optimal action selection. Based on the network output, select the action a* with the highest probability or the largest Q-value. k,t The node selected as the optimal next-hop forwarding node for the data packet is then determined. This decision is transmitted back to the CPU of the routing processing board via the high-speed inter-board interface.

[0050] Step S45: High-speed packet forwarding. The CPU of the routing processing board parses the decision result and instructs its FPGA to forward the data packets from the specified output port at high speed and line speed according to the decision.

[0051] Example 2: LEO Satellite Network Routing System Based on Multi-Agent Reinforcement Learning and Heterogeneous Spaceborne Computing This embodiment provides a routing system deployed on LEO satellite nodes as the hardware foundation for implementing the method described in Embodiment 1. Figure 2As shown, the system mainly consists of three components: a routing processing board 100, a high-performance computing board 200, and a power supply unit 300.

[0052] 1. Router processing board 100 This board is the core of the data plane, responsible for the high-speed processing and forwarding of all incoming and outgoing data packets. In one specific embodiment, it includes: Central Processing Unit (CPU) 101: For example, a radiation-hardened PowerPC or ARM architecture processor is selected. The CPU runs an embedded real-time operating system and is responsible for: ① collecting low-level link and queue status information from FPGA 102; ② running the protocol stack to construct control beacons containing information such as neighbor congestion status; ③ sending the constructed observation vectors to the high-performance computing board 200 through the high-speed interface 103 and receiving the decision results returned by it; ④ issuing forwarding instructions to FPGA 102.

[0053] Field Programmable Gate Array (FPGA) 102: For example, radiation-hardened FPGAs such as Xilinx Virtex series or Intel Agilex series are selected. The FPGA is mainly responsible for line-speed data plane processing, including: ① Physical layer interface control, interfacing with the communication terminal; ② Data frame parsing and encapsulation; ③ Queue management and scheduling, real-time monitoring of queue length; ④ Executing hardware-level switching of data packets from the input port to the output port according to instructions issued by the CPU, ensuring nanosecond-level latency.

[0054] High-speed inter-board interface 103: For example, an Aurora protocol or SpaceWire link implemented by the GTH / GTY high-speed transceiver of FPGA 102. This interface is used for low-latency, high-bandwidth data and control signaling interaction with the high-performance computing board 200.

[0055] 2. High-performance computing board 200. This board is the core of the intelligent plane, responsible for performing complex AI model inference. In one specific embodiment, it includes: System-on-a-Chip (SoC) 201: For example, a Xilinx Zynq UltraScale+ MPSoC is selected, which integrates an ARM CPU and FPGA logic. The ARM CPU in the SoC is responsible for: ① communicating with the routing processing board 100 through the high-speed inter-board interface 103 to receive observation vectors and return decision results; ② task scheduling, managing the inference task queue of GPU 202; ③ system monitoring and health management.

[0056] Graphics Processing Unit (GPU) 202: For example, NVIDIA Jetson series or GPU modules hardened for space applications may be selected. The GPU is the core of AI inference computing power, responsible for loading and running pre-trained executor policy network models, and performing large-scale parallel matrix operations on the input observation vectors to quickly output routing decisions.

[0057] Memory 203 includes DDR or GDDR video memory tightly coupled to the GPU for storing neural network model parameters and intermediate calculation results; and large-capacity non-volatile memory (such as SSD or Flash) for storing the operating system, AI model files and application software.

[0058] 3. Power Supply Unit 300 This unit provides a stable and efficient power supply for the entire onboard intelligent processing unit. It includes a highly reliable, radiation-resistant DC-DC converter module 301, which converts the primary power supply provided by the satellite platform (such as 28V or 100V bus voltage) into various secondary voltages required by each board (such as core voltage 1.0V, IO voltage 3.3V, DDR voltage 1.8V, etc.), and has overvoltage, overcurrent, and undervoltage lockout protection functions.

[0059] System Connections and Operation: The power supply unit 300 is connected to the routing processing board 100 and the high-performance computing board 200 via an onboard power network, supplying power to all active devices. The routing processing board 100 and the high-performance computing board 200 are physically connected via a high-speed inter-board interface 103 (e.g., a high-speed differential signal pair on a board-to-board connector). During operation, the CPU 101 and FPGA 102 of the routing processing board 100 sense the network status in real time. The FPGA 102 is responsible for high-speed transmission and reception of data packets, while the CPU 101 is responsible for state aggregation and constructing observation vectors, and sending decision tasks to the high-performance computing board 200 via the high-speed interface 103. The SoC 201 of the high-performance computing board 200 schedules the GPU 202 to complete model inference and returns the decision results. The CPU 101 of the routing processing board 100 executes the final forwarding according to the result instructions to the FPGA 102. This clear functional division and tightly coupled hardware connection ensures the system's superior performance in terms of real-time performance, computing power, and reliability.

[0060] Preferred embodiments and application scenarios Scenario 1: Multi-path dynamic cooperative routing and congestion management under large-scale emergencies. In this scenario, an earthquake occurs in a certain area, and a large amount of emergency data (high-definition video, voice, sensor data) floods into the satellite S covering the disaster area. A S A The system of this invention, which is equipped with this invention, can sense each outgoing link (L) in real time. AB L AC L ADThe queue length and downstream congestion status are considered. The MAPPO strategy model on its high-performance computing board, having learned during training to avoid transferring congestion downstream, does not indiscriminately direct all traffic to the traditional shortest path. Instead, it dynamically allocates traffic to paths with lower congestion levels or better downstream link quality, such as selecting L... AC When the network state changes, for example, L... AC When congestion also begins, S A The model can quickly detect and adjust, redirecting new traffic to restore normal L. AD The entire process is completely autonomous and real-time, requiring no ground intervention, effectively preventing network crashes and ensuring uninterrupted emergency communication. This scenario fully demonstrates the advantages of this invention in terms of dynamic adaptability, cooperative congestion avoidance, and distributed robustness.

[0061] Scenario 2: Efficient and Reliable Backhaul of Large-Scale Earth Observation Data. In this scenario, a remote sensing satellite (SObs) needs to backhaul several terabytes of observation data to the ground. The SObs system, based on this invention, makes intelligent decisions based on observation vectors (including the estimated sustained bandwidth of each ISL, link stability, and upcoming satellite-to-ground link windows). The strategy model may choose to relay a portion of the data blocks through the currently stable, high-bandwidth ISL, while reserving a large amount of data for the upcoming direct satellite-to-ground link window. Once the satellite-to-ground link is established, the system immediately prioritizes it, utilizing its high bandwidth for full-scale downhaul. If an in-use ISL experiences an increased bit error rate due to interference, the system can quickly switch traffic to other reliable links. This scenario demonstrates the comprehensive capabilities of this invention in mission planning, opportunistic link utilization, and reliability assurance.

[0062] like Figure 2 As shown, the spaceborne intelligent processing unit proposed in this invention adopts a heterogeneous computing architecture of CPU+GPU+FPGA / SoC. This structure overcomes the bottleneck of insufficient computing power in traditional single-CPU spaceborne computers.

[0063] CPU (High-performance CPU): Responsible for overall task scheduling, control flow management, operating system operation, and some non-computation-intensive protocol processing.

[0064] GPU (High-performance GPU, Heterogeneous Series): With its powerful parallel processing capabilities, it is specifically designed to accelerate the computation of a large number of matrix operations and nonlinear activation functions in deep reinforcement learning models (such as MAPPO enforcer networks), ensuring the real-time nature of intelligent decision-making.

[0065] FPGA (Programmable Gate Array, such as the Xilinx V7 series) / Zynq-SoC: FPGAs are used on routing processing boards to implement high-speed, low-latency line-rate packet forwarding, protocol parsing, and queue management. On high-performance computing boards, Zynq-SoC combines the control flexibility of ARM processors with the programmable hardware acceleration capabilities of FPGAs, and can be used for data preprocessing, hardware implementation of specific algorithm modules, or working in conjunction with GPUs.

[0066] This heterogeneous structure (its deployment pattern in satellite constellations can be referenced) Figure 4 The "processing board" module inside each satellite enables complex intelligent routing algorithms to be executed efficiently and in real time in a spaceborne environment with limited power consumption, size, and weight (SWaP).

[0067] like Figure 5 As shown in the performance comparison chart (a), the network using the proposed solution (MAPPO algorithm) exhibits a significantly lower average packet loss rate than traditional single-agent reinforcement learning (PPO) algorithms and shortest path-based algorithms at different packet injection rates. The advantages of this invention are particularly pronounced under high-load scenarios (e.g., 500 pps), effectively avoiding buffer overflows and packet loss caused by local congestion and ensuring the reliability of data transmission. Figure 5 As shown in (b), the present invention also demonstrates superior performance in terms of end-to-end latency. Through implicit cooperation among agents, data streams are dynamically scheduled to paths with lower congestion and better link quality, avoiding excessive queuing on a single congested node or link, thereby significantly reducing the average end-to-end transmission time of data packets and meeting the stringent requirements of low-latency communication services.

[0068] In summary, this invention provides a complete and efficient intelligent routing solution for LEO satellite networks by combining multi-agent reinforcement learning (especially the MAPPO algorithm) with a heterogeneous spaceborne computing system of CPU+GPU+FPGA / SoC. It not only endows the network with dynamic adaptation and global collaboration capabilities at the algorithmic level, but also ensures that this complex intelligence can operate reliably and in real-time in the harsh space environment at the hardware level. Compared with existing technologies, this invention demonstrates significant progress in reducing end-to-end latency, minimizing packet loss, improving network resource utilization, and enhancing system robustness.

[0069] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A routing method for low Earth orbit satellite networks based on multi-agent reinforcement learning and heterogeneous spaceborne computing, characterized in that, include: Offline intensive training phase: Step A1: Define each satellite node in the low Earth orbit (LEO) satellite network as an agent and build a training environment to simulate the dynamic environment of the network; Step A2: In the training environment, a multi-agent reinforcement learning algorithm is used to conduct centralized collaborative training on all agents, so that each agent learns an executor policy model for outputting routing decisions based on local observation information; wherein, during the training process, a centralized commentator network that can obtain global state information is used to evaluate the global value of the joint actions of all agents, and the parameters of each executor policy model are updated based on the global value, so that each agent learns a cooperative routing strategy that can jointly optimize the global performance of the network; Online distributed execution phase: Step B1: Deploy the trained policy models of each executor to the corresponding satellite node's onboard heterogeneous computing system; Step B2: When each satellite node is in orbit, it uses its onboard heterogeneous computing system to perceive local network status information in real time and construct a local observation vector. Step B3: Input the local observation vector into the executor policy model deployed locally for inference, obtain the decision results of the current data packet toward each candidate next-hop node, and select the optimal next-hop forwarding action accordingly; Step B4: Based on the optimal next-hop forwarding action, the data packet is forwarded through the onboard heterogeneous computing system.

2. The low Earth orbit satellite network routing method based on multi-agent reinforcement learning and heterogeneous spaceborne computing according to claim 1, characterized in that, The multi-agent reinforcement learning algorithm used in step A2 is a multi-agent proximal policy optimization algorithm, and step A2 further includes: Step A21: Initialize the agent policy network for each agent and a centralized critic network; the agent policy network is used to output the action probability distribution based on local observations; the centralized critic network is used to perform value evaluation based on global state information or the joint actions of all agents; Step A22: Each agent interacts with the simulation environment based on the current agent policy network to generate experience trajectory data containing each agent's local observations, selected actions, global rewards, and the global state at the next time step; Step A23: Process the experience trajectory data using a centralized critic network to calculate the advantage function value of each agent's actions; Step A24: Based on the advantage function value, update the parameters of each executor's policy network by optimizing the preset objective function, and at the same time update the parameters of the centralized critic network by minimizing the error between the critic network's predicted value and the actual return; Step A25: Repeat steps A22 to A24 until each agent policy network converges, and the trained agent policy model is obtained.

3. The low Earth orbit satellite network routing method based on multi-agent reinforcement learning and heterogeneous spaceborne computing according to claim 1 or 2, characterized in that, The local observation information in step B2 includes at least one of the following: the real-time queue length of each outgoing link of this node, the average delay of each outgoing link of this node, the average packet loss rate of each outgoing link of this node, the neighbor congestion status information obtained from neighboring nodes, and the quality of service requirement field in the data packet to be forwarded.

4. The low Earth orbit satellite network routing method based on multi-agent reinforcement learning and heterogeneous spaceborne computing according to claim 1 or 2, characterized in that, The spaceborne heterogeneous computing system includes a routing processing board and a high-performance computing board connected via a high-speed inter-board interface; steps B2 to B4 further include: Step B21: The field-programmable gate array on the routing processing board receives and parses data packets, and simultaneously monitors the queue status of each outgoing link in real time; Step B22: The central processing unit on the routing processing board obtains the queue status from the field programmable gate array, and constructs the local observation vector by combining the link statistics information it maintains and the congestion status obtained from neighboring nodes. Step B23: The central processing unit on the routing processing board sends the local observation vector to the high-performance computing board through the high-speed inter-board interface; Step B24: The graphics processing unit or on-chip system on the high-performance computing board loads and runs the executor policy model, performs parallel computation on the received local observation vectors, completes model inference, and outputs the optimal next-hop forwarding action; Step B25: The high-performance computing board returns the optimal next-hop forwarding action to the central processing unit of the routing processing board through the high-speed inter-board interface; Step B26: The central processing unit of the routing processing board instructs the field-programmable gate array to encapsulate and send the data packet from the designated port according to the optimal next-hop forwarding action.

5. The low Earth orbit satellite network routing method based on multi-agent reinforcement learning and heterogeneous spaceborne computing according to claim 4, characterized in that, The model inference in step B24 further includes: using the local observation vector as input to the executor policy network, performing a nonlinear transformation through a multi-layer neural network, and outputting the action probability distribution of all candidate next-hop nodes; the optimal next-hop forwarding action is the action with the highest probability in the action probability distribution.

6. A low Earth orbit satellite network routing system based on multi-agent reinforcement learning and heterogeneous spaceborne computing, characterized in that, Deployed on each satellite node of the LEO satellite network for performing the method of any one of claims 1 to 5, the system comprises: The routing processing board includes at least: Field-programmable gate arrays (FPGAs) are used to implement line-rate parsing of data packets, protocol processing, real-time monitoring of queue status, and data packet switching and forwarding based on routing decisions. A central processing unit, coupled to the field-programmable gate array (FPGA), is used to acquire state information from the FPGA and construct local observation vectors, control communication with the high-performance computing board, and parse and issue routing decision commands to the FPGA; the high-performance computing board, connected to the routing processing board via a high-speed inter-board interface, includes at least: A graphics processing unit and / or a system-on-a-chip (SoC) are used to load and run a pre-trained executor policy model, perform inference based on local observation vectors obtained from the routing processing board, and output the optimal next-hop forwarding action; a power supply unit is connected to the routing processing board and the high-performance computing board respectively to provide power supply.

7. The low Earth orbit satellite network routing system based on multi-agent reinforcement learning and heterogeneous spaceborne computing according to claim 6, characterized in that, The pre-trained executor policy model is obtained through offline centralized collaborative training on the ground using a multi-agent proximal policy optimization algorithm. The training process utilizes a centralized commentator network capable of acquiring global state information to guide the updates of each executor policy network, thereby enabling each satellite node to learn a distributed routing strategy that can collaboratively optimize network end-to-end latency and packet loss rate.

8. The low Earth orbit satellite network routing system based on multi-agent reinforcement learning and heterogeneous spaceborne computing according to claim 6 or 7, characterized in that, The on-chip system on the high-performance computing board is also used to manage communication protocol conversion with the routing processing board, task queue scheduling, and monitoring the working status of the graphics processing unit.

9. The low Earth orbit satellite network routing system based on multi-agent reinforcement learning and heterogeneous spaceborne computing according to claim 6 or 7, characterized in that, The high-speed inter-board interface is a serial communication link based on a high-speed transceiver, used to provide low-latency data and control signaling interaction between the routing processing board and the high-performance computing board.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 5.