Multi-layer satellite network routing method and system based on multi-agent deep reinforcement learning

Through the multi-agent deep reinforcement learning method of centralized training in mid-orbit satellites and distributed execution in low-orbit satellites, the routing complexity problem caused by dynamic topological changes in LEO satellite networks is solved, and the dynamic optimization and performance improvement of the satellite network is achieved.

CN120263258APending Publication Date: 2025-07-04SHANGHAI ADVANCED RES INST CHINESE ACADEMY OF SCI
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510155528.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing satellite network routing algorithms are difficult to adapt to highly dynamic topological changes in LEO satellite networks, resulting in increased complexity of network access and routing problems, especially the lack of global information guidance in multi-layer satellite networks, resulting in local optimal problems.

Method used

Multi-agent deep reinforcement learning method is adopted to optimize routing strategies by centralized training in mid-orbit satellites and distributed execution in low-orbit satellites, combining local and global parameters. Specific steps include: low-orbit satellites are based on distributed multi-agent reinforcement learning training executor network, medium-orbit satellites train global evaluation network based on local parameters, and issuing global parameters to guide the routing decisions of low-orbit satellites.

Benefits of technology

The dynamic optimization of the satellite network is achieved, the transmission efficiency and system performance are significantly improved, local optimization problems are avoided, and communication needs under complex dynamic topology are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263258A_ABST
    Figure CN120263258A_ABST
Patent Text Reader

Abstract

The invention provides a multi-layer satellite network routing method and system based on multi-agent deep reinforcement learning. The method comprises the following steps that: a plurality of low-orbit satellites train respective executor networks by adopting a distributed multi-agent reinforcement learning algorithm to realize routing decision; and one medium-orbit satellite trains a global evaluation network according to the local parameters of the performer network uploaded by the plurality of low-orbit satellites, and provides the global parameters of the trained global evaluation network to each low-orbit satellite, so that the low-orbit satellites update respective local evaluation networks according to the global parameters. And updating respective executor networks according to the updated local evaluation network. According to the multi-layer satellite network routing method and system based on multi-agent deep reinforcement learning, dynamic optimization of a routing strategy can be realized in a medium and low orbit satellite network, and the overall performance of the satellite network is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of wireless communication, and particularly to a multi-layer satellite network routing method and system based on multi-agent deep reinforcement learning. Background Art

[0002] In June 2023, the International Telecommunication Union (ITU) completed the "Recommendation on the Framework and Overall Objectives of IMT for 2030 and Beyond". As a programmatic document for 6G, this recommendation proposed six typical 6G scenarios, namely immersive communication, ultra-massive connection, ultra-high reliability and low latency communication, ubiquitous connection, communication AI integration, and communication sensing integration, as well as four design principles, namely sustainability, ubiquitous intelligence, security / privacy / resilience, and connecting unconnected users. In particular, both the typical 6G scenarios and design principles proposed in this recommendation emphasize the beautiful vision of achieving ubiquitous connection and wide-area coverage. Due to its inherent limitations, it is difficult for traditional terrestrial networks to achieve global three-dimensional coverage. Considering the wide-area coverage advantage of satellite networks, realizing unified networking of satellite-terrestrial integration has become an inevitable choice for future networks to provide seamless connection.

[0003] Satellite communication systems have many advantages compared to terrestrial communication systems, including global coverage characteristics and one-to-many broadcast capabilities. In addition, satellite communication has a simple access method, large capacity, and faster network construction speed. Compared with terrestrial communication, the on-board processing ability of satellite networks can increase system flexibility, and multiple spot beam coverages can optimize bandwidth efficiency. In addition, satellite communication is more convenient in terms of device and maintenance, and communication loss is not limited by the length of the communication distance. When terrestrial networks fail due to natural disasters, satellite networks can also serve as an effective alternative. Therefore, satellite communication can make up for the deficiencies of terrestrial networks in many fields and complete accurate, efficient, and fast information transmission. Satellite communication systems are a very feasible solution and are attracting the attention of many countries, becoming an indispensable part of the Next Generation Network (NGN).

[0004] According to different criteria, today's satellite systems can be divided in various ways. Classified according to the satellite orbit height from low to high, the current satellites can be distinguished into Low Earth Orbit (LEO) satellites, Medium Earth Orbit (MEO) satellites, and Geostationary Earth Orbit (GEO) satellites. The LEO satellite orbit height is from 500 km to 2000 km, and the running period is 2 to 4 hours; the orbit height of MEO satellites is between 2000 km and 20000 km, and the running period is 4 to 12 hours; the orbit height of GEO satellites is 35786.6 km, and the running period is 23 hours 56 minutes 04 seconds, that is, one sidereal day. Compared with GEO satellite networks and MEO satellite networks, the LEO satellite network consists of a constellation of a large number of satellites. These satellites operate at an altitude of 500 to 2000 kilometers from the Earth's surface and orbit the Earth along the orbit at high speed. Multiple satellites are distributed on each orbital plane, and these orbital planes are parallel to each other, thus forming a complete LEO constellation system.

[0005] Due to its relatively low operating altitude, LEO satellites have the communication advantages of low latency and high rate, and are very suitable for supporting high-speed Internet access, navigation, and various space missions. In addition, the global coverage ability of LEO satellites is significantly better than that of ground networks, making them widely used in systems such as Starlink and OneWeb. However, due to the high-speed movement, LEO satellites can only cover a certain area on the ground for a short time (about 5 to 12 minutes), resulting in frequent changes in their network topology. This high degree of dynamicity is a significant difference between LEO satellite networks and ground communication networks, and also makes network access and routing problems more challenging. Therefore, it is particularly necessary to design an efficient and adaptable routing algorithm, especially in LEO satellite networks.

[0006] Machine Learning (ML) is an artificial intelligence technology that enables computers to learn and improve from data through training algorithms without explicit programming. It is mainly divided into two types: deep learning and reinforcement learning. Among them, Deep Reinforcement Learning (DRL) combines deep learning and reinforcement learning, uses deep neural networks to learn policy functions, effectively solves the problem of high dimensionality in the state space of reinforcement learning, and optimizes the parameters of neural networks through backpropagation algorithms, thus achieving more efficient and accurate decision-making. Currently, deep reinforcement learning has been widely applied in fields such as robot control, game intelligence, and natural language processing, and has also been preliminarily explored in the network field. For example, adaptive routing is performed through deep reinforcement learning algorithms to improve network optimization effects; it is applied to intrusion detection and threat analysis in network security to automatically identify and prevent attacks and improve security; in resource management, it is used for dynamic bandwidth allocation, load balancing, and traffic management, significantly optimizing network performance.

[0007] The great success of DRL has prompted researchers to turn their attention to the multi-agent field. They boldly attempt to integrate DRL methods into the Multi-Agent System (MAS) with the intention of completing numerous complex tasks in multi-agent environments, which has given rise to Multi-agent Deep Reinforcement Learning (MDRL). After several years of development and innovation, numerous algorithms, rules, and frameworks have emerged in MDRL and have been widely applied in various real-world fields. The development context from single to multi, from simple to complex, and from low-dimensional to high-dimensional indicates that MDRL is gradually becoming the hottest research and application direction in the field of machine learning and even artificial intelligence, with extremely high research value and significance.

[0008] Routing algorithms can be divided into two main categories: static routing algorithms and dynamic routing algorithms. Static routing adopts a network topology strategy based on predictable satellite orbit patterns to mitigate the impact of satellite mobility. Subsequently, offline routing calculations are performed through static topology algorithms to generate routing tables. Common static routing algorithms include methods based on virtual topologies and virtual nodes.

[0009] To address the lack of flexibility in static routing, recent dynamic routing algorithms, especially those that utilize deep reinforcement learning in networks with multiple agents, can dynamically recalculate routes based on the topological structure and collected link state information to adapt to changes in the network.

[0010] Considering that different satellites require globally optimal routing information, multi-agent deep reinforcement learning is introduced to assist the routing optimization strategy of multi-layer satellite networks, which is applied to low-earth orbit and medium-earth orbit satellite networks to enhance the network update and unification of low-earth orbit satellite networks. Multi-agent deep reinforcement learning adopts the centralized training and distributed execution technology, centrally trains and updates the global network parameters, and distributes the updated parameters. The distributed execution agents obtain environmental feedback based on their actions to ensure that each satellite has better global attributes. However, most of the existing research work has considered the single-layer satellite network assisted by deep reinforcement learning without global information guidance, and has not considered the routing method of multi-layer satellite networks assisted by multi-agent reinforcement learning. Summary of the Invention

[0011] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a routing method and system for multi-layer satellite networks based on multi-agent deep reinforcement learning, which can achieve dynamic optimization of routing strategies in medium and low-earth orbit satellite networks and significantly improve the overall performance of satellite networks.

[0012] In a first aspect, the present invention provides a routing method for multi-layer satellite networks based on multi-agent deep reinforcement learning. The method includes the following steps: multiple low-earth orbit satellites use a distributed multi-agent reinforcement learning algorithm to train their respective executor networks to make routing decisions; a medium-earth orbit satellite trains a global evaluation network based on the local parameters of the executor networks uploaded by the multiple low-earth orbit satellites, and provides the global parameters of the trained global evaluation network to each low-earth orbit satellite, so that the low-earth orbit satellites update their respective local evaluation networks according to the global parameters and update their respective executor networks according to the updated local evaluation networks.

[0013] In an implementation manner of the first aspect, each low-earth orbit satellite uses a distributed multi-agent reinforcement learning algorithm to train its respective executor network to make routing decisions, including the following steps:

[0014] Divide the satellite topology of the medium and low-earth orbit satellite network into multiple time slices;

[0015] In the static virtual topology corresponding to each time slice, each low-earth orbit satellite establishes inter-satellite links with adjacent satellites, and the adjacent satellites include the two satellites before and after in the same orbit and the two satellites in adjacent left and right orbits;

[0016] Each low-earth orbit satellite obtains the state information of the adjacent satellites based on the inter-satellite links;

[0017] Use the state information as satellite network environment information to input into the multi-agent reinforcement learning algorithm, and calculate routing actions based on the executor network.

[0018] In one implementation of the first aspect, each low-earth orbit satellite obtains the state information based on a partially observable Markov decision process; the state information includes the distance from adjacent satellites, congestion information, and energy consumption information.

[0019] In one implementation of the first aspect, the local parameters include local reward parameters and next moment state parameters.

[0020] In one implementation of the first aspect, it further includes making a routing decision based on the trained executor network according to the real-time satellite network environment information and real-time satellite topology.

[0021] In a second aspect, the present invention provides a multi-agent deep reinforcement learning-based multi-layer satellite network routing system. The multi-agent deep reinforcement learning-based multi-layer satellite network routing system includes a plurality of low-earth orbit satellites and a medium-earth orbit satellite;

[0022] Each low-earth orbit satellite is used to train its own executor network by using a distributed multi-agent reinforcement learning algorithm to achieve routing decision-making;

[0023] The medium-earth orbit satellite is used to train a global evaluation network according to the local parameters of the executor network uploaded by the plurality of low-earth orbit satellites, and provide the global parameters of the trained global evaluation network to each low-earth orbit satellite, so that the low-earth orbit satellite updates its local evaluation network according to the global parameters, and updates its executor network according to the updated local evaluation network.

[0024] In one implementation of the second aspect, each low-earth orbit satellite uses a distributed multi-agent reinforcement learning algorithm to train its own executor network to achieve routing decision-making, including the following steps:

[0025] Divide the satellite topology of the medium-low earth orbit satellite network into multiple time slices;

[0026] In the static virtual topology corresponding to each time slice, each low-earth orbit satellite establishes an inter-satellite link with adjacent satellites. The adjacent satellites include the two satellites in front and behind on the same orbit and the two satellites on adjacent left and right orbits;

[0027] Each low-earth orbit satellite obtains the state information of the adjacent satellites based on the inter-satellite link;

[0028] Input the state information as satellite network environment information into the multi-agent reinforcement learning algorithm, and calculate the routing action based on the executor network.

[0029] In one implementation of the second aspect, each low-earth orbit satellite obtains the state information based on a partially observable Markov decision process; the state information includes the distance from adjacent satellites, congestion information, and energy consumption information.

[0030] In one implementation of the second aspect, the local parameters include local reward parameters and next moment state parameters.

[0031] In one implementation of the second aspect, the LEO satellite is further configured to complete routing decisions based on the trained executor network according to real-time satellite network environment information and real-time satellite topology.

[0032] As described above, the multi-agent deep reinforcement learning-based multi-layer satellite network routing method and system of the present invention have the following beneficial effects:

[0033] (1) Deploy the multi-agent reinforcement learning technology in both the medium-earth orbit and low-earth orbit satellite networks simultaneously, and combine the methods of centralized training and distributed execution, so as to achieve the dynamic optimization of satellite routing;

[0034] (2) With the goal of minimizing the global transmission delay of the system, under the constraints of packet loss rate and transmission rate, adopt a multi-agent reinforcement learning framework to maximize the local cumulative reward for LEO satellites;

[0035] (3) By deploying a centralized training mechanism in the medium-earth orbit satellite, integrating parameters such as local rewards uploaded by the low-earth orbit satellite, and sending the optimized global model parameters to the low-earth orbit satellite, providing more global information guidance, thereby significantly improving the training efficiency, optimizing the local model performance, and effectively avoiding the local optimum problem that may occur in distributed decision-making;

[0036] (4) Can significantly improve the overall performance of the satellite network and meet the communication requirements under complex dynamic topologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 Shows a schematic structural diagram of the multi-agent deep reinforcement learning-based multi-layer satellite network routing system of the present invention in an embodiment;

[0038] Figure 2 Shows a flowchart of the multi-agent deep reinforcement learning-based multi-layer satellite network routing method of the present invention in an embodiment;

[0039] Figure 3 Shows a schematic structural diagram of the multi-agent deep reinforcement learning-based multi-layer satellite network routing system of the present invention in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The following specific examples illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0041] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0042] Building a low-earth orbit satellite constellation network is the key to building an integrated space-air-ground network. In the communication transmission system of a multi-layer satellite network, low-earth orbit satellites, as relay nodes for transmission, play an important role in signal transmission. By distributively training a multi-agent reinforcement learning model on low-earth orbit satellites and uploading local partial reward and other parameters to medium-earth orbit satellites, after centralized training by the medium-earth orbit satellites, the optimized model parameters are sent down to the low-earth orbit satellites to guide the actions of the low-earth orbit satellites. The medium-earth orbit satellites provide global information to avoid the local optimum problem and further improve the routing strategy design of the low-earth orbit satellites. Therefore, the communication method and system of the multi-layer satellite network of the present invention adopt a combination of centralized training and distributed execution to optimize the communication performance of the multi-layer satellite network.

[0043] Next, the technical solutions in the embodiments of the present invention will be described in detail with reference to the accompanying drawings in the embodiments of the present invention.

[0044] As Figure 1 shown, the routing of the multi-layer satellite network based on multi-agent deep reinforcement learning of the present invention consists of a medium-earth orbit satellite layer and a low-earth orbit satellite layer. In the low-earth orbit satellite layer, each satellite is equipped with computing resources and a small server, and establishes communication with adjacent satellites above, below, left, and right through laser links; while the medium-earth orbit satellites serve as a higher-level computing and coordination center. The multi-agent reinforcement learning algorithm is distributively deployed on the medium-earth orbit satellites and the low-earth orbit satellites. The low-earth orbit satellites are responsible for training the executor network and locally optimizing their routing decisions; the medium-earth orbit satellites centrally update the evaluator network and form a global evaluation network by integrating local reward and other parameters uploaded by the low-earth orbit satellites. Subsequently, the medium-earth orbit satellites send down the global parameters of the updated global evaluation network to all low-earth orbit satellites to guide their local training and execution, so as to ensure that the executor network has global information during the routing optimization process, avoid falling into the local optimum, and at the same time improve the transmission efficiency and performance of the entire system.

[0045] As Figure 2 shown, in one embodiment, the multi-layer satellite network routing method based on multi-agent deep reinforcement learning of the present invention includes steps S1-S2.

[0046] Step S1: Multiple low-earth orbit satellites use a distributed multi-agent reinforcement learning algorithm to train their respective executor networks to make routing decisions.

[0047] Specifically, the Decentralized Partially Observable Markov Decision Process (Dec-POMDP) method can well describe the collaborative decision-making problem of multiple agents under limited information. Specifically, in Dec-POMDP, multiple agents interact with the environment together, and each agent can only observe the local environmental state. Agents make decisions under incomplete information through a cooperative strategy to optimize the global goal. In this framework, each agent independently selects an action, and the environment then updates the global state according to the actions of all agents and generates a corresponding reward signal. The goal of the agent is to find a globally optimal strategy through coordinated cooperation to maximize the long-term cumulative global reward. Therefore, the satellite network system based on Dec-POMDP has better adaptability when facing complex network scenarios. In the present invention, each satellite can be regarded as an independent agent, makes routing decisions based on the locally observable information, and optimizes the routing strategy of the entire network through cooperation with other satellites.

[0048] Among them, as an independent agent, each low-earth orbit satellite, based on Dec-POMDP, uses the deployed multi-agent reinforcement learning algorithm to train the executor network with local information reinforcement learning. The low-earth orbit satellite agent makes routing decisions by real-time monitoring the network state around it, and optimizes the executor network parameters according to the reward value and the state parameters of the next moment feedback by the environment.

[0049] In one embodiment, each low-earth orbit satellite uses a distributed multi-agent reinforcement learning algorithm to train its respective executor network to make routing decisions, including the following steps:

[0050] a) Divide the satellite topology of the medium and low-earth orbit satellite network into multiple time slices.

[0051] b) In the static virtual topology corresponding to each time slice, each low-earth orbit satellite establishes an inter-satellite link with adjacent satellites, and the adjacent satellites include the two satellites before and after in the same orbit and the two satellites in adjacent left and right orbits.

[0052] c) Each low-earth orbit satellite obtains the status information of the adjacent satellite based on the inter-satellite link. Among them, each low-earth orbit satellite obtains the status information based on the partially observable Markov decision process; the status information includes the distance from the adjacent satellite, congestion information, and energy consumption information.

[0053] d) Use the status information as the satellite network environment information and input it into the multi-agent reinforcement learning algorithm, and calculate the routing action based on the executor network. Introducing the multi-agent reinforcement learning algorithm into the multi-layer satellite network system can improve the intelligent decision-making performance of each independent satellite, realize the optimal path search under the dynamically changing satellite topology state at each moment, so as to minimize the average transmission delay while meeting the packet loss rate and transmission rate limitations.

[0054] Step S2: A medium-earth orbit satellite trains the global evaluation network according to the local parameters of the executor network uploaded by the multiple low-earth orbit satellites, and provides the global parameters of the trained global evaluation network to each low-earth orbit satellite, so that the low-earth orbit satellite updates its respective local evaluation network according to the global parameters, and updates its respective executor network according to the updated local evaluation network.

[0055] Specifically, the medium-earth orbit satellite receives the local parameters sent by the low-earth orbit satellite based on the inter-satellite link, that is, the local reward parameter and the next moment state parameter, and performs centralized training to optimize the global model, so as to provide global optimization guidance for the low-earth orbit satellite. Among them, the medium-earth orbit satellite avoids falling into the local optimal solution by integrating the local parameters in the distributed learning. After the medium-earth orbit satellite completes the centralized training, it sends the updated global parameters to the low-earth orbit satellite. The low-earth orbit satellite updates its respective local evaluation network according to the global parameters, and updates its respective executor network according to the updated local evaluation network, so as to further optimize the routing strategy.

[0056] When the model training reaches the preset convergence condition, the multi-layer satellite network routing method based on multi-agent deep reinforcement learning of the present invention further includes completing the routing decision based on the trained executor network according to the real-time satellite network environment information and the real-time satellite topology, so as to realize the dynamic optimization of satellite routing.

[0057] The protection scope of the multi-layer satellite network routing method based on multi-agent deep reinforcement learning described in the embodiments of the present invention is not limited to the execution order of the steps listed in this embodiment. Any solution realized by adding or subtracting steps of the prior art and replacing steps according to the principle of the present invention is included in the protection scope of the present invention.

[0058] An embodiment of the present invention further provides a multi - layer satellite network routing system based on multi - agent deep reinforcement learning. The multi - layer satellite network routing system based on multi - agent deep reinforcement learning can implement the multi - layer satellite network routing method of the present invention. However, the implementation device of the multi - layer satellite network routing system based on multi - agent deep reinforcement learning of the present invention includes, but is not limited to, the structure of the multi - layer satellite network routing system listed in this embodiment. Any structural deformation and replacement of the prior art made according to the principle of the present invention are included in the protection scope of the present invention.

[0059] In one embodiment, the multi - layer satellite network routing system based on multi - agent deep reinforcement learning of the present invention includes a plurality of low - earth orbit satellites and a medium - earth orbit satellite.

[0060] Each low - earth orbit satellite is used to train its own executor network by using a distributed multi - agent reinforcement learning algorithm to achieve routing decisions.

[0061] The medium - earth orbit satellite is used to train a global evaluation network according to the local parameters of the executor networks uploaded by the plurality of low - earth orbit satellites, and provide the global parameters of the trained global evaluation network to each low - earth orbit satellite, so that the low - earth orbit satellite updates its local evaluation network according to the global parameters, and updates its executor network according to the updated local evaluation network.

[0062] The multi - layer satellite network routing system based on multi - agent deep reinforcement learning of the present invention will be further elaborated below through specific embodiments.

[0063] As Figure 3 shown, the multi - layer satellite network routing system based on multi - agent deep reinforcement learning of the present invention is composed of a layer of low - earth orbit satellite network and a layer of medium - earth orbit satellite network. In the low - earth orbit satellite layer, each satellite is equipped with computing power resources and a small server, and establishes communication with adjacent satellites above, below, left, and right through laser links. The multi - agent reinforcement learning algorithm is distributedly deployed on the low - earth orbit satellites. Each low - earth orbit satellite obtains local state information such as the distance and congestion situation of adjacent satellites based on a partially observable Markov decision process (Dec - POMDP), inputs the observed state information into the local executor network to calculate action decisions, and updates the parameters of the local executor network through temporal difference residuals according to the reward information.

[0064] Meanwhile, the medium-earth orbit (MEO) satellite conducts centralized training on the global evaluation network based on parameter information such as local rewards uploaded by the low-earth orbit (LEO) satellites. Specifically, each LEO satellite regularly uploads the parameters such as the locally trained rewards to the MEO satellite. The MEO satellite uses the local parameters uploaded by each LEO satellite to update the global evaluation network and then distributes the updated global parameters to the LEO satellites. This collaborative mechanism of centralized training and distributed execution ensures that the action decisions of the LEO satellites can not only adapt to local environmental changes but also obtain global optimization guidance, thus avoiding the local optimum problem and improving the overall performance of the system.

[0065] By introducing multi-agent reinforcement learning technology into the multi-layer satellite network, the present invention enables the LEO satellites to achieve distributed learning in local environmental observation, action selection, and reward feedback. At the same time, through the centralized training of the MEO satellite, global guidance information is provided for the LEO satellites. This not only significantly improves the model convergence speed, reduces energy consumption, but also optimizes the routing strategy, enhancing the system stability and transmission efficiency.

[0066] In summary, in the multi-layer satellite network communication method based on multi-agent reinforcement learning of the present invention, the satellite network system adopts the actor-critic strategy and the architecture of centralized training and distributed execution. The LEO satellites achieve local optimization through distributed learning; the MEO satellite integrates the parameters uploaded by the LEO satellites through centralized training, optimizes the global evaluation network, and distributes it to the LEO satellites. This collaborative mechanism provides a new optimization guidance for the routing selection of the LEO satellites, effectively reducing the end-to-end delay and packet loss rate, accelerating the model convergence process, and enhancing the overall transmission rate and performance of the network.

[0067] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A routing method for a multi-layer satellite network based on multi-agent deep reinforcement learning, characterized in that: The method includes the following steps: Multiple low Earth orbit (LEO) satellites use a distributed multi-agent reinforcement learning algorithm to train their respective executor networks to achieve routing decisions; A medium Earth orbit (MEO) satellite trains a global evaluation network based on the local parameters of the executor networks uploaded by the multiple LEO satellites, and provides the global parameters of the trained global evaluation network to each LEO satellite, so that the LEO satellites update their respective local evaluation networks according to the global parameters, and update their respective executor networks according to the updated local evaluation networks.

2. The multi-layer satellite network routing method based on multi-agent deep reinforcement learning according to claim 1, characterized in that: Each LEO satellite uses a distributed multi-agent reinforcement learning algorithm to train its respective executor network to achieve routing decisions, including the following steps: Divide the satellite topology of the LEO-MEO satellite network into multiple time slices; In the static virtual topology corresponding to each time slice, each LEO satellite establishes inter-satellite links with adjacent satellites, and the adjacent satellites include the two front and rear satellites in the same orbit and the two satellites in adjacent left and right orbits; Each LEO satellite obtains the state information of the adjacent satellites based on the inter-satellite links; Use the state information as satellite network environment information to input into the multi-agent reinforcement learning algorithm, and calculate routing actions based on the executor network.

3. The multi-layer satellite network routing method based on multi-agent deep reinforcement learning according to claim 2, wherein: Each LEO satellite obtains the state information based on a partially observable Markov decision process; the state information includes the distance from adjacent satellites, congestion information, and energy consumption information.

4. The multi-layer satellite network routing method based on multi-agent deep reinforcement learning according to claim 1, wherein: The local parameters include local reward parameters and next moment state parameters.

5. The routing method for a multi-layer satellite network based on multi-agent deep reinforcement learning according to claim 1, characterized in that: It also includes completing routing decisions based on the trained executor network according to real-time satellite network environment information and real-time satellite topology.

6. A multi-layer satellite network routing system based on multi-agent deep reinforcement learning, characterized in that: The multi-layer satellite network routing system based on multi-agent deep reinforcement learning includes multiple LEO satellites and one MEO satellite; Each LEO satellite is used to use a distributed multi-agent reinforcement learning algorithm to train its respective executor network to achieve routing decisions; The MEO satellite is used to train a global evaluation network based on the local parameters of the executor networks uploaded by the multiple LEO satellites, and provide the global parameters of the trained global evaluation network to each LEO satellite, so that the LEO satellites update their respective local evaluation networks according to the global parameters, and update their respective executor networks according to the updated local evaluation networks.

7. The multi-layer satellite network routing system based on multi-agent deep reinforcement learning according to claim 6, characterized in that: Each LEO satellite uses a distributed multi-agent reinforcement learning algorithm to train its respective executor network to achieve routing decisions, including the following steps: Divide the satellite topology of the LEO-MEO satellite network into multiple time slices; In the static virtual topology corresponding to each time slice, each LEO satellite establishes inter-satellite links with adjacent satellites, and the adjacent satellites include the two front and rear satellites in the same orbit and the two satellites in adjacent left and right orbits; Each LEO satellite obtains the state information of the adjacent satellites based on the inter-satellite links; Use the state information as satellite network environment information to input into the multi-agent reinforcement learning algorithm, and calculate routing actions based on the executor network.

8. The multi-layer satellite network routing system based on multi-agent deep reinforcement learning according to claim 7, wherein: Each LEO satellite obtains the state information based on a partially observable Markov decision process; the state information includes the distance from adjacent satellites, congestion information, and energy consumption information.

9. The multi-layer satellite network routing system based on multi-agent deep reinforcement learning according to claim 6, characterized in that: The local parameters include a local reward parameter and a next moment state parameter.

10. The multi-layer satellite network routing system based on multi-agent deep reinforcement learning according to claim 6, wherein: The low-earth orbit satellite is further configured to complete a routing decision based on the trained executor network according to real-time satellite network environment information and real-time satellite topology.

Citation Information

Cited By

  • Inter-satellite link dynamic control method and device of multilayer satellite constellation and storage medium

    CN121441366A

  • Federal learning-based routing policy network training method and device

    CN121509301A

  • Training Method and Apparatus for Routing Policy Networks Based on Federated Learning

    CN121509301B

  • Internet of Things low earth orbit satellite communication data interaction generation method and system

    CN121966693A