Multi-layer satellite network routing method and system for deep reinforcement learning and federated learning
By combining deep reinforcement learning and federated learning technologies in satellite networks, dynamic routing decisions are achieved, solving the problem that traditional routing algorithms are difficult to adapt to network changes, and significantly improving the performance and efficiency of satellite networks.
Patent Information
- Application Number
- CN202510155550.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-06
AI Technical Summary
The existing satellite network routing algorithms are difficult to achieve dynamic optimization, especially in LEO satellite networks. Due to frequent changes in topology, traditional static routing algorithms cannot effectively adapt to network changes.
Combining deep reinforcement learning and federated learning technology, multiple low-orbit satellites use distributed deep reinforcement learning algorithms to train executor networks and evaluation networks. The mid-orbit satellite generates a global evaluation network through federated learning algorithms, and issues global evaluation parameters to update the evaluation network and executor network of low-orbit satellites.
Dynamic optimization of satellite routing is achieved, the overall performance of the satellite network is significantly improved, the end-to-end delay and packet loss rate are reduced, the model convergence process is accelerated, and the transmission rate is improved.
Smart Images

Figure CN120110976A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wireless communications, and in particular to a multi-layer satellite network routing method and system for deep reinforcement learning and federated learning. Background Art
[0002] In June 2023, the International Telecommunication Union (ITU) completed the "Recommendation on the Framework and Overall Objectives of IMT for 2030 and Beyond". As a programmatic document for 6G, the proposal proposes six typical 6G scenarios: immersive communication, ultra-large-scale connection, ultra-high reliability and low-latency communication, ubiquitous connection, communication AI integration, and communication perception integration, as well as four design principles: sustainability, ubiquitous intelligence, security / privacy / resilience, and connecting unconnected users. In particular, the typical 6G scenarios and design principles proposed in the proposal emphasize the beautiful vision of achieving ubiquitous connectivity and wide-area coverage. Traditional terrestrial networks are difficult to achieve global three-dimensional coverage due to their inherent limitations. Considering the wide-area coverage advantage of satellite networks, realizing a unified network that integrates satellites and ground has become an inevitable choice for future networks to provide seamless connectivity.
[0003] Satellite communication systems have many advantages over terrestrial communication systems, including global coverage and point-to-multipoint broadcasting capabilities. In addition, satellite communication has simple access methods, large capacity, and faster network construction. Compared with terrestrial communications, the on-board processing capabilities of satellite networks can increase system flexibility, and multiple spot beam coverage can optimize bandwidth efficiency. In addition, satellite communications are more convenient in terms of installation and maintenance, and communication losses are not limited by the length of communication distance. Satellite networks can also serve as an effective alternative when terrestrial networks fail due to natural disasters. Therefore, satellite communications can make up for the shortcomings of terrestrial networks in many fields and complete accurate, efficient and fast information transmission. Satellite communication systems are a very feasible solution and are attracting the attention of many countries, becoming an indispensable part of the Next Generation Network (NGN).
[0004] According to different standards, today's satellite systems can be divided into many categories. According to the satellite orbit altitude from low to high, the current satellites can be divided into low earth orbit satellites (LEO), medium earth orbit satellites (MEO) and geostationary earth orbit satellites (GEO). The orbit altitude of LEO satellites is 500km to 2000km, and the operation period is 2 to 4 hours; the orbit altitude of MEO satellites is 2000km to 20000km, and the operation period is 4 to 12 hours; the orbit altitude of GEO satellites is 35786.6km, and the operation period is 23 hours, 56 minutes and 04 seconds, that is, one sidereal day. Compared with the GEO satellite network and the MEO satellite network, the LEO satellite network is composed of a constellation of a large number of satellites. These satellites operate at an altitude of 500 to 2000 kilometers above the earth's surface and orbit the earth at a high speed. Multiple satellites are distributed on each orbital plane, and these orbital planes are parallel to each other, thus forming a complete LEO constellation system.
[0005] Unlike GEO satellite networks and MEO satellite networks, LEO satellite networks consist of a constellation of multiple small satellites. These satellites operate at altitudes of 500 to 2000 kilometers above the Earth's surface, and each satellite orbits the Earth along a set of orbital planes. There are multiple satellites on each orbital plane that successively follow the Earth's orbit, and all orbital planes are parallel to each other. The LEO satellite network is close to the Earth and is very suitable for high-speed, low-latency communications, navigation, and space missions. In addition, LEO satellites have low transmission delays and extensive global coverage, so they have always been a hot research area, such as Starlink and OneWeb. Since LEO satellites operate at high speeds, it takes about 5 to 12 minutes to cover an area each time, which leads to frequent changes in satellite topology and makes network access highly dynamic, which is the main difference from terrestrial communication networks. Considering routing issues in satellite communications is very challenging because it is not a simple process to apply terrestrial routing protocols to satellite networks. Therefore, it is necessary to design a novel and appropriate routing algorithm, especially in LEO satellite networks.
[0006] Machine Learning (ML) is an artificial intelligence technology that enables computers to learn and improve from data without explicit programming through training algorithms. It is mainly divided into two types: deep learning and reinforcement learning. Among them, deep reinforcement learning (DRL) combines deep learning and reinforcement learning, using deep neural networks to learn policy functions, effectively solving the problem of too high spatial dimensions in reinforcement learning, and optimizing neural network parameters through back-propagation algorithms to achieve more efficient and accurate decision-making. At present, deep reinforcement learning has been widely used in fields such as robot control, game intelligence, and natural language processing, and has been initially explored in the network field, such as using deep reinforcement learning algorithms for adaptive routing to improve network optimization effects; it is applied to intrusion detection and threat analysis in network security to automatically identify and prevent attacks and improve security; in terms of resource management, it is used to dynamically allocate bandwidth, load balancing, and traffic management to significantly optimize network performance.
[0007] The huge success of DRL has prompted researchers to turn their attention to the field of multi-agents. They boldly tried to integrate the DRL method into the Multi-Agent System (MAS) in order to complete many complex tasks in the multi-agent environment. This gave birth to Multi-agent Deep Reinforcement Learning (MDRL). After several years of development and innovation, MDRL has produced many algorithms, rules, and frameworks, and has been widely used in various real-world fields. The development from single to multiple, from simple to complex, and from low-dimensional to high-dimensional shows that MDRL is gradually becoming the hottest research and application direction in the field of machine learning and even artificial intelligence, with extremely high research value and significance.
[0008] Routing algorithms can be divided into two categories: static routing algorithms and dynamic routing algorithms. Static routing is a network topology strategy based on predictable satellite orbit patterns to mitigate the impact of satellite mobility. Subsequently, the static topology algorithm performs offline routing calculations to generate routing tables. Common static routing algorithms include virtual topology-based and virtual node-based methods.
[0009] To address the lack of flexibility of static routing, recent dynamic routing algorithms, especially those leveraging deep reinforcement learning in similar fields, are able to dynamically recalculate routes based on the topology and collected link state information to adapt to changes in the network.
[0010] Considering the similarities between different satellites, federated learning is introduced to assist the routing optimization strategy of multi-layer satellite networks. Deep reinforcement learning is combined with federated learning and applied to low-orbit and medium-orbit satellite networks to enhance the network update and unification of low-orbit satellite networks. According to the network parameters obtained from low-orbit satellites by federated learning, the global network parameters are updated centrally and sent down after the update to ensure that each satellite has better global properties. However, most of the existing research works consider deep reinforcement learning to assist single-layer satellite networks, and do not consider multi-layer satellite network communication methods based on deep reinforcement learning and federated learning. Summary of the invention
[0011] In view of the shortcomings of the prior art mentioned above, the purpose of the present invention is to provide a multi-layer satellite network routing method and system based on deep reinforcement learning and federated learning, which realizes dynamic optimization of satellite routing based on federated learning technology and deep reinforcement learning technology, and significantly improves the overall performance of the satellite network.
[0012] In a first aspect, the present invention provides a multi-layer satellite network routing method of deep reinforcement learning and federated learning, the method comprising the following steps: a plurality of low-orbit satellites adopt a distributed deep reinforcement learning algorithm to train their respective executor networks to implement routing decisions, and train their respective evaluation networks to implement local evaluation of the executor networks; a medium-orbit satellite generates a global evaluation network based on a federated learning algorithm according to the evaluation network parameters of the evaluation network uploaded by the plurality of low-orbit satellites, and provides the global evaluation parameters of the trained global evaluation network to each low-orbit satellite, so that the low-orbit satellites update their respective evaluation networks according to the global evaluation parameters, and update their respective executor networks according to the updated evaluation networks.
[0013] In an implementation of the first aspect, multiple low-orbit satellites use a distributed deep reinforcement learning algorithm to train their respective actor networks to implement routing decisions, and train their respective evaluation networks to implement local evaluation of the actor networks, including the following steps:
[0014] Divide the satellite topology of the medium and low orbit satellite network into multiple time slices;
[0015] In the static virtual topology corresponding to each time slice, each low-orbit satellite establishes an inter-satellite link with an adjacent satellite, wherein the adjacent satellites include two satellites in front and behind the same orbit and two satellites in adjacent left and right orbits;
[0016] Each low-orbit satellite obtains the status information of the adjacent satellite based on the inter-satellite link;
[0017] The state information is input into the deep reinforcement learning algorithm as satellite network environment information, local update parameters of the executor network are calculated, and local update parameters of the evaluation network are calculated.
[0018] In an implementation of the first aspect, each low-orbit satellite obtains the status information based on a partially observable Markov decision process; the status information includes the distance to adjacent satellites, congestion information, and energy consumption information.
[0019] In an implementation of the first aspect, the federated learning algorithm uses a weighted average algorithm to merge evaluation parameters of multiple low-orbit satellites into global evaluation parameters.
[0020] In an implementation of the first aspect, it also includes completing a routing decision based on a trained executor network according to real-time satellite network environment information and real-time satellite topology.
[0021] In a second aspect, the present invention provides a multi-layer satellite network routing system of deep reinforcement learning and federated learning, the system comprising a plurality of low-orbit satellites and a medium-orbit satellite;
[0022] The plurality of low-orbit satellites are used to train respective actor networks using a distributed deep reinforcement learning algorithm to implement routing decisions, and train respective evaluation networks to implement local evaluation of the actor networks;
[0023] The medium-orbit satellite is used to generate a global evaluation network based on a federated learning algorithm according to the evaluation network parameters of the evaluation network uploaded by the multiple low-orbit satellites, and provide the global evaluation parameters of the trained global evaluation network to each low-orbit satellite, so that the low-orbit satellites update their respective evaluation networks according to the global evaluation parameters, and update their respective executor networks according to the updated evaluation networks.
[0024] In an implementation of the second aspect, multiple low-orbit satellites use a distributed deep reinforcement learning algorithm to train their respective actor networks to implement routing decisions, and train their respective evaluation networks to implement local evaluation of the actor networks, including the following steps:
[0025] Divide the satellite topology of the medium and low orbit satellite network into multiple time slices;
[0026] In the static virtual topology corresponding to each time slice, each low-orbit satellite establishes an inter-satellite link with an adjacent satellite, wherein the adjacent satellites include two satellites in front and behind the same orbit and two satellites in adjacent left and right orbits;
[0027] Each low-orbit satellite obtains the status information of the adjacent satellite based on the inter-satellite link;
[0028] The state information is input into the deep reinforcement learning algorithm as satellite network environment information, local update parameters of the executor network are calculated, and local update parameters of the evaluation network are calculated.
[0029] In an implementation of the second aspect, each low-orbit satellite obtains the status information based on a partially observable Markov decision process; the status information includes the distance to adjacent satellites, congestion information, and energy consumption information.
[0030] In an implementation of the second aspect, the federated learning algorithm uses a weighted average algorithm to merge evaluation parameters of multiple low-orbit satellites into global evaluation parameters.
[0031] In an implementation of the second aspect, the low-orbit satellite is also used to complete routing decisions based on a trained executor network according to real-time satellite network environment information and real-time satellite topology.
[0032] As described above, the multi-layer satellite network routing method and system of deep reinforcement learning and federated learning of the present invention have the following beneficial effects:
[0033] (1) Dynamic optimization of satellite routing is achieved based on federated learning technology and deep reinforcement learning technology, which significantly improves the overall performance of satellite networks;
[0034] (2) With the goal of minimizing the global transmission delay of the system, deep reinforcement learning is used to maximize the cumulative reward of a single low-orbit satellite under the constraints of packet loss rate and transmission rate. At the same time, federated learning is introduced to integrate and distribute the locally trained low-orbit satellite network parameters to ensure that the low-orbit satellite has more global information, thereby greatly improving the training efficiency and optimizing the parameter performance of the local model;
[0035] (3) The deep reinforcement learning algorithm is deployed on the low-orbit satellite network based on the actor-evaluator strategy, while the federated learning algorithm is deployed on the medium-orbit satellite network based on the federated average algorithm. The combination of deep reinforcement learning and federated learning technology provides new guidance for the routing selection of low-orbit satellites, reduces end-to-end delay and packet loss rate, accelerates the model convergence process, and improves the transmission rate.
[0036] (4) It is possible to achieve the optimal path search under the satellite topology state at each moment, thereby minimizing the average transmission delay while meeting the packet loss rate and transmission rate constraints. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a schematic diagram showing the structure of a multi-layer satellite network routing system of deep reinforcement learning and federated learning in one embodiment of the present invention;
[0038] Figure 2 Shown is a flowchart of a multi-layer satellite network routing method of deep reinforcement learning and federated learning in one embodiment of the present invention;
[0039] Figure 3Shown is a schematic diagram of the structure of a multi-layer satellite network routing system for deep reinforcement learning and federated learning in one embodiment of the present invention. DETAILED DESCRIPTION
[0040] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0041] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0042] Building a low-orbit satellite constellation network is the key to building an integrated air-space-ground network. In the communication transmission system of the multi-layer satellite network, the low-orbit satellite, as a relay node for transmission, plays an important role in the transmission of signals, and optimizes the local network through its own deep reinforcement learning algorithm. The medium-orbit satellite receives the satellite network parameters uploaded by the low-orbit satellite, and uses the federated learning algorithm to update and distribute the global parameters to achieve the optimal design of the low-orbit satellite routing. Among them, the low-orbit satellite updates the local rating network and executor network through environmental observation, action selection, reward feedback. The medium-orbit satellite receives the parameters uploaded by the low-orbit satellite, aggregates and updates the parameters, and distributes them to each low-orbit satellite as a global rating network. Therefore, the multi-layer satellite network communication method and system of the present invention achieves the improvement of the low-orbit satellite routing strategy through the local parameter calculation of deep reinforcement learning and the global parameter integration of federated learning.
[0043] like Figure 1As shown, the multi-layer satellite network routing system of deep reinforcement learning and federated learning of the present invention is composed of a layer of low-orbit satellite network and a layer of medium-orbit satellite network. In the low-orbit satellite network layer, each low-orbit satellite is equipped with computing resources and small servers, and will establish communication laser links with satellites around it. The present invention deploys the deep reinforcement learning algorithm in a distributed manner on each low-orbit satellite, and updates the parameters of the evaluation network and the parameters of the executor network locally. At the same time, the federated learning algorithm is deployed on the medium-orbit satellite. Each low-orbit satellite layer uploads its own evaluation network parameters, integrates the parameters through the federated learning algorithm, forms a global evaluation network, and then sends the global evaluation network parameters to the low-orbit satellite to update the evaluation network of the low-orbit satellite, thereby ensuring that the executor network has global guidance and will not easily fall into the local optimum, thereby improving the overall system performance.
[0044] The technical solutions in the embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0045] like Figure 2 As shown, in one embodiment, the multi-layer satellite network routing method of deep reinforcement learning and federated learning of the present invention includes steps S1 and S2.
[0046] Step S1: Multiple low-orbit satellites use a distributed deep reinforcement learning algorithm to train their respective executor networks to implement routing decisions, and train their respective evaluation networks to implement local evaluation of the executor networks.
[0047] Specifically, deep reinforcement learning algorithms can be used in the field of network decision optimization, and most problems can be expressed as Markov decision processes (MDPs). Deep reinforcement learning is a method in which an agent continuously interacts with the environment to learn strategies. It learns through repeated interactions with the goal of maximizing cumulative rewards. During the learning process, the agent's decision follows the Markov decision process. In this random process, the agent interacts with the environment in a series of time steps. At each time step, the agent chooses an action to perform, and the environment then moves to a new state and provides feedback based on a reward signal. The goal of the agent is to find a strategy that maximizes its accumulated rewards in the long term.
[0048] The Decentralized Partially Observable Markov Decision Process (Dec-POMDP) method can well describe the collaborative decision-making problem of multiple agents under limited information. In the present invention, as an independent agent, the low-orbit satellite adopts a partially observable Markov decision process, collects the status information of the current LEO satellite network according to the observable range, makes routing decisions based on the deployed deep reinforcement learning algorithm, and obtains the rewards returned by the environment and the environmental state of the next stage. In the training stage, the agent randomly samples a batch of experience tuples from the experience buffer, and updates the weights of each low-orbit satellite agent neural network by calculating the loss value and gradient until the executor model converges.
[0049] In one embodiment, multiple low-orbit satellites use a distributed deep reinforcement learning algorithm to train their respective actor networks to implement routing decisions, and train their respective evaluation networks to implement local evaluation of the actor networks, including the following steps:
[0050] 11) Divide the satellite topology of the medium and low orbit satellite network into multiple time slices.
[0051] 12) In the static virtual topology corresponding to each time slice, each low-orbit satellite establishes an inter-satellite link with adjacent satellites, wherein the adjacent satellites include the two preceding and following satellites in the same orbit and the two adjacent left and right satellites in the same orbit.
[0052] 13) Each low-orbit satellite obtains the state information of the adjacent satellite based on the intersatellite link. Each low-orbit satellite obtains the state information based on a partially observable Markov decision process; the state information includes the distance to the adjacent satellite, congestion information, and energy consumption information.
[0053] 14) Input the state information as satellite network environment information into the deep reinforcement learning algorithm, calculate the local update parameters of the executor network, and calculate the local update parameters of the evaluation network. Each low-orbit satellite inputs the observed satellite network environment information into the deep reinforcement learning algorithm, obtains the next node action of the corresponding destination node, obtains the corresponding reward feedback in the environment, and the local evaluation network and executor network are sequentially updated according to the time difference residual, thereby realizing the corresponding network training.
[0054] Step S2: A medium-orbit satellite generates a global evaluation network based on a federated learning algorithm according to the evaluation network parameters of the evaluation network uploaded by the multiple low-orbit satellites, and provides the global evaluation parameters of the trained global evaluation network to each low-orbit satellite, so that the low-orbit satellites update their respective evaluation networks according to the global evaluation parameters, and update their respective executor networks according to the updated evaluation networks.
[0055] Specifically, considering the potential application scenarios of federated learning, the present invention introduces federated learning technology into a three-dimensional multi-layer satellite network. The topological structure of the satellite network will change regularly and flexibly over time. At the same time, the resources on the satellite will also change with the communication transmission. The distributed satellite architecture will face the situation where the distributed nodes lack global information and fall into the local optimum. In this case, the information uploaded by the low-orbit satellite is integrated through federated learning to improve the global performance of the system. The medium-orbit satellite distributes the parameters of the global evaluation model to all low-orbit satellites through the intersatellite link between the medium-orbit satellite and the low-orbit satellite. After receiving the global model parameters from the medium-orbit satellite, the low-orbit satellite trains the local evaluation model and updates the executor model according to the updated evaluation model. After the training is completed, the parameters of the local evaluation model are uploaded after the decentralized process, otherwise the local evaluation model parameters are directly converged or terminated through the global model.
[0056] In one embodiment, the federated learning algorithm uses a weighted average algorithm to merge the evaluation parameters of multiple low-orbit satellites into a global evaluation parameter.
[0057] After the executor network is trained, the multi-layer satellite network routing method of deep reinforcement learning and federated learning of the present invention also includes completing routing decisions based on the trained executor network according to real-time satellite network environment information and real-time satellite topology.
[0058] The protection scope of the multi-layer satellite network routing method of deep reinforcement learning and federated learning described in the embodiment of the present invention is not limited to the execution order of the steps listed in this embodiment. All solutions implemented by adding, reducing, and replacing steps in the prior art based on the principles of the present invention are included in the protection scope of the present invention.
[0059] An embodiment of the present invention also provides a multi-layer satellite network routing system of deep reinforcement learning and federated learning. The multi-layer satellite network routing system of deep reinforcement learning and federated learning can implement the multi-layer satellite network routing method of deep reinforcement learning and federated learning described in the present invention. However, the implementation device of the multi-layer satellite network routing system of deep reinforcement learning and federated learning described in the present invention includes but is not limited to the structure of the multi-layer satellite network routing system of deep reinforcement learning and federated learning listed in this embodiment. All structural deformations and replacements of the prior art made according to the principles of the present invention are included in the protection scope of the present invention.
[0060] In one embodiment, the multi-layer satellite network routing system of deep reinforcement learning and federated learning of the present invention includes multiple low-orbit satellites and one medium-orbit satellite.
[0061] The plurality of low-orbit satellites are used to train respective actor networks using a distributed deep reinforcement learning algorithm to implement routing decisions, and train respective evaluation networks to implement local evaluation of the actor networks;
[0062] The medium-orbit satellite is used to generate a global evaluation network based on a federated learning algorithm according to the evaluation network parameters of the evaluation network uploaded by the multiple low-orbit satellites, and provide the global evaluation parameters of the trained global evaluation network to each low-orbit satellite, so that the low-orbit satellites update their respective evaluation networks according to the global evaluation parameters, and update their respective executor networks according to the updated evaluation networks.
[0063] The multi-layer satellite network routing system of deep reinforcement learning and federated learning of the present invention is further explained below through specific embodiments.
[0064] like Figure 3 As shown, the multi-layer satellite network routing system of deep reinforcement learning and federated learning of the present invention is composed of a layer of low-orbit satellite network and a layer of medium-orbit satellite network. In the low-orbit satellite layer, each satellite is equipped with computing resources and small servers, and will establish communication laser links with satellites around it. The deep reinforcement learning algorithm is distributed and deployed on each low-orbit satellite. Based on the partially observable Markov model, each low-orbit satellite obtains the distance to the adjacent satellite and the congestion of the adjacent satellite through the device. The low-orbit satellite transmits the observed information to the local executor network to obtain the next action, and guides the policy gradient to learn through the time difference residual, and updates the parameters of the evaluation network and the parameters of the executor network locally.
[0065] At the same time, the federated learning algorithm is deployed in the medium-orbit satellite layer. Each low-orbit satellite in the low-orbit satellite layer uploads its own evaluation network parameters, and the parameters are integrated through the federated learning algorithm to form a global evaluation network. Then the global evaluation network parameters are sent down to update the local evaluation network, thereby ensuring that the local executor network has global guidance and will not easily enter the local optimum. At the same time, by introducing the federated learning algorithm, the model training process is accelerated, energy consumption is saved, and the overall performance of the system is improved.
[0066] Therefore, the multi-layer satellite network routing system of deep reinforcement learning and federated learning of the present invention introduces deep reinforcement learning and federated learning into the multi-layer satellite network system to improve the decision-making performance of each independent satellite intelligence, and realizes the optimal path search under the satellite topology state at each moment, thereby minimizing the average transmission delay while meeting the packet loss rate and transmission rate limits.
[0067] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.
Claims
1. A multi-layer satellite network routing method based on deep reinforcement learning and federated learning, characterized in that: The method comprises the following steps: Multiple low-orbit satellites use a distributed deep reinforcement learning algorithm to train their respective actor networks to implement routing decisions, and train their respective evaluation networks to implement local evaluation of the actor networks; A medium-orbit satellite generates a global evaluation network based on a federated learning algorithm according to the evaluation network parameters of the evaluation network uploaded by the multiple low-orbit satellites, and provides the global evaluation parameters of the trained global evaluation network to each low-orbit satellite, so that the low-orbit satellites update their respective evaluation networks according to the global evaluation parameters, and update their respective executor networks according to the updated evaluation networks.
2. The multi-layer satellite network routing method of deep reinforcement learning and federated learning according to claim 1, characterized in that: Multiple low-orbit satellites use a distributed deep reinforcement learning algorithm to train their respective actor networks to implement routing decisions, and train their respective evaluation networks to implement local evaluation of the actor networks, including the following steps: Divide the satellite topology of the medium and low orbit satellite network into multiple time slices; In the static virtual topology corresponding to each time slice, each low-orbit satellite establishes an inter-satellite link with an adjacent satellite, wherein the adjacent satellites include two satellites in front and behind the same orbit and two satellites in adjacent left and right orbits; Each low-orbit satellite obtains the status information of the adjacent satellite based on the inter-satellite link; The state information is input into the deep reinforcement learning algorithm as satellite network environment information, local update parameters of the executor network are calculated, and local update parameters of the evaluation network are calculated.
3. The multi-layer satellite network routing method of deep reinforcement learning and federated learning according to claim 2, characterized in that: Each low-orbit satellite obtains the state information based on a partially observable Markov decision process; the state information includes the distance to adjacent satellites, congestion information, and energy consumption information.
4. The multi-layer satellite network routing method of deep reinforcement learning and federated learning according to claim 1, characterized in that: The federated learning algorithm adopts a weighted average algorithm to merge the evaluation parameters of multiple low-orbit satellites into a global evaluation parameter.
5. The multi-layer satellite network routing method of deep reinforcement learning and federated learning according to claim 1, characterized in that: It also includes completing routing decisions based on the trained executor network according to real-time satellite network environment information and real-time satellite topology.
6. A multi-layer satellite network routing system based on deep reinforcement learning and federated learning, characterized in that: The system includes a plurality of low-orbit satellites and a medium-orbit satellite; The plurality of low-orbit satellites are used to train respective actor networks using a distributed deep reinforcement learning algorithm to implement routing decisions, and train respective evaluation networks to implement local evaluation of the actor networks; The medium-orbit satellite is used to generate a global evaluation network based on a federated learning algorithm according to the evaluation network parameters of the evaluation network uploaded by the multiple low-orbit satellites, and provide the global evaluation parameters of the trained global evaluation network to each low-orbit satellite, so that the low-orbit satellites update their respective evaluation networks according to the global evaluation parameters, and update their respective executor networks according to the updated evaluation networks.
7. The multi-layer satellite network routing system of deep reinforcement learning and federated learning according to claim 6, characterized in that: Multiple low-orbit satellites use a distributed deep reinforcement learning algorithm to train their respective actor networks to implement routing decisions, and train their respective evaluation networks to implement local evaluation of the actor networks, including the following steps: Divide the satellite topology of the medium and low orbit satellite network into multiple time slices; In the static virtual topology corresponding to each time slice, each low-orbit satellite establishes an inter-satellite link with an adjacent satellite, wherein the adjacent satellites include two satellites in front and behind the same orbit and two satellites in adjacent left and right orbits; Each low-orbit satellite obtains the status information of the adjacent satellite based on the inter-satellite link; The state information is input into the deep reinforcement learning algorithm as satellite network environment information, local update parameters of the executor network are calculated, and local update parameters of the evaluation network are calculated.
8. The multi-layer satellite network routing system of deep reinforcement learning and federated learning according to claim 7, characterized in that: Each low-orbit satellite obtains the state information based on a partially observable Markov decision process; the state information includes the distance to adjacent satellites, congestion information, and energy consumption information.
9. The multi-layer satellite network routing system of deep reinforcement learning and federated learning according to claim 6, characterized in that: The federated learning algorithm adopts a weighted average algorithm to merge the evaluation parameters of multiple low-orbit satellites into a global evaluation parameter.
10. The multi-layer satellite network routing system of deep reinforcement learning and federated learning according to claim 6, characterized in that: The low-orbit satellite is also used to complete routing decisions based on the trained executor network according to real-time satellite network environment information and real-time satellite topology.