A Method for Optimizing On-Orbit Routing Decision and Wavelength Allocation in Sub-Domains
By adopting distributed routing decision control architecture and multi-agent deep reinforcement learning optimization methods in satellite optical networks, the high dynamics and resource waste of routing and wavelength allocation in satellite networks are solved, and network performance with low blocking rate and high throughput is achieved.
Patent Information
- Application Number
- CN202211728694.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-12-30
AI Technical Summary
In existing satellite optical networks, routing algorithms are difficult to effectively handle the high dynamics and node density of satellite networks, resulting in failure of local traffic accumulation and communication request resource reservation; at the same time, unreasonable wavelength allocation will lead to wavelength fragmentation and waste of resources.
Using a distributed routing decision-making control architecture based on LEO satellite network, satellite constellations are divided into multiple domains, and software-defined network controllers are set up in each domain to make routing and wavelength allocation decisions. Using a routing decision optimization method of independent stratified multiagent deep reinforcement learning, routing paths and wavelength allocation are optimized through graph neural networks and pheromone mechanisms.
It realizes distributed routing management and control of satellite networks, reduces the blocking rate, improves the high throughput and resource utilization of the network, and adapts to the high dynamics of satellite networks.
Smart Images

Figure CN116232425B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of satellite communication technology, and particularly to a method for optimizing domain-based on-orbit routing decision and wavelength allocation Background Art
[0002] With the rapid increase in the demand for instant communication services and video-on-demand services, and the rapid development of Internet of Things technology. In the era of Internet of Everything, there are not only communication needs between people, but also communication between things has become more frequent and important. In addition, the communication needs in remote mountainous areas and the ocean have gradually emerged. Traditional terrestrial networks are limited by capacity and coverage area and can no longer fully meet the requirements of large data transmission and high-reliability access services. Therefore, satellite networks have become increasingly important in the future network development due to their unique high coverage, long distance, and multi-access capabilities. Among them, with the production and use of small and lightweight lasers on satellites, satellite optical networks have become an indispensable part of satellite networks. However, ordinary wavelength-division multiplexing optical networks have the problem of large wavelength granularity, resulting in a waste of a large amount of communication resources. To solve this problem, a series of routing and wavelength allocation algorithms have been proposed, which have greatly improved the utilization rate of wavelength resources in satellite optical networks
[0003] Routing and wavelength allocation algorithms are the key technologies that determine the performance of satellite optical networks. Routing refers to the path sequence from the source node to the destination node in the network topology. Compared with terrestrial communication networks, satellite network topologies are highly dynamic and the node distribution is more dense and regular than that of terrestrial network topologies. Therefore, the use of conventional routing algorithms will lead to local traffic accumulation in satellite topologies, resulting in the failure of communication request resource reservation. Moreover, since satellite constellation links will be disconnected over time, existing terrestrial communication network routing algorithms are not applicable to satellite communication, and it is necessary to study satellite communication network routing algorithms based on terrestrial communication network routing algorithms. Therefore, with the development of satellite networks, satellite routing algorithms have gradually become a research direction for scholars at home and abroad. In terms of wavelength allocation, unreasonable wavelength allocation methods will lead to a large number of wavelength fragments in satellite optical networks. The wavelength capacity of these wavelength fragments cannot meet the needs of a single service, but the excessive accumulation of wavelength fragments will also lead to excessive waste of wavelength resources. Therefore, wavelength allocation is also one of the important issues in optical networks
[0004] At present, based on the rapid development of the single-layer satellite constellation network structure of LEO in recent decades, satellite optical networks can undertake the access and transmission functions of the global Internet of Things. An optical satellite network composed of dozens of satellites can provide almost global coverage. Low-Earth orbit satellites provide sufficient ground coverage by forming a constellation. In addition, low-Earth orbit satellites can access remote ground networks that cannot be served by traditional ground networks through microwave beams. In a low-Earth orbit constellation, satellites are connected by laser beams to form permanent or non-permanent laser links. In addition, with the development and application of software-defined network technology, data transmission and computing are decoupled, which can relieve the pressure of insufficient satellite computing power. The advantages of optical satellite networks have promoted the development of civil, military, and commercial communications.
[0005] The optimization algorithm needs to have the characteristics of high robustness and easy modification to be suitable for application to the highly dynamic network topology of satellites. For the routing and wavelength assignment problems in satellite optical networks, current routing strategies have problems such as satellite inter-link loss and high end-to-end delay. Nowadays, there are more and more studies on solving network routing problems. Most of these studies aim to reduce the blocking rate, reduce the delay, and improve the convergence speed to improve network performance. Summary of the Invention
[0006] The purpose of the embodiments of the present invention is to provide a domain-based on-board routing decision and wavelength assignment optimization method to achieve distributed routing control of satellite network controllers, low blocking rates for routing services, and at the same time ensure high throughput of satellite networks. The specific technical solutions are as follows:
[0007] On the one hand, the present invention provides a distributed routing decision control architecture. The distributed routing decision control architecture based on the LEO satellite network aims to solve the problem that in the LEO satellite network, due to the large network scale nodes, strong network dynamics, great difficulty in real-time global state statistics of the network, and excessive overhead of traditional flooding methods, it is impossible to optimize large-scale LEO satellite networks using traditional centralized routing and wavelength assignment optimization methods in reality. The distributed routing decision control architecture based on the LEO satellite network is mainly implemented in the satellite constellation network of LEO. The arrangement of LEO satellites forms the form of a Manhattan network. Except for edge nodes, each satellite node has four links connected to other satellites.
[0008] The reason for adopting the LEO constellation in this patent is that LEO satellites have the advantages of small propagation delay, low bit error rate, and the ability to cover polar regions, which is beneficial to avoiding the disadvantages of GEO satellites being unable to cover polar regions and weak on-board processing capabilities of LEO satellites. Moreover, this model also adopts redundant design because satellite nodes need to have redundancy to prevent node failures during communication from affecting the performance of the entire network.
[0009] The distributed routing decision control architecture based on the LEO satellite network, based on the LEO satellite constellation network, proposes a distributed routing decision control architecture based on the LEO satellite network. The entire low-earth orbit satellite constellation is divided into several domains according to the Manhattan street network, and a software-defined network controller is arranged in each domain as the routing domain wavelength allocation decision controller for that domain. When a service arrives, the domain to which it belongs will make a purely distributed intra-domain routing or cross-domain routing decision based on the router addresses of the source node and destination node of the service. In the controller within each domain, several pre-schemes of the routing path and wavelength between each pair of nodes within that domain are stored.
[0010] On the one hand, the implementation of the present invention provides a routing decision optimization based on independent hierarchical multi-agent deep reinforcement learning. The routing decision optimization based on independent hierarchical multi-agent deep reinforcement learning mainly includes that there are two interdependent routing decision modules in each domain of the network, namely the intra-domain routing controller and the intra-domain routing controller; the intra-domain routing controller is only responsible for the optimization of intra-domain routing services. For cross-domain routing services, an inter-domain routing controller is required; the inter-domain routing controller is responsible for the optimization of cross-domain routing services. It does not directly issue routing decisions itself, but designates a cross-domain routing link (i.e., a destination node within an edge domain) for the intra-domain routing controller, transmits the routing service to the adjacent domain, and allocates wavelengths using the first-match method.
[0011] The routing decision based on independent hierarchical multi-agent deep reinforcement learning is implemented within each domain of the low-earth orbit satellite constellation. The working principle within each of its domains is as Figure 1 shown. The low-earth orbit satellite routing decision task is abstracted into a Markov decision process, that is, a Markov chain of state -> action -> environmental change -> reward feedback; first, the intra-domain routing controller observes the network state within the domain to which it belongs; its network state S(B, D, λ, P) includes the link distances D of all links within the domain to which it belongs, the remaining wavelength capacity B, the Doppler frequency shift λ, and the routing occupancy mark P; the inter-domain routing controller observes the inter-domain link network state of the domain to which it belongs and the fuzzy distribution of the load on the intra-domain links S c (B d , B, D, λ, P, R), B dIt represents the comprehensive resource distribution within each domain after dividing the in-domain satellite network into several large domains; B represents the wavelength resource distribution of the cross-domain links; D represents the link distance distribution of the cross-domain links; λ represents the Doppler frequency shift distribution of the cross-domain links; P represents the pheromone concentration distribution of the cross-domain links, representing the quality of the cross-domain links between adjacent domains, and R represents the routing table marking bits to be marked for adjacent domains; when a routing service arrives, the controller will first determine whether it is a cross-domain routing service. When it is determined to be an in-domain routing service, only the in-domain routing controller is responsible for routing decision output; first, the in-domain routing controller observes the in-domain network status; marks several alternative routing schemes at the routing occupancy marking bits of the network observation status, and inputs each marked alternative routing scheme into the graph neural network. According to Equation (1), the output value of the graph neural network is the score for different routing paths.
[0012]
[0013] The in-domain routing controller selects the routing path with the highest current score as the routing strategy for output according to the greedy formula; the low-earth orbit satellite network performs routing forwarding according to the strategy, allocates wavelengths using the first-match method, and feeds back the reward value R to the in-domain routing controller according to the actual forwarding situation. As shown in Equation (2), it is the ratio of the minimum wavelength resource of the selected path to the wavelength resource and the maximum wavelength resource in the domain.
[0014]
[0015] When it is determined to be a cross-domain routing service, the in-domain routing controller and the cross-domain routing controller respectively obtain and observe the network status of the in-domain links and the cross-domain links. First, the cross-domain routing controller reads the cross-domain routing table according to the destination node of the routing service, marks the cross-domain links that match the cross-domain routing table at the R position. The cross-domain controller calculates the action state value of each cross-domain link according to Equation (1) and outputs the link with the highest action value as the selection of the cross-domain link. After the cross-domain controller outputs the cross-domain routing decision, the environment does not immediately feedback the reward value and the subsequent network status to the cross-domain routing controller. Instead, it first hands the cross-domain routing decision to the in-domain routing controller. According to the in-domain edge satellite network nodes output by the cross-domain routing controller, as the sub-destination nodes of the in-domain routing, an in-domain routing is generated; after both the cross-domain routing controller and the in-domain routing controller output their respective routing decisions, the joint routing decision is handed to the satellite network; the satellite network performs cross-domain forwarding of the routing service according to the routing decision, allocates wavelengths using the first-match method, and according to the routing forwarding result, allocates wavelengths using the first-match method; its cross-domain reward is R c , as shown in Equation (3):
[0016]
[0017] Its cross-domain reward is not only related to the cross-domain routing decision itself, but also to the quality of the intra-domain routing decision after the cross-domain routing decision;
[0018] In the intra-domain routing decision module, the neural network adopted is the message passing neural network, which is a branch of the graph neural network. Since in the routing and wavelength assignment problem, the problem is described in the form of a graph, using the graph neural network is more conducive to routing and wavelength assignment optimization. The patent uses the link as the neural network entity and performs the message passing process between all links. Figure 3 Shows the information passing process. The message passing neural network iterates over all links in the LEO network topology. For each link, its features are combined with the features of adjacent links (resource distribution, delay distribution), and the messages calculated for each link are aggregated with the information of its adjacent links by an element-wise addition method (according to Equations 4 and 5). Then, a recurrent neural network (RNN) is used to update the hidden state of the link with the new aggregated information. The algorithm receives the input link features and outputs a q value.
[0019]
[0020]
[0021] On the one hand, an embodiment of the present invention provides a cross-domain link pheromone mechanism based on a satellite network. Considering the scalability of satellite network control (the control architecture can adapt to a sufficiently large satellite constellation), a flat SDN deployment architecture is adopted. Each SDN controller can only observe the link status of its own domain. The SDN controllers only communicate asynchronously with adjacent controllers to jointly maintain the distribution of cross-domain link pheromones to which they both belong. In this solution, the pheromone of a cross-domain link represents the function calculation of several network factors that affect the routing success rate (link load, whether the link is connected, link Doppler frequency shift). Pheromone mechanism: In the network, pheromones are only distributed on cross-domain links (the purpose is to guide adjacent SDN controllers to select higher-quality cross-domain links as routing links), and the controller will update the pheromone distribution of the cross-domain links it belongs to at certain time intervals. Pheromone update rules: 1. In each update interval, if a link adjacent to a cross-domain link has a link disconnection or a Doppler frequency shift exceeding the threshold, the cross-domain link will increase a certain amount of pheromone. 2. In each update interval, if a route passes through a domain, all cross-domain links within the domain will increase a certain amount of pheromone. 3. In each update interval, if a route passes through a domain, the cross-domain link through which it passes will additionally increase a certain amount of pheromone. 4. After each time interval, all cross-domain links evaporate pheromones proportionally to avoid infinite accumulation of pheromones. Obviously, the fewer the pheromones of a cross-domain link, the higher the possible link quality. 5. After a routing failure, the pheromones of the selected cross-domain link will increase. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is a schematic diagram of routing for distributed routing decision-making
[0023] Figure 2 is a schematic diagram of routing decision-making for independent hierarchical multi-agent deep reinforcement learning
[0024] Figure 3 is a schematic diagram of cross-domain routing and pheromone update DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0026] Figure 1 The schematic diagram of routing for distributed routing decision-making includes:
[0027] The satellite node location management strategy based on the LEO satellite network proposes a distributed routing decision control architecture based on the LEO satellite network on the basis of the LEO satellite constellation network. The entire low-earth orbit satellite constellation is divided into several domains according to the Manhattan street network, and a software-defined network controller is arranged in each domain as the routing domain wavelength allocation decision controller of the domain. When a service arrives, the domain to which it belongs will make a pure distributed intra-domain routing or inter-domain routing decision according to the router addresses of the source node and the destination node of the service. The controller in each domain stores several pre-schemes of the routing paths and wavelengths between each pair of nodes in the domain. The satellite network performs routing forwarding according to the routing policy issued by the routing controller and allocates wavelengths using the first-match method. The first-match method is used to allocate wavelengths.
[0028] Figure 2 It is a schematic diagram of routing decision-making for independent hierarchical multi-agent deep reinforcement learning, including:
[0029] The routing decision task of low-earth orbit satellites is abstracted into a Markov decision process, that is, a Markov chain of state -> action -> environmental change -> reward feedback; first, the intra-domain routing controller observes the network state within the domain to which it belongs; its network state S(B, D, λ, P) includes the link distances D of all links within the domain, the remaining wavelength capacity B, the Doppler frequency shift λ, and the routing occupancy marker P; the inter-domain routing controller observes the inter-domain link network state of the domain to which it belongs and the fuzzy distribution of the load of the intra-domain links S c (B d , B, D, λ, P, R), B d represents the comprehensive resource distribution within each domain after dividing the intra-domain satellite network into several large domains; B represents the wavelength resource distribution of the inter-domain links; D represents the link distance distribution of the inter-domain links; λ represents the Doppler frequency shift distribution of the inter-domain links; P represents the pheromone concentration distribution of the inter-domain links, representing the quality of the inter-domain links of adjacent domains, and R represents the routing table marker bit to be marked for adjacent domains; when a routing service arrives, the controller will first determine whether it is an inter-domain routing service. When it is determined to be an intra-domain routing service, only the intra-domain routing controller is responsible for routing decision output; first, the intra-domain routing controller observes the intra-domain network state; marks several alternative routing schemes at the routing occupancy marker bit of the network observation state, and inputs each marked alternative routing scheme into the graph neural network. According to Equation (1), the output value of the graph neural network is the score of different routing paths;
[0030]
[0031] The in-domain routing controller selects the routing path with the highest current score as the routing policy for output according to the greedy formula; the low-earth orbit satellite network performs routing forwarding according to the policy, allocates wavelengths using the first-match method, allocates wavelengths using the first-match method, and feeds back the reward value R to the in-domain routing controller according to the actual forwarding situation. As shown in Equation (2), it is the ratio of the minimum remaining wavelength resource of the selected path to the maximum wavelength resource in the domain;
[0032]
[0033] When it is judged as a cross-domain routing service, the in-domain routing controller and the cross-domain routing controller respectively obtain and observe the network status of the in-domain link and the cross-domain link. First, the cross-domain routing controller reads the cross-domain routing table according to the destination node of the routing service, and marks the links that match the cross-domain routing table in the R position for the cross-domain links. The cross-domain controller calculates the action state value of each cross-domain link according to Equation (1), and outputs the link with the highest action value as the selection of the cross-domain link. After the cross-domain controller outputs the cross-domain routing decision, the environment will not immediately feedback the reward value and the subsequent network status to the cross-domain routing controller. Instead, it first hands the cross-domain routing decision to the in-domain routing controller. According to the in-domain edge satellite network nodes output by the cross-domain routing controller, as the sub-destination nodes of the in-domain routing, an in-domain routing is generated; after both the cross-domain routing controller and the in-domain routing controller output their respective routing decisions, the combined routing decision is handed to the satellite network; the satellite network performs cross-domain forwarding of the routing service according to the routing decision, and allocates wavelengths using the first-match method according to the routing forwarding, and the result of allocating wavelengths using the first-match method; its cross-domain reward is R c , as shown in Equation (3):
[0034]
[0035] Figure 3 It is a schematic diagram of cross-domain routing and pheromone update, including:
[0036] When a service arrives, the affiliated domain will make a pure distributed intra-domain routing or cross-domain routing decision based on the router addresses of the source node and the destination node of the service. When it is determined to be a cross-domain routing service, the intra-domain routing controller and the cross-domain routing controller respectively obtain and observe the network status of the intra-domain link and the cross-domain link. First, the cross-domain routing controller reads the cross-domain routing table according to the destination node of the routing service, generates a cross-domain routing decision, and hands the cross-domain routing decision to the intra-domain routing controller. According to the intra-domain edge satellite network nodes output by the cross-domain routing controller, as the sub-destination nodes of the intra-domain routing, an intra-domain routing is generated; after both the cross-domain routing controller and the intra-domain routing controller output their respective routing decisions, the joint routing decision is handed to the satellite network; the satellite network forwards the routing service across domains according to the routing decision, and then feeds back the routing reward to the controller based on the network status change after the forwarding routing and updates the cross-domain link pheromone.
Claims
1. A method for optimizing on - satellite routing decision and wavelength allocation in different domains, characterized in that, The method includes the following steps: Step 1: Divide the entire low-earth orbit satellite constellation into several domains according to the Manhattan street network, and deploy a software-defined network controller in each domain as the routing domain wavelength allocation decision controller for that domain; when a service arrives, the domain to which it belongs will make a purely distributed intra-domain routing or cross-domain routing decision based on the router addresses of the source node and destination node of the service. In the controller within each domain, several pre-schemes of routing paths and wavelengths between each pair of nodes within the domain are stored. Step 2: The software-defined network controller in each domain adopts routing decision optimization based on independent hierarchical multi-agent deep reinforcement learning; it is characterized in that a hierarchical routing decision architecture is proposed; the software-defined network controller in each domain is split into two different modules, namely the intra-domain routing controller and the inter-domain routing controller; the intra-domain routing controller is only responsible for optimizing intra-domain routing services, and for cross-domain routing services, the inter-domain routing controller is required; the inter-domain routing controller is responsible for optimizing cross-domain routing services. It does not directly issue routing decisions itself, but designates a cross-domain routing link for the intra-domain routing controller and transmits the routing service through this cross-domain routing link to the adjacent domain. The routing decision based on independent hierarchical multi-agent deep reinforcement learning is implemented in each domain of the low-earth orbit satellite constellation. The working principle in each domain includes abstracting the low-earth orbit satellite routing decision task into a Markov decision process, that is, a Markov chain of state -> action -> environmental change -> reward feedback. First, the in-domain routing controller observes the network state within its domain; its network state includes the link distance D, remaining wavelength capacity B, Doppler frequency shift λ, and routing occupancy marked as P of all links within the domain. The inter-domain routing controller observes the inter-domain link network state of its domain and the load fuzzy distribution of intra-domain links , represents the comprehensive resource distribution within each domain after dividing the intra-domain satellite network into several large domains; B represents the wavelength resource distribution of the cross-domain link; D represents the link distance distribution of the cross-domain link; λ represents the Doppler frequency shift distribution of the cross-domain link; P represents the pheromone concentration distribution of the cross-domain link, represents the cross-domain link quality of the adjacent domain, and R represents the identifier on the cross-domain link according to the cross-domain routing table of the adjacent domain. When a routing service arrives, the controller will first determine whether it is a cross-domain routing service. When it is determined to be an in-domain routing service, only the in-domain routing controller is responsible for routing decision output. First, the in-domain routing controller observes the in-domain network state; marks several alternative routing schemes at the routing occupancy marker bits of the network observation state, and inputs each marked alternative routing scheme into the value graph neural network. According to Equation (1), the output value of the graph neural network is the score for different routing paths; (1) The intra-domain routing controller selects the routing path with the highest current score as the routing strategy for output according to the greedy formula; the low-earth orbit satellite network performs routing forwarding according to the strategy, allocates wavelengths using the first-match method, and feeds back the reward value R' to the intra-domain routing controller according to the actual forwarding situation. As shown in Equation (2), it is the ratio of the minimum remaining wavelength resource of the selected path to the maximum wavelength resource in the domain. (2) When it is determined to be a cross-domain routing service, the in-domain routing controller and the cross-domain routing controller respectively obtain and observe the network status of the in-domain link and the cross-domain link. First, the cross-domain routing controller reads the cross-domain routing table according to the destination node of the routing service, and marks the links that match the cross-domain routing table in the R position of the cross-domain link; the cross-domain controller calculates the action state value of each cross-domain link according to formula (1), and outputs the link with the highest action value as the selection of the cross-domain link. After the cross-domain controller outputs the cross-domain routing decision, the environment will not immediately feedback the reward value and the subsequent network status to the cross-domain routing controller. Instead, it first hands the cross-domain routing decision to the in-domain routing controller, and uses the in-domain edge satellite network node output by the cross-domain routing controller as the sub-destination node of the in-domain routing to generate an in-domain routing; after both the cross-domain routing controller and the in-domain routing controller output their respective routing decisions, the combined routing decision is handed to the satellite Internet; the satellite Internet performs cross-domain forwarding of the routing service according to the routing decision, and allocates the wavelength result using the first-match method according to the routing forwarding; its cross-domain reward is , as shown in formula (3): (3) Its cross-domain reward is not only related to the cross-domain routing decision itself, but also related to the quality of the intra-domain routing decision after the cross-domain routing decision. Step 3: After each cross-domain routing decision and forwarding, it is necessary to update the cross-domain link pheromone according to the routing forwarding and the wavelength allocation situation using the first-match method; considering that satellite network control can adapt to large-scale low-earth orbit satellite constellations, a flat SDN deployment architecture is adopted. Each SDN controller can only observe the link status of its own domain, and the SDN controllers only jointly maintain the distribution of the cross-domain link pheromone belonging to them through asynchronous communication with adjacent controllers; the pheromone of the cross-domain link is a function calculation of several network factors that affect the routing success rate, including link load, whether the link is connected, and link Doppler frequency shift; the pheromone mechanism means that pheromone is only distributed on the cross-domain links in the network, and the pheromone is used to guide adjacent SDN controllers to select a cross-domain link with higher quality as the routing link, and the controller will update the distribution of the cross-domain link pheromone it belongs to at a certain time interval; the specific update rules of the pheromone are as follows:
1. Within each update interval, if there is a link disconnection in the links adjacent to a cross-domain link, then this cross-domain link will increase a certain amount of pheromone.
2. Within each update interval, if there is a Doppler frequency shift exceeding the threshold in the links adjacent to a cross-domain link, then this cross-domain link will increase a certain amount of pheromone.
3. Within each update interval, if a route passes through a domain, then all cross-domain links within the domain will increase a certain amount of pheromone.
4. During each update interval, if a route passes through a domain, a certain amount of pheromone will be additionally added to the cross-domain links it passes through; 5. After each time interval, all cross-domain links evaporate pheromone proportionally to avoid infinite accumulation of pheromone; the less pheromone a cross-domain link has, the higher its link quality; 6. After a routing failure, the pheromone of the selected cross-domain link will increase; Step 4: Repeat the second and third steps to train the cross-domain routing controller and the intra-domain routing controller until their reward curves converge.