Multi-domain routing selection method facing satellite-ground cooperative architecture, medium and equipment
By building a satellite-ground collaboration architecture in the satellite network, improving the ant colony algorithm and DQN model optimization routing strategy, the problems of slow routing convergence and unrefined weight parameters in the satellite network are solved, and satellite communication with high throughput and low latency are achieved.
Patent Information
- Application Number
- CN202510483226.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
AI Technical Summary
The existing satellite network routing algorithm converges slowly in dynamic environments and has large network overhead. The ant colony optimization algorithm does not design the weight parameters of bandwidth, delay and packet loss rate, resulting in unsatisfactory path planning.
The satellite-oriented collaboration architecture is adopted, combined with MEO satellites, LEO satellites and ground station agents to form a network to build a hierarchical network architecture. By improving the ant colony algorithm and DQN model to optimize the routing strategy, distinguish the weight parameters of elephant flow and mouse flow, use the DQN model to train the heuristic function of the ant colony algorithm, and combine the Dijkstra algorithm to calculate cross-domain routing.
It improves the throughput of the satellite network, reduces the end-to-end delay and packet loss rate, improves the communication service level of the satellite network, and avoids network congestion.
Smart Images

Figure CN120342934A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of satellite network routing algorithms, and more particularly to a multi-domain routing selection method, medium and device for a satellite-ground cooperation architecture. Background Art
[0002] In terms of satellite network routing optimization algorithms, the most commonly used method is the Open Shortest Path First (OSPF) routing protocol based on the shortest path algorithm. However, in a dynamic satellite network, the OSPF protocol needs to update the routing table according to the real-time network state, resulting in slow routing convergence and significant network overhead problems.
[0003] Routing solutions in satellite networks usually adopt fully centralized or distributed methods. However, the centralized routing control method has large overhead and response delay in collecting global information, resulting in the calculated routing table lacking timeliness and flexibility; the fully distributed method lacks a global perspective and is prone to falling into local optima.
[0004] The routing strategy of satellite Internet can generally be modeled as an optimization problem. Due to different QoS requirements of different services and different network states in different environments, optimizing only a certain parameter can no longer meet the requirements. In the prior art, when using the ant colony optimization algorithm to conduct multi-objective optimization and comprehensive modeling of multiple factors, the weight parameters of bandwidth, delay and packet loss rate are not finely designed. Based on setting the weight of bandwidth greater than the weights of delay and packet loss rate for elephant flows, different weight ratios may lead to different paths planned by the ant colony algorithm, and the same is true for mouse flows, and the effectiveness and convergence speed of the algorithm are not ideal. Summary of the Invention
[0005] The present invention provides a multi-domain routing selection method, medium and device for a satellite-ground cooperation architecture, which can, while meeting the QoS requirements of user services, improve throughput, reduce end-to-end delay and packet loss rate, improve algorithm convergence, and avoid satellite network congestion, thereby enhancing the communication service level of satellite networks, and can at least solve one of the above technical problems.
[0006] To solve the above technical problems, the present invention adopts the following technical solutions:
[0007] A multi-domain routing selection method for a satellite-ground cooperation architecture includes the following steps:
[0008] S1. Combine MEO satellites, LEO satellites and ground station agents for networking, apply SDN to the satellite network, construct a hierarchical network architecture, and achieve global control and cross-domain routing of the satellite network;
[0009] S2. Classify the traffic reaching the in-domain edge LEO satellite nodes by threshold, and establish a multi-objective optimization model based on the throughput, delay, and packet loss rate of the satellite network;
[0010] S3. Improve the ant colony algorithm for the multi-objective optimization model in S2;
[0011] S4. Use the bandwidth, delay, and packet loss rate weight parameters in the heuristic function of the improved ant colony algorithm to train the DQN model for routing planning, and solve to obtain the optimal multi-objective optimization routing strategy for in-domain LEO satellites based on separating large and small flows;
[0012] S5. Divide the cross-domain partition routing into intra-region routing and inter-region routing, and adopt different algorithms for routing strategy formulation within and between regions to form a complete cross-domain routing table.
[0013] Furthermore, S1 further includes:
[0014] S1.1. Divide the LEO satellites into regions so that the number of in-orbit satellites and the number of satellites in the orbital plane in each region are close to the set threshold of the number of in-orbit satellites in the region;
[0015] S1.2. Deploy multiple in-domain domain controllers DCs in the entire heterogeneous network composed of MEO satellites and LEO satellites. The in-domain domain controllers interact with the ground station agents for the current domain and numerous in-domain satellite nodes SNs to issue routing tables;
[0016] S1.3. As the control layer, the DC manages the in-domain SNs, transfers the computing center to the ground station agent for network deployment and resource scheduling. Each domain is supervised by a dedicated DC. Within the domain, the DC continuously monitors the link information, location, and network topology in the domain and transmits this information to the ground station agent of the current domain. The ground station agent performs in-domain routing calculation and forms a routing strategy. The SNs act as the forwarding layer, are connected to each other through ISLs, and process data packets according to the commands of the domain controllers DCs. The data packet includes at least voice data, video data, file data, and picture data.
[0017] Furthermore, S2 further includes:
[0018] S2.1. Let the total data stream bytes received at time t1 be Bt1, the total data stream bytes received at time t2 be Bt2, and the current link maximum bandwidth be BW max , and the ratio flow of the data stream bytes to the current link maximum bandwidth is calculated as:
[0019]
[0020] When the ratio flow is greater than 20%, the data flow is classified as an elephant flow; otherwise, the data flow is classified as a mouse flow.
[0021] S2.2. With the goal of reducing the delay and packet loss rate of traffic transmission and improving the network throughput, the total cost is defined as:
[0022]
[0023] R L = H L,A (4)
[0024]
[0025] Among them, R L and R N are the costs of the elephant flow and the mouse flow respectively, H t is the throughput within the time interval t, t is the statistical time interval, O is the amount of data received, H L,A is the average throughput of the large flow, D N is the delay of the small flow, L N is the packet loss rate of the small flow;
[0026] For the ant colony optimization algorithm combined with the DQN model, the optimization goal is expressed as:
[0027]
[0028] Furthermore, in S3, the improvement of the ant colony algorithm includes the following process:
[0029] S3.1. Optimize the heuristic function of the ant colony algorithm to:
[0030]
[0031] Among them:
[0032] η ij is the comprehensive score of the link between node i and the next node j, which is used as the heuristic function. The higher the score value, the greater the probability that the next ant will choose this link;
[0033] BW ij , Delay ij and Loss ij are the remaining bandwidth, delay and packet loss rate on the link between node i and node j respectively;
[0034] b, d and l are the weights corresponding to the remaining bandwidth, delay and packet loss rate respectively;
[0035] The weights b, d, and l of the remaining bandwidth, delay, and packet loss rate are set with different scoring weight values according to the high-bandwidth characteristics of the elephant flow and the low-latency requirements of the mouse flow. When the data flow is identified as an elephant flow, the weight b of the remaining bandwidth is greater than d and l. When the data flow is identified as a mouse flow, the weights d of the delay and l of the packet loss rate are both greater than the weight b of the remaining bandwidth;
[0036] S3.2. Incorporate the elite ant system, improve the pheromone update rule of the ant colony algorithm, and further optimize the ant colony algorithm. The corresponding pheromone update rule of the optimized system is as follows:
[0037]
[0038] Among them:
[0039] τ ij is the pheromone between node i and node j. As the number of iteration rounds increases, the pheromone will continuously evaporate and accumulate;
[0040] ρ is the evaporation coefficient of the pheromone, ρ ∈ [0, 1];
[0041] 1 - ρ is the existence coefficient of the pheromone;
[0042] Δτ ij is the pheromone increment of the path, and the calculation formula is:
[0043]
[0044] is the pheromone increment generated by the elite ants;
[0045] σ is the number of elite ants;
[0046] Q is the total amount of pheromone left by the ants when passing through the path, and is set as a constant;
[0047] BW uij is the used bandwidth of the link between node i and node j;
[0048] The preference of the ants to select the link is directly proportional to the remaining bandwidth, inversely proportional to the delay and packet loss rate, and changes with the size of the traffic. After the ants complete one cycle, the pheromone concentration is updated according to the above formulas (8) - (10).
[0049] Furthermore, in the above S3, to improve the global search ability and convergence speed of the improved ant colony algorithm, the evaporation coefficient ρ is taken in stages:
[0050]
[0051] Among them, I cur represents the current iteration number of the algorithm, Imax Represents the maximum number of iterations set by the algorithm.
[0052] Furthermore, in S4, the total amount of pheromone Q value is calculated using the CNN convolutional neural network, and the state space State, reward function Reward, and action space Actor in the DQN model are mapped as follows:
[0053] State space S:
[0054] The state space is represented as a three-dimensional array or matrix with a shape of n×3, where each row corresponds to the link state of a network, and there are a total of n network link states:
[0055] S = [ B i , D i , L i ], i ∈ [1, n] (12)
[0056] Among them, B, D, and L are the remaining bandwidth matrix, delay matrix, and packet loss rate matrix of the network topology, respectively;
[0057] Reward function R:
[0058] For large flows, high throughput needs to be obtained, and the reward function is shown in Equation (4);
[0059] For small flows, low latency and low packet loss rate need to be obtained, and the reward function is shown in Equation (5);
[0060] Action space A:
[0061] Executing action a means adjusting the weights of the ant colony algorithm according to the current network state. Each action corresponds to a specific weight adjustment, and the expression of the action space is:
[0062] A = [a1, a2, a3, …, a k ] (13)
[0063] Among them, a k = [b, d, l];
[0064] Two different action spaces are designed for elephant flows and mouse flows respectively:
[0065] Let the action space A S of the mouse flow be:
[0066]
[0067] Let the action space A L of the elephant flow be:
[0068]
[0069] Further, in step S5, the satellite inter-domain routing table is calculated by the ground station agent based on the link delay between the border routing node pairs of adjacent regions obtained through the ephemeris information using the Dijkstra algorithm, and is distributed to each border routing node satellite in the adjacent regions in a pre-stored manner. The satellite intra-domain routing table is calculated using an improved ant colony algorithm optimized by the DQN model.
[0070] Further, the ground station agent returns the formed routing policy to the DC, and the DC inserts the routing policy into the segment routing list, and starts forwarding data from the edge LEO satellite node, and reaches the destination of the data packet through intra-region routing and inter-region routing.
[0071] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the steps of the above multi-domain routing selection method for a satellite-ground cooperation architecture.
[0072] A computer device includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the above multi-domain routing selection method for a satellite-ground cooperation architecture.
[0073] The beneficial effects of the present invention are reflected in:
[0074] 1. By adopting the idea of "satellite zoning and satellite-ground cooperation", a network is formed by combining MEO satellites, LEO satellites and ground station agents. The MEO satellite acts as a controller, responsible for issuing the routing policy of the intra-domain LEO satellite. The LEO satellite is responsible for data forwarding, and the ground station agent is responsible for formulating the routing policy, which improves the flexibility of the system framework.
[0075] 2. The satellite intra-domain routing table is calculated using the new algorithm proposed in this method, that is, the traditional ant colony algorithm is redesigned. In the heuristic function of the designed ant colony algorithm, the weight parameters of bandwidth, delay and packet loss rate are set using the DQN algorithm to train convolutional neural networks (CNNs) for large flows and small flows in the satellite network respectively, realizing the refined design of the weight parameters, as well as more accurate and more real-time differential scheduling of elephant flows and mouse flows, accelerating the convergence speed of the multi-parameter joint optimization solution algorithm, increasing the throughput of the satellite network, reducing the delay and packet loss rate, and meeting the requirements of high throughput for large flows and low latency for small flows in the satellite network. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application.
[0077] Figure 1It is a schematic diagram of the overall process of the multi-domain routing method according to an embodiment of the present invention.
[0078] Figure 2 It is an architecture diagram of the multi-domain routing method according to an embodiment of the present invention.
[0079] Figure 3 It is the in-domain system framework of the multi-domain routing method according to an embodiment of the present invention.
[0080] Figure 4 It is the in-domain flowchart of the multi-domain routing method according to an embodiment of the present invention.
[0081] Figure 5 It is the in-domain experimental simulation topology structure diagram of the multi-domain routing method according to an embodiment of the present invention.
[0082] Figure 6 It is a comparison chart of the throughput of Dijkstra, ECMP in the domain and the algorithm proposed in the embodiment of the present invention.
[0083] Figure 7 It is a comparison chart of the end-to-end average delay of Dijkstra, ECMP in the domain and the algorithm proposed in the embodiment of the present invention.
[0084] Figure 8 It is a comparison chart of the end-to-end average packet loss rate of Dijkstra, ECMP in the domain and the algorithm proposed in the embodiment of the present invention.
[0085] Figure 9 It is a structural block diagram of a computer device according to an embodiment of the present invention. Detailed implementation manners
[0086] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0087] It should be noted that the meaning of "and / or" appearing throughout the text includes three parallel solutions. Taking "A and / or B" as an example, it includes solution A, or solution B, or a solution where A and B are satisfied simultaneously. In addition, "a plurality of" means two or more. In addition, the technical solutions between the embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions appears to be contradictory or unable to be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0088] The explanations of the key technical terms involved in the present invention are as follows:
[0089] SDN: By separating the control plane and data plane of the network, it realizes the programmability, flexibility, and controllability of the network, thus solving the problems existing in the traditional network architecture. This technology simplifies network management and configuration work, improves network scalability and innovation, and provides important support for the innovation, upgrade, and optimization of future networks.
[0090] DQN: With the rapid development of artificial intelligence technology based on deep neural networks, it has been widely applied in fields such as natural language processing, image recognition, and game strategy calculation. With the in-depth research of neural networks and the improvement of computer hardware foundation, artificial intelligence models have great potential. Compared with before, they can learn more and more complex strategies, and the training and execution efficiency have also been greatly improved. Deep reinforcement learning algorithms are divided into value-based and action-probability-based. Q-Learning is a value-based reinforcement learning algorithm that uses a Q-table to store the values of each action in each state and selects actions based on the Q-values. However, Q-Learning is only applicable to discrete problems, while most problems in the real world are high-dimensional or continuous problems. Therefore, DQN is proposed to solve the problem of high-dimensional input.
[0091] Ant Colony Algorithm: A path planning algorithm proposed by Italian scholar M. Dorigo in 1991, which utilizes the mechanism of pheromone in the process of ants foraging in nature, that is, the rules of pheromone accumulation and pheromone evaporation followed by ants during foraging and homing. It has strong robustness, high efficiency, and good adaptability, and is commonly used to solve the traveling salesman problem and the path planning problem of mobile robots. However, it also has problems such as being easily trapped in local optima and slow convergence speed.
[0092] See Figures 1-4 , the embodiment of the present invention provides a multi-domain routing selection method for a satellite-ground cooperation architecture, including the following steps:
[0093] S1. Based on the idea of "satellite partitioning and satellite-ground cooperation", combine MEO satellites, LEO satellites, and ground station agents to form a network, apply SDN to the satellite network, construct a hierarchical network architecture, and achieve global control and cross-domain routing of the satellite network;
[0094] S2. Classify the traffic arriving at the edge LEO satellite nodes within the domain by threshold, and establish a multi-objective optimization model based on the throughput, delay, and packet loss rate of the satellite network;
[0095] S3. Improve the ant colony algorithm for the multi-objective optimization model in S2;
[0096] S4. Use the bandwidth, delay, and packet loss rate weight parameters in the heuristic function of the improved ant colony algorithm to train the DQN model for routing planning, and solve to obtain the optimal multi-objective optimization routing strategy based on separating large and small flows within the LEO satellite domain;
[0097] S5. Divide the cross-domain partition routing into intra-region routing and inter-region routing, and use different algorithms for formulating routing strategies within the region and between regions to form a complete cross-domain routing table.
[0098] See Figures 1-4 , in this embodiment, the S1 further includes:
[0099] S1.1. Divide the LEO satellites into regions so that the number of satellites within the orbit and the number of satellites on the orbital plane in each region are close to the set threshold of the number of satellites within the orbit in the region;
[0100] S1.2. Deploy multiple domain controllers DCs within the satellite domain in the entire heterogeneous network composed of MEO satellites and LEO satellites. The domain controllers within the satellite domain interact with the ground station agents for the current domain and numerous satellite nodes SNs within the domain to issue routing tables;
[0101] S1.3. As the control layer, the DC manages the SNs within the domain, transfers the computing center to the ground station agent for network deployment and resource scheduling. Each domain is supervised by a dedicated DC. Within the domain, the DC continuously monitors the link information, location, and network topology within the domain, and transmits this information to the ground station agent of the current domain. The ground station agent performs intra-domain routing calculation and forms a routing strategy. The SNs serve as the forwarding layer, are connected to each other through ISLs, and process data packets according to the commands of the domain controllers DCs. The data packet includes at least voice data, video data, file data, and picture data.
[0102] See Figures 1-4 , in this embodiment, the S2 further includes:
[0103] S2.1. Let the total data stream bytes received at time t1 be Bt1, the total data stream bytes received at time t2 be Bt2, and the maximum bandwidth of the current link be BW max , and the ratio flow of the data stream bytes to the maximum bandwidth of the current link is calculated as:
[0104]
[0105] When the ratio flow is greater than 20%, classify the data stream as an elephant flow, otherwise classify the data stream as a mouse flow;
[0106] S2.2. With the goal of reducing the delay and packet loss rate of traffic transmission and improving the throughput of the network, define the total cost as:
[0107]
[0108] R L = H L,A (4)
[0109]
[0110] wherein, R L and R N are the costs of the elephant flow and the mouse flow respectively, H t is the throughput within the time interval t, t is the statistical time interval, O is the amount of data received, H L,A is the average throughput of the large flow, D N is the delay of the small flow, L N is the packet loss rate of the small flow;
[0111] For the ant colony optimization algorithm combined with the DQN model, the optimization objective is expressed as:
[0112]
[0113] See Figures 1-4 , in this embodiment, in S3, the improvement of the ant colony algorithm includes the following process:
[0114] S3.1. Optimize the heuristic function of the ant colony algorithm to:
[0115]
[0116] wherein:
[0117] η ij is the comprehensive link score between node i and the next node j. As a heuristic function, the higher the score value, the greater the probability that the next ant will select this link;
[0118] BW ij , Delay ij and Loss ij are the remaining bandwidth, delay and packet loss rate on the link between node i and node j respectively;
[0119] b, d and l are the weights corresponding to the remaining bandwidth, delay and packet loss rate respectively;
[0120] The weights b, d and l of the remaining bandwidth, delay and packet loss rate are set with different scoring weight values according to the high bandwidth requirement characteristics of the elephant flow and the low latency requirement characteristics of the mouse flow. When the data flow is identified as an elephant flow, the weight b of the remaining bandwidth is greater than d and l. When the data flow is identified as a mouse flow, the weights d of the delay and l of the packet loss rate are both greater than the weight b of the remaining bandwidth;
[0121] S3.2. Incorporate the elite ant system, improve the pheromone update rule of the ant colony algorithm, and further optimize the ant colony algorithm. The corresponding pheromone update rule of the optimized system is as follows:
[0122]
[0123] Where:
[0124] τ ij is the pheromone between node i and node j. As the number of iterations increases, the pheromone will continuously evaporate and accumulate;
[0125] ρ is the evaporation coefficient of the pheromone, and ρ ∈ [0, 1];
[0126] 1 - ρ is the existence coefficient of the pheromone;
[0127] Δτ ij is the pheromone increment of the path, and the calculation formula is:
[0128]
[0129] is the pheromone increment generated by the elite ants;
[0130] σ is the number of elite ants;
[0131] Q is the total amount of pheromone left by the ants passing through the path, and is set as a constant;
[0132] BW uij is the used bandwidth of the link between node i and node j;
[0133] The preference of the ants to select the link is proportional to the remaining bandwidth, inversely proportional to the delay and packet loss rate, and changes with the size of the traffic. After the ants complete one cycle, the pheromone concentration is updated according to the above formulas (8) - (10).
[0134] See Figures 1-4 , in this embodiment, in S3, to improve the global search ability and convergence speed of the improved ant colony algorithm, the evaporation coefficient ρ is taken in stages:
[0135]
[0136] Where I cur represents the current iteration number of the algorithm, and I max represents the maximum number of iterations set by the algorithm.
[0137] The specific process of the ant colony optimization algorithm is as follows:
[0138] A. Initialize the parameters, set the maximum number of iterations to N, the maximum number of ants to K, and initialize the pheromone concentration τ ij to the total pheromone amount Q, and set the link weight factor, dependence coefficient, link preference coefficient, and pheromone evaporation coefficient of the elephant flow and the mouse flow. Initialize the taboo list tabu to the source node, the path library Route to be empty, and the current iteration number n = 0;
[0139] B. Detect large and small flows, and assign values to b, d, and l according to the flow category respectively;
[0140] C. Update the current iteration number n = n + 1, and the current number of ants k = 0;
[0141] D. Update the current number of ants k = k + 1;
[0142] E. According to the probability transfer formula, select the switch of the next node. If the destination node is reached, store the path nodes in the taboo list tabu;
[0143] F. If k < K, return to step D. Otherwise, save the path with the highest comprehensive score of the links in the taboo list tabu to the path library Route, empty the taboo list tabu, and update the pheromone concentration according to the pheromone update formula;
[0144] G. If n < N, return to step C;
[0145] H. Select the path with the highest comprehensive score of the links from the path library Route as the optimal solution for output.
[0146] See Figures 1-4 , in this embodiment, in S4, use the CNN convolutional neural network to calculate the total pheromone amount Q value, and map the state space State, reward function Reward, and action space Actor in the DQN model as follows:
[0147] State space S:
[0148] Represent the state space as a three-dimensional array or matrix with a shape of n×3, where each row corresponds to the link state of a network, and there are a total of n network link states:
[0149] S = [ B i , D i , L i , i∈[1,n] (12)
[0150] where B, D, and L are the remaining bandwidth matrix, delay matrix, and packet loss rate matrix of the network topology respectively;
[0151] Reward function R:
[0152] For large flows, high throughput needs to be achieved, and the reward function is shown in Equation (4);
[0153] For small flows, low latency and low packet loss rate need to be achieved, and the reward function is shown in Equation (5);
[0154] Action space A:
[0155] Executing action a means adjusting the weights of the ant colony algorithm according to the current network state. Each action corresponds to a specific weight adjustment. The expression of the action space is:
[0156] A = [a1, a2, a3, …, a k (13)
[0157] where a k = [b, d, l];
[0158] Design two different action spaces for elephant flows and mouse flows respectively:
[0159] Let the action space of mouse flows A S be:
[0160]
[0161] Let the action space of elephant flows A L be:
[0162]
[0163] The training steps of the ant colony optimization algorithm model combined with DQN are as follows:
[0164] Input: G (experimental network topology), B (remaining bandwidth matrix), D (delay matrix), L (packet loss rate matrix)
[0165] Output: Optimal large flow path, optimal small flow path
[0166] 1) Threshold traffic classification
[0167] 2) IF a new flow flow arrives at Source (starting node)
[0168] Determine whether it is a large or small flow through formula (1);
[0169] 3) IF flow ≤ 20%
[0170] 4) flow is classified as a mouse flow, denoted as F S ;
[0171] 5) ELSE
[0172] 6) flow is classified as an elephant flow, denoted as F L ;
[0173] 7) When both F exist at the Source S and F L , the two Agents are trained simultaneously;
[0174] 8) Initialize the environment, and use the link state s = {B, D, L} of the network as the input. Given hyperparameters such as the learning rate α and the discount factor Υ.
[0175] 9) for episode = 1 to M do:
[0176] 10) for i = 1 to T do
[0177] 11) Select a according to the greedy policy t . According to different flows, the actions a of the two Agents t are selected from Equation (14) and Equation (15) respectively;
[0178] 12) Update the weight parameters of the heuristic function of the improved ant colony optimization algorithm according to the selected action a t , and calculate the QoS path from step A to step H respectively, and issue the flow table.
[0179] 13) Calculate the rewards R of the elephant flow and the mouse flow according to Equation (4) and Equation (5) respectively L and R N .
[0180] 14) Collect network link information through the SDN controller to obtain the next moment state. BufferL, BufferS ← Store the experience data.
[0181] 15) Update the Q values of the two Agents respectively according to Q(s t+1 , a t+1 ) = Q(s t , a t ) + α Q [R + γ Q (Q(s t+1 , a t+1 ) - Q(s t , a t ))];
[0182] 16) s = s t+1 . Update the action a = a according to different action spaces respectively t+1 .
[0183] 17) i = i + 1
[0184] 18) end for
[0185] 19) end for
[0186] See Figures 1-4 In this embodiment, in step S5, the satellite inter-domain routing table is calculated by the ground station agent based on the link delay between the border routing node pairs of adjacent regions obtained from the ephemeris information using the Dijkstra algorithm, and is distributed to each border routing node satellite in the adjacent regions by means of pre-storage. The intra-satellite domain routing table is calculated using an improved ant colony algorithm optimized by the DQN model.
[0187] The specific steps of cross-domain partition routing are as follows:
[0188] Input: Source node V s , destination node V d , region number R, intra-region routing table N, inter-region routing table I, inter-region link L b
[0189] Output: Complete path P
[0190] a. R s ← R[V s / / Obtain the region number of the source node
[0191] b. R d ← R[V d / / Obtain the region number of the destination node
[0192] c. if R s == R d then / / If the source node and the destination node are in the same region
[0193] d. P ← N[V s [V d / / Use the intra-region routing table to obtain the large and small flow paths
[0194] e. else
[0195] f. P ← [] / / Initialize the path
[0196] g. P region ← I[R s [R d / / Use the inter-region routing table to obtain the shortest inter-region path
[0197] h. R current ← R s / / Set the current region as the source node region
[0198] i. V current ← V s / / Set the current node as the source node
[0199] j. for Rnext in P region do / / Traverse the path between regions
[0200] k.V candicate ←GETBORDERNODES(R current ,R next ,L b ) / / Get the border nodes connecting the current region and the next-hop region
[0201] l.V currentd ←SELECTNEARESTNODE(V currents ,V candicate ,N) / / Select the candidate node with the highest comprehensive link score with the current node
[0202] m.P intra ←N[V currents [V currenta / / Get the path within the current region
[0203] n.P←P∪P intra / / Add the path to the total path
[0204] o.V nexts ←GETNEXTNODE(V currentd ,R current, ,R nexx ) / / The corresponding border node of the next region is used as the source node of the next region
[0205] p.border_link←L b [V currentd [V nexts / / Get the link between border nodes
[0206] q.P←P∪border_link / / Add the border link to the path
[0207] r.V currents ←V nexts / / Update the current source node
[0208] s.R current ←R next / / Update the current region
[0209] t.end for
[0210] u.P intra ←N[V currents [V d / / Get the path within the destination region
[0211] v.P←P∪P intra / / Add the path to the total path
[0212] w.end if
[0213] x.e.return P / / Return the complete path
[0214] See Figures 1-4 , in this embodiment, the ground station agent returns the formed routing policy to the DC, and the DC inserts the routing policy (the SID relative to the node) into the segment routing list, and starts forwarding data from the edge LEO satellite node, and reaches the packet destination through intra-region routing and inter-region routing.
[0215] To verify the application of the present invention, a specific experiment will be provided below to further illustrate the multi-domain routing selection method of this satellite-ground cooperation architecture:
[0216] Build a DQN-optimized ant colony intelligent routing algorithm system based on SDN as Figure 3 shown. This experiment uses the Gym 0.21.0 platform of OpenAI to build a simulated SDN environment. This algorithm is implemented based on the Keras 2.7+TensorFlow 2.7 framework. In the Mininet 2.22 experimental simulation platform, a Fat-tree topology structure as Figure 5 shown is established through python 2.7. This network consists of 1 controller, 15 programmable BMv2 software switches and 12 hosts. The controller and Mininet both run on a physical machine installed with the Ubuntu 20.04 operating system (GPU: NVIDIA GeForce RTX2070).
[0217] The maximum number of iterations of the improved ant colony optimization algorithm is set to 60, the maximum number of ants is set to 20, the dependence coefficient is set to 1, the link preference coefficient is set to 2, the pheromone evaporation coefficient is initially set to 0.8, and the total pheromone amount is set to 100.
[0218] The learning rates of the two ground station agents are set to α = 0.001, and the discount factor Υ = 0.95. Take k = 6 in the action space A. Among them: (0.10, 0.50, 0.40); (0.15, 0.45, 0.40); (0.20, 0.35, 0.45); (0.30, 0.40, 0.40); (0.05, 0.30, 0.65); (0.05, 0.65, 0.30); (0.40, 0.30, 0.30); (0.50, 0.30, 0.20); (0.60, 0.15, 0.25); (0.75, 0.15, 0.10); (0.90, 0.05, 0.05); (0.96, 0.02, 0.02).
[0219] To verify the effectiveness of the algorithm, two different iperf clients are started to generate 100 small flows and 10 large flows respectively, and traffic is sent from h1 to h12 simultaneously.
[0220] The throughput comparison of Dijkstra, ECMP, and the algorithm proposed in this invention is as Figure 6 shown. As the transmission rate increases, the throughput of Dijkstra and ECMP algorithms rises slowly because the Dijkstra and ECMP algorithms do not consider the link load information and forward multiple large flows to the same path, resulting in link congestion. In addition, Dijkstra forwards both small and large flows to the same path, and the link congestion problem is more serious than that of ECMP. The algorithm proposed in this invention plans paths for small and large flows respectively, and updates the routing strategy by continuously learning historical experience and training, selects the optimal forwarding paths for small and large flows respectively, alleviates link congestion, and realizes network load balancing.
[0221] The comparison of the end-to-end average delay and end-to-end average packet loss rate of Dijkstra, ECMP, and the algorithm proposed in this invention is as Figure 7 and Figure 8 shown.
[0222] The above experimental results show that the algorithm proposed in this invention has better performance in terms of throughput, end-to-end delay, packet loss rate, and algorithm convergence while meeting the QoS requirements of user services, effectively avoiding satellite network congestion, improving the communication service level of the satellite network, and ensuring better performance of the satellite network.
[0223] The embodiment of this invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor, causes the processor to execute the steps of the multi-domain routing selection method for the satellite-ground cooperation architecture as described above.
[0224] See Figure 9 , the embodiment of this invention also provides a computer device including a memory and a processor, the memory storing a computer program, which when executed by the processor, causes the processor to execute the steps of the multi-domain routing selection method for the satellite-ground cooperation architecture as described above.
[0225] An embodiment of the present invention also provides a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the steps of the multi-domain routing selection method for the satellite-ground cooperation architecture as described above.
[0226] It can be understood that the system, device, and storage medium provided in the embodiments of the present invention correspond to the method provided in the embodiments of the present invention. For the explanations, examples, and beneficial effects of the relevant content, reference can be made to the corresponding parts in the multi-domain routing selection method for the satellite-ground cooperation architecture described above.
[0227] It should be noted that those of ordinary skill in the art can understand that all or part of the steps implemented in the embodiments of the present invention can be implemented in whole or in part through software, hardware, firmware, or any combination thereof. When implemented using hardware, it can be implemented in whole or in part in the form of purchasing standard parts or modified parts. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)).
[0228] In summary, to solve the problems of satellite network congestion and low satellite network communication service level, the present invention provides a multi-domain routing selection method for a space-ground cooperation architecture. Specifically, on the one hand, by adopting the idea of "satellite zoning and space-ground collaboration", a network is formed by combining MEO satellites, LEO satellites and ground station agents. The MEO satellite acts as a controller, responsible for issuing routing policies for LEO satellites within the domain. The LEO satellite is responsible for data forwarding, and the ground station agent is responsible for formulating routing policies, which improves the flexibility of the system framework. On the other hand, the routing table within the satellite domain is calculated using the algorithm proposed in the present invention. The traditional ant colony algorithm is redesigned. In the heuristic function of the designed ant colony algorithm, the weight parameters of bandwidth, delay and packet loss rate are set using a convolutional neural network (CNN) for large flows and small flows in the satellite network respectively by the DQN algorithm, realizing the refined design of the weight parameters and more accurate and real-time differential scheduling of elephant flows and mouse flows, accelerating the convergence speed of the optimization solution algorithm for joint multi-parameters, increasing the throughput of the satellite network, reducing the delay and packet loss rate, and meeting the requirements of high throughput for large flows and low latency for small flows in the satellite network.
[0229] It should be understood that the examples and embodiments described herein are for illustrative purposes only and are not intended to limit the present invention. Those skilled in the art can make various modifications or changes based on it. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-domain routing method for a satellite-ground cooperation architecture, characterized in that, It includes the following steps: S1. Combine MEO satellites, LEO satellites and ground station agents for networking, apply SDN to the satellite network, construct a hierarchical network architecture, and achieve global control and cross-domain routing of the satellite network; S2. Classify the traffic arriving at the edge LEO satellite nodes within the domain by threshold, and establish a multi-objective optimization model based on the throughput, delay and packet loss rate of the satellite network; S3. Improve the ant colony algorithm for the multi-objective optimization model in S2; S4. Use the bandwidth, delay and packet loss rate weight parameters in the heuristic function of the improved ant colony algorithm to train the DQN model for routing planning, and solve to obtain the optimal multi-objective optimization routing strategy based on separating large and small flows within the LEO satellite domain; S5. Divide cross-domain partition routing into intra-region routing and inter-region routing, adopt different algorithms for formulating routing strategies within and between regions, and form a complete cross-domain routing table.
2. The multi-domain routing method for a space-ground cooperation architecture according to claim 1, characterized in that The S1 further includes: S1.
1. Divide the LEO satellites into regions so that the number of in-orbit satellites and the number of satellites in the orbital plane in each region are close to the set threshold of the number of in-orbit satellites in the region; S1.
2. Deploy multiple domain controllers DCs within the satellite domain in the entire heterogeneous network composed of MEO satellites and LEO satellites. The domain controllers within the satellite domain interact with the ground station agents for the current domain and numerous satellite nodes SNs within the domain, and issue routing tables; S1.
3. The DC serves as the control layer to manage the SNs within the domain, transfer the computing center to the ground station agent for network deployment and resource scheduling. Each domain is supervised by a dedicated DC. Within the domain, the DC continuously monitors the link information, location and network topology within the domain, and transmits this information to the ground station agent of the current domain. The ground station agent performs intra-domain routing calculation and forms a routing strategy. The SNs serve as the forwarding layer, are connected to each other through ISLs, and process data packets according to the commands of the domain controllers DCs. The data packet includes at least voice data, video data, file data and picture data.
3. The multi-domain routing method for a space-ground cooperation architecture according to claim 1, wherein The S2 further includes: S2.
1. Let the total data stream bytes received at time t1 be Bt1, the total data stream bytes received at time t2 be Bt2, and the current maximum link bandwidth be BW max , and the ratio flow of the data stream bytes to the current maximum link bandwidth is calculated as: When the ratio flow is greater than 20%, classify the data flow as an elephant flow, otherwise classify the data flow as a mouse flow; S2.
2. Aiming at reducing the delay and packet loss rate of traffic transmission and improving the throughput of the network, define the total cost as: wherein, R L and R N are the costs of the elephant flow and the mouse flow respectively, H t is the throughput within the time interval t, t is the statistical time interval, O is the amount of data received, H L,A is the average throughput of the large flow, D N is the delay of the small flow, L N is the packet loss rate of the small flow; For the ant colony optimization algorithm combined with the DQN model, the optimization objective is expressed as:
4. The multi-domain routing method for a space-ground cooperation architecture according to claim 1, characterized in that In S3, the improvement of the ant colony algorithm includes the following process: S3.
1. Optimize the heuristic function of the ant colony algorithm to: Where: η ij is the comprehensive score of the link between node i and the next node j. As a heuristic function, the higher the score value, the greater the probability that the next ant will select this link; BW ij , Delay ij and Loss ij are respectively the remaining bandwidth, delay, and packet loss rate on the link between node i and node j; b, d and l are the weights corresponding to the remaining bandwidth, delay and packet loss rate respectively; The weights b, d and l of the remaining bandwidth, delay and packet loss rate are set with different scoring weight values according to the high-bandwidth characteristics of elephant flows and the low-latency requirements of mouse flows. When the data flow is identified as an elephant flow, the weight b of the remaining bandwidth is greater than d and l. When the data flow is identified as a mouse flow, the weights d of the delay and l of the packet loss rate are both greater than the weight b of the remaining bandwidth; S3.2 Incorporate the elite ant system, improve the pheromone update rule of the ant colony algorithm, and further optimize the ant colony algorithm. The corresponding pheromone update rule of the optimization system is as follows: Where: τ ij It is the pheromone between node i and node j. As the number of iterations increases, the pheromone will continuously evaporate and accumulate; ρ is the evaporation coefficient of pheromone, ρ ∈ [0, 1]; 1 - ρ is the existence coefficient of pheromone; Δτ ij is the increment of path pheromone, and its calculation formula is: is the pheromone increment generated by elite ants; σ is the number of elite ants; Q is the total amount of pheromone left by the ant when passing through the path, which is set as a constant; BW uij is the used bandwidth of the link between node i and node j; The preference of the ant to select a link is proportional to the remaining bandwidth and inversely proportional to the delay and packet loss rate, and changes with the size of the traffic. After the ant completes one cycle, the pheromone concentration is updated according to the above formulas (8) - (10).
5. The multi-domain routing method for a space-ground cooperation architecture according to claim 4, characterized in that, In the above S3, to improve the global search ability and convergence speed of the improved ant colony algorithm, the evaporation coefficient ρ is taken in stages: Among them, I cur represents the current iteration number of the algorithm, and I max represents the maximum number of iterations set by the algorithm.
6. The multi-domain routing method for a space-ground cooperation architecture according to claim 1, characterized in that In the above S4, use the CNN convolutional neural network to calculate the value of the total pheromone Q, and map the state space State, reward function Reward, and action space Actor in the DQN model as follows: State space S: The state space is represented as a three-dimensional array or matrix with a shape of n×3, where each row corresponds to the link state of a network, and there are a total of n network link states: S = [ B i , D i , L i , i ∈ [1, n] (12) Where B, D, and L are the remaining bandwidth matrix, delay matrix, and packet loss rate matrix of the network topology respectively; Reward function R: For large flows, high throughput needs to be obtained, and the reward function is shown in formula (4); For small flows, low delay and low packet loss rate need to be obtained, and the reward function is shown in formula (5); Action space A: Executing action a means adjusting the weights of the ant colony algorithm according to the current network state. Each action corresponds to a specific weight adjustment, and the expression of the action space is: A = [a1, a2, a3, …, a k (13) where a k = [b, d, l]; Design two different action spaces for elephant flows and mouse flows respectively: Let the action space of the mouse flow be A S be:[[]] Let the action space \(A\) of the elephant flow L be 7. The multi-domain routing method for a space-ground cooperation architecture according to claim 1, characterized in that In the above S5, the satellite inter-domain routing table is calculated by the ground station agent using the Dijkstra algorithm through the link delay between the boundary routing node pairs of adjacent regions obtained from the ephemeris information, and is distributed to each boundary routing node satellite in the adjacent region by pre-storing. The satellite intra-domain routing table is calculated using the improved ant colony algorithm optimized by the DQN model.
8. The multi-domain routing method for a space-ground cooperation architecture according to claim 7, characterized in that, The ground station agent returns the formed routing strategy to the DC, and the DC inserts the routing strategy into the segment routing list, and starts forwarding data from the edge LEO satellite node, and reaches the packet destination through intra-region routing and inter-region routing.
9. A computer-readable storage medium, characterized in that, A computer program is stored, and when the computer program is executed by a processor, the processor is caused to execute the steps of the multi-domain routing selection method for a satellite-ground cooperation architecture as described in any one of claims 1 - 8.
10. Computer device, characterized in that, It includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the multi-domain routing selection method for a satellite-ground cooperation architecture as described in any one of claims 1 - 8.
Citation Information
Cited By
Computing task unloading method based on heaven and earth cross-domain multi-slice computing resource management
CN121711710A