Data stream transmission scheduling method based on distributed SDN satellite network
By adopting distributed SDN and deep reinforcement learning methods in large-scale low-earth orbit satellite networks, the problems of data transmission routing complexity and congestion in giant constellations are solved, and load balancing and delay optimization are achieved.
Patent Information
- Application Number
- CN202510173766.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-13
AI Technical Summary
Large-scale low Earth Orbit (LEO) satellite networks face problems such as high routing computing complexity, network congestion, path state diversity and time-varying when data transmission is carried out. Existing routing algorithms are difficult to directly apply to giant constellations.
A data stream transmission scheduling method based on distributed SDN satellite network is proposed, and a data stream transmission scheduling algorithm with deep reinforcement learning is adopted, including an inter-cluster routing algorithm based on estimating propagation delay and an in-cluster routing algorithm based on DQN to achieve load balancing and reduce end-to-end delay.
The load balancing routing of data transmission streams for large-scale satellite networks is realized, improving the load balancing effect of the network and the end-to-end delay performance of the transmission streams.
Smart Images

Figure CN120150786A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of satellite routing, and particularly relates to a data flow transmission scheduling method based on a distributed SDN satellite network. Background Art
[0002] Emerging megaconstellations utilize thousands of Low Earth Orbit (LEO) satellites, which can ensure low-latency, high-throughput, and high-transmission-efficiency data transmission globally, and are particularly crucial for users in remote areas where ground infrastructure is difficult to reach. With its extensive coverage and abundant bandwidth resources, this technology can not only effectively fill the gaps in ground network services but also become a key link in realizing the integration of terrestrial, maritime, and space communication networks. The low Earth orbit satellite network is deployed at an altitude of 160 to 2000 kilometers from the Earth. Compared with Geostationary Earth Orbit (GEO) satellites, the low-orbit satellite network can achieve lower latency and higher data transmission capabilities with lower power consumption, and can provide stable communication services for global users. In this context, megaconstellations, as an emerging communication network model, have gradually attracted extensive attention in the academic community.
[0003] Although low-orbit megaconstellations have many advantages, they still face a series of technical challenges in practical applications. Compared with small satellite networks, LEO megaconstellations have characteristics such as higher dynamics, more frequent network topology switching, and more satellite nodes, and thus also face the following challenges. First, due to the larger number of satellites and satellite links, there are thousands of satellites in the LEO megaconstellation that need to transmit data simultaneously, resulting in a significant increase in routing calculation complexity and rerouting signaling overhead. Second, due to the uneven distribution of ground users and the bursty characteristics of traffic, the satellite link traffic between some cities will be too large, leading to congestion in the satellite network during data transmission, which in turn affects user service performance. Finally, since the arrival time distribution of data flows in the satellite network is unknown, the path states such as throughput and delay are diverse and time-varying, which requires the routing algorithm to be able to intelligently adapt to environmental changes.
[0004] Currently, routing algorithms for low-orbit satellites have received extensive attention in the academic community, but there is still relatively little discussion on megaconstellations. Most of the existing research on satellite network routing algorithms focuses on small-scale satellite constellations with about 100 constellation satellites. However, due to the huge number of nodes and the huge overhead of link state collection in large-scale satellite networks, the routing algorithms for small-scale satellite constellations are difficult to directly apply to megaconstellations, and there is an urgent need to solve the data transmission and load balancing problems in large-scale satellite networks. Summary of the Invention
[0005] The main objective of the present invention is to propose a data flow transmission scheduling method based on a distributed SDN satellite network, aiming at the data transmission problem in large-scale satellite networks, designing a data flow transmission scheduling algorithm based on deep reinforcement learning, realizing load-balanced routing of data transmission flows in large-scale satellite networks, and improving the overall load-balancing effect of the network and the end-to-end delay performance of transmission flows.
[0006] To achieve the above objective, the present invention provides a data flow transmission scheduling method based on a distributed SDN satellite network, and the method includes the following steps:
[0007] Step S10, establish a distributed SDN satellite network model according to the application background of the large-scale satellite network, and construct a data flow load-balanced routing problem of the large-scale satellite network under the distributed SDN satellite network model;
[0008] Step S20, for the data flow load-balanced routing problem of the large-scale satellite network, adopt the data transmission strategy of the distributed SDN satellite network to achieve the load-balancing effect and reduce the end-to-end delay.
[0009] A further technical solution of the present invention is that in step S20, for the data flow load-balanced routing problem of the large-scale satellite network, adopting the data transmission strategy of the distributed SDN satellite network to achieve the load-balancing effect and reduce the end-to-end delay includes two stages: an inter-cluster routing algorithm based on estimated propagation delay and an intra-cluster routing algorithm based on DQN. Among them,
[0010] The inter-cluster routing algorithm based on estimated propagation delay combines the geographical location of the target cluster and the load conditions of adjacent clusters, introduces the estimated remaining propagation delay as a decision-making index, and adaptively selects a cluster with a lighter load and closer to the destination as the forwarding direction of the next-hop cluster;
[0011] In the intra-cluster routing algorithm based on DQN, the optimization problem is transformed into a Markov decision process. The state space and action space of the intelligent agent are defined based on the distributed SDN satellite network, and a reward function is constructed according to the optimization objective.
[0012] A further technical solution of the present invention is that in step S10, the steps of establishing a distributed SDN satellite network model according to the application background of the large-scale satellite network include:
[0013] Construct a LEO mega-constellation network model, and construct a satellite communication model considering the actual channel conditions. According to the topological structure of the mega-constellation, divide the LEO mega-constellation network into multiple clusters, where satellite nodes always belong to the same cluster to avoid cumbersome management switching problems.
[0014] A further technical solution of the present invention is that in the step of constructing the LEO giant constellation network model, a Walker-Delta giant constellation is considered, where:
[0015] The topology of the giant constellation network is represented as where is the set of satellite nodes, and ε = {e uv |u≠v} represents the set of ISLs established between satellite nodes; the undirected edge is denoted as uv, and the source-destination edge is denoted as; since the satellites are distributed on the spherical surface, is used to represent the coordinates of satellite u, where r represents the altitude of the satellite, and θ i and represent the polar angle and azimuth angle of the satellite respectively; each satellite will establish 4 ISLs, including 2 inter-orbit ISLs and 2 intra-orbit ISLs, and generally, the inter-orbit ISLs are not established between the ascending satellite and the descending satellite; for each ISL, e uv = 1 indicates that there is an ISL link connection between satellite u and satellite v, otherwise it represents no ISL connection; the LEO giant constellation includes N = N p ×M p satellites, where N p is the number of orbital planes, M p is the number of satellites in each plane, all orbits have the same inclination angle α, the satellite orbital altitude is r, and the spacing along the equator is equal; Mp satellites are evenly distributed on each plane, and the phase deviation between the satellites on adjacent planes is Δf = 2πF / (N p M p ), where, is the phase factor; then the Walker-Delta constellation can be represented by α:N p M p / N p / F.
[0016] A further technical solution of the present invention is that in the step of constructing the satellite communication model considering the actual channel conditions, it is assumed that the noise is additive white Gaussian noise; ||uv|| is used to represent the distance between satellite u and satellite v, and through the derivation in the spherical coordinate system, we can obtain:
[0017]
[0018] where the coordinates of satellite u and satellite v are respectively and
[0019] The maximum line-of-sight distance refers to the maximum straight-line distance between two points where they can directly see each other without physical obstacles blocking, which determines the maximum visible range between satellites; for the maximum line-of-sight distance between two points on the Earth's surface, when calculating the maximum line-of-sight distance, factors such as the curvature of the Earth and the heights of the observation point and the target point need to be considered, and the following simplified formula can be used for estimation:
[0020]
[0021] where d represents the maximum line-of-sight distance, R is the radius of the Earth, h 1 and h 2 are the heights of the two observation points from the ground respectively;
[0022] Assume that I * (u, v) is the maximum line-of-sight distance between satellite u and satellite v. When ||uv|| > I * (u, v), an ISL cannot be established; then the FSPL can be expressed as:
[0023]
[0024] where f is the frequency of light and c is the speed of light;
[0025] Assume that the wireless channel is symmetric. We can define the signal-to-noise ratio of a satellite pair as:
[0026]
[0027] In the formula, P t is the transmit power, G t is the transmit antenna gain, G t is the receive antenna gain, k B is the Boltzmann constant, B is the channel bandwidth in hertz, and T is the thermal noise in kelvin;
[0028] Assume that satellite u chooses to communicate with v at the maximum data rate in a non-interfering environment. Then the link capacity C uv can be calculated as:
[0029] C uv = Blog(1 + SNR uv ) (5).
[0030] A further technical solution of the present invention is that in the step of dividing the LEO mega-constellation network into multiple clusters according to the topological structure of the mega-constellation, where satellite nodes always belong to the same cluster to avoid cumbersome management handover problems, the clusters are divided into a rectangular area according to adjacent orbital satellites, which contains n × m satellites, 2 ≤ n ≤ N p , 2 ≤ m ≤ M p; These areas are relatively fixed to reduce the update frequency and the routing overhead for each update; the geocentric satellite within the cluster serves as the control satellite and is equipped with an SDN network controller. Satellites within the cluster will regularly report their own load information to the SDN network controller, and the SDN controller will make routing decisions for routing events within the cluster based on the collected link information; meanwhile, control satellites will regularly exchange the load information of adjacent clusters through ISLs;
[0031] The cluster set of the entire satellite network consists of 1 ≤ k ≤ K, where K is the number of clusters, and the central control satellite corresponding to each cluster is represented by CH k It is equipped with an SDN network controller, which can monitor the link status information of each satellite within the cluster and perform routing planning for the network traffic within the cluster; other satellites are data transmission satellites, responsible for transmitting the data traffic in the cluster and regularly reporting the link status to the control satellite.
[0032] A further technical solution of the present invention is that in step S10, the steps of constructing the data flow load balancing routing problem of a large-scale satellite network under the distributed SDN satellite network model include:
[0033] Define the set of flows in the network as 1 ≤ i ≤ I; where f i represents the traffic demand in the network, represents the starting satellite of flow f i , represents the ending satellite of flow f i , volume i represents the data volume of flow f i , and I is the number of flows in the network; for each there will be multiple feasible paths in the satellite network to achieve traffic transmission; represents the set of alternative traffic transmission paths for flow f i , where represents an alternative path; let the binary variable represent whether path p i of flow f i contains ISLe uv , represents e uv in path p i , otherwise the number of flows passing through e uv can be calculated through ; since there may be multiple data flows transmitting simultaneously in an ISL, it is assumed here that the traffic on the same link fairly shares the link capacity, and the flow f i in cluster c kThe bandwidth within is:
[0034]
[0035] where p i,k is the path of flow f i within cluster c k ;
[0036] And the bandwidth of flow f i can be calculated as:
[0037]
[0038] The latency of flow f i can be calculated as:
[0039]
[0040] The first term in Equation (8) represents the transmission time of the flow, and the second term represents the propagation delay of the multi-hop ISL;
[0041] Based on the above model, the routing problem in the giant constellation is formulated as a non-linear programming problem. Each control satellite equipped with an SDN controller in the cluster needs to formulate an optimal routing strategy for the flows entering the cluster to maximize the system throughput while satisfying the end-to-end delay constraints of each flow; the optimization problem can be expressed as follows:
[0042]
[0043] where the constraint condition C1 is the bandwidth of flow f i , which is the minimum value of the bandwidth allocated by each cluster SDN controller. The constraint condition C2 is the protection constraint condition of the flow, which ensures the traffic conservation of each end satellite; the constraint condition C3 sets the delay constraint α and hopes that the latency i of each flow f does not exceed α i , and α i is the QoS constraint of flow f i , and is in a certain proportion to the shortest path latency of flow f i .
[0044] A further technical solution of the present invention is that in the step of combining the geographical location of the target cluster and the load conditions of adjacent clusters in the inter-cluster routing algorithm based on the estimated propagation delay, introducing the estimated remaining propagation delay as a decision metric and adaptively selecting the next-hop cluster with a lighter load and closer to the destination as the forwarding direction, it includes:
[0045] Assume that the nodes in the network know their own geographical location coordinates and have obtained the geographical location information of the target node. For flow f i, in order to calculate the current cluster c k to the target cluster c d of the estimated remaining propagation delay, it is necessary to share the control satellites CH j corresponding to the four adjacent clusters c j , j ∈ {1, 2, 3, 4}, and the spatial coordinates and obtain in advance the spatial coordinates of the control satellite CH d corresponding to the target cluster c d Then the estimated remaining propagation delay can be calculated as follows:
[0046]
[0047] where ||CH d CH j || is the spatial distance between the target control satellite CH d and the adjacent control satellite CH j , which can be calculated by Equation (1);
[0048] Define the decision-making index of cluster c k as and can be calculated by the following formula:
[0049]
[0050] Assume that the current cluster c k has four adjacent clusters c j , j ∈ {1, 2, 3, 4} available for forwarding. Among them, one cluster is the previous hop of the inter-cluster routing. Then the forwarding direction of the candidate next-hop cluster is c j′ , j′ ∈ {1, 2, 3}; The inter-cluster routing algorithm based on the estimated propagation delay will calculate the backlog decision-making index of the adjacent clusters and select the next-hop cluster with the smallest decision-making index as the forwarding direction of the inter-cluster routing:
[0051]
[0052] A further technical solution of the present invention is that in the intra-cluster routing algorithm based on DQN, when transforming the optimization problem into a Markov decision process, defining the state space, action space of the intelligent agent, and constructing the reward function based on the distributed SDN satellite network, for the flow f k entering the cluster c i , define the corresponding control satellite CH k as the state space, action space, and reward function of the intelligent agent as follows:
[0053] (1) State space:
[0054] Let Denote the state space, the binary tuple s i,k =[M k , δ i,k represents the state observed by the control satellite CH k after it enters the cluster in the flow f i ; M k =[m uv n×n represents the traffic distribution matrix in the cluster c k where m uv is the traffic in the cluster, and the calculation formula is:
[0055]
[0056] represents the routing request of the flow f i in the cluster c k ; where represents the starting satellite of the flow f i in the cluster c k ; l i,k represents the forwarding direction of the flow f i in the cluster c k ; volume i represents the data volume size of the flow f i ;
[0057] (2) Action space:
[0058] When the control satellite CH k receives the flow f k entering the cluster c i and observes the state s i,k , it will execute the action a i,k ; The action space k that CH can choose comes from the set of candidate equivalent paths i,k constructed according to the routing request δ That is Considering that there may be multiple equivalent paths between each source-destination node pair in the cluster, the agent will select m optional paths between the source-destination node pairs from the cluster according to the number of hops as the set of feasible paths;
[0059] (3) Reward function:
[0060] Define the reward obtained by the control satellite CH k after executing the action a i,k in the state s i,k as r i,k ; Since the goal of the proposed optimization problem is to maximize the system throughput while satisfying the end-to-end delay constraints of each data stream, the reward function is based on the bandwidth size of the allocated path, and the reward r is defined as follows: i,k as follows:
[0061]
[0062] where represents the control satellite CH k the allocated path delay, and α k represents the delay limit within the cluster; when the path allocated by the control satellite within the cluster satisfies the delay constraint, the DRL agent will receive a positive reward, and the larger the bandwidth of the allocated path, the greater the reward; when the path allocated by the control satellite within the cluster does not satisfy the delay constraint, the DRL agent will receive a penalty, and the smaller the bandwidth of the allocated path, the greater the penalty; the Markov decision problem is described as (s, a, p, r), and the optimization goal of the constructed optimization problem is to maximize the total throughput of the giant constellation. This optimization problem is then solved by an algorithm based on reinforcement learning.
[0063] A further technical solution of the present invention is that the in-cluster routing algorithm based on DQN includes:
[0064] (1) Action selection:
[0065] When the flow f i enters the cluster c k the agent CH k will observe the link information within the cluster, and combined with the current state s i,k the agent CH k will construct an equivalent path set as the action space, and select an action a from the action space i,k based on the soft-∈-greedy policy, that is, with a probability of 1 - ∈, greedily select an action that maximizes the Q value, or with a probability of ∈, randomly select a random action:
[0066]
[0067] (2) Exploration strategy:
[0068] The soft-∈-greedy strategy is adopted to guide the action selection of the control satellite. In this strategy, a probability that gradually decreases as the number of training episodes of the agent increases is defined to randomly select an action. The specific formula is expressed as follows:
[0069]
[0070] where ∈ is a positive integer, representing the probability that the agent randomly selects an action; at the beginning of training, a larger ∈ means that the agent has a greater probability of randomly selecting an action from the action space, so as to widely explore possible optimal action options.
[0071] (3) Neural network:
[0072] Split the maximization operator in the DQN network into two independent steps of action selection and action evaluation, and use the current Q-network parameter θ i to select the optimal action, and the target Q-network parameter θ i - To further stabilize the training process, the DQN network uses two neural networks: one is the evaluation network, which is used to predict the Q value of the current state; the other is the target network, which is used to calculate the target Q value; the parameters of the target network are periodically copied from the evaluation network instead of being updated in real time, which can effectively reduce the fluctuation of the target Q value and improve the stability of training. Two sets of parameters are used to separate the process of action selection and action evaluation, thereby reducing the risk of overestimation: each agent is set with its own independent estimation network Q k (s i,k , a i,k ; θ k ) and the target network where the parameters of the two networks are θ k and The experience replay of each agent is also independently initialized and uses random samples to update the network parameters.
[0073] After executing the action a i,k , the agent CH k will receive the reward r i,k , and observe the next state s i+1,k ; based on the above information, the agent will convert the process (s i,k , a i,k , r i,k , s i+1,k ) into the experience replay pool R, and then randomly draw a batch of samples and perform learning;
[0074] At the end of each decision-making, the gradient descent method is used to fit the neural network, and the parameter θ is updated by minimizing the objective function k ; the objective function L i,k is given by the following formula:
[0075] L i,k =(y i,k -Q k (s i,k , a i,k ; θ k ))2 (17);
[0076] y i,k is the target value, the maximum expected value that may be obtained after performing the action, and can be calculated by the following formula:
[0077]
[0078] where γ is the discount factor of future rewards, θ k and are the parameters of the Q estimation network and the Q target network respectively;
[0079] Then, the parameters of the neural network are updated by gradient descent:
[0080]
[0081] where α is the learning rate, which is used to balance the learning speed and accuracy.
[0082] The beneficial effects of the data flow transmission scheduling method based on the distributed SDN satellite network of the present invention are:
[0083] (1) A distributed SDN satellite network model is established according to the application background of the large-scale satellite network, and the data flow load balancing routing problem of the large-scale satellite network is constructed under this model;
[0084] (2) For this optimization problem, a data transmission strategy based on the distributed SDN satellite network is proposed, and this strategy can achieve better load balancing effect and smaller end-to-end delay. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the structures shown in these drawings.
[0086] Figure 1 is a schematic flowchart of a preferred embodiment of the data flow transmission scheduling method based on the distributed SDN satellite network of the present invention;
[0087] Figure 2 is a schematic diagram of the distributed SDN satellite network model;
[0088] Figure 3 is a schematic diagram of the inter-cluster routing algorithm;
[0089] Figure 4It is a framework diagram of a load - balancing routing algorithm based on a distributed SDN.
[0090] The realization, functional features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0091] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0092] The present invention proposes a data - stream transmission scheduling method based on a distributed SDN satellite network. Please refer to Figure 1 , the preferred embodiment of the data - stream transmission scheduling method based on the distributed SDN satellite network of the present invention includes the following steps:
[0093] Step S10: Establish a distributed SDN satellite network model according to the application background of a large - scale satellite network, and construct a data - stream load - balancing routing problem of the large - scale satellite network under the distributed SDN satellite network model;
[0094] Step S20: For the data - stream load - balancing routing problem of the large - scale satellite network, adopt the data - transmission strategy of the distributed SDN satellite network to achieve the load - balancing effect and reduce the end - to - end delay.
[0095] In this embodiment, in the scenario of a large - scale satellite network, aiming at the problem of unbalanced data - transmission load in a large - scale satellite network, a data - transmission scheduling strategy based on deep reinforcement learning is designed. The purpose is to achieve load - balancing routing of the data - transmission flow of the large - scale satellite network. This strategy design needs to make a trade - off among the transmission delay of the flow, the calculation delay, the overall energy consumption of the satellite, the overall load - balancing performance of the network, and meeting the QoS communication requirements. Such a joint design is challenging and has application prospects.
[0096] First, considering the large number of satellites in a large-scale satellite network, centralized routing control can lead to significant propagation delays and huge signaling overheads. In this embodiment, a distributed SDN (Soft Defined Network) satellite network model is designed, which can effectively address the routing problems brought about by the increasing network scale of large-scale satellite networks. Secondly, based on the actual channel and communication scenarios of satellite links, the present invention constructs an optimization equation for maximizing the system throughput of transmission flows and satisfying the QoS (Quality of Service) of the flows. Finally, due to the uncertainty of flow arrival times and the highly dynamic nature of giant constellations, this problem is difficult to solve using traditional algorithms. The present invention transforms this optimization problem into a competition problem for bandwidth among different flows and designs a data transmission scheduling strategy based on deep reinforcement learning. The proposed algorithm includes two stages: inter-cluster routing based on estimated propagation delay and intra-cluster routing based on the DQN (Deep Q Network) network. The algorithm designed in this embodiment has significant advantages in aspects such as the load balancing effect of the network, the end-to-end delay of flows, and the QoS performance of users.
[0097] The beneficial effects of this embodiment are as follows: (1) A distributed SDN satellite network model is established according to the application background of large-scale satellite networks, and the data flow load balancing routing problem of large-scale satellite networks is constructed under this model; (2) For this optimization problem, a data transmission strategy based on a distributed SDN satellite network is proposed, and this strategy can achieve a better load balancing effect and a smaller end-to-end delay.
[0098] Further, in this embodiment, in step S20, for the data flow load balancing routing problem of large-scale satellite networks, a data transmission strategy of a distributed SDN satellite network is adopted to achieve a load balancing effect and reduce the end-to-end delay, which includes two stages: an inter-cluster routing algorithm based on estimated propagation delay and an intra-cluster routing algorithm based on DQN.
[0099] Among them, the inter-cluster routing algorithm based on estimated propagation delay combines the geographical location of the target cluster and the load conditions of adjacent clusters, introduces the estimated remaining propagation delay as a decision-making metric, and adaptively selects a cluster with a lighter load and closer to the destination as the forwarding direction of the next-hop cluster.
[0100] In the intra-cluster routing algorithm based on DQN, the optimization problem is transformed into a Markov decision process. The state space and action space of the agent are defined based on the distributed SDN satellite network, and a reward function is constructed according to the optimization objective.
[0101] In step S10, the steps of establishing a distributed SDN satellite network model according to the application background of large-scale satellite networks include:
[0102] Construct a LEO mega-constellation network model and a satellite communication model considering actual channel conditions. According to the topological structure of the mega-constellation, divide the LEO mega-constellation network into multiple clusters, where satellite nodes always belong to the same cluster to avoid cumbersome management handover problems.
[0103] In this embodiment, a LEO mega-constellation network model is first constructed, and a satellite communication model considering actual channel conditions is constructed. According to the satellite network topological structure of the mega-constellation, the LEO mega-constellation network is divided into multiple clusters. The geographical center satellite within the cluster is the control satellite and is equipped with an SDN network controller to perform centralized control over the traffic path planning within the cluster. The geographical center satellite within the cluster is the control satellite and is equipped with an SDN network controller. Satellites within the cluster will regularly report their own load information to the SDN network controller, and the SDN controller will make routing decisions for routing events within the cluster based on the collected link information. At the same time, the control satellites will regularly exchange the load information within adjacent clusters through ISLs. Through this distributed SDN design of dividing clusters, centralized control over the traffic path planning within the cluster can be achieved. The distributed SDN satellite network model can combine the advantages of the distributed scheme in satellite routing. By regularly exchanging load information between control satellites of different clusters through ISLs, it can achieve a rapid response to the load changes of large-scale satellite networks; and the advantages of the centralized method, implementing active load balancing control over the traffic within the cluster.
[0104] Then, the load balancing routing problem to be solved in this embodiment is described. The routing problem in the giant constellation is formulated as a non-linear programming problem. Each control satellite equipped with an SDN controller in a cluster needs to formulate an optimal routing strategy for the flows entering the cluster to maximize the system throughput while satisfying the end-to-end delay constraints of each flow. However, due to multiple challenges, it is still not an easy task to design an algorithm to solve the above problem. Since the optimization variable is a binary integer, the problem can be reduced to a bin-packing problem and a shortest path problem and proven to be a Binary Integer Programming (BIP) problem, which has been proven to be an NP-hard problem. Secondly, since the low-orbit satellite orbits are in a high-speed running state, this dynamics leads to topological fluctuations and frequent visibility changes between satellites, which will cause the visibility matrix of satellite links to be updated frequently. Therefore, the optimization problem needs to be effectively solved within each short interval when the network maintains a short-term static topology. Finally, although the running trajectories of LEO satellites in the satellite constellation are predictable, in the actual user flow arrival model, the specific arrival times of user flows are uncertain. Therefore, it is difficult to pre-determine and allocate high-bandwidth paths, and it is difficult to use traditional algorithms such as greedy algorithms and heuristic algorithms to solve.
[0105] Considering that there are a large number of equivalent paths with similar delays in the giant constellation, and transmitting flows using equivalent paths will not cause a significant increase in delay. Therefore, the competition problem of different flows for bandwidth in the optimization problem can be transformed into reducing the bandwidth competition during flow transmission by scheduling routing paths to different equivalent paths. The proposed transmission strategy for the giant constellation based on a distributed SDN network includes two stages: an inter-cluster routing algorithm based on estimated propagation delay and an intra-cluster routing algorithm based on DQN. The inter-cluster routing algorithm based on estimated propagation delay combines the geographical location of the target cluster and the load conditions of adjacent clusters, introduces the estimated remaining propagation delay as a decision metric, and adaptively selects a cluster with a lighter load and closer to the destination as the forwarding direction of the next-hop cluster; the intra-cluster routing algorithm uses a learning-based active load algorithm to schedule flows to appropriate equivalent paths to reduce bandwidth competition under the premise that the arrival distribution of flows cannot be known in advance, so as to achieve high-bandwidth routing transmission and meet the corresponding delay requirements, and realize the long-term optimization of the load.
[0106] Specifically, in the inter-cluster routing algorithm based on the estimated propagation delay, based on the topological structure characteristics of the distributed SDN satellite network, in order to minimize the signaling overhead caused by large-scale collection of link states between clusters as much as possible, this algorithm combines the geographical location of the target cluster and the load conditions of adjacent clusters, introduces the estimated remaining propagation delay as the backlog decision metric for the cluster, and adaptively selects the cluster with lighter load and closer distance to the destination as the forwarding direction of the next-hop cluster. In the intra-cluster routing algorithm based on DQN, the optimization problem is transformed into a Markov decision process. The state space and action space of the agent are defined based on the distributed SDN satellite network, and the reward function is constructed according to the optimization goal. The control satellite, as an independent agent, learns by observing the state within the cluster and allocates high-bandwidth paths for each flow entering the cluster, and meets the corresponding delay constraints to achieve long-term balanced optimization of the load.
[0107] The data flow transmission scheduling method based on the distributed SDN satellite network of the present invention is further elaborated in detail below.
[0108] I. System Model
[0109] (1) Network Model:
[0110] The present invention considers a Walker-Delta giant constellation. The giant constellation network topology is represented as where is the set of satellite nodes, and ε = {e uv |u ≠ v} represents the set of established ISLs between satellite nodes. In the present invention, the undirected edge is denoted as uv, and the source-destination edge is denoted as (u, v). Since the satellites are distributed on the spherical surface, is used to represent the coordinates of satellite u, where r represents the height of the satellite, θ i and represent the polar angle and azimuth angle of the satellite respectively. Each satellite will establish 4 ISLs, including 2 inter-orbit ISLs and 2 intra-orbit ISLs. Among them, the inter-orbit ISLs are generally not established between the ascending satellite and the descending satellite. For each ISL, e uv = 1 indicates that there is an ISL link connection between satellite u and satellite v, otherwise it represents no ISL connection. The LEO giant constellation includes N = N p × M p satellites, where N p is the number of orbital planes, M p is the number of satellites in each plane. All orbits have the same inclination angle α, the satellite orbit height is r, and the spacing along the equator is equal. M p satellites are evenly distributed on each plane, and the phase deviation between the satellites on adjacent planes is Δf = 2πF / (N p Mp ), where is the phase factor. Then the Walker-Delta constellation can be represented by α:N p M p / N p / F.
[0111] (2) Satellite communication model:
[0112] Since satellite-to-satellite communication occurs in a free space environment, it is mainly affected by free space path loss (FSPL) and thermal noise power. Assume the noise is additive white Gaussian noise. Let ||uv|| represent the distance between satellite u and satellite v. Through the derivation in spherical coordinates, we can obtain:
[0113]
[0114] where the coordinates of satellite u and satellite v are respectively and
[0115] The maximum line-of-sight (LOS) distance refers to the maximum straight-line distance between two points where they can directly see each other without physical obstacles blocking. It determines the maximum visible range between satellites. For the maximum LOS distance between two points on the Earth's surface, when calculating the maximum LOS distance, factors such as the curvature of the Earth and the heights of the observation point and the target point need to be considered. The following simplified formula can be used for estimation:
[0116]
[0117] where d represents the maximum LOS distance, R is the radius of the Earth, h 1 and h 2 are the heights of the two observation points from the ground respectively.
[0118] Assume I * (u, v) is the maximum LOS distance between satellite u and satellite v. When ||uv|| > I * (u, v), an ISL cannot be established. Then the FSPL can be expressed as:
[0119]
[0120] where f is the frequency of light and c is the speed of light.
[0121] Assume the wireless channel is symmetric. We can define the signal-to-noise ratio of a satellite pair as:
[0122]
[0123] In the formula, P tis the transmission power, G t is the transmitting antenna gain, G t is the receiving antenna gain, k B is the Boltzmann constant, B is the channel bandwidth in Hertz, and T is the thermal noise in Kelvin.
[0124] Assume that satellite u chooses to communicate with v at the maximum data rate in a non-interfering environment, then the link capacity C uv can be calculated as:
[0125] C uv = Blog(1 + SNR uv ) (5);
[0126] (3) Distributed SDN network model:
[0127] According to the topology of the mega-constellation, the LEO mega-constellation network is divided into multiple clusters, where satellite nodes always belong to the same cluster to avoid cumbersome management handover problems. The clusters are divided into a rectangular area according to adjacent orbital satellites, which contains n×m satellites, 2 ≤ n ≤ N p , 2 ≤ m ≤ M p . These areas are relatively fixed to reduce the update frequency and reduce the routing overhead of each update, as Figure 2 shown. The geocentric satellite within the cluster is the control satellite and is equipped with an SDN network controller. The satellites within the cluster will regularly report their own load information to the SDN network controller, and the SDN controller will make routing decisions for the routing events within the cluster based on the collected link information. At the same time, the control satellites will regularly exchange the load information within adjacent clusters through ISL. Through this distributed SDN design of dividing clusters, centralized control can be performed on the traffic path planning within the cluster. The distributed SDN-satellite network model can combine the advantages of the distributed scheme in satellite routing. By regularly exchanging load information through ISL between different cluster control satellites, it can achieve a rapid response to the load changes of large-scale satellite networks; and the advantages of the centralized method, actively performing load balancing control on the traffic within the cluster.
[0128] The set of clusters of the entire satellite network is represented by 1 ≤ k ≤ K, where K is the number of clusters, as Figure 2 shown. The central control satellite corresponding to each cluster is represented by CH k , which is equipped with an SDN network controller. It can monitor the link status information of each satellite within the cluster and perform routing planning on the network traffic within the cluster; other satellites are data transmission satellites, responsible for transmitting the data traffic within the cluster and regularly reporting the link status to the control satellite.
[0129] (4) Optimization problem:
[0130] Define the set of flows in the network as 1 ≤ i ≤ I. Where f i represents the traffic demand in the network, represents the starting satellite of flow f i and represents the ending satellite of flow f i and volume i represents the data volume of flow f i and I is the number of flows in the network. For each there are multiple feasible paths in the satellite network to achieve traffic transmission. represents the set of alternative traffic transmission paths for flow f i where represents an alternative path. Let the binary variable represent whether path p i of flow f i contains ISLe uv , represents e uv in path p i otherwise the number of flows passing through e uv can be calculated by . Since there may be multiple data flows transmitted simultaneously in an ISL, it is assumed that the traffic on the same link fairly shares the link capacity, and the bandwidth of flow f i in cluster c k can be calculated as:
[0131]
[0132] where p i,k is the path of flow f i in cluster c k .
[0133] And the bandwidth of flow f i can be calculated as:
[0134]
[0135] The delay of flow f i can be calculated as:
[0136]
[0137] The first term in Equation (8) represents the transmission time of the flow, and the second term represents the propagation delay of the multi-hop ISL.
[0138] Based on the above model, the routing problem in the megaconstellation in the present invention is formulated as a non-linear programming problem. Each control satellite equipped with an SDN controller in each cluster needs to formulate an optimal routing strategy for the flows entering the cluster, so as to maximize the system throughput while satisfying the end-to-end delay constraints of each flow. The optimization problem can be expressed as follows:
[0139]
[0140] where the constraint condition C1 is the bandwidth of flow f i , which is the minimum value of the bandwidth allocated to each cluster's SDN controller. The constraint condition C2 is the protection constraint condition of the flow, which ensures the conservation of traffic at each end satellite; the constraint condition C3 sets the delay constraint α and hopes that the delay i of each flow f does not exceed α i , and α i is the QoS constraint of flow f i , and is in a certain proportion to the shortest path delay of flow f i .
[0141] II. Problem Analysis and Solutions
[0142] (1) Problem Analysis
[0143] For problem P1, representing the routing strategy determined by the SDN network controllers equipped in each cluster, P1 can be solved by finding the optimal path strategy. However, due to multiple challenges, it is still not an easy task to design an algorithm to solve the above problem. Since the optimization variable is a binary integer, the problem can be proven to be a binary integer programming problem by being reduced to a bin-packing problem and a shortest path problem, and has been proven to be an NP-hard problem. Solving it directly as an integer programming problem is very costly. For example, the running time of the solver for a single snapshot on a ThinkPad X1 Carbon laptop ranges from several minutes to several hours, especially when the constellation scale and the number of users increase, the solving time will be longer. Secondly, since the low-orbit satellite orbits are in a high-speed running state, this dynamics leads to topological fluctuations and frequent visibility changes between satellites, which will cause the visibility matrix of the satellite links to be updated frequently. Therefore, the optimization problem needs to be effectively solved within each short interval when the network maintains a short-term static topology. Finally, although the running trajectories of the LEO satellites in the satellite constellation are predictable, in the actual user flow arrival model, the specific arrival times of the user flows are uncertain. Therefore, it is difficult to pre-determine and allocate high-bandwidth paths, and it is difficult to use traditional algorithms such as greedy algorithms and heuristic algorithms to solve. Therefore, the subsequent focus is on using different methods to solve the optimal path strategy.
[0144] Considering that there are a large number of equivalent paths with similar delays in the mega-constellation, and using equivalent path convection for transmission will not cause a significant increase in delay. Therefore, the competition problem of different flows for bandwidth in the optimization problem can be transformed into reducing the bandwidth competition during the flow transmission process by scheduling the routing paths to different equivalent paths. The proposed transmission strategy for the mega-constellation based on the distributed SDN network includes two stages: the inter-cluster routing algorithm based on the estimated propagation delay and the intra-cluster routing algorithm based on DQN. The inter-cluster routing algorithm based on the estimated propagation delay combines the geographical location of the target cluster and the load conditions of adjacent clusters, introduces the estimated remaining propagation delay as a decision metric, and adaptively selects the cluster with lighter load and closer to the destination as the forwarding direction of the next-hop cluster; the intra-cluster routing algorithm adopts a learning-based active load algorithm to schedule the flow to a suitable equivalent path to reduce bandwidth competition under the premise that the arrival distribution of the flow cannot be known in advance, so as to achieve high-bandwidth routing transmission and meet the corresponding delay requirements, and realize the long-term optimization of the load.
[0145] (2) Inter-cluster routing algorithm based on the estimated propagation delay
[0146] This subsection will introduce the first stage of the transmission strategy based on distributed SDN. When the data flow is transmitted between clusters, it is necessary for the local control satellite CH k to determine the forwarding direction of the next-hop cluster. The present invention designs an inter-cluster routing algorithm based on the estimated propagation delay. Based on the topological structure characteristics of the distributed SDN satellite network, in order to minimize the signaling overhead caused by large-scale collection of link states between clusters as much as possible, this algorithm combines the geographical location of the target cluster and the load conditions of adjacent clusters, introduces the estimated remaining propagation delay as the backlog decision metric of the cluster, and adaptively selects the cluster with lighter load and closer to the destination as the forwarding direction of the next-hop cluster.
[0147] However, in the LE0 mega-constellation, due to the large propagation delay between satellites, only considering the queue length as the backlog value may lead to loops and detours in the routing path, and thus result in high delay of the routing path. Therefore, in order to improve the delay performance of inter-cluster routing and route the flow to the cluster with lighter load (such as the cluster over non-hotspot areas such as the ocean and mountains), the estimated remaining propagation delay is introduced as the backlog decision metric for inter-cluster routing to avoid unnecessary long paths and loops in the inter-cluster routing selection process. The definition of the remaining propagation delay will be introduced below.
[0148] Assume that the nodes in the network know their own geographical location coordinates and have obtained the geographical location information of the target node. For flow f i , in order to calculate the current cluster c k to the target cluster cd The estimated remaining propagation delay needs to be shared by four adjacent clusters c j , j∈{1, 2, 3, 4} corresponds to the control satellite CH j The spatial coordinates of And get the target cluster c in advance d Corresponding control satellite CH d The spatial coordinates of Then the estimated residual propagation delay can be calculated as follows:
[0149]
[0150] where ||CH d CH j ||Control satellite CH for target d and adjacent control satellite CH j The spatial distance can be calculated by the formula (1) derived above.
[0151] Define cluster c k The decision indicator is And can be calculated by the following formula:
[0152]
[0153] like Figure 3 As shown, assuming that the current cluster c k There are four adjacent clusters c j , j∈{1, 2, 3, 4} is available for forwarding, and one of the clusters is the previous hop of the inter-cluster routing, then the forwarding direction of the candidate next-hop cluster is c j′ , j′∈{1, 2, 3}. The inter-cluster routing algorithm based on estimated propagation delay calculates the backlog decision index of the adjacent clusters and selects the next-hop cluster with the smallest decision index as the forwarding direction of the inter-cluster routing:
[0154]
[0155] By introducing the estimated propagation delay into the backlog decision index of the adjacent clusters, the algorithm can select the cluster that is closer to the target satellite and has a lighter load as the next-hop cluster's relay direction. This adaptive selection mechanism allows the algorithm to perceive the congestion information of the surrounding clusters and bypass the congested area to achieve active load balancing; at the same time, due to the introduction of the estimated propagation delay, the algorithm can effectively avoid routing detours between clusters and reduce the propagation delay of the transmission flow.
[0156] The pseudo code based on the estimated propagation delay algorithm is shown below.
[0157]
[0158] (3) Markov decision process formation
[0159] Since the optimization objective of the optimization problem proposed in this invention is to maximize the total throughput of the large constellation transmission flow under the condition of satisfying the QoS constraints of the flow, and this optimization problem is not a standard convex optimization problem and cannot be directly solved, therefore, in this subsection, the optimization problem is first transformed into a Markov decision (Markov Decision Process, MDP) problem, where the control satellites in the cluster act as independent agents. The proposed MDP problem consists of four parts: a) a finite state set, b) a finite action set, c) a state transition process, which describes how the current state and action affect the future state, and d) a reward function.
[0160] In the intra-cluster routing algorithm based on the deep Q-network proposed in this section, the control satellites, as independent agents, learn by observing the state within the cluster and allocate high-bandwidth paths for each flow entering the cluster, and satisfy the corresponding delay constraints to achieve long-term balanced optimization of the load. For the flow f k entering the cluster c i , the corresponding control satellite CH k is defined as the state space, action space, and reward function of the agent as follows.
[0161] (1) State space
[0162] Let represent the state space, and the binary tuple s i,k = [M k , δ i,k represents the state observed by the control satellite CH k after the flow f i enters the cluster. M k = [m uv n×n represents the traffic distribution matrix within the cluster c k , where m uv is the traffic in the cluster, and the calculation formula is:
[0163]
[0164] represents the routing request of the flow f i within the cluster c k . Among them, represents the starting satellite of the flow f i within the cluster c k ; l i,k represents the forwarding direction of the flow f i within the cluster c k ; volume i represents the flow f i The size of the data volume.
[0165] (2) Action space
[0166] Control satellite CH k Upon receiving the flow f k entering cluster c i and observing the state s i,k it will execute the action a i,k . CH k The available action space is derived from the set of candidate equivalent paths i,k constructed according to the routing request δ That is Considering that there may be multiple equivalent paths between each source-destination node pair in the cluster, the agent will select m alternative paths between the source-destination node pairs in the cluster according to the hop count as the set of feasible paths.
[0167] (3) Reward function
[0168] Define the reward obtained by the control satellite CH k when executing the action a i,k in the state s i,k as r i,k . Since the goal of the proposed optimization problem is to maximize the system throughput while satisfying the end-to-end delay constraint of each data flow, the reward function is based on the bandwidth size of the allocated path, and the reward r i,k is defined as follows:
[0169]
[0170] where represents the path delay allocated by the control satellite CH k and α k represents the delay limit within the cluster. When the path allocated by the control satellite within the cluster satisfies the delay constraint, the DRL agent will obtain a positive reward, and the greater the bandwidth of the allocated path, the greater the reward; when the path allocated by the control satellite within the cluster does not satisfy the delay constraint, the DRL agent will obtain a penalty, and the smaller the bandwidth of the allocated path, the greater the penalty. The MDP problem is described as (s, a, p, r), and the optimization goal of the constructed optimization problem is to maximize the total throughput of the giant constellation. This optimization problem is then solved by an algorithm based on reinforcement learning.
[0171] (4) Intra-cluster routing algorithm based on DQN:
[0172] (1) Action selection
[0173] When the flow f i enters the cluster c kAfter that, the agent CH k The link information in the cluster will be observed, combined with the current state s i,k , Agent CH k A set of equivalent paths will be constructed As the action space, and based on the soft-∈-greedy strategy from the action space Select action a i,k , that is, greedily choose an action with the highest Q value with a probability of 1-∈, or randomly choose a random action with a probability of ∈:
[0174]
[0175] (2) Exploration strategy
[0176] The present invention adopts a soft-∈-greedy strategy to guide the action selection of the control satellite. In this strategy, a probability that gradually decreases as the number of agent training episodes increases is defined to randomly select actions. This can avoid the agent simply switching between the best action and the random action, which is beneficial to increase the probability of action exploration in the early stage of training, and is more inclined to select the best action in the later stage of training, so as to speed up the convergence speed of agent training and reduce sudden strategy changes, making the learning process more stable. The specific formula is expressed as follows:
[0177]
[0178] Where ∈ is a positive integer, representing the probability of the agent randomly selecting an action. At the beginning of training, a larger ∈ means that the agent has a greater probability of randomly selecting an action in the action space, thereby widely exploring the possible best action options. The purpose of this stage is to allow the agent to discover potential optimal routing paths. As training progresses, the agent gradually deepens its exploration of the action space. Since randomly selecting actions will affect the convergence performance of learning, and the action value function fitted by the neural network is more accurate at this time, the probability of randomly selecting actions needs to be smaller to speed up the convergence process.
[0179] (3) Neural Network:
[0180] The present invention splits the maximization operator in DQN into two independent steps: action selection and action evaluation, and uses the current Q network parameter θ i Select the optimal action, target Q network parameter θ i -Evaluate the selected optimal action. To further stabilize the training process, the DQN network of the present invention uses two neural networks: one is the Evaluation Network, which is used to predict the Q value of the current state; the other is the Target Network, which is used to calculate the target Q value. The parameters of the target network are copied from the evaluation network regularly instead of being updated in real time, which can effectively reduce the fluctuation of the target Q value and improve the stability of training. Two sets of parameters are used to separate the process of action selection and action evaluation, thus reducing the risk of overestimation. Each agent is set with its independent estimation network Q k (s i,k ,a i,k ;θ k ) and the target network where the parameters of the two networks are θ k and The experience replay of each agent is also independently initialized and uses random samples to update the network parameters.
[0181] After executing the action a i,k , the agent CH k will receive the reward r i,k , and observe the next state s i+1,k . Based on the above information, the agent will store the transition process (s i,k , a i,k , r i,k , s i+1,k ) into the experience replay pool and then randomly draw a batch of samples and perform learning.
[0182] At the end of each decision-making, the gradient descent method is used to fit the neural network, and the parameter θ k is updated by minimizing the objective function. i,k The objective function L
[0183] L i,k =(y i,k -Q k (s i,k , a i,k ;θ k )) 2 (17);
[0184] y i,k is the target value, which is the expected maximum value that may be obtained after executing the action, and can be calculated by the following formula:
[0185]
[0186] where γ is the discount factor of future rewards, θ k and They are the parameters of the Q estimation network and the Q target network respectively.
[0187] Then, the parameters of the neural network are updated by gradient descent:
[0188]
[0189] where α is the learning rate, which is used to balance the learning speed and accuracy.
[0190] Based on the above network description, the present invention is trained with a distributed SDN satellite network as the basic model, where each control satellite can be described as a separate agent, and a training network based on deep reinforcement learning can be completed through an intra-cluster routing algorithm based on DQN. The intra-cluster routing algorithm based on DQN is described in Algorithm 2 in the table.
[0191] First, the algorithm initializes the state-action function, exploration probability, and network state. The lines from 1 to 15 are the loop part, starting learning from a random state. In each learning cycle, a new transmission flow is generated for transmission. Then, for each agent, when a flow enters the cluster, it will execute an action according to the observed state, obtain the corresponding reward value, and obtain the next state. Then, the quadruple (s i,k , a i,k , r i,k , s i+1,k ) is stored in the experience replay pool, and a neural network model is trained by randomly obtaining a batch of experience replay data, calculating the loss function and updating the parameters of the training network, and updating the parameters of the target network every once in a while. After multiple learning cycles of fitting, the network parameters θ k of the training network can be obtained, and they are used as the output of the algorithm training.
[0192]
[0193] (5) Overall process of the data flow transmission strategy based on the distributed SDN satellite network:
[0194] Based on the descriptions in the previous sections, the giant constellation transmission strategy based on the distributed SDN network proposed in this paper includes two stages: an inter-cluster routing algorithm based on estimated propagation delay and an intra-cluster routing algorithm based on DQN. The overall framework diagram of the algorithm is as Figure 4 shown. When the flow f i is newly generated in the cluster c k or enters the cluster c kWhen this happens, the transport layer updates the network traffic matrix within the cluster and reports it to the control layer regularly; the control layer runs Algorithm 1, an inter-cluster routing algorithm based on estimated propagation delay, based on the network traffic matrix within the cluster and the load status of adjacent clusters, to obtain the inter-cluster routing and forwarding direction, which is used as one of the inputs for training the network; the training network is trained by Algorithm 2, an intra-cluster routing algorithm based on DQN. The training network constructs a set of equivalent paths for intra-cluster routing as the action space of the agent according to the results of Algorithm 1, an inter-cluster routing algorithm based on estimated propagation delay, and obtains the output of the intra-cluster routing path based on the collected link status; then the routing decision of the control layer is output to the transport layer to complete an intra-cluster routing plan. Then, flow f i will enter the next cluster and repeat the above process until it is transmitted to the target satellite cluster to complete the routing path of the transmission flow. Through this load balancing routing algorithm based on distributed SDN, theoretically, it can solve the problem of excessive signaling overhead caused by using a central node to centrally control the network in a large-scale satellite network, as well as the local optimization problem of using a fully distributed solution.
[0195] The beneficial effects of the data flow transmission scheduling method based on a distributed SDN satellite network of the present invention are as follows:
[0196] (1) A distributed SDN satellite network model is established according to the application background of a large-scale satellite network, and the data flow load balancing routing problem of a large-scale satellite network is constructed under this model;
[0197] (2) For this optimization problem, a data transmission strategy based on a distributed SDN satellite network is proposed, and this strategy can achieve a better load balancing effect and a smaller end-to-end delay.
[0198] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural transformation made under the concept of the present invention by using the content of the specification and drawings of the present invention, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present invention.
Claims
1. A data stream transmission scheduling method based on a distributed SDN satellite network, characterized in that: The method comprises the following steps: Step S10, establishing a distributed SDN satellite network model according to the application background of the large-scale satellite network, and constructing a data flow load balancing routing problem of the large-scale satellite network under the distributed SDN satellite network model; Step S20, to solve the problem of data flow load balancing routing in large-scale satellite networks, a distributed SDN satellite network data transmission strategy is adopted to achieve load balancing effect and reduce end-to-end delay.
2. The data stream transmission scheduling method based on a distributed SDN satellite network according to claim 1 is characterized in that: The step S20, for the data flow load balancing routing problem of large-scale satellite networks, adopts the data transmission strategy of distributed SDN satellite networks to achieve load balancing effect and reduce end-to-end delay. It includes two stages: an inter-cluster routing algorithm based on estimated propagation delay and an intra-cluster routing algorithm based on DQN, wherein: The inter-cluster routing algorithm based on estimated propagation delay combines the geographical location of the target cluster and the load of the adjacent clusters, introduces the estimated residual propagation delay as a decision indicator, and adaptively selects a cluster with a lighter load and closer to the destination as the forwarding direction of the next-hop cluster; In the DQN-based intra-cluster routing algorithm, the optimization problem is transformed into a Markov decision process, the state space and action space of the agent are defined based on the distributed SDN satellite network, and the reward function is constructed according to the optimization goal.
3. The data stream transmission scheduling method based on a distributed SDN satellite network according to claim 1 is characterized in that: In step S10, the steps of establishing a distributed SDN satellite network model according to the application background of a large-scale satellite network include: A LE0 mega-constellation network model is constructed, and a satellite communication model that takes actual channel conditions into consideration is constructed. According to the topological structure of the mega-constellation, the LE0 mega-constellation network is divided into multiple clusters, where satellite nodes always belong to the same cluster to avoid cumbersome management switching problems.
4. The data stream transmission scheduling method based on the distributed SDN satellite network according to claim 3 is characterized in that: In the step of constructing the LE0 giant constellation network model, a Walker-Delta giant constellation is considered, where: The network topology of the mega-constellation is represented as Where v is the set of satellite nodes, ε={e uv |u≠v} represents the ISL set established between satellite nodes; the undirected edge is denoted as uv, and the source-destination edge is denoted as; since the satellites are distributed on the sphere, to represent the coordinates of satellite u, where r represents the altitude of the satellite, θ i and Respectively represent the polar angle and azimuth of the satellite; each satellite will establish 4 ISLs, including 2 inter-orbit ISLs and 2 intra-orbit ISLs, among which the inter-orbit ISL is generally not established between the ascending satellite and the descending satellite; for each ISL, e uv =1 indicates that there is an ISL link connection between satellite u and satellite v, otherwise it means there is no ISL connection; LE0 giant constellation includes N = N p ×M p Satellites, of which N p is the number of orbital planes, M p is the number of satellites in each plane. All orbits have the same inclination α, the satellite orbit height is r, and the spacing along the equator is equal; Mp satellites are evenly distributed in each plane, and the phase deviation between satellites in adjacent planes is Δf=2πF / (N p M p ), where is the phase factor; then the Walker-Delta constellation can be represented by α:N p M p / N p / F means.
5. The data stream transmission scheduling method based on the distributed SDN satellite network according to claim 4 is characterized in that: In the step of constructing the satellite communication model considering the actual channel conditions, it is assumed that the noise is additive Gaussian white noise; ||uv|| represents the distance between satellite u and satellite v, and through the derivation of the spherical coordinate system, it can be obtained: The coordinates of satellite u and satellite v are and The maximum line-of-sight distance refers to the maximum straight-line distance between two points that can directly see each other without physical obstacles. It determines the maximum visible range between satellites. When calculating the maximum line-of-sight distance between two points on the earth's surface, factors such as the curvature of the earth and the height of the observation point and the target point need to be considered. The following simplified formula can be used for estimation: Where d represents the maximum viewing distance, R is the radius of the earth, and h1 and h2 are the heights from the two observation points to the ground respectively; Hypothesis I * (u, v) is the maximum line-of-sight distance between satellite u and satellite v. When ||uv||>I * (u, v) cannot establish ISL; then FSPL can be expressed as: Where f is the frequency of light and c is the speed of light; Assuming the wireless channel is symmetric, we can define the signal-to-noise ratio of a satellite pair as: Where P t is the transmission power, G t is the transmitting antenna gain, G t is the receiving antenna gain, k B is the Boltzmann constant, B is the channel bandwidth in Hertz, and T is the thermal noise in Kelvin; Assuming that satellite u chooses to communicate with v at the maximum data rate in an interference-free environment, the link capacity C uv It can be calculated as: C uv =Blog(1+SNR uv ) (5).
6. The data stream transmission scheduling method based on the distributed SDN satellite network according to claim 5 is characterized in that: In the step of dividing the LE0 giant constellation network into multiple clusters according to the topological structure of the giant constellation, wherein the satellite nodes always belong to the same cluster to avoid cumbersome management switching problems, the cluster is divided into a rectangular area according to adjacent orbital satellites, which contains n×m satellites, 2≤n≤N p , 2≤m≤M p ; These areas are relatively fixed to reduce the update frequency and the routing overhead of each update; the geographical center satellite in the cluster is the control satellite and is equipped with an SDN network controller. The satellites in the cluster will regularly report their load information to the SDN network controller, and the SDN controller will make routing decisions for routing events in the cluster based on the collected link information; at the same time, the control satellites will regularly exchange load information in adjacent clusters through ISL; The cluster set of the entire satellite network is represented by C = {c k }, 1≤k≤K, where K is the number of clusters, and the central control satellite corresponding to each cluster is represented by CH k It is said that it is equipped with an SDN network controller, which can monitor the link status information of each satellite in the cluster and plan the routing of the network traffic within the cluster; other satellites are data transmission satellites, which are responsible for transmitting data traffic in the cluster and regularly reporting the link status to the control satellite.
7. The data stream transmission scheduling method based on a distributed SDN satellite network according to claim 6 is characterized in that: In step S10, the steps of constructing a data flow load balancing routing problem of a large-scale satellite network under the distributed SDN satellite network model include: The set of flows in the network is defined as where f i Represents the traffic demand in the network, Represents the flow f i The starting satellite of Represents the flow f i The terminal satellite, volume i Represents the flow f i The amount of data, I is the number of flows in the network; for each There will be multiple feasible paths for traffic transmission in the satellite network; Represents the flow f i A set of available traffic transmission paths, where Represents an alternative path: Let the binary variable Represents the flow f i The path p i Does it include ISL e uv , Represents e uv In the path p i In, otherwise By e uv The number of flows can be To calculate; Since there may be multiple data flows transmitted simultaneously in an ISL, it is assumed here that the traffic on the same link shares the link capacity fairly, and the flow f can be calculated i In cluster c k The bandwidth within is: where p i,k For flow i In cluster c k Path within; And we can calculate the flow f i The bandwidth is: Flow i The delay can be calculated as: The first term in (8) represents the transmission time of the flow, and the second term represents the propagation delay of the multi-hop ISL; Based on the above model, the routing problem in the giant constellation is formulated as a nonlinear programming problem. The control satellite of each cluster equipped with an SDN controller needs to formulate the best routing strategy for the flow entering the cluster to maximize the system throughput while satisfying the end-to-end delay constraints of each flow. The optimization problem can be expressed as follows: The constraint C1 is the flow f i The bandwidth of each cluster SDN controller is allocated the minimum bandwidth. Constraint C2 is the flow protection constraint, which ensures the flow conservation of each terminal satellite; Constraint C3 sets the delay constraint α and hopes that each flow f i Delay No more than α i , α i For flow i QoS constraints, and flow f i The shortest path delay is proportional to a certain value.
8. The data stream transmission scheduling method based on a distributed SDN satellite network according to claim 7 is characterized in that: The inter-cluster routing algorithm based on estimated propagation delay combines the geographical location of the target cluster and the load of the adjacent cluster, introduces the estimated residual propagation delay as a decision indicator, and adaptively selects a cluster with a lighter load and closer to the destination as the forwarding direction of the next-hop cluster, including: Assuming that all nodes in the network know their own geographic coordinates and have obtained the geographic location information of the target node, for flow f i , in order to calculate the current cluster c k To the target cluster c d The estimated remaining propagation delay needs to be shared by four adjacent clusters c j , j∈{1, 2, 3, 4} corresponds to the control satellite CH j The spatial coordinates of And get the target cluster c in advance d Corresponding control satellite CH d The spatial coordinates of Then the estimated residual propagation delay can be calculated as follows: where ||CH d CH j ||Control satellite CH for target d and adjacent control satellite CH j The spatial distance can be calculated by formula (1); Define cluster c k The decision indicator is And can be calculated by the following formula: Assume that the current cluster c k There are four adjacent clusters c j , j∈{1,2,3,4} is available for forwarding, and one of the clusters is the previous hop of the inter-cluster routing, then the forwarding direction of the candidate next-hop cluster is c j′ , j′∈{1,2,3}; The inter-cluster routing algorithm based on estimated propagation delay calculates the backlog decision index of the adjacent clusters and selects the next-hop cluster with the smallest decision index as the forwarding direction of the inter-cluster routing:
9. The data stream transmission scheduling method based on a distributed SDN satellite network according to claim 8 is characterized in that: In the DQN-based cluster routing algorithm, the optimization problem is transformed into a Markov decision process, the state space and action space of the agent are defined based on the distributed SDN satellite network, and the reward function is constructed according to the optimization goal. k The flow i , define the corresponding control satellite CH k The state space, action space, and reward function of the agent are as follows: (1) State space: Let S denote the state space, and the tuple s i,k =[M k , δ i,k ] indicates control satellite CH k In the flow i The state observed after entering the cluster; M k =[m uv ] n×n Represents cluster c k The traffic distribution matrix in uv is the traffic in the cluster, and the calculation formula is: Represents the flow f i In cluster c k Routing requests within Represents the flow f i In cluster c k The starting satellite within i,k Representative representative flow f i In cluster c k Internal forwarding direction; volume i Represents the flow f i The amount of data; (2) Action Space: Control satellite CH k Upon receiving the entry into cluster c k The flow i And observe the state s i,k When , action a will be executed i,k ; CH k The space of available actions Derived from the routing request δ i,k The set of candidate equivalent paths constructed Right now Considering that there may be multiple equivalent paths between each source-destination node pair in the cluster, the intelligent body will select m optional paths between the source-destination node pair from the cluster according to the number of hops as the feasible path set; (3) Reward function: Define control satellite CH k In status i,k Next, perform action a i,k The reward obtained is r i,k ; Since the goal of the proposed optimization problem is to maximize the system throughput while satisfying the end-to-end delay constraint of each data flow, the reward function is based on the bandwidth size of the allocated path, and the reward r is defined as i,k as follows: in Stands for control satellite CH k The assigned path delay, α k represents the delay constraint within the cluster; when the path allocated by the control satellite within the cluster meets the delay constraint, the DRL agent will receive a positive reward, and the larger the allocated path bandwidth, the greater the reward; when the path allocated by the control satellite within the cluster does not meet the delay constraint, the DRL agent will receive a penalty, and the smaller the allocated path bandwidth, the greater the penalty; the Markov decision problem is described as (s, a, p, r), and the optimization goal of the constructed optimization problem is to maximize the total throughput of the giant constellation. This optimization problem is then solved by an algorithm based on reinforcement learning.
10. The data stream transmission scheduling method based on the distributed SDN satellite network according to claim 9 is characterized in that: The DQN-based intra-cluster routing algorithm includes: (1) Action selection: When the flow i Enter cluster c k After that, the agent CH k The link information in the cluster will be observed, combined with the current state s i,k , Agent CH k A set of equivalent paths will be constructed As the action space, and based on the soft-∈-greedy strategy from the action space Select action a i,k , that is, greedily choose an action with the highest Q value with a probability of 1-∈, or randomly choose a random action with a probability of ∈: (2) Exploration strategy: The soft-∈-greedy strategy is used to guide the action selection of the control satellite. In this strategy, a probability that gradually decreases as the number of training episodes of the agent increases is defined to randomly select actions. The specific formula is expressed as follows: Among them, ∈ is a positive integer, representing the probability of the agent randomly selecting an action; at the beginning of training, a larger ∈ means that the agent has a greater probability of randomly selecting an action in the action space, thereby widely exploring possible optimal action options. (3) Neural Network: The maximization operator in the DQN network is split into two independent steps: action selection and action evaluation, and the current Q network parameter θ i Select the optimal action, target Q network parameter θ i - To further stabilize the training process, the DQN network uses two neural networks: one is the evaluation network, which is used to predict the Q value of the current state; the other is the target network, which is used to calculate the target Q value; the parameters of the target network are regularly copied from the evaluation network instead of being updated in real time, which can effectively reduce the fluctuation of the target Q value and improve the stability of training. Two sets of parameters are used to separate the process of action selection and action evaluation, thereby reducing the risk of over-estimation; each agent is set with its own independent estimation network Q k (s i,k , a i,k θ k ) and the target network The parameters of the two networks are θ k as well as The experience replay for each agent is also initialized independently and uses random samples to update the network parameters. In performing action a i,k After that, the agent CH k Will be rewarded i,k , and observe the next state s i+1,k ; Based on the above information, the agent will transform the process (s i,k , a i,k , r i,k ,s i+1,k ) is stored in the experience replay pool R, and then a batch of samples are randomly selected and learned; At the end of each decision, the neural network is fitted using the gradient descent method and the parameters θ are updated by minimizing the objective function k ; Objective function L i,k Given by: L i,k =(y i,k -Q k (s i,k ,a i,k ;θ k )) 2 (17); y i,k is the target value, is the expected maximum value that may be obtained after executing the action, and can be calculated by the following formula: Where γ is the discount factor for future rewards, θ k and are the parameters of the Q estimation network and Q target network respectively; Then update the parameters of the neural network by gradient descent: Where α is the learning rate, which is used to balance learning speed and accuracy.
Citation Information
Cited By
Dynamic routing and flow balancing method and system for regionalized network topology
CN120915372A
A dynamic routing and traffic balancing method and system for a regionalized network topology
CN120915372B
Load balancing routing method and system based on dynamic region segmentation and medium
CN122027006A
A large-scale constellation load balancing routing method, system, and medium based on dynamic region segmentation.
CN122027006B