Satellite Internet resource scheduling method based on deep deterministic policy gradient algorithm
By applying deep deterministic policy gradient algorithms and TEG shunt mechanisms in satellite networks, combined with deep reinforcement learning to optimize SFC traffic engineering, the challenges of resource scheduling and traffic allocation in satellite networks are solved, and more efficient resource utilization and fairer rate allocation are achieved.
Patent Information
- Application Number
- CN202411089778.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-09
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-08-09
AI Technical Summary
In satellite networks, how to effectively schedule network resources in a highly dynamic topological environment to achieve fairer rate allocation and more efficient traffic engineering, especially in predefined VNF sequential routing and multipath transmission in SFC.
A satellite Internet resource scheduling method based on deep deterministic strategy gradient algorithm is proposed. By building a model that maximizes the minimum flow rate, combining the time evolution graph (TEG) shunt mechanism and deep reinforcement learning (DDPG) algorithm, SFC traffic engineering and resource allocation are optimized.
It realizes more efficient resource scheduling and traffic allocation in highly dynamic satellite networks, improves bandwidth resource utilization, significantly improves the minimum flow rate of the network, and shows excellent performance in multi-path and multi-SFC scenarios.
Smart Images

Figure CN119094449B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a satellite Internet resource scheduling method based on a deep deterministic policy gradient algorithm. Background Art
[0002] With the advancement of space technology and the expansion of commercial satellite networks, traditional satellite systems can no longer meet the rapidly growing user needs, which has prompted researchers to turn to the development of new satellite network systems. As an emerging network architecture, SDSN provides efficient, centralized control and flexible management capabilities for satellite networks by separating the control and data forwarding functions of satellite networks, heralding the future development direction of satellite networks. In addition, NFV (network function virtualization) technology realizes the softwareization of network functions, thereby reducing dependence on dedicated hardware, significantly reducing capital and operating costs, and improving network scalability. By connecting a series of VNFs in series in a determined order, the SFC formed further enhances the flexibility and efficiency of network services.
[0003] In satellite network environments, due to its highly dynamic network topology, service interruption and continuity issues have become increasingly prominent, which puts higher demands on QoS. In response to this challenge, dynamically managing the routing of SFC through the SDSN controller to achieve flexible traffic diversion has become an effective solution. This not only ensures effective collaboration between the SDN controller and the NFV manager, but also efficiently allocates and manages the resources required for SFC and optimizes overall network performance.
[0004] In the ground network, in order to achieve finer-grained control of traffic segmentation, a decomposition method using ADMM and weighted minimum mean square error method is proposed to maximize the minimum rate of all transmission flows. A scalable distributed method based on ADMM is proposed to minimize the cost of routing and VNF (virtual network function) deployment, and encourage processing traffic in local subnets, ensure the privacy of administrators' sensitive information and reduce implementation costs. Two deep reinforcement learning techniques for traffic engineering problems are proposed, namely traffic engineering-aware exploration and priority experience replay based on AC framework. Compared with several baseline methods, this method significantly reduces end-to-end delay, continuously improves total utility, and is robust to network changes. An improved deep reinforcement learning method is also proposed to control the split ratio of traffic in multiple paths, which can effectively cope with the dynamic changes of traffic load.
[0005] In addition, considering the importance of traffic engineering in optimizing network resource utilization, some people have studied the traffic diversion problem in the integrated space-ground network and proposed a line generation method, which effectively solves the problem of how to maximize the total amount of data received by the ground data processing center under multiple resource constraints and traffic restrictions. This method achieves efficient solution of the optimization solution by iterating the main problem and sub-problems, providing an acceleration method for solving large-scale systems.
[0006] Although research has made progress, when considering the predefined VNF sequential routing in SFC and the highly dynamic topology of satellite networks, how to ensure the effective scheduling of network resources and achieve fairer rate distribution still requires in-depth exploration of the SFC deployment based on traffic engineering in satellite networks. Summary of the invention
[0007] In view of the above problems existing in the prior art, the purpose of the present invention is to propose a TEG diversion mechanism based on TEG routing, construct a maximized minimum flow model through the TEG flow conservation constraint of traffic segmentation, and propose a SFC traffic engineering method based on DDPG to solve the problem.
[0008] In order to solve the above technical problems, the present invention adopts the following technical solution: a satellite Internet resource scheduling method based on a deep deterministic policy gradient algorithm, comprising the following steps:
[0009] S1: Modeling a satellite network SDSN with multiple satellite nodes and multiple service function chains SFC. Let SDSN be represented by a directed graph G(V,E), where V is the set of nodes, E is the set of links, and E includes the physical link set E vv and the storage link set E between satellites in adjacent time slots v ; The time evolution graph TEG divides the total time into T time slots, the length of each time slot is η, let t∈T represent the index of the time slot, in TEG, (i t ,j t )∈E vv Represents the physical link between two different satellites, with (i t ,i t +1 )∈E v Represents the storage link between consecutive time slots of the same satellite.
[0010] Let K = {1, 2, ..., k, ...} represent the SFC request set in SDSN. For k ∈ K, represents an ordered set of VNFs, Represents the nth VNF of the kth SFC.
[0011] The last VNF is represented as m represents the number of VNFs; and They represent the source node and destination node of the kth SFC respectively.
[0012] S2: Assuming that in the same time slot, It can only be deployed on one satellite node. The VNF deployment constraints are modeled as:
[0013]
[0014] where i t represents the replica of satellite node i in the tth time slot, Deploy indicator variables for VNFs, Indicates The location where the VNF is deployed is node i t , otherwise it means The VNF is deployed at a location other than node i t superior.
[0015] Define y k is the flow rate of the kth SFC, and defines the variable is the kth SFC on link (i t ,j d )∈E, where Indicates that the kth SFC is in link (i t ,j t )∈E vv The flow rate, Indicates that the kth SFC is in link (i t ,i t+1 )∈E v The flow rate is expressed as:
[0016]
[0017] in Indicates that the kth SFC passes through the After the VNF, the link (i t ,j d )∈E, Indicates an auxiliary VNF located on the source node, identifying the SFC flow that has not been processed by any VNF.
[0018] S3: The computational resource constraints of satellite nodes are modeled as:
[0019]
[0020] in Indicates that in the link (j d,i t ) Traffic rate after VNF processing, Represents node i t The computing resource capacity; Indicates the computing resource requirements per unit data flow rate.
[0021] In addition, the bandwidth resource capacity of the link is expressed as:
[0022]
[0023] in Indicates link (i t ,j t ) bandwidth resource capacity, Indicates that the kth SFC is in link (i t ,j t )∈E vv flow rate.
[0024] S4: Construct flow conservation constraints for the TEG flow diversion model in different situations.
[0025] S5: Set the minimum flow rate y that maximizes all SFCs min To optimize the goal:
[0026]
[0027] Non-negativity constraints are imposed on the flow rate of the SFC and the flow rate of each stage of each SFC and are modeled as:
[0028] y k ≥0
[0029]
[0030] S6: The controller of the SDSN network is used as an intelligent agent to centrally control the SFC deployment, and the optimization problem is modeled as an MDP model suitable for deep reinforcement learning. The state space, action space and reward function are defined.
[0031] S7: Solve the MDP model constructed in S6 based on the DDPG model to obtain the optimal SFC deployment.
[0032] The DDPG model includes an Actor network, a Critic network, and an experience replay pool. The parameters of the current network π(s) of the Actor network and its target network π′(s) are θ π and θ π′ , the parameters of the current network Q(s,a) of the Critic network and its target network Q′(s,a) are θ Q and θQ′ The Actor network is responsible for action screening and strategy formulation, and updates the Actor network and Critic network parameters according to the strategy gradient ascent and the gradient descent of the loss function respectively. The Critic network is responsible for evaluating the generated strategy, and the experience replay pool is used to store the state s t , action a t , r t and the next state s t+1 The experience tuple formed.
[0033] Initialize the parameters of the Actor network and Critic network and the experience replay pool, and update the parameters of the Actor network and Critic network in each iteration.
[0034] Input s to the DDPG model t , get a from the Actor network t , a t Apply to t In the process, the SDSN controller performs SFC deployment, and then obtains r t And update t+1 , will (s t ,a t ,r t ,s t+1 ) is saved as an experience tuple in the experience replay pool. When the experience replay pool is full, the earliest experience tuple will be replaced by a new experience tuple. A small batch of experience tuples are randomly selected from the experience replay pool for training. Under the goal of maximizing the expected cumulative discounted reward, the parameters of the Actor network and the Critic network are updated through gradient solution. The expected cumulative discounted reward refers to the expected value of the sum of all possible rewards in the future starting from the current state during the strategy execution process. The state-action value function Q π (s(t), a(t)) is used to estimate the expected cumulative discounted reward for a given state and action, which is the core goal of the Critic network optimization. By maximizing the state-action value function, the model not only focuses on the current reward, but also on the future long-term benefits, so that the strategy can achieve the overall optimality.
[0035]
[0036] The parameter updates of the Actor network and the Critic network are used to guide the SDSN controller to better perform SFC deployment in the next iteration until the number of training rounds reaches the set maximum value and the expected cumulative discounted reward of the optimization task is maximized, indicating that the training is completed and the optimal SFC deployment is obtained.
[0037] Furthermore, in S4, the flow conservation constraints of the TEG flow splitting model are divided into four cases:
[0038] Source Node The sum of the outflow rates should be y k :
[0039]
[0040] in, Indicates the physical link (i t ,j t ) outbound rate that is not processed by VNF, Indicates that in the storage link (i t ,i t+1 ) is the outbound rate that is not processed by the VNF.
[0041] Destination Node The flow conservation constraint is expressed as:
[0042]
[0043] in, Indicates the physical link Pass The inflow rate after treatment, Indicates storage link Pass Inflow rate after treatment.
[0044] Deploy Node i t The inflow rate and outflow rate are both y k :
[0045]
[0046] in, Indicates that in the physical link (j t ,i t ) The inflow rate after treatment, Indicates that in the storage link (i t-1 ,i t ) treated inflow rate; Indicates that in the physical link (i t , j t ) The outflow rate after treatment, Indicates that in the storage link (i t ,i t+1 ) The outflow rate after treatment, Indicates the VNF deployment indicator variable.
[0047] Forwarding node it The inflow rate is equal to the outflow rate:
[0048]
[0049] in, Indicates that in the physical link (j t ,i t ) The inflow rate after treatment, Indicates that in the storage link (i t-1 ,i t ) Inflow rate after treatment.
[0050] Furthermore, in S6, the state space, action space and reward function are respectively:
[0051] State Space:
[0052] s t ={y(t),E(t),C(t),B(t)}
[0053] Among them, y(t) represents the flow rate status requested by SFC, E(t) represents the network topology status of SDSN, C(t) and B(t) represent the remaining computing resources of the node and the remaining bandwidth resources of the physical link respectively.
[0054] Action Space:
[0055] a t ={X(t),Y(t)}
[0056] in, Represents a collection of VNF deployment actions. express With satellite t The relationship between Represents a set of traffic distribution actions, Indicates that the kth SFC passes Then assigned to the physical link (i t , j t ) flow rate, Indicates that the kth SFC passes Then assigned to the storage link (i t ,i t+1 ) flow rate.
[0057] Reward function:
[0058]
[0059] where y k (t) represents the traffic rate of the kth SFC in the tth time slot.
[0060]
[0061] Furthermore, in S7, the loss function L(θ Q )for:
[0062] L(θ Q )=E[(y t -Q(s t ,a t |θ Q )) 2 ]
[0063] y t =r t +γQ′(s t+1 ,π′(s t+1 |θ π′ )|θ Q′ )
[0064] Among them, y t is the target value, π′(s t+1 |θ π′ ) indicates that according to θ π′ and t+1 To decide the best action to take, Q′(s t+1 ,π′(s t+1 |θ π′ )|θ Q′ ) is the estimated value of the cumulative discounted reward expected from the actions taken by Q′(s,a) on future states, and γ represents the discount factor.
[0065] Furthermore, in S7, the Actor network parameters are updated in the following manner:
[0066]
[0067] Among them, α represents the learning rate of the Actor network, H represents the number of mini-batch experience tuples, s h They represent solving the gradient of the Q function with respect to a and solving the gradient of the policy function π with respect to θ. π The gradient of , the state of the h-th mini-batch sample.
[0068] The critic network parameter update is expressed as:
[0069]
[0070] Among them, β represents the learning rate of the Critic network, y h , a hThey represent the target values of the h-th mini-batch samples, respectively, and solve the Q function with respect to θ Q The gradient of , the action of the h-th mini-batch sample.
[0071] Furthermore, in S7, the target networks of the Actor network and the Critic network are updated as follows:
[0072] θ π′ =τθ π +(1-τ)θ π′
[0073] θ Q′ =τθ Q +(1-τ)θ Q′
[0074] Where τ represents the soft update factor.
[0075] Compared with the prior art, the present invention has at least the following advantages:
[0076] 1. When satellite networks allow multi-path transmission and TEG modeling allows traffic temporary storage operations to be considered as a feasible link, a satellite SFC deployment strategy suitable for highly dynamic conditions is designed using traffic engineering and deep reinforcement learning. By considering the temporary storage operation as a storage link, TEG diversion can be both simultaneous and cross-time slot diversion. A TEG diversion mechanism is then proposed, which allows the SFC flow to be split into multiple sub-flows that traverse different paths, thereby improving the utilization of bandwidth resources. The TEG flow conservation constraints in the diversion mode are improved, and a SFC traffic engineering model that maximizes the minimum flow rate is constructed by combining resource constraints and flow rate non-negativity constraints. Considering that traditional algorithms cannot cope with high-dimensional and complex SFC deployment decisions, especially in larger-scale networks, the optimal deployment decision cannot be obtained. The optimization solution based on heuristics or meta-heuristics relies on prior knowledge and is prone to fall into local optimal solutions. Therefore, in the face of the ever-changing LEO satellite topology and environment, we use deep reinforcement learning to solve it.
[0077] 2. The simulation results show that in the deployment scenarios of multiple SFCs and different resource capacity settings, compared with TEG routing and static routing, the TEG diversion mechanism can effectively improve the minimum flow rate. Compared with the baseline algorithm, the proposed algorithm has better performance and convergence speed than the baseline algorithm. As the number of SFCs increases, compared with DPG-SFC-TE and DQN-SFC-TE, the algorithm can still effectively utilize the bandwidth resources in the network. Compared with the baseline strategy, this strategy significantly improves the minimum flow rate of the network, indicating the effectiveness of the SFC deployment strategy based on traffic engineering in improving the overall performance of the satellite network system. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 It is the SDSN network model.
[0079] Figure 2 It is a SFC traffic engineering algorithm framework based on DDPG.
[0080] Figure 3 is the minimum flow rate under different routing mechanisms.
[0081] Figure 4 is the minimum flow rate under different bandwidth resource capacities.
[0082] Figure 5 is the minimum flow rate under different computing resource capacities. DETAILED DESCRIPTION
[0083] The present invention is described in further detail below.
[0084] In the present invention, the main optimization goal is to improve the minimum flow rate of the entire system. In view of this goal, the multi-agent architecture is not an ideal choice because it mainly aims to improve the performance of a single agent and does not directly meet the requirements of improving the overall minimum standard. In addition, although algorithms that combine continuous and discrete action spaces, such as parameterized deep Q networks, show potential, they still face the limitation of poor performance in independent practice. Regarding the traditional DQN algorithm, since it is mainly suitable for dealing with discrete problems with low action space dimensions, it is not suitable for continuous action domains. This is because DQN relies on finding specific actions that can maximize the action value function in each step, and in the continuous action space, this search process needs to be achieved through iterative optimization. Based on these factors, the SFC traffic engineering algorithm based on DDPG is used to solve the problem.
[0085] A satellite Internet resource scheduling method based on a deep deterministic policy gradient algorithm comprises the following steps:
[0086] Firstly, the time evolution graph (TEG) is used to represent the network resources of low earth orbit (LEO) satellites. Then, the traffic engineering model of service function chain (SFC) is constructed. Finally, deep reinforcement learning method is used to solve it.
[0087] The overall implementation process of the strategy: build a service function chain SFC diversion model, take TEG routing as the benchmark, and optimize the mechanism and solve the problem on this basis. Specifically, a TEG-based on-board SFC diversion mechanism is proposed, and the multi-time slot related SFC minimum flow rate maximization problem is established, and the deep deterministic policy gradient DDPG is used to solve the problem.
[0088] S1: Modeling a satellite network SDSN with multiple satellite nodes and multiple service function chains SFC. Let SDSN be represented by a directed graph G(V,E), where V is the set of nodes, E is the set of links, and E includes the physical link set E vv and the storage link set E between satellites in adjacent time slots v ; The time evolution graph TEG divides the total time into T time slots, the length of each time slot is η, let t∈T represent the index of the time slot, in TEG, (i t ,j t )∈E vv Represents the physical link between two different satellites, with (i t ,i t +1 )∈E v represents the storage link between consecutive time slots of the same satellite; assuming that the storage resources of LEO WeChat are infinite, but the computing resources are limited, use Represents node i t ∈V computing resource capacity, when two satellites are in contact with each other, Represents the physical links between different satellites (i t ,j t ) bandwidth resource capacity, if satellite node i t and j t If there is no physical link between
[0089] Let K = {1, 2, ..., k, ...} represent the SFC request set in SDSN. For k ∈ K, represents an ordered set of VNFs, represents the nth VNF of the kth SFC;
[0090] The last VNF is represented as m represents the number of VNFs; and They represent the source node and destination node of the kth SFC respectively.
[0091] S2: Assuming that in the same time slot, It can only be deployed on one satellite node. The VNF deployment constraints are modeled as:
[0092]
[0093] where i t represents the replica of satellite node i in the tth time slot, Deploy indicator variables for VNFs, Indicates The location where the VNF is deployed is node i t, otherwise it means The VNF is deployed at a location other than node i t superior;
[0094] Define y k is the flow rate of the kth SFC. Since the flow in TEG is allowed to split into multiple sub-flows traversing different paths, define the variable is the kth SFC on link (i t ,j d )∈E, where Indicates that the kth SFC is in link (i t ,j t )∈E vv The flow rate, Indicates that the kth SFC is in link (i t ,i t+1 )∈E v The flow rate is expressed as:
[0095]
[0096] in Indicates that the kth SFC passes through the After the VNF, the link (i t ,j d )∈E, Indicates an auxiliary VNF located on the source node, identifying the SFC flow that has not been processed by any VNF.
[0097] S3: The computational resource constraints of satellite nodes are modeled as:
[0098]
[0099] in Indicates that in the link (j d ,i t ) Traffic rate after VNF processing, Represents node i t The computing resource capacity; Indicates the computing resource requirements per unit data flow rate;
[0100] In addition, the bandwidth resource capacity of the link is expressed as:
[0101]
[0102] in Indicates link (i t ,j t ) bandwidth resource capacity, Indicates that the kth SFC is in link (i t ,j t )∈E vv flow rate.
[0103] S4: Construct flow conservation constraints for the TEG flow diversion model in different situations.
[0104] S5: To ensure that each SFC in SDSN can at least obtain this rate level, set the minimum flow rate y that maximizes all SFCs min To optimize the goal:
[0105]
[0106] In addition, considering the non-negativity of flow rate, non-negativity constraints are imposed on the flow rate of SFC and the flow rate of each stage of each SFC and are modeled as:
[0107] y k ≥0
[0108]
[0109] S6: The controller of the SDSN network is used as an intelligent agent to centrally control the SFC deployment, and the optimization problem is modeled as an MDP model suitable for deep reinforcement learning. The state space, action space and reward function are defined.
[0110] S7: Solve the MDP model constructed in S6 based on the DDPG model to obtain the optimal SFC deployment.
[0111] The DDPG model includes an Actor network, a Critic network, and an experience replay pool. The parameters of the current network π(s) of the Actor network and its target network π′(s) are θ π and θ π′ , the parameters of the current network Q(s,a) of the Critic network and its target network Q′(s,a) are θ Q and θ Q′ The Actor network is responsible for action screening and strategy formulation, and updates the Actor network and Critic network parameters according to the strategy gradient ascent and the gradient descent of the loss function respectively. The Critic network is responsible for evaluating the generated strategy, and the experience replay pool is used to store the state s t , action a t , r t and the next state s t+1 The experience tuple formed.
[0112] Initialize the parameters of the Actor network and Critic network and the experience replay pool, use the DDPG model to train the SFC deployment strategy, and update the parameters of the Actor network and Critic network in each iteration;
[0113] Input s to the DDPG model t , get a from the Actor network t , a t Apply to t In the process, the SDSN controller performs SFC deployment, and then obtains r t And update t+1 , will (s t ,a t ,r t ,s t+1 ) is saved as an experience tuple in the experience replay pool. When the experience replay pool is full, the earliest experience tuple will be replaced by a new experience tuple. A small batch of experience tuples are randomly selected from the experience replay pool for training. Under the goal of maximizing the expected cumulative discounted reward, the parameters of the Actor network and the Critic network are updated through gradient solution. The expected cumulative discounted reward refers to the expected value of the sum of all possible rewards in the future starting from the current state during the strategy execution process. The state-action value function Q π (s(t), a(t)) is used to estimate the expected cumulative discounted reward for a given state and action, which is the core goal of the Critic network optimization. By maximizing the state-action value function, the model not only focuses on the current reward, but also on the future long-term benefits, so that the strategy can achieve the overall optimality.
[0114]
[0115] The parameter updates of the Actor network and the Critic network are used to guide the SDSN controller to better perform SFC deployment in the next iteration until the number of training rounds reaches the set maximum value and the expected cumulative discounted reward of the optimization task is maximized, indicating that the training is completed and the optimal SFC deployment is obtained.
[0116] Specifically, the flow conservation constraints of the TEG diversion model are divided into four cases:
[0117] Source Node The sum of the outflow rates should be y k :
[0118]
[0119] in, Indicates the physical link (i t ,j t) outbound rate that is not processed by VNF, Indicates that in the storage link (i t ,i t+1 ) is the outbound rate that is not processed by the VNF.
[0120] Destination Node The flow conservation constraint is expressed as:
[0121]
[0122] in, Indicates the physical link Pass The inflow rate after treatment, Indicates storage link Pass Inflow rate after treatment.
[0123] Deploy Node i t The inflow rate and outflow rate are both y k :
[0124]
[0125] in, Indicates that in the physical link (j t ,i t ) The inflow rate after treatment, Indicates that in the storage link (i t-1 ,i t ) treated inflow rate; Indicates that in the physical link (i t , j t ) The outflow rate after treatment, Indicates that in the storage link (i t , j t+1 ) The outflow rate after treatment, Represents the VNF deployment indicator variable, Indicates The location where the VNF is deployed is node i t , otherwise it means The VNF is deployed at a location other than node i t superior.
[0126] Forwarding node i t The inflow rate is equal to the outflow rate:
[0127]
[0128] in, Indicates that in the physical link (j t ,i t ) The inflow rate after treatment, Indicates that in the storage link (i t-1 ,i t ) Inflow rate after treatment.
[0129] Specifically, in S6, the state space, action space and reward function are respectively:
[0130] State Space:
[0131] s t ={y(t),E(t),C(t),B(t)}
[0132] Among them, y(t) represents the flow rate status requested by SFC, E(t) represents the network topology status of SDSN, C(t) and B(t) represent the remaining computing resources of the node and the remaining bandwidth resources of the physical link respectively.
[0133] Action Space:
[0134] a t ={X(t),Y(t)}
[0135] in, Represents a collection of VNF deployment actions. express With satellite t The relationship between Represents a set of traffic distribution actions, Indicates that the kth SFC passes Then assigned to the physical link (i t , j t ) flow rate, Indicates that the kth SFC passes Then assigned to the storage link (i t ,i t+1 ) flow rate.
[0136] Reward function:
[0137]
[0138] where y k (t) represents the traffic rate of the kth SFC in the tth time slot.
[0139] In a given state, the agent uses the strategy π, i.e., a t =π(s t) to determine the action, in order to evaluate the effect of the current strategy, given a t and t The state-action value function Q π (s(t), a(t)):
[0140]
[0141] Among them, E[x] represents the expected value of x, Q π (s(t), a(t)) represents the expected value of the cumulative reward at time t, which includes the immediate reward r(t) and the future discounted reward γQ π (s(t+1), a(t+1)), γ represents the discount factor, and its value range is 0≤γ≤1. The larger the γ, the greater the impact of future rewards on the current ones, which means that this strategy focuses more on long-term benefits.
[0142] The current optimal strategy is expressed as:
[0143]
[0144] Specifically, in S7, the loss function L(θ Q )for:
[0145] L(θ Q )=E|(y t -Q(s t , a t |θ Q )) 2 ]
[0146] y t =r t +yQ′(s t+1 ,π′(s t+1 |θ π′ )|θ Q′ )
[0147] Among them, y t is the target value, π′(s t+1 |θ π′ ) indicates that according to θ π′ and t+1 To decide the best action to take, Q′(s t+1 ,π′(s t+1 |θ π′ )|θ Q′ ) is the estimated value of the cumulative discounted reward expected from the actions taken by Q′(s,a) on future states, and γ represents the discount factor.
[0148] Specifically, in S7, the method of updating the Actor network parameters is as follows:
[0149]
[0150] Among them, α represents the learning rate of the Actor network, H represents the number of mini-batch experience tuples, s h They represent solving the gradient of the Q function with respect to a and solving the gradient of the policy function π with respect to θ. π The gradient of , the state of the h-th mini-batch sample.
[0151] The critic network parameter update is expressed as:
[0152]
[0153] Among them, β represents the learning rate of the Critic network, y h , a h They represent the target values of the h-th mini-batch samples, respectively, and solve the Q function with respect to θ Q The gradient of , the action of the h-th mini-batch sample.
[0154] Specifically, in S7, the target networks of the Actor network and the Critic network are updated as follows: This slower updating method is used instead of directly copying parameters to further increase the stability of the learning process.
[0155] θ π′ =τθ π +(1-τ)θ π′
[0156] θ Q′ =τθ Q +(1-τ)θ Q′
[0157] Where τ represents the soft update factor.
[0158] The main objective of the present invention is to evaluate the performance of the SFC deployment strategy based on traffic engineering in the SDSN network under different system parameter settings through comparative simulation. For the topology of the satellite network, the Walker satellite network generated by the STK software is also simulated. In this network, 36 LEO satellites are evenly distributed on 6 orbital planes, each orbital plane contains 6 satellites, the orbital inclination is 90 degrees, and the orbital altitude is 1000 kilometers. In the simulation, the time slot length is set to 30 seconds and the number of time slots is set to 60. First, this paper compares the performance of the proposed method with two benchmark methods, namely, the service function chain traffic engineering method based on deterministic policy gradient and the service function chain traffic engineering method using DQN. DPG is a policy gradient method that directly learns a deterministic policy, that is, given a state, the method will directly output a certain behavior. DQN is a value-based learning method that learns the expected reward value of taking different actions under a given state, and indirectly determines which action to take through the reward value. For the sake of convenience, the present invention expresses the proposed method as DDPG-SFC-TE, expresses the DPG-based service function chain traffic engineering method as DPG-SFC, and expresses the DQN-based service function chain traffic engineering method as DQN-SFC-TE.
[0159] Then, on the basis of adopting the proposed method, the present invention also compares the performance of the proposed TEG diversion mechanism with the TEG routing mechanism and the static routing mechanism. The TEG diversion mechanism refers to the routing of multiple paths between adjacent VNF node pairs, and this mechanism can effectively utilize the bandwidth resources in the network. TEG routing means that in the case of TEG modeling, the on-off status of the physical link can be predicted in advance so as to dynamically adjust the route. Static routing refers to the routing selection of SFC according to the network topology of the current time slot. In the simulation of the present invention, the deployment of multiple SFCs and the deployment scenarios of different resource capacity settings are considered. The specific network simulation parameters are shown in Table 1.
[0160] Table 1 Network simulation parameters
[0161]
[0162] The DDPG-SFC-TE method adopted in the present invention is an SFC deployment strategy for minimum flow rate optimization. Therefore, the minimum flow rate is the main indicator for evaluating the performance of the method, that is, the larger the minimum flow rate, the better the performance of the method. First, the present invention verifies and compares the convergence performance of the proposed DDPG-SFC-TE method with the benchmark methods DPG-SFC-TE and DQN-SFC-TE. In the current simulation, the number of SFC deployments is set to 20. The DDPG-SFC-TE method can achieve the highest minimum flow rate, followed by DPG-SFC-TE, while the minimum flow rate of DQN-SFC-TE is the lowest. This is because DDPG-SFC-TE further incorporates deep learning technology, target network and experience replay mechanism, enhances learning efficiency and performance, and is more effective than DPG-SFC-TE. In contrast, the performance of DQN-SFC-TE is worse mainly because the DQN method itself is more suitable for processing discrete problems, while the SFC traffic engineering problem combines the characteristics of continuous and discrete, and its applicability is obviously not as good as the other two methods. Comparing the convergence speed of the three methods, the DDPG-SFC-TE method converges at about 200 iterations, the DPG-SFC-TE method converges at about 250 iterations, and the DQN-SFC-TE method converges at about 320 iterations. The convergence speed of DQN-SFC-TE is the slowest among the three, mainly because the DQN method needs to discretize the continuous action space, which increases the complexity of learning.
[0163] The minimum flow rate of the three methods under different numbers of SFCs is compared. First, the minimum flow rate will decrease as the number of SFCs increases. This is because the increase in the number of SFCs will increase the network load, causing the available bandwidth resources of the links in the network to become scarce. In addition, under different numbers of SFCs, DDPG-SFC-TE can achieve a higher minimum flow rate than DPG-SFC-TE and DQN-SFC-TE. This shows that the DDPG-SFC-TE method can more effectively utilize the bandwidth resources in the network and maintain high performance indicators even when the network load is high. However, DQN-SFC-TE still has the worst performance among the three, which also proves the limitations of deep reinforcement learning methods based on value functions in dealing with continuous problems.
[0164] Compare the performance of switching between different routing mechanisms based on the proposed method. Figure 3 The changing trends of the minimum flow rate with the change of the number of SFCs under different routing mechanisms are compared. Figure 3The most obvious thing is that the minimum flow rate of the three mechanisms decreases as the number of SFCs increases. This is because as the number of SFCs increases, the demand for available bandwidth resources of network links increases, which in turn makes resources more scarce. Then, by comparing the minimum flow rates between different mechanisms, the TEG diversion mechanism is the most obvious in optimizing the minimum flow rate, followed by the TEG-based routing mechanism, and the worst is the routing mechanism based on static topology. This is because TEG diversion can make more full use of the bandwidth resources of the physical link and achieve more detailed traffic distribution. The TEG routing strategy can dynamically adjust the routing and find a better transmission path. In contrast, static routing lags behind in performance because it is limited by the fixed topology and only routes through a single path.
[0165] Figure 4 The minimum flow rates of the three mechanisms under different bandwidth resource capacity conditions are compared. In order to ensure that the capacity of computing resources does not become a limiting factor for the increase of bandwidth resources, the average computing resource capacity of the node is set to 400 MIPS in the current simulation. With the increase of the average bandwidth resource capacity of the link, the minimum flow rates of the three mechanisms have been improved to varying degrees, among which the TEG diversion mechanism has the largest improvement, the TEG routing has a higher improvement, and the static routing has the slowest improvement. When bandwidth resources are less, the network congestion will be more serious. The TEG-based mechanism will bring a certain degree of flexibility, while the diversion mechanism can divide the traffic at a finer granularity, so the performance is the best, while the performance of static routing based on static topology is the worst among the three.
[0166] Figure 5 The performance of the three mechanisms is compared under different computing resource capacities. When the average computing resource capacity of the node increases, the minimum flow rate of the three mechanisms also increases accordingly. The performance of the three mechanisms is also the highest for TEG diversion, followed by TEG routing, and the worst for static routing. It is worth noting that, unlike the case of bandwidth resource capacity, the increase in computing resource capacity leads to a gradual decrease in the increase in the minimum flow rate. This is because the link bandwidth resource capacity of the current simulation setting is uniformly distributed at [200,400] Mbps. When the computing resource capacity increases, the minimum flow rate increases slowly due to the limitation of the link bandwidth resource capacity.
[0167] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution of the present invention, which should be included in the scope of the claims of the present invention.
Claims
1. A satellite Internet resource scheduling method based on a deep deterministic policy gradient algorithm, characterized in that: The steps include: S1: Modeling a satellite network SDSN with multiple satellite nodes and multiple service function chains SFC. Let SDSN be represented by a directed graph G(V,E), where V is the set of nodes, E is the set of links, and E includes the physical link set E vv and the storage link set E between satellites in adjacent time slots v ; The time evolution graph TEG divides the total time into T time slots, the length of each time slot is η, let t∈T represent the index of the time slot, in TEG, (i t ,j t )∈E vv Represents the physical link between two different satellites, with (i t ,i t+1 )∈E v represents the storage link between consecutive time slots of the same satellite; Let K = {1, 2, ..., k, ...} represent the SFC request set in SDSN. For k ∈ K, represents an ordered set of VNFs, represents the nth VNF of the kth SFC; The last VNF is represented as m represents the number of VNFs; and They represent the source node and destination node of the kth SFC respectively; S2: Assuming that in the same time slot, It can only be deployed on one satellite node. The VNF deployment constraints are modeled as: where i t represents the replica of satellite node i in the tth time slot, Deploy indicator variables for VNFs, Indicates The location where the VNF is deployed is node i t , otherwise it means The VNF is deployed at a location other than node i t superior; Define y k is the flow rate of the kth SFC, and defines the variable is the kth SFC on link (i t ,j d )∈E, where Indicates that the kth SFC is in link (i t ,j t )∈E vv The flow rate, Indicates that the kth SFC is in link (i t ,i t+1 )∈E v The flow rate is expressed as: in Indicates that the kth SFC passes through the After the VNF, the link (i t ,j d )∈E, Indicates an auxiliary VNF located on the source node, identifying the SFC flow that has not been processed by any VNF; S3: The computational resource constraints of satellite nodes are modeled as: in Indicates that in the link (j d ,j t ) Traffic rate after VNF processing, Represents node i t The computing resource capacity; Indicates the computing resource requirements per unit data flow rate; In addition, the bandwidth resource capacity of the link is expressed as: in Indicates link (i t ,j t ) bandwidth resource capacity, Indicates that the kth SFC is in link (i t ,j t )∈E vv The flow rate; S4: Construct the flow conservation constraints on the TEG flow diversion model according to different situations; S5: Set the minimum flow rate y that maximizes all SFCs min To optimize the goal: Non-negativity constraints are imposed on the flow rate of the SFC and the flow rate of each stage of each SFC and are modeled as: y k ≥0 S6: The controller of the SDSN network is used as an intelligent agent to centrally control the SFC deployment, and the optimization problem is modeled as an MDP model suitable for deep reinforcement learning. The state space, action space and reward function are defined. S7: Solve the MDP model constructed in S6 based on the DDPG model to obtain the optimal SFC deployment; The DDPG model includes an Actor network, a Critic network, and an experience replay pool. The parameters of the current network π(s) of the Actor network and its target network π′(s) are θ π and θ π ′, the parameters of the current network Q(s,a) of the Critic network and its target network Q′(s,a) are θ Q and θ Q′ The Actor network is responsible for action screening and strategy formulation, and updates the Actor network and Critic network parameters according to the strategy gradient ascent and the loss function gradient descent respectively. The Critic network is responsible for evaluating the generated strategy, and the experience replay pool is used to store the state s t , action a t , r t and the next state s t+1 The experience tuples constituted by Initialize the parameters of the Actor network and Critic network and the experience replay pool, and update the parameters of the Actor network and Critic network in each iteration; Input s to the DDPG model t , get a from the Actor network t , a t Apply to t In the process, the SDSN controller performs SFC deployment, and then obtains r t And update t+1 , will (s t ,a t ,r t ,s t+1 ) is saved as an experience tuple in the experience replay pool. When the experience replay pool is full, the earliest experience tuple will be replaced by the new experience tuple; A small batch of experience tuples are randomly selected from the experience replay pool for training. Under the goal of maximizing the expected cumulative discounted reward, the parameters of the Actor network and the Critic network are updated through gradient solving. The expected cumulative discounted reward refers to the expected value of the sum of all possible rewards in the future starting from the current state during the strategy execution process; the state-action value function Q π (s(t), a(t)) is used to estimate the expected cumulative discounted reward for a given state and action; Q π (s(t),a(t))=E[r(t)+γQ π (s(t+1),a(t+1))] Among them, r(t) represents the immediate reward, γQ π (s(t+1), a(t+1)) represents the future discounted reward, γ represents the discount factor, and the value range is 0≤γ≤1; The parameter updates of the Actor network and the Critic network are used to guide the SDSN controller to better perform SFC deployment in the next iteration until the number of training rounds reaches the set maximum value and the expected cumulative discounted reward of the optimization task is maximized, indicating that the training is completed and the optimal SFC deployment is obtained.
2. A satellite Internet resource scheduling method based on a deep deterministic policy gradient algorithm as claimed in claim 1, characterized in that: In S4, the flow conservation constraints of the TEG flow splitting model are divided into four cases: Source Node The sum of the outflow rates should be y k : in, Indicates the physical link (i t ,j t ) outbound rate that is not processed by VNF, Indicates that in the storage link (i t ,i t+1 ) outbound rate that is not processed by VNF; Destination Node The flow conservation constraint is expressed as: in, Indicates the physical link Pass The inflow rate after treatment, Indicates storage link Pass treated inflow rate; Deploy Node i t The inflow rate and outflow rate are both y k : in, Indicates that in the physical link (j t ,i t ) The inflow rate after treatment, Indicates that in the storage link (i t-1 ,i t ) treated inflow rate; Indicates that in the physical link (i t , j t ) The outflow rate after treatment, Indicates that in the storage link (i t ,i t+1 ) The outflow rate after treatment, Represents the VNF deployment indicator variable; Forwarding node i t The inflow rate is equal to the outflow rate: in, Indicates that in the physical link (j t ,i t ) The inflow rate after treatment, Indicates that in the storage link (i t-1 ,i t ) The inflow rate after treatment.
3. A satellite Internet resource scheduling method based on a deep deterministic policy gradient algorithm as claimed in claim 1, characterized in that: In S6, the state space, action space and reward function are: State Space: s t ={y(t),E(t),C(t),B(t)} Among them, y(t) represents the flow rate status requested by SFC, E(t) represents the network topology status of SDSN, C(t) and B(t) represent the remaining computing resources of the node and the remaining bandwidth resources of the physical link respectively; Action Space: a t ={X(t),Y(t)} in, Represents a collection of VNF deployment actions. express With satellite t The relationship between Represents a set of traffic distribution actions, Indicates that the kth SFC passes Then assigned to the physical link (i t , j t ) flow rate, Indicates that the kth SFC passes Then assigned to the storage link (i t ,i t+1 ) flow rate; Reward function: where y k (t) represents the traffic rate of the kth SFC in the tth time slot.
4. A satellite Internet resource scheduling method based on a deep deterministic policy gradient algorithm as claimed in claim 3, characterized in that: In S7, the loss function L(θ Q )for: L(θ Q )=E[(y t -Q(s t ,a t |θ Q )) 2 ] y t =r t +γQ′(s t+1 ,π(s t+1 |θ π′ )|θ Q′ ) Among them, y t is the target value, π′(s t+1 |θ π′ ) indicates that according to θ π′ and t+1 To decide the best action to take, Q′(s t+1 ,π′(s t+1 |θ π′ )|θ Q′ ) is the estimated value of the cumulative discounted reward expected from the actions taken by Q′(s,a) on future states, and γ represents the discount factor.
5. A satellite Internet resource scheduling method based on a deep deterministic policy gradient algorithm as claimed in claim 4, characterized in that: In S7, the Actor network parameters are updated as follows: Among them, α represents the learning rate of the Actor network, H represents the number of mini-batch experience tuples, s h They represent solving the gradient of the Q function with respect to a and solving the gradient of the policy function π with respect to θ. π The gradient of , the state of the h-th mini-batch sample; The critic network parameter update is expressed as: Among them, β represents the learning rate of the Critic network, y h , a h They represent the target values of the h-th mini-batch samples, respectively, and solve the Q function with respect to θ Q The gradient of , the action of the h-th mini-batch sample.
6. A satellite Internet resource scheduling method based on a deep deterministic policy gradient algorithm as claimed in claim 5, characterized in that: In S7, the target networks of the Actor network and the Critic network are updated as follows: i π′ =tθ π +(1-τ)θ π′ i Q′ =tθ Q +(1-τ)θ Q′ Where τ represents the soft update factor.
Citation Information
Patent Citations
Software defined satellite network virtual network function migration method
CN114710196A