A Satellite Internet Resource Scheduling Method Based on Deep Reinforcement Learning
By establishing a multi-layer converged network model in the satellite Internet, combining LSTM and DPPO algorithms, VNF backup strategy is optimized, and the problem of increasing backup AoI caused by link instability is solved, rapid failure recovery and efficient resource utilization is achieved, and the requirements of delay sensitivity of different services are adapted to the needs of different services.
Patent Information
- Application Number
- CN202411089606.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-09
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-08-09
AI Technical Summary
In the satellite Internet environment, link instability leads to an increase in the virtual network function (VNF) backup age (AoI), making it difficult to quickly recover failures, and the prior art has failed to effectively solve the problems of delay sensitivity and resource utilization efficiency of different services.
Using a method based on deep reinforcement learning, a software-defined multi-layer satellite fusion network model is established, combined with LSTM and DPPO algorithms, and by predicting the probability of VNF state transition, optimizing backup strategies, introducing adaptive bandwidth adjustment and migration mechanisms, optimizing resource allocation, ensuring rapid failure recovery and business continuity.
It effectively reduces the delay of VNF failure interruption, improves network reliability and resource utilization efficiency, adapts to dynamic environmental changes, ensures business reliability and reasonable allocation of resources, and improves prediction accuracy.
Smart Images

Figure CN119155264B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of satellite Internet, and particularly relates to a satellite Internet resource scheduling method based on deep reinforcement learning. Background Art
[0002] A satellite is a complex system in space, which consists of many components. If any one of these components fails due to software failure, CPU overload, or exceeding the temperature threshold, etc., it may cause the satellite to malfunction. Coupled with many factors in the space environment, such as solar wind, solar flares, sunspots, etc., they will also affect the normal operation of the satellite. Once a satellite node fails, it will directly affect the satellite communication function, resulting in a large amount of network latency and even service interruption. Usually, when a failure occurs in the SFC, the user will automatically retry to continue the connection. If the orchestrator can mitigate the impact of the failure before the recognizable time interval, the user will not experience the service interruption caused by the failure. Most failures can be predicted by monitoring the changes in the status information (such as resource overload and temperature) of the VNF on the node and using DRL, and then prepare for fault recovery according to the prediction results. However, migrating the VNF when a VNF failure occurs will generate a large amount of interruption latency, which seriously affects the service reliability. Using a predictive migration scheme to predict faults in advance can reduce the migration latency to a certain extent, but the migrations with prediction errors will also generate additional latency and resource waste. Therefore, in order to further ensure the reliability of the SFC while reasonably utilizing network resources, the present invention completes the fault recovery of the VNF through the method of pre-backup, and improves the performance loss caused by prediction errors. Specifically speaking, for satellite nodes that are prone to failure, backup VNFs are deployed on other servers in advance, and the backup VNFs are immediately started when the satellite nodes fail, and at the same time, the original node VNFs are deleted to save resources. If the VNF backup is not prepared in time due to prediction errors, the failed VNF needs to be immediately migrated urgently to reduce some interruption latency.
[0003] Existing literature has studied the rapid recovery of VNFs when ground network nodes fail, but most of them focus on stateless backup placement and availability optimization. Among them, in order to improve resource utilization efficiency, a new sharing mechanism for redundancy and multi-tenant technologies is proposed, and an availability-aware VNF deployment solution is proposed to calculate the modified availability of each VNF based on a given shared redundancy block. Existing methods also propose a backup resource allocation model that considers the importance of functional and backup server failure probabilities and intermediate servers, and a reliability-aware resource allocation algorithm to reduce the cost of redundant backup VNFs. The above studies do not consider the synchronous backup of VNF status information, and offline backup schemes are difficult to apply to dynamic network environments, so rapid recovery may not be possible when a VNF fails. To solve the problem of rapid recovery of VNFs by predicting failures, algorithms such as machine learning are used to predict failures and formulate backup plans in advance. To ensure service continuity, a master-slave VNF structure is used, and an active path recovery strategy is adopted to rapidly recover the VNF.
[0004] However, most existing studies are based on ground networks and cannot be directly applied to complex SGINs. On the one hand, different from fixed ground base stations, the relative movement speed between LEO satellites in different orbits is very fast, so the satellite server where the VNF backup is located may move away from the original VNF as the satellite moves; on the other hand, the interruption of the physical link increases the AoI of the backup VNF, resulting in more recovery time required when starting the VNF backup. Excessive recovery time will cause service interruption and it is difficult to ensure service quality. In addition, existing studies also rarely consider that different services have different delay sensitivities, and different decision weights should be set for the backups of services with different delay sensitivities. Since the available network resources on satellites are scarcer than those on ground networks, excessive VNF backups may occupy a large amount of network resources. Therefore, different VNF backup decisions should be made for VNFs with different delay sensitivities and different failure severities to achieve a compromise between delay and network resources. Summary of the Invention
[0005] Aiming at the above problems existing in the prior art, the technical problem to be solved by the present invention is: the unstable link in the SGIN environment easily leads to an increase in the AoI of the VNF backup, making it difficult to ensure rapid recovery when the VNF fails.
[0006] To solve the above technical problem, the present invention adopts the following technical solution: A satellite Internet resource scheduling method based on deep reinforcement learning, including the following steps:
[0007] S1: Combining the high coverage of MEO satellites and the high dynamics of LEO satellites, a satellite-ground integrated network SDMSGIN model under software-defined multi-layer satellites is established. The physical layer topology of the SDMSGIN model is represented as a weighted undirected graph G P =(N P , L P ), where N P is the set of physical nodes n, and L P is the set of physical links L nn′ (n, n′∈N P ). Define L nn′ (t) to represent the connection status of the physical link L nn′ at time slot t. Define the remaining computing resource capacity of each physical node n at time slot t as the remaining storage resource capacity as and the remaining bandwidth capacity in each physical link L nn′ as
[0008] The virtual application layer where the SFC is located is abstracted as a weighted directed graph G V =(F V , L V ), where F V is the set of VNFs, and L V is the set of virtual links. Define the set of SFCs as U, and define to represent the set of k VNFs that make up the u-th SFC. Define L u as the virtual link between two adjacent VNFs in the u-th SFC set, where the broadband requirement of the virtual link is Define the binary variable to represent the mapping situation of the virtual link. Define 's computing resource requirement as the storage resource requirement as
[0009] S2: A VNF state transition and fault recovery model is established to solve the problem of the increasing AoI of VNF backup caused by unstable links. To specifically describe how to help a faulty VNF recover quickly through VNF backup, an AoI-aware VNF backup model is established to adjust bandwidth resources to optimize AoI.
[0010] S3: In the AoI-aware VNF backup model, a VNF backup migration mechanism is introduced to solve the problem that the VNF backup moves away from the original VNF due to the relative movement between network nodes. An adaptive bandwidth adjustment strategy is introduced to allocate more bandwidth to more urgent situations to quickly synchronize the status data.
[0011] S4: Introduce a delay-sensitive factor as a weight to make an adaptive optimization strategy for services with different delay sensitivities. Under the tolerable delay constraint and resource constraint, an optimization problem is established with the goal of maximizing the VNF backup benefit.
[0012] S5: Propose a method combining LSTM and DPPO to solve the optimization problem proposed in S4. First, transform the optimization problem into an MDP, and define a quadruple (S, A, P, R) to represent the state space, action space, state transition probability, and reward function respectively;
[0013] Use LSTM to predict the VNF state transition probability, and use the state sequence before VNF time slot t to train LSTM. When the maximum number of iterations is reached, the trained LSTM is obtained and a new state is output.
[0014] Use the distributed reinforcement learning algorithm DDPO suitable for handling complex environments and coping with large-scale state spaces to optimize the backup strategy. When DDPO reaches the maximum number of iterations, the obtained strategy at this time is the optimal strategy.
[0015] Furthermore, in S2, the process of establishing the VNF state transition and fault recovery model is as follows:
[0016] Set the states of VNFs to normal state, warning state, and severe state. The transition probabilities between the three states of VNFs are:
[0017] Represents the probability that the th VNF in the normal state still remains in the normal state at time slot t, is the probability of entering the warning state. If this VNF enters the warning state, with probability it is in the warning state, the probability of completing the repair and restoring to the normal state. Let the binary variables and respectively represent the normal state, warning state, and severe state of the th VNF at time slot t. If the th VNF is in one of the three states at time slot t, the binary variable corresponding to this state is 1, otherwise it is 0.
[0018] Let represent the time when the th VNF fails and is in the warning state at time slot t, and define to represent the backup trigger waiting time threshold. If The server should immediately back up the VNF to other available nodes and transfer the status data in real time. If the VNF in the warning state returns to the normal state after and transfers with a probability of , the backup of the VNF is deleted to save node resources. The VNF in the warning state reaches the critical state with a probability of . The VNF in the critical state cannot return to other states and will stay in this state with a probability of 1. Obviously, if the VNF is not backed up at this time, it will cause service interruption, and the VNF can only be urgently migrated to other nodes. If the VNF is transferred to the critical state after being backed up in the above warning stage, it will return to the normal state with a probability of 1.
[0019] Since the delay sensitivity of different types of services is inconsistent, different backup trigger waiting time thresholds need to be designed for different services. Specifically, services with high delay sensitivity require a lower to prevent service degradation. Define the delay sensitivity factor of as . The higher the delay sensitivity of the service, the larger its value. Therefore, the value of is expressed as:
[0020]
[0021] where e represents the natural constant, and T w represents the basic waiting time.
[0022] Furthermore, in S2, the process of establishing the AoI-aware VNF backup model is as follows: The AoI at the end of the t-th time slot is expressed as:
[0023]
[0024] where is the AoI at the end of the t-th time slot, t p is the end time of the p-th time slot, p is a natural number, τ is the length of the time slot, is the size of the status data packet that the node n where the t-th time slot is located needs to send to the node n' where its backup * is located, is the bandwidth occupied by the logical link for real-time transmission of the VNF status.
[0025] The cumulative amount of untransmitted data in the t-th time slot is expressed as: is expressed as:
[0026]
[0027] where is the forwarding hop count between the original VNF and the backup VNF, is the backup interruption duration parameter indicating VNF backup * The sum of the interruption durations between the backup at node n and the original VNF in time slot t.
[0028] Furthermore, in the above S3, the adaptive bandwidth adjustment strategy is that the bandwidth adjustment factor is defined as Expressed as:
[0029]
[0030] where is the delay sensitivity factor, is the probability that the VNF status transfers to a severe fault predicted by the agent.
[0031] The AoI at the end of the t-th time slot is expressed as:
[0032]
[0033] Furthermore, in the above S3, the VNF backup migration mechanism is:
[0034] The AoI at the end of the t-th time slot is expressed as:
[0035]
[0036] where represents the VNF backup on node n * migrates to node n' in time slot t.
[0037] If the VNF is not backed up in advance due to prediction errors, a large amount of interruption delay will be generated when migrating the VNF emergently. Only by increasing the migration bandwidth urgently can this delay be reduced. Since the information difference between the node to be migrated and the original node is the size of the entire VNF data to be migrated, migrating this data emergently will generate interruption delay. Therefore, this migration delay is equivalent to AoI, and a binary variable is set represents the occurrence of this emergency migration, otherwise Then the AoI is finally expressed as:
[0038]
[0039] where is the total amount of data to be transmitted for migrating this VNF;
[0040]
[0041] where is the storage resource requirement for backup.
[0042] Furthermore, in S4, the process of establishing an optimization problem with maximizing the VNF backup benefit as the optimization objective is as follows:
[0043] Assume that the migration of VNF backup is completed instantaneously, each VNF can have at most one backup at the same moment, and in order to ensure that the traffic is not split, introduce another constraint 1: Among them, and respectively represent the deployment of VNF backup and the deployment of the status synchronization link;
[0044] Each VNF and its backup cannot be in the same node, introduce another constraint 2: Among them, is the VNF deployment situation.
[0045] Introduce resource constraints:
[0046]
[0047] Among them, are respectively the computing resource requirements of the backup and the computing resource requirements of are respectively the storage resource requirements of the backup and the storage resource requirements of is the remaining computing resource capacity of each physical node n at time slot t, is the remaining storage resource capacity;
[0048] The sum of the virtual links between VNFs of adjacent nodes, the logical synchronization links for transmitting status, and the bandwidth occupied by migrating backup VNFs is less than the remaining available bandwidth capacity of this physical link:
[0049]
[0050] Among them, is the mapping situation of the virtual link and is the virtual link the broadband requirement of, is the remaining bandwidth capacity in each physical link;
[0051] Define a binary variable If has been backed up at time slot t and backup recovery can be performed at this time slot, then otherwise it is 0: Among them, IF(.) is an indicator function, when holds, IF(.) is 1, otherwise IF(.) is 0;
[0052] Introduce the tolerable delay constraint: Among them is the severe state of the VNF in the t-th time slot;
[0053] Define the optimization objective function It consists of two parts: the benefit Gain(t) and the cost st(t):
[0054] Among them, is the severe state of the VNF in the (t + 1)-th time slot, is the warning state of the VNF in the t-th time slot.
[0055] For the cost part of the objective function, it is divided into three parts: Cost(t) = Cost 1 (t) + Cost 2 (t) + Cost 3 (t)
[0056]
[0057] Among them, is the probability that the VNF in the (t + 1)-th time slot completes the recovery from the warning state to the normal state, is the normal state of the VNF in the (t + 1)-th time slot, is the normal state of the VNF in the t-th time slot;
[0058]
[0059] Among them, is the probability that the VNF in the (t + 1)-th time slot is in the warning state, is the warning state of the VNF in the (t + 1)-th time slot, is the warning state of the VNF in the t-th time slot;
[0060]
[0061] Among them, is the probability that the VNF in the (t + 1)-th time slot enters the severe state from the warning state.
[0062] Furthermore, in the above S5, the state space, action space, state transition probability, and reward function are respectively:
[0063] State space: Define s(t) = {Res(t), L(t), P(t)} ∈ S to represent the state space of the SDMSGIN network in the t-th time slot, where represents the state space of the VNF resource demand, L(t) = {l nn′ (t)|n ∈ N p} represents the state space of the physical link, represents the state space of the three key state transition probabilities of the VNF;
[0064] Action space: Define a(t) = {X(t), Y(t), Z(t)} ∈ A, where and represents the action space of the VNF backup decision and the state synchronization link mapping decision, is the action space of the backup migration decision.
[0065] State transition probability: The transition probability P represents the probability that the agent transfers to the next state s(t+1) after taking the action a(t) in the state s(t), denoted as p(s(t+1)|s(t),a(t)).
[0066] Reward function: To maximize the reward R, the reward at time slot t is represented by the optimization objective:
[0067] If the constraint conditions in P1 are not satisfied, then: r(t) = -1 / ζ, where ζ represents an infinitesimal quantity approaching 0, so as to impose a large penalty value on the decision that violates the constraint. To consider future benefits while considering the immediate reward, the reward discount factor at time slot t is set to γ ∈ (0,1), then its cumulative discounted reward is: In the formula, k is the number of iterations. Obviously, the more iterations, the smaller the cumulative reward. It represents the decision to take the action a(t) in the state s(t). The policy value function is used to evaluate the quality of the policy at time slot t: Q π (s(t), a(t)) = E[R π (t)|s(t), a(t)]
[0068] where, Q π (s(t),a(t)) is the policy value function, and E[.] is the expectation.
[0069] It is iteratively represented by the Bellman equation: Q π (s(t), a(t)) = E[r(t) + γQ π (s(t+1), a(t+1))], where r(t) is the reward function.
[0070] The optimal decision π * is represented as:
[0071] Furthermore, in the above S5, the process of obtaining the optimal decision is as follows:
[0072] Input: State sequence {s(1), s(2),…s(t)}, G P =(N P ,LP ), G V = (F V , L V ), discount factor γ, target function update frequency F, number of threads N thread , global network iteration times M all , local network iteration times M part , learning rate α 1 and α 2 ;
[0073] Output: optimal policy π * ;
[0074] 1) Initialize LSTM network parameters, initialize global Actor network parameters and global Critic network parameters θ, initialize the Actor network parameters of the i-th agent as the Critic network parameters of the i-th agent as θ i ;
[0075] 2) Let ecod = 1
[0076] 3) Let episode = 1;
[0077] 4) Let thread = 1;
[0078] 5) Initialize the environment S and obtain the initial state s' from the SDN controller;
[0079] 6) t = 1 to T;
[0080] 7) Sequentially extract the state sequence {s(1), s(2), … s(t)} from the state data before time t;
[0081] 8) Input {s(1), s(2), … s(t)} into LSTM;
[0082] 9) Calculate the probability distribution parameters (μ, σ 2 ) of state s(t + 1) through the LSTM network and the fully connected layer;
[0083] 10) Sample to determine state s(t + 1) according to the probability distribution parameters (μ, σ 2 ) of state s(t + 1);
[0084] 11) Calculate the VNF backup trigger threshold according to the prediction result of the state transition probability of the VNF
[0085]
[0086] where e represents the natural constant and Tw represents the base waiting time;
[0087] 12) The SDN controller observes the environmental state s and selects an action a(t) from the policy π(s i (t)|a i (t)) of the local Actor network;
[0088] 13) If the tolerable delay constraint, resource constraint, other constraint one, and other constraint two are all satisfied simultaneously, then
[0089] execute the action a(t), obtain the reward r(t), transfer to the new state s(t + 1), and execute the next step;
[0090] Otherwise, the reward value r(t) = -1 / ζ, re - select the action a(t) from the local Actor network, and execute 12);
[0091] 14) Judge whether thread > N thread holds. If so, execute the next step; otherwise, set thread = thread + 1 and return to 5);
[0092] 15) Update the global Actor network parameter and the global Critic network parameter θ according to the following formula;
[0093]
[0094] φ = φ + α 1 Δφ
[0095]
[0096] θ = θ + α 2 Δθ
[0097] where, is the gradient of the global Actor network, Δθ is the gradient of the global Critic network, α 1 and α 2 are the learning rates of the global Actor network and the global Critic network respectively, L clip (.) and L(.) are the clipped loss function and the loss function respectively;
[0098] 16) Update the state sequence {s(1), s(2), … s(t), s(t + 1)};
[0099] 17) Update the LSTM network parameter using r(t) and s(t + 1);
[0100] 18) Judge whether episode > M partIs it established? If so, proceed to the next step; otherwise, set episode = episode + 1 and return to 4);
[0101] 19) Determine whether ecod > M all Is it established? If so, proceed to the next step; otherwise, set ecod = ecod + 1 and return to 3);
[0102] 20) Output the optimal strategy.
[0103] Relationship between the global Actor network and the local Actor network:
[0104] Corresponding relationship: The global Actor network is a centralized network that maintains a set of parameters, which are the "summarization" or "average" of the parameters of multiple local Actor networks. Each local Actor network is a copy of the global Actor network and is used to make decisions and learn independently in different environment instances or different threads. Synchronization: The local Actor network synchronizes its parameters from the global Actor network regularly to ensure that they are up-to-date. This helps to reduce the bias caused by the parameter differences between different local networks. Asynchronous update: In each local environment, the local Actor network interacts with the environment independently and calculates gradients based on its experience (state, action, reward, etc.). These gradients are then sent to the global Actor network to update the parameters of the global network. The update of the global network is asynchronous, that is, it does not need to wait for all local networks to complete their updates.
[0105] Relationship between the global Critic network and the local Critic network:
[0106] Corresponding relationship: Similar to the relationship between the global Actor network and the local Actor network, the global Critic network also maintains a set of parameters, which are the summarization or average of the parameters of multiple local Critic networks. Each local Critic network evaluates the value of the action selected by the corresponding local Actor network and calculates the loss and gradients. Synchronization: Similar to the local Actor network, the local Critic network also synchronizes its parameters from the global Critic network regularly to maintain the consistency and accuracy of its evaluation. Asynchronous update: The local Critic network calculates the loss and gradients based on the experience of the local Actor network (including state, action, and reward) and sends these gradients to the global Critic network for asynchronous update. This update mechanism helps the global Critic network to more accurately evaluate the action values in different states.
[0107] The global Actor network and the global Critic network are centralized, and they summarize the information and parameters of multiple local networks (local Actor networks and local Critic networks) respectively. The local networks are responsible for making decisions and learning independently in their respective environments, and feedback the learned experience (through gradients) to the global network for updating. This combination of global and local helps to improve the parallelism and learning efficiency of the algorithm.
[0108] Compared with the prior art, the present invention has at least the following advantages:
[0109] The method of the present invention effectively solves the challenges of VNF backup in SGIN, and improves the reliability of the network and the utilization efficiency of resources. Specifically, it reduces the interruption delay: by predicting faults and making early backups, and quickly starting the backup VNF when a fault occurs, it significantly reduces the interruption delay caused by VNF faults. It improves service reliability: through the pre-backup and backup migration mechanisms, it ensures the continuity and reliability of services even under the dynamic changes in the SGIN environment. It optimizes resource utilization: through the AoI awareness model and optimization algorithms, it reasonably allocates network resources, avoiding waste of resources while ensuring service quality. It has strong adaptability: the algorithm can adapt to the dynamic changes in the SGIN environment and adjust the backup strategy in real time to cope with the changing network conditions. It improves prediction accuracy: by predicting the VNF status through the LSTM network, it improves the prediction accuracy of VNF faults, thus making more accurate backup decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0110] Figure 1 It is a field diagram of the SDMSGIN network.
[0111] Figure 2 It is a VNF state transition and fault recovery framework.
[0112] Figure 3 It is a schematic diagram of the training process of implementing the LSTM algorithm.
[0113] Figure 4 The effective VNF backup ratio under different algorithms.
[0114] Figure 5 The average interruption delay of VNF under different bandwidth adjustment strategies.
[0115] Figure 6 The average interruption delay of VNF under different backup migration mechanisms. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0116] The present invention will be further described in detail below.
[0117] First, to address the problem of increased AoI in VNF backup due to unstable links, a VNF state transition and fault recovery model was established, and an AoI-aware VNF backup model was established to adjust bandwidth resources to optimize AoI. Second, to solve the problem of the relative movement between network nodes causing the VNF backup to move away from the original VNF, a VNF backup pre-migration mechanism was introduced. Finally, to make adaptive optimization strategies for services with different delay sensitivities, a delay-sensitive factor was introduced as a weight, an optimization problem was established with the goal of maximizing the VNF backup benefit, and an algorithm combining LSTM and DPPO was proposed for solution. The simulation results show that the proposed scheme can reduce the interruption delay of VNFs with severe failures while reasonably utilizing network resources.
[0118] A satellite Internet resource scheduling method based on deep reinforcement learning is described as follows:
[0119] First, by combining the high coverage of MEO satellites and the high dynamics of LEO satellites, a software-defined multi-layer satellite-ground integrated network (SDMSGIN) model was established to create a highly dynamic and widely covered network environment for VNF backup.
[0120] Next, for the link instability problem in SDMSGIN, through the real-time monitoring of the SDN controller, the VNF status and link conditions were accurately detected, and the backup strategy was dynamically adjusted to meet the rapid recovery requirements when VNF failures occur. In addition, considering that the relative movement of nodes may cause the distance between the backup node and the original node to increase, a VNF backup migration mechanism was introduced to ensure that the backup always stays on the optimal node, thereby minimizing the fault recovery time.
[0121] Second, to solve the optimization problems of resource allocation and different service delay sensitivities during the backup process, an AoI-aware model was constructed. By taking the AoI optimization after backup success as the benefit and the possible self-repair and emergency migration interruption during the backup process as the cost, corresponding optimization weights were assigned to services with different delay sensitivities. Then, with the goal of maximizing the decision benefit after subtracting the cost, an optimization problem was established under the tolerable delay and resource constraints.
[0122] Finally, in order to solve the problem of uncertain state transition probability when VNF fails, a state prediction algorithm based on LSTM is designed to predict the state information. In order to solve the problem of long model training time and high overhead of single-agent VNF backup algorithm, an AoI-aware VNF hierarchical backup algorithm based on DPPO is proposed. The VNF backup strategy is optimized in real time by predicting the backup needs of VNF through multi-agents, and the optimal VNF backup strategy is obtained through continuous training.
[0123] 1. SDMSGIN model
[0124] Compared with LEO satellites, MEO satellites have a larger delay but a larger coverage area, and then have more stable satellite-to-ground links and inter-layer links between different satellite layers. Therefore, the present invention combines multiple layers of MEO satellites and LEO satellites and integrates the ground network to establish the SDMSGIN model, such as Figure 1 The ground central control server is responsible for the central management of the entire SGIN, including the deployment, configuration and optimization of SFC, as well as the scheduling of network resources such as bandwidth, computing and storage; GEO satellites with ultra-large coverage can serve as auxiliary controllers and work together with the ground central control server to manage SDMSGIN; MEO satellites and LEO satellites are equipped with servers that can process and forward data.
[0125] To model SDMSGIN, its physical layer topology is represented as a weighted undirected graph G P =(N P ,L P ), where N P is the set of physical nodes n, L P is the physical link L nn′ (n,n′∈N P ). Due to the high-speed movement of the satellite, the physical link is periodically disconnected and connected, so L is defined nn′ (t) represents time slot t physical link L nn′ The connection state, where L nn′ (t) = 1 indicates physical link L nn′ In the connected state at time slot t, L nn′ (t) = 0 indicates that the physical link L nn′ In order to describe the network resource usage, the remaining computing resource capacity of each physical node n in time slot t is defined as The remaining capacity of storage resources is Each physical link L nn′ The remaining bandwidth capacity is
[0126] The virtual application layer where SFC is located is abstracted as a weighted directed graph GV = (F V , L V ), where F V is the set of VNFs, and L V is the set of virtual links. Define the SFC set as U, and define to represent the set of k VNFs that make up the u-th SFC. Define L u as the set of virtual links between adjacent VNFs in the u-th SFC , where the bandwidth requirement of the virtual link is Define the binary variable to represent the mapping situation of the virtual link. When the virtual link is mapped to the physical link L nn′ at time slot t otherwise Define the binary variable to represent the deployment situation of the VNF. If is deployed at node n at time slot t, then otherwise Define 's computing resource requirement as the storage resource requirement as
[0127] 2. VNF State Transition and Fault Recovery Model
[0128] The state transition and fault recovery framework is as Figure 2 shown. The three states of the VNF are set as the normal state, the warning state, and the severe state. Among them, the normal state means that the result caused by the failure of the VNF is at a tolerable level; the warning state means that the failure of the VNF has started to cause a recognizable performance degradation. If it remains in the warning state for a period of time, it means that the failure may not be restored to normal in a short time. At this time, the VNF needs to be backed up to other available nodes, and the status information is transmitted in real time; the severe state means that the failure degree of the VNF is already very serious and cannot be restored to the normal state within the tolerable time. It is necessary to immediately start the VNF backup instance, and in order to reduce the load of this node, it is necessary to delete the failed VNF instance.
[0129] The transition probabilities between the three states of the VNF are defined as: represents the probability that the VNF in the normal state remains in the normal state in the next time slot t, is the probability of entering the warning state. If the VNF enters the warning state, with probability it is in the warning state, the probability of completing the repair and restoring to the normal state. Represents the probability of entering the severe state. Let the binary variables and represent the three states of normal state, warning state, and severe state respectively. If in this state, it is equal to 1, otherwise it is equal to 0. Represents the time when the VNF is in the warning state due to a failure, and define Represents the backup trigger waiting time threshold. If Indicates that the server should immediately back up the VNF to other available nodes and transfer the status data in real time. If the VNF in the warning state is returns to the normal state (with a probability of transition of ), then delete the backup of the VNF to save node resources. The VNF in the warning state reaches the severe state with a probability of . The VNF that enters the severe state cannot recover to other states and will stay in this state with a probability of 1 (absorbing state). Obviously, at this time, if the VNF is not backed up, it will cause service interruption and can only be urgently migrated to other nodes. If the VNF is transferred to the severe state after backing up the VNF in the above warning stage, it will return to the normal state with a probability of 1.
[0130] Since the delay sensitivity of different types of services is inconsistent, therefore, different backup trigger waiting time thresholds need to be designed for different services. Specifically, services with high delay sensitivity require a lower to prevent service degradation. Define the delay sensitivity factor of service as The higher the delay sensitivity of the service, the larger its value. Therefore, the value of can be expressed as:
[0131]
[0132] where e represents the natural constant, which is introduced to weaken the excessive influence of probability on the backup trigger waiting time threshold, and Tw represents the base waiting time.
[0133] 3. Establish an AoI-aware VNF backup model
[0134] Since the VNF generates status data such as configuration information and event logs during operation, in order to enable the VNF instance to quickly recover and start from a failure, an additional logical synchronization link is required to transfer the status information in real time to synchronize the status data between the VNF instances of the backup node and the original node in real time. Define the backup of as * Assume that the VNF data packet is generated and starts to be transmitted at the beginning of each time slot. Let the node n where is located at the beginning of the t time slot need to back up its* The size of the status data packet sent by the node n' where it is located is Let the bandwidth occupied by the logical link for real-time transmission of VNF status be Use AoI to represent the time difference between the synchronization of the VNF backup instance and the primary VNF instance, that is, the difference between the time when the VNF backup instance receives the data packet and the generation time of the data packet. Let t p be the end time of the p-th (p is a natural number) time slot, and let the length of a time slot be τ. The AoI at the end of the t-th time slot is expressed as:
[0135]
[0136] Among them, is the AoI at the end of the t-th time slot, t p is the end time of the p-th time slot, p is a natural number, τ is the length of the time slot, is the t-th time slot The node n where it is located needs to send its backup * The size of the status data packet sent by the node n' where it is located, is the bandwidth occupied by the logical link for real-time transmission of VNF status;
[0137] Since the link on-off situation can be predicted through the network topology, the time when the link is disconnected can be avoided by adjusting the sending of data packets. Therefore, it is assumed that data packets will not be sent during the disconnection process, that is, it is assumed that there is no packet loss. In this case, the actual time available for data transmission in each time slot is Then in each time slot, the amount of data that the backup node can receive is Considering that there may be multiple physical links between the original VNF and the backup VNF, the amount of data not transmitted in each time slot is Among them is the number of forwarding hops between the original VNF and the backup VNF. By the end of the p-th time slot, the total amount of data not transmitted will be the sum of the amounts of data not transmitted in all previous time slots. Therefore, at the t-th time slot The accumulated amount of data not transmitted is expressed as:
[0138]
[0139] Among them, is the number of forwarding hops between the original VNF and the backup VNF, The backup interruption duration parameter represents the VNF backup * The sum of the interruption durations between the node n and the original VNF in the time slot t.
[0140] 4. Adaptive Bandwidth Adjustment Strategy
[0141] For more urgent situations, network resources should be sacrificed first to avoid service interruption. Therefore, the delay-sensitive factor The larger the service, the greater the bandwidth should be allocated, and the probability of transfer to a severe fault predicted by the agent The larger this probability is, the greater the bandwidth should be. Therefore, an adaptive bandwidth adjustment strategy is proposed, that is, more urgent situations are allocated a larger bandwidth to quickly synchronize the status data. Define the bandwidth adjustment factor as Expressed as:
[0142]
[0143] where is the delay-sensitive factor is the probability of transfer of the VNF status to a severe fault predicted by the agent;
[0144] After introducing the adaptive bandwidth adjustment strategy, the AoI at the end of the t-th time slot is expressed as:
[0145]
[0146] 5. VNF Backup Migration Mechanism
[0147] Due to the relative movement of the satellite, the interruption of the link will cause an increase in AoI. The relative movement between the backup node and the original VNF node may cause an increase in the number of physical links passed through, thus increasing the forwarding hops, and this situation will also increase AoI. To solve the above problems, a VNF backup migration mechanism is introduced, that is, the VNF backup instance with too long cumulative interruption time should be pre-migrated to other nodes to reduce the routing forwarding hops. Although the VNF instance can be pre-migrated, the synchronization of the status data will change the AoI Denote the VNF backup on node n Migrate to node n at time slot t ′ The size of the data volume.
[0148] After introducing the VNF backup migration mechanism, the AoI at the end of the t-th time slot is expressed as:
[0149]
[0150] where Denote the VNF backup on node n Migrate to node n at time slot t ′ ;
[0151] If the VNF is not backed up in advance due to prediction errors, a large amount of interruption delay will be generated when the VNF is urgently migrated. Only by urgently increasing the migration bandwidth can this delay be reduced. Since the information difference between the node to be migrated and the original node is the size of the entire VNF data to be migrated, urgently migrating this data will generate interruption delay. Therefore, this migration delay is equivalent to AoI, and a binary variable is used to represent the situation where this urgent migration occurs. Otherwise the final expression of AoI is:
[0152]
[0153] where is the total amount of data to be transmitted for migrating this VNF, including the sum of the VNF status data and the data size required for re-instantiating the
[0154] VNF;
[0155]
[0156] where is the storage resource requirement for backup.
[0157] 6. Establishing an optimization problem
[0158] The backup of the failed VNF transferred from the warning state to the severe state is an effective backup, while the backup existing when it returns to the normal state is an invalid backup. The invalid backup will occupy too much network resources, but if there is no prior backup when the VNF transfers to the severe state, it will cause service interruption. The present invention takes the optimization of AoI under the effective backup of the VNF as the benefit, and the optimization of AoI and service interruption under the invalid backup as the cost. Under the constraints of network resources and tolerable delay, the backup of the VNF and the benefits and costs of optimizing AoI are jointly considered to establish an optimization problem.
[0159] The backup of the failed VNF transferred from the warning state to the severe state is an effective backup, while the backup existing when it returns to the normal state is an invalid backup. The invalid backup will occupy too much network resources, but if there is no prior backup when the VNF transfers to the severe state, it will cause service interruption. Taking the optimization of AoI under the effective backup of the VNF as the benefit, and the optimization of AoI and service interruption under the invalid backup as the cost, under the constraints of network resources and tolerable delay, the backup of the VNF and the benefits and costs of optimizing AoI are jointly considered to establish an optimization problem.
[0160] Define binary variables and to represent the deployment of the VNF backup and the deployment of the status synchronization link respectively. If the backup of is mapped to node n during time slot t, then On the contrary, it is Similarly, if during time slot t, and its backup* the state synchronization link between them is embedded into the physical link L nn′ , then Otherwise, Since the migration of the VNF backup is pre-emptive and the migration process of the VNF backup does not affect the original VNF's data processing, it is assumed that the migration of the VNF backup is completed instantaneously. Each VNF can have at most one backup at the same time, and in order to ensure that the traffic is not split, another constraint is introduced:
[0161]
[0162] Among them, and respectively represent the deployment of the VNF backup and the deployment of the state synchronization link;
[0163] If a VNF fails on a certain node server, then all VNFs of this type on this node will stop working. Each VNF and its backup cannot be in the same node, and another constraint is introduced:
[0164]
[0165] Among them, is the VNF deployment situation.
[0166] Once the failure of the VNF reaches a serious level, the backup instance of the VNF should be immediately started. Therefore, there should be remaining resource capacity for the backup instance on its node and the passing links. In each time slot, the computing resource capacity and storage resource capacity occupied by all VNF instances and backup instances on the node should not exceed the remaining available resource capacity, and a resource constraint is introduced:
[0167]
[0168] Among them, are respectively the computing resource requirements of the backup and the computing resource requirements of are respectively the storage resource requirements of the backup and the storage resource requirements of is the remaining computing resource capacity of each physical node n in time slot t, is the remaining storage resource capacity;
[0169] The sum of the virtual links between VNFs of adjacent nodes, the logical synchronization links for transmitting status, and the bandwidth occupied by migrating backup VNFs is less than the remaining available bandwidth capacity of the physical link:
[0170]
[0171] Among them, is the mapping situation of the virtual link ; is the broadband requirement of the virtual link ; is the remaining bandwidth capacity in each physical link;
[0172] Define the binary variable If has been backed up at time slot t and backup recovery can be performed at this time slot, then otherwise it is 0:
[0173]
[0174] Among them, IF(.) is the indicator function. When holds, IF(.) is 1, otherwise IF(.) is 0;
[0175] When the VNF in the warning state transfers to the severe state, if a VNF backup instance has been deployed in advance at this time, the time required to start this VNF backup instance and transfer the remaining status data is the size of the VNF backup AoI; if there is no VNF backup instance, the time required to urgently migrate the VNF is the size of the equivalent AoI. Both of these situations will generate interruption delay. In order to prevent the interruption delay from being too large and causing a violation of the service level agreement, it is necessary to complete the fault recovery of the VNF by starting a backup instance or urgently migrating the VNF within the tolerable delay , and introduce the tolerable delay constraint: Among them, is the severe state of the VNF at time slot t.
[0176] Take as the tolerable coefficient because different delay-sensitive services have different tolerable delays. The more delay-sensitive the service is, the smaller its tolerable delay should be, and the less delay-sensitive the service is, the larger its tolerable delay should be.
[0177] In order for the VNF to recover quickly when a failure occurs, the AoI of the VNF backup should be minimized to reduce the time for synchronizing the remaining status data. However, only when the effective backup and fault recovery are performed under the correct prediction of the VNF state transition probability will there be a benefit in optimizing the AoI. Otherwise, the optimization of the AoI in the case of incorrect prediction will be a cost for occupying additional network resources. Therefore, define the optimization objective function It consists of two parts: the gain Gain(t) and the cost st(t):
[0178] For the gain part of the objective function, it should be the AoI when the VNF completes the backup in the warning state at time slot t and transfers to the severe state at time slot t+1. At this time, the smaller the AoI of the objective function, the greater the gain. In addition, for different latency-sensitive services, minimizing the AoI has different gains, and the higher the latency sensitivity ( the greater), the greater the gain. Therefore, the gain part of the objective function is expressed as:
[0179]
[0180] Among them, is the severe state of the VNF at time slot t+1, is the warning state of the VNF at time slot t;
[0181] For the cost part of the objective function, it is divided into three parts: Cost(t) = Cost 1 (t) + Cost 2 (t) + Cost 3 (t).
[0182] The first part of the cost, Cost 1 (t), occurs when the VNF backs up the VNF in the warning state at time slot t, but due to a prediction error, the VNF is successfully repaired and transferred to the normal state at time slot t+1. Then, the AoI optimization for this VNF before was the cost of occupying additional network resources. In addition, to reduce the impact of unexpectedness on the solution of the optimization objective, the lower the probability of this situation occurring, the lower the cost should be. Therefore, Cost 1 (t) is expressed as:
[0183]
[0184] Among them, is the probability that the VNF completes the recovery from the warning state to the normal state at time slot t+1, is the normal state of the VNF at time slot t+1, is the normal state of the VNF at time slot t;
[0185] The second part of the cost, Cost 2(t) backs up the VNF when the VNF is in a warning state at time slot t and is still in a warning state at time slot t + 1. At this time, the backup and the optimization for AoI belong to premature backup and also belong to the cost of occupying less network resources. However, the priority of uninterrupted service should be greater than the occupation of network resources. To avoid sacrificing reliability to reduce the consumption of network resources, the weight of this cost should be further weakened, taking the square of the probability of this situation occurring. Therefore, Cost 2 (t) is expressed as:
[0186]
[0187] Among them, is the probability that the VNF is in a warning state at time slot t + 1, is the warning state of the VNF at time slot t + 1, is the warning state of the VNF at time slot t;
[0188] The third part of the cost, Cost 3 (t) means that the VNF is in a warning state at time slot t but the VNF is not backed up, but transfers to a severe state at time slot t + 1. If the VNF is not backed up in the warning state but transfers to the warning state due to a prediction error, the equivalent under the emergency migration VNF interruption delay will be generated at this time Similarly, in order to reduce the impact of unexpectedness on the optimization. The impact of the solution is weighted using this state transition probability, so Cost 3 (t) is expressed as:
[0189]
[0190] Among them, is the probability that the VNF transfers from the warning state to the severe state at time slot t + 1.
[0191] 7. MDP model
[0192] The selection decision of VNF backup, the migration decision of backup, and the bandwidth adjustment decision not only affect the network performance at the current moment but also affect the future state and decision. To capture the sequentiality and cumulative effect of this decision and provide a basis for finding the long-term optimal strategy, the optimization problem is transformed into an MDP model, and a quadruple (S, A, P, R) is defined to represent the state space, action space, state transition probability, and reward function respectively.
[0193] State space: Define s(t) = {Res(t), L(t), P(t)} ∈ S to represent the SDMSGIN network state space at time slot t, where represents the state space of VNF resource requirements, L(t) = {l nn′ (t)|n ∈ N p} represents the state space of the physical link, represents the state space of the three key state transition probabilities of the VNF;
[0194] Action space: Define a(t) = {X(t), Y(t), Z(t)} ∈ A, where and represents the action space of the VNF backup decision and the state synchronization link mapping decision, is the action space of the backup migration decision.
[0195] State transition probability: The transition probability P represents the probability that the agent transfers to the next state s(t + 1) after taking the action a(t) in the state s(t), denoted as p(s(t + 1)|s(t), a(t)).
[0196] Reward function: To maximize the reward R, the reward at time slot t is represented by the optimization objective:
[0197] If the constraint conditions in P1 are not satisfied, then: r(t) = -1 / ζ, where ζ represents an infinitesimal quantity approaching 0; The reward discount factor at time slot t is set to γ ∈ (0, 1), then its cumulative discounted reward is: In the formula, k is the number of iterations. Obviously, the more iterations, the smaller the cumulative reward. It represents the decision to take the action a(t) in the state s(t). The policy value function is used to evaluate the quality of the policy at time slot t: Q π (s(t), a(t)) = E[R π (t)|s(t), a(t)], where Q π (s(t), a(t)) is the policy value function, and E[.] is the expectation.
[0198] It is iteratively represented by the Bellman equation: Q π (s(t), a(t)) = E[r(t) + γQ π (s(t + 1), a(t + 1))], where r(t) is the reward function.
[0199] The optimal decision π* is represented as:
[0200] 7. State prediction algorithm based on LSTM
[0201] Such as Figure 3The figure shows the framework of the probability distribution parameters output by the policy network, demonstrating the training process of the policy network using the LSTM network. Specifically, to capture the long-term dependencies in the data, first, the LSTM network intercepts a sequence of past states {s(1), s(2),... s(t)}, where these states are mainly the state transition probabilities P(t) of the VNF observed by the agent in the historical dynamic network environment. This sequence is sequentially input into the LSTM network. After being processed repeatedly, the LSTM network outputs a sequence of hidden states, where each element represents the output of an LSTM cell. Then, these outputs pass through a fully connected layer to generate the probability distribution parameters of state s(t+1). Finally, the probability distribution output by the output layer includes the mean μ and variance σ 2 , where the mean μ represents the average of the predicted state, reflecting the most likely situation of the state s(t+1) predicted by the LSTM network based on the input data, and the variance σ 2 represents the prediction uncertainty or the degree of dispersion of the predicted values around the mean. The smaller the variance, the more confident the model is in the predicted state transition probability, and the lower the uncertainty of the prediction result. The larger the variance, the higher the prediction uncertainty. During the training phase, the agent selects the backup and migration strategies of the agent by sampling from this distribution to ensure the randomness and exploration of strategy selection. As the training progresses, the variance σ 2 , gradually decreases, indicating that the network is gradually stabilizing during learning. After the training is completed, for the obtained VNF state prediction results, the agent can determine the backup and migration operations of the VNF based on the mean of the decision, while considering the variance to evaluate the uncertainty and risk of the decision. The training process of the VNF state prediction algorithm based on LSTM is shown in Algorithm 1.
[0202] The training process of the policy network introducing LSTM
[0203]
[0204] 8. VNF Backup Algorithm Based on DPPO
[0205] Regarding each VNF as an agent, each is equipped with a pair of Actor network and Critic network. Let the parameters of the Actor network of the i-th agent be The parameters of the Critic network of the i-th agent are θ i; First, each agent explores in the network environment, collects the network resource information required by the VNF, the status data of the physical links of the SDMSGIN, and the three states of the VNF, and performs VNF backup behavior based on the collected status data and calculates the immediate reward. Through interaction with the network environment, the local PPO (Proximal Policy Optimization algorithm) network of each agent can gradually optimize the VNF backup strategy based on the obtained reward information. For each agent, the Critic part in its PPO network uses the state value function and the advantage function to evaluate the current output strategy of the Actor network. The expected cumulative reward that the i-th agent can obtain by adopting the mapping strategy in the current state, that is, the state value function, is expressed as:
[0206] V(s i (t))=E[r i (t)+γV(s i (t+1))]
[0207] When the agent is in state s i (t) and takes action a i (t), it can obtain a cumulative expected value, that is, the state-action value function, which is expressed as:
[0208] Q(s i (t), a i (t))=E[r i (t)+γQ(s i (t+1), a i (t+1))]
[0209] And use the advantage function A π (s i (t), a i (t)) to evaluate the quality of the action a i (t) made by agent i for the current state s i (t):
[0210] A π (s i (t), a i (t))=Q(s i (t), a i (t))-V(s i (t))
[0211] For the advantage estimation of policy gradient reinforcement learning, there are usually methods such as TD advantage estimation, Q-value advantage estimation, and maximum entropy advantage estimation. However, considering the dynamic environmental factors in this model, such as failure transfer probability and bandwidth adjustment, in order to effectively update the advantage estimation in such a dynamic environment, help the agent better understand the impact of environmental changes on long-term rewards, and thus make better decisions, the more flexible generalized advantage estimation is used. Then, the advantage function of agent i is expressed as:
[0212]
[0213] In the formula, T is the step length of one training, and the parameter λ is the discount decay factor. When making the VNF backup strategy and how to adjust the bandwidth, the long-term rewards need to be considered. Especially when considering the optimization of AoI, GAE improves the efficiency and stability of the policy gradient algorithm by adjusting the trade-off between bias and variance, allowing the agent to find a balance between immediate rewards and long-term rewards, which helps to formulate better backup strategies and bandwidth adjustment strategies. And δ i (t) is the TD error, which is expressed as:
[0214] δ i (t) = r i (t) + γ(V(s i (t + 1)) - V(s i (t))
[0215] Through the continuous evaluation of the Critic network of agent i, the TD error will gradually become smaller. To measure the difference between the predicted value and the target value in this process, a loss function is introduced:
[0216]
[0217] To limit the update range of the new policy, thereby reducing the sensitivity of the learning rate and improving the stability of the algorithm, the DPPO algorithm adds the action ratio ρ(φ i ) parameter of the new and old policies to represent the importance sampling weight. Let the new policy be π new (s i (t)|a i (t), φ i ), and the old policy be π old (s i (t)|a i (t), φ i ). The small ratio ρ(φ i ) < 1 indicates that the agent is more inclined to choose the old policy, otherwise it chooses the new policy. The action ratio is expressed as follows:
[0218] Excessive policy updates can lead to unstable learning. To ensure that the agent does not lose its learning direction due to excessive changes when exploring backup policies and bandwidth adjustment decisions, a PPO algorithm that clips and optimizes the objective is used to prevent the new policy from deviating significantly from the old policy. The clipped loss function is expressed as:
[0219] L clip (φ i ) = E[min(A(s i (t), a i (t)) clip(ρ(φ i ), 1 - δ, 1 + δ), ρ(φ i ) A(s i (t), a i (t)))]
[0220] In the formula, δ is a hyperparameter that clips the policy update amplitude, and clip(ρ(φ i ), 1 - δ, 1 + δ) represents clipping the ratio of the new and old policies outside the interval (1 - δ, 1 + δ). The clipping method is expressed as:
[0221]
[0222] The parameters of the global PPO network are updated by the loss functions of the local Critic and Actor networks. Among them, the gradients and parameters of the global Actor network, as well as the gradients Δθ and parameters θ of the global Critic network, are updated as follows:
[0223]
[0224] φ = φ + α 1 Δφ
[0225]
[0226] θ = θ + α 2 Δθ
[0227] Among them, is the gradient of the global Actor network, Δθ is the gradient of the global Critic network, α 1 and α 2 are the learning rates of the global Actor network and the global Critic network respectively. L clip (.) and L(.) are the clipped loss function and the global network parameters after each round of training of the loss function respectively. The specific training process of the algorithm is shown in Algorithm 2.
[0228]
[0229] 9. Simulation and Evaluation
[0230] To evaluate the performance of the method of the present invention, experimental simulations were carried out on a computer with an AMD Ryzen 7 processor and 16GB of memory. First, a satellite network topology was built in the STK software, and the generated data was saved. Then, a custom environment, that is, the system model part, was built in the ycharm2022 software, and the data generated in STK was imported into the custom environment. Then, the LSTM was built using the Pytorch toolkit. Finally, various experimental results were obtained by continuously training the agent using the DPPO algorithm.
[0231] (1) First, 6 polar orbits with a height of 1000 km were built in STK, each orbit having a LEO satellite constellation of 6 satellites, and 3 MEO satellite constellations with an orbital height of 10000 km and an inclination of 550, each orbit having 4 satellites. And 12 ground stations were set up. After simulating and running for 1800 s, link visibility report data was obtained, and this data was applied to the custom environment built by the Py charm software. The custom environment was built using the network system simulation parameters shown in Table 1. The LSTM model is responsible for capturing the time series dependencies of the VNF states, while the DPPO algorithm optimizes the network backup and resource allocation strategies based on these prediction results. The present invention combines the LSTM and DPPO algorithms for the design of VNF state prediction and backup strategy optimization, and adopts distributed reinforcement learning to handle large-scale state spaces and action spaces, constructs a data set, and after the agent runs in the environment 20 times, the generated data trajectory is saved as a batch processing data set in the buffer. In the LSTM-DPPO algorithm, the LSTM network configuration includes 2 hidden layers, with 100 hidden units in each layer. The DPPO part uses a two-layer fully connected network, with the number of neurons in each layer being 256 to adapt to complex decision-making requirements. The learning rates of the Critic and Actor networks are set to 0.002, the clipping parameter is set to 0.2, and the target network update frequency is 5. The overall configuration of the algorithm aims to balance learning efficiency and prediction accuracy while maintaining the adaptability to dynamic network environments.
[0232] Table 1 Simulation Parameters
[0233]
[0234] (2) Performance Evaluation
[0235] When the learning rate is set to 0.0005, the cumulative reward value is too small, resulting in a too small step size for model parameter update. This makes the model move slowly in the parameter space, leading to a slower convergence rate of the reward value and difficulty in converging to a higher reward value. When the learning rate is 0.002, this value is too large, causing the step size of parameter update to be too large. Although it speeds up the convergence rate of the algorithm, the excessive parameter update amplitude causes deviation from the optimal solution and is very unstable. When the learning rate is 0.001, it is relatively stable and can converge to a higher reward value.
[0236] The LSTM-DPPO algorithm converges faster and more stably, while the LSTM-PPO algorithm and the LSTM-SAC algorithm converge slower and have relatively larger fluctuations. This is because the DPPO algorithm uses multiple agents to distributively and comprehensively explore the network environment. Although it will result in a higher algorithm space complexity, the parallel processing makes its algorithm time complexity lower, so that it can more efficiently find a broader solution space.
[0237] When the VNF failure that occurs transfers from the warning state to the severe state, if the VNF is backed up in advance for fast fault recovery, it belongs to an effective backup. And the higher the proportion of effective VNF backups within a certain period of time, the higher the accuracy of the algorithm's prediction of severe VNF failures. Figure 4 By comparing the effective VNF backup ratios of the LSTM-DPPO algorithm, the LSTM-PPO algorithm, and the LSTM-SAC algorithm, it can be seen that the effective VNF backup ratio under the LSTM-DPPO algorithm is significantly higher than that of the LSTM-PPO algorithm and the LSTM-SAC algorithm. And as the number of VNFs increases, the decline in its performance is also significantly lower than that of the LSTM-PPO algorithm and the LSTM-SAC algorithm. The LSTM-DPPO algorithm improves the accuracy by approximately 11.06% - 17.93% compared to the LSTM-PPO algorithm. This shows that the LSTM-DPPO algorithm can predict the state transition probability of VNFs more accurately, thus being able to make better VNF backup decisions.
[0238] Figure 5The average interruption delay of VNFs with different degrees of delay sensitivity under the adaptive bandwidth adjustment strategy proposed in the present invention, the fixed bandwidth allocation strategy, and the random bandwidth allocation strategy is compared. It can be clearly seen that the average interruption delay of VNFs under the adaptive bandwidth adjustment strategy proposed in the present invention is significantly reduced, which is approximately 41.43%-46.37% lower than that of the fixed bandwidth allocation strategy, and the reduction is greater for VNFs with higher delay sensitivity. On the one hand, the adaptive bandwidth adjustment strategy proposed in the present invention adjusts the bandwidth of the state synchronization link according to different state transition probabilities and different delay sensitivities. For VNFs with a greater probability of transitioning to a severe state predicted by LSTM, a larger bandwidth will be allocated to reduce the backup AoI, so that the interruption delay can be reduced when the VNF recovers from a failure. On the other hand, when LSTM fails to predict the situation where a VNF in a warning state transitions to a severe state and no backup is performed, the VNF and state data need to be urgently migrated at this time. The adaptive bandwidth adjustment strategy also adjusts the bandwidth required for the urgent migration of the VNF according to different degrees of delay sensitivity. The higher the delay sensitivity, the larger the allocated bandwidth, so the reduced interruption delay is also more.
[0239] Figure 6 The average interruption delay of VNFs under the backup migration mechanism proposed in the present invention, the non-backup migration mechanism, and the backup random migration mechanism is compared. It can be seen that the average interruption delay of VNFs under the backup migration mechanism proposed in the present invention is approximately 37.23%-43.62% lower than that of the non-backup migration mechanism. This is because in the face of unstable links and relative movement of nodes in the SGIN environment, which will lead to an increase in the backup AoI of VNFs, the backup VNF is migrated closer to the original VNF or to a location with less interruption time of the state synchronization link, so that the backup AoI is reduced. And this migration mechanism also adopts the adaptive bandwidth adjustment strategy, so that while ensuring service reliability, it does not occupy too much network resources.
[0240] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A satellite Internet resource scheduling method based on deep reinforcement learning, characterized in that: The steps include: S1: Combining the high coverage of MEO satellites and the high dynamics of LEO satellites, a satellite-ground fusion network SDMSGIN model under software-defined multi-layer satellites is established. The physical layer topology of the SDMSGIN model is represented as a weighted undirected graph G P =(N P ,L P ), where N P is the set of physical nodes n, L P is the physical link L nn′ (n,n′∈N P ), define L nn′ (t) represents the time slot t of the physical link L nn′ The remaining computing resource capacity of each physical node n in time slot t is defined as The remaining capacity of storage resources is Each physical link L nn′ The remaining bandwidth capacity is The virtual application layer where SFC is located is abstracted as a weighted directed graph G V =(F V ,L V ), where F V is the VNF set, L V is a set of virtual links, define the SFC set as U, and define Denotes the set of k VNFs that make up the u-th SFC, and defines L u is the virtual link between two adjacent VNFs in the uth SFC A collection of Virtual Link The bandwidth requirement is Defining binary variables Indicates the mapping of virtual links and defines The computing resource requirements are The storage resource requirement is S2: A VNF state transfer and fault recovery model was established to solve the problem of increased AoI of VNF backup caused by link instability. An AoI-aware VNF backup model was established to adjust bandwidth resources to optimize AoI. S3: Introduce a VNF backup migration mechanism in the AoI-aware VNF backup model to solve the problem that the VNF backup is far away from the original VNF due to the relative movement between network nodes, and introduce an adaptive bandwidth adjustment strategy to allocate a larger bandwidth in more urgent situations; S4: Introduce the delay sensitivity factor as a weight to make adaptive optimization strategies for services with different delay sensitivities. Under the constraints of tolerable delay and resources, establish an optimization problem with maximizing the VNF backup benefit as the optimization goal. S5: A method combining LSTM and DPPO is proposed to solve the optimization problem proposed in S4. First, the optimization problem is transformed into MDP, and a four-tuple (S, A, P, R) is defined to represent the state space, action space, state transition probability and reward function respectively. The VNF state transition probability is predicted through LSTM, and the LSTM is trained using the state sequence before the VNF time slot t. When the maximum number of iterations is reached, the trained LSTM is obtained and the new state is output; The distributed reinforcement learning algorithm DDPO is used to optimize the backup strategy. When DDPO reaches the maximum number of iterations, the strategy obtained at this time is the optimal strategy.
2. A satellite Internet resource scheduling method based on deep reinforcement learning as claimed in claim 1, characterized in that: In S2, the process of establishing the VNF state transfer and fault recovery model is as follows: The states of VNF are set as normal state, warning state and critical state; the transition probabilities between the three states of VNF are: In normal state The probability that a VNF remains in a normal state at time slot t is is the probability of entering the warning state. If the VNF enters the warning state, the probability In warning state, The probability of completing the repair and returning to normal state; let the binary variable and Respectively represent The normal state, warning state and critical state of a VNF in time slot t. If a VNF is in one of the three states at time slot t, the binary variable corresponding to the state is 1, otherwise it is 0; set up Indicates The time that a VNF is in the warning state when a fault occurs in time slot t, and define Indicates the backup trigger waiting time threshold; if The server should immediately back up the VNF to other available nodes and transmit status data in real time. After that, it returns to normal state. If the probability of transfer is high, the backup of the VNF is deleted to save node resources; the VNF in the warning state is The probability of reaching a critical state is 1. A VNF that enters a critical state cannot be restored to other states and will stay in this state with a probability of 1. If the VNF has been backed up and transferred to a critical state in the above warning stage, it will be restored to a normal state with a probability of 1. definition The delay sensitivity factor is The more delay-sensitive the service is, the larger the value is. The value of is expressed as: Where e represents a natural constant, T w Indicates the basic waiting time.
3. A satellite Internet resource scheduling method based on deep reinforcement learning as claimed in claim 2, characterized in that: In S2, the process of establishing the AoI-aware VNF backup model is as follows: The AoI at the end of the tth time slot is expressed as: in, is the AoI at the end of the t-th time slot, t p is the end time of the pth time slot, p is a natural number, τ is the length of the time slot, is time slot t The node n where the backup is needed The size of the state data packet sent by node n′, Occupies bandwidth for logical links used to deliver VNF status in real time; t time slot The accumulated amount of data not transferred It is expressed as: in, is the number of forwarding hops between the original VNF and the backup VNF, The backup interruption duration parameter indicates VNF backup The sum of the interruption durations between node n and the original VNF in time slot t.
4. A satellite Internet resource scheduling method based on deep reinforcement learning as claimed in claim 3, characterized in that: In S3, the adaptive bandwidth adjustment strategy is: Define the bandwidth adjustment factor as Expressed as: in is the delay sensitivity factor, The probability of the VNF state transitioning to a serious fault predicted by the agent; After introducing the adaptive bandwidth adjustment strategy, the AoI at the end of the tth time slot is expressed as:
5. A satellite Internet resource scheduling method based on deep reinforcement learning as claimed in claim 4, characterized in that: In S3, the VNF backup migration mechanism is: After the VNF backup migration mechanism is introduced, the AoI at the end of the tth time slot is expressed as: in, Indicates VNF backup on node n The amount of data migrated to node n′ at time slot t; If the VNF is not backed up in advance due to prediction errors, the emergency migration of the VNF will cause a lot of interruption delays. The only way to reduce this delay is to increase the migration bandwidth urgently. Since the information difference between the node to be migrated and the original node is the size of the entire VNF data to be migrated, the emergency migration of this data will cause interruption delays. Therefore, this migration delay is equivalent to AoI, and a binary variable is set Indicates that the emergency migration occurs, otherwise The AoI is finally expressed as: in, The total amount of data that needs to be transferred to migrate the VNF; in Storage resource requirements for backup.
6. A satellite Internet resource scheduling method based on deep reinforcement learning as claimed in claim 5, characterized in that: In S4, the process of establishing the optimization problem with maximizing the VNF backup benefit as the optimization goal is: Assume that the migration of VNF backup is completed instantly, each VNF can have at most one backup at the same time, and to ensure that traffic is not split, introduce other constraints: in and They represent the deployment of VNF backup and the deployment of state synchronization link respectively; Each VNF and its backup cannot be in the same node, which introduces another constraint: in VNF deployment status; Introducing resource constraints: in, They are The computing resource requirements for backup and The computing resource requirements, They are Backup storage resource requirements and Storage resource requirements, For each physical node n, the remaining computing resource capacity in time slot t is: The remaining capacity of the storage resource; The sum of the bandwidth occupied by the virtual links between VNFs of adjacent nodes, the logical synchronization links for transmitting status, and the migrated backup VNF is less than the remaining available bandwidth capacity of the physical link: in, Virtual Link The mapping situation, Virtual Link The broadband demand is the remaining bandwidth capacity in each physical link; Defining binary variables if Backup has been performed in time slot t and backup recovery can be performed in this time slot. Otherwise 0: Among them, IF(.) is an indicator function. If true, IF(.) is 1, otherwise IF(.) is 0; Introducing tolerable delay constraints: in, is the critical state of the VNF in time slot t; Define the optimization objective function It consists of two parts: gain (t) and cost (t): in, is the critical state of the VNF at time slot t+1, is the warning status of the VNF in time slot t; For the cost part of the objective function, it is divided into three parts: Cost(t)=Cost1(t)+Cost2(t)+Cost3(t); in, is the probability that VNF completes recovery from warning state to normal state in time slot t+1, is the normal state of VNF at time slot t+1, is the normal state of the VNF in time slot t; in, is the probability that the VNF is in the warning state at time slot t+1, is the warning status of the VNF in time slot t+1, is the warning status of the VNF in time slot t; in, is the probability that the VNF enters the severe state from the warning state at the t+1 time slot.
7. A satellite Internet resource scheduling method based on deep reinforcement learning as claimed in claim 6, characterized in that: In S5, the state space, action space, state transition probability and reward function are respectively: State space: Define s(t) = {Res(t), L(t), P(t)}∈S to represent the SDMSGIN network state space of time slot t, where The state space representing the VNF resource requirements, L(t) = {l nn′ (t)|n∈N p } represents the state space of the physical link, The state space representing the three key state transition probabilities of the VNF; Action space: define a(t) = {X(t), Y(t), Z(t)}∈A, where and represents the action space of VNF backup decision and state synchronization link mapping decision, is the action space of backup migration decisions; State transition probability: The transition probability P represents the probability that the agent will transfer to the next state s(t+1) after taking action a(t) in state s(t), expressed as p(d(t+1)|s(t),a(t); Reward function: To maximize the reward R, the reward for time slot t is expressed by the optimization objective: If the constraints in P1 are not met, then: r(t) = -1 / ζ, where ζ represents an infinitesimal quantity that tends to 0; the reward discount factor for time slot t is set to γ∈(0,1), then its cumulative discounted reward is: Where k is the number of iterations. Obviously, the more iterations, the smaller the cumulative reward. It represents the decision of taking action a(t) in state s(t). The strategy value function is used to judge the quality of the strategy in time slot t: Q π (s(t), a(t)) = E[R π (t)|s(t),a(t)], where Q π (s(t), a(t)) is the strategy value function, E[.] is the expectation; It is expressed by the iteration of Bellman equation: Q π (s(t), a(t)) = E[r(t) + γQ π (s(t+1), a(t+1))], where r(t) is the reward function; Optimal decision π * Expressed as:
8. A satellite Internet resource scheduling method based on deep reinforcement learning as claimed in claim 7, characterized in that: In S5, the process of obtaining the optimal decision is: Input: state sequence {s(1), s(2), …s(t)}, G P =(N P ,L P ), G V =(F V ,L V ), discount factor γ, objective function update frequency F, number of threads N thread , the number of global network iterations M all , the number of local network iterations M part , learning rates α1 and α2; Output: optimal strategy π * ; 1) Initialize LSTM network parameters and global Actor network parameters And the global Critic network parameters θ, initialize the Actor network parameters of the i-th agent as The parameter of the Critic network of the i-th agent is θ i ; 2) Set ecod = 1 3) Set episode = 1; 4) Set thread = 1; 5) Initialize the environment S and obtain the initial state s' from the SDN controller; 6) t = 1 to T; 7) Sequentially extract the state sequence {s(1), s(2), …s(t)} from the state data before time t; 8) Input {s(1), s(2), …s(t)} into LSTM; 9) Calculate the probability distribution parameters (μ, σ) of state s(t+1) through the LSTM network and the fully connected layer 2 ); 10) According to the probability distribution parameters (μ, σ) of state s(t+1) 2 ) Sampling determines the state s(t+1); 11) Calculate the VNF backup trigger threshold based on the VNF state transition probability prediction result Where, e represents the natural constant, Tw represents the basic waiting time; 12) The SDN controller observes the environment state s and selects the strategy π(s) from the local actor network i (t)|a i (t)) select action a(t); 13) If the tolerable delay constraint, resource constraint, other constraint 1 and other constraint 2 are satisfied at the same time, then Execute action a(t), obtain reward r(t), transfer to new state s(t+1), and execute the next step; Otherwise, the reward value r(t) = -1 / ζ, and the action a(t) is reselected from the local Actor network and 12) is executed; 14) Determine thread>N thread Is it true? If yes, execute the next step, otherwise set thread = thread + 1 and return 5); 15) Update the global Actor network parameters according to the following formula and global critic network parameter θ; φ=φ+α1Δφ θ=θ+α2Δθ in, is the gradient of the global Actor network, Δθ is the gradient of the global Critic network, α1 and α2 are the learning rates of the global Actor network and the global Critic network respectively, L clip (.) and L(.) are the loss function and loss function after pruning, respectively; 16) Update the state sequence {s(1), s(2), …s(t), s(t+1)}; 17) Use r(t) and s(t+1) to update the LSTM network parameters; 18) Determine episode>M part Is it true? If yes, go to the next step, otherwise set episode = episode + 1 and return to 4); 19) Determine if ecod>M all Is it true? If yes, execute the next step, otherwise set ecod = ecod + 1 and return to 3); 20) Output the optimal strategy.
Citation Information
Patent Citations
Virtual network function reliability deployment method based on backup revenue and remapping
CN110190987A
Service function chain reliable deployment method based on deep reinforcement learning
CN111147307A