Method, device and medium for parallel routing optimization of delay-sensitive service function chains
By using the SA-EDDPG algorithm with a self-attention mechanism in edge computing networks, VNF placement and traffic routing are optimized, solving the SFC parallel routing problem for latency-sensitive applications in edge computing networks, and achieving low latency and efficient resource utilization.
Patent Information
- Application Number
- CN202310532713.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-10
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2043-05-10
AI Technical Summary
In edge computing networks, the existing SFC parallel routing algorithm struggles to effectively handle the needs of latency-sensitive applications, especially under resource-limited and dynamic network conditions. It cannot simultaneously optimize both discrete and continuous actions, resulting in high latency and resource consumption.
The enhanced deep deterministic policy gradient algorithm (SA-EDDPG) based on the self-attention mechanism is adopted. By building a Markov decision process model and combining offline training and online operation, it optimizes VNF placement and traffic routing, solves the SFC parallel routing problem, and meets latency and resource constraints.
While meeting latency and resource constraints, it significantly reduces server and link resource consumption, improves SFC parallel routing efficiency, reduces latency, and adapts to network status changes.
Smart Images

Figure CN116566891B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of edge computing technology, and in particular to a method, device, and medium for optimizing parallel routing of delay-sensitive service function chains. Background Art
[0002] In recent years, with the rapid development of 5G and artificial intelligence technologies, a large number of latency-sensitive and compute-intensive applications have emerged in networks. Traditional cloud computing models are no longer able to effectively meet these demands. Consequently, the concept of edge computing has been proposed. Edge computing can fully utilize resource-constrained edge devices to provide computing, storage, and communication services close to end users, significantly reducing network latency and alleviating pressure on data center networks. With the rise of edge computing applications, a large number of network function elements (NFs) providing network services have been placed at the network edge. Therefore, how to flexibly and efficiently manage these network functions, enabling them to be dynamically created, deleted, and migrated to meet users' diverse service needs within the limited resources of the network edge, has become a pressing issue in the development of edge computing.
[0003] Traditionally, network functions are implemented using dedicated network devices, an approach that is costly and lacks scalability. Network Function Virtualization (NFV) has emerged as an emerging solution. Unlike traditional approaches, NFV implements network functions using virtualization technology on general-purpose servers. Through NFV's orchestration and scheduling mechanisms, service requests in edge computing can be fulfilled through a Service Function Chain (SFC), an ordered set of Virtual Network Functions (VNFs). This solves the network function management challenges in edge computing and reduces service latency and operating expenses.
[0004] However, a key issue with using SFC is that each VNF in the SFC needs to be placed on a physical node, and traffic flowing through the SFC must be processed sequentially by each VNF, which can result in high latency and fail to meet the needs of latency-sensitive applications in edge computing networks. Processing service requests through parallel SFCs has become a solution to this problem. To implement this approach, each traffic needs to be split into multiple sub-flows, and VNFs with the same functionality need to be repeatedly instantiated on different server nodes. In this way, duplicate instances of each VNF can process the split sub-flows in parallel to share the processing load, thereby reducing the processing latency of the service. In addition, traffic splitting and repeated instantiation of VNFs consume a certain amount of resources. Due to the limited resources in edge computing networks, considering resource consumption is crucial to ensuring cost-effectiveness for network operators.
[0005] There has been some research on the SFC parallel routing problem. Because this problem is NP-hard, researchers often propose heuristic or approximate algorithms to address it. However, most of these studies assume known network traffic and perform SFC parallel routing based on this assumption, ignoring the dynamic changes in network state. Reinforcement learning (RL) can automatically adjust policies based on historical experience and environmental feedback to make optimal decisions in complex and uncertain environments. Therefore, some studies have used RL algorithms to perform SFC parallel routing in dynamic network environments. However, when the network scale is large, traditional RL algorithms struggle to accurately describe complex network state changes, becoming inefficient. Deep reinforcement learning (DRL) combines deep learning and RL, using deep neural networks (DNNs) to handle the complex state changes of large-scale networks. It has become an emerging technology for solving the online SFC parallel routing problem. However, most research using DRL methods has been conducted in data center networks and has not considered services with low latency requirements. Therefore, it is possible to consider using DRL methods for SFC parallel routing in dynamic edge computing networks, but some challenges remain. On the one hand, the SFC parallel routing process involves both discrete actions (such as VNF placement) and continuous actions (such as traffic segmentation). Existing DRL-based SFC routing algorithms cannot effectively handle both types of actions simultaneously. On the other hand, the location distribution of server nodes in edge computing networks may be relatively dispersed, which affects the efficiency of SFC routing. Summary of the Invention
[0006] The purpose of the present invention is to provide a method, device and medium for optimizing parallel routing of delay-sensitive service function chains, which minimizes the resource consumption of servers and links while meeting delay and resource constraints, solves the online routing problem of SFC, and the discrete and continuous action problems, and improves the efficiency of SFC parallel routing.
[0007] The purpose of the present invention can be achieved by the following technical solutions:
[0008] A method for parallel routing optimization of a delay-sensitive service function chain comprises the following steps:
[0009] Construct the parallel routing problem of service function chains in edge computing networks based on network function virtualization. Considering the real-time network status changes, the Markov decision process is used to model the parallel routing problem of service function chains and obtain the MDP model.
[0010] An enhanced deep deterministic policy gradient algorithm based on the self-attention mechanism is used to solve the MDP model. The solution process includes two parts: offline training and online operation. The offline training part obtains a trained SA-EDDPG model. The online operation part determines the optimal routing solution by running the trained SA-EDDPG model and performs parallel routing of the service function chain.
[0011] The network model of the edge computing network is:
[0012] The underlying physical network is represented as a connectivity graph G = (N, L), where N is the set of server nodes, including various devices with computing capabilities at the edge of the network, and L is the set of physical links connecting two server nodes. For each server node n∈N, use C n Indicates its resource capacity. For each link l∈L, use B l and D l Represent its available bandwidth capacity and communication delay respectively;
[0013] Each service function chain request is defined as a series of ordered virtual network functions (VNFs). Network traffic needs to be routed through each VNF in the service function chain in turn. For the traffic routed through the service function chain, it is assumed that it can be split and divided into I sub-flows, each sub-flow is represented by i. R represents the set of service function chain requests arriving in real time. For each service function chain request r(F,ε r ,D r )∈R, F represents the set of VNFs required by the service function chain request, ε r Indicates the flow rate of the routing traffic through the service function link, D r Represents the delay limit requested by the service function chain; for each sub-flow i requested by the service function chain, the set of VNFs required by sub-flow i is represented as F i ={f1,f2,…,f m ,…,f |i|}, where f m is the mth VNF required by subflow i, |i| represents the number of VNFs contained in subflow i; each VNF service instance f m ∈F i There is a resource requirement, use To express.
[0014] The method of modeling the service function chain parallel routing problem using the Markov decision process includes the following steps:
[0015] S1: Determine the constraints;
[0016] S11: Determine the node capacity constraint. The node capacity constraint ensures that the total resources consumed by the VNFs required by the request do not exceed the resource limit of the server node n to be deployed, which is expressed as:
[0017]
[0018] in, is a 0-1 variable. If the mth VNF f required by subflow i of request r m If it is placed on physical node n, the value of this variable is 1, otherwise its value is 0;
[0019] S12: Determine a link capacity constraint, which ensures that the total bandwidth required by all requests through the link l∈L does not exceed the link's bandwidth capacity B l , expressed as:
[0020]
[0021] in, is a 0-1 variable that takes the value 1 if a link l is used to deliver subflow i of request r, and takes the value 0 otherwise; Is a continuous variable, indicating the traffic ratio divided by the sub-flow i of request r, that is, the traffic split ratio, and its value range is
[0022] S13: Determine the delay constraint: For a request r, its delay consists of two parts: the processing delay on the server node and the communication delay when the link transmits traffic, that is, the total delay D of request r total Expressed as:
[0023]
[0024] in, Represents the server node n∈N on f m ∈i’s processing delay, the maximum delay among all sub-flows i is used as the total delay of request r;
[0025] For each request r, if it can be successfully received, the total delay D of request r is total Cannot exceed its delay limit D r :
[0026]
[0027] S14: Determine a placement constraint that ensures that only one physical server is selected to place the mth VNF required by subflow i of request r, and that all VNFs in subflow i of request r can be served, expressed as:
[0028]
[0029]
[0030] S2: Determine that the optimization goal of the service function chain parallel routing problem is to minimize the joint resource consumption, that is, to minimize the resource consumption of the server and the bandwidth consumption of the link, where
[0031] Resource consumption of all servers U N Expressed as:
[0032]
[0033] Bandwidth consumption of all links U L Expressed as:
[0034]
[0035] Then, the optimization objective of joint resource consumption is expressed as:
[0036]
[0037] Where η1 and η2 are the weights of server resource consumption and link bandwidth consumption, respectively, satisfying η1, η2∈(0,1) and η1+η2=1.
[0038] The MDP model is defined as a five-tuple in and are the state space and action space respectively, is the state transition probability distribution, is the reward function, γ∈[0,1] is the discount factor for future rewards, which is defined as follows:
[0039] State space: The state at time t Defined as a vector G t =(C t ,B t ,D t ), G t Used to represent the characteristics of the underlying physical network at time t, where Represents the currently available resources of all server nodes, Represents the current available bandwidth of all physical links, Indicates the delay of all physical links;
[0040] Action space: The action space of an agent is a set Each action a∈A represents VNF placement and traffic routing, and the action a at time t is defined as in, It is a discrete action, indicating whether the m-th VNF of subflow i of request r is placed on server node n at time t; It is a discrete action used to indicate whether the subflow i requesting r at time t is routed through the physical link l; It is a continuous action, used to represent the traffic split ratio of request r at time t;
[0041] State transition: The state transition is represented as (s t ,a t ,r t ,s t+1 ), where s t is the network status at the current time t, a t is the action used to process VNF placement and traffic routing in subflow i of request r, r t and s t+1 Execute action a t The immediate reward obtained after and the network state at the next moment t+1, for each state State transition probability p(s t+1 ∣s t ,a t ) represents the agent in the network state s t Next, perform action a t After that, the network state transitions to s t+1 probability;
[0042] Reward function: Based on the optimization goal of the service function chain parallel routing problem, the reward function Defined as the negative of the total resource consumption of the server and link:
[0043]
[0044] The offline training process includes the following steps:
[0045] Step 1) The agent interacts with the environment to generate training data, where the environment refers to the edge computing network based on network function virtualization. The agent first observes the current network state s t , and the network state s is converted into t Transition to state s′ based on neighbor node information t , the agent performs action a based on its current strategy t , the environment is based on the current state s′ t and the received action a t , update the status to s t+1 , and feeds back a reward signal r to the agent t , the agent receives rt Update your own strategy to make better decisions at the next moment t+1, and generate training data (s t ,a t ,r t ,s t+1 );
[0046] Step 2) The training data (s t ,a t ,r t ,s t+1 ) is stored in the experience replay pool;
[0047] Step 3) When the state transition samples accumulate to a preset number, a batch of samples is randomly selected from the experience replay pool and input into the SA-EDDPG model;
[0048] Step 4) Train the SA-EDDPG model, where the input of the SA-EDDPG model is the underlying physical network G and the set of service requests R, and the output is VNFf m The placement location and traffic routing path of the network adopts a dual actor-critic network structure, which contains a total of four neural networks, namely the main actor network μ(s|θ μ ), the main critic network Q(s,a|θ Q ), target actor network μ′(s|θ μ′ ) and the target critic network Q′(s,a|θ Q′ ), where θ μ ,θ Q ,θ μ′ and θ Q′ is a parameter in the neural network. The actor network is responsible for generating VNF placement and traffic routing actions under a given network state. The critic network is responsible for evaluating the actions generated by the actor network. During the training process, the main actor network updates the parameter θ through the policy gradient method. μ , the master critic network updates the parameters θ by the gradient descent method based on the temporal difference error Q , the target actor network and the target critic network use soft update to update the parameters θ μ′ and θ Q′ ;
[0049] The online operation process includes the following steps:
[0050] Step 5) Select the SA-EDDPG model trained during the offline training process for online routing of the service function chain;
[0051] Step 6) Set the network status to s tInput into the trained SA-EDDPG model and use the model to evaluate each action a t performance to obtain corresponding rewards, and at the same time, the data (s t ,a t ,r t ,s t+1 ) stored in the experience replay pool In, it is used to update the SA-EDDPG model;
[0052] Step 7) Perform VNF placement and traffic routing actions on the underlying physical network that can obtain the highest reward.
[0053] In the SA-EDDPG model,
[0054] The actor network has five layers, namely the input layer, an attention layer, two hidden layers and the output layer. The main actor network μ(s|θ μ ) and the target actor network μ′(s|θ μ′ ) has the same neural network structure, where the input layer is the network state vector s t ; The attention layer converts the state vector s of each node t Transformed into a vector s that considers all node information t '; Both fully connected hidden layers contain 256 neurons to process state information; the output of the output layer is action a t , where the output layer is divided into two parts, used to obtain discrete and continuous action decisions respectively. Specifically, the FC1 output layer is defined to obtain discrete VNF and link placement decisions, i.e. and Define the FC2 output layer to obtain the continuous flow split ratio, that is, The FC1 output layer uses the Sigmoid activation function to obtain the corresponding action value, that is, 0 or 1, and the FC2 output layer directly outputs the original continuous action value. And add noise or limit to get the action value
[0055] The critic network uses DNN to approximate the action value function Q(s,a). It has five layers, namely the input layer, an attention layer, two hidden layers and the output layer. The main critic network Q(s,a|θ Q ) and the target critic network Q′(s,a|θ Q′ ) have the same neural network structure, where the input layer is the current network state s t and the action a output by the actor network t ; The attention layer converts the network state vector s t Convert to s′ t; The two hidden layers contain 256 neurons to process state and action information; the output layer is used to output the value of the action value function Q(s,a); the critic network uses the activation function ReLU to introduce nonlinear features.
[0056] The training process of the SA-EDDPG model includes updating the parameters of four neural networks, specifically:
[0057] For the master critic network Q(s,a|θ Q ), using the TD error in the DQN algorithm to update the parameter θ Q , the loss function of the main critic network is calculated as follows:
[0058]
[0059] Where M is the batch size of sampling, y i is the target value, y i The calculation formula is as follows:
[0060] y i =r i +γQ′(s i+1 ,μ′(s i+1 |θ μ′ )|θ Q′ )
[0061] For the main actor network μ(s|θ μ ), update the parameter θ using the sampling policy gradient μ , the update formula of the main actor network is as follows:
[0062]
[0063] For the target critic network Q′(s,a|θ Q′ ) and the target actor network μ′(s|θ μ′ ), use the soft update method to update their respective parameters, the formula is as follows:
[0064] θ Q′ ←τθ Q +(1-τ)θ Q′
[0065] θ μ′ ←τθ μ +(1-τ)θ μ′
[0066] Among them, the parameter τ<<1.
[0067] The training process of the SA-EDDPG model specifically includes the following steps:
[0068] Step 4-1) Initialize the master critic network Q(s,a|θ) Q ), the main actor network μ(s|θ μ ), target critic network Q′(s,a|θ Q′ ) and the target actor network μ′(s|θ μ′ ) network weight parameters, initialize the experience replay pool
[0069] Step 4-2) In each training round, perform the following steps:
[0070] Step 4-2-1) Reset the network environment, and the agent obtains the initial network state according to the reset environment;
[0071] Step 4-2-2) Initialize the random process and add random noise to each output action;
[0072] Step 4-2-3) For each time slot t in this training round, perform the following iterations:
[0073] Step 4-2-3-1) Based on the self-attention mechanism, the input network state s t Based on the information of neighbor nodes, it is transformed into s t ';
[0074] Step 4-2-3-2) Based on s t ′, the agent uses the formula Select action a t and execute, where is random noise, perform action a t Afterwards, you will receive a reward t At the same time, the network status changes from s t Transition to s t+1 ;
[0075] Step 4-2-3-3) The state transfer data (s) obtained by the interaction between the agent and the environment t ,a t ,r t ,s t+1 ) is stored in the experience replay pool middle;
[0076] Step 4-2-3-4) From the Experience Replay Pool Randomly select M samples to train the SA-EDDPG model;
[0077] Step 4-2-3-5) Update the main network and target network: calculate the target value y i , update the parameters θ of the master critic network using the loss function of the master critic network Q, update the parameters θ of the main actor network using the update formula of the main actor network μ , update the parameters of the target critic network and target actor network based on the soft update formula;
[0078] Step 4-3) When the preset number of training rounds is reached, the training is completed.
[0079] A device for parallel routing optimization of a delay-sensitive service function chain is implemented based on the method described above, comprising:
[0080] The parallel routing problem construction module is used to construct the service function chain parallel routing problem in the edge computing network based on network function virtualization. Considering the real-time network state changes, the Markov decision process is used to model the service function chain parallel routing problem and obtain the MDP model;
[0081] The SA-EDDPG algorithm solving module is used to solve the MDP model using an enhanced deep deterministic policy gradient algorithm based on the self-attention mechanism. The solving process includes two parts: offline training and online operation. The offline training part obtains a trained SA-EDDPG model, and the online operation part determines the optimal routing solution by running the trained SA-EDDPG model and performs parallel routing of the service function chain.
[0082] A device for parallel routing optimization of a delay-sensitive service function chain includes a memory, a processor, and a program stored in the memory. When the processor executes the program, the method described above is implemented.
[0083] A storage medium stores a program thereon, and when the program is executed, the method described above is implemented.
[0084] Compared with the prior art, the present invention has the following beneficial effects:
[0085] (1) This paper proposes an enhanced deep deterministic policy gradient (EDPG) algorithm, which can effectively process both discrete and continuous actions by improving the structure of the DNN in the DDPG algorithm.
[0086] (2) The present invention introduces a self-attention mechanism into the EDDPG algorithm. By calculating the attention value between server nodes in the edge network, the self-attention mechanism enables the intelligent agent in the DRL to focus its attention on more valuable server nodes when performing actions, which can reduce the attention to irrelevant server nodes, thereby reducing unnecessary exploration of the intelligent agent, helping to solve the problem of SFC parallel routing on overly dispersed edge server nodes, accelerating the training and convergence time of DNN, and improving the efficiency of SFC parallel routing.
[0087] (3) The method of the present invention can minimize the resource consumption of the server and the link while satisfying the delay and resource constraints, and has lower delay and smaller resource consumption than other algorithms.
[0088] (4) The present invention takes into account the real-time network status changes and uses Markov decision process (MDP) to model the SFC parallel routing problem. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] Figure 1 is a flow chart of the method of the present invention;
[0090] Figure 2 Schematic diagram of the solution process of the enhanced deep deterministic policy gradient algorithm based on the self-attention mechanism;
[0091] Figure 3 A comparison chart of rewards for different algorithms during training in one embodiment;
[0092] Figure 4 A comparison chart of acceptance rates of different algorithms under different numbers of requests in one embodiment;
[0093] Figure 5 A comparison chart of total resource consumption of different algorithms under different numbers of requests in one embodiment;
[0094] Figure 6 This is a comparison chart of average latency of different algorithms under different numbers of requests in one embodiment;
[0095] Figure 7 This is a comparison chart of average delays of different algorithms under different numbers of nodes in an embodiment. DETAILED DESCRIPTION
[0096] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0097] This embodiment provides a method for optimizing parallel routing of delay-sensitive service function chains. Figure 1As shown, the following steps are included:
[0098] S1: Construct the service function chain parallel routing problem in the edge computing network based on network function virtualization. Considering the real-time network status changes, the Markov decision process is used to model the service function chain parallel routing problem and obtain the MDP model.
[0099] S11: Constructing the SFC parallel routing problem in NFV-based edge computing networks to minimize the resource consumption of servers and links while meeting latency and resource constraints.
[0100] First, for ease of reading, this embodiment provides a parameter explanation table as shown in Table 1, which is universal in this embodiment.
[0101] Table 1 Parameter explanation
[0102]
[0103]
[0104] The underlying physical network is represented as a connectivity graph G = (N, L), where N is the set of server nodes (including various devices with computing capabilities at the edge of the network), and L is the set of physical links connecting two server nodes. For each server node n∈N, use C n Represents its resource capacity, i.e., available computing and storage resources. Similarly, for each link l∈L, use B l and D l They represent the available bandwidth capacity and communication delay respectively.
[0105] Each SFC request is defined as a series of ordered VNFs. Network traffic needs to be routed through the VNFs in the SFC in sequence. For the traffic routed through the SFC, it is assumed that it is divisible and can be divided into I sub-flows, each of which is represented by i. R is used to represent the set of SFC requests arriving in real time. For each SFC request r(F,ε r ,D r )∈R, F represents the set of VNFs required by the SFC request, ε r It is used to indicate the flow rate of traffic passing through the SFC route, D r represents the delay constraint of the SFC request. In addition, for each sub-flow i of the SFC request, the set of VNFs required by sub-flow i is represented as F i ={f1,f2,...,f m ,...,f |i|}, where f mis the mth VNF required by subflow i, and |i| represents the number of VNFs contained in subflow i. Each VNF service instance f m ∈F i There is a resource requirement, use Ci fm To express.
[0106] First, the constraints considered in the SFC parallel routing problem are determined.
[0107] ① Node capacity constraint. Node capacity constraint can ensure that the total resources consumed by the VNFs required by the request do not exceed the resource limit of the server node n to be deployed. Its mathematical form is shown in formula (1):
[0108]
[0109] in, Is a 0-1 variable. Its specific meaning is that if the mth VNFf required by the subflow i of the request r m If the node is placed on physical node n, the value of this variable is 1, otherwise its value is 0.
[0110] ② Link capacity constraint. Link capacity constraint ensures that the total bandwidth required by all requests through link l∈L does not exceed the link’s bandwidth capacity B l , specifically expressed as follows:
[0111]
[0112] In formula (2) is a 0-1 variable. Its specific meaning is that if a link l is used to transmit the subflow i of the request r, the value of the variable is 1, otherwise its value is 0. is a continuous variable, which represents the traffic ratio divided by the sub-flow i of the request r (hereinafter referred to as the traffic split ratio), and its value range is [0,1].
[0113] ③ Delay constraint. For a request r, its delay consists of two parts: the processing delay on the server node and the communication delay when the link transmits traffic. Therefore, the total delay D of request r is total It is expressed as the following formula:
[0114]
[0115] in, Represents the server node n∈N on f m ∈ i. Request r contains I sub-flows, and the delay of each sub-flow i may be different. Therefore, the maximum delay of all sub-flows i is used as the total delay of request r.
[0116] For each request r, if it can be successfully received, then the total delay D of request r is total Cannot exceed its delay limit D r , the specific expression is shown in formula (4):
[0117]
[0118] ④ Placement constraint. The placement constraint ensures that only one physical server is selected to place the mth VNF required by subflow i of request r. In addition, the placement constraint also ensures that all VNFs in subflow i of request r can be served. The specific constraint is expressed as the following formula:
[0119]
[0120]
[0121] The goal of the SFC parallel routing problem is to jointly minimize the resource consumption of the server and the bandwidth consumption of the link.
[0122] For all server resource consumption U N , which can be expressed by formula (7):
[0123]
[0124] For all links, the resource consumption U L , which can be expressed by formula (8):
[0125]
[0126] Then, the optimization problem of joint resource consumption is finally expressed as follows:
[0127]
[0128]
[0129] Where η1 and η2 are the weights of server resource consumption and link bandwidth consumption, respectively. They satisfy η1, η2∈(0,1) and η1+η2=1.
[0130] S12: Considering the real-time network status changes, the SFC parallel routing problem is modeled using Markov decision process (MDP).
[0131] With the above-mentioned modeling of the SFC parallel routing problem, this embodiment continues to describe how to convert it into an MDP model. In general, the MDP model is defined as a five-tuple in and are the state space and action space respectively, is the state transition probability distribution, is the reward function, and γ∈[0,1] is the discount factor for future rewards. In order to solve the real-time network state changes caused by the random arrival and departure of service requests, this embodiment introduces a discrete time period T. The five-tuple related to the SFC parallel routing problem The definition is as follows:
[0132] State space: The state at time t Defined as a vector G t =(C t ,B t ,D t ). Specifically, G t It is used to represent the characteristics of the underlying physical network at time t, where Represents the currently available resources of all server nodes, Represents the current available bandwidth of all physical links, Indicates the delay of all physical links.
[0133] Action space: The action space of an agent is a set Each action a∈A represents VNF placement and traffic routing. Specifically, the action a at time t is defined as in, It is a discrete action, which specifically means whether the m-th VNF of subflow i of request r is placed on server node n at time t. It is also a discrete action, used to indicate whether the subflow i requesting r at time t is routed through the physical link l. It is a continuous action used to represent the traffic split ratio of request r at time t.
[0134] State transition: The state transition of MDP is represented as (s t ,a t ,r t ,s t+1 ), where s t is the network status at the current time t, a t is the action used to process VNF placement and traffic routing in subflow i of request r, r t and s t+1 Execute action a t The immediate reward obtained after and the network state at the next moment t+1. For each state State transition probability p(s t+1 ∣s t ,a t) represents the agent in the network state s t Next, perform action a t After that, the network state transitions to s t+1 probability.
[0135] Reward function: Generally, the goal of reinforcement learning is to maximize the return, that is, to maximize the cumulative discounted reward. In this embodiment, the optimization goal of the SFC routing problem is to minimize the resource consumption of the server and the bandwidth consumption of the link. Therefore, the reward function It should be defined as the negative of the total resource consumption of the server and the link, as shown in formula (10):
[0136]
[0137] S2: The Self-Attention Mechanism-based Enhanced Deep Deterministic Policy Gradient (SA-EDDPG) algorithm solves the MDP model and addresses the online routing problem of SFC, simultaneously handling both discrete and continuous actions. Furthermore, the self-attention mechanism is introduced into the DNN structure to reduce attention to irrelevant server nodes, thereby reducing unnecessary exploration of the agent, accelerating DNN training and convergence time, and improving the efficiency of SFC parallel routing.
[0138] Figure 2 This is the architecture of the SA-EDDPG algorithm proposed in the present invention, which includes two main processes: an offline training process (steps 1 to 4) and an online operation process (steps 5 to 7).
[0139] ① Training process. The offline training process of this architecture mainly includes Figure 2 Steps 1-4. The goal of the training process is to generate a SA-EDDPG model.
[0140] Step 1: The agent interacts with the environment to generate training data. In this process, the environment refers to the NFV-based edge computing network. The agent first observes the current network state s t , and then the network state s is transformed into t Transition to state s′ based on neighbor node information t Next, the agent performs action a based on its current policy t The environment is based on the current state s′ t and the received action a t , update the status to s t+1 , and feeds back a reward signal r to the agent t Then, the agent receives r tUpdate your own strategy to make better decisions at the next moment t+1. The whole process is repeated in this way, and a large amount of training data (s t ,a t ,r t ,s t+1 In order to break the correlation between adjacent samples in the reinforcement learning algorithm and improve the utilization efficiency of training samples, this embodiment introduces the experience replay technology.
[0141] Step 2: Transform the training data (s t ,a t ,r t ,s t+1 ) is stored in the experience replay pool.
[0142] Step 3: When the state transition samples accumulate to a certain number (depending on the size of the experience replay pool), a small batch of samples is randomly selected from the experience replay pool and transmitted to the SA-EDDPG model.
[0143] Step 4: Train the SA-EDDPG model. SA-EDDPG uses a dual actor-critic (AC) network structure, so it contains a total of four neural networks, namely the main actor network μ(s|θ μ ), the main critic network Q(s,a|θ Q ), target actor network μ′(s|θ μ′ ) and the target critic network Q′(s,a|θ Q′ ), where θ μ ,θ Q ,θ μ′ and θ Q′ is a parameter in the neural network. The actor network is responsible for generating VNF placement and traffic routing actions under a given network state, and the critic network is responsible for evaluating the actions generated by the actor network. During the training process, the main actor network updates the parameter θ through the policy gradient method. μ , the main critic network updates the parameters θ by using the gradient descent method based on the temporal difference error (TD error) Q , while the target actor network and target critic network use soft update to update the parameters θ μ′ and θ Q′ .
[0144] ② Operation process. The online operation process of this architecture includes Figure 2 Follow steps 5-7 in the previous step.
[0145] Step 5: Select the SA-EDDPG model that has been trained during the training process to perform online routing of SFC.
[0146] Step 6: Set the network status t Input into the trained SA-EDDPG model, and then use the model to evaluate each action a t To further update the SA-EDDPG model, this embodiment converts the data (s t ,a t ,r t ,s t+1 ) stored in the experience replay pool middle.
[0147] Step 7: Execute the VNF placement and traffic routing actions that maximize reward on the underlying physical network. Furthermore, if the accuracy of a trained model significantly decreases, the model needs to be retrained. Specifically, this is done by copying the previously trained SA-EDDPG model and training it using the newly acquired training data. Then, after the training process is complete, the old model is replaced with the new SA-EDDPG model to improve its performance.
[0148] The SA-EDDPG algorithm proposed in this paper adopts an actor-critic network structure, and both the actor network and the critic network contain a main network and a target network. Therefore, the SA-EDDPG algorithm contains a total of four neural networks. The structures of these neural networks are introduced in detail below.
[0149] The actor network is a policy-based deep reinforcement learning method that uses DNN to learn a deterministic policy and directly uses the policy to generate deterministic actions without sampling from random policies. The input layer of the actor network is the network state vector s t In order to obtain the information of neighbor nodes, this embodiment also introduces an attention layer in the DNN structure of the actor network. Through the attention layer, the state vector s of each node t Can be transformed into a vector s′ considering all node information t The attention layer is followed by two fully connected hidden layers, each containing 256 neurons to process state information. The output layer of the actor network is action a t .
[0150] Different from the output layer in traditional actor networks, this embodiment divides the last layer of the actor network into two parts, which are used to obtain discrete and continuous action decisions respectively. Specifically, the FC1 output layer is defined to obtain discrete VNF and link placement decisions (i.e. and ), define the FC2 output layer to obtain the continuous flow split ratio (i.e. ). The FC1 output layer uses the Sigmoid activation function to obtain the corresponding action value (i.e. 0 or 1). The FC2 output layer directly outputs the original continuous action value Then add noise or limit to get the action value Therefore, the neural network structure of the actor network has five layers, namely the input layer, an attention layer, two hidden layers and the output layer, and the main actor network μ(s|θ μ ) and the target actor network μ′(s|θ μ′ ) have the same neural network structure.
[0151] The critic network is a deep reinforcement learning method based on value function, which uses DNN to approximate the action value function Q(s,a). The input of the critic network is the current network state s t and the action a output by the actor network t Similarly, an attention layer is introduced into the DNN of the critic network, which converts the network state vector s t Convert to s t ′. After the attention layer, there are two hidden layers with 256 neurons, which are used to process state and action information. The output layer of the critic network is used to output the value of the action value function Q(s,a). In addition, the critic network uses the activation function ReLU to introduce nonlinear features, thereby enhancing the processing power of the DNN. Therefore, the neural network structure of the critic network also contains five layers, and the main critic network Q(s,a|θ Q ) and the target critic network Q′(s,a|θ Q′ ) also has the same neural network structure.
[0152] As mentioned above, the SA-EDDPG algorithm contains four neural networks. The operation of the SA-EDDPG algorithm involves the parameter updates of these four neural networks, which are explained below.
[0153] For the master critic network Q(s,a|θ Q ), its parameter θ Q The update utilizes the TD error in the DQN algorithm. Therefore, the loss function of the main critic network is calculated as follows:
[0154]
[0155] Where M is the batch size of the sample, y i is the target value. i The calculation formula is as follows:
[0156]
[0157] For the main actor network μ(s|θμ ), its parameter θ μ The update of utilises the sampling policy gradient method. The update formula of the main actor network is as follows:
[0158]
[0159] Target critic network Q′(s,a|θ Q′ ) and the target actor network μ′(s|θ μ′ ) use soft update to update their parameters. The specific formula is as follows:
[0160]
[0161] Among them, the parameter τ<<1, which makes the parameter update of the target network slow and smooth, and improves the stability of learning.
[0162] Table 2 describes the training process of the SA-EDDPG-based SFC parallel routing method. Specifically, the input of the model is the underlying physical network G and the set of service requests R, and the output is the VNFf m The placement of the and the routing path of the traffic. First, initialize the main critic network Q(s,a|θ Q ), the main actor network μ(s|θ μ ), target critic network Q′(s,a|θ Q′ ) and the target actor network μ′(s|θ μ′ )'s network weight parameters (lines 1-2), and then initialize the experience replay pool (Line 3). In each training round, the network environment should be reset before training begins, and the agent obtains the initial network state according to the reset environment (Line 5). In order to balance the exploration and utilization of the agent, this embodiment introduces a random process in the training process of the SA-EDDPG algorithm. It can be done by adding random noise to each output action (line 6). For each time slot t, the self-attention mechanism is first introduced to transform the network state s t Based on the information of neighbor nodes, it is transformed into s t '(line 8). Then based on s' t The agent uses the formula Select action a t ,in is random noise, which allows the agent to better explore potential optimal actions (line 9). t Afterwards, you will get reward r t At the same time, the network status changes from s t Transition to s t+1Then, the state transition data (s t ,a t ,r t ,s t+1 ) is stored in the experience replay pool (lines 10-11). Afterwards, from Randomly select M samples from to train the SA-EDDPG model (line 12).
[0163] Next, update the main network and the target network. First, calculate the target value y according to formula (12) i (Line 13). Then, the parameters θ of the master critic network are updated using formula (11) Q (Line 14). Specifically, the updating process is to minimize the mean square error between the target value and the Q value of the main critic network using the gradient descent method. Then, the parameters θ of the main actor network are updated using formula (13) μ (Line 15). For the target critic network and target actor network, update their parameters θ according to formula (14) respectively Q′ From formula (14), we can see that, unlike the DQN algorithm that periodically copies parameters from the main network to update the target network, the SA-EDDPG algorithm adopts a soft update method, which makes the update of the target network more stable.
[0164] Table 2 SFC routing algorithm based on SA-EDDPG
[0165]
[0166]
[0167] In this embodiment, the performance of the SA-EDDPG algorithm was evaluated through simulation experiments. First, relevant parameters of the simulation experiment were set. Then, SA-EDDPG was compared with other existing algorithms using various indicators, and the experimental results were analyzed.
[0168] (1) Experimental setup
[0169] In this example, a physical network with 50 nodes is generated using the NetworkX tool, with a link between each pair of nodes. Other network parameters are generated based on previous research. The capacity of each node in the network is randomly selected from the range [30, 100]. The bandwidth of each link in the physical network is randomly assigned to [100, 1000] Mbps, while its latency is randomly assigned to [1, 10] ms. Each SFC contains a varying number of VNFs, ranging from 2 to 8. The number of resources requested by the VNFs follows a uniform distribution of [1, 5]. The bandwidth requirements of the SFC follow a uniform distribution of [10, 50] Mbps.
[0170] All simulation experiments were conducted on a computer equipped with an Intel(R) Core(TM) i7-11700 CPU @ 2.50GHz and an NVIDIA GeForce RTX 3050 GPU. In addition, all experiments in this example were completed using Python 3.7 and PyTorch 1.8. As mentioned above, the deep neural network structure consists of an input layer, an attention layer, two hidden layers, and an output layer, with each hidden layer having 256 neurons. To process both discrete and continuous actions, two different output layers were used in the actor network, with a Sigmoid activation function. Reluctant Unit (ReLU) was used as the activation function for the other layers of the actor network and the critic network. Other experimental parameter settings are shown in Table 3.
[0171] Table 3 Experimental parameters
[0172]
[0173]
[0174] (2) Performance Analysis
[0175] In order to evaluate the performance of the SA-EDDPG algorithm, this embodiment introduces the following two algorithms for comparison.
[0176] SFC-DOP: SFC-DOP is an online SFC orchestration algorithm based on DDPG. It includes VNF placement and traffic routing modules, enabling dynamic SFC orchestration in complex and dynamic networks. Furthermore, it effectively addresses the continuous actions involved in SFC orchestration. However, it does not consider information about other server nodes when placing VNFs.
[0177] A-DDPG: A-DDPG uses a DDPG-based DRL algorithm to solve dynamic VNF placement and traffic routing problems. Unlike SFC-DOP, A-DDPG incorporates an attention mechanism to consider the impact of neighboring node states on VNF placement decisions, accelerating algorithm training and convergence time. However, compared to the SA-EDDPG algorithm of our present invention, it cannot simultaneously handle discrete and continuous actions.
[0178] Figure 3 The reward values for all algorithms during training are shown. As can be seen from the figure, the reward values of each algorithm gradually converge as the number of rounds increases. In particular, it can be observed that the SA-EDDPG algorithm consistently achieves the highest reward and converges faster than the other two algorithms. Specifically, SFC-DOP uses the DDPG algorithm to obtain the optimal policy. However, SFC-DOP does not consider the impact of information from other server nodes on VNF placement, resulting in the lowest reward and the slowest convergence. In comparison, A-DDPG significantly outperforms SFC-DOP in terms of reward value and convergence speed. Thanks to the introduction of an attention mechanism, A-DDPG considers additional state information from neighboring nodes when selecting actions. This allows the agent to focus on more valuable nodes during decision-making, resulting in faster positive rewards during training and faster training and convergence times. However, A-DDPG performs worse than SA-EDDPG because it does not simultaneously consider both discrete and continuous decision actions. As previously mentioned, the SA-EDDPG algorithm proposed in this paper improves the neural network structure, allowing it to simultaneously handle discrete VNF placement decisions and continuous traffic segmentation decisions. As a result, SA-EDDPG agents are more intelligent in dynamic NFV-based edge computing networks.
[0179] Figure 4The request acceptance rates for all algorithms under different request numbers are shown. It can be observed that the request acceptance rates of all three algorithms show a downward trend as the number of requests increases. This is because as the number of service requests increases, the underlying physical network resources gradually become occupied, resulting in the rejection of subsequent requests. More specifically, SFC-DOP's request acceptance rate is significantly lower than the other two algorithms because its reward function does not consider the impact of underlying server and link resources on the SFC routing strategy. In contrast, A-DDPG and the proposed SA-EDDPG algorithm take into account the impact of underlying resources on SFC routing. Furthermore, thanks to the introduction of a self-attention mechanism, A-DDPG and SA-EDDPG algorithms can obtain information about other server nodes, enabling the agent to better capture the dynamic changes in the state of the underlying physical network resources, thereby achieving a higher request acceptance rate. However, SA-EDDPG's request acceptance rate remains higher than A-DDPG's because it can simultaneously handle both discrete and continuous actions in the SFC routing problem, thereby more efficiently utilizing the underlying network resources to accept requests. Compared to the other two algorithms, SA-EDDPG's request acceptance rate increases by 15% and 7%, respectively.
[0180] Figure 5 The total resource consumption of the three algorithms under different request numbers, including server resource consumption and link bandwidth consumption, is shown in the figure. As the number of requests increases, the total resource consumption of the three algorithms also increases. Compared with the other algorithms, SA-EDDPG consumes the least total resources. Specifically, it consumes 38% less than SFC-DOP and A-DDPG, respectively. This is because SA-EDDPG incorporates a resource optimization incentive mechanism into its reward function, which incentivizes the agent to select actions that consume less resources, thereby reducing total resource consumption. In contrast, A-DDPG consumes more total resources than SA-EDDPG because it cannot effectively handle the discrete decision actions in the SFC routing problem, resulting in resource waste. SFC-DOP consumes the most total resources, partly because it does not consider resource optimization when performing SFC routing and partly because, compared with the other two algorithms, it does not fully utilize information from other server nodes, which prevents the agent from effectively adapting to the NFV-based edge computing network environment.
[0181] Figure 6 The average latency of all algorithms under different numbers of requests is shown. As the number of requests increases, the average latency of the three algorithms also increases to a certain extent. Figure 4As shown in the figure, the request acceptance rate decreases as the number of requests increases, which leads to a decrease in the number of accepted requests. Therefore, the growth rate of the average delay of all algorithms has slowed down. Specifically, the average delay of the SFC-DOP algorithm is significantly higher than that of the A-DDPG algorithm and the SA-EDDPG algorithm. This is because the SFC-DOP algorithm does not take into account the status information of neighboring nodes, which increases unnecessary exploration of the agent, resulting in higher delay. By introducing the attention mechanism, the A-DDPG algorithm can make full use of the information of other server nodes, reduce the inefficient exploration of the agent, and thus reduce the average delay. Compared with the other two algorithms, the SA-EDDPG algorithm proposed in the present invention has the lowest average delay, which is 40% and 24% lower than the SFC-DOP and A-DDPG algorithms respectively. This is because the SA-EDDPG algorithm adopts a parallel routing strategy, which can effectively reduce the delay of service requests.
[0182] Figure 7 The average latency of all algorithms for different numbers of physical nodes is shown. The average latency of all three algorithms decreases with increasing node numbers. This is because as the number of physical nodes increases, the network becomes more resource-rich and has sufficient suitable paths to accommodate service requests. Compared to the other two algorithms, the SFC-DOP algorithm has the highest average latency. This is because as the number of nodes increases, the network topology becomes more complex, and the SFC-DOP algorithm ignores the state information of other physical nodes, causing the agent to spend more time exploring, resulting in higher latency. In contrast, the A-DDPG and SA-EDDPG algorithms both have lower average latency than the SFC-DOP algorithm because they both incorporate attention mechanisms, which enable the agent to focus on more valuable nodes when executing actions. Furthermore, the SA-EDDPG algorithm performs parallel routing of SFCs through traffic splitting and repeated VNF instantiation. It also utilizes an improved neural network structure to simultaneously process both discrete and continuous actions, resulting in the lowest average latency. Furthermore, the experimental results demonstrate that the SA-EDDPG algorithm can achieve lower latency while consuming fewer resources.
[0183] This embodiment primarily discusses the SFC parallel routing problem in NFV-based edge computing networks. To meet the needs of latency-sensitive applications in edge computing networks, this paper proposes an optimization method for SFC parallel routing through traffic segmentation and repeated VNF instantiation. Specifically, a model for the SFC parallel routing problem is first constructed. Considering the resource constraints in edge computing networks, the goal is to minimize server and link resource consumption while satisfying latency and resource constraints. To address the dynamic changes in network state and the distributed nature of edge server nodes, a DRL-based SA-EDDPG algorithm is proposed to address this problem. By improving the neural network structure and utilizing a self-attention mechanism, this algorithm can simultaneously handle both discrete and continuous actions in the SFC parallel routing problem, reducing unnecessary exploration by intelligent agents and thus improving the efficiency of SFC parallel routing. Finally, this embodiment conducts extensive simulation experiments to evaluate the performance of the SA-EDDPG algorithm in terms of request acceptance rate, total resource consumption, and latency. Experimental results demonstrate that, compared to existing algorithms, the SA-EDDPG algorithm proposed in this paper effectively improves the acceptance rate of service requests and reduces service request latency and total resource consumption.
[0184] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0185] The above describes in detail the preferred embodiments of the present invention. It should be understood that those skilled in the art can make numerous modifications and variations based on the concepts of the present invention without inventive effort. Therefore, any technical solutions that can be derived by those skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.
Claims
1. A method for parallel routing optimization of delay-sensitive service function chains, characterized in that: The following steps are involved: Construct the service function chain parallel routing problem in the edge computing network based on network function virtualization, consider the real-time network state changes, use Markov decision process to model the service function chain parallel routing problem, and obtain the MDP model; the MDP model is defined as a five-tuple in and are the state space and action space respectively, is the state transition probability distribution, is the reward function, γ∈[0,1] is the discount factor for future rewards, which is defined as follows: State space: The state at time t Defined as a vector G t =(C t ,B t ,D t ), G t Used to represent the characteristics of the underlying physical network at time t, where Represents the currently available resources of all server nodes, Represents the current available bandwidth of all physical links, Indicates the delay of all physical links; Action space: The action space of an agent is a set Each action a∈A represents VNF placement and traffic routing, and the action a at time t is defined as in, It is a discrete action, indicating whether the m-th VNF of subflow i of request r is placed on server node n at time t; It is a discrete action used to indicate whether the subflow i requesting r at time t is routed through the physical link l; It is a continuous action, used to represent the traffic split ratio of request r at time t; State transition: The state transition is represented as (s t ,a t ,r t ,s t+1 ), where s t is the network status at the current time t, a t is the action used to process VNF placement and traffic routing in subflow i of request r, r t and s t+1 Execute action a t The immediate reward obtained after and the network state at the next moment t+1, for each state State transition probability p(s t+1 ∣s t ,a t ) represents the agent in the network state s t Next, perform action a t After that, the network state transitions to s t+1 probability; Reward function: Based on the optimization goal of the service function chain parallel routing problem, the reward function Defined as the negative of the total resource consumption of the server and link: An enhanced deep deterministic policy gradient algorithm based on the self-attention mechanism solves the MDP model. The solution process includes two parts: offline training and online operation. The offline training part obtains a trained SA-EDDPG model. The online operation part determines the optimal routing solution by running the trained SA-EDDPG model and performs parallel routing of the service function chain. The offline training process includes the following steps: Step 1) The agent interacts with the environment to generate training data, where the environment refers to the edge computing network based on network function virtualization. The agent first observes the current network state s t , and the network state s is converted into t Transition to state s based on neighbor node information t ′, the agent performs action a based on its current strategy t , the environment is based on the current state s t ′ and the received action a t , update the status to s t+1 , and feeds back a reward signal r to the agent t , the agent receives r t Update your own strategy to make better decisions at the next moment t+1, and generate training data (s t ,a t ,r t ,s t+1 ); Step 2) The training data (s t ,a t ,r t ,s t+1 ) is stored in the experience replay pool; Step 3) When the state transition samples accumulate to a preset number, a batch of samples is randomly selected from the experience replay pool and input into the SA-EDDPG model; Step 4) Train the SA-EDDPG model, where the input of the SA-EDDPG model is the underlying physical network G and the set of service requests R, and the output is VNFf m The placement and routing paths of traffic, The online operation process includes the following steps: Step 5) Select the SA-EDDPG model trained during the offline training process for online routing of the service function chain; Step 6) Set the network status to s t Input into the trained SA-EDDPG model and use the model to evaluate each action a t performance to obtain corresponding rewards, and at the same time, the data (s t ,a t ,r t ,s t+1 ) stored in the experience replay pool In, it is used to update the SA-EDDPG model; Step 7) Perform VNF placement and traffic routing actions on the underlying physical network that can obtain the highest reward.
2. The method for parallel routing optimization of delay-sensitive service function chains according to claim 1, characterized in that: The network model of the edge computing network is: The underlying physical network is represented as a connectivity graph G = (N, L), where N is the set of server nodes, including various devices with computing capabilities at the edge of the network, and L is the set of physical links connecting two server nodes. For each server node n∈N, use C n Indicates its resource capacity. For each link l∈L, use B l and D l Represent its available bandwidth capacity and communication delay respectively; Each service function chain request is defined as a series of ordered virtual network functions (VNFs). Network traffic needs to be routed through each VNF in the service function chain in turn. For the traffic routed through the service function chain, it is assumed that it can be split and divided into I sub-flows, each sub-flow is represented by i. R represents the set of service function chain requests arriving in real time. For each service function chain request r(F,ε r ,D r )∈R, F represents the set of VNFs required by the service function chain request, ε r Indicates the flow rate of the routing traffic through the service function link, D r Represents the delay limit requested by the service function chain; for each sub-flow i requested by the service function chain, the set of VNFs required by sub-flow i is represented as F i ={f1,f2,…,f m ,...,f |i| }, where f m is the mth VNF required by subflow i, |i| represents the number of VNFs contained in subflow i; each VNF service instance f m ∈F i There is a resource requirement, use To express.
3. The method for parallel routing optimization of delay-sensitive service function chains according to claim 2, characterized in that: The method of modeling the service function chain parallel routing problem using the Markov decision process includes the following steps: S1: Determine the constraints; S11: Determine the node capacity constraint. The node capacity constraint ensures that the total resources consumed by the VNFs required by the request do not exceed the resource limit of the server node n to be deployed, which is expressed as: in, is a 0-1 variable. If the mth VNFf required by subflow i of request r m If it is placed on physical node n, the value of this variable is 1, otherwise its value is 0; S12: Determine a link capacity constraint, which ensures that the total bandwidth required by all requests through the link l∈L does not exceed the link's bandwidth capacity B l , expressed as: in, is a 0-1 variable that takes the value 1 if a link l is used to deliver subflow i of request r, and takes the value 0 otherwise; Is a continuous variable, indicating the traffic ratio divided by the sub-flow i of request r, that is, the traffic split ratio, and its value range is S13: Determine the delay constraint: For a request r, its delay consists of two parts: the processing delay on the server node and the communication delay when the link transmits traffic, that is, the total delay D of request r total Expressed as: in, Represents the server node n∈N on f m ∈i’s processing delay, the maximum delay among all sub-flows i is used as the total delay of request r; For each request r, if it can be successfully received, the total delay D of request r is total Cannot exceed its delay limit D r : S14: Determine a placement constraint that ensures that only one physical server is selected to place the mth VNF required by subflow i of request r, and that all VNFs in subflow i of request r can be served, expressed as: S2: Determine that the optimization goal of the service function chain parallel routing problem is to minimize the joint resource consumption, that is, to minimize the resource consumption of the server and the bandwidth consumption of the link, where Resource consumption of all servers U N Expressed as: Bandwidth consumption of all links U L Expressed as: Then, the optimization objective of joint resource consumption is expressed as: Where η1 and η2 are the weights of server resource consumption and link bandwidth consumption, respectively, satisfying η1, η2∈(0,1) and η1+η2=1.
4. The method for parallel routing optimization of delay-sensitive service function chains according to claim 3, characterized in that: The SA-EDDPG model adopts a dual actor-critic network structure, which contains four neural networks: the main actor network μ(s|θ μ ), the main critic network Q(s,a|θ Q ), target actor network μ′(s|θ μ′ ) and the target critic network Q′(s,a|θ Q′ ), where θ μ ,θ Q ,θ μ′ and θ Q′ is a parameter in the neural network. The actor network is responsible for generating VNF placement and traffic routing actions under a given network state. The critic network is responsible for evaluating the actions generated by the actor network. During the training process, the main actor network updates the parameter θ through the policy gradient method. μ , the master critic network updates the parameters θ by the gradient descent method based on the temporal difference error Q , the target actor network and the target critic network use soft update to update the parameters θ μ′ and θ Q′ .
5. The method for parallel routing optimization of delay-sensitive service function chains according to claim 4, characterized in that: In the SA-EDDPG model, The actor network has five layers, namely the input layer, an attention layer, two hidden layers and the output layer. The main actor network μ(s|θ μ ) and the target actor network μ′(s|θ μ′ ) has the same neural network structure, where the input layer is the network state vector s t ; The attention layer converts the state vector s of each node t Transformed into a vector s that considers all node information t '; Both fully connected hidden layers contain 256 neurons to process state information; the output of the output layer is action a t , where the output layer is divided into two parts, used to obtain discrete and continuous action decisions respectively. Specifically, the FC1 output layer is defined to obtain discrete VNF and link placement decisions, that is, and Define the FC2 output layer to obtain the continuous flow split ratio, that is, The FC1 output layer uses the Sigmoid activation function to obtain the corresponding action value, that is, 0 or 1, and the FC2 output layer directly outputs the original continuous action value. And add noise or limit to get the action value The critic network uses DNN to approximate the action value function Q(s,a). It has five layers, namely the input layer, an attention layer, two hidden layers and the output layer. The main critic network Q(s,a|θ Q ) and the target critic network Q′(s,a|θ Q′ ) have the same neural network structure, where the input layer is the current network state s t and the action a output by the actor network t ; The attention layer converts the network state vector s t Convert to s t ′; The two hidden layers contain 256 neurons to process state and action information; the output layer is used to output the value of the action value function Q(s,a); the critic network uses the activation function ReLU to introduce nonlinear features.
6. The method for parallel routing optimization of delay-sensitive service function chains according to claim 1, characterized in that: The training process of the SA-EDDPG model includes updating the parameters of four neural networks, specifically: For the master critic network Q(s,a|θ Q ), using the TD error in the DQN algorithm to update the parameter θ Q , the loss function of the main critic network is calculated as follows: Where M is the batch size of sampling, y i is the target value, y i The calculation formula is as follows: y i =r i +γQ′(s i+1 ,μ′(s i+1 |θ μ′ )|θ Q′ ) For the main actor network μ(s|θ μ ), update the parameter θ using the sampling policy gradient μ , the update formula of the main actor network is as follows: For the target critic network Q′(s,a|θ Q′ ) and the target actor network μ′(s|θ μ′ ), use the soft update method to update their respective parameters, the formula is as follows: i Q′ ←tth Q +(1-τ)θ Q′ i μ′ ←tth μ +(1-τ)θ μ′ Among them, the parameter τ<<1.
7. The method for parallel routing optimization of delay-sensitive service function chains according to claim 1, characterized in that: The training process of the SA-EDDPG model specifically includes the following steps: Step 4-1) Initialize the master critic network Q(s,a|θ) Q ), the main actor network μ(s|θ μ ), target critic network Q′(s,a|θ Q′ ) and the target actor network μ′(s|θ μ′ ) network weight parameters, initialize the experience replay pool Step 4-2) In each training round, perform the following steps: Step 4-2-1) Reset the network environment, and the agent obtains the initial network state according to the reset environment; Step 4-2-2) Initialize the random process and add random noise to each output action; Step 4-2-3) For each time slot t in this training round, perform the following iterations: Step 4-2-3-1) Based on the self-attention mechanism, the input network state s t Based on the information of neighbor nodes, it is transformed into s t '; Step 4-2-3-2) Based on s′ t The agent uses the formula Select action a t and execute, where is random noise, perform action a t Afterwards, you will receive a reward t At the same time, the network status changes from s t Transition to s t+1 ; Step 4-2-3-3) The state transfer data (s) obtained by the interaction between the agent and the environment t ,a t ,r t ,s t+1 ) is stored in the experience replay pool middle; Step 4-2-3-4) From the Experience Replay Pool Randomly select M samples to train the SA-EDDPG model; Step 4-2-3-5) Update the main network and target network: calculate the target value y i , update the parameters θ of the master critic network using the loss function of the master critic network Q , update the parameters θ of the main actor network using the update formula of the main actor network μ , update the parameters of the target critic network and target actor network based on the soft update formula; Step 4-3) When the preset number of training rounds is reached, the training is completed.
8. A device for parallel routing optimization of a delay-sensitive service function chain, implemented based on the method according to any one of claims 1 to 7, characterized in that: include: The parallel routing problem construction module is used to construct the service function chain parallel routing problem in the edge computing network based on network function virtualization. Considering the real-time network state changes, the Markov decision process is used to model the service function chain parallel routing problem and obtain the MDP model; The SA-EDDPG algorithm solving module is used to solve the MDP model using an enhanced deep deterministic policy gradient algorithm based on the self-attention mechanism. The solving process includes two parts: offline training and online operation. The offline training part obtains a trained SA-EDDPG model, and the online operation part determines the optimal routing solution by running the trained SA-EDDPG model and performs parallel routing of the service function chain.
9. A device for parallel routing optimization of a delay-sensitive service function chain, comprising a memory, a processor, and a program stored in the memory, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A storage medium having a program stored thereon, characterized in that: When the program is executed, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Dynamic service function chain arrangement method and system based on deep reinforcement learning
CN114172937A
Service function chain dynamic reconstruction method in space-air-ground integrated scene
CN115361288A