Context-Aware Service Chain Embedding Method in Space-Air-Ground Integrated Networks

By using graph neural network and deep reinforcement learning technology to embed service chains in the integrated air-space and earth network, the problem of difficulty in adapting to high dynamic characteristics of the existing technology is solved, the system stability and performance improvement is achieved, and the business function chain deployment is optimized.

CN118157746BActive Publication Date: 2025-06-27HARBIN INST OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410358289.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-27
Publication Date
2025-06-27
Estimated Expiration
2044-03-27

AI Technical Summary

Technical Problem

The existing service function chain embedding method is difficult to adapt to the high dynamic characteristics of the space-space integrated network, especially the link instability between the space-based network, the space-based network and the ground base station nodes. It is an offline method and does not support online learning and optimization.

Method used

Deep reinforcement learning technology based on graph neural network and context perception is adopted to realize online update of service chain functional chain (SFC) deployment strategy, adapting to the high dynamic characteristics of the space-space integrated network.

Benefits of technology

Through online dynamic update strategies, we can improve the stability and performance of the system, optimize the success rate of business function chain deployment, reduce resource waste, reduce operational costs, and improve operators' long-term average returns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118157746B_ABST
    Figure CN118157746B_ABST
Patent Text Reader

Abstract

Method for Embedding Service Chain Based on Context Awareness in Space-Air-Ground Integrated Network. The present invention relates to a method for embedding a service chain. The present invention includes the following steps: 1. Construct an SFC embedding description model for the space-air-ground integrated network and analyze dynamic resource constraint conditions; 2. Convert the SFC embedding model into an MDP process and introduce it into the deep reinforcement learning model; 3. Construct a graph neural network and a context awareness mechanism in the deep reinforcement learning framework; 4. Train the deep reinforcement learning model and test the embedding effect; 5. Repeat steps 2-4 until all SFC embeddings are completed. The present invention can realize online update of the SFC deployment strategy by using deep reinforcement learning technology based on graph neural network and context awareness, so as to adapt to the high dynamic characteristics of the space-air-ground integrated network. The present invention belongs to the field of communication technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for embedding a service chain, in particular to a method for embedding a service chain based on context awareness in a space-air-ground integrated network. The present invention belongs to the field of communication technologies. Background Art

[0002] In recent years, the development cost of satellite equipment has been continuously decreasing, enabling the transmission connectivity of communication services to extend to the non-ground part, which can solve problems such as deployment, coverage, and capacity commonly faced by ground networks. At the same time, aerial unmanned aerial vehicles (UAVs) can assist ground and satellite networks in data collection, perform computing tasks, and relay data to other nodes for further processing or aggregation. The construction of a space-air-ground integrated network needs to adapt to the market demand for communication technologies in the past few decades, continuously improve its performance, capacity, reduce latency, and optimize the management of various resources in the network.

[0003] One of the challenges to be faced is that the way consumers use the network has changed nowadays, and a large number of heterogeneous services have emerged, such as large-scale Internet of Things, streaming services (short videos), or telemedicine. To achieve diverse services in a space-air-ground integrated network, it is necessary to rely on network virtualization technology, namely virtual network functions (VNFs). It converts network functions from dedicated hardware into software-based modules, which can significantly reduce the capital expenditure and operating costs required for the space-air-ground integrated network to meet various differentiated services. Each network service is usually implemented in the form of a network service function chain (SFC). The SFC consists of a group of VNFs that are executed in a strict order, and network services can be deployed or deleted on commercial servers of network nodes as needed, making the network service deployment more flexible and agile.

[0004] The existing service function chain embedding methods mainly have the following defects:

[0005] 1. The existing service function chain embedding strategies are difficult to adapt to the high dynamic characteristics of a space-air-ground integrated network, especially the link instability between satellite nodes in the space-based network, UAV nodes in the air-based network, and ground base station nodes. This instability will have a significant impact on system performance;

[0006] 2. The existing service function chain embedding strategies are offline methods, which establish the overall optimal solution obtained through overall optimization after knowing all service requests, and do not support online learning and optimization. Summary of the Invention

[0007] To solve the above problems existing in the prior art, the present invention further proposes a service function chain embedding method based on context awareness in a space-air-ground integrated network. By using graph neural network and context awareness-based deep reinforcement learning technology, the online update of the SFC deployment strategy can be realized, so as to adapt to the high dynamic characteristics of the space-air-ground integrated network.

[0008] The technical solution adopted by the present invention to solve the above problems is as follows:

[0009] The present invention includes the following steps:

[0010] Step 1: Construct an SFC embedding model for the space-air-ground integrated network and analyze dynamic resource constraints;

[0011] Step 101: Construct the motion models of each platform in the space-air-ground integrated network;

[0012] Step 102: Construct an SFC transmission service model in the space-air integrated network;

[0013] Step 103: Evaluate the indicators of the embedding effect during the SFC embedding process;

[0014] Step 2: Convert the SFC embedding model into an MDP process and introduce it into the deep reinforcement learning model;

[0015] Step 201: Extract the SFC requests to be processed.

[0016] Step 202: Represent the network links, configurations, and the current SFC as states in the learning model;

[0017] Step 203: Represent the mapping selection for embedding the SFC as an action in the learning model;

[0018] Step 204: Represent whether the embedding is successful as a reward in the learning model;

[0019] Step 3: Construct a graph neural network and a context awareness mechanism in the deep reinforcement learning framework;

[0020] Step 301: Process the network links and configurations using the graph neural network model;

[0021] Step 302: Process each embedding process using the Seq2Seq model encoder;

[0022] Step 4: Train the deep reinforcement learning model and test the embedding effect,

[0023] Step 401: Use the A3C model to construct a parallel training process.

[0024] Step 5: Repeat steps 2-4 until all SFC embeddings are completed.

[0025] The beneficial effects of the present invention are as follows:

[0026] 1. The service function chain embedding method provided by the present invention can realize the online dynamic update of policies by using graph neural networks and context-aware deep reinforcement learning techniques. This means that it can adapt to the highly dynamic characteristics of the space-air-ground integrated network, such as the link instability between satellite nodes, drone nodes, and ground base station nodes, thereby improving the stability and performance of the system.

[0027] 2. The intelligent method DRL-PTSE used in the present invention utilizes graph neural networks and Seq2Seq models to extract network features, and captures the laws of future topological changes through feedback to intelligently generate SFC deployment strategies. Such an intelligent deployment strategy can optimize the success rate of service function chain deployment, reduce resource waste caused by unstable links, reduce capital expenditure and operating costs, and improve the long-term average revenue of operators. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is a schematic diagram of the scenario for deploying service function chains in the space-air-ground integrated network of the present invention;

[0029] Figure 2 is a learning model of the service chain embedding method based on the combination of graph neural networks and context awareness;

[0030] Figure 3 is a schematic diagram of the space-air-ground integrated network simulation scenario;

[0031] Figure 4 is a graph of the loss change during the model training process;

[0032] Figure 5 is a comparison graph of the change in the SFC acceptance rate after multiple trainings;

[0033] Figure 6 is a comparison graph of the change in the long-term average revenue of SFC after multiple online trainings;

[0034] Figure 7 is a comparison graph of the change in the SFC acceptance rate of different methods in a dynamic scenario;

[0035] Figure 8 is a comparison graph of the change in the long-term average revenue of different methods in a dynamic scenario. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] Embodiment 1: The context-aware service chain embedding method in the space-air-ground integrated network described in this embodiment includes the following steps:

[0037] Step 1: Construct an SFC embedding model for the space-air-ground integrated network and analyze dynamic resource constraints;

[0038] Step 2: Convert the SFC embedding model into an MDP process and input it into the deep reinforcement learning model;

[0039] Step 3: Construct a graph neural network and a context awareness mechanism in the deep reinforcement learning framework;

[0040] Step 4: Train the deep reinforcement learning model and test the embedding effect;

[0041] Step 5: Repeat Steps 2 - 4 until all SFC embeddings are completed.

[0042] The process of this embedding method is as follows:

[0043] Input the current substrate network G S =(N S , E S ); The virtual network G f of the SFC request f=(N f , E f ); Output the placement success flag flag; Action set a f ; Routing set p f . Initialize the action set Routing set Initialize the actor network with the parameters θ obtained from training; Input the initial state f of the virtual network G of the SFC request f into the encoder. For each NFV embedding step k = 1, ···, |f|, execute: Obtain the state S of the physical network G at the current step. Use GCN to extract the feature Z S of G t,k ; Run the decoder to select an action a k ; If revoke all previous embeddings;

[0044] If k == 1: Based on the action a k place on the physical node ; Add the action a k to the set a f ; Update the state S of the physical network G and continue the loop;

[0045] Execute the Dijkstra algorithm to find a route p k between a k-1 and a k ; If the route p k exists: Based on the action a k place Deploy to physical nodes ; According to the route p k Place the virtual link onto the physical link ; Add the action a k to the set a f ; Add the route p k to the set p f ; Update the state of the physical network G S Otherwise: Cancel all previous embeddings. Otherwise: Cancel all previous embeddings.

[0046] Embodiment 2: In the space-air-ground integrated network, the service chain embedding method based on context awareness described in this embodiment includes, in step 1:

[0047] Step 101: Construct the motion models of each platform in the space-air-ground integrated network;

[0048] As shown in the attached Figure 1 space-air-ground integrated network (SAGIN), consider describing the underlying network composed of all communication devices as an undirected graph G S =(N S , E S ), where the set N S contains all physical facility nodes in SAGIN, and the set E S contains all wired and wireless links in the network. This undirected graph can be represented in the form of an adjacency matrix A S . The set N S can be further divided into 3 subsets and representing the space-based network nodes, the space-based network nodes, and the ground-based network nodes of SAGIN respectively. Therefore, there is Through the functions and represent the remaining and maximum node resources provided by the substrate network respectively. Denote the set of node resources provided by the substrate network (such as the number of cores of the central processing unit, the memory size, the external storage space, etc.) as K. Therefore, for the physical node n ∈ N of the substrate network S , its remaining resources can be expressed as Its maximum resources are expressed as For the link resources on the substrate network, we use the functions and to represent the remaining amount and the maximum amount provided respectively. The difference is that the input parameter of this function is the link e ∈ E in the substrate network S . Without loss of generality, the present invention takes the bandwidth B as the link resource in the substrate network. For the link e ∈ E S , its remaining resources can be expressed as Its maximum resource is expressed as

[0049] SAGIN operates in a time-slot mode over a long period. We characterize the change in its structure through snapshots of the network topology within a time slot. Let the integer t represent the current time of the system, which operates in the P-th time unit δ t , so we have t = Pδ t . A three-dimensional Cartesian coordinate system is constructed within the SAGIN service coverage area, and based on this, the motion models of the nodes in the network are described. The space-based network consists of a LEO constellation, and the number of satellites orbiting is rotating in a given orbital plane. Assume the satellite is at a fixed altitude H in the predetermined orbital plane S and moving at a constant speed V S . In the orbital plane, the satellites are evenly spaced in a cycle. We represent the position of the satellite nodes in the space-based network through the function q S (·). For the satellite , its position changing discretely over time can be expressed as:

[0050]

[0051] where represents the initial position of the satellite in the orbit, and represent the components of the constant motion speed V of the satellite on the x-axis and y-axis of the reference coordinate system respectively. Assume the access range for each satellite to build a communication link is R S . The space-based network consists of fixed-wing unmanned aerial vehicles (UAVs) and the communication and computing devices deployed on them, and the number of nodes is S . Assume the UAV is flying at a fixed altitude H within the area . We represent the position of the UAV nodes in the space-based network through the function q A (·). For the UAV A , its position changing discretely over time can be expressed as:

[0052] q A (n) = [x(t), y A (t), H A . A .

[0053] where [x A (·), y A (·), H A represents the flight planned trajectory of the UAV, which consists of a sequence of position coordinates of the UAV with respect to time. Assume the access range for each UAV to build a communication link is RA The ground network consists of ground base stations and communication computing devices deployed on them, and the number of nodes is The base stations in the ground network The positions are modeled using a homogeneous Poisson point process (PPPs), and the function q G (·) represents its position q G (n) = [x, y, 0].

[0054] For the network topology model in SAGIN, it is necessary to assume that the communication characteristics of each link are sufficient to adapt to the link types (VLEO-to-ground, VLEO-to-UAV, UAV-to-ground, etc.) and the movement patterns of the two end nodes. Among them, it is assumed that the topology between satellite nodes in the space-based network is a chain topology on a single orbit. The base stations in the ground network are connected by wire to form a fixed topology, which is constructed using the Waxman random topology model. The various topology changes in SAGIN result in changes in the availability of physical links. For this, we use the previously defined adjacency matrix A S (e S , t) = A S ([n1, n2], t) to represent whether there is a physical link e between nodes n1 and n2 in the base network at time t S . A S (t) will be dynamically updated with the topology changes of SAGIN to reflect the current network topology. For example, the construction of the communication link between a satellite and a ground base station follows an opportunistic access model. When D(q S (n1), q G (n2)) ≤ R S , there is a physical link between satellite node n1 ∈ N S and ground base station n2 ∈ N G , that is, A S ([n1, n2], t) = 1, otherwise A S ([n1, n2], t) = 0.

[0055] Step 102 constructs the SFC transmission service model in the space-air-ground integrated network;

[0056] In a network scenario, a data stream usually needs to pass through multiple network service devices, such as IDS / IPS, firewalls, etc., before finally reaching the destination. A service function chain is an abstraction of an ordered set of service functions, which performs a series of service processes on IP packets, link frames, or data streams on the network based on classification and policies. The basis for implementing service functions is the virtualized network functions provided on the underlying network devices, such as Proxy, IPS, Optimizer, Firewall, Billing, NAT, Transcoder, etc. Each SFC deployed on the underlying network corresponds to a service category and consists of a chained structure of VNFs with different functions. To meet various differentiated requirements of users in SAGIN, service providers deploy a set of SFCs composed of multiple VNFs in sequence, denoted as F. For each SFC request f ∈ F, it is denoted as where denotes the i th -th VNF of the SFC request f. Each SFC request f can be modeled as a weighted directed graph G f =(N f , E f ), where N f denotes the set composed of multiple VNFs of this SFC request, and E f denotes the set of virtual links. The arrival time of this SFC request f, that is, the start service time of the request, is denoted as The request service completion time is The arrivals of all SFC requests in the set F are modeled using a Poisson process. For an SFC request f in the set F, deploying VNFs in the underlying network means requesting resources from the physical facilities in the underlying network. We use the function C f (·) to represent the amount of resources requested by the VNF. Therefore, for n ∈ N f , C f (n)=[C f,1 ,···, C f,k ,···, C f,|K| represents the resource vector requested by this VNF, where C f,k represents the amount of k ∈ K resources requested by the VNF n of the SFC request f. Correspondingly, for the virtual link e ∈ E f , the link resources requested on it can be represented as C f (e)=[B f . The VNF embedding process needs to complete the mapping of the virtual nodes N f in the virtual network G f of the SFC request f to the physical nodes N S in the underlying network G S , that is, N f →NS . Define a group of binary decision variables If it represents the virtual node n f ∈N f , f ∈ F and the physical node n s ∈N S have a mapping at time t, otherwise Since the VNF should be uniquely deployed in a physical node of a substrate network, there are constraints on the mapping relationship The total amount of resources k required for the SFC request f mapped to the physical node n s should not exceed the remaining resources of the physical node, and the constraint needs to be satisfied The VNF embedding process also needs to complete the virtual network G of the SFC request f f in the virtual links E f to the physical links E in the substrate network G S i.e., E S → E f → E S . Define a group of binary decision variables If it represents the virtual link e f ∈ E f , f ∈ F and the physical link e S ∈ E S have a mapping at time t, otherwise The total amount of resources k required for the SFC request f mapped to the physical link e S should not exceed the remaining resources of the physical link, so there is a constraint If the SFC request f is accepted, the path mapped to the physical network should traverse the required VNFs in the order specified in the request. We represent the in - and out - link sets of the physical node n S ∈ N S as I(n S ) and O(n S ), and represent the source and destination VNFs of the virtual link e f respectively, and there is a constraint

[0057]

[0058] Step 103 constructs the metrics for evaluating the embedding effect during the SFC embedding process;

[0059] An important metric that this invention focuses on for the SFC embedding method is the proportion of SFC requests that are accepted by the SAGIN and operate normally among all SFC requests. A certain SFC request f in the set may fail to be deployed due to insufficient allocated resources or network topology changes, and this request cannot be executed. The problem of SFC embedding caused by insufficient resources has been solved through the previous constraints. Regarding the impact of network topology changes on SFC embedding, we define the function Ace(f) to represent the identifier that the SFC request f can operate normally within its life cycle. This function is defined as:

[0060]

[0061] Therefore, the acceptance rate of the SFC service is obtained through calculation as The SAGIN operator uses the "pay-as-you-go" model widely used in the cloud platform to price the provided services for the SFC service. Therefore, the revenue obtained for each SFC request f is defined as:

[0062]

[0063] where μ is the vector composed of the unit prices of various resources in the physical nodes, which corresponds one-to-one with the request resource vector obtained by the function C f (·). Similarly, η is the unit price of resources in the physical link, that is, the unit price of the transmission bandwidth in the link. If the SFC service is interrupted due to topology changes during the life cycle, no revenue will be obtained. Therefore, combined with the above formula, the long-term average revenue of the SAGIN operator over time can be calculated as:

[0064]

[0065] where represents the SFCs whose arrival time is before τ. The goal of this invention is to find an SFC embedding strategy that can maximize the long-term cumulative revenue

[0066] Specific Embodiment 3: The method for embedding a service chain based on context awareness in the space-air-ground integrated network described in this embodiment, step 2 includes:

[0067] Step 201 extracts the SFC requests to be processed.

[0068] Step 202 represents the network links, configurations, and the current SFC as states in the learning model;

[0069] The present invention uses a deep reinforcement learning framework to describe the SFC embedding method, and finally generates an embedding strategy for virtual nodes and virtual links, which is executed by the agent in the framework. The interaction between the agent and the SAGIN environment is defined as a Markov decision process (MDP). We construct the mapping of VNFs in the virtual node sequence of each SFC request f as an MDP with a finite horizon. The agent sequentially selects a physical node n from the substrate network S to place the virtual node n f , until the mapping of each virtual node in f is completed. The state representation in the reinforcement learning framework defines the information that the agent can obtain from the environment and serves as the original input for the upcoming feature extraction phase. In the SFC embedding problem, the state should include the current substrate network status and the processing status of the SFC request. Define the state of SAGIN at the k th th virtual node during the deployment of the current SFC request at time t as is the set of states of the substrate network when deploying this VNF The dimension of each sub-state in this set is The set of states of the SFC when deploying this VNF is

[0070] Step 203 represents the mapping selection for embedding the SFC as an action in the learning model;

[0071] The action of the model is an effective mapping process executed by the agent, in which the virtual network request is mapped to a subset of the substrate network through the SFC embedding strategy. As the scale of nodes and links increases, the number of subgraphs of the substrate network topology grows exponentially. If the entire SFC is directly embedded as one of the subgraphs, the action space of the agent will have a huge computational cost, bringing a huge learning cost to the reinforcement learning framework. Corresponding to the definition of the state space, the embedding process of the SFC is decomposed into the sequential embedding of the virtual node sequence therein. At this time, in each step, the agent only needs to focus on the virtual node that needs to be placed currently. If there is enough idle resource remaining in the physical nodes of the substrate network, this virtual node will be embedded according to the agent's strategy, otherwise it will be forced to transition to the termination state. Therefore, during the deployment of the current SFC request, the action set at the k th th virtual node is defined as

[0072] Step 204 represents whether the embedding is successful as a reward in the learning model;

[0073] In reinforcement learning, an agent needs to continuously receive a reward function from the external environment to improve its performance. We use the reward function to tell the agent how good the relative performance of the current action is. To maximize the estimate of the cumulative discounted reward, the agent may give up the action with the best current reward to obtain better long-term performance. When an SFC request is fully mapped, that is, no resource constraints are violated and the links during the life cycle are stable, this mapping strategy will be considered good, and a positive reward will be returned to strengthen the probability of the current strategy being executed. For this purpose, we set that when the last is deployed for the SFC request f, the agent obtains a reward of

[0074]

[0075] During the middle of the mapping process when it is successfully deployed, the agent can obtain a small reward r t,k = ξRev(f). If the remaining resources cannot meet the SFC requirements or cannot ensure the stability of the links during its life cycle, a negative reward r t,k = -ξRev(f) is obtained to avoid making a failed action again, where ξ is the reward coefficient. During the training process, let the agent search for alternative decisions and control the optimization direction of the SFC embedding strategy through the reward function.

[0076] Specific implementation method 4: The method for service function chain embedding based on context awareness in the space-air-ground integrated network described in this implementation method, step 3 includes:

[0077] Step 301 Process the network links and configurations using a graph neural network model;

[0078] Step 302 Process each embedding process using the Seq2Seq model encoder;

[0079] The model architecture used in the present invention is as Figure 2 shown, and it consists of three parts: 1) Embedding of the physical network; 2) Embedding of the SFC request and 3) Policy generation. To explore the characteristics of the network topology, a graph neural network (GCN) based on semi-supervised learning is used to extract features of the physical network. In each NFV placement step, the current state of the physical network is fed into the GCN layer to learn a new representation matrix where U gcn is the number of units in the GCN layer. The arithmetic operation of GCN is simply formalized as where σ is the activation function and W is the trainable parameter. is an approximate graph convolutional filter similar to the convolutional neural network (CNN). Among them At the same time, there is The base network G that adds self - connections using the re - normalization trick S is the adjacency matrix, where Λ is the identity matrix. To capture the ordered requirements of SFC requests, we use a Seq2Seq model encoder implemented by a gated recurrent unit (GRU) network. It takes an input sequence and generates a hidden state e k . For a given placement step k, the GRU cell takes the sum of the current input and the hidden state e of the previous step k-1 as input, and then outputs the hidden state of the current step The specific calculation process of GRU is as follows:

[0080]

[0081] where r k , z k , represent the reset gate, update gate, and candidate hidden state respectively. {W, V, b} are the parameters of the corresponding cell. σ(·) is the activation function, and the symbol ⊙ represents element - wise multiplication. An attention decoder of the Seq2Seq model is used to generate appropriate actions. The decoder receives the last state of the encoder According to the current decoder state d k and the previous action a k-1 , it generates an output a of size |f| in order, a = {a1, ···, a k , ···, a |f|}.

[0082] where

[0083] To infer a reasonable placement order for SFC requests, a context - based attention mechanism is used to calculate the correlation between the input sequence and the output sequence. The context vector is calculated as a weighted sum of e j as The weight α j for each e k,j is defined as Here, score(d k , e j ) measures the matching degree between the input near position j and the output at position k, and it is calculated as and W a are trainable variables. To obtain the probability distribution of candidate actions, we intuitively combine the current decoder state d k , the current context vector c k and the flattened output Z of the GCN layer t,kConnect them and then convert them into a fully connected layer so that the final output is consistent with the number of physical network nodes. The agent performs an action according to the conditional probability to select a physical node to deploy the current VNF, and the conditional probability is

[0084] where Finally, use the Dijkstra algorithm to find the shortest path connecting a k and a k-1 in the physical network. If there exists a path p k that meets the virtual connection bandwidth requirement and is stable during the life cycle, then the current VNF is successfully placed on a k ; otherwise, the current SFC cannot be placed, and the previously occupied physical resources will be released.

[0085] Specific Embodiment 5: In the integrated space-air-ground network, the service chain embedding method based on context awareness, Step 4 uses the A3C model to construct a parallel training process:

[0086] To accelerate the training speed of the model and improve the robustness of the model, the asynchronous actor-critic algorithm (A3C) is adopted. A3C is a "master-worker" parallel training architecture, which consists of multiple worker agents and a master agent. Each agent contains two networks: 1) The actor network maintains a parameterized placement policy π θ (a k |S t,k ) and generates an action according to the current state; 2) The evaluation network uses the Q value V ω (S t,k , a k ) to evaluate whether the performance of the action is good. After obtaining the global shared parameters from the master agent, each agent u ∈ U initializes its own parameters θ ω and u ω u , explores the environment to collect experiences, and then sends them to the master agent to update the global shared parameters. At each placement step k, each agent selects an action a according to the random policy k , and interacts with the environment to obtain a reward r k and the next state S t,k+1 . The evaluation network evaluates the new state A(a k , S t,k ) = r k +γV ω (S t,k+1 ) - V ω (S t,k) where γ ∈ (0, 1) is the discount factor. After processing a complete SFC embedding, the critic network updates its parameters by minimizing the square of the TD error loss. ε θ represents the learning rate of the actor network, logπ θ (a k |S t,k )A(a k ,S t,k ) represents the cross-entropy loss weighted by the TD error.

[0087] The complete process of training by using the SFC embedding method DRL-PTSE to completely embed all SFCs is as follows:

[0088] Place all SFCs using the A3C parallel training DRL-PTSE method in the system, input the current substrate network G S =(N S ,E S ); the set of SFC requests F; the maximum number of training iterations maxIter; output the SFC embedding policy π θ ;

[0089] Initialize the iteration number iter ← 0; initialize the network parameters θ and ω of the master agent; initialize the respective independent environments of all worker agents U;

[0090] Repeat when iter < maxIter: For each agent u ∈ U, execute: For each SFC request f ∈ F, execute: Synchronize the network parameters θ u ← θ, ω u ← ω;

[0091] Sample the trajectory using the formula Deploy f through the DRL-PTSE method, and use the formula A(a k ,S t,k ) = r k + γV ω (S t,k+1 ) - V ω (S t,k ) to obtain the reward r k and calculate the TD error;

[0092] Update the parameters ω of the critic network using the formula ;

[0093] Adjust the policy π of the actor using the formula ; θ (a k |S t,k ).

[0094] Embodiment Six: To verify the effectiveness of the present invention, we constructed a SAGIN simulation scenario with 100 heterogeneous network nodes, and the area size is 100km * 100km. Within the scope of this area, there are 2 space nodes 1 aerial node and 97 ground base stations The orbital plane height where the space nodes are located is set to H S = 350km, and the moving speed is V S = 8.1km / s, and its communication coverage range is R S = 278km. The fixed height where the aerial node is located is H A = 10km, and its aircraft planned trajectory is generated by interpolating the marked points randomly placed in the area with the simulation time. The number of marked points is randomly generated within the range of [3, 5] with an average distribution probability, and its communication coverage range is R A = 7km. The positions of the ground base stations are randomly generated after being modeled by the homogeneous Poisson point process (PPPs), and the network topology is randomly generated by the Waxman model. An example diagram of a SAGIN scenario after random generation is as shown in the appendix Figure 3 shown, where the dotted line is the flight trajectory of the unmanned aerial vehicle, and the solid line is the sub-satellite point movement trajectory of the satellite. In the simulation of the SAGIN service model, we set the node resources and link bandwidth in the physical network to be evenly distributed from 50 to 100 units. In each simulation, the SFC requests arrive at the system in sequence according to the Poisson process. The service arrival rate is set within the range of [0.1, 0.22] (services per time unit). The total number of SFC requests is set within the range of [600 - 1000]. Specifically, each SFC request consists of 2 - 15 VNFs with different quantities distributed evenly, and its lifetime follows an exponential distribution with an average of 400. The node and link resource requirements of the SFC requests are evenly distributed from 2 to 30

[0095] Appendix Figure 4 describes the changes in the actor loss and critic loss during the training process of this method. In 20,000 training steps, the actor loss and critic loss converge to local optimal values of 0.006 and 0.02 respectively, which indicates that the deep neural network in DRL - PTSE has a good approximation effect

[0096] In the appendix as shown in Figure 5 and the appendix Figure 6In the verification experiments shown, we compared DRL-PTSE with other intelligent methods, namely PSO-VNE, PG-SEQ2SEQ, and MCTS. The above methods are called static benchmark methods, that is, when new service requests arrive, these methods directly perform mapping and scheduling of VNFs in the SFC without combining the dynamic characteristics of the network topology. PSO-VNE is a heuristic method based on particle swarm optimization. PG-SEQ2SEQ uses a continuous decision-making scheme based on reinforcement learning. MCTS is a method based on traditional reinforcement learning that uses Monte Carlo tree search to make SFC placement decisions. From the online learning process in the figure, we can see that the proposed method outperforms the baseline methods in terms of the performance and stability of the acceptance rate and long-term average revenue.

[0097] To further intuitively compare the performance of the present invention during the system operation, taking the complete deployment process as an example, we show Figure 7 in [the specific situation] the change in the SFC request reception rate as the SFC requests are accessed in sequence, Figure 8 and show the change in the long-term average revenue. Since the algorithm proposed by this method optimizes the stability of SFC deployment, the embedding success rate is higher than that of other baseline methods. As time goes by, the SFC acceptance rate rises fastest. The long-term average revenue rises rapidly at first and then reaches equilibrium. Under the condition of a fixed SFC request arrival rate, the revenue reaches a stable state as time goes by until all SFC services end. In the stable state, the proposed DRL-PTSE method improves the average revenue by approximately 45%, 61%, and 141% compared with MCTS, PSO-VNE, and PG-SEQ2SEQ, respectively.

[0098] The above are only the preferred embodiments of the present invention and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the above-disclosed technical content to form equivalent embodiments with equivalent changes within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention and is based on the technical essence of the present invention, any simple modification, equivalent replacement, and improvement made to the above embodiments still fall within the protection scope of the technical solution of the present invention.

Claims

1. A context-aware service chain embedding method in an air-ground integrated network, characterized in that: The method comprises the following steps: Step 1: Construct the SFC embedding model of the air-ground integrated network and analyze the dynamic resource constraints; Step 2: Convert the SFC embedding model into an MDP process and bring it into the deep reinforcement learning model; Step 3: Build graph neural network and context-aware mechanism in deep reinforcement learning framework; Step 4: Train the deep reinforcement learning model and test the embedding effect; Step 5. Repeat steps 2-4 until all SFCs are embedded; The step 1 comprises the following steps: Step 101: Construct the motion model of each platform in the air-ground integrated network, specifically: In the integrated air-ground network, the underlying network composed of all communication devices is described as an undirected graph G S =(N S ,E S ), where the set N S Contains all physical facility nodes in SAGIN, set E S Contains all wired and wireless links in the network. The undirected graph is represented by the adjacency matrix A S The form of representation; set N S Further divided into 3 subsets and They represent SAGIN's air-based network nodes, space-based network nodes, and ground-based network nodes respectively. Through the function and They represent the remaining and maximum node resources provided by the base network respectively; the node resource set provided by the base network is denoted as K; for the physical node n∈N of the base network S , and its remaining resources are expressed as Its maximum resource is expressed as The bandwidth B is regarded as the link resource in the base network; for link e∈E S , and its remaining resources are expressed as Its maximum resource is expressed as SAGIN runs in time slot mode for a long time, and the changes of its structure are characterized by snapshots of the network topology within the time slot. Let the integer t represent the current time of the system, and run in the Pth time unit δ t , we have t = Pδ t ; Construct three-dimensional Cartesian coordinates in the SAGIN service coverage area, and use them as a basis to describe the motion model of nodes in the network; The air-based network consists of a LEO constellation with a number of satellites in orbit Rotating in a given orbital plane; assuming that the satellite At a fixed height H in the predetermined orbital plane S and at a constant speed V S Motion; In the orbital plane, the satellites circulate at equal intervals; through the function q S (·) represents the location of the satellite node in the air-based network. Its position changing discretely over time is expressed as: in represents the initial position of the satellite in orbit, and They represent the constant speed of the satellite V S The components on the x-axis and y-axis of the reference coordinate system; Assume that the access range of each satellite to build a communication link is R S ; The space-based network consists of fixed-wing drones and communication computing equipment deployed on them, with the number of nodes being Assume drone The fixed height H in the area A Flying; through the function q A (·) indicates the location of the drone node in the space-based network. Its position changing discretely over time is expressed as: q A (n)=[x A (t),y A (t),H A ]; where [x A (·),y A (·),H A ] represents the flight plan trajectory of the UAV, which is composed of the position coordinate sequence of the UAV with respect to time; assuming that the access range of each UAV to build a communication link is R A ; The ground-based network consists of ground base stations and communication computing equipment deployed on them, with the number of nodes being Base stations in ground-based networks The location of is modeled using homogeneous Poisson point processes (PPPs) through the function q G (·) indicates its position q G (n) = [x, y, 0]; For the network topology model in SAGIN, use the adjacency matrix A S (e S ,t)=A S ([n1,n2],t) indicates whether there is a physical link e between nodes n1 and n2 in the base network at time t. S ; A S (t) will be dynamically updated as the SAGIN topology changes to reflect the current network topology; Step 102: constructing an SFC transmission service model in the air-space integrated network, specifically: Each SFC deployed on the base network corresponds to a service category, which is composed of a chain structure of VNFs with different functions. A series of SFC sets composed of multiple VNFs in sequence is represented as F. For each SFC request f∈F, it is represented as in Represented as the i-th SFC request f th VNFs; each SFC request f is modeled as a weighted directed graph G f =(N f ,E f ), where N f It represents the set of multiple VNFs requested by the SFC, E f is represented as a set of virtual links; the arrival time of the SFC request f, i.e., the start service time of the request, is represented as The requested service will be completed in The arrival of all SFC requests in set F is modeled using a Poisson process. For an SFC request f in set F, deploying a VNF in the base network is to request resources from the physical facilities in the base network. f (·) represents the amount of resources requested by the nth VNF, for n∈N f , C f (n)=[C f,1 ,···,C f,k ,···,C f,|K| ] represents the resource vector requested by the VNF, where C f,k k∈K represents the amount of resources requested by the VNF of SFC request f; correspondingly, for the virtual link e∈E f , the link resource requested is represented by C f (e) = [B f ]; The VNF embedding process needs to complete the virtual network G of SFC request f f Virtual node N in f To the base network G S Physical node N in S Mapping, that is, N f →N S ; Define a binary decision variable group if It represents virtual node n f ∈N f ,f∈F and physical node n s ∈N S At time t there exists a mapping, otherwise Since VNF should be deployed only in a physical node of a base network, there are constraints on the mapping relationship. Mapped to physical node n s The total amount of resources k required by the SFC request f on the node should not exceed the remaining resources of the physical node, and the constraint The VNF embedding process also needs to complete the virtual network G of SFC request f. f Virtual link E in f To the base network G S The physical link E S Mapping, that is, E f →E S ; Define a binary decision variable group if It indicates the virtual link e f ∈E f ,f∈F and physical link e S ∈E S At time t there exists a mapping, otherwise Mapping to physical link e S The total amount of resources k required by the SFC request f on the physical link should not exceed the remaining resources of the physical link. If the SFC request f is accepted, the path mapped into the physical network should traverse the required VNFs in the order specified in the request; S ∈N S The inbound and outbound link sets are represented as I(n S ) and O(n S ), and They represent virtual links e f The source and destination VNFs have constraints Step 103: construct an index for evaluating the embedding effect during the SFC embedding process, specifically: A certain SFC request f in the set may fail to be deployed due to insufficient allocated resources or changes in network topology, and the request cannot be executed. The problem of SFC embedding caused by insufficient resources has been solved by the previous constraints. In view of the impact of network topology changes on SFC embedding, a function Ace(f) is defined to indicate the normal operation of the SFC request f during its life cycle. The function is defined as: The acceptance rate of SFC business is calculated as SAGIN operators adopt the "pay-as-you-go" model widely used in cloud platforms to price the SFC services provided. The revenue obtained for each SFC request f is defined as: Where μ is a vector of unit prices of various resources in the physical node, and is related to the function C f (·) corresponds to the requested resource vector obtained by one-to-one; η is the unit price of resources in the physical link, that is, the unit price of the transmission bandwidth in the link; if the SFC service is interrupted due to topology changes during the life cycle, no profit will be obtained. Combined with the above formula, the long-term average profit of the SAGIN operator over time is obtained: in represents the SFCs arriving before τ, and finds an SFC embedding strategy that maximizes the long-term average benefit The step 2 includes: Step 201, extracting the SFC request to be processed; Step 202, network links, configurations, and current SFCs are represented as states in the learning model; Step 203, representing the mapping selection embedded in the SFC as an action in the learning model; Step 204: The embedding success or failure is represented as a reward in the learning model.

2. The context-aware service chain embedding method in an air-ground integrated network according to claim 1, characterized in that: Step 3 includes: Step 301, processing the network links and configurations using a graph neural network model; Step 302: Use the Seq2Seq model encoder to process each step of the embedding process.

3. The context-aware service chain embedding method in an air-ground integrated network according to claim 1, characterized in that: In step 4, we use the A3C model to build a parallel training process.

Citation Information

Patent Citations

  • Service function chain dynamic reconstruction method in space-air-ground integrated scene

    CN115361288A