Training method of virtual network embedding strategy and virtual network embedding method
By employing the hierarchical decision-making mechanism and meta-learning training method of graph neural networks, the problems of low search efficiency and insufficient generalization ability in virtual network embedding methods are solved, achieving efficient and flexible virtual network embedding and improving the acceptance rate and cost-benefit ratio of virtual network requests.
Patent Information
- Application Number
- CN202510852632.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-31
AI Technical Summary
Existing reinforcement learning-based virtual network embedding methods suffer from low search efficiency and insufficient generalization ability. In particular, when dealing with virtual network requests of different sizes, it is difficult to achieve a balanced learning of global optimal solutions and cross-size policy knowledge.
By employing a hierarchical decision-making mechanism and meta-learning training method based on graph neural networks, model-independent meta-learning and fine-tuning training are performed through generated trajectory data. Combined with a hierarchical decoder and a compatibility scoring layer, the matching decision between virtual nodes and physical nodes is optimized, thereby improving the search space and generalization ability.
It significantly improves the search efficiency and flexibility of virtual network embedding, enhances the understanding of network status and the evaluation accuracy of node matching, realizes high-quality embedding of virtual network requests of different sizes, and improves the virtual network request acceptance rate and cost-benefit ratio.
Smart Images

Figure CN120880924A_ABST
Abstract
Description
Technical Field
[0001] This disclosure belongs to the field of virtual network technology, specifically relating to a training method for a virtual network embedding strategy and a virtual network embedding method. Background Technology
[0002] Virtual Network Embedding (VNE) refers to embedding Virtual Network Requests (VNRs) into the underlying physical network to meet the network service requirements of multiple tenants while satisfying various resource constraints of the physical network. Each VNR has a different virtual network topology and attributes, customized according to the specific functional requirements of its corresponding tenant.
[0003] Virtual Novel Elementary (VNE) is an NP-hard combinatorial optimization problem. Recently, reinforcement learning (RL) has shown significant potential in effectively solving VNE problems. Current RL-based approaches model the construction process of VNEs as a Markov decision process (MDP). Existing RL-based methods typically employ unidirectional action design from the MDP to sequentially execute the mapping from virtual nodes to physical nodes, and train a single general policy to handle VNEs of different sizes by building various neural network policy models.
[0004] However, current approaches have the following drawbacks: the unidirectional action design of MDPs severely limits the agent's search capabilities, significantly reduces the available action space, hinders the efficiency of exploring the solution space, and may result in the inability to obtain a globally optimal solution. Furthermore, existing methods use ordinary RL methods to train a single general policy, ignoring the unique complexities of VNRs of different sizes. Treating VNRs of different sizes equally in this way poses challenges to achieving balanced learning of policy knowledge across sizes and generalization capabilities across VNRs of different sizes. Summary of the Invention
[0005] This disclosure proposes a virtual network embedding scheme to address the problems of low search efficiency and insufficient generalization ability in existing schemes.
[0006] A first aspect of this disclosure provides a method for training a virtual network embedding strategy, comprising:
[0007] Retrieve historical virtual network requests;
[0008] In the Markov decision process framework, trajectory data is generated through a hierarchical decision-making mechanism of a graph neural network. This includes: encoding the topology and resource characteristics of the virtual and physical networks using a graph neural network; then executing pairing actions between virtual and physical nodes through the hierarchical decision-making mechanism; updating the network state and recording data points containing state, actions, and immediate rewards; iterating until the virtual network embedding is completed to form the trajectory data. The hierarchical decision-making refers to first generating selection decisions for virtual nodes, and then generating selection decisions for physical nodes in response to the selected virtual nodes.
[0009] Then, based on the trajectory data, model-independent meta-learning training and fine-tuning training are performed to generate the virtual network embedding strategy. The meta-learning training stage includes training meta-strategy parameters using trajectory data of virtual network requests of multiple sizes. The fine-tuning training stage includes optimizing the embedding strategy parameters with the meta-strategy parameters as initial values for virtual network requests of untrained sizes, and outputting the optimized embedding strategy parameters as the deployment strategy for virtual network embedding.
[0010] In some embodiments of this disclosure, the method further includes:
[0011] Train a virtual network request prediction model to predict request arrival characteristics within future time periods based on historical virtual network request data;
[0012] The predicted arrival characteristics are used as additional features input to a graph neural network encoder to enhance the state representation in order to generate trajectory data.
[0013] The arrival characteristics include at least one of the following:
[0014] Number of virtual network requests arriving per unit time;
[0015] Resource demand distribution, including the mean and variance of CPU and bandwidth;
[0016] The scale distribution of virtual networks includes the distribution of the number of nodes and the distribution of the number of links.
[0017] In some embodiments of this disclosure, the graph neural network includes a sequentially connected feature construction module, a weighted encoder module, and a hierarchical decoder module, wherein:
[0018] The feature construction module outputs physical network features and virtual network features to the encoder module;
[0019] The encoder module uses the same set of graph convolutional weights to process the dual network features and outputs physical node embeddings and virtual node embeddings to the decoder module.
[0020] The hierarchical decoder module includes an upper-layer policy unit and a lower-layer policy unit.
[0021] The upper-layer strategy unit generates a virtual node selection probability distribution based on virtual node embedding and global physical network representation.
[0022] The lower-level strategy unit generates a physical node selection probability distribution based on physical node embedding, global representation of the virtual network, and selected virtual nodes.
[0023] In some embodiments of this disclosure, the global representation is obtained by graph attention pooling of the corresponding node embeddings.
[0024] In some embodiments of this disclosure, both the upper-layer policy unit and the lower-layer policy unit include a compatibility scoring layer and a shielding Softmax layer.
[0025] The compatibility scoring layer is used to calculate the matching degree between the virtual node and the physical node;
[0026] The shielded Softmax layer is used to filter invalid node selection decisions based on the available resource status of physical nodes and output a probability distribution. The invalid node selection decision refers to an allocation scheme where the resource demand exceeds the available resources of the physical node.
[0027] In some embodiments of this disclosure, the state representation in the Markov decision process framework is defined as:
[0028]
[0029] in, The virtual network request topology and resource requirements are represented at time step t.
[0030] This represents the physical network topology and resource status at time step t.
[0031] The state s t It dynamically reflects the real-time resource distribution of the virtual network and the physical network currently awaiting processing;
[0032] The action a t This is represented as generating a mapping pair between the virtual node and the physical node:
[0033] a t =(n v n p ),
[0034] Where n v Refers to virtual nodes, generated by the aforementioned upper-layer strategy, n p Refers to physical nodes, generated by the underlying strategy;
[0035] The instant reward is defined as:
[0036]
[0037] Where R(s) t a t This indicates that at time step t, the agent is in state s. t And execute action a t The instant reward obtained afterward
[0038] G v This represents a virtual network request, including its topology and resource requirements.
[0039] Indicates virtual network request G v The set of all virtual nodes in the set.
[0040] Indicates the number of virtual nodes.
[0041] R2C(G v ) indicates a request for virtual network G v The calculated benefit-cost ratio, where benefit refers to the revenue generated by serving the virtual network request, and cost refers to the physical resource cost consumed by embedding the virtual network request.
[0042] In some embodiments of this disclosure, the meta-learning training includes iteratively executing inner and outer loops until convergence, outputting a meta-policy π. φ ,in:
[0043] The inner loop includes processing each task. Using its trajectory data D i Calculate gradient updates:
[0044]
[0045] Where, θ i φ represents the task-specific policy parameter for the i-th task, φ represents the meta-policy parameter shared across tasks, and α is a hyperparameter representing the inner loop step size. L(φ) represents the gradient calculation with respect to φ, where L(φ) is the policy optimization objective function.
[0046] The outer loop includes aggregating the losses of multiple tasks based on the following formula to update the meta-policy parameters:
[0047]
[0048] Where φ is the meta-policy parameter, and β is the hyperparameter representing the outer loop step size. It's a mission. It is the policy loss of task i.
[0049] In some embodiments of this disclosure, the fine-tuning training includes:
[0050] Copy the meta-policy parameter φ to initialize the policy parameter θnew of the untrained virtual network request;
[0051] Using the trajectory data requested by the untrained virtual network, the policy parameters are updated only through the inner loop mechanism.
[0052] In some embodiments of this disclosure, the policy optimization objective function L(φ) is defined as:
[0053]
[0054] in, It is the estimate of the dominance function.
[0055] It is the ratio of the probability of actions under the current strategy to that under the old strategy.
[0056] ∈ is the clipping hyperparameter used to constrain r. φ Within the interval [1-∈, 1+∈],
[0057] Clip(·) is the clipping function.
[0058] A second aspect of this disclosure provides a virtual network embedding method, including:
[0059] Obtain historical virtual network requests and train a virtual network embedding strategy based on the method described in the first aspect of this disclosure;
[0060] Receive a virtual network request, and in response to the request, perform virtual network embedding based on the policy;
[0061] Based on the virtual network embedding, update the state and resource configuration of the target network, and repeat the above steps.
[0062] In summary, the training method and virtual network embedding method of the virtual network embedding strategy provided in the embodiments of this disclosure, by adopting hierarchical decision-making MDP modeling, simultaneously explore both the virtual node space and the physical node space, enabling the RL agent to explore a broader search space, significantly improving search efficiency, and thus enhancing the quality and flexibility of VNE solution. By simultaneously considering the matching degree of virtual and physical nodes, it discovers high-quality embedding schemes with better global coordination that are difficult to achieve with traditional unidirectional decision-making, thereby effectively improving VNR acceptance rate and cost-benefit ratio. Meanwhile, the GNN-based policy network, through its powerful graph representation learning ability and specially designed compatibility scoring layer, enhances the understanding of network state and the accuracy of node matching potential evaluation. Combined with the masked softmax mechanism, it ensures the effectiveness of decision-making and training efficiency. Furthermore, the introduction of Meta-RL training endows the model with excellent generalization and adaptability, enabling it to master general embedding knowledge across different VNR scales and quickly adapt to new, unseen VNR sizes with a very small number of samples, which is crucial for handling dynamic, heterogeneous VNR streams. Furthermore, by integrating multidimensional predictive features into the state representation of a graph neural network, this disclosure enables reinforcement learning agents to not only optimize the current state's benefits but also satisfy future resource constraints when making decisions, thereby overcoming the short-sighted decision-making limitations of traditional RL. Attached Figure Description
[0063] The features and advantages of this disclosure will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the scope of this disclosure in any way.
[0064] Figure 1 This is a diagram illustrating the problem of virtual network embedding.
[0065] Figure 2 This is the training framework for the VNE strategy provided in this disclosure;
[0066] Figure 3 This is a flowchart illustrating a training method for a virtual network embedding strategy according to some embodiments of the present disclosure;
[0067] Figure 4 This is a flowchart illustrating a virtual network embedding method according to some embodiments of the present disclosure;
[0068] Figure 5 This is a performance comparison test result of our method and mainstream virtual network embedding algorithms on WX500. Detailed Implementation
[0069] In the following detailed description, numerous specific details of this disclosure are set forth by way of example in order to provide a thorough understanding of the relevant disclosure. However, it will be apparent to those skilled in the art that this disclosure may be practiced without these details. It should be understood that the terms “system,” “apparatus,” “unit,” and / or “module” used in this disclosure are a method of distinguishing different parts, elements, sections, or components at different levels in a sequential arrangement. However, these terms may be replaced by other expressions if they can achieve the same purpose.
[0070] It should be understood that when a device, unit, or module is referred to as being "on," "connected to," or "coupled to" another device, unit, or module, it may be directly connected to or coupled to, or communicate with, other devices, units, or modules, or there may be intermediate devices, units, or modules present, unless the context explicitly indicates otherwise. For example, the term "and / or" as used in this disclosure includes any one and all combinations of one or more of the associated listed items.
[0071] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure. As shown in this specification and claims, unless the context clearly indicates otherwise, words such as "a," "an," "an," and / or "the" do not specifically refer to the singular and may include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified features, integrals, steps, operations, elements, and / or components, and such expressions do not constitute an exclusive list, in which other features, integrals, steps, operations, elements, and / or components may also be included.
[0072] Referring to the following description and accompanying drawings, these and other features and characteristics, operating methods, functions of related structural elements, combinations of parts, and economics of manufacture of this disclosure can be better understood, wherein the description and drawings form part of the specification. However, it is clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this disclosure. It is understood that the drawings are not drawn to scale.
[0073] Various structural diagrams are used in this disclosure to illustrate various variations of embodiments according to this disclosure. It should be understood that the preceding or following structures are not intended to limit this disclosure. The scope of protection of this disclosure is defined by the claims.
[0074] Network virtualization is a technology that effectively separates network infrastructure from the user services it provides. It allows a single physical network infrastructure to be divided into multiple independent, isolated virtual networks. These virtual networks can be configured, managed, and used independently. It can improve resource utilization and reduce deployment costs, and is widely used in various fields such as 5G networks, cloud computing, and edge computing. A significant challenge facing network virtualization is the Virtual Network Embedding (VNE) problem, an NP-hard combinatorial optimization problem. Figure 1 This is a diagram illustrating a virtual network embedding problem. (Example) Figure 1 As shown, in the VNE problem, it is necessary to effectively embed VNRs into the underlying physical network to meet the network service requirements of multiple tenants while satisfying various resource constraints. Each Virtual Network Request (VNR) has different topology and attributes, and is customized according to the specific functional requirements of its corresponding tenant. Effectively allocating resources to these VNRs is crucial for improving service quality and the revenue of Internet service providers.
[0075] like Figure 1 As shown, in a real-world network system, continuously arriving user service requests are represented as VNRs. These VNRs are mapped onto the physical network managed by the ISP, called VNEs. The physical network is constructed as a weighted undirected graph. Where N p It is a set of physical nodes, L p It is a set of physical edges. Each physical node n p ∈N p Equipped with various resource capacities in It is a set of node resource types, for each physical link. p ∈L p Has bandwidth capacity B(l) p In this paper, we consider multi-dimensional node resources, including central processing units (CPUs), storage resources, and graphics processing units (GPUs). Similarly, each VNR is modeled as a weighted undirected graph. Where N v It is a virtual node set, E v It is a virtual edge set, d v This indicates the lifecycle of the VNR. Once the VNR is accepted, it will be maintained. v Each time slot. Each virtual node Represents a system with multiple resource requirements Virtual machines, each virtual link Indicates bandwidth requirement B(l) v ).
[0076] The optimization objective of VNEs is to minimize the embedding cost of allocating VNRs to the physical network, thereby improving request acceptance rates and ISP profits. To evaluate the quality of the solution, the revenue-to-cost ratio (R2C) is defined as a key metric as follows:
[0077]
[0078] Among them REV(G) v ) indicates VNR G v Income, COST (G v ) represents the embedding cost generated by the solution.
[0079] Solving VNE problems is challenging, involving addressing complex optimization challenges, dynamic changes, and real-time requirements. First, the VNE solution space is vast, encompassing a wide range of combinations between VNRs and the underlying physical network. Therefore, a comprehensive exploration of this broad solution space is crucial for identifying high-quality solutions. Second, the integration of different VNR topologies and their associated resource requirements is dynamic due to specific user needs. Different sized VNRs further exhibit varying degrees of complexity, rendering a single, general strategy insufficient to effectively address the inherent variability in this scenario. Third, the time sensitivity of network systems demands the rapid provision of solutions to meet operational requirements.
[0080] Recently, machine learning techniques have been used to optimize the VNE (Virtual Novelty Provider) solution process, leading to faster and more efficient solutions. Among these, reinforcement learning (RL) has demonstrated significant potential in effectively solving VNE problems. RL models the solution construction process for each VNR (Virtual Novelty Provider) as a Markov decision process (MDP). Existing RL-based methods typically employ unidirectional action design in MDPs (i.e., sequentially executing the mapping from virtual nodes to physical nodes) and build policy models for various neural networks, training a single general policy to handle VNRs of different sizes.
[0081] However, existing RL-based VNE methods still suffer from several significant problems. First, the unidirectional action design of MDPs severely limits the agent's search capabilities, significantly reducing the available action space and hindering the efficiency of exploring the solution space. Second, existing methods employ conventional RL to train a single general policy, ignoring the unique complexities of VNRs of different sizes. Treating VNRs of different sizes equally in this way presents challenges for achieving balanced learning of policy knowledge across sizes and generalization ability across different VNRs. Third, if we directly train multiple policies for VNRs of different sizes, we will encounter problems such as slow adaptation to unseen distributions and learning imbalance. These problems inevitably affect the overall performance.
[0082] In view of this, this disclosure provides a flexible and generalizable RL-based VNE solution scheme, aiming to improve the search and generalization capabilities of RL-based VNE solution frameworks, while achieving rapid adaptation to unseen distributions, thereby improving network virtualization performance. The VNE solution scheme includes a VNE policy training framework (such as...) Figure 2 (as shown) and a VNE solution method based on the training framework.
[0083] Figure 3 This is a flowchart illustrating a training method for a virtual network embedding strategy according to some embodiments of the present disclosure. In some embodiments, the training method for the virtual network embedding strategy is executed by a training server, which is used to train the virtual network embedding strategy. The method includes the following steps:
[0084] S310, retrieve historical virtual network requests.
[0085] S320, within the Markov decision process framework, trajectory data is generated through a hierarchical decision-making mechanism of a graph neural network, including: encoding the topology and resource characteristics of the virtual and physical networks using a graph neural network; then executing pairing actions between virtual and physical nodes through the hierarchical decision-making mechanism; updating the network state and recording data points containing state, actions, and immediate rewards; iterating until the virtual network embedding is completed to form the trajectory data. The hierarchical decision-making refers to first generating selection decisions for virtual nodes, and then generating selection decisions for physical nodes in response to the selected virtual nodes.
[0086] This disclosure models the solution construction process for each virtual network request as a Markov decision process based on bidirectional mapping actions, thereby achieving joint selection of virtual and physical nodes. Specifically, in each decision time step t, the environmental state s is observed. t Then, the agent takes action a according to its strategy π. t ~π(·|st Then, the environment will provide a reward R(s). t a t And according to the transition probability function, it is transformed into a new state s. t+1 ~P(s) t a t During the interaction, the trajectory stores state-action pairs and corresponding rewards. Specifically,
[0087] The state represents the network system state at a specific decision timestamp t, including the current VNR and physical network status, i.e.
[0088] An action is defined as selecting a pair of virtual nodes to be placed and a hosted physical node, denoted as 'a'. t =(n v n p ),in and This bidirectional mapping action introduces increased flexibility and significantly expands the search space, allowing for a more comprehensive exploration of possible solutions.
[0089] State transition refers to changing the virtual node n v Placed on physical node n p And the process of routing virtual links. Based on the selected bidirectional mapping action, the environment attempts to route virtual node n v Placed on physical node n p If the node is successfully placed, link routing will be performed according to a breadth-first search algorithm, which finds the link from n. p To host n v The virtual node's neighboring physical nodes are determined by the shortest physical path that meets the bandwidth requirements. If node placement and link routing are successful, the available resources of the physical network will be updated according to the VNR requirement. Otherwise, the current VNR is rejected.
[0090] The reward design is used to guide the RL agent in learning solution strategies, and is defined as follows:
[0091]
[0092] Implicit rewards are designed to encourage successful placement (using) Rewards) and penalties for failure (using) (Penalty). Once the VNR embedding is complete, it will return. As a reward.
[0093] The strategy is described by the parameter θ, which is based on the given state s. t Select action a t Characterized as a probability distribution:
[0094] πθ (a t |s t )=P(a t |s t )
[0095] The discount factor λ = ∈ (0,1) balances the importance of immediate rewards and future rewards.
[0096] In general, the optimization objective of RL is to maximize the expected return over time step T, i.e., the cumulative discount reward:
[0097]
[0098] Some embodiments of this disclosure are based on graph neural networks to specifically perform the pairing action between virtual nodes and physical nodes. For example... Figure 2 As shown: The graph neural network described in this disclosure includes at least a feature builder, a residual graph neural network encoder, and a hierarchical decoder. Specifically:
[0099] Feature constructor:
[0100] Used to determine the current state (current physical network condition) and VNR processing status The feature builder constructs the encoder's feature inputs. It considers the resource requirements of various nodes in both physical and virtual networks, and aggregates bandwidth resource requirements into node features.
[0101]
[0102] in, These are characteristics of a virtual network. For each virtual node, its characteristics include not only various types of node resource requirements (such as CPU, storage, GPU, etc.), but also the bandwidth information of the virtual links connected to it. (e.g., maximum, average, and total bandwidth requirements). Additionally, it includes a flag bit. This is used to indicate whether the virtual node has been embedded.
[0103] These are physical network characteristics. Similarly, for each physical node, its characteristics include multiple types of available node resources. Aggregated available link bandwidth information and a selection flag. This is used to indicate whether the physical node has been selected as the host.
[0104] Residual graph neural network encoder:
[0105] This is used to encode features from both virtual and physical networks into latent representations. First, an initial node representation is obtained using a Multilayer Perceptron (MLP). and Then, multiple Graph Convolutional Networks (GCN) layers are used as Graph Neural Network (GNN) modules to obtain the embedded representations of virtual and physical nodes. and
[0106]
[0107] Layered decoder :
[0108] It breaks down complex pairing decisions into two simpler, interdependent steps by employing a two-layer strategy:
[0109] The upper-level ranking strategy πH is responsible for selecting which virtual node is prioritized. It first obtains the global representation of the physical network through graph attention pooling. Unlike simple mean pooling, graph attention pooling assigns different importance weights to different nodes in the graph based on the attention mechanism, thereby generating a more representative and expressive graph-level global representation. Specifically, for any node j in the graph, its attention weight a... j The final graph representation G is calculated using a learnable attention network and normalized by the Softmax function, and is the representation of all nodes z. j The weighted sum. Its calculation process can be expressed as:
[0110]
[0111] Where W is a learnable weight matrix and a is a learnable attention vector. This yields the global representation of the physical network. Then, a compatibility scoring network (CSN, usually an MLP) is used to calculate the representation of each virtual node that was not placed. With this global representation The matching degree. The compatibility scoring layer aims to address the shortcomings of traditional methods in feature interaction and dynamic pairwise probability distribution generation. It represents virtual nodes. As a query, physical node representation As the key, calculate the pairwise compatibility score. Probability of generation:
[0112] • Lower-level placement strategy πL: Select a virtual node in the upper-level strategy. (It is represented as) After that, the strategy is responsible for deciding which physical node to place it on. It also uses the Graph Attention Pooling (GAP) mechanism to aggregate the global representation of the virtual network. Then calculate the representation of each available physical node. With virtual network context information ( and The degree of matching,
[0113] Probability of generation and placement:
[0114]
[0115] This disclosure describes a method for obtaining placed nodes during training by using random sampling based on probability, to allow for thorough exploration. During inference, a greedy algorithm is used to select the physical node with the highest probability to place the current virtual node, thereby obtaining a high-quality solution.
[0116] S330, then perform model-independent meta-learning training and fine-tuning training based on the trajectory data to generate the virtual network embedding strategy. The meta-learning training stage includes training meta-strategy parameters using trajectory data of virtual network requests of multiple sizes. The fine-tuning training stage includes optimizing the embedding strategy parameters with the meta-strategy parameters as initial values for virtual network requests of untrained sizes, and outputting the optimized embedding strategy parameters as the deployment strategy for virtual network embedding.
[0117] This disclosure treats VNRs of different sizes as different tasks and formalizes them as following a distribution. This study investigates multiple Markov Decision Processes (MDPs). Model-Agnostic Meta-Learning (MAML) is employed as the meta-reinforcement learning method. MAML aims to facilitate the learning of meta-policies, enabling rapid fine-tuning on new tasks with a small number of training samples to improve adaptability and generalization.
[0118] The training method consists of two phases: First, during the meta-learning process, an inner and outer loop are iteratively executed to derive a meta-policy π with cross-task general knowledge. φ Secondly, during the fine-tuning process, task-specific experience is used to fine-tune the meta-policy, and a set of size-specific policies θ is obtained through an inner loop. i .
[0119] In the inner loop, the meta-policy π φ According to specific tasks And a limited amount of task-specific trajectory data D i Update to suit the corresponding tasks:
[0120]
[0121] Where, θ i φ represents the task-specific policy parameter for the i-th task, φ represents the meta-policy parameter shared across tasks, and α is a hyperparameter representing the inner loop step size. Let L(φ) represent the gradient calculation with respect to φ, where L(φ) is the policy optimization objective function.
[0122] In some embodiments of this disclosure, the policy optimization objective function L(·) follows the Proximal Policy Optimization (PPO) algorithm:
[0123]
[0124] in, This indicates the estimated advantage of taking action. It is the current strategy π φ Compared to the previous strategy The ratio between them. The clip function restricts r with the hyperparameter ∈. φ Improve the stability of policy updates within the range of [1-∈, 1+∈].
[0125] In the outer loop, the goal is to find a meta-policy π. φ It can learn the balancing policy knowledge required for VNR of different scales and demonstrates excellent generalization ability in different tasks, enabling it to quickly learn the optimal policy for each specific task:
[0126]
[0127] Update ψ using the meta-learning rate β based on the meta-gradient. The meta-gradient is calculated as the average gradient of the updated task-specific policy.
[0128]
[0129] Where φ is the meta-policy parameter, and β is the hyperparameter representing the outer loop step size. It's a mission. It is the policy loss of task i.
[0130] This training method effectively balances policy knowledge learning across different tasks and improves generalization ability for VNRs of different sizes.
[0131] Some embodiments of this disclosure also make VNE decisions more forward-looking by predicting future VNR trends and resource requirements, thereby optimizing long-term resource utilization.
[0132] Specifically, it integrates VNR demands decided in historical windows to construct input features. Then, a graph neural network is used to predict the size distribution and resource demands of VNRs that may arrive in the future. Next, the predicted future VNR load information (such as the expected total CPU demand and hotspots for link bandwidth demand in the next k time steps) is integrated into the state representation of the reinforcement learning agent. Within meta-learning frameworks such as MAML, the meta-policy learned by the agent must not only adapt to the current task (current VNR) but also utilize predicted information to learn an embedding policy that better balances current and future demands. During the inner loop update, the task-specific policy θ... i It optimizes based on the state containing predictive information. When updating the meta-policy φ in the outer loop, it evaluates the average performance of this forward-looking decision across multiple tasks (including different predictive scenarios). Shifting from passively accepting VNRs to proactively planning resources effectively avoids future resource bottlenecks caused by short-sighted decisions, improves long-term VNR acceptance rates and the overall utilization efficiency of physical network resources, and provides decision support for network operators in capacity planning and upgrades.
[0133] Figure 4 This is a flowchart illustrating a virtual network embedding method according to some embodiments of the present disclosure. In some embodiments, the virtual network embedding method is performed by a virtual network embedding server, the virtual network embedding server being deployed with... Figure 3 The virtual network embedding strategy trained by the method described in s310-S330. The method includes the following steps:
[0134] S410: Obtain historical virtual network requests and train a virtual network embedding strategy based on the methods described in S310-S330.
[0135] S420, Receive a virtual network request, and in response to the request, perform virtual network embedding based on the policy.
[0136] S430, based on the virtual network embedding, update the state and resource configuration of the target network, repeating S410-S430.
[0137] To further verify the effectiveness of this method in large-scale network systems, an embodiment of this disclosure compares the performance of this method with mainstream virtual network embedding algorithms in large networks. Following previous work, the comparative experiment is based on a random Waxman topology (named WX500) containing 500 nodes and nearly 13,000 links. In this comparative experiment, the size distribution of the virtual network embedding (VNE) is increased to a uniform distribution of 2 to 20, and the arrival rate η of the virtual network request (VNR) is set to 3. For training the method, meta-learning is performed in the first 40 simulations, followed by fine-tuning with 10 simulations. The verification results are as follows: Figure 5 As shown. Figure 5 In this document, this method is labeled as FlagVNE.
[0138] Depend on Figure 5 As can be seen, FlagVNE and NRM-VNE achieved the best and worst performance, respectively. Compared with other baseline algorithms, FlagVNE demonstrates a significant performance advantage, proving its effectiveness in large-scale network systems.
[0139] In summary, the training method and virtual network embedding method of the virtual network embedding strategy provided in the embodiments of this disclosure, by adopting hierarchical decision-making MDP modeling, simultaneously explore both the virtual node space and the physical node space, enabling the RL agent to explore a broader search space, significantly improving search efficiency, and thus enhancing the quality and flexibility of VNE solution. By simultaneously considering the matching degree of virtual and physical nodes, it discovers high-quality embedding schemes with better global coordination that are difficult to achieve with traditional unidirectional decision-making, thereby effectively improving VNR acceptance rate and cost-benefit ratio. Meanwhile, the GNN-based policy network, through its powerful graph representation learning ability and specially designed compatibility scoring layer, enhances the understanding of network state and the accuracy of node matching potential evaluation. Combined with the masked softmax mechanism, it ensures the effectiveness of decision-making and training efficiency. Furthermore, the introduction of Meta-RL training endows the model with excellent generalization and adaptability, enabling it to master general embedding knowledge across different VNR scales and quickly adapt to new, unseen VNR sizes with a very small number of samples, which is crucial for handling dynamic, heterogeneous VNR streams. Furthermore, by integrating multidimensional predictive features into the state representation of a graph neural network, this disclosure enables reinforcement learning agents to not only optimize the current state's benefits but also satisfy future resource constraints when making decisions, thereby overcoming the short-sighted decision-making limitations of traditional RL.
[0140] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding descriptions in the foregoing device embodiments, and will not be repeated here.
[0141] Although the subject matter described herein is provided in the general context of execution on a computer system in conjunction with an operating system and applications, those skilled in the art will recognize that other implementations can also be executed in conjunction with other types of program modules. Generally, program modules include routines, programs, components, data structures, and other types of structures that perform specific tasks or implement specific abstract data types. Those skilled in the art will understand that the subject matter described herein can be practiced using other computer system configurations, including handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframes, etc., and can also be used in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules may reside on both local and remote memory storage devices.
[0142] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0143] It should be understood that the specific embodiments described above are merely illustrative or explanatory of the principles of this disclosure and do not constitute a limitation thereof. Therefore, any modifications, equivalent substitutions, improvements, etc., made without departing from the spirit and scope of this disclosure should be included within the protection scope of this disclosure. Furthermore, the appended claims are intended to cover all variations and modifications falling within the scope and boundaries of the appended claims, or equivalent forms of such scope and boundaries.
Claims
1. A training method for a virtual network embedding strategy, characterized in that, include: Retrieve historical virtual network requests; In the Markov decision process framework, trajectory data is generated through a hierarchical decision-making mechanism of a graph neural network. This includes: encoding the topology and resource characteristics of the virtual and physical networks using a graph neural network; then executing pairing actions between virtual and physical nodes through the hierarchical decision-making mechanism; updating the network state and recording data points containing state, actions, and immediate rewards; iterating until the virtual network embedding is completed to form the trajectory data. The hierarchical decision-making refers to first generating selection decisions for virtual nodes, and then generating selection decisions for physical nodes in response to the selected virtual nodes. Then, based on the trajectory data, model-independent meta-learning training and fine-tuning training are performed to generate the virtual network embedding strategy. The meta-learning training stage includes training meta-strategy parameters using trajectory data of virtual network requests of multiple sizes. The fine-tuning training stage includes optimizing the embedding strategy parameters with the meta-strategy parameters as initial values for virtual network requests of untrained sizes, and outputting the optimized embedding strategy parameters as the deployment strategy for virtual network embedding.
2. The method according to claim 1, characterized in that, The method further includes: Train a virtual network request prediction model to predict request arrival characteristics within future time periods based on historical virtual network request data; The predicted arrival characteristics are used as additional features input to a graph neural network encoder to enhance the state representation in order to generate trajectory data. The arrival characteristics include at least one of the following: Number of virtual network requests arriving per unit time; Resource demand distribution, including the mean and variance of CPU and bandwidth; The scale distribution of virtual networks includes the distribution of the number of nodes and the distribution of the number of links.
3. The method according to claim 1, characterized in that: The graph neural network includes a sequentially connected feature construction module, a weighted encoder module, and a hierarchical decoder module, wherein: The feature construction module outputs physical network features and virtual network features to the encoder module; The encoder module uses the same set of graph convolutional weights to process the dual network features and outputs physical node embeddings and virtual node embeddings to the decoder module. The hierarchical decoder module includes an upper-layer policy unit and a lower-layer policy unit. The upper-layer strategy unit generates a virtual node selection probability distribution based on virtual node embedding and global physical network representation. The lower-level strategy unit generates a physical node selection probability distribution based on physical node embedding, global representation of the virtual network, and selected virtual nodes.
4. The method according to claim 3, characterized in that: The global representation is obtained by graph attention pooling of the corresponding node embeddings.
5. The method according to claim 3, characterized in that: Both the upper-layer strategy unit and the lower-layer strategy unit include a compatibility scoring layer and a shielding Softmax layer. The compatibility scoring layer is used to calculate the matching degree between the virtual node and the physical node; The shielded Softmax layer is used to filter invalid node selection decisions based on the available resource status of physical nodes and output a probability distribution. The invalid node selection decision refers to an allocation scheme where the resource demand exceeds the available resources of the physical node.
6. The method according to claim 1, characterized in that... : The state representation in the Markov decision process framework is defined as follows: in, The virtual network request topology and resource requirements are represented at time step t. This represents the physical network topology and resource status at time step t. The state s t It dynamically reflects the real-time resource distribution of the virtual network and the physical network currently awaiting processing; The action a t This is represented as generating a mapping pair between the virtual node and the physical node: a t =(n v ,n p ), Where n v Refers to virtual nodes, generated by the aforementioned upper-layer strategy, n p Refers to physical nodes, generated by the underlying strategy; The instant reward is defined as: Where R(s) t a t This indicates that at time step t, the agent is in state s. t And execute action a t The instant reward obtained afterward G v This represents a virtual network request, including its topology and resource requirements. Indicates virtual network request G v The set of all virtual nodes in the set. Indicates the number of virtual nodes. R2C(G v ): For virtual network request G v The calculated benefit-cost ratio, where benefit refers to the revenue generated by serving the virtual network request, and cost refers to the physical resource cost consumed by embedding the virtual network request.
7. The method according to claim 1, characterized in that, The meta-learning training includes iteratively executing the inner and outer loops until convergence, outputting the meta-policy π. φ ,in: The inner loop includes processing each task. Using its trajectory data D i Calculate gradient updates: Where, θ i φ represents the task-specific policy parameter for the i-th task, φ represents the meta-policy parameter shared across tasks, and α is a hyperparameter representing the inner loop step size. L(φ) represents the gradient calculation with respect to φ, where L(φ) is the policy optimization objective function. The outer loop includes aggregating the losses of multiple tasks based on the following formula to update the meta-policy parameters: Where φ is the meta-policy parameter, and β is the hyperparameter representing the outer loop step size. It's a mission. It is the policy loss of task i.
8. The method according to claim 7, characterized in that, The fine-tuning training includes: Copy the meta-policy parameter φ to initialize the policy parameter θnew of the untrained virtual network request; Using the trajectory data requested by the untrained virtual network, the policy parameters are updated only through the inner loop mechanism.
9. The method according to claim 7, characterized in that, The objective function L(φ) for strategy optimization is defined as: in, It is the estimate of the dominance function. It is the ratio of the probability of actions under the current strategy to that under the old strategy. ∈ is the clipping hyperparameter used to constrain r. φ Within the interval [1-∈, 1+∈], Clip(·) is the clipping function.
10. A virtual network embedding method, characterized in that, include: Obtain historical virtual network requests and train a virtual network embedding strategy based on the method described in any one of claims 1-9; Receive a virtual network request, and in response to the request, perform virtual network embedding based on the policy; Based on the virtual network embedding, update the state and resource configuration of the target network, and repeat the above steps.