Multi-threaded virtual network mapping and performance adjustment method based on entropy weight method
Through the multi-threaded virtual network mapping method based on the entropy weight method, combined with service quality perception and multi-threaded asynchronous advantage actor and critic algorithm, the problem of unbalanced mapping results and single service quality constraints in the virtual network mapping is solved, and efficient service quality assurance and resource utilization adjustment are achieved.
Patent Information
- Application Number
- CN202310386539.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-12
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-04-12
AI Technical Summary
The mapping results in virtual network mapping are unbalanced and the service quality constraints are single. The existing algorithm ignores the time sensitivity of QoS, resulting in unbalanced service quality of network traffic.
The multi-threaded virtual network mapping method based on the entropy weight method is adopted to classify virtual network requests through service quality perception, differentiated service quality relaxation coefficients are set, and the multi-threaded asynchronous advantage actor critic algorithm is used to train node mapping strategies to evaluate the network and perform node and link mapping. The preliminary mapping results were analyzed by entropy weight method, and the best relaxation ratio evaluation score was obtained, and the service acceptance rate, resource utilization rate and QoS loss rate were dynamically adjusted.
It realizes that on the basis of ensuring the success rate of requests, provides service quality assurance at different levels, dynamically adjusts the service acceptance rate, resource utilization rate and QoS loss rate, and improves the efficiency and service quality of virtual network mapping.
Smart Images

Figure CN116405385B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network mapping, and in particular to a multi-threaded virtual network mapping and performance adjustment method based on an entropy weight method. Background Art
[0002] In a network virtualization environment, network terminal users' demands for applications are presented in the form of virtual network requests (VNRs). SP is responsible for sending VNRs. Its purpose is to hope that InP can allocate enough underlying network resources to meet the demand.
[0003] The application of machine learning (ML) algorithms in virtual network embedding (VNE) has also achieved great results. Machine learning algorithms can process large amounts of data collected at any time, automatically learn statistical signals in data processing, and make analysis or predictions. In the early days of the Internet, the traditional Internet model could only be a "best effort" delivery method because designers did not pay attention to the use of differentiated quality of service (QoS). Therefore, in the later operation of the Internet, a large number of data packets were often lost, the Internet delay was large, transmission program errors and throughput imbalance occurred. VNE algorithms with differentiated quality of service (QoS) may have high real-time requirements and are therefore time-sensitive. Existing algorithms usually ignore this problem. About 90% of network traffic is generated based on QoS-sensitive applications. QoS has a wide range of application processes, such as IP calls (VoIP), real-time network and video conferences (Skype, WebEx) and online games. Although the popularity has increased, different applications have different sensitivities to QoS. In order to cope with frequently changing network requirements, different mapping schemes should be set for different types of network requests to dynamically solve the virtual network mapping problem.
[0004] Reinforcement learning is a classic application in computer knowledge application, showing great advantages in dealing with complex problems. Reinforcement learning uses agents to interact with the real environment through "trial and error" to seek the best method, so as to achieve the goal of maximizing profits through agents. Existing QoS-aware VNE algorithms only consider latency, so the QoS metric index is low; and each virtual network request is treated uniformly, resulting in the lowest request reception and the highest resource utilization.
[0005] Existing reinforcement learning algorithms combined with VNE usually need to use a training strategy evaluation network like training neural networks during model training. Error back propagation (BP) practice calculation is currently the most widely used neural network training calculation, but Policy Gradient often uses round updates to reduce the uncertainty generated when traditional policy gradient schemes and neural networks are combined, and various deep decision gradient schemes (such as DDPG, SVG, etc.) also introduce empirical reasons to reduce the correlation between training results. However, the experience replay method faces two difficulties: 1. Each information exchange process between agents and the environment must consume a lot of memory and computing power. 2. The experience replay method also requires agents to complete learning through an off-policy method, and the information generated by the old policy can be updated through the off-policy method. Summary of the invention
[0006] The present invention provides a multi-threaded virtual network mapping and performance adjustment method based on an entropy weight method, so as to solve the problems of unbalanced mapping results and single service quality constraint in virtual network mapping.
[0007] The embodiment of the present invention provides a multi-threaded virtual network mapping and performance adjustment method based on an entropy weight method, comprising the following steps: obtaining multiple virtual network requests, grading the multiple virtual network requests according to service quality perception, obtaining multiple levels of virtual network requests, and setting differentiated service quality relaxation coefficients for the multiple levels of virtual network requests; inputting the multiple levels of virtual network requests into a trained node mapping strategy evaluation network for node and link mapping, and obtaining preliminary mapping results of the multiple levels of virtual network requests, wherein the node mapping strategy evaluation network uses the resource conditions of the underlying physical nodes, the virtual network requests, and the mapping data between the underlying physical nodes and the virtual network request nodes as training data, and is trained using a multi-threaded asynchronous dominant actor critic algorithm; analyzing the preliminary mapping results using the entropy weight method to obtain the optimal relaxation ratio evaluation scores of the multiple levels of virtual network requests, and using the highest score among the optimal relaxation ratio evaluation scores as the differentiated service quality relaxation coefficients of the virtual network requests at each level.
[0008] Optionally, in one embodiment of the present invention, before performing node and link mapping on the multiple levels of virtual network requests, the method further includes: performing a temporary link pruning operation on the trained node mapping policy evaluation network to remove physical links in the trained node mapping policy evaluation network that do not meet the virtual link requirements, and replying to the node mapping policy evaluation network after the mapping is completed.
[0009] Optionally, in one embodiment of the present invention, the node mapping strategy evaluation network includes an actor network and a critic network, the actor network and the critic network both include a convolutional layer, a fully connected layer and an output layer, and the actor network and the critic network share the same architecture except for the output layer; wherein the input of the convolutional layer of the actor network and the critic network is the state matrix of the resource status of the underlying physical nodes and the virtual network request, the output of the convolutional layer is the suitability of each underlying physical node for the virtual network request after considering its own attributes, the output of the fully connected layer is the probability distribution of each underlying physical node being mapped by the virtual network request, the output of the output layer of the actor network is a parameterized strategy, the output of the critic network is a parameterized estimated value, and the node mapping strategy evaluation network obtains a preliminary mapping result of the virtual network request based on the parameterized strategy and the parameterized estimated value.
[0010] Optionally, in one embodiment of the present invention, the method further includes: screening the underlying physical nodes that carry virtual mapping resources, extracting multiple node characteristic attributes of the screened underlying physical nodes to obtain state characteristic values, normalizing and screening the state characteristic values, and using the state characteristic values to construct a state matrix of the underlying physical nodes.
[0011] Optionally, in one embodiment of the present invention, in the node mapping strategy evaluation network, an advantage function is used instead of a state-action-value function, wherein the advantage function is:
[0012]
[0013] Among them, s is the estimated state, a is the action, is the state function that estimates the cumulative payoff in state s, is the state-action-value function.
[0014] Optionally, in one embodiment of the present invention, training the actor network based on the advantage function includes:
[0015]
[0016] Among them, β is the decay parameter, α is the actor network learning rate, and π θ (s t ,a t ) is a parameterized strategy, is the advantage function, H(.) is the entropy of each strategy step, θ is the gradient of the parameter set, is the gradient vector;
[0017] The critic network is trained by a temporal difference method, including:
[0018]
[0019] Among them, θ v is the critic network value function, α′ is the critic network learning rate, is the critic network gradient vector, r t is the reward at time t, γ is the discount rate, is the state function that estimates the cumulative benefit in the next state of state s, is the state function for estimating the cumulative return under state s.
[0020] Optionally, in one embodiment of the present invention, when the node mapping strategy evaluation network is trained, the actor network and the critic network are trained in parallel.
[0021] Optionally, in one embodiment of the present invention, grading the multiple virtual network requests based on service quality perception includes: considering the business needs of multiple networks and grading the multiple virtual network requests according to differentiated service quality requirements in actual different virtual network requests.
[0022] Optionally, in one embodiment of the present invention, the preliminary mapping result is analyzed by an entropy weight method to obtain optimal relaxation ratio evaluation scores of the virtual network requests of the multiple levels, and the highest score of the optimal relaxation ratio evaluation scores is used as the differentiated service quality relaxation coefficient of the virtual network requests of each level, including:
[0023] Extracting the node resource utilization rate and the link resource utilization rate in the preliminary mapping result as positive indicators, and extracting the delay loss rate, jitter loss rate and packet loss rate loss rate in the preliminary mapping result as negative indicators;
[0024] Convert the negative indicators into positive indicators, and standardize the dimensions of the positive indicators and the negative indicators, so as to convert the original decision matrix into a standardized matrix;
[0025] The information entropy of each indicator is calculated according to the definition of information entropy in the entropy weight method, and the weight of each indicator is calculated based on the information entropy. The optimal relaxation ratio evaluation scores of the virtual network requests of the multiple levels are obtained based on the weights and the standardized matrix, and the highest score among the optimal relaxation ratio evaluation scores is used as the differentiated service quality relaxation coefficient of the virtual network requests of each level.
[0026] Optionally, in one embodiment of the present invention, the preliminary mapping results are analyzed by entropy weight method, and then the method further includes: updating the ratio of the weights of delay loss rate, jitter loss rate, and packet loss rate in the reward function of the node mapping strategy evaluation network according to the weights to update the service quality loss rate, wherein the reward function of the node mapping strategy evaluation network is obtained by the sum of network utilization and service quality achievement rate.
[0027] The multi-threaded virtual network mapping and performance adjustment method based on the entropy weight method of the embodiment of the present invention studies the VNE problem with different QoS requirements, and designs differentiated QoS VNE algorithms from three aspects: VNE cost, network bandwidth and latency according to the functional requirements of different users. The QoS index is taken into account in the virtual network request, and different levels of service quality assurance are provided for VNRs with different QoS requirements, thereby ensuring the service quality of VNR on the basis of ensuring the request success rate. After obtaining the preliminary mapping results, the TOPSIS entropy weight method is used to conduct a comprehensive analysis of the preliminary mapping results of the virtual network request, so as to achieve a balance between dynamic adjustment of the service acceptance rate, resource utilization and QoS loss rate.
[0028] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0030] Figure 1 A flowchart of a multi-threaded virtual network mapping and performance adjustment method based on an entropy weight method provided according to an embodiment of the present invention;
[0031] Figure 2 A schematic diagram of a node mapping strategy evaluation network construction provided according to an embodiment of the present invention;
[0032] Figure 3 A schematic diagram of a node mapping strategy evaluation network structure provided according to an embodiment of the present invention;
[0033] Figure 4 A schematic diagram of an agent gradient strategy process provided according to an embodiment of the present invention;
[0034] Figure 5 A schematic diagram of network performance considering different QoS indicators provided according to an embodiment of the present invention;
[0035] Figure 6A schematic diagram of changes in different indicators under 1.1-2 times QoS relaxation provided according to an embodiment of the present invention. DETAILED DESCRIPTION
[0036] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0037] The following describes the multi-threaded virtual network mapping and performance adjustment method based on the entropy weight method of the embodiment of the present invention with reference to the accompanying drawings. In view of the problem of unbalanced mapping results and single service quality constraints in the virtual network mapping mentioned in the background technology center, the present invention provides a multi-threaded virtual network mapping and performance adjustment method based on the entropy weight method, in which multiple virtual network requests are graded according to the service quality perception, and differentiated service quality relaxation coefficients are set for the graded virtual network requests, and multiple levels of virtual network requests are input into the trained node mapping strategy evaluation network for node and link mapping, and the preliminary mapping results of multiple levels of virtual network requests are obtained. The preliminary mapping results are analyzed by the entropy weight method to obtain the optimal relaxation ratio evaluation scores of multiple levels of virtual network requests, and the highest score in the optimal relaxation ratio evaluation score is used as the differentiated service quality relaxation coefficient of each level of virtual network requests. Thus, the problem of unbalanced mapping results and single service quality constraints in virtual network mapping is solved.
[0038] Specifically, Figure 1 The present invention provides a flowchart of a multi-threaded virtual network mapping and performance adjustment method based on an entropy weight method according to an embodiment of the present invention.
[0039] like Figure 1 As shown, the multi-threaded virtual network mapping and performance adjustment method based on the entropy weight method includes the following steps:
[0040] In step S101, multiple virtual network requests are obtained, the multiple virtual network requests are graded according to service quality perception to obtain multiple levels of virtual network requests, and differentiated service quality relaxation coefficients are set for the multiple levels of virtual network requests.
[0041] An embodiment of the present invention provides a differentiated service quality-aware virtual network mapping algorithm, which takes QoS indicators into account in virtual network requests and provides different levels of service quality assurance for VNRs with different QoS requirements, thereby ensuring the service quality of VNRs while ensuring the request success rate.
[0042] Optionally, in one embodiment of the present invention, multiple virtual network requests are graded based on service quality perception, including: considering the business needs of multiple networks and grading multiple virtual network requests according to differentiated service quality requirements in actual different virtual network requests.
[0043] Specifically, virtual network requests are layered according to service quality perception, and the business needs of three types of networks, BE (Best-effort), AF (Expedited Forwarding), and EF (Assured Forwarding), are considered to classify the QoS requirements in different VNRs according to the actual situation.
[0044] Set differentiated weights for virtual network requests and proportionally increase the tolerance for latency, jitter, and packet loss.
[0045] In order to achieve a balance between QoS guarantee and acceptance rate, the QoS relaxation ratio γ and charging ratio η are set, and the QoS index (D ν , J v , P v ) is relaxed proportionally to (γD ν , γJ v ,γP v ), then the link mapping satisfies:
[0046] D s ≤γD ν ,J s ≤γJ ν ,P s ≤γP ν
[0047] Among them, D ν , J v , P v They are delay, jitter and packet loss rate respectively.
[0048] As shown in Table 1, this is the result of grading according to the actual QoS requirements in different VNRs.
[0049] Table 1
[0050] QoS Level application bandwidth Latency Jitter Lost bag 1 Online Games 05Mbp 50ms 10ms 01% 2 Streaming Communications 2Mbps 100ms 10ms 0.1% 3 Voice Call 0.5Mbp 150ms 30ms 1% 4 Media Broadcast 4Mbps 200ms 50ms 1% 5 High bandwidth applications 5Mbps - - - 6 Low bandwidth general business 0.5Mbp - - -
[0051] In step S102, multiple levels of virtual network requests are input into the trained node mapping strategy evaluation network for node and link mapping to obtain preliminary mapping results of multiple levels of virtual network requests, wherein the node mapping strategy evaluation network uses the resource status of the underlying physical nodes, the virtual network requests and the mapping data between the underlying physical nodes and the virtual network request nodes as training data, and is trained using a multi-threaded asynchronous advantage actor-critic algorithm.
[0052] In the past, DRL training relied on powerful graphics devices (such as GPUs). This invention will adopt a general asynchronous concurrent reinforcement learning architecture: the Asynchronous Advantage Actor-Critic (A3C) algorithm, which can be carried on a multi-core CPU, and the waiting delay of the learning agent of experience sampling is greatly shortened. Experiments have also confirmed that the virtual network mapping algorithm using the asynchronous advantage AC can increase the request acceptance and resource utilization in the underlying physical network system at the same time complexity.
[0053] like Figure 2 As shown in the figure, before virtual node mapping, the node mapping strategy evaluation network is trained by a multi-threaded asynchronous advantage actor-critic algorithm. The reward of the mapping training algorithm is obtained by the sum of network utilization and service quality achievement rate. The training data of the node mapping strategy evaluation network includes the resource status of the underlying physical nodes of each link, the virtual network request, and the mapping data of each virtual network request node. When mapping, node mapping is performed first, and link mapping is completed according to the results of node mapping.
[0054] Optionally, in one embodiment of the present invention, before performing node and link mapping on multiple levels of virtual network requests, the method also includes: performing a temporary link pruning operation on the trained node mapping policy evaluation network to remove physical links in the trained node mapping policy evaluation network that do not meet the virtual link requirements, and replying to the node mapping policy evaluation network after the mapping is completed.
[0055] In order to reduce the risk of overfitting and improve the link mapping success rate, the link model is simplified and the links between nodes that do not meet the QoS requirements are temporarily deleted. That is, the physical links that do not meet the virtual link requirements are pruned before link mapping through pruning operations, thereby temporarily forming a new underlying network, so that each down-hop link in the network meets the bandwidth requirements of the virtual link, increasing resource utilization and improving the link mapping success rate. The underlying network is also restored to the initial network after the mapping is completed, and this operation is repeated when a new virtual network arrives.
[0056] Optionally, in one embodiment of the present invention, the node mapping strategy evaluation network includes an actor network and a critic network, both of which include a convolutional layer, a fully connected layer and an output layer, and the actor network and the critic network share the same architecture except for the output layer; wherein the input of the convolutional layer of the actor network and the critic network is the state matrix of the resource status of the underlying physical nodes and the virtual network request, the output of the convolutional layer is the suitability of each underlying physical node for the virtual network request after considering its own attributes, the output of the fully connected layer is the probability distribution of each underlying physical node being mapped by the virtual network request, the output of the output layer of the actor network is a parameterized strategy, the output of the critic network is a parameterized estimated value, and the node mapping strategy evaluation network obtains a preliminary mapping result of the virtual network request based on the parameterized strategy and the parameterized estimated value.
[0057] like Figure 3 As shown in the figure, the specific structure of the node mapping strategy evaluation network is shown. The feature vector of the state matrix of the resource status of the underlying physical node and the virtual network request are input into the convolution layer. The convolution layer outputs a representation of the suitability of the underlying physical node for the virtual network request node to be mapped after considering its own attributes. Then, a fully connected layer (softmax) is used to output a probability distribution of the number of physical nodes, indicating the possibility of each node being used in the next mapping step with the node mapping probability after convolution. During the training process, the parameters of the node mapping strategy evaluation network are optimized by the asynchronous dominant actor critic algorithm.
[0058] The node mapping strategy evaluates the network training model, models the virtual node mapping as a QoS-aware Markov strategy process, and introduces reinforcement learning technology to learn an optimal node mapping strategy. The Markov strategy process contains four elements: action, state, strategy and reward. The present invention maps a physical node selected from the virtual node as an action. The state is a description of the current environment, and the state will also change as the strategy changes after the action is taken. In this algorithm, the node feature state matrix extracted from the underlying network is regarded as the current state. The strategy is based on the conditional probability distribution of the node mapping under the current physical network state, and is obtained by evaluating the network state according to the node mapping strategy. The reward is reflected by the difference between resource utilization and QoS loss, and the environment is based on the reward given to the strategy after the action is given.
[0059] During the training process, the advantage function is used instead of the simple state-action-value function to reduce the variance of the training experience. Secondly, a parallel training scheme is adopted to improve the sampling efficiency.
[0060] Optionally, in one embodiment of the present invention, in the node mapping strategy evaluation network, an advantage function is used instead of a state-action-value function, wherein the advantage function is:
[0061]
[0062] Among them, s is the estimated state, a is the action, is the state function that estimates the cumulative payoff in state s, is the state-action-value function.
[0063] Optionally, in one embodiment of the present invention, the method also includes: screening the underlying physical nodes that carry virtual mapping resources, extracting multiple node characteristic attributes of the screened underlying physical nodes to obtain state characteristic values, normalizing and screening the state characteristic values, and using the state characteristic values to construct a state matrix of the underlying physical nodes.
[0064] It can be understood that the underlying physical nodes that carry virtual mapping resources are selected from the underlying physical nodes and classified. First, the nine characteristic attributes extracted from the underlying physical nodes are used to describe their status. In order to facilitate their training, the characteristic attributes need to be normalized, and the vector of available resources of each physical node is convolved, and then the special sample data is removed, so that the state matrix of the underlying physical nodes composed of all physical nodes can be obtained.
[0065] Specifically, when performing feature extraction, node attributes such as the degree of proximity between nodes, the centrality of the remaining available resources of nodes, and the sum of the adjacent bandwidths of nodes can be used as feature extraction attributes before node mapping.
[0066] The state matrix contains the physical state of the underlying nodes, including node computing resources, link resources, and the characteristics between nodes. After convolution, the state matrix is used to output a probability distribution of the number of physical network nodes through a softmax operation. The A3C algorithm is then used to optimize the parameters of the neural network and finally output the mapping probability. The larger the probability, the better the effect of the node as a virtual network mapping.
[0067] Specifically, the node mapping strategy evaluation network is initialized, and multiple rounds of iterative training are performed through the A3C algorithm. Each round of iterative training will form a new mapping result according to the change of the underlying state, and at the same time form a new evaluation value. The evaluation value can affect the direction of the next iteration and converge to the expected value. Among them, the process of one round of iterative training is: each iteration is a mapping result formed according to the resource state of the current physical node. After the environment executes the policy sampling action generated by the learning agent, it sends the corresponding reward signal to the agent to complete a single sampling process. Using the advantage function instead of a simple state and action value function can reduce the variance of the training experience, parallel training, and improve the sampling rate. The node mapping strategy evaluation network constructs two networks: 1. Actor network. Generates embedded strategies. 2. Critic network generates estimates under different states to help calculate the advantage function. The two networks share a similar architecture except for the output layer. After each successful iteration, the reward and gradient calculation process will be performed again. When all node mappings in the virtual network request are completed, the training is terminated.
[0068] Optionally, in one embodiment of the present invention, training the actor network based on the advantage function includes:
[0069]
[0070] Among them, β is the decay parameter, α is the actor network learning rate, and π θ (s t ,a t ) is a parameterized strategy, is the advantage function, H(.) is the entropy of each strategy step, θ is the gradient of the parameter set, is the gradient vector;
[0071] The critic network is trained via a temporal difference method, including:
[0072]
[0073] Among them, θ v is the critic network value function, α′ is the critic network learning rate, ▽ θv is the critic network gradient vector, r t is the reward at time t, γ is the discount rate, is the state function that estimates the cumulative benefit in the next state of state s, is the state function for estimating the cumulative return under state s.
[0074] One round of iterative training includes training an actor network and a critic network. Figure 3 As shown, the former is used to generate a set of parameterized strategies π θ , which is used to generate a set of parameterized estimates And helps calculate the advantage function. The gradient return of the traditional policy gradient algorithm is:
[0075]
[0076] in, is the state-action value function, which can be used to predict the strategy π in state S θ The estimated long-term benefits of the resulting action. This produces a large variance, making the training system unstable. When a conditional value is relatively high, that is, no matter how the result of the condition is operated, Q is at a relatively high level, which will ignore the difference between the operations. Therefore, the difference between the expected return and the average state value described by the advantage function can be expressed as follows:
[0077]
[0078] in, is a state function that estimates the cumulative reward in state S. The advantage function represents the advantage of the current action a compared to the "average action" derived by the corresponding policy in a specific state s, and can reduce the variance during training without changing the bias. With the help of the advantage function, the update of the actor network can follow the following policy gradient training method:
[0079]
[0080] Among them, H(.) is the entropy of the main strategy at each step. These entropy items can be processed more regularly and explored by using a uniform distribution method to avoid the local optimum of the algorithm. During the training phase, the attenuation parameter β is gradually reduced from the peak value in the setting.
[0081] The critic network can be evaluated more effectively through the temporal difference method. The temporal difference method uses the approximate value function of the next step instead of the exact value function, which can minimize the square loss between the actor network and the critic network. The training process of the critic network is:
[0082]
[0083] Through multiple rounds of training, the node mapping strategy is adjusted to evaluate the network parameters. The learning agent first needs to operate in the environment and then perform training operations, which results in the experience gained from the environment being usually slow. To overcome these problems, A3C speeds up the training process by practicing in parallel, thereby improving its robustness. The central agent is responsible for collecting the work experience gained by the twelve working agents in their own environments. It is responsible for real-time training in a certain cycle and changing the network parameters, such as Figure 4 shown.
[0084] In step S103, the preliminary mapping results are analyzed by entropy weight method to obtain optimal relaxation ratio evaluation scores of virtual network requests at multiple levels, and the highest score among the optimal relaxation ratio evaluation scores is used as the differentiated service quality relaxation coefficient of virtual network requests at each level.
[0085] After obtaining the preliminary mapping results, the embodiment of the present invention uses the TOPSIS entropy weight method to comprehensively analyze the preliminary mapping results of the virtual network request. By constructing a standardized matrix, setting feedback values, and calculating the weights of different indicators and the optimal evaluation scores of each relaxation coefficient, a dynamic adjustment of the service acceptance rate, resource utilization, and QoS loss rate can be achieved.
[0086] Specifically, the entropy weight method is used to analyze the preliminary mapping results to conduct a TOPSIS entropy weight method comprehensive evaluation of the mapping results after using the original reward function to complete the mapping of various virtual network requests and setting 1-2 times QoS relaxation.
[0087] Optionally, in one embodiment of the present invention, the preliminary mapping results are analyzed by an entropy weight method to obtain optimal relaxation ratio evaluation scores for virtual network requests of multiple levels, and the highest score in the optimal relaxation ratio evaluation scores is used as the differentiated service quality relaxation coefficient for virtual network requests of each level, including: extracting the node resource utilization and link resource utilization in the preliminary mapping results as positive indicators, and extracting the delay loss rate, jitter loss rate and packet loss rate in the preliminary mapping results as negative indicators; converting the negative indicators into positive indicators, and dimensionalizing the positive and negative indicators, and converting the original decision matrix into a standardized matrix; calculating the information entropy of each indicator according to the definition of information entropy in the entropy weight method, and calculating the weight of each indicator based on the information entropy, obtaining the optimal relaxation ratio evaluation scores for virtual network requests of multiple levels based on the weights and the standardized matrix, and using the highest score in the optimal relaxation ratio evaluation scores as the differentiated service quality relaxation coefficient for virtual network requests of each level.
[0088] Specifically, a standardized matrix is first established, and the node resource utilization rate, link resource utilization rate, etc. are set as positive indicators, while the delay loss rate, jitter loss rate, and packet loss rate are set as negative indicators to achieve positive data conversion, and then the standardized matrix is dimensionally standardized. The information entropy of the indicator is calculated according to the definition of information entropy in the entropy weight method, and the weight of each indicator is calculated according to the obtained information entropy to obtain the optimal relaxation ratio evaluation score of each type of VNRs. The ratio setting with the highest score will be considered as the optimal relaxation ratio. Figure 5 As shown in Figure 2, the network performance under different QoS indicators is shown. Figure 5 (a) is the network performance under the acceptance rate, Figure 5 (b) is the network performance under link resource utilization, Figure 5 (c) is the network performance under node resource utilization, Figure 5 (d) is the long-term profit ratio of each algorithm. Figure 6 As shown in Figure 2, the changes of different indicators under 1.1-2 times QoS relaxation are shown, among which, Figure 6 (a) is the change of service acceptance rate indicator, Figure 6 (b) is the change of link utilization. Figure 6 (c) is the change of node resource utilization. Figure 6 (d) is the change of delay loss rate, Figure 6 (e) is the change of jitter loss rate, Figure 6 (f) is the change of packet loss rate.
[0089] The specific steps include:
[0090] Construct a standardized matrix: extract the classified VNRs mapping results to obtain node resource utilization, link resource utilization, delay loss rate, jitter loss rate and packet loss rate. The original decision matrix is as follows:
[0091]
[0092] Matrix A indicates that the number of evaluation objects (i.e., the values of the mapping indicator groups under each relaxation ratio) is n, and the number of evaluation indicators for each object is m. In this scheme, the evaluation indicators are node resource utilization, link resource utilization, delay loss rate, jitter loss rate, and packet loss rate. Among these five indicators, node resource utilization and link resource utilization are positive indicators, and delay loss rate, jitter loss rate, and packet loss rate are negative indicators. Therefore, the following formula is used to convert positive indicators:
[0093]
[0094] Then the five indicators are dimensionally standardized, and the formula is as follows:
[0095]
[0096] Finally, it is converted into a standardized matrix Z. The standardized matrix is as follows:
[0097]
[0098] Calculate the weights of different indicators:
[0099] According to the definition of information entropy in the entropy weight method, the information entropy of the jth indicator is calculated:
[0100]
[0101]
[0102] The weight of each indicator is calculated based on the information entropy obtained. The formula is as follows:
[0103]
[0104] Different types of relaxation ratios are set for the classified VNRs. Similarly, further, the best relaxation ratio evaluation scores for each type of VNRs are obtained, and the ratio setting with the highest score will be considered as the best relaxation ratio. Further, after setting the evaluation indicators and calculating the weights, the standardized matrix Y can be constructed:
[0105]
[0106] sty ij =w j z ij
[0107] This gives the optimal relaxation ratio evaluation score S for each type of VNRs. i , the ratio setting with the highest score will be considered as the optimal relaxation ratio. i The calculation formula is as follows:
[0108]
[0109] in, and It is the Euclidean distance between each value and the positive and negative ideal solutions.
[0110] Optionally, in one embodiment of the present invention, the preliminary mapping results are analyzed by the entropy weight method, and then the method further includes: updating the ratio of the weights of the delay loss rate, jitter loss rate, and packet loss rate in the reward function of the node mapping strategy evaluation network according to the weights to update the service quality loss rate, wherein the reward function of the node mapping strategy evaluation network is obtained by the sum of the network utilization rate and the service quality achievement rate.
[0111] According to the weight calculation results, the weight ratio of the delay loss rate, jitter loss rate, and packet loss rate in the reward function is updated to update the service quality loss rate L to:
[0112]
[0113] Where a, b, and c are the weights of the delay loss rate, jitter loss rate, and packet loss rate loss rate, respectively.
[0114] According to TOPSIS, the optimal relaxation values of the classified VNR are calculated and ranked to obtain the non-differentiated (uniform) relaxation. The same method is used to classify the data into VNRs, and then the calculation results of the differentiated relaxation evaluation scores of each category are obtained.
[0115] The optimal differentiation relaxation ratio is set for the classified VNRs, as shown in Table 2:
[0116] Table 2
[0117] QoS Level 1 2 3 4 Relaxation ratio m n p q
[0118] The QoS of classified virtual network requests is relaxed by 1.1-2 times, and the weights of delay loss rate, jitter loss rate, and packet loss rate are calculated using the TOPSIS entropy weight method. The resource utilization rate and QoS loss rate of each type of classified virtual network request are averaged and the acceptance rate of each type of virtual network request is recorded. The data is subjected to the TOPSIS entropy weight method to generate a standardized matrix, calculate the information entropy value and weight value, and generate evaluation indicators, information entropy value e, information utility value d, and weight.
[0119] Through the above introduction, the embodiment of the present invention includes two parts. The first step is to train the node mapping strategy evaluation network through a multi-threaded asynchronous dominant actor critic algorithm, set differentiated service quality relaxation coefficients for multi-level virtual network requests, temporarily modify the base network by pruning operation before link mapping, and obtain the preliminary mapping results of the virtual network request by using the node mapping strategy evaluation network. The second step is to use the TOPSIS entropy weight method to comprehensively analyze the preliminary mapping results of the virtual network request. The experimental results show that the present invention achieves a balance between meeting multiple QoS metrics and improving resource utilization. The multi-threaded asynchronous dominant actor critic (A3C) learning algorithm used effectively trains the proposed network model, and uses the training model to complete the mapping. While optimizing the node mapping, the link mapping is optimized, and good results are achieved in terms of mapping success rate, long-term yield, and node and link utilization.
[0120] In a specific embodiment, the present invention divides 2000 VNRs into six types according to Table 1, wherein the number of online games, streaming media communications, voice calls, and media broadcasts is 400 respectively, and the number of high-bandwidth and low-bandwidth general services is 200, accounting for 20% in total. The number of nodes is set to 100, and the packet loss rate of nodes in the physical network is set using a random sorting allocation method, and the delay and jitter are set for the physical link. For example, when the delay of the physical link is related to the size of the broadband resource, the larger the network bandwidth, the smaller the delay, so this algorithm generates random numbers within the range of 10-50, sorts the physical links in the order of bandwidth size, and then allocates time delays for them respectively. After obtaining the preliminary mapping results, the QoS relaxation of the classified virtual network requests is implemented by 1.1-2 times, and the weight calculation of the delay loss rate, jitter loss rate, and packet loss rate loss rate is performed using the TOPSIS entropy weight method. The resource utilization rate and QoS loss rate of each type of classified virtual network request are averaged and output, and the acceptance rate of each type of virtual network request is recorded at the same time.
[0121] According to the multi-threaded virtual network mapping and performance adjustment method based on the entropy weight method proposed in the embodiment of the present invention, the VNE problem with different QoS requirements is studied, and differentiated QoS VNE algorithms are designed from three aspects: VNE cost, network bandwidth and latency according to the functional requirements of different users. The QoS index is taken into account in the virtual network request, and different levels of service quality assurance are provided for VNRs with different QoS requirements, so as to ensure the service quality of VNR on the basis of ensuring the request success rate. After obtaining the preliminary mapping results, the TOPSIS entropy weight method is used to conduct a comprehensive analysis of the preliminary mapping results of the virtual network request, so as to achieve a balance between dynamic adjustment of the service acceptance rate, resource utilization and QoS loss rate.
[0122] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0123] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "N" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0124] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present invention belong.
Claims
1. A multi-threaded virtual network mapping and performance adjustment method based on entropy weight method, characterized in that: The following steps are involved: Acquire multiple virtual network requests, classify the multiple virtual network requests according to service quality perception to obtain multiple levels of virtual network requests, and set differentiated service quality relaxation coefficients for the multiple levels of virtual network requests; Inputting the multiple levels of virtual network requests into a trained node mapping strategy evaluation network to perform node and link mapping, and obtaining preliminary mapping results of the multiple levels of virtual network requests, wherein the node mapping strategy evaluation network is trained using a multi-threaded asynchronous advantage actor-critic algorithm using resource conditions of underlying physical nodes, virtual network requests, and mapping data between underlying physical nodes and virtual network request nodes as training data; The preliminary mapping results are analyzed by the entropy weight method to obtain the optimal relaxation ratio evaluation scores of the virtual network requests at the multiple levels, and the highest score among the optimal relaxation ratio evaluation scores is used as the differentiated service quality relaxation coefficient of the virtual network requests at each level.
2. The method according to claim 1, characterized in that Before performing node and link mapping on the virtual network requests of the multiple levels, the method further includes: A temporary link pruning operation is performed on the trained node mapping strategy evaluation network to remove physical links that do not meet virtual link requirements in the trained node mapping strategy evaluation network, and the node mapping strategy evaluation network is restored after the mapping is completed.
3. The method according to claim 1, characterized in that The node mapping strategy evaluation network includes an actor network and a critic network, and both the actor network and the critic network include a convolutional layer, a fully connected layer and an output layer, and the actor network and the critic network share the same architecture except for the output layer; wherein the input of the convolutional layer of the actor network and the critic network is the state matrix of the resource status of the underlying physical nodes and the virtual network request, the output of the convolutional layer is the suitability of each underlying physical node for the virtual network request after considering its own attributes, the output of the fully connected layer is the probability distribution of each underlying physical node being mapped by the virtual network request, the output of the output layer of the actor network is a parameterized strategy, the output of the critic network is a parameterized estimated value, and the node mapping strategy evaluation network obtains a preliminary mapping result of the virtual network request based on the parameterized strategy and the parameterized estimated value.
4. The method according to claim 3, characterized in that The method further comprises: The underlying physical nodes carrying virtual mapping resources are screened, multiple node characteristic attributes of the screened underlying physical nodes are extracted to obtain state characteristic values, the state characteristic values are normalized and screened, and then the state matrix of the underlying physical nodes is constructed using the state characteristic values.
5. The method according to claim 1, characterized in that In the node mapping strategy evaluation network, an advantage function is used to replace the state-action-value function, wherein the advantage function is: Among them, s is the estimated state, a is the action, is the state function that estimates the cumulative payoff in state s, is the state-action-value function.
6. The method according to claim 5, characterized in that The actor network is trained based on an advantage function, including: Among them, β is the decay parameter, α is the actor network learning rate, and π θ (s t ,a t ) is a parameterized strategy, is the advantage function, H(.) is the entropy of each strategy step, θ is the gradient of the parameter set, is the gradient vector; The critic network is trained by a temporal difference method, including: Among them, θ v is the critic network value function, α′ is the critic network learning rate, is the critic network gradient vector, r t is the reward at time t, γ is the discount rate, is the state function that estimates the cumulative benefit in the next state of state s, is the state function for estimating the cumulative return under state s.
7. The method according to claim 6 or 3, characterized in that: During the training of the node mapping strategy evaluation network, the actor network and the critic network are trained in parallel.
8. The method according to claim 1, characterized in that The grading of the plurality of virtual network requests according to the service quality perception includes: Taking into account the business requirements of various networks, the multiple virtual network requests are graded according to the differentiated service quality requirements in the actual different virtual network requests.
9. The method according to claim 5, characterized in that The preliminary mapping result is analyzed by an entropy weight method to obtain optimal relaxation ratio evaluation scores of the virtual network requests of the multiple levels, and the highest score among the optimal relaxation ratio evaluation scores is used as the differentiated service quality relaxation coefficient of the virtual network requests of each level, including: Extracting the node resource utilization rate and the link resource utilization rate in the preliminary mapping result as positive indicators, and extracting the delay loss rate, jitter loss rate and packet loss rate loss rate in the preliminary mapping result as negative indicators; Convert the negative indicators into positive indicators, and standardize the dimensions of the positive indicators and the negative indicators, so as to convert the original decision matrix into a standardized matrix; The information entropy of each indicator is calculated according to the definition of information entropy in the entropy weight method, and the weight of each indicator is calculated based on the information entropy. The optimal relaxation ratio evaluation scores of the virtual network requests of the multiple levels are obtained based on the weights and the standardized matrix, and the highest score among the optimal relaxation ratio evaluation scores is used as the differentiated service quality relaxation coefficient of the virtual network requests of each level.
10. The method according to claim 9, characterized in that The preliminary mapping results are analyzed by entropy weight method, and then: The weight ratios of the delay loss rate, jitter loss rate, and packet loss rate in the reward function of the node mapping strategy evaluation network are updated according to the weights to update the service quality loss rate, wherein the reward function of the node mapping strategy evaluation network is obtained by the sum of the network utilization rate and the service quality achievement rate.