A Virtual Machine Optimization Scheduling Method for Cloud Computing

By constructing the Markov decision-making process and using reinforcement learning technology to train the auxiliary decision matrix, the problem of difficult to comprehensively consider the system energy efficiency and robustness in the multi-objective scheduling optimization of virtual machines is solved, and the effect of outputting the global optimal scheduling scheme according to decision makers' preferences is achieved.

CN115016889BActive Publication Date: 2025-06-13EAST CHINA UNIV OF SCI & TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210423376.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-21
Publication Date
2025-06-13
Estimated Expiration
2042-04-21

AI Technical Summary

Technical Problem

The prior art is difficult to concentrate non-inferior solutions from the optimization of multi-objective scheduling of virtual machines, comprehensively consider system energy efficiency and system robustness, and automatically decide on the optimal scheduling scheme according to decision makers' preferences.

Method used

By converting the non-inferior solution set into an objective function set, combining the system state transfer relationship of the resource scheduling process, Markov decision making process (MDP), using reinforcement learning technology to train the auxiliary decision matrix, and finally constructing a trade-off decision matrix and outputting a global optimal scheduling scheme according to user preferences.

Benefits of technology

In the multi-objective scheduling optimization of virtual machines, we have comprehensively considered system energy efficiency and system robustness, and flexibly selected the optimal scheduling scheme according to decision makers' preferences, adapting to infrastructure layer cloud platforms of different architecture types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115016889B_ABST
    Figure CN115016889B_ABST
Patent Text Reader

Abstract

The present invention relates to a virtual machine optimization scheduling method for cloud computing. The method takes the multi-objective scheduling optimization result of virtual machines as input parameters. First, the method establishes an original decision matrix through objective transformation, and then combines the system state transition relationship of the resource scheduling process to establish a Markov decision process MDP corresponding to the resource scheduling process. Furthermore, an auxiliary decision matrix with virtual machine migration cost information is obtained through reinforcement learning technology training. Finally, a trade-off decision matrix is constructed using the original decision matrix and the auxiliary decision matrix, and according to the user's preference information for target attributes, a global optimal scheduling scheme is output. Compared with the prior art, the present invention has the advantages that it not only considers the steady-state target information before the execution of the virtual machine scheduling scheme, such as energy consumption, quality of service, resource utilization rate, etc., but also takes into account the potential migration costs that may be caused to subsequent resource integration after the execution of the virtual machine scheduling scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud computing resource scheduling, and in particular to a virtual machine optimization scheduling method for cloud computing. Background Art

[0002] Virtual machine scheduling in a cloud computing environment refers to, according to a certain scheduling strategy, allocating virtual resources requested by different tenants to multiple computing nodes in a data center at a specified time sequence, and reasonably integrating resources according to the load status during the operation of an application, in order to obtain better system execution performance.

[0003] Virtual machine scheduling is an NP-complete problem, involving the optimization of multiple objectives such as the energy consumption, resource loss, and tenant service quality of a data center, and there are mutual constraints and conflicts between multiple scheduling objectives. Therefore, there is no globally optimal solution for single-objective scheduling, and the solution result is a set of non-dominated solutions that are a compromise of multiple objectives. Traditional heuristic scheduling frameworks and models lack the ability to efficiently search for the globally optimal solution and depend on the type of scheduling problem, and cannot adapt to the changing cloud computing application environment. And typical scheduling problems can be directly mapped to bin-packing problems. Multi-objective evolutionary algorithms (such as NSGA-II, MOEA / D, SPEA2, etc.) can design coding methods suitable for different bin-packing problems, have good global optimization capabilities, and can be easily combined with other optimization strategies (thus solving their own defects such as lack of feedback mechanism and slow convergence speed), and have natural advantages in solving the multi-objective scheduling optimization problem of virtual machines.

[0004] However, the result of a multi-objective evolutionary algorithm is a decision set, and no method for selecting a specific decision from the decision set is given. Without using other auxiliary decision-making mechanisms, cloud service providers can only randomly select a scheduling scheme from the non-dominated decision set. In addition, the non-dominated solution set only contains steady-state target information before virtual machine scheduling, such as energy consumption, service quality, resource utilization rate, etc., and does not reflect the migration cost that may be caused to subsequent resource integration after virtual machine placement. Therefore, how to design a virtual machine multi-objective scheduling trade-off decision-making mechanism that takes into account both system energy efficiency and system robustness for the non-dominated solution set is a technical problem commonly faced by infrastructure-as-a-service layer platforms. Summary of the Invention

[0005] The purpose of the present invention is to provide a virtual machine optimization scheduling method for cloud computing in order to overcome the defects existing in the above-mentioned prior art.

[0006] The purpose of the present invention can be achieved by the following technical solutions:

[0007] According to one aspect of the present invention, there is provided a method for optimizing the scheduling of virtual machines for cloud computing. The method takes the multi-objective scheduling optimization result of virtual machines as an input parameter. First, the method establishes an original decision matrix through objective transformation, and then combines the system state transition relationship of the resource scheduling process to establish a Markov decision process MDP (Markov Decision Process) corresponding to the resource scheduling process. Furthermore, the method trains through reinforcement learning technology to obtain an auxiliary decision matrix with virtual machine migration cost information. Finally, the method constructs a trade-off decision matrix using the original decision matrix and the auxiliary decision matrix, and outputs a globally optimal scheduling scheme according to the user's preference information for target attributes.

[0008] As a preferred technical solution, the method specifically includes the following steps:

[0009] Step S1: Establish an original decision matrix based on the non-dominated solution set Convert the non-dominated solution set X=(x 1 ,x 2 ,...,x j ,...,x n ) T into a set of objective functions according to the objective function

[0010] Step S2: Establish an auxiliary decision matrix Q;

[0011] Step S3: Train the MDP model using reinforcement learning technology until the reward matrix Q-Value converges, where the reward matrix is the auxiliary decision matrix;

[0012] Step S4: Establish a trade-off decision matrix based on the original decision matrix and the auxiliary decision matrix Q;

[0013] Step S5: Construct a weighted normalized decision matrix based on user preferences

[0014] Step S6: Define the ideal point such that each attribute value in the ideal point is the optimal value in the decision set, where each attribute value in the ideal point is the minimum value of each element in the weighted normalized decision matrix;

[0015] Step S7: Output the globally optimal solution based on the ideal point in the weighted normalized decision matrix.

[0016] As a preferred technical solution, step S2 specifically includes:

[0017] 201) Determine the measurement standard of the migration cost M, and use the migration cost M as the virtual machine scheduling scheme xj The corresponding reward function;

[0018] 202) Taking the migration cost M as the system robustness index, an MDP model is established based on the system state transition relationship in the virtual machine scheduling process.

[0019] As an optimal technical solution, the migration cost M uses the migration time of each virtual machine vm j for measurement, and is specifically expressed as:

[0020]

[0021] Wherein, represents the time required for the virtual machine vm j to complete the migration; M j represents the memory request amount when the virtual machine vm j is migrated; B j represents the available bandwidth of the physical host.

[0022] As an optimal technical solution, the MDP model is defined by a quadruple: M = (S, A, P sa , R);

[0023] Wherein, S is the state space, and there is s ∈ S, s t represents the state received by the Agent at time step t;

[0024] A is the action space, and there is a ∈ A, a t represents the action executed by the Agent at time step t;

[0025] P sa represents the probability distribution of the Agent transferring to other states s ∈ S after the action a ∈ A acts in the current state s ∈ S;

[0026] R is the reward function;

[0027] Among them, the state space S is defined as: expressing the CPU utilization rate of the i-th physical host at time t as s ti , then the state space of the computing node cluster at time step t is expressed as S t = (s t1 , s t2 ,..., s tn ), and n represents the number of physical hosts;

[0028] The action space A is defined as follows: Each non-dominated solution represents a virtual machine scheduling scheme. Using the method based on distributed support vector machines, the Pareto optimal solution set of the multi-objective resource scheduling problem is divided into three categories: energy consumption priority actions, quality of service priority actions, and resource utilization efficiency priority actions, that is, A = {energy consumption priority, quality of service priority, resource utilization efficiency priority}.

[0029] As a preferred technical solution, the state space S adopts neural network technology, and combines neural network technology to perform dimensionality reduction and aggregation processing on the state set of the server.

[0030] As a preferred technical solution, in the auxiliary decision matrix in step S3, the reward function value is negatively correlated with the migration cost. The greater the migration cost, the smaller the expected reward value;

[0031] The MDP model in step S3 uses the Double Q-Learning algorithm of reinforcement learning technology to train and solve it.

[0032] As a preferred technical solution, step S4 specifically includes:

[0033] 401) Calculate the system cluster state (s j , s j1 ,..., s j2 ) after executing the virtual machine placement scheme; jn )

[0034] 402) Calculate the action a corresponding to the execution of the x j scheduling scheme;

[0035] 403) Obtain the reward value Reward j1 , s j2 ,..., s jn ) corresponding to the execution of the action a according to the auxiliary decision matrix; ji ;

[0036] 404) Add Reward ji as a new objective function value of the non-dominated solution x j to the decision matrix to form a trade-off decision matrix

[0037] As a preferred technical solution, step S5 specifically includes:

[0038] 501) If the decision maker has no specific preference for the objective attributes of virtual machine scheduling, use the entropy weight method to automatically determine the objective weights of each objective;

[0039] 502) For the trade-off decision matrix Normalize the attribute values in it to construct a normalized decision matrix

[0040] 503) Construct a weighted normalized decision matrix by combining the preference values of the decision maker for each objective or the system default objective weight values

[0041] Among them, the weighted normalized decision matrix The representation is defined as:

[0042] c ij = w j × b ij (2)

[0043] Among them, w j Represents the preference weight set by the decision maker for the j-th objective.

[0044] As a preferred technical solution, step S7 specifically includes:

[0045] Calculate the distance of all scheduling schemes in the weighted normalized decision matrix from the ideal point ;

[0046] Output the point with the shortest distance and use it as the final trade-off solution.

[0047] Compared with the prior art, the present invention has the following advantages:

[0048] 1) The present invention takes the multi-objective scheduling optimization result of the virtual machine as the decision set, comprehensively considers the system energy efficiency and system robustness, and can flexibly select the optimal scheduling scheme according to the decision maker's preference;

[0049] 2) The present invention not only considers the steady-state target information before the execution of the virtual machine scheduling scheme, such as energy consumption, quality of service, resource utilization rate, etc., but also takes into account the potential migration cost that may be caused to the subsequent resource integration after the execution of the virtual machine scheduling scheme;

[0050] 3) The present invention allows the decision maker to flexibly set the weights of each objective according to his own preference, with better flexibility;

[0051] 4) The present invention is applicable to infrastructure layer cloud platforms of different architecture types, and is decoupled from multi-objective optimization algorithms and specific optimization objectives, with strong adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 Is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0054] As Figure 1 shown, a trade-off decision-making method for the virtual machine multi-objective scheduling optimization problem of the present invention takes the non-dominated virtual machine multi-objective scheduling optimization solution set as an input parameter. First, the non-dominated solution set is converted into an objective function set (original decision matrix) through objective function transformation; then, combined with the system state transition relationship in the resource scheduling process, the state space, action space, and reward function are defined, and then the MDP corresponding to the resource scheduling process is constructed. The Q-value matrix (auxiliary decision matrix) is trained to convergence through reinforcement learning technology; finally, a trade-off decision matrix is constructed using the original decision matrix and the auxiliary decision matrix, and based on the user's preference information for the target attributes, a trade-off decision-making scheme based on the ideal point is output. The specific steps are as follows:

[0055] S1. Establish the original decision matrix based on the non-dominated solution set of virtual machine multi-objective scheduling optimization: Assume that the non-dominated solution set X is represented as an n-dimensional column vector: X = (x 1 , x 2 ,..., x j ,..., x n ), T where x j is the j-th non-dominated solution, and n is the number of non-dominated solutions; the objective function set F is represented as F = (p(x j ), q(x j ), u(x j )), where p(x j ), q(x j ), u(x j ) represent the objective functions of system energy consumption, virtual machine service quality, and resource utilization efficiency, respectively; then the non-dominated objective function set corresponding to the non-dominated solution set X can be expressed as: as follows;

[0056]

[0057] S2. Establish the MDP based on the system state transition process in the virtual machine scheduling process: The MDP model is defined as a quadruple M = (S, A, P sa , R). Among them, S is the state space, representing the state space S t of the computing node cluster at time step t = (s t1 , st2 ,..., s tn ), s ti represents the CPU utilization rate of the i-th computing node at time t; A is the action space. Based on the distributed support vector machine method, the Pareto optimal solution set of the multi-objective resource scheduling problem is divided into three categories: energy consumption priority actions, quality of service priority actions, and resource utilization efficiency priority actions, that is, A = {energy consumption priority, quality of service priority, resource utilization efficiency priority}; P sa represents the probability distribution of other states s ∈ S that the Agent will transfer to after the action a ∈ A acts in the current state s ∈ S; R is the reward function, as shown in formula (1);

[0058] S3. Use Double Q-Learning to train the Q-Value matrix of the MDP in step S2 until convergence, and call the converged Q-Value matrix the auxiliary decision matrix, as shown in Table 1;

[0059] Table 1

[0060]

[0061] S4. Establish a trade-off decision matrix based on the original decision matrix and the auxiliary decision matrix: First, calculate the system cluster state (s j after executing the virtual machine placement scheme, s j1 , s j2 ,..., s jn ) and the action space a ∈ {energy consumption priority, quality of service priority, robustness priority} to which x j belongs; then obtain the reward value Reward j1 corresponding to the action a executed according to the state (s j2 , s jn ) from the auxiliary decision matrix; finally, by using Reward ji as a new objective function value of the non-dominated solution x ji and adding it to the decision matrix j , a trade-off decision matrix is formed as follows; as shown below;

[0062]

[0063] S6. Construct a weighted normalized decision matrix based on user preferences: First, normalize the attribute values in the trade-off decision matrix to construct a normalized decision matrix , and then combine the user's preference values for each objective to construct a weighted normalized decision matrix , where c ij = wj ×b ij ,w j represents the weight set by the decision maker for the j-th objective; if the decision maker has no preference for the attributes of each optimization objective of virtual machine scheduling, the entropy weight method is used to determine the objective weights of each objective, so as to automatically transform the non-preference decision problem into a preference decision problem;

[0064] S6. Define the ideal point such that each attribute value in the ideal point is the optimal value in the decision set;

[0065] S7. Calculate the weighted normalized decision matrix Calculate the Euclidean distance (negative correlation coefficient) between each point in and the ideal point , and output the decision point closest to (with the largest negative correlation coefficient) as the final trade-off decision scheme.

[0066] The technical problem to be solved by the present invention is how to comprehensively consider system energy efficiency and system robustness from the non-dominated solution set of virtual machine multi-objective scheduling optimization, and automatically make an optimal scheduling decision in combination with the decision maker's preference. The technical solution adopted is: taking the non-dominated solution set of virtual machine multi-objective scheduling optimization as the input parameter, first converting the non-dominated solution set into an objective function set (original decision matrix) through objective function transformation; then combining the system state transition relationship in the resource scheduling process, defining the state space, action space and reward function, and further constructing the MDP corresponding to the resource scheduling process, training the Q-value matrix (auxiliary decision matrix) to convergence through reinforcement learning technology; finally, constructing a trade-off decision matrix using the original decision matrix and the auxiliary decision matrix, and outputting a trade-off decision scheme based on the ideal point according to the user's preference information for the objective attributes.

[0067] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.​​

Claims

1. A virtual machine optimization scheduling method for cloud computing, characterized in that, this method takes the multi-objective scheduling optimization result of the virtual machine as the input parameter. First, the original decision matrix is established through objective transformation, and then combined with the system state transition relationship in the resource scheduling process, the Markov decision process MDP corresponding to the resource scheduling process is established. Furthermore, the auxiliary decision matrix with virtual machine migration cost information is obtained through reinforcement learning technology training. Finally, the trade-off decision matrix is constructed using the original decision matrix and the auxiliary decision matrix, and the global optimal scheduling scheme is output according to the user's preference information for the target attributes; The specific steps of the method are as follows: Step S1. Establish an original decision matrix based on the non-dominated solution set Let the non-dominated solution set be X = (x 1 , x 2 ,..., x j ,..., x n ) T Convert it into a set of objective functions according to the objective function Step S2: Establish the auxiliary decision matrix Q; Step S3: Use reinforcement learning technology to train the MDP model until the reward matrix Q-Value converges, where the reward matrix is the auxiliary decision matrix; Step S4. Based on the original decision matrix and the auxiliary decision matrix Q, establish a trade-off decision matrix; Step S5: Construct a weighted normalized decision matrix based on user preferences Step S6: Define the ideal point Make the ideal point such that each attribute value in it is the optimal value in the decision set, where each attribute value of the ideal point is the minimum value of each element in the weighted normalized decision matrix; Step S7: Output the global optimal solution based on the ideal point in the weighted criterion decision matrix Output the global optimal solution; The specific content of step S2 includes: 201) Determine the metric for measuring the migration cost M, and use the migration cost M as the reward function corresponding to the virtual machine scheduling scheme x j corresponding to; 202) Taking the migration cost M as the system robustness index, establish the MDP model based on the system state transition relationship in the virtual machine scheduling process; The migration cost M is measured using the migration time of each virtual machine vm j and is specifically expressed as: Among them, represents the virtual machine vm j The time required to complete the migration; M j represents the virtual machine vm j The memory request volume during migration; B j represents the available bandwidth of the physical host; The described MDP model is defined as a quadruple: M = (S, A, P sa , R); Among them, S is the state space, and there is s ∈ S, s t represents the state received by the Agent at time step t; Let \(A\) be the action space, where \(a\in A\), and \(a\) t represents the action executed by the Agent at time step \(t\); P sa represents the probability distribution of other states s ∈ S to which the Agent will transfer after the action a ∈ A is applied in the current state s ∈ S; R is the reward function; Among them, the state space S is defined as: representing the CPU utilization rate of the i-th physical host at time t as s ti , then the state space of the computing node cluster at time step t is represented as S t =(s t1 , s t2 ,..., s tn ), where n represents the number of physical hosts; The action space A is defined as: each non-dominated solution represents a virtual machine scheduling scheme. The Pareto optimal solution set of the multi-objective resource scheduling problem is divided into three categories: energy consumption priority actions, quality of service priority actions, and resource utilization efficiency priority actions by using the method based on the distributed support vector machine, that is, A = {energy consumption priority, quality of service priority, resource utilization efficiency priority}; The specific content of step S4 includes: 401) Calculate according to x j The system cluster state (s j1 , s j2 ,..., s jn ) after executing the virtual machine placement plan; 402) Calculate the execution of x j The action a corresponding to the scheduling scheme; 403) Obtain the reward value Reward corresponding to the execution of action a for the states (s j1 , s j2 ,..., s jn ) ji ; (404) Add Reward ji as a new objective function value of the non-inferior solution x j to the decision matrix to form a trade-off decision matrix 2. The virtual machine optimization scheduling method for cloud computing according to claim 1, characterized in that, The state space S adopts neural network technology, and combines neural network technology to perform dimensionality reduction and aggregation processing on the state set of the server.

3. The virtual machine optimization scheduling method for cloud computing according to claim 1, characterized in that, In the auxiliary decision matrix in step S3, the reward function value is negatively correlated with the migration cost. The greater the migration cost, the smaller the expected reward value; The MDP model in step S3 is trained and solved using the reinforcement learning technology Double Q-Learning algorithm.

4. The virtual machine optimization scheduling method for cloud computing according to claim 1, characterized in that, The specific content of step S5 includes: 501) If the decision maker has no specific preference for the target attributes of the virtual machine scheduling, use the entropy weight method to automatically determine the objective weights of each target; 502) Normalize the attribute values in the trade-off decision matrix to construct a normalized decision matrix 503) Construct a weighted normalized decision matrix by combining the preference values of the decision maker for each objective or the default objective weights of the system Among them, the weighted normalized decision matrix The representation method is defined as: c ij = w j × b ij (2) Among them, w j represents the preference weight set by the decision maker for the j-th objective.

5. The virtual machine optimization scheduling method for cloud computing according to claim 1, characterized in that, The specific content of step S7 includes: Calculate the distances of all scheduling schemes in the weighted normalized decision matrix from the ideal point ; Output distance The nearest point and use it as the final trade-off solution.

Citation Information

Patent Citations

  • Cloud network service resource scheduling method

    CN110399223A

  • Virtual machine resource scheduling method based on reinforcement learning

    CN111143036A