Virtual machine placement method based on confidence upper bound and adaptive learning rate improved Q-learning
By using an improved Q-learning method based on confidence bounds and adaptive learning rates, we constructed an IFMT virtual machine selection algorithm and an improved Q-learning model. This algorithm solves the problems of low efficiency and slow convergence of virtual machine selection and placement algorithms in complex cloud environments, achieves efficient and flexible virtual machine placement, and improves resource utilization and cloud service quality.
Patent Information
- Application Number
- CN202510707916.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-05-29
AI Technical Summary
Existing virtual machine selection and placement algorithms have difficulty making optimal decisions in the face of complex and ever-changing cloud computing environments. Traditional Q-learning algorithms have low learning efficiency and slow convergence speed, making it difficult to find the best placement strategy in large-scale cloud computing environments.
An improved Q-learning method based on confidence bound and adaptive learning rate is adopted. The IFMT virtual machine selection algorithm is constructed by improving the influence coefficient and migration time. Combining the confidence bound technology and adaptive learning rate adjustment, an improved Q-learning virtual machine placement model is constructed to optimize the virtual machine selection and placement strategy.
It improves the accuracy and efficiency of virtual machine selection, quickly recovers overloaded host loads, optimizes resource utilization, reduces energy costs, and improves the reliability and economy of cloud services.
Smart Images

Figure CN120610778A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a virtual machine placement method, in particular to a virtual machine placement method based on improved Q-learning with a confidence upper bound and an adaptive learning rate, and belongs to the technical field of virtual machine placement. Background Art
[0002] An appropriate virtual machine selection strategy is fundamental to achieving efficient cloud computing resource allocation and stable system operation. During actual operation, selecting appropriate virtual machines for migration or scheduling can effectively prevent host overload, improve resource utilization, and ensure user service quality. Conversely, improper virtual machine selection can lead to unbalanced resource allocation, with some hosts overloaded and others idle. This not only reduces overall resource utilization but can also cause system performance degradation, severely impacting user experience. Therefore, developing scientific and efficient virtual machine selection algorithms has long been a key research priority in cloud computing.
[0003] Although existing VM selection algorithms have been effective in some scenarios, they have gradually revealed many drawbacks as cloud computing environments become increasingly complex and diverse. They generally focus on a single factor, such as the VM's resource usage, while ignoring the time required for VM migration and its impact on other VMs and the host. This single-dimensional consideration makes it difficult for the algorithm to make optimal decisions in complex and changing cloud computing environments. Furthermore, in most current research, VM selection strategies focus on the resource correlation between VMs, while research on the relationship between hosts and VMs is relatively lacking. For example, one algorithm evaluates the differences between different VMs by calculating the VM's impact coefficient on host overload. Based on this, it identifies VMs with the most significant impact on host overload and helps overloaded hosts quickly return to normal load. However, this algorithm uses cosine similarity to measure the directional similarity between two vectors. This approach only considers whether the vectors are in the same direction, but ignores the magnitude and linear relationship, and cannot fully reflect the actual relationship between the vectors. Furthermore, this algorithm does not fully consider migration time. When selecting VMs for migration, excessive migration time may negatively impact overall system performance.
[0004] At the same time, with the rapid development of the cloud computing era, the importance of virtual machine placement strategy has become increasingly prominent. Virtual machine placement is directly related to the resource utilization, energy consumption and performance of the data center. Virtual machine selection also affects the effect of virtual machine placement. By optimizing the placement of virtual machines, we can effectively reduce resource waste, improve resource utilization, reduce energy consumption costs, and provide users with more efficient, reliable and economical cloud services.
[0005] Although existing VM placement algorithms can meet basic requirements in some common scenarios, with the increasing complexity of cloud computing environments and the increasing diversity of application scenarios, traditional VM placement algorithms have gradually exposed a series of problems that need to be addressed. Traditional VM placement algorithms mostly rely on pre-set rules or simple heuristic strategies. Such methods lack flexibility in the face of dynamically changing cloud computing loads and are difficult to adapt to real-time fluctuations in resource demand, which can easily lead to irrational resource allocation and low resource utilization. In recent years, VM placement algorithms based on reinforcement learning have gradually attracted attention. Among them, the Q-learning algorithm has shown certain advantages in the field of VM placement due to its characteristic of making decisions based on state-action value functions. It can optimize VM placement strategies through continuous trial-and-error learning. However, the Q-learning algorithm suffers from low learning efficiency and slow convergence. In large-scale cloud computing environments, faced with a large number of states and action spaces, it is difficult to quickly find the optimal placement strategy. Its fixed learning rate and exploration strategy also make it difficult for the algorithm to balance exploring new strategies and leveraging existing experience in complex and changing environments. It is easy to get stuck in local optimal solutions and fail to fully discover the optimal VM placement solution.
[0006] In summary, it is necessary to improve the virtual machine placement method of Q-learning based on confidence bounds and adaptive learning rates. Summary of the Invention
[0007] A brief overview of the present invention is provided below to provide a basic understanding of certain aspects of the present invention. It should be understood that this overview is not an exhaustive overview of the present invention. It is not intended to identify key or important aspects of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is simply to present certain concepts in a simplified form as a prelude to the more detailed description discussed later.
[0008] In view of this, in order to solve the problems of low efficiency and limitations of traditional virtual machine placement methods in the prior art, the present invention provides a virtual machine placement method based on improved Q-learning with confidence upper bound and adaptive learning rate.
[0009] The technical solution is as follows: Based on the confidence upper bound and adaptive learning rate, the virtual machine placement method of Q-learning is improved, including the following steps:
[0010] S1. Based on the improved impact coefficient and migration time, an IFMT virtual machine selection algorithm is constructed to select virtual machines and obtain the most suitable virtual machine for migration;
[0011] S2. Using confidence bounds and adaptive learning rate adjustment, we set the state space, action space, and reward function, build an improved Q-learning virtual machine placement model, and find the optimal placement host.
[0012] S3. By improving the Q-learning virtual machine placement model, the best placement host is used to place the most suitable virtual machine for migration, thus realizing the virtual machine placement of the virtual machine consolidation strategy.
[0013] Furthermore, in said S1, the influence coefficient is improved Expressed as:
[0014] (1)
[0015] (2)
[0016] in, To overload the host, Number the overloaded host. Overloaded host No. virtual machines, Indicates an overloaded host and its first virtual machines The Pearson correlation coefficient between the resource usage of Indicates an overloaded host CPU usage, Indicates an overloaded host No. virtual machines The CPU usage, d is the resource dimension, and Represents overloaded hosts and its first virtual machines various resources, and Represents overloaded hosts and its first virtual machines The mean value of each resource;
[0017] The time required to migrate the virtual machine Expressed as:
[0018] (3)
[0019] in, Represents a virtual machine The current amount of memory used, Indicates an overloaded host The remaining network bandwidth available;
[0020] The IFMT virtual machine selection algorithm is expressed as:
[0021] (4)
[0022] Filtered by IFMT virtual machine selection algorithm The virtual machine with the largest value is selected as the most suitable virtual machine for migration, and the resource utilization of the overloaded host is updated at the same time.
[0023] Furthermore, in S2, the state space Expressed as:
[0024] (5)
[0025] (6)
[0026] in, is the average CPU utilization of all active hosts, is the number of overloaded hosts, is the number of underloaded hosts, Overloaded host Total CPU capacity, is the number of active hosts;
[0027] Action Space Expressed as:
[0028] (7)
[0029] in, For possible hosts, For all possible hosts A collection of
[0030] Reward Function Expressed as:
[0031] (8)
[0032] (9)
[0033] (10)
[0034] in, is the energy consumption of the host in the current state space and action space, which is expressed as a linear function of utilization. Indicates the SLA breach status after selecting a host. is the penalty coefficient for SLA breach, Express dissatisfaction situation, is the maximum power consumption of the host, The power consumption of the host in idle state;
[0035] Through the confidence upper bound technology, the Q-learning update formula and UCB formula are constructed. Under the UCB strategy, the criterion formula for selecting actions is constructed, that is, at each time step, according to the reward function Evaluate the effect of the current state action, update the Q value of adaptive learning, and complete the construction of the virtual machine placement model for improved Q-learning;
[0036] The Q-learning update formula is expressed as:
[0037] (11)
[0038] in, is the learning rate, which represents the step size of Q value update, γ is the discount factor, For the next state The maximum Q value of , which is used to update the Q value of the current state space and action space pair, For action;
[0039] The UCB formula is expressed as:
[0040] (12)
[0041] Among them, c is the coefficient of control exploration, is the state space The total number of visits, is the current state space Select Action the number of times;
[0042] The criterion formula for selecting an action is expressed as:
[0043] (13)
[0044] The adaptive learning rate adjustment formula is expressed as:
[0045] (14).
[0046] The beneficial effects of the present invention are as follows: the present invention provides a virtual machine placement method based on an improved Q-learning with a confidence upper bound and an adaptive learning rate. Through a virtual machine selection strategy algorithm based on an improved influence coefficient and migration time, the virtual machine with the most significant impact on the overloaded host and the shortest migration time is accurately identified, which is conducive to quickly restoring the load of the overloaded host; through the confidence upper bound technology, exploration and utilization are balanced, and the learning rate is adjusted according to the dynamic changes of learning, thereby improving the convergence speed and flexibility of the algorithm to optimize virtual machine placement decisions, find the best placement host, and finally use the improved Q-learning virtual machine placement model to implement virtual machine placement. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0048] Figure 1 A flowchart of the virtual machine placement method for improving Q-learning based on confidence bounds and adaptive learning rates;
[0049] Figure 2 This is a pseudo code diagram of the IFMT algorithm;
[0050] Figure 3 A pseudo-code diagram of the virtual machine placement model running process to improve Q-learning;
[0051] Figure 4 Schematic diagram of the structural principle of the virtual machine placement model for improving Q-learning. DETAILED DESCRIPTION
[0052] To make the technical solutions and advantages of the embodiments of the present invention more clearly understood, exemplary embodiments of the present invention are further described in detail below with reference to the accompanying drawings. It should be noted that the embodiments described are only a portion of the embodiments of the present invention, and are not an exhaustive list of all embodiments. It should be noted that the embodiments of the present invention and the features thereof may be combined with each other unless they conflict.
[0053] refer to Figure 1-Figure 4 This embodiment is described in detail. The virtual machine placement method based on improving the Q-learning based on the confidence upper bound and the adaptive learning rate specifically includes the following steps:
[0054] S1. Based on the improved impact coefficient and migration time, an IFMT virtual machine selection algorithm is constructed to select virtual machines and obtain the most suitable virtual machine for migration;
[0055] S2. Using confidence bounds and adaptive learning rate adjustment, we set the state space, action space, and reward function, build an improved Q-learning virtual machine placement model, and find the optimal placement host.
[0056] S3. By improving the Q-learning virtual machine placement model, the best placement host is used to place the most suitable virtual machine for migration, thus realizing the virtual machine placement of the virtual machine consolidation strategy.
[0057] Furthermore, in said S1, the influence coefficient is improved Expressed as:
[0058] (1)
[0059] (2)
[0060] in, To overload the host, Number the overloaded host. Overloaded host No. virtual machines, Indicates an overloaded host and its first virtual machines The Pearson correlation coefficient between the resource usage of Indicates an overloaded host CPU usage, Indicates an overloaded host No. virtual machines The CPU usage, d is the resource dimension, and Represents overloaded hosts and its first virtual machines various resources, and Represents overloaded hosts and its first virtual machines The mean value of each resource;
[0061] The time required to migrate the virtual machine Expressed as:
[0062] (3)
[0063] in, Represents a virtual machine The current amount of memory used, Indicates an overloaded host The remaining network bandwidth available;
[0064] The IFMT virtual machine selection algorithm is expressed as:
[0065] (4)
[0066] Filtered by IFMT virtual machine selection algorithm The virtual machine with the largest value is selected as the most suitable virtual machine for migration, and the resource utilization of the overloaded host is updated at the same time;
[0067] Specifically, according to the definition of the impact coefficient, within a specific resource dimension, a larger impact coefficient indicates that the resource consumption pattern of the overloaded host is closer to the resource demand pattern of the virtual machine, indicating that the virtual machine has a more significant impact on the overloaded host. Therefore, it is necessary to select virtual machines with higher precision to screen out the most suitable virtual machines for migration.
[0068] The IFMT virtual machine selection algorithm focuses on balancing the dual impact of virtual machine resource usage and migration time on host overload. Its core goal is to accurately identify the most suitable virtual machines for migration on overloaded physical hosts, namely those that have the most significant impact on the overloaded host and have the shortest migration time. Using the IFMT virtual machine selection algorithm, the host can maintain high resource utilization efficiency while reducing the number of virtual machine migration operations, and quickly and effectively adjust the host load to an appropriate state. Figure 2 ,IFMT virtual machine selection algorithm first initializes the host resource utilization list, and obtains the virtual machine list of the current host, obtains the resource utilization list of each host, calculates the average resource utilization of the host and obtains the resource utilization situation at the current moment. Next, the overload host detection and virtual machine selection policy resource utilization list of each virtual machine on the host is obtained, the average resource utilization of the virtual machine is calculated, and the resource utilization at the current moment is obtained. Based on the above data, the algorithm calculates the Pearson correlation coefficient (Pearson) to measure the correlation between the host and virtual machine resource utilization to calculate the influence coefficient (IF), which is used to evaluate the impact of the virtual machine on the host resource utilization. After estimating the time required for virtual machine migration (Migrate_time), the ratio of the influence coefficient to the migration time (IFMT) is calculated as the indicator for selecting the virtual machine to be migrated. Under the condition that the current resource utilization is greater than the target resource utilization and the virtual machine list is not empty, the virtual machine vm with the largest IFMT value is found and added to the list of virtual machines to be migrated. At the same time, the resource utilization of the host is updated. Finally, the algorithm returns the list of virtual machines to be migrated (selected_vms), completing the entire process of virtual machine selection.
[0069] Furthermore, in S2, the state space Expressed as:
[0070] (5)
[0071] (6)
[0072] in, is the average CPU utilization of all active hosts, is the number of overloaded hosts, is the number of underloaded hosts, Overloaded host Total CPU capacity, is the number of active hosts;
[0073] Action Space Expressed as:
[0074] (7)
[0075] in, For possible hosts, For all possible hosts A collection of
[0076] Reward Function Expressed as:
[0077] (8)
[0078] (9)
[0079] (10)
[0080] in, is the energy consumption of the host in the current state space and action space, which is expressed as a linear function of utilization. Indicates the SLA breach status after selecting a host. is the penalty coefficient for SLA breach, which is usually set to a large constant (such as 10) to emphasize the severity of SLA breach. Express dissatisfaction situation, is the maximum host energy consumption, for;
[0081] Through the confidence upper bound technology, the Q-learning update formula and UCB formula are constructed. Under the UCB strategy, the criterion formula for selecting actions is constructed, that is, at each time step, according to the reward function Evaluate the effectiveness of the current state action, balance energy consumption and SLA violation penalties in the reward function, update the adaptive learning Q value to learn the optimal strategy, and complete the construction of the improved Q-learning virtual machine placement model;
[0082] The Q-learning update formula is expressed as:
[0083] (11)
[0084] in, is the learning rate, which represents the step size of the Q value update, and γ is the discount factor used to weigh the importance of current rewards and future rewards. For the next state The maximum Q value of , which is used to update the Q value of the current state space and action space pair, For action;
[0085] The UCB formula is expressed as:
[0086] (12)
[0087] Among them, c is the coefficient of control exploration, is the state space The total number of visits, is the current state space Select Action The number of times, which is used to balance the exploration of known information and new information;
[0088] The criterion formula for selecting an action is expressed as:
[0089] (13)
[0090] The adaptive learning rate adjustment formula is expressed as:
[0091] (14)
[0092] Specifically, the state space describes the resource utilization of the cloud environment and mainly includes three parts;
[0093] The action space represents the choice of assigning virtual machines to hosts, which includes all possible hosts The set of each virtual machine allocation problem, the action is to select a possible host (i.e. a∈A) is allocated. In adaptive learning, this selection is based on the current state space The optimal adaptive learning Q value or exploration strategy (such as UCB and ϵ-greedy strategy) under θ is determined;
[0094] For the Q-Learning algorithm, the learning rate This parameter is a key parameter that determines the degree to which newly acquired information affects the current Q value. Different virtual machine placement models for improved Q-learning have different learning rate adjustment methods. This application designs an adaptive learning rate adjustment formula, and to avoid division by 0, the denominator in the formula is increased by 1.
[0095] In the initial stage, when a specific state and corresponding action are visited for the first time, the number of visits is 0, and the learning rate is Set to 1, which means that the newly acquired information will completely replace the current Q value. As the number of times the state and action are visited increases, the learning rate The learning rate will gradually decrease. The decrease in learning rate means that the impact of newly acquired information on the current Q value will gradually decrease. By adjusting the learning rate in the above way, the improved Q-learning virtual machine placement model can dynamically balance the impact of new information and existing information at different learning stages, thereby improving learning efficiency and stability. This method of using adaptive learning rate can achieve rapid exploration in the initial stage of the model and maintain the robustness of learning in the later stage to prevent overfitting.
[0096] refer to Figure 3 , first initialize the Q value list and learning rate , discount factor γ, exploration rate ε and coefficient c for controlling exploration. For each virtual machine vm to be allocated, the current system state is obtained, including the average resource utilization of all active hosts, the number of overloaded hosts and the number of underloaded hosts. Then, based on the conditions, it is determined whether to randomly select an action with probability ϵ or to use the UCB strategy to select an action. The UCB strategy calculates the UCB value of each action in the current state, selects the action with the largest UCB value, and tries to allocate the virtual machine to the corresponding host according to the selected action. If the host cannot accommodate the virtual machine, it tries to randomly select a host that can accommodate this virtual machine and has the lowest energy consumption for placement, or activates a host that has not yet been active for placement. After the placement is completed, the algorithm obtains the next system state and calculates the reward obtained from this allocation. The Q value is updated according to the Q-Learning update formula, and the learning rate is dynamically adjusted according to the number of action visits, and the number of state visits and action visits are updated.
[0097] refer to Figure 4, according to the current state of the cloud data center composed of hosts, an action is executed, that is, the virtual machine to be migrated is allocated to the appropriate host. Then the system provides feedback reward Reward based on the execution result. This placement action will change the environment state and make the entire environment reach a new state.
[0098] Although the present invention has been described with respect to a limited number of embodiments, it will be apparent to those skilled in the art, having benefit of the foregoing description, that other embodiments are contemplated within the scope of the invention thus described. Furthermore, it should be noted that the language used in this specification has been selected primarily for readability and didactic purposes, rather than for the purpose of explaining or limiting the subject matter of the present invention. Consequently, many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the appended claims. The disclosure of the present invention is intended to be illustrative rather than restrictive of the scope of the invention, which is defined by the appended claims.
Claims
1. A virtual machine placement method based on improved Q-learning based on confidence upper bound and adaptive learning rate, characterized in that: The following steps are involved: S1. Based on the improved impact coefficient and migration time, an IFMT virtual machine selection algorithm is constructed to select virtual machines and obtain the most suitable virtual machine for migration; S2. Using confidence bounds and adaptive learning rate adjustment, we set the state space, action space, and reward function, build an improved Q-learning virtual machine placement model, and find the optimal placement host. S3. By improving the Q-learning virtual machine placement model, the best placement host is used to place the most suitable virtual machine for migration, thus realizing the virtual machine placement of the virtual machine consolidation strategy.
2. The virtual machine placement method based on improved Q-learning with confidence upper bound and adaptive learning rate according to claim 1, characterized in that: In S1, the impact coefficient is improved Expressed as: (1) (2) in, To overload the host, Number the overloaded host. Overloaded host No. virtual machines, Indicates an overloaded host and its first virtual machines The Pearson correlation coefficient between the resource usage of Indicates an overloaded host CPU usage, Indicates an overloaded host No. virtual machines The CPU usage, d is the resource dimension, and Represents overloaded hosts and its first virtual machines various resources, and Represents overloaded hosts and its first virtual machines The mean value of each resource; The time required to migrate the virtual machine Expressed as: (3) in, Represents a virtual machine The current amount of memory used, Indicates an overloaded host The remaining network bandwidth available; The IFMT virtual machine selection algorithm is expressed as: (4) Filtered by IFMT virtual machine selection algorithm The virtual machine with the largest value is selected as the most suitable virtual machine for migration, and the resource utilization of the overloaded host is updated at the same time.
3. The virtual machine placement method based on improved Q-learning with confidence upper bound and adaptive learning rate according to claim 2, characterized in that: In S2, the state space Expressed as: (5) (6) in, is the average CPU utilization of all active hosts, is the number of overloaded hosts, is the number of underloaded hosts, Overloaded host Total CPU capacity, is the number of active hosts; Action Space Expressed as: (7) in, For possible hosts, For all possible hosts A collection of Reward Function Expressed as: (8) (9) (10) in, is the energy consumption of the host in the current state space and action space, which is expressed as a linear function of utilization. Indicates the SLA breach status after selecting a host. is the penalty coefficient for SLA breach, Express dissatisfaction situation, is the maximum power consumption of the host, The power consumption of the host in idle state; Through the confidence upper bound technology, the Q-learning update formula and UCB formula are constructed. Under the UCB strategy, the criterion formula for selecting actions is constructed, that is, at each time step, according to the reward function Evaluate the effect of the current state action, update the Q value of adaptive learning, and complete the construction of the virtual machine placement model for improved Q-learning; The Q-learning update formula is expressed as: (11) in, is the learning rate, which represents the step size of Q value update, γ is the discount factor, For the next state The maximum Q value of , which is used to update the Q value of the current state space and action space pair, For action; The UCB formula is expressed as: (12) Among them, c is the coefficient of control exploration, is the state space The total number of visits, is the current state space Select Action the number of times; The criterion formula for selecting an action is expressed as: (13) The adaptive learning rate adjustment formula is expressed as: (14)。
Citation Information
Patent Citations
Optimal decision-making method based on improved Q-learning
CN112598137A
Virtual machine optimization scheduling method for cloud computing
CN115016889A
System and method for fair and economical resource partitioning using virtual hypervisor
US20110185064A1