Cloud computing resource scheduling method and system based on elastic telescopic double-layer scheduling framework
By introducing a flexible scalable two-layer scheduling framework based on deep reinforcement learning in cloud computing, the coordination problem of task scheduling and resource elastic expansion in cloud computing is solved, efficient task scheduling and resource allocation are achieved, cost reduction and service quality is improved.
Patent Information
- Application Number
- CN202510637214.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The existing technology is difficult to effectively coordinate task scheduling and resource elastic expansion in cloud computing, resulting in insufficient or over-supply of resource supply, affecting service quality and increasing costs.
A flexible scaling double-layer scheduling framework based on deep reinforcement learning is proposed. By defining a task allocation scheduler and a virtual machine automatic scaling scheduler, the two-layer scheduling algorithm works together to achieve efficient task scheduling and elastic resource allocation.
On the premise of ensuring user satisfaction and QoS requirements, reduce the waste and additional costs of virtual machine rental resources, and improve task success rate and resource utilization rate.
Smart Images

Figure CN120216204A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cloud computing resource scheduling, and particularly to a cloud computing resource scheduling method and system based on an elastic scaling double-layer scheduling framework. Background Art
[0002] The statements in this section merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.
[0003] Due to the reliability and flexibility of cloud computing, more and more companies are migrating their applications to the cloud. Companies rent appropriate cloud resources from infrastructure service (IaaS) providers to build their own software application services at a lower cost. For these software operation service (SaaS) providers, they have a dual function in the cloud market. On the one hand, as service providers for end-users, SaaS applications must process service requests and execute user tasks to meet different service requirements. On the other hand, as users of IaaS providers, they need to pay to rent virtual machine instances from the cloud platform to deploy their SaaS applications. In this business model, the goal of SaaS providers is to ensure the quality of their SaaS application services while minimizing the rental cost of virtual machine instances. Therefore, how to effectively select a virtual machine rental plan and allocate the tasks submitted by users to appropriate resources is a major challenge faced by SaaS providers.
[0004] Since the workload of transactional applications changes dynamically over time, the demand for cloud resources will also change accordingly. When the resources currently rented by a SaaS provider are less than the resources required to process the application workload, it will lead to insufficient resource supply, which will cause the application performance to decline, thus affecting the service quality. On the contrary, when the resources used by a SaaS provider exceed the resources required by the application workload, this will lead to over-supply of resources, resulting in unnecessary costs. Due to the elastic characteristics of cloud computing, the resources used by SaaS applications can be flexibly scaled up or down.
[0005] The prior art can be roughly divided into two categories of related research: task scheduling and resource elastic scaling. Machine learning algorithms solve the problems of task scheduling and resource allocation in cloud computing by learning and predicting task patterns, optimizing resource allocation strategies, and dynamically adapting to environmental changes. Resource elastic scaling is an important feature of cloud computing, which enables cloud platforms or SaaS applications to flexibly and dynamically allocate or release resources at any time to respond to fluctuations in user demand. According to different scaling methods, elastic resource provisioning can be divided into vertical scaling and horizontal scaling. Since horizontal scaling is more targeted at the common resource problems in SaaS applications, resource horizontal scaling strategies based on rules or thresholds and resource horizontal scaling strategies based on prediction mechanisms are proposed; in addition, platforms such as Amazon AWS and Google Cloud provide Auto Scaling, which automatically increases or decreases the number of VM instances according to the load on the computing instances; the Auto Scaling function in Amazon AWS automatically adjusts the number of Amazon EC2 instances according to the load while ensuring high availability and performance; the task scheduler of Google Cloud reasonably allocates tasks to the most suitable VM instances for execution according to resource requirements and load; however, the above technologies often perform resource elastic scaling independently and cannot work well in coordination with task scheduling.
[0006] The cloud computing environment is highly dynamic and complex, and factors such as task load, resource requirements, system failures, and resource availability often change. Task scheduling usually involves multiple optimization goals, such as resource utilization, response time, task completion time, energy efficiency, and cost control. Traditional scheduling methods often have difficulty finding a balance among multiple goals. On the one hand, the resource pool in cloud computing usually changes dynamically; on the other hand, the resource requirements, execution time, priority, etc. of tasks may be uncertain. At the same time, the elastic scaling of cloud platforms is a core feature that allows automatic adjustment of resources according to changes in the workload. How to work in coordination with the automatic scaling mechanism during the task scheduling process is an important issue in task scheduling in cloud computing. Summary of the Invention
[0007] To overcome the deficiencies of the above prior art, the present invention provides a cloud computing resource scheduling method and system based on an elastic scaling double-layer scheduling framework, and innovatively proposes an elastic scaling double-layer scheduling framework based on deep reinforcement learning, realizing efficient task scheduling and elastic resource allocation.
[0008] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions: In a first aspect, the present invention provides a cloud computing resource scheduling method based on an elastic scaling double-layer scheduling framework, including: Based on the three - layer cloud computing market, a two - layer scheduling framework is established; the three - layer cloud computing market includes infrastructure service providers, software operation service providers, and end - users; the two - layer scheduling framework includes a task allocation scheduler and a virtual machine auto - scaling scheduler; The task allocation scheduler defines a first agent, and by collecting the task characteristics and virtual machine usage data of the current cloud computing platform, determines the state, action, and reward function of the first agent; The virtual machine auto - scaling scheduler defines a second agent, and by collecting the workload and resource utilization data of the current cloud computing platform, determines the state, action, and reward function of the second agent; The two - layer scheduling algorithm is used to solve the task allocation decision of the first agent, and the two - layer scheduling algorithm is used to solve the virtual machine adjustment decision of the second agent.
[0009] In a further technical solution, the task response time in the two - layer scheduling framework is expressed as:
[0010] where, represents the task response time, represents the task execution time, represents the task waiting time.
[0011] In a further technical solution, the state of the first agent is expressed as:
[0012] The action is expressed as:
[0013] where, represents the state space of the th task of the task allocation scheduler, represents the number of virtual machine instances in the current resource pool, represents the waiting time when allocated to the th virtual machine instance, represents the action of allocating the th task in the task allocation scheduler to a virtual machine, represents the th virtual machine.
[0014] In a further technical solution, the reward function of the first agent is expressed as:
[0015] where, represents the reward function value of the th task of the task allocation scheduler, Indicates the total time the task stays on the virtual machine instance, Indicates the actual execution time of the task.
[0016] Further technical solution, the state of the second agent is represented as:
[0017] The action is represented as:
[0018] Among them, Indicates the state when the virtual machine auto-scaling scheduler executes the th virtual machine lease adjustment plan, Indicates the number of tasks submitted by users in the recent period, Indicates the number of virtual machine instances rented in the recent period, Is a specific time within the workload cycle, And Respectively represent the average response time and success rate of all tasks during this period, Is The resource utilization rate of all virtual machine instances in the resource pool during the Indicates the maximum range of virtual machine instances that can be increased or decreased.
[0019] Further technical solution, the reward function of the second agent is defined as:
[0020] Among them, Indicates the resource utilization rate during the period , Indicates the penalty that needs to be given as a reward when the success rate is lower than a certain threshold, Indicates the weight of the penalty for the reward.
[0021] Further technical solution, the two-layer scheduling algorithm consists of two layers of DQN algorithms. The first layer of DQN algorithm is responsible for task allocation, and the second layer of DQN algorithm is responsible for virtual machine adjustment.
[0022] Second, the present invention provides a cloud computing resource scheduling system based on an elastic scaling two-layer scheduling framework, including: A framework construction module, which is configured to: establish a two-layer scheduling framework based on a three-layer cloud computing market; the three-layer cloud computing market includes an infrastructure service provider, a software operation service provider, and an end user; the two-layer scheduling framework includes a task allocation scheduler and a virtual machine auto-scaling scheduler; The first scheduling module is configured to: The task allocation scheduler defines a first agent, and determines the state, actions, and reward function of the first agent by collecting the task characteristics and virtual machine usage data of the current cloud computing platform; The second scheduling module is configured to: The virtual machine auto-scaling scheduler defines a second agent, and determines the state, actions, and reward function of the second agent by collecting the workload and resource utilization data of the current cloud computing platform; The scheduling decision-making module is configured to: Solve the task allocation decision of the first agent by using a two-layer scheduling algorithm, and solve the virtual machine adjustment decision of the second agent by using a two-layer scheduling algorithm.
[0023] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in the cloud computing resource scheduling method based on the elastic scaling two-layer scheduling framework described in the first aspect are implemented.
[0024] In a fourth aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in the cloud computing resource scheduling method based on the elastic scaling two-layer scheduling framework described in the first aspect are implemented.
[0025] The above one or more technical solutions have the following beneficial effects: In order to reduce the resource waste and additional costs caused by renting virtual machines while ensuring user satisfaction and QoS requirements, the present invention proposes an elastic scaling two-layer scheduling framework based on deep reinforcement learning to handle the problems of reasonable task scheduling and resource elastic scaling based on deep reinforcement learning. On the one hand, select the most suitable virtual machine in the current environment, aiming to reduce the task completion time and improve the user experience while meeting user needs. On the other hand, elastically adjust the resources according to the task load of the environment, and use deep reinforcement learning to dynamically increase or decrease the rental of a certain number of virtual machines, and reduce the cost consumption of the cloud service provider for renting virtual machines while ensuring the task QoS requirements.
[0026] The present invention proposes an elastic scaling double-layer scheduling framework based on deep reinforcement learning, and provides an intelligent double-layer scheduling algorithm for SaaS providers. Through the collaborative cooperation of the QoS-aware task scheduling function and the workload-aware virtual machine automatic scaling function, efficient task scheduling and elastic resource allocation are achieved. The first layer provides the QoS-aware task scheduling function, using the DQN algorithm to select the most suitable virtual machine instance from the resource pool for the current task, minimizing the average task response time while ensuring the QoS requirements of the task. The second layer uses the DQN algorithm to provide the workload-aware virtual machine automatic scaling function, which is responsible for dynamically adjusting the number of leased virtual machine instances according to the load changes of the tasks and the attributes of the currently leased virtual machines. When the virtual machine instances in the resource pool are saturated, the number of virtual machine instances can be appropriately reduced to avoid resource waste and reduce the leasing cost. Conversely, when the number of tasks is small during a certain period, more virtual machine instances can be leased appropriately to ensure the success rate of the tasks and improve user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0028] Figure 1 It is a block diagram of the internal scheduling process in a three-layer cloud computing market based on the double-layer scheduling framework in an embodiment of the present invention; Figure 2 It is a block diagram of the process of the double-layer scheduling algorithm in an embodiment of the present invention; Figure 3 It is a graph showing the changes in the number of tasks and the number of virtual machines in the double-layer scheduling framework in a simulated scenario in an embodiment of the present invention; Figure 4 It is a graph showing the changes in the number of tasks and the number of virtual machines in the double-layer scheduling framework in an actual scenario in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0030] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0031] In the case of no conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0032] Explanation of professional terms: SaaS: Software as a Service, software operation service; IaaS: Infrastructure as a Service, infrastructure service; VM: Virtual Machine, virtual machine; QoS: Quality of Service, quality of service; DQN: Deep Q-Network, deep Q network; DRL: Deep Reinforcement Learning, deep reinforcement learning; DNN: Deep Neural Network, deep neural network.
[0033] Embodiment 1 As Figure 1 shown, this embodiment discloses a cloud computing resource scheduling method based on an elastic scaling double-layer scheduling framework. The method includes the following steps: S1: Based on a three-layer cloud computing market, establish a double-layer scheduling framework; the three-layer cloud computing market includes an infrastructure service provider, a software operation service provider, and an end user; the double-layer scheduling framework includes a task assignment scheduler and a virtual machine automatic scaling scheduler; In this embodiment, consider a typical three-layer cloud computing market, which involves three roles: IaaS provider, SaaS provider, and end user. The IaaS provider uses virtualization technology to abstract the underlying physical resources, and takes the on-demand rental service of virtual machine instances as the basic unit to provide users with the on-demand rental service of virtual machine instances. The SaaS provider rents appropriate types and quantities of virtual machine instances from the IaaS provider according to actual needs, constructs a resource pool and deploys applications to provide convenient services for end users. End users can submit task requests with QoS requirements at any time and place. From the perspective of the SaaS provider, effectively managing the rented resources and allocating tasks to appropriate virtual machine instances has become their main consideration.
[0034] As Figure 1As shown in the figure, the present invention designs a two - layer scheduling framework for SaaS applications, including a virtual machine auto - scaling scheduler and a task assignment scheduler. The two are interdependent and collaborative to achieve their respective goals. Specifically, the auto - scaling scheduler is workload - aware and is responsible for adjusting the virtual machine leasing plan according to resource supply and demand. The task assignment scheduler is QoS - aware and can assign tasks submitted by end - users to appropriate virtual machine instances in the resource pool. By adopting appropriate scheduling strategies and algorithms, and conducting real - time monitoring and dynamic adjustment, efficient virtual machine auto - scaling and task scheduling can be achieved, thereby improving application performance and resource utilization.
[0035] The workload of a SaaS application consists of independent tasks submitted by numerous users, and tasks arrive at the application in real - time. For most transactional SaaS applications, the task workload varies at different times but usually has a certain regularity over a long period (such as a day or a week). Therefore, it is assumed that within a given short - time period the total number of tasks submitted by users is , and the long period with a certain regularity is called . The changes in task load and the attributes of each task cannot be known in advance. The task types submitted by end - users are divided into two categories: memory - optimized and compute - optimized. Based on the above, each user task can be represented as:
[0036] where represents the th user task, represents the serial number of the th task, represents the arrival time of the th task, represents the type of the th task, represents the length of the th task, represents the QoS requirement of the th task.
[0037] In the cloud computing market, IaaS providers typically offer various types of virtual machine instances and classify them according to their configurations and computing capabilities. Common types include general-purpose, graphics processing, compute-optimized, etc. For different types of tasks, different types of virtual machine instances are selected for optimization to achieve optimal performance. The execution speeds of different types of virtual machines for the same task can vary significantly. Therefore, a suitable match can reduce the task execution time. In the study, it is simplified to only two types of virtual machine instances: memory-optimized and compute-optimized. Assume that the resource pool contains a limited number of virtual machine instances rented from the IaaS provider. With the help of an auto-scaling scheduler, the SaaS provider can flexibly adjust the virtual machine rental plan according to the changes in the task workload. The number of virtual machine instances rented within a time period is , and the definition of each virtual machine instance is as follows:
[0038] Among them, represents the th virtual machine, represents the serial number of the th virtual machine, represents the type of the th virtual machine, represents the rental price of the th virtual machine at , represents the processing speed of the th virtual machine for memory-optimized tasks, represents the processing speed of the th virtual machine for compute-optimized tasks.
[0039] For a time period with a cycle of , the rental cost of the virtual machine can be calculated as:
[0040] Among them, represents the rental cost of the virtual machine.
[0041] The task assignment scheduler assigns tasks to specified virtual machine instances according to the scheduling algorithm. Assume that when a new task arrives, it will be quickly processed by the task assignment scheduler and assigned to a specific virtual machine instance. When the assigned virtual machine instance is idle, the task will be executed immediately; conversely, if the specified virtual machine instance is already occupied, the task will enter the waiting queue of that virtual machine instance. The tasks in the queue are executed according to the first-come-first-served principle, ensuring that the virtual machine instance executes tasks in the order of arrival, which means that the task will remain in a waiting state until all earlier tasks in the queue are completed. At the same time, a virtual machine instance can only execute a single task at a time, and the task will not be interrupted during processing.
[0042] Define the task response time as the total time the task stays in the system, which can be expressed as:
[0043] where represents the task execution time, represents the task waiting time.
[0044] Assume the task is assigned to the virtual machine instance, then the execution time is calculated according to the task attributes and the assigned virtual machine instance, and is expressed as:
[0045] The waiting time of the task on the virtual machine instance is determined by the number of tasks in the waiting queue, so can be defined as:
[0046] where represents the number of tasks in the waiting queue of the virtual machine instance , represents the idle time of the virtual machine instance executing the task , represents the task arrival time.
[0047]
[0048] where represents the idle time of the virtual machine instance executing the task , represents the task on the virtual machine instance execution time, represents the task Arrival time
[0049] In the field of cloud computing task scheduling, task response time is a key QoS metric for transactional SaaS applications. For a specific task, if its total time in the system is within the user-acceptable response time, the task is judged to be successful; otherwise, the task is judged to be a failure. On a virtual machine instance, the task The formula for successful execution is as follows:
[0050] where represents the QoS requirement of task
[0051] To maximize the advantages of SaaS providers, the two-layer scheduling framework needs to reduce the total rental cost of virtual machine instances while increasing the task success rate. To achieve this goal, a workload-aware virtual machine auto-scaling scheduler and a QoS-aware task allocation scheduler need to work together. This way, the number of rented virtual machine instances can be flexibly adjusted and the allocation of tasks to virtual machines can be optimized.
[0052] S2: The task allocation scheduler defines the first agent, and determines the state, action, and reward function of the first agent by collecting the task characteristics and virtual machine usage data of the current cloud computing platform; In this embodiment, the QoS-aware task allocation scheduler in the two-layer scheduling framework focuses on task allocation, that is, how to select the most suitable virtual machine instance from the resource pool for the current task. The following is a reinforcement learning model (RL Model) for solving the online task scheduling problem.
[0053] State and action: The task allocation scheduler records the task characteristics and the usage of virtual machine instances in the resource pool, which can be expressed as:
[0054] where represents the state space of the th task of the task allocation scheduler, represents the number of virtual machine instances in the current resource pool, represents the waiting time when allocated to the th virtual machine instance.
[0055] The action of the th task in the task allocation scheduler to allocate the task to a virtual machine
[0056] Among them, represents the th virtual machine.
[0057] Reward function: For transactional SaaS applications, task response time is the main indicator determining application performance and user satisfaction. Therefore, set the reward function as follows:
[0058] Among them, represents the reward function value of the th task of the task assignment scheduler, represents the total time the task stays on the virtual machine instance, represents the actual execution time of the task. Therefore, the smaller the percentage of waiting time in the total time, the better the allocation result.
[0059] S3: The virtual machine auto-scaling scheduler defines a second agent, and determines the state, action, and reward function of the second agent by collecting the workload and resource utilization data of the current cloud computing platform; In this embodiment, in the two-layer scheduling framework, the workload-aware virtual machine auto-scaling scheduler is responsible for continuously adjusting the number of leased virtual machine instances according to the change of task workload and the utilization rate of currently leased resources, and regularly adjusting the virtual machine leasing plan to achieve elastic resource allocation. When the number of tasks is small within a period of time, the number of virtual machine instances can be appropriately reduced to avoid resource waste and reduce the leasing cost. On the contrary, when the virtual machine instances in the resource pool are saturated, leasing more virtual machine instances can ensure the success rate of user tasks and improve user satisfaction. The reinforcement learning model (RL Model) of the virtual machine auto-scaling scheduler is introduced as follows.
[0060] State and action: The state when the virtual machine auto-scaling scheduler executes the th virtual machine leasing adjustment plan can be expressed as:
[0061] Among them, represents the number of tasks submitted by users in the recent period, represents the number of leased virtual machine instances in the recent period, is a specific time within the workload cycle, and respectively represent the average response time and success rate of all tasks during this period, is the resource utilization rate of all virtual machine instances in the resource pool during the
[0062] In the action selection stage, the agent selects to increase or decrease the number of leased virtual machine instances according to the current state. The action set is the range of changes in the number of virtual machine instances, which can be expressed as:
[0063] where, represents the maximum range of virtual machine instances that can be increased or decreased. For example, the action means increasing virtual machine instances, and the action means deleting virtual machine instances.
[0064] Reward function: Adjusting the number of leased virtual machine instances through the virtual machine auto-scaling scheduler will not only affect resource utilization but also affect the response time of tasks. The optimization goal of the two-layer scheduling framework is to reduce the cost of leased virtual machine instances while ensuring service quality. Therefore, the design of the reward function needs to consider both the success rate of tasks and the utilization rate of virtual machine instances, and the reward function is defined as follows:
[0065] where, represents the resource utilization rate during period , represents the penalty that needs to be given as a reward when the success rate is lower than a certain threshold, represents the weight of the penalty on the reward, and the penalty is defined as:
[0066] where, represents the task success rate during period .
[0067] S4: Use the two-layer scheduling algorithm to solve the task allocation decision of the first agent, and use the two-layer scheduling algorithm to solve the virtual machine adjustment decision of the second agent.
[0068] In a three-tier cloud market, the task workload of SaaS applications is unpredictable and dynamically changing. Assigning continuously arriving tasks submitted by users to a limited number of virtual machine instances and adjusting the number of leased virtual machine instances according to resource usage and workload changes are two major challenges faced by SaaS providers. SaaS providers can only monitor the status of the virtual machine instances they lease in the resource pool and the current task workload, but cannot predict the future arrival time and resource requirements of tasks. Therefore, intelligent online scheduling algorithms that can adaptively solve such complex problems are needed. Compared with traditional task scheduling that relies on predefined rules and static models, DRL algorithms can more effectively adapt to changes in dynamic environments without any prior knowledge of the environment.
[0069] For the design of a QoS-aware task assignment scheduler and a workload-aware virtual machine auto-scaling scheduler, a two-layer scheduling algorithm for SaaS providers based on DRL is proposed. By continuous learning and improvement, it can significantly improve cloud resource utilization and system performance. The two-layer scheduling algorithm for SaaS providers based on DRL consists of two layers of DQN algorithms. The first layer of DQN algorithm is responsible for task assignment, and the second layer of DQN algorithm is responsible for dynamically adjusting the number of leased virtual machine instances. These two DQN layers interact with each other, and each DQN algorithm contains two collaborative stages: the online decision-making stage and the offline training stage.
[0070] Furthermore, in the two-layer scheduling framework, the task assignment scheduler and the virtual machine auto-scaling scheduler jointly optimize the system resource utilization and the execution efficiency of tasks through a real-time collaboration and feedback mechanism. The task scheduler drives the virtual machine demand. The task scheduler dynamically assigns tasks to available virtual machines according to the status of the current task queue. When the virtual machine scheduler detects a shortage of resources, it will trigger a request to expand the capacity of new virtual machines to relieve the resource bottleneck; the virtual machine scheduler feeds back the resource supply. The virtual machine scheduler monitors the real-time status of the resource pool and dynamically scales the number of virtual machines. When expanding, the information of newly leased virtual machines will be synchronized to the task scheduler to update its available resource list. When scaling down, the virtual machine scheduler needs to ensure in advance that the executing tasks on the virtual machines are released to ensure seamless switching; closed-loop feedback optimization. The allocation strategy of the task scheduler directly affects the utilization rate of virtual machines, which in turn triggers the response of the virtual machine scheduler. If the task scheduler evenly distributes tasks to each node, the virtual machine scheduler can maintain a stable scale; if the task scheduler causes some nodes to be overloaded, the virtual machine scheduler will quickly supplement resources and guide the task scheduler to reallocate tasks. The two form a dynamic balance through periodic metric synchronization to avoid resource waste or performance degradation. Through this hierarchical cooperation mechanism, the system can adapt to dynamic loads and at the same time meet the core requirements of high performance, high availability and low cost.
[0071] Such as Figure 2As shown, in the online decision-making stage: This stage is responsible for making task-to-virtual machine allocation decisions and virtual machine adjustment decisions for the task allocation scheduler and the virtual machine auto-scaling scheduler respectively. Specifically, for the task allocation scheduler, when a task submitted by a new user arrives at the system, the DRL agent will observe the current state of the environment and use the deep neural network DNN to calculate the Q-values of all leased virtual machine instances in the current state. A virtual machine instance will be allocated according to the "ϵ-greedy" policy in the resource pool to execute this task. In addition, the DRL agent will receive an immediate reward for the task scheduling decision. For the virtual machine auto-scaling scheduler, the attributes of tasks and virtual machine instances over a period of time will be recorded and used as part of the state. The DRL agent adjusts the number of virtual machine instances according to the current state and the output value of the DNN.
[0072] Offline training stage: The deep neural network is the core component in DRL responsible for approximating the value function, policy function, and model function. The deep neural network uses a multi-layer neuron structure and non-linear activation functions to process high-dimensional complex state spaces. Due to its strong generalization ability and flexible adaptability, the DNN has significant advantages in solving the proposed problems. In this embodiment, two different DNNs are constructed for the two-layer DQN algorithm, and the two DNNs are different in the number of network layers and the number of neurons in each layer. In the offline training stage, past selections and corresponding results are used to train the underlying DNN to continuously correct and update the state-action value function. To improve the stability and performance of the DQN algorithm, experience replay and fixed target network techniques are adopted. The experience replay technique stores past experiences in a replay buffer and randomly samples them to train the agent. This method breaks the connection between consecutive experiences and improves the efficiency and stability of the learning process. The fixed target network technique involves maintaining a separate target network whose parameters are updated less frequently than the main network. The target network is used to generate target Q-values. This technique can reduce oscillations and divergences during the training process.
[0073] Furthermore, due to the fully connected form of the deep neural network DNN, the solution process may slightly lead to overfitting or getting stuck in local optima. To prevent the algorithm from overfitting or underfitting, an L2 regularization term is added to the loss function of the network. The role of the loss function is to ensure that the output of the model is consistent with the expected target. Regularization improves the generalization ability of the model by adding penalties to the loss function, greatly enhancing the robustness of the DQN algorithm.
[0074] As Figure 1As shown in the figure, the two-layer scheduling framework is divided into three parts, namely, end users, IaaS providers and SaaS providers. Among them, end users are responsible for submitting tasks, IaaS providers are responsible for providing physical equipment, and SaaS providers divide the physical equipment provided by IaaS providers into resource pools consisting of several virtual machine instances, and allocate tasks submitted by users; the controller is responsible for managing all parts of SaaS, including the user portal, task allocation scheduler, virtual machine automatic scaling scheduler and resource pool.
[0075] The specific steps are: the user submits the task, and the task allocation scheduler generates status according to the task ( ), the first agent according to the state Get the action assigned to a virtual machine instance ( ), and then according to the arrival and Get reward value ( ), the first agent according to the reward value Learning is carried out to continuously evolve and obtain the optimal task allocation strategy. In the process of task allocation, the second-layer virtual machine automatic scaling scheduler is executed once every set time period (set to 5 minutes in this embodiment). The second agent collects changes in multiple aspects of tasks within the set time period and then generates a state. ( ), the second agent obtains the state Select the number of virtual machines to increase or decrease within the next set time period ( ), and finally according to the obtained and Get reward value ( ), similarly, the second agent learns according to the reward value to continuously evolve and obtain the optimal resource adjustment strategy.
[0076] It should be noted that , , Refers to all states, actions, and reward values of the first agent in the reinforcement learning process, not to a specific task; similarly, , and It refers to all states, actions, and reward values of the second agent during the reinforcement learning process, not a specific time period.
[0077] The test experiment is described below.
[0078] In the test experiment of the present invention, simulated workload and actual workload are used for experiments respectively. At each auto-scaling decision moment, that is, every five minutes, the second layer of the two-layer scheduling framework will generate a virtual machine adjustment decision to change the number of virtual machine instances in the resource pool. First, the two-layer scheduling algorithm is applied to process the simulated workload.
[0079] Figure 3 Shows the experimental results of the two-layer scheduling framework in a simulated environment for 8 days. The x-axis represents the number of days. The y-axis represents the changes in the number of tasks and virtual machines within every 5 minutes. This figure describes the changes in the task workload and the number of leased virtual machine instances using the two-layer scheduling framework. As Figure 3 shown, the learning in the first three days usually involves 8 to 10 virtual machine instances. On the 4th to 5th days of learning, the number of virtual machine instances begins to gradually adapt to the trend of the load, and the change in virtual machine instances begins to adapt to the change in tasks. In the last three days, the change in the number of virtual machine instances has fully adapted to the change in the workload. When the load is low, the number of virtual machine instances is generally 3 to 5; when the load is high, the number of virtual machine instances is generally 8 to 10.
[0080] Table 1 Comparison results of the two-layer scheduling framework in a simulated environment
[0081] Table 1 presents the specific experimental results of each algorithm in a simulated environment. It can be found that although the average response time of the two-layer DQN is slightly longer, compared with the single-layer DQN-10VM algorithm with a fixed 10 virtual machine instances, the two-layer DQN algorithm has a higher success rate and lower leasing cost, and it reduces the leasing cost by approximately 43%. Compared with the single-layer DQN algorithm with a fixed 3 virtual machine instances, the two-layer DQN algorithm is significantly superior to the DQN-3VM algorithm in terms of average response time and success rate, and is also superior to the DQN-3VM algorithm in terms of leasing cost. In the comparison of the two-layer algorithms, the DQN algorithm is used in the first layer and the second layer respectively. The two-layer DQN algorithm not only improves the average response time and success rate, but also saves 24.8% and 21.1% in leasing cost compared with the FCFS-DQN algorithm and the DQN-ARIMA algorithm respectively.
[0082] To further verify the performance of the two-layer scheduling framework, a real NASA dataset is used for testing. The NASA workload trace used contains all HTTP requests issued by NASA servers located at the Kennedy Space Center in Florida during August 1995. Table 2 records the experimental results of the two-layer DQN algorithm and other comparison algorithms throughout the month.
[0083] Table 2 Comparison Results of the Double-Layer Scheduling Framework in the Real Environment
[0084] As can be seen from Table 2, in terms of the average task response time and success rate, the performance of the single-layer DQN-10VM algorithm with 10 fixed virtual machine instances and the double-layer DQN algorithm is significantly better than that of other comparison algorithms. However, the single-layer DQN algorithm cannot dynamically adjust the number of leased virtual machine instances according to the change of the workload, resulting in resource waste and increased leasing costs. The double-layer DQN algorithm can balance the optimization goals of QoS requirements and costs. Although the average response time and success rate of the double-layer DQN algorithm are slightly lower than those of the single-layer DQN algorithm with 10 fixed virtual machine instances, the cost is saved by 43.3%.
[0085] Since there are many tasks in a month, Figure 4 only the changes in the workload and the number of leased virtual machine instances in the first 8 days of using the double-layer DQN algorithm in the actual scenario are shown. Similar to Figure 3 the learning process shown, the change in the number of leased virtual machine instances also gradually adapts to the workload trend.
[0086] Efficient task scheduling and elastic resource supply are two major challenges faced by application providers. In view of this problem, the present invention proposes an intelligent double-layer scheduling framework. The algorithm uses double-layer scheduling based on DRL, where the first-layer DQN algorithm is responsible for allocating tasks to appropriate virtual machine instances, and the second-layer DQN algorithm is responsible for dynamically adjusting the number of virtual machine instances to reduce the cost of leasing virtual machine instances. The single-layer DQN algorithm and the double-layer DQN algorithm are respectively compared with other comparison algorithms. The experimental results show that this method is effective in ensuring a high success rate and reducing leasing costs.
[0087] In some other embodiments, for the second-layer elastic resource expansion, instead of horizontal resource expansion, the resource configuration of a single instance is dynamically adjusted through vertical resource expansion, or a predictive scaling technology is adopted to predict future loads through historical data, adjust resources in advance, reduce response latency, and avoid over-provisioning. Therefore, the elastic scaling double-layer scheduling framework based on deep reinforcement learning proposed by the present invention may be modified into an elastic resource scheduler based on vertical resource expansion or predictive resource scaling, so as to implement the scheduling framework based on deep reinforcement learning proposed by the present invention from another perspective.
[0088] Embodiment 2 This embodiment discloses a cloud computing resource scheduling system based on an elastic scaling double-layer scheduling framework, including: A framework construction module, which is configured to: establish a two-layer scheduling framework based on a three-layer cloud computing market; the three-layer cloud computing market includes infrastructure service providers, software operation service providers, and end users; the two-layer scheduling framework includes a task allocation scheduler and a virtual machine automatic scaling scheduler; A first scheduling module, which is configured to: the task allocation scheduler defines a first agent, and determines the state, actions, and reward function of the first agent by collecting the task characteristics and virtual machine usage data of the current cloud computing platform; A second scheduling module, which is configured to: the virtual machine automatic scaling scheduler defines a second agent, and determines the state, actions, and reward function of the second agent by collecting the workload and resource utilization data of the current cloud computing platform; A scheduling decision module, which is configured to: use a two-layer scheduling algorithm to solve the task allocation decision of the first agent, and use a two-layer scheduling algorithm to solve the virtual machine adjustment decision of the second agent.
[0089] Embodiment III The purpose of this embodiment is to provide a computing device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method in Embodiment I are implemented.
[0090] Embodiment IV The purpose of this embodiment is to provide a computer-readable storage medium. A computer-readable storage medium stores a computer program, and when the program is executed by a processor, the steps of the method in Embodiment I are executed.
[0091] The steps involved in the devices in the above Embodiments III and IV correspond to those in Method Embodiment I. For specific implementation manners, reference may be made to the relevant description part of Embodiment I. The term "computer-readable storage medium" should be understood to include a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0092] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by a computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. The present invention is not limited to any specific combination of hardware and software.
[0093] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0094] Although the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A cloud computing resource scheduling method based on an elastically scalable two-layer scheduling framework, characterized in that: include: Based on the three-tier cloud computing market, a two-tier scheduling framework is established; The three-layer cloud computing market includes infrastructure service providers, software operation service providers and end users; the two-layer scheduling framework includes a task allocation scheduler and a virtual machine automatic scaling scheduler; The task allocation scheduler defines a first agent, and determines the state, action, and reward function of the first agent by collecting task characteristics and virtual machine usage data of the current cloud computing platform; The virtual machine automatic scaling scheduler defines a second agent, and determines the state, action and reward function of the second agent by collecting the workload and resource utilization data of the current cloud computing platform; A two-layer scheduling algorithm is used to solve the task allocation decision of the first agent, and a two-layer scheduling algorithm is used to solve the virtual machine adjustment decision of the second agent.
2. The cloud computing resource scheduling method based on the elastic scaling double-layer scheduling framework according to claim 1 is characterized in that: The task response time in the two-layer scheduling framework is expressed as: in, represents the task response time, Indicates the task execution time. Indicates the task waiting time.
3. The cloud computing resource scheduling method based on the elastic scaling double-layer scheduling framework according to claim 1 is characterized in that: The state of the first agent is expressed as: The action is represented by: in, Represents the task allocation scheduler The state space of a task, Represents the number of virtual machine instances in the current resource pool. Indicates that the The waiting time for each virtual machine instance is Indicates the first The action of assigning a task to a virtual machine, Indicates virtual machines.
4. The cloud computing resource scheduling method based on the elastic scaling double-layer scheduling framework according to claim 3 is characterized in that: The reward function of the first agent is expressed as: in, Represents the task allocation scheduler The reward function value of a task, Indicates the total time that the task stays on the virtual machine instance. Indicates the actual execution time of the task.
5. The cloud computing resource scheduling method based on the elastic scaling double-layer scheduling framework according to claim 1, characterized in that: The state of the second agent is expressed as: The action is represented by: in, Indicates that the virtual machine auto-scaling scheduler executes The status of the secondary VM lease adjustment plan. Indicates the number of tasks submitted by users in the recent period. Indicates the number of virtual machine instances rented in the recent period. is a specific time within the workload cycle, and Respectively represent the average response time and success rate of all tasks in the time period, for The resource utilization of all virtual machine instances in the resource pool during the period, Indicates the maximum range of virtual machine instances that can be increased or decreased.
6. The cloud computing resource scheduling method based on the elastic scaling double-layer scheduling framework according to claim 5, characterized in that: The reward function of the second agent is defined as: in, Representation cycle Resource utilization during the period, Indicates the penalty that needs to be given when the success rate is below a certain threshold. Represents the weight of punishment on reward.
7. The cloud computing resource scheduling method based on the elastic scaling double-layer scheduling framework according to claim 1, characterized in that: The two-layer scheduling algorithm consists of two layers of DQN algorithms, the first layer of DQN algorithm is responsible for task allocation, and the second layer of DQN algorithm is responsible for virtual machine adjustment.
8. A cloud computing resource scheduling system based on an elastically scalable two-layer scheduling framework, characterized in that: include: A framework building module, which is configured to: establish a two-tier scheduling framework based on a three-tier cloud computing market; The three-layer cloud computing market includes infrastructure service providers, software operation service providers and end users; the two-layer scheduling framework includes a task allocation scheduler and a virtual machine automatic scaling scheduler; The first scheduling module is configured as follows: a task allocation scheduler defines a first agent, and determines the state, action and reward function of the first agent by collecting task characteristics and virtual machine usage data of the current cloud computing platform; The second scheduling module is configured as follows: the virtual machine automatic scaling scheduler defines a second agent, and determines the state, action and reward function of the second agent by collecting workload and resource utilization data of the current cloud computing platform; The scheduling decision module is configured to: use a double-layer scheduling algorithm to solve the task allocation decision of the first intelligent agent, and use a double-layer scheduling algorithm to solve the virtual machine adjustment decision of the second intelligent agent.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the cloud computing resource scheduling method based on an elastically scalable two-layer scheduling framework as described in any one of claims 1 to 7 are implemented.
10. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the cloud computing resource scheduling method based on an elastically scalable two-layer scheduling framework are implemented as described in any one of claims 1-7.
Citation Information
Patent Citations
Virtual computing resource dynamic management system of cloud computing service platform
CN102681899A
Cloud computing based virtual machine two-level optimization scheduling management platform
CN106161640A
Cloud resource adaptive configuration method and system based on depth deterministic strategy
CN113641445A
Container cluster resource scheduling method and system based on deep reinforcement learning
CN114443249A
Cloud resource dynamic scaling method and system
CN116126534A
Cited By
Virtual machine resource scheduling method and device, equipment and storage medium
CN120803616A
Kubernetes-based stateful elastic computing framework implementation method and system
CN121455604A