Computing resource scheduling method and device, electronic equipment and storage medium
By using the average resource utilization rate on the cloud computing platform to determine the resource balancing benchmark and performing hot migration of virtual machines, the problem of load imbalance after cold start upgrade of computing nodes within the host group is solved, achieving efficient resource utilization and business continuity.
Patent Information
- Application Number
- CN202510946098.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-11-04
AI Technical Summary
In cloud computing platforms, after multiple computing nodes within a host group complete a cold start upgrade, some computing nodes cannot adaptively achieve load balancing, resulting in low efficiency in the allocation of computing resources.
The system determines the resource balance benchmark by calculating the average resource utilization rate, automatically identifies computing nodes with excessive resource concentration, and performs hot migration operations to migrate virtual machines from the source computing node with high resource utilization to the target computing node with low utilization until resource balance is achieved.
It achieves load balancing among computing nodes, avoids resource waste, reduces business interruption time, improves business continuity and resource utilization efficiency, reduces maintenance workload, and enhances system manageability and stability.
Smart Images

Figure CN120892183A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of cloud computing, in particular to a computing resource scheduling method and device, an electronic device and a storage medium. BACKGROUND
[0002] With the rapid development of cloud computing technology, infrastructure as a service has become a core component of enterprise IT architecture, providing flexible and scalable computing resource services for users. In the infrastructure as a service environment, in order to maintain the security and functionality of the system, it is necessary to periodically upgrade or patch the computing nodes. However, the traditional upgrade method is often accompanied by service interruption, especially cold upgrade, which requires downtime operation, which may seriously affect user experience and business continuity.
[0003] In a cloud computing platform, computing resources are usually managed in the form of host groups, each host group containing multiple computing nodes that jointly carry a large number of virtual machines and applications. However, in the prior art, when multiple computing nodes in a host group complete cold start upgrade, the system faces a significant technical challenge: some computing nodes cannot adaptively complete load balancing, resulting in low efficiency of computing resource allocation.
[0004] In view of the above problems, no effective solution has been proposed so far. SUMMARY
[0005] The embodiments of the present application provide a computing resource scheduling method and device, an electronic device and a storage medium, to at least solve the technical problem in the prior art that after multiple computing nodes of a host group complete cold start upgrade, some computing nodes cannot adaptively complete load balancing, resulting in low efficiency of computing resource allocation.
[0006] According to an aspect of an embodiment of the present application, a computing resource scheduling method is provided, comprising: determining a resource balancing reference value corresponding to N computing nodes according to an average value of resource usage rates of the N computing nodes after the N computing nodes complete upgrade, wherein N is an integer greater than 1; sorting the N computing nodes according to the resource usage rates of each computing node, and determining a source computing node and a target computing node from the N computing nodes according to the sorting result, wherein the resource usage rate of the source computing node is higher than that of the target computing node; selecting a virtual machine to be migrated on the source computing node according to the resource usage rate of the source computing node and the resource balancing reference value; and performing a hot migration operation to migrate the virtual machine to be migrated from the source computing node to the target computing node.
[0007] Optionally, after each execution of the hot migration operation, the N computing nodes are reordered according to the latest resource usage of each computing node, and the determination of the source computing node and the target computing node, the selection of the virtual machine to be migrated, and the hot migration operation are performed again according to the reordering result until the resource balance reference value corresponding to the N computing nodes is updated to be less than the preset threshold.
[0008] Optionally, the resource usage of each computing node includes at least one of the following performance indicators:
[0009] The first performance indicator is used to represent the CPU usage of each computing node.
[0010] The second performance indicator is used to represent the memory usage of each computing node.
[0011] The third performance indicator is used to represent the node ratio of each computing node, wherein the node ratio is used to reflect the usage degree of each computing node relative to the maximum resource capacity of the computing node.
[0012] Optionally, the resource balance reference value corresponding to the N computing nodes is determined according to the average value of the resource usage of the N computing nodes, including: in a predetermined time interval, at least the CPU usage, the memory usage and the node ratio are monitored and recorded in real time from the N computing nodes; the collected different types of resource usage are standardized; wherein the standardization is used to convert all data into data under the same measurement scale; the weight of each resource usage is obtained, wherein the weight of each resource usage is used to reflect the relative importance of each resource in the process of determining the resource balance reference value; according to the standardized resource usage of the N computing nodes and the weight of each resource usage, the weighted average value of various resource usages is calculated to obtain the resource balance reference value.
[0013] Optionally, the virtual machine to be migrated on the source computing node is selected according to the resource usage of the source computing node and the resource balance reference value, including: detecting the difference between the resource usage of the source computing node and the resource balance reference value; according to the difference, the resource usage data of the source computing node and the resource consumption data of each virtual machine on the source computing node, one or more virtual machines on the source computing node are selected as the virtual machine to be migrated, wherein the selection process of the virtual machine to be migrated follows one or a combination of the minimum resource impact principle, the maximum resource release principle or the service priority principle.
[0014] Optionally, a minimum resource impact principle is used to select a virtual machine that has the minimum impact on the source computing node and the target computing node after migration; a maximum resource release principle is used to select a virtual machine that has the maximum resource occupation or is closest to the part of the source computing node resource load exceeding the resource balance reference value for migration; and a service priority principle is used to preferentially migrate a virtual machine of a low-priority service.
[0015] Optionally, the live migration operation further includes performing a rollback operation when the migration fails, wherein the rollback operation is used to restore the virtual machine that fails in migration to a state before migration.
[0016] According to another aspect of the embodiments of the present application, a computing resource scheduling apparatus is further provided, and the apparatus includes: a first determining unit configured to determine a resource balance reference value corresponding to N computing nodes according to an average value of resource usage rates of the N computing nodes after the N computing nodes complete upgrading, wherein N is an integer greater than 1; a second determining unit configured to sort the N computing nodes according to the resource usage rates of the computing nodes, and determine a source computing node and a target computing node from the N computing nodes according to a sorting result, wherein the resource usage rate of the source computing node is higher than that of the target computing node; a selecting unit configured to select a virtual machine that needs to be migrated on the source computing node according to the resource usage rate of the source computing node and the resource balance reference value; and a migrating unit configured to perform a live migration operation to migrate the virtual machine that needs to be migrated from the source computing node to the target computing node.
[0017] According to another aspect of the embodiments of the present application, a computer readable storage medium is further provided, and the computer readable storage medium stores a computer program, wherein when the computer program is running, the computer readable storage medium makes a device on which the computer readable storage medium is located perform the computing resource scheduling method.
[0018] According to another aspect of the embodiments of the present application, an electronic device is further provided, and the electronic device includes one or more processors and a memory, and the memory is configured to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the computing resource scheduling method.
[0019] In the present application, after the N computing nodes complete the upgrade, the resource balancing reference value corresponding to the N computing nodes is determined according to the average of the resource utilization rates of the N computing nodes, where N is an integer greater than 1. Then, the N computing nodes are sorted according to the resource utilization rate of each computing node, and the source computing node and the target computing node are determined from the N computing nodes according to the sorting result, where the resource utilization rate of the source computing node is higher than that of the target computing node. Subsequently, the virtual machine to be migrated on the source computing node is selected according to the resource utilization rate of the source computing node and the resource balancing reference value. Finally, the hot migration operation is performed to migrate the virtual machine to be migrated from the source computing node to the target computing node.
[0020] From the above, according to the technical solution of the present application, the resource balancing reference value is determined by calculating the average of the resource utilization rates, and then according to the reference value and the actual resource utilization of each node, the scheduling system can automatically identify the computing node (source computing node) where the resources are excessively concentrated. Subsequently, the appropriate virtual machine is selected for hot migration, thereby balancing the load among the computing nodes, avoiding resource waste, and maximizing the utilization of resources. The present application adopts the method of automatic sorting and determining the source computing node and the target computing node, which can dynamically migrate the resource-intensive virtual machine from the overloaded computing node to the computing node with relatively loose resources, effectively reducing the difference in resource pressure between nodes, and achieving load balancing.
[0021] In addition, the hot migration of the virtual machine can transfer between computing nodes without stopping the service, reducing the service interruption time caused by resource reallocation, improving the business continuity, especially in critical business scenarios, which can significantly improve the service quality and user experience. Moreover, according to the technical solution of the present application, the entire resource balancing process is automated, and the resource reallocation after the computing node upgrade can be completed without human intervention, greatly reducing the operation and maintenance workload, reducing the errors caused by human operation, and improving the manageability and stability of the system. The automatic balancing mechanism can quickly respond to changes in the resource utilization of the computing node, and even in the case of fluctuating node state after upgrade, it can quickly adjust to keep the resources in a balanced state, thereby improving the overall response speed and reliability of the cloud computing platform.
[0022] In summary, the technical solution provided by the present application effectively solves the problem of uneven resource allocation after the cold start upgrade of the computing nodes in the host group through automatic resource balancing operation, improves the resource utilization efficiency and business continuity, reduces the operation and maintenance burden, and enhances the response speed and reliability of the cloud computing system. BRIEF DESCRIPTION OF DRAWINGS
[0023] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and together with the description serve to explain the application. In the drawings:
[0024] Figure 1 is a flow chart of an optional computing resource scheduling method according to an embodiment of the application;
[0025] Figure 2 is a flow chart of an optional virtual machine migration method according to an embodiment of the application;
[0026] Figure 3 is a schematic diagram of an optional computing resource scheduling apparatus according to an embodiment of the application. DETAILED DESCRIPTION
[0027] In order to enable persons skilled in the art to better understand the application scheme, the technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only a part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those of ordinary skill in the art without creative work should fall within the protection scope of the application.
[0028] It should be noted that the terms "first", "second", and the like in the specification and claims of the application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0029] It should be noted that the information collected by the present application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards in the relevant region, necessary security measures are taken, and the public order is not violated, and appropriate operation portal is provided for user to choose authorization or refusal. For example, the system and related users or institutions are provided with an interface, and before obtaining the relevant information, the interface needs to send a request to the aforementioned user or institution, and after receiving the consent information feedback from the aforementioned user or institution, the relevant information is obtained.
[0030] The following is an explanation of some professional terms that appear in the embodiments of the present application:
[0031] Infrastructure cloud: a cloud computing service model that provides virtualized computing resources as a service. Users can access and use these resources through the Internet without the need to purchase and maintain physical hardware. Infrastructure cloud services usually include the following aspects: servers (cloud service providers provide virtual servers, users can configure CPU, memory and storage resources according to needs), storage (provides network-connected storage services, users can store data and expand as needed), network (provider provides network resources, including virtual private network (VPN), firewall and load balancing, etc.), operating system (users can choose to install different operating systems to meet specific application requirements), security (provides security measures such as identity authentication, data encryption and network security, etc.).
[0032] Cold patch: in the infrastructure cloud environment, in order to ensure system security, stability and the latest function, version upgrade needs to be carried out on the existing environment. Version upgrade can be divided into two categories: hot upgrade and cold upgrade. Hot upgrade allows updating applications without restarting the system, which has the least impact on business operation, but not all upgrades can use this method. Cold upgrade needs to change the system state or fix global variables, and restart the system to take effect, which requires more planning and coordination to minimize the impact on users.
[0033] Rolling upgrade: a method of upgrading systems or applications step by step without interrupting service, which can minimize the impact on users during the upgrade process, especially suitable for cloud services and distributed systems.
[0034] According to the embodiments of the present application, an embodiment of a computing resource scheduling method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0035] In an alternative embodiment, a scheduling system can be used as the execution subject of the computing resource scheduling method in the embodiments of the present application. The scheduling system can be a software system or a combination of software and hardware embedded system. Moreover, those skilled in the art should know that, in addition to using the scheduling system as the execution subject, the execution subject of the computing resource scheduling method of the present application can also be other devices, apparatuses, etc., and the present application does not particularly limit the specific forms of the execution subject.
[0036] Figure 1 is a flowchart of an alternative computing resource scheduling method according to the embodiments of the present application, as shown in Figure 1 , the method comprises the following steps:
[0037] Step S101: After the N computing nodes complete the upgrade, the resource balancing reference value corresponding to the N computing nodes is determined according to the average value of the resource usage rates of the N computing nodes.
[0038] In step S101, N is an integer greater than 1.
[0039] Optionally, after the N computing nodes complete the upgrade (especially the cold upgrade), since the upgrade process can involve the restart of the nodes and the update of other resource states, this can cause certain differences in the resource usage of the nodes at the initial running time after the upgrade. In order to ensure that the cloud computing platform can run efficiently and stably, while avoiding the situation that some nodes are overloaded and other nodes are idle, measures need to be taken immediately after the upgrade to automatically balance the resources.
[0040] First, according to the technical solution of the present application, after the N computing nodes complete the upgrade, the scheduling system automatically starts collecting the resource usage information of the N computing nodes. The resource usage information includes but is not limited to CPU usage rate, memory usage rate, disk I / O usage rate, network bandwidth usage rate and other key indicators. For each computing node, the scheduling system records the resource usage rate data within a period of time (such as the last hour), and takes the average value to obtain the comprehensive resource usage rate of the node.
[0041] Secondly, the scheduling system aggregates the resource usage rates of the N computing nodes and calculates the average value thereof. This average value reflects the average level of the resource usage of the entire computing cluster at the initial stage after the upgrade.
[0042] Optionally, the resource balancing reference value is an ideal value or range, which is used to guide the subsequent resource balancing operation. For example, the scheduling system can define the resource balancing reference value by taking the average of the resource usage of the N computing nodes as the basis, considering a certain tolerance range (such as ± 5%), and defining the resource balancing reference value. The setting of this reference value should not only consider the maximization of resource utilization, but also ensure that the performance of a single node will not be degraded or unstable due to excessive resource consumption.
[0043] After the resource balancing reference value is defined, the scheduling system will use it as a standard to evaluate the resource usage of each computing node, identify computing nodes with high or low resource usage, and then perform corresponding resource adjustments, such as virtual machine migration, until the resource usage of all computing nodes tends to approach this reference value.
[0044] In step S102, the N computing nodes are sorted according to the resource usage of each computing node, and the source computing node and the target computing node are determined from the N computing nodes according to the sorting result.
[0045] In step S102, the resource usage of the source computing node is higher than that of the target computing node.
[0046] Optionally, after the N computing nodes are upgraded, the scheduling system starts to collect resource usage data of each computing node in real time or periodically, including CPU, memory, disk I / O, and network bandwidth, etc. These data reflect the current load state of the node. According to the collected resource usage data, the scheduling system can sort the N computing nodes, and the sorting standard is usually based on the high and low of the resource usage. In this way, the node with the highest resource usage will be at the forefront of the sequence, and the node with the lowest usage will be at the end of the sequence.
[0047] After sorting, the scheduling system will select the source computing node and the target computing node from the two ends of the sequence, respectively. The source computing node usually refers to the node whose resource usage is significantly higher than the resource balancing reference value, while the target computing node refers to the node whose resource usage is lower than the resource balancing reference value. Among them, the main purpose of determining the source computing node is to select virtual machines from this node for migration to reduce its resource usage; the target computing node is the node that is ready to receive these virtual machines, the purpose is to improve its utilization and balance the load of the entire host group.
[0048] It should be noted that when determining the source computing node and the target computing node, not only the absolute value of the resource usage should be considered, but also the resource type (such as the migration priority of CPU-intensive resources and memory-intensive resources may be different) and the current running state of the node (such as whether there is a critical task in progress).
[0049] Assuming there are N = 12 computing nodes in a cloud computing platform, after upgrading, the system collects and calculates the resource usage of these nodes. The sorting result shows that the resource usage of node A is the highest, reaching 80%, while the resource usage of node Z is the lowest, only 30%. The scheduling system can mark node A as the source computing node because its resource usage is much higher than the resource balance benchmark value (assuming 50%). At the same time, the scheduling system can determine node Z with lower resource usage as the target computing node to receive virtual machines migrated from node A.
[0050] Then, the scheduling system will analyze the resource consumption of virtual machines on node A and select those virtual machines that can best alleviate the resource pressure of node A for hot migration to node Z. The migration process will take into account the resource capacity of node Z to ensure that node Z will not be suddenly overloaded after migration.
[0051] Through such sorting and node determination mechanism, the scheduling system can effectively identify computing nodes with uneven resource usage and adjust resource allocation through hot migration of virtual machines to ultimately achieve automatic balancing of resources. This mechanism not only improves resource utilization, but also enhances system stability and response capability, providing higher quality cloud computing services for users.
[0052] Step S103, according to the resource usage of the source computing node and the resource balance benchmark value, selecting the virtual machines on the source computing node that need to be migrated.
[0053] Step S104, performing hot migration operation to migrate the virtual machines that need to be migrated from the source computing node to the target computing node.
[0054] Optionally, in cloud computing resource management, achieving resource balancing of computing nodes is the key to ensuring system performance and stability. After the system identifies the source computing node with high resource usage and the target computing node with low resource usage, the next step is to intelligently select the virtual machines on the source computing node that need to be migrated to achieve resource balancing. This selection process needs to be based on the resource usage of the source computing node and the resource balance benchmark value to ensure the rationality and efficiency of the migration decision.
[0055] For example, the scheduling system first needs to conduct an in-depth analysis of the resource usage of the source computing node, which includes but is not limited to CPU usage, memory usage, disk I / O, and network bandwidth, and other core indicators. In the analysis, the scheduling system can pay special attention to those resource types whose usage exceeds the resource balance benchmark value, for example, if the CPU usage of the source computing node is as high as 90%, and the resource balance benchmark value is set to 65%, then the CPU-intensive virtual machine becomes the primary migration object. After determining the resource type that needs to be focused on, the scheduling system further analyzes the virtual machines running on the source computing node to identify which virtual machines consume the most of the specified resource type (such as CPU).
[0056] Optionally, in selecting virtual machines, in addition to considering resource consumption, the scheduling system also needs to evaluate the business importance of the virtual machine, the migration cost (including migration time, data transmission volume, etc.), and the accommodation capacity of the target computing node, to ensure the comprehensiveness and feasibility of the migration decision. The scheduling system can use certain algorithms or strategies, such as prioritizing the migration of virtual machines with the highest resource consumption and the least business impact, or selecting based on the comprehensive score of resource usage and business demand.
[0057] Suppose in the above scenario, the CPU usage of the source computing node A reaches 85%, and the resource balance benchmark value is set to 65%. When the system analyzes the virtual machines on node A, it finds that virtual machine V1 has the highest CPU consumption, accounting for 30% of the total CPU usage of the node; while virtual machine V2 also has a high CPU consumption, but it is carrying a critical business application and should not be migrated easily. The scheduling system can prioritize virtual machine V1 as the migration object because it significantly consumes CPU resources and migrating it can effectively reduce the CPU usage of node A. After selecting virtual machine V1, the scheduling system can further evaluate its feasibility of migrating to a target computing node (such as node Z), including the current CPU usage of node Z, the migration cost of virtual machine V1, and the impact on business continuity. If the evaluation result shows that migrating virtual machine V1 to node Z is feasible, the scheduling system will automatically perform a hot migration operation to move V1 from node A to node Z, and monitor the resource usage of node A to confirm whether the migration effectively reduces its load.
[0058] It should be noted that the scheduling system can also continuously monitor the resource usage of node A, and if it finds that the CPU usage is still higher than the benchmark value, it can select more virtual machines for migration until node A reaches the resource balance standard.
[0059] Through this intelligent selection mechanism based on resource usage and benchmark values, the scheduling system can accurately identify and migrate virtual machines that cause resource overload of the source computing nodes, automatically adjust resource allocation through live migration without affecting business continuity and user experience, and achieve resource balancing of computing nodes within the host group. This method not only improves resource utilization efficiency, but also effectively avoids performance bottlenecks and potential failures caused by node overload, enhancing the stability and response capability of the cloud computing platform.
[0060] In an optional embodiment, after each execution of the live migration operation, the N computing nodes are reordered according to the latest resource usage of each computing node, and the determination operation of the source computing node and the target computing node, the selection operation of the virtual machines to be migrated, and the live migration operation are performed again according to the reordered results until the resource balancing benchmark value corresponding to the N computing nodes is updated to be less than the preset threshold.
[0061] Optionally, in cloud computing resource management, the resource usage state is not static and unchangeable, but dynamically changes over time, business traffic, system load, and other factors. Therefore, achieving resource balancing also requires a dynamic process, that is, after a round of resource adjustment (such as virtual machine live migration) is performed, the scheduling system needs to reevaluate the resource usage of all computing nodes to confirm whether the balancing goal is achieved and decide whether further adjustment is needed. This dynamic balancing concept runs through the entire resource balancing operation process until the resource usage of all nodes reaches or approaches the preset balancing benchmark value.
[0062] In an optional embodiment, Figure 2 is a flowchart of an optional virtual machine migration method according to an embodiment of the present application, as Figure 2 shown, comprising the following steps:
[0063] Step 1. Calculate the resource situation in the host group, including CPU usage, memory usage, node ratio, etc. Further calculate the resource usage average to determine the resource balancing benchmark line;
[0064] Step 2. Sort the computing nodes in the host group according to the resource usage, determine the source computing nodes to be migrated and the target computing nodes;
[0065] Step 3. Perform the migration operation of the target computing nodes in turn, first determine the migrated virtual machines in the nodes according to the CPU / memory / ratio difference between the nodes and the balancing benchmark, and perform the migration operation;
[0066] Step 4. After each migration, repeat steps 2 and 3 until all migrated computing nodes in the host group complete the migration operation;
[0067] Step 5. Check the resource status of the host group, confirm whether it has reached below the balance baseline, if "no", repeat steps 2, 3, 4; if "yes", the resource balancing operation is completed.
[0068] For example, after the source computing node and the target computing node are initially determined, and the hot migration of the virtual machine is performed, the scheduling system does not immediately stop, but continues to collect the latest resource usage data of all computing nodes. Here, "latest" refers to the time point after the hot migration operation, because the migration operation itself can also cause temporary fluctuations in resource usage.
[0069] The scheduling system reorders the N computing nodes according to the new resource usage data. The ordering principle is still the high-low order of resource usage, in order to identify new resource overload nodes and resource loose nodes. After reordering, the scheduling system will again check the relationship between the resource usage of each computing node and the resource balancing baseline value. If the resource usage of a computing node is still significantly higher than the baseline value, the node can be marked again as a new source computing node; conversely, the computing node with resource usage lower than the baseline value can become a new target computing node. This process can be repeated multiple times until the resource usage of all computing nodes is fully adjusted to reach or approach the balancing baseline value.
[0070] In addition, in each cycle, the scheduling system needs to reanalyze the resource usage of the source computing node and identify which virtual machines are most suitable for migration. The selection criteria can include resource consumption, business importance, migration cost, and target node acceptance capacity of the virtual machine, and other factors.
[0071] Optionally, the resource balancing baseline value is not fixed, and as the overall resource usage of the system changes, the baseline value will also be adjusted accordingly. For example, when the scheduling system judges that the resource usage of all computing nodes is very close to the balanced state, it can further narrow the range of the resource balancing baseline value to pursue higher resource utilization accuracy and stability.
[0072] Optionally, if after a round of cycle, the resource usage of all computing nodes falls within the preset threshold range of the updated resource balancing baseline value, the system will consider that the resource balancing operation is completed and stop the cycle.
[0073] Through this cyclic iteration resource balancing method, the scheduling system can continuously fine-tune the resource allocation state of the computing nodes, even in the face of complex resource usage dynamics, it can ensure efficient resource utilization and load balancing between nodes, and ultimately achieve the purpose of optimizing the overall cloud computing platform performance and user experience.
[0074] In an optional embodiment, the resource usage of each computing node includes at least one of the following performance indicators:
[0075] a first performance indicator for characterizing the CPU usage of each computing node;
[0076] a second performance indicator for characterizing the memory usage of each computing node;
[0077] a third performance indicator for characterizing the node ratio of each computing node, wherein the node ratio is used to reflect the usage degree of each computing node relative to the maximum resource capacity of the computing node.
[0078] Optionally, the CPU usage refers to the usage of central processing units on the computing node, usually expressed in percentage. The higher the CPU usage at any given time, the more stressed the computing resources are. In a resource-balancing scenario, monitoring and analyzing CPU usage helps the system identify which nodes are performing a large number of computing tasks, which may lead to overload, and thus requires adjusting resource allocation, such as migrating virtual machines to other nodes with lower CPU usage, to achieve balanced use of CPU resources.
[0079] Optionally, the memory usage refers to the consumption degree of available memory resources on the computing node, also expressed in percentage. Memory is a key resource for executing programs and storing data, and high memory usage may mean that virtual machines or applications on the node are processing a large amount of data, or there are problems such as memory leaks. By monitoring memory usage, the system can identify nodes with tight memory resources and migrate virtual machines with high memory consumption to nodes with more abundant memory resources, achieving balanced use of memory resources.
[0080] Optionally, the node ratio is a comprehensive indicator that reflects the usage degree of the computing node relative to its maximum resource capacity. It not only considers the usage of a single CPU or memory, but also considers the usage of all resources to reflect the overall load state of the node. The calculation of the node ratio can be a weighted average of all key resource usage rates, with weights allocated according to the type and importance of the resources, for example, if CPU and memory have a greater impact on node performance, their weights in the calculation of the node ratio may be higher.
[0081] In resource-balancing operations, analysis of the node ratio helps the system assess the resource state of the node from a more comprehensive perspective, ensuring that virtual machine migration is not limited to optimizing CPU or memory, but considers balancing all key resources, avoiding incomplete resource balancing or secondary load problems caused by neglecting the usage of certain resources.
[0082] By considering multiple performance indicators such as CPU usage, memory usage, and node ratio, the scheduling system can more comprehensively and accurately assess the resource status of the computing nodes, thereby formulating more reasonable and effective resource balancing strategies to ensure the stable operation of the cloud computing platform and efficient use of resources.
[0083] In an optional embodiment, in the process of determining the resource balancing reference value corresponding to the N computing nodes according to the average value of the resource usage of the N computing nodes, the scheduling system can monitor and record multiple resource usages including CPU usage, memory usage, and node ratio in real time from the N computing nodes within a predetermined time interval. Then, the scheduling system can standardize the collected resource usages of different types; wherein the standardization processing is used to convert all data into data under the same measurement scale. The scheduling system can also obtain the weight of each resource usage, wherein the weight of each resource usage reflects the relative importance of each resource in determining the resource balancing reference value. Finally, the scheduling system calculates the weighted average value of various resource usages according to the standardized resource usages of the N computing nodes and the weight of each resource usage, and obtains the resource balancing reference value.
[0084] Optionally, the scheduling system can automatically capture multiple resource usage indicators from the N computing nodes in real time within a predetermined time interval (such as every minute, every 5 minutes), including CPU usage, memory usage, and node ratio reflecting the overall resource usage of the node. These data reflect the resource load state of the computing node at that moment. Of course, the monitoring is not limited to the above three performance indicators, but can be extended to more dimensions such as disk I / O usage, network bandwidth usage, etc. to cover all possible resource types that may affect node performance.
[0085] Optionally, since the units and magnitudes of different types of resource usage may be different (such as CPU usage may be a percentage, while disk I / O may be bytes / second), in order to be able to compare and analyze these data on a unified scale, the scheduling system will standardize the collected raw data. Optional standardization methods include linear scaling, i.e. scaling all data to the same interval (such as [0, 1]); another is Z-score-based standardization, which normalizes data by calculating the deviation of each data point from the average value and the standard deviation. Regardless of which method is used, the ultimate goal of standardization is to enable various resource usage data to be analyzed under the same comparison standard.
[0086] In addition, the scheduling system can assign a weight to each resource usage according to the characteristics of the cloud computing platform, business requirements, and historical data analysis. The greater the weight, the more significant the impact of the resource on node performance, and the more important its role in the calculation of the resource balance benchmark. The weight of resource usage is not fixed and can be dynamically adjusted as the business scenario and resource requirements change. For example, in a high-concurrency online transaction processing scenario, the weights of CPU and network bandwidth may increase, while in a cloud computing storage and retrieval task, the weights of memory and disk I / O may be higher.
[0087] The scheduling system multiplies the standardized resource usage data by its corresponding weight and then averages all the weighted data for all computing nodes to obtain the resource balance benchmark. This benchmark reflects the average resource usage state that the system expects each computing node to maintain under the current business environment. The resource balance benchmark calculated in this way not only takes into account the actual values of resource usage, but also fully reflects the importance and impact of different resource types in the cloud computing environment, ensuring the scientificity and effectiveness of the resource balance strategy.
[0088] Suppose there are N = 8 computing nodes in a cloud computing platform, and the scheduling system collects resource usage data every 5 minutes, including CPU usage, memory usage, and node ratio. The system assigns a weight of 0.5 to CPU usage, 0.3 to memory usage, and 0.2 to node ratio based on business requirements. In a certain data collection period, the scheduling system observes the standardized CPU usage of [0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1], the standardized memory usage of [0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1, 0.0], and the standardized node ratio of [0.6, 0.5, 0.4, 0.3, 0.2, 0.1, 0.0, 0.0].
[0089] The scheduling system calculates the weighted average of these standardized data according to the weights, obtaining a weighted average CPU usage of (0.8*0.5+0.7*0.5+...+0.1*0.5) / 8 = 0.45, a weighted average memory usage of (0.7*0.3+0.6*0.3+...+0.0*0.3) / 8 = 0.21, and a weighted average node ratio of (0.6*0.2+0.5*0.2+...+0.0*0.2) / 8 = 0.1.
[0090] Resource balancing benchmark value determination: Based on the above three weighted average values, the final resource balancing benchmark value calculation result is obtained. For example, if the above weight distribution is adopted, the resource balancing benchmark value can be a simple average or a weighted average of the three weighted average values, depending on the specific design strategy. In this example, if the calculation is simplified, the resource balancing benchmark value can be close to (0.45+0.21+0.1) / 3≈0.25, or more accurate considering the weight distribution calculation, the resource balancing benchmark value will be calculated by the scheduling system according to the actual weighted average value.
[0091] Through a series of operations such as real-time monitoring, data standardization, weight distribution and weighted average value calculation, the scheduling system can dynamically define the resource balancing benchmark value, ensure that the resource allocation strategy always meets the needs of the current business scenario, effectively avoid resource waste and system overload, and improve the operation efficiency of the cloud computing platform and user experience.
[0092] In an optional embodiment, in the process of selecting virtual machines to be migrated from the source computing node according to the resource usage of the source computing node and the resource balancing benchmark value, the scheduling system can detect the difference between the resource usage of the source computing node and the resource balancing benchmark value, and then the scheduling system can select one or more virtual machines from the source computing node as the virtual machines to be migrated according to the difference, the resource usage data of the source computing node and the resource consumption data of each virtual machine on the source computing node, wherein the selection process of the virtual machines to be migrated follows one or a combination of the minimum resource impact principle, the maximum resource release principle or the business priority principle.
[0093] Optionally, the minimum resource impact principle is used to select virtual machines that have the least impact on the resource changes of the source computing node and the target computing node after migration; the maximum resource release principle is used to select virtual machines with the largest resource occupation or closest to the part of the source computing node resource load exceeding the resource balancing benchmark value for migration; and the business priority principle is used to preferentially migrate virtual machines of low priority services.
[0094] Optionally, in a cloud computing environment, an important part of achieving resource balancing is to select appropriate virtual machines for migration on overloaded source computing nodes. This selection process needs to follow several strategies to ensure that resource balancing is met while considering business continuity and resource management efficiency. These strategies include the minimum resource impact principle, the maximum resource release principle, and the business priority principle, each of which focuses on different optimization goals but serves the common goal of resource balancing and business stability.
[0095] Optionally, the minimum resource impact principle focuses on minimizing the impact of virtual machine migration on existing system resources, aiming to ensure that the migration operation does not trigger new resource pressure or instability. When implementing this principle, the scheduling system evaluates the potential resource changes that each virtual machine to be migrated can cause on the source and target compute nodes. This involves predicting the changes in resource usage of both nodes after the migration, ensuring that the migration operation does not cause a sharp increase in resource usage on either node. When using this principle for resource scheduling, the scheduling system tends to select virtual machines that have the least impact on the resource usage of both nodes after migration. This often means selecting virtual machines that have moderate resource occupancy and less business impact.
[0096] Optionally, compared to the minimum resource impact principle, the maximum resource release principle has a more direct goal - to reduce the resource usage of the source compute node as quickly as possible through the migration operation. The implementation includes: the scheduling system analyzes the resource occupancy of each virtual machine on the source compute node, especially for critical resources such as CPU, memory, disk I / O, etc. The scheduling system will preferentially select virtual machines with the largest resource occupancy or those closest to the part of the source node resource load exceeding the benchmark value for migration. This can quickly reduce the resource usage of the node and make it closer to the resource balance benchmark value.
[0097] Optionally, the business priority principle focuses on the impact of virtual machine migration on business, ensuring that critical business is not disturbed. The implementation process includes: the scheduling system assigns a business priority value to each virtual machine based on factors such as business type, urgency, customer level, etc. When using this principle for resource scheduling, the scheduling system tends to migrate virtual machines of low-priority businesses first to minimize the impact on high-priority, critical businesses.
[0098] In addition, in most cases, the scheduling system does not rely solely on a single principle, but combines them to achieve the comprehensive optimal resource balance goal. For example, when the resource usage of the source compute node is far beyond the benchmark value, the system may first select the virtual machine with the largest resource occupancy according to the maximum resource release principle, and then consider the minimum resource impact principle to ensure the stability of subsequent migration operations. At the same time, throughout the process, the business priority principle is used to prioritize virtual machines of low-impact businesses to maintain business continuity and user experience.
[0099] Suppose in a cloud computing platform, the CPU usage of source compute node A reaches 85%, which is higher than the resource balance benchmark value of 70%. When selecting virtual machines for migration, the system first follows the maximum resource release principle and identifies that virtual machine VM1 is the highest CPU consumer, occupying about 30% of the resources. Then, the system checks the business priority of VM1 and finds that it belongs to a non-critical business, further verifying that it is suitable as a migration object.
[0100] After the migration of VM1 is completed, the system evaluates the impact of the remaining virtual machines on the source computing node A and the target computing node B after migration according to the principle of minimum resource impact, and finally selects the virtual machine VM2 with moderate resource occupation and less business impact for migration to prevent the node B from being overloaded.
[0101] During the whole process, if a virtual machine (such as VM3) that also meets the maximum resource release principle but has a high business priority is encountered, the system will temporarily suspend it as a migration object and instead look for the next best migration target.
[0102] Through the combination of the above strategies, the scheduling system can flexibly and efficiently adjust resource allocation in a complex and variable cloud computing environment, achieving resource balancing while maintaining business stability and user experience, and exhibiting a high level of intelligence and automation.
[0103] In an optional embodiment, the hot migration operation further includes performing a rollback operation when a migration failure occurs, wherein the rollback operation is used to restore the virtual machine that fails to migrate to a pre-migration state.
[0104] Optionally, when the hot migration fails, the rollback operation can quickly restore the state of the virtual machine to the pre-migration state, avoiding business interruption or data loss caused by migration failure. The rollback mechanism helps to prevent computing node resource abnormalities caused by migration failure, maintain the stable state of the system, and prevent the spread of faults. By ensuring the atomicity of the migration operation (either all succeed or all fail and rollback), the rollback mechanism enhances the overall reliability of the cloud computing platform and improves user experience.
[0105] Optionally, before the hot migration begins, the scheduling system will save the state snapshot of the virtual machine on its original computing node, including memory data, virtual machine configuration parameters, network settings, etc., to ensure that these information can be restored in the event of migration failure. During the hot migration process, the scheduling system will monitor the migration progress and state in real time, and trigger the rollback operation immediately upon detecting a migration failure. The rollback operation will be based on the previously saved state snapshot to quickly restore the virtual machine to the pre-migration physical computing node and restore all its runtime states, ensuring that the virtual machine's services are not affected. After the rollback operation is completed, the scheduling system will update the resource usage of each computing node in real time, and adjust other migration plans or resource allocation strategies as necessary to cope with the resource changes that may be caused by rollback.
[0106] According to another aspect of the embodiments of the present application, a computing resource scheduling device is also provided, wherein, Figure 3 is a schematic diagram of an optional computing resource scheduling device according to an embodiment of the present application, as Figure 3As shown, the apparatus comprises: a first determining unit 301 configured to determine a resource balance reference value corresponding to the N computing nodes according to an average value of resource usage rates of the N computing nodes after the N computing nodes complete upgrading, wherein N is an integer greater than 1; a second determining unit 302 configured to sort the N computing nodes according to the resource usage rates of each computing node, and determine a source computing node and a target computing node from the N computing nodes according to a sorting result, wherein the resource usage rate of the source computing node is higher than that of the target computing node; a selecting unit 303 configured to select a virtual machine to be migrated on the source computing node according to the resource usage rate of the source computing node and the resource balance reference value; and a migrating unit 304 configured to perform a hot migration operation to migrate the virtual machine to be migrated from the source computing node to the target computing node.
[0107] Optionally, the computing resource scheduling apparatus further comprises a processing unit configured to re-sort the N computing nodes according to the latest resource usage rates of each computing node after each time the hot migration operation is performed, and perform again the determining operation of the source computing node and the target computing node, the selecting operation of the virtual machine to be migrated, and the hot migration operation according to a re-sorting result until the resource balance reference value corresponding to the N computing nodes is updated to be less than a preset threshold.
[0108] Optionally, the resource usage rate of each computing node comprises at least one of the following performance indicators:
[0109] The first performance indicator is used to represent a CPU usage rate of each computing node.
[0110] The second performance indicator is used to represent a memory usage rate of each computing node.
[0111] The third performance indicator is used to represent a node multiple rate of each computing node, wherein the node multiple rate is used to reflect a usage degree of each computing node relative to a maximum resource capacity of the computing node.
[0112] Optionally, the first determining unit 301 comprises: a monitoring sub-unit configured to monitor and record in real time a plurality of resource usage rates including at least the CPU usage rate, the memory usage rate, and the node multiple rate from the N computing nodes within a predetermined time interval; a first processing sub-unit configured to perform standardization processing on the collected different types of resource usage rates; wherein the standardization processing is used to convert all data into data under the same measurement scale; an acquisition sub-unit configured to acquire a weight of each resource usage rate, wherein the weight of each resource usage rate is used to reflect a relative importance of each resource in determining the resource balance reference value; and a calculation sub-unit configured to calculate a weighted average value of various resource usage rates according to the standardized resource usage rates of the N computing nodes and the weight of each resource usage rate to obtain the resource balance reference value.
[0113] Optionally, the selecting unit 303 comprises: a detecting subunit, configured to detect a difference between the resource usage of the source computing node and the resource balance reference value; and a selecting subunit, configured to select one or more virtual machines from the source computing node as the virtual machines to be migrated according to the difference, the resource usage data of the source computing node, and the resource consumption data of each virtual machine on the source computing node, wherein the selection of the virtual machines to be migrated follows one or a combination of a minimum resource impact principle, a maximum resource release principle, or a service priority principle.
[0114] Optionally, the minimum resource impact principle is used to select the virtual machine that has the minimum impact on the source computing node and the target computing node after migration; the maximum resource release principle is used to select the virtual machine that has the maximum resource occupation or is closest to the part of the source computing node resource load exceeding the resource balance reference value for migration; and the service priority principle is used to preferentially migrate the virtual machine of a low-priority service.
[0115] Optionally, the hot migration operation further comprises performing a rollback operation when the migration fails, wherein the rollback operation is used to restore the virtual machine that fails in migration to a state before migration.
[0116] According to another aspect of the embodiments of the present application, a computer readable storage medium is further provided, wherein the computer readable storage medium stores a computer program, and the computer program, when executed, causes a device where the computer readable storage medium is located to perform the above-mentioned computing resource scheduling method.
[0117] According to another aspect of the embodiments of the present application, an electronic device is further provided, wherein the electronic device comprises one or more processors and a memory, the memory is configured to store one or more programs, and the one or more programs, when executed by the one or more processors, cause the one or more processors to perform the above-mentioned computing resource scheduling method.
[0118] The serial numbers of the above-mentioned embodiments of the present application are only for description, and do not represent the advantages or disadvantages of the embodiments.
[0119] In the above-mentioned embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0120] In several embodiments provided in the present application, it should be understood that the disclosed technology can be implemented by other ways. Among them, the above-described device embodiments are only schematic, for example, the division of the units can be a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, units or modules, and can be electrical or other forms.
[0121] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0122] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0123] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part of the prior art that contributes to the technical solutions or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0124] The above is only the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.
Claims
1. A method for scheduling computing resources, characterized in that, include: After N computing nodes have been upgraded, the resource balancing benchmark value corresponding to the N computing nodes is determined based on the average resource utilization rate of the N computing nodes, where N is an integer greater than 1. The N computing nodes are sorted according to the resource utilization rate of each computing node, and the source computing node and the target computing node are determined from the N computing nodes according to the sorting result, wherein the resource utilization rate of the source computing node is higher than that of the target computing node. Based on the resource utilization rate of the source compute node and the resource balancing benchmark value, select the virtual machines that need to be migrated on the source compute node; Perform a hot migration operation to migrate the virtual machine to be migrated from the source compute node to the target compute node.
2. The method according to claim 1, characterized in that, The method further includes: After each hot migration operation, the N compute nodes are reordered according to the latest resource utilization of each compute node, and the source compute node and target compute node determination operation, the virtual machine to be migrated selection operation, and the hot migration operation are executed again according to the reordering result, until the resource balancing benchmark value corresponding to the N compute nodes is updated to less than a preset threshold.
3. The method according to claim 1, characterized in that, The resource utilization of each computing node includes at least one of the following performance metrics: The first performance metric is used to characterize the CPU utilization of each computing node; The second performance metric is used to characterize the memory utilization rate of each computing node; The third performance metric is used to characterize the node multiplier of each computing node, wherein the node multiplier is used to reflect the degree of utilization of each computing node relative to the maximum resource capacity of that computing node.
4. The method according to claim 1, characterized in that, Based on the average resource utilization of N computing nodes, determine the resource balancing benchmark value corresponding to the N computing nodes, including: Within a predetermined time interval, the utilization rates of various resources, including at least CPU utilization, memory utilization, and node multiplier, are monitored and recorded in real time from the N computing nodes. The collected utilization rates of different types of resources are standardized; wherein, the standardization process is used to convert all data into data under the same measurement scale; Obtain the weight of each resource utilization rate, wherein the weight of each resource utilization rate is used to reflect the relative importance of each resource in the process of determining the resource balance benchmark value; Based on the standardized resource utilization rates of the N computing nodes and the weight of each resource utilization rate, the weighted average of the various resource utilization rates is calculated to obtain the resource balance benchmark value.
5. The method according to claim 1, characterized in that, Based on the resource utilization rate of the source compute node and the resource balancing benchmark value, select the virtual machines to be migrated on the source compute node, including: The difference between the resource utilization rate of the source computing node and the resource balance benchmark value is detected; Based on the difference, the resource usage data of the source compute node, and the resource consumption data of each virtual machine on the source compute node, one or more virtual machines are selected from the source compute node as virtual machines to be migrated. The selection process of the virtual machines to be migrated follows one or a combination of the principles of minimum resource impact, maximum resource release, or business priority.
6. The method according to claim 5, characterized in that, The minimum resource impact principle is used to select the virtual machine that has the least impact on the resource changes of the source computing node and the target computing node after migration; the maximum resource release principle is used to select the virtual machine with the largest resource consumption or the part of the source computing node's resource load that exceeds the resource balancing benchmark value for migration; the service priority principle is used to prioritize the migration of virtual machines with low-priority services.
7. The method according to any one of claims 1 to 6, characterized in that, The hot migration operation also includes performing a rollback operation when a migration failure occurs, wherein the rollback operation is used to restore the virtual machine that failed to migrate to its state before migration.
8. A computing resource scheduling device, characterized in that, include: The first determining unit is configured to, after the N computing nodes have completed their upgrades, determine the resource balancing benchmark value corresponding to the N computing nodes based on the average resource utilization rate of the N computing nodes, wherein... N is an integer greater than 1; The second determining unit is used to sort the N computing nodes according to the resource utilization rate of each computing node, and determine the source computing node and the target computing node from the N computing nodes according to the sorting result, wherein the resource utilization rate of the source computing node is higher than the resource utilization rate of the target computing node. The selection unit is used to select the virtual machines to be migrated on the source computing node based on the resource utilization rate of the source computing node and the resource balancing benchmark value. The migration unit is used to perform a hot migration operation to migrate the virtual machine to be migrated from the source compute node to the target compute node.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed, the device in which the computer-readable storage medium is located performs the computing resource scheduling method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, It includes one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to perform the computing resource scheduling method according to any one of claims 1 to 7.