Resource scheduling methods, resource scheduling devices, resource scheduling equipment and storage media

By training a predictive model to predict the temperature of core components and allocate tasks, the problem of uneven temperature in the server cluster was solved, the temperature of core components was balanced, and the reliability and stability of business applications were improved.

CN115344394BActive Publication Date: 2026-04-03SHAANXI LANGCHAO YINGXIN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-31
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, uneven temperatures of core components in different servers within a server cluster can cause some core components to overheat, increasing the risk of damage and affecting the reliable operation of business applications.

Method used

By training a prediction model and utilizing data on the type, usage rate, and physical location of core components, the temperature impact of tasks to be assigned can be predicted, thereby achieving resource scheduling with balanced core component temperatures and preventing task assignments from exceeding safe temperature thresholds.

Benefits of technology

This reduces the probability of core components overheating and damage, and improves the reliability and operational stability of business applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115344394B_ABST
    Figure CN115344394B_ABST
Patent Text Reader

Abstract

This application relates to the field of server technology, specifically disclosing a resource scheduling method, resource scheduling device, resource scheduling equipment, and storage medium. By pre-training a prediction model using the type of core components, the utilization rate of core components, and the physical location of the servers where the core components are located as input data, and the temperature of the core components as output data, a prediction model is obtained. During resource scheduling, for tasks to be assigned, the prediction model simulates the predicted first temperature of the core components of the candidate servers when assigning the tasks to the candidate servers. Under the premise that the predicted first temperature does not exceed the safe temperature threshold of the candidate servers, the tasks to be assigned are evenly distributed among the candidate servers. This achieves resource scheduling based on the goal of balancing the temperature of core components, reducing the probability of damage caused by excessively high core component temperatures and improving the reliability of business application operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of server technology, and in particular to a resource scheduling method, resource scheduling device, resource scheduling equipment, and storage medium. Background Technology

[0002] The goal of distributed resource scheduling is to achieve a more consistent utilization of physical resources in a server cluster. The physical resources involved include the central processing unit (CPU), graphics processing unit (GPU), and memory.

[0003] Research has revealed a significant correlation between the operating status of core server components (such as CPU, GPU, and memory) and their temperature. Excessive temperature makes these core components more susceptible to damage, leading to tasks that rely on them failing to execute, virtual machine crashes, and service interruptions in hosted business applications. In practical applications, due to differences in the physical location and resource types of servers, the same resource utilization has varying impacts on different server hardware resources. A key manifestation of this is that even if resource utilization is balanced across servers, the temperature rise of core components on different servers can differ, posing a risk that overheating of certain components could damage these core components, compromising the reliable operation of business applications. Summary of the Invention

[0004] The purpose of this application is to provide a resource scheduling method, resource scheduling device, resource scheduling equipment, and storage medium to achieve resource scheduling based on the goal of core component temperature balancing, reduce the probability of damage caused by excessive temperature of core components, and improve the reliability of business application operation.

[0005] To address the aforementioned technical problems, this application provides a resource scheduling method, comprising:

[0006] The prediction model is trained by taking the type of core component, the usage rate of the core component, and the physical location of the server where the core component is located as input data and the temperature of the core component as output data.

[0007] For the task to be assigned, the prediction model simulates the first temperature prediction value of the core component of the candidate server when the task to be assigned is assigned to the candidate server.

[0008] Provided that the first temperature prediction value does not exceed the safe temperature threshold of the candidate server, the tasks to be assigned will be evenly distributed among the candidate servers.

[0009] The core components include computing components, storage components, and network communication components.

[0010] Optionally, the type of task to be assigned includes at least one of business computing tasks, virtual machine scheduling tasks, or container scheduling tasks.

[0011] Optional, also includes:

[0012] When the temperature of the core components of the source server exceeds the safe temperature threshold of the source server, the task to be assigned is determined to be migrated on the source server.

[0013] Optionally, when the temperature of the core components of the source server is detected to exceed the safe temperature threshold of the source server, the task to be assigned is determined to be migrated on the source server, specifically as follows:

[0014] When the device running the virtualization management system detects that the temperature of the core components of the source server exceeds the safe temperature threshold of the source server, it determines the virtual machine to be migrated on the source server.

[0015] Optionally, when the device running the virtualization management system detects that the temperature of the core components of the source server exceeds the safe temperature threshold of the source server, the virtual machine to be migrated is determined on the source server, specifically as follows:

[0016] The prediction model is used to predict the virtual machines running on the source server one by one, simulating the second temperature prediction value of the core components of the source server after different combinations of virtual machines are migrated from the source server, until the combination of virtual machines that reduces the second temperature prediction value to below the safe temperature threshold of the source server and minimizes the migration cost is determined, and this combination is determined as the virtual machine to be migrated.

[0017] Optionally, the step of evenly distributing the tasks to be assigned to the candidate servers, provided that the first temperature prediction value does not exceed the safe temperature threshold of the candidate servers, specifically involves:

[0018] Using the prediction model, predictions are made for servers with low real-time temperature detection values ​​one by one. The model simulates the combination of candidate servers where the predicted first temperature value of each candidate server does not exceed the safe temperature threshold of each candidate server after the virtual machine to be migrated is migrated to each candidate server and the migration cost is minimized. The virtual machine to be migrated is then migrated to the combination of candidate servers.

[0019] Optional, also includes:

[0020] When the temperature difference of the same type of core component on different servers exceeds the temperature difference threshold, the task to be assigned is determined on the server with the higher real-time temperature detection value.

[0021] Optionally, the utilization rate of core components can be obtained in-band through the host operating system running on the server.

[0022] Optionally, the physical location of the server containing the core component is specifically the coordinates of the center of the server containing the core component in a three-dimensional coordinate system established with the location of the cooling equipment in the data center as the origin.

[0023] Optionally, the temperature of the core components can be obtained out-of-band via the baseboard management controller of the server.

[0024] Optionally, a preset number of servers with the lowest real-time temperature detection value can be selected as the candidate servers, and the preset number of servers is consistent with the number of tasks to be assigned.

[0025] Optional, also includes:

[0026] The prediction model is updated based on the measured real-time utilization rate of the core components, the real-time temperature detection value of the core components, the type of the corresponding core components, and the physical location of the server where the core components are located.

[0027] To address the aforementioned technical problems, this application also provides a resource scheduling device, comprising:

[0028] The training unit is used to train a prediction model in advance, taking the type of core component, the usage rate of the core component, and the physical location of the server where the core component is located as input data and the temperature of the core component as output data.

[0029] The allocation prediction unit is used to simulate the first temperature prediction value of the core component of the candidate server when the task to be allocated is assigned to the candidate server using the prediction model.

[0030] The allocation unit is used to evenly allocate the task to be allocated to the candidate server, provided that the first temperature prediction value does not exceed the safe temperature threshold of the candidate server.

[0031] The core components include computing components, storage components, and network communication components.

[0032] To address the aforementioned technical problems, this application also provides a resource scheduling device, comprising:

[0033] Memory, used to store computer programs;

[0034] A processor for executing the computer program, which, when executed by the processor, implements the steps of the resource scheduling method as described in any of the preceding descriptions.

[0035] To address the aforementioned technical problems, this application also provides a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the resource scheduling method described in any of the preceding claims.

[0036] The resource scheduling method provided in this application trains a prediction model by using the type of core component, the utilization rate of the core component, and the physical location of the server where the core component is located as input data, and the temperature of the core component as output data. When performing resource scheduling, for the task to be assigned, the prediction model simulates the first predicted temperature value of the core component of the candidate server when the task to be assigned is assigned to the candidate server. Under the premise that the first predicted temperature value does not exceed the safe temperature threshold of the candidate server, the task to be assigned is evenly distributed among the candidate servers. This achieves resource scheduling based on the goal of balancing the core component temperature, reduces the probability of damage caused by excessive temperature of the core component, and improves the reliability of business application operation.

[0037] This application also provides a resource scheduling device, resource scheduling equipment, and storage medium, which have the above-mentioned beneficial effects, and will not be described in detail here. Attached Figure Description

[0038] To more clearly illustrate the technical solutions of the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0039] Figure 1 A flowchart illustrating a resource scheduling method provided in an embodiment of this application;

[0040] Figure 2 This is a schematic diagram of the structure of a resource scheduling device provided in an embodiment of this application;

[0041] Figure 3 This is a schematic diagram of the structure of a resource scheduling device provided in an embodiment of this application. Detailed Implementation

[0042] The core of this application is to provide a resource scheduling method, resource scheduling device, resource scheduling equipment, and storage medium to achieve resource scheduling based on the goal of core component temperature balancing, thereby reducing the probability of damage caused by excessive temperature of core components and improving the reliability of business application operation.

[0043] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0044] Example 1

[0045] Figure 1 This is a flowchart of a resource scheduling method provided in an embodiment of this application.

[0046] like Figure 1 As shown, the resource scheduling method provided in this application includes:

[0047] S101: The prediction model is trained in advance using the type of core component, the utilization rate of the core component, and the physical location of the server where the core component is located as input data, and the temperature of the core component as output data.

[0048] S102: The first predicted temperature of the core components of the candidate server when the candidate task is assigned to the candidate server by simulating the assignment of the candidate task through a prediction model.

[0049] S103: Provided that the first temperature prediction value does not exceed the safe temperature threshold of the candidate server, the tasks to be assigned will be evenly distributed among the candidate servers.

[0050] The core components include computing components, storage components, and network communication components.

[0051] In practical implementation, the resource scheduling method provided in this application can be implemented on a server cluster basis, based on one or more resource scheduling servers within the server cluster. The core components to be monitored in this application embodiment can be directed to some or all of the servers in the server cluster.

[0052] The core components in this application embodiment include computing components, storage components, and network communication components. The computing components include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), and other heterogeneous computing devices. Storage components may include memory modules, non-volatile storage media, etc. Network communication components include network interface cards (NICs). Furthermore, different core components may have different models. For ease of management, in this application embodiment, the type of the core component can be used as training data, or further, the model number of the core component can be used as training data.

[0053] Practical applications show a strong correlation between the temperature of core components and their failure rate; excessively high temperatures lead to an increased failure rate. Similarly, the utilization rate of core components is strongly correlated with their failure rate; excessive utilization also increases the failure rate. Furthermore, the ambient temperature of the server environment is strongly correlated with the temperature of its components; higher external temperatures result in higher core component temperatures. Finally, the server's ambient temperature is strongly correlated with its location within the data center, as air conditioning (such as air-cooled systems) provides uneven cooling to different areas of the data center.

[0054] For S101, a neural network model is used as the core decision engine for distributed resource scheduling. Data such as the type of the corresponding core components, the utilization rate of the core components, the physical location of the server where the core components are located, and the temperature of the core components are collected in advance. The neural network model is trained with the type of the core components, the utilization rate of the core components, and the physical location of the server where the core components are located as input data and the temperature of the core components as output data, to obtain the prediction model.

[0055] The utilization rate of core components can be obtained in-band via the host operating system running on the server. Specifically, the utilization rate of each core component can be sampled at a first preset interval (e.g., 10 minutes) for a first preset duration (e.g., 15 days) to obtain sufficient core component utilization data. The resources consumed by tasks, running virtual machines, or containers are usually fixed, so the utilization rate of core components can be obtained by directly reading the monitoring parameters of the core components, or by detecting the type and number of tasks being executed by the core components.

[0056] Data center cooling equipment is a set of facilities established to ensure the temperature and humidity environment required for the operation of IT equipment. It mainly includes the air conditioning equipment used in the computer room and the equipment providing the cooling source. Computer room air conditioning is a crucial piece of data center cooling equipment. Based on different cooling methods, computer room air conditioning is divided into room-level, row-level, and rack-level. Based on different cooling sources, computer room air conditioning can be divided into air-cooled, water-cooled, and chilled water types. For data centers, the temperature of the core server components is strongly correlated with the relative position of the air conditioning (especially air-cooled air conditioning). Therefore, the physical location of the server containing the core components can be considered as the coordinates of the center of the server within a three-dimensional coordinate system established with the location of the cooling equipment in the computer room as the origin. Specifically, a three-dimensional coordinate system is established in the computer room along its length, width, and height, with the location of the cooling equipment as the origin, and the geometric center of each server is used as the coordinate point.

[0057] The Baseboard Management Controller (BMC) is a small, independent operating system separate from the server system. It's a chip integrated on the motherboard, though some products connect it via PCIe or other means. Externally, it appears as a standard RJ45 network port and has its own dedicated firmware system. It's a fundamental core functional subsystem of the server, responsible for core functions such as hardware status management, operating system management, health status management, and power consumption management. In this embodiment, temperature data of the server's internal core components can be obtained based on the BMC. Therefore, the temperature of the core components can be obtained out-of-band through the server's BMC.

[0058] For S102 and S103, when a task to be assigned is generated, to achieve resource scheduling based on core component temperature balancing, one or more target servers are selected from the candidate servers to execute the assigned task. The criterion for selecting a target server is that after assigning the task to the target server, the predicted first temperature value of the target server's core component will not exceed the target server's safe temperature threshold, as predicted by the prediction model. This safe temperature threshold corresponds to the server; furthermore, the safe temperature threshold can also correspond one-to-one with the type of core component.

[0059] The types of tasks to be assigned include, but are not limited to, at least one of the following: business computing tasks, virtual machine scheduling tasks, or container scheduling tasks. Virtual machine scheduling tasks refer to creating virtual machines on the target server or migrating virtual machines from another server to the target server. Container scheduling tasks refer to creating containers on the target server or migrating containers from another server to the target virtual machine. It should be noted that, regardless of whether it is a business computing task, virtual machine creation / migration, or container creation / migration, unless otherwise specified, they are all tasks that occupy their own processes equally and are subject to unified allocation and scheduling. Based on this, if the tasks to be assigned have no priority requirements, resource scheduling is performed according to the order in which they were created; if the tasks to be assigned have priority requirements, the processing order of the tasks is adjusted according to their priority.

[0060] Understandably, if there is only one task to be assigned, only one target server needs to be selected to execute it. If there are multiple tasks to be assigned, one or more target servers need to be selected to execute them, such as assigning all tasks to one target server, or assigning each task to a different target server, and so on. To quickly select target servers, candidate servers can be determined first. Specifically, a preset number of servers with the lowest real-time temperature readings can be used as candidate servers, with the preset number matching the number of tasks to be assigned. Based on the current number of tasks to be assigned and the preset number of servers with the lowest real-time temperature readings as candidate servers, a predictive model can be used to simulate assigning all tasks to the candidate server with the lowest real-time temperature reading. The first predicted temperature value of the candidate server is then determined, and it is checked whether it exceeds the safe temperature threshold of that candidate server. If it does not exceed the threshold, all tasks can be assigned to that candidate server for execution; if it exceeds the threshold, then the two candidate servers with the lowest real-time temperature readings are selected, and so on. Alternatively, based on the number of tasks to be assigned and the preset number of servers with the lowest real-time temperature detection values, the tasks to be assigned can be directly assigned to different servers. If the prediction model simulates that the first temperature prediction value of the core component of a server exceeds the safe temperature threshold of that server after adding tasks, the assigned server can be changed or a new server can be selected for assignment.

[0061] At the same time, the selected candidate server should have the core components that meet the execution requirements of the task to be assigned.

[0062] To ensure the accuracy of the prediction model, the resource scheduling method provided in this application embodiment may further include: updating the prediction model based on the measured real-time utilization rate of the core components, the real-time temperature detection value of the core components, the type of the corresponding core components, and the physical location of the server where the core components are located. Since insufficient training data may be used in step S101, different types of core components may be added to the cluster, servers may be added / replaced in the data center, or servers may be relocated, the prediction model may be unable to accurately predict the temperature of the core components. Therefore, during resource scheduling, the real-time utilization rate of each core component, the real-time temperature detection value of the core components, the type of the corresponding core components, and the physical location of the server where the core components are located can be monitored periodically to update the model parameters of the prediction model and improve prediction accuracy.

[0063] The resource scheduling method provided in this application pre-trains a prediction model using the type of core component, the utilization rate of the core component, and the physical location of the server where the core component is located as input data, and the temperature of the core component as output data. During resource scheduling, for the task to be assigned, the prediction model simulates the first predicted temperature value of the core component of the candidate server when the task to be assigned is assigned to the candidate server. Under the premise that the first predicted temperature value does not exceed the safe temperature threshold of the candidate server, the task to be assigned is evenly distributed among the candidate servers. This achieves resource scheduling based on the goal of balancing the core component temperature, reduces the probability of damage caused by excessive temperature of the core component, and improves the reliability of business application operation.

[0064] Example 2

[0065] As described in Embodiment 1 of this application, due to various reasons, the prediction model cannot guarantee 100% accuracy in predicting the temperature of core components. Inaccurate predictions can still lead to large temperature differences among core components in the cluster, and damage to some core components due to excessively high temperatures. Based on the above embodiments, the resource scheduling method provided in this application further includes: when the temperature of the core component of the source server exceeds the safe temperature threshold of the source server, determining the tasks to be migrated on the source server.

[0066] In practice, real-time temperature readings of core components can be acquired periodically. This acquisition can be done out-of-band by the server's baseboard management controller. Besides updating the prediction model, the real-time temperature readings are also used to determine which tasks need to be migrated. These tasks include, but are not limited to, business computing tasks, virtual machines, and containers.

[0067] When the temperature of the core components of the source server exceeds the safe temperature threshold of the source server, the task to be migrated is determined on the source server. Specifically, when the device running the virtualization management system detects that the temperature of the core components of the source server exceeds the safe temperature threshold of the source server, the virtual machine to be migrated is determined on the source server.

[0068] Alternatively, when the temperature of the core components of the source server is detected to exceed the safe temperature threshold of the source server, the tasks to be migrated are determined on the source server. Specifically, when the device running the virtualization management system detects that the temperature of the core components of the source server exceeds the safe temperature threshold of the source server, the containers to be migrated are determined on the source server.

[0069] Taking virtual machine migration as an example, when the device of the virtualization management system detects that the temperature of the core component of the source server exceeds the safe temperature threshold of the source server, the virtual machine to be migrated is determined on the source server. Specifically, the prediction model is used to predict the temperature of each virtual machine running on the source server, and different combinations of virtual machines are simulated to determine the second temperature prediction value of the core component of the source server after being migrated from the source server, until the combination of virtual machines that reduces the second temperature prediction value to below the safe temperature threshold of the source server and minimizes the migration cost is determined, and this combination is determined as the virtual machine to be migrated.

[0070] In practice, to reduce the temperature of overheated core components on the source server to within a safe temperature threshold, the decision is made regarding which virtual machines (VMs) to migrate from the source server. Each VM running on the source server is analyzed and predicted. After subtracting the CPU, memory, and other core component resource usage of that VM from the source server's specifications, a predictive model is used to predict the second temperature of each core component. If the second temperature prediction falls below the corresponding safe temperature threshold, only one VM needs to be migrated. If migrating only one VM is insufficient, two VMs are exhaustively selected for prediction analysis through permutations and combinations. If migrating only two VMs is insufficient, three VMs are exhaustively selected for prediction analysis through permutations and combinations, and so on, until a VM combination that reduces the second temperature prediction below the source server's safe temperature threshold is found. If multiple VM combinations satisfy the objective, the VM combination with the lowest migration cost is determined based on the number of VMs within each combination, the priority of the services executed by each VM, and the real-time requirements of the services executed by each VM. This combination is then selected as the VMs to be migrated.

[0071] Similarly, S103: Under the premise that the first temperature prediction value does not exceed the safe temperature threshold of the candidate server, the task to be assigned is evenly distributed to the candidate server. Specifically, it can be: using the prediction model, predicting the server with low real-time temperature detection value one by one, simulating the candidate server combination that minimizes the migration cost and ensures that the first temperature prediction value of each candidate server does not exceed the safe temperature threshold of each candidate server after the virtual machine to be migrated to each candidate server, so as to migrate the virtual machine to be migrated to the candidate server combination.

[0072] In practice, a server with a lower core component temperature can be initially selected as a candidate server. The resource usage of the virtual machine to be migrated is increased on this candidate server. A predictive model is used to obtain the first predicted temperature value of each core component on the candidate server. If the first predicted temperature value does not exceed the corresponding safe temperature threshold of the candidate server, then this candidate server can be used as the target server, and the virtual machine to be migrated can be migrated to the target server. If one candidate server cannot meet the goal of ensuring that the first predicted temperature value of each core component does not exceed the corresponding safe temperature threshold after increasing the resource usage of the virtual machine to be migrated, then two candidate servers are selected, and the servers to be migrated are evenly distributed among them for temperature prediction… and so on, with the maximum number of target servers matching the number of virtual machines to be migrated. If multiple candidate server combinations that meet the objectives are found, the combination with the lowest migration cost can be identified based on factors such as the core components of the candidate server requiring the assigned tasks, the priority of the tasks performed on the candidate server, and the real-time requirements of the tasks performed on the candidate server. This combination will then be used as the target server for the virtual machine to be migrated.

[0073] The same principle applies if the task to be assigned is a business computing task or a container. If considering migrating or assigning different types of tasks simultaneously, prioritize the different types of tasks based on business priority, migration efficiency, resource utilization, etc., and then select the task to be assigned. Additionally, based on the types of core components contained in the candidate servers with lower core component temperatures, determine the types of tasks to be migrated from the source servers with higher core component temperatures.

[0074] Example 3

[0075] When there are significant temperature differences between the same type of core components on different servers, it indicates an imbalance in resource scheduling within the cluster, specifically regarding the temperature of the core components. Therefore, based on the above embodiments, the resource scheduling method provided in this application further includes: when the temperature difference between the same type of core components on different servers exceeds a temperature difference threshold, determining the server with the higher real-time temperature detection value as the server to be assigned tasks.

[0076] In practice, even if the real-time temperature readings of all core components are within their corresponding safe temperature thresholds, the system determines whether there is an imbalance in resource scheduling based on core component temperatures by calculating the temperature differences between the same type of core components on different servers and checking if these differences exceed a pre-set threshold. Specifically, the real-time monitored temperatures of the same type of core components on different servers can be sorted within the cluster. The temperature difference between the core component with the highest and lowest real-time monitored temperatures is compared to a temperature difference threshold. If this difference exceeds the threshold, the server with the highest real-time monitored temperature is selected as the source server, and the tasks to be migrated are determined. Simultaneously, the temperature of the second-highest ranked core component is checked, and so on.

[0077] The above details various embodiments of the resource scheduling method. Based on this, this application also discloses a resource scheduling device, equipment, and storage medium corresponding to the above method.

[0078] Example 4

[0079] Figure 2 This is a schematic diagram of the structure of a resource scheduling device provided in an embodiment of this application.

[0080] like Figure 2 As shown, the resource scheduling device provided in this application embodiment includes:

[0081] Training unit 201 is used to train a prediction model in advance with the type of core component, the usage rate of the core component and the physical location of the server where the core component is located as input data and the temperature of the core component as output data.

[0082] The allocation prediction unit 202 is used to simulate the first temperature prediction value of the core component of the candidate server when the task to be allocated is assigned to the candidate server by the prediction model.

[0083] The allocation unit 203 is used to evenly distribute the tasks to be allocated to the candidate servers, provided that the first temperature prediction value does not exceed the safe temperature threshold of the candidate servers.

[0084] The core components include computing components, storage components, and network communication components.

[0085] Furthermore, the types of tasks to be assigned include at least one of business computing tasks, virtual machine scheduling tasks, or container scheduling tasks.

[0086] Furthermore, the resource scheduling device provided in this application embodiment also includes:

[0087] The over-temperature monitoring unit is used to determine the tasks to be migrated on the source server when the temperature of the core components of the source server exceeds the safe temperature threshold of the source server.

[0088] Furthermore, when the over-temperature monitoring unit detects that the temperature of the core components of the source server exceeds the safe temperature threshold of the source server, it determines the tasks to be migrated on the source server, specifically as follows:

[0089] When the device running the virtualization management system detects that the temperature of the core components of the source server exceeds the safe temperature threshold of the source server, it determines the virtual machine to be migrated on the source server.

[0090] Furthermore, when the device running the virtualization management system detects that the temperature of the core components of the source server exceeds the safe temperature threshold of the source server, the over-temperature monitoring unit determines the virtual machine to be migrated on the source server, specifically as follows:

[0091] The prediction model is used to predict the virtual machines running on the source server one by one, simulating the second temperature prediction value of the core components of the source server after different combinations of virtual machines are migrated from the source server, until the combination of virtual machines that reduces the second temperature prediction value to below the safe temperature threshold of the source server and minimizes the migration cost is determined, and then the virtual machines to be migrated are determined.

[0092] Furthermore, provided that the first predicted temperature does not exceed the safe temperature threshold of the candidate servers, the allocation unit 203 evenly distributes the tasks to be allocated to the candidate servers, specifically as follows:

[0093] Using a predictive model, predictions are made for servers with low real-time temperature readings. The simulation aims to find the optimal combination of servers where the predicted first temperature of each server does not exceed the safe temperature threshold of the server after the virtual machine to be migrated is moved to the optimal combination of servers.

[0094] Furthermore, the resource scheduling device provided in this application embodiment also includes:

[0095] The temperature difference monitoring unit is used to determine the server with the higher real-time temperature detection value to assign tasks when the temperature difference value of the same type of core component on different servers exceeds the temperature difference threshold.

[0096] Furthermore, the utilization rate of core components is specifically obtained through the host operating system running on the server in an in-band manner.

[0097] Furthermore, the physical location of the server containing the core component is specifically the coordinate position of the center of the server containing the core component in a three-dimensional coordinate system established with the location of the cooling equipment in the data center as the origin.

[0098] Furthermore, the temperature of the core components is obtained out-of-band via the baseboard management controller of the server.

[0099] Furthermore, a preset number of servers with the lowest real-time temperature readings are selected as candidate servers, with the preset number matching the number of tasks to be assigned.

[0100] Furthermore, the resource scheduling device provided in this application embodiment also includes:

[0101] The update unit is used to update the prediction model based on the measured real-time utilization rate of the core components, the real-time temperature detection value of the core components, the type of the corresponding core components, and the physical location of the server where the core components are located.

[0102] Since the embodiments of the apparatus and the embodiments of the method correspond to each other, please refer to the description of the embodiments of the method for the embodiments of the apparatus, which will not be repeated here.

[0103] Example 5

[0104] Figure 3 This is a schematic diagram of the structure of a resource scheduling device provided in an embodiment of this application.

[0105] like Figure 3 As shown, the resource scheduling device provided in this application embodiment includes:

[0106] Memory 310 is used to store computer program 311;

[0107] Processor 320 is configured to execute computer program 311, which, when executed by processor 320, implements the steps of the resource scheduling method as described in any of the above embodiments.

[0108] The processor 320 may include one or more processing cores, such as a 3-core processor or an 8-core processor. The processor 320 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor 320 may also include a main processor and a coprocessor. The main processor, also known as a Central Processing Unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 320 may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 320 may also include an Artificial Intelligence (AI) processor, which handles computational operations related to machine learning.

[0109] The memory 310 may include one or more storage media, which may be non-transitory. The memory 310 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In this embodiment, the memory 310 is used to store at least the following computer program 311, wherein, after being loaded and executed by the processor 320, the computer program 311 is able to implement the relevant steps in the resource scheduling method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 310 may also include an operating system 312 and data 313, and the storage method may be temporary storage or permanent storage. The operating system 312 may be Windows. The data 313 may include, but is not limited to, the data involved in the above methods.

[0110] In some embodiments, the resource scheduling device may further include a display screen 330, a power supply 340, a communication interface 350, an input / output interface 360, a sensor 370, and a communication bus 380.

[0111] Those skilled in the art will understand that Figure 3 The structure shown does not constitute a limitation on the resource scheduling device and may include more or fewer components than illustrated.

[0112] The resource scheduling device provided in this application includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the resource scheduling method as described above, with the same effect.

[0113] It should be noted that the device and equipment embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or modules may be electrical, mechanical, or other forms. Modules described as separate components may or may not be physically separate, and components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0114] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0115] If the integrated modules are implemented as software functional modules and sold or used as independent products, they can be stored in a storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods described in the various embodiments of this application.

[0116] Therefore, embodiments of this application also provide a storage medium storing a computer program, which, when executed by a processor, implements steps such as a resource scheduling method.

[0117] The storage medium can include various media that can store program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0118] The computer program contained in the storage medium provided in this embodiment can implement the steps of the resource scheduling method described above when executed by the processor, with the same effect.

[0119] The resource scheduling method, apparatus, device, and storage medium provided in this application have been described in detail above. The various embodiments in the specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus, device, and storage medium disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to in the method section. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

[0120] It should also be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

Claims

1. A resource scheduling method, characterized in that, include: The prediction model is trained by taking the type of core component, the usage rate of the core component, and the physical location of the server where the core component is located as input data and the temperature of the core component as output data. For the task to be assigned, the prediction model simulates the first temperature prediction value of the core component of the candidate server when the task to be assigned is assigned to the candidate server. Provided that the first temperature prediction value does not exceed the safe temperature threshold of the candidate server, the tasks to be assigned will be evenly distributed among the candidate servers. The core components include computing components, storage components, and network communication components; the physical location of the server containing the core components is specifically the coordinate position of the center of the server containing the core components in a three-dimensional coordinate system established with the location of the cooling equipment in the computer room as the origin.

2. The resource scheduling method according to claim 1, characterized in that, The types of tasks to be assigned include at least one of business computing tasks, virtual machine scheduling tasks, or container scheduling tasks.

3. The resource scheduling method according to claim 1, characterized in that, Also includes: When the temperature of the core components of the source server exceeds the safe temperature threshold of the source server, the task to be assigned is determined to be migrated on the source server.

4. The resource scheduling method according to claim 3, characterized in that, When the temperature of the core components of the source server exceeds the safe temperature threshold of the source server, the task to be assigned is determined to be migrated on the source server, specifically as follows: When the device running the virtualization management system detects that the temperature of the core components of the source server exceeds the safe temperature threshold of the source server, it determines the virtual machine to be migrated on the source server.

5. The resource scheduling method according to claim 4, characterized in that, When the device running the virtualization management system detects that the temperature of the core component of the source server exceeds the safe temperature threshold of the source server, the virtual machine to be migrated is determined on the source server, specifically as follows: The prediction model is used to predict the virtual machines running on the source server one by one, simulating the second temperature prediction value of the core components of the source server after different combinations of virtual machines are migrated from the source server, until the combination of virtual machines that reduces the second temperature prediction value to below the safe temperature threshold of the source server and minimizes the migration cost is determined, and this combination is determined as the virtual machine to be migrated.

6. The resource scheduling method according to claim 4, characterized in that, The step of evenly distributing the tasks to be assigned to the candidate servers, provided that the first predicted temperature value does not exceed the safe temperature threshold of the candidate servers, specifically involves: Using the prediction model, predictions are made for servers with low real-time temperature detection values ​​one by one. The model simulates the combination of candidate servers where the predicted first temperature value of each candidate server does not exceed the safe temperature threshold of each candidate server after the virtual machine to be migrated is migrated to each candidate server and the migration cost is minimized. The virtual machine to be migrated is then migrated to the combination of candidate servers.

7. The resource scheduling method according to claim 1, characterized in that, Also includes: When the temperature difference of the same type of core component on different servers exceeds the temperature difference threshold, the task to be assigned is determined on the server with the higher real-time temperature detection value.

8. The resource scheduling method according to claim 1, characterized in that, The utilization rate of core components is specifically obtained through the host operating system running on the server in an in-band manner.

9. The resource scheduling method according to claim 1, characterized in that, The temperature of the core components is obtained out-of-band through the baseboard management controller of the server.

10. The resource scheduling method according to claim 1, characterized in that, The server with the lowest real-time temperature detection value is selected as the candidate server, and the preset number of servers is consistent with the number of tasks to be assigned.

11. The resource scheduling method according to claim 1, characterized in that, Also includes: The prediction model is updated based on the measured real-time utilization rate of the core components, the real-time temperature detection value of the core components, the type of the corresponding core components, and the physical location of the server where the core components are located.

12. A resource scheduling device, characterized in that, include: The training unit is used to train a prediction model in advance, taking the type of core component, the usage rate of the core component, and the physical location of the server where the core component is located as input data and the temperature of the core component as output data. The allocation prediction unit is used to simulate the first temperature prediction value of the core component of the candidate server when the task to be allocated is assigned to the candidate server using the prediction model. The allocation unit is used to evenly allocate the task to be allocated to the candidate server, provided that the first temperature prediction value does not exceed the safe temperature threshold of the candidate server. The core components include computing components, storage components, and network communication components; the physical location of the server containing the core components is specifically the coordinate position of the center of the server containing the core components in a three-dimensional coordinate system established with the location of the cooling equipment in the computer room as the origin.

13. A resource scheduling device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program, which, when executed by the processor, implements the steps of the resource scheduling method as described in any one of claims 1 to 11.

14. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the resource scheduling method as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Data center task scheduling method based on dynamic temperature prediction model

    CN104317654A

  • Virtual machine migration planning and scheduling method and system based on temperature prediction and medium

    CN111625321A