Cluster resource scheduling method and device, electronic equipment and storage medium

By obtaining the resource characteristics and labels of candidate clusters, the problem of resource waste in cluster resource management is solved, efficient resource allocation and utilization is achieved, and the performance loss of nested solutions is reduced.

CN120371518APending Publication Date: 2025-07-25INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510462707.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

When cluster resource management with different technical architectures in the prior art, cross-cluster scheduling management cannot be managed, resulting in resource waste and performance loss, especially in the nested solution that the basic services cannot be destroyed when they are running, resulting in the problem of resource allocation.

Method used

By obtaining the resource attribute characteristics, resource status characteristics and resource tags of the candidate cluster from the cluster resource information database, the scheduling score is calculated, and the candidate cluster can be scheduled to run the jobs to be run based on the score, avoiding some of the resources that are permanently run and reducing the performance loss caused by cluster nesting.

Benefits of technology

It realizes that there is no need to run some resources on a permanent basis in cluster resource management with different technical architectures, and reasonably allocates resources, improving the utilization rate of cluster resources and system efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371518A_ABST
    Figure CN120371518A_ABST
Patent Text Reader

Abstract

Embodiments of the invention disclose a cluster resource scheduling method and apparatus, an electronic device and a storage medium. The method comprises the steps of obtaining resource attribute features, resource state features and resource tags of a candidate cluster from a cluster resource information base in response to a job operation request of a to-be-operated job; calculating schedulable scores of the candidate clusters based on job operation characteristics of the to-be-operated job and resource attribute characteristics, resource state characteristics and resource labels of the candidate clusters; and scheduling the candidate cluster to run the to-be-run job based on the schedulable score. The embodiment of the invention can improve the resource utilization rate of cluster resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud computing, and in particular, to a method and apparatus for scheduling cluster resources, an electronic device, and a storage medium. Background Art

[0002] When the prior art manages and allocates multiple sets of cluster resources with different technical architectures, since each cluster with a different technical architecture has a complete resource scheduling and management function, cross-cluster scheduling and management cannot be performed between multiple sets of clusters. A cluster nesting scheme needs to be adopted, that is, one set of clusters is nested within another set of clusters. As shown in Figure 1 shown, such a two-layer nesting method will cause waste of resources in an additional container layer. And since the basic services of the second set of containers nested within the first set of containers need to run permanently, even when there is no job running in the second set of containers, the corresponding first cluster containers cannot be destroyed because the corresponding second cluster basic services cannot be destroyed. Therefore, this part of the resources cannot be allocated to other applications in the first cluster, which will further cause waste of resources. Summary of the Invention

[0003] Embodiments of the present invention provide a method and apparatus for scheduling cluster resources, an electronic device, and a storage medium, which can improve the resource utilization rate of cluster resources.

[0004] In a first aspect, an embodiment of the present invention provides a method for scheduling cluster resources, including:

[0005] Obtaining the resource attribute characteristics, resource status characteristics, and resource labels of candidate clusters from a cluster resource information library in response to a job running request of a to-be-run job;

[0006] Calculating a schedulable score of the candidate clusters based on the job running characteristics of the to-be-run job and the resource attribute characteristics, resource status characteristics, and resource labels of the candidate clusters; and

[0007] Scheduling the candidate clusters to run the to-be-run job based on the schedulable score;

[0008] Wherein, the resource attribute characteristics of the candidate clusters include the hardware performance and geographical location of the candidate clusters, the resource status characteristics of the candidate clusters include the amount of idle resources and queued resources of the candidate clusters, and the resource labels of the candidate clusters include specific job labels for running specific jobs.

[0009] In a second aspect, an embodiment of the present invention provides a device for scheduling cluster resources, including:

[0010] A feature and label acquisition module, configured to obtain the resource attribute characteristics, resource status characteristics, and resource labels of candidate clusters from a cluster resource information library in response to a job running request of a to-be-run job;

[0011] A schedulable score calculation module, configured to calculate a schedulable score of the candidate cluster based on the job running characteristics of the to-be-run job, as well as the resource attribute characteristics, resource status characteristics, and resource tags of the candidate cluster; and

[0012] A scheduling and running module, configured to schedule the candidate cluster to run the to-be-run job based on the schedulable score;

[0013] Wherein, the resource attribute characteristics of the candidate cluster include the hardware performance and the geographical location of the candidate cluster, the resource status characteristics of the candidate cluster include the amount of idle resources and the amount of queued resources of the candidate cluster, and the resource tags of the candidate cluster include specific job tags for running specific jobs.

[0014] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the cluster resource scheduling method described in any one of the embodiments of the present invention is implemented.

[0015] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the cluster resource scheduling method described in any one of the embodiments of the present invention is implemented.

[0016] A cluster resource scheduling method, device, electronic device, and storage medium provided by an embodiment of the present invention, by obtaining the resource attribute characteristics, resource status characteristics, and resource tags of a candidate cluster from a cluster resource information library, calculating and determining the schedulable score of the candidate cluster based on the obtained information, and further scheduling the candidate cluster to run a to-be-run job based on the schedulable score of the candidate cluster, can enable the management and allocation of cluster resources with different technical architectures without the need to permanently run some cluster resources, reduce the performance loss brought by the cluster nesting solution, thereby reasonably and optimally allocate cluster resources with different technical architectures, and achieve the efficient operation of the cluster system. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is a schematic diagram of a nesting structure of a cluster nesting solution in the prior art;

[0019] Figure 2It is a flowchart of a cluster resource scheduling method provided by an embodiment of the present invention;

[0020] Figure 3 It is another flowchart of a cluster resource scheduling method provided by an embodiment of the present invention;

[0021] Figure 4 It is another flowchart of a cluster resource scheduling method provided by an embodiment of the present invention;

[0022] Figure 5 It is a structural diagram of a cluster resource scheduling device provided by an embodiment of the present invention;

[0023] Figure 6 It is a structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0024] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0026] Figure 2 It is a flowchart of a cluster resource scheduling method provided by an embodiment of the present invention. This embodiment is applicable to the scenario of simultaneously scheduling cluster resources of multiple different technical architectures. This method can be executed by the cluster resource scheduling device provided by the embodiment of the present invention, and the device can be implemented in a software and / or hardware manner. In a specific embodiment, the device can be integrated in an electronic device, such as a computer, a server, etc. The following embodiments will be described by taking the integration of the device in an electronic device as an example. Refer to Figure 2, the method may specifically include the following steps:

[0027] Step 201, obtain the resource attribute characteristics, resource status characteristics, and resource labels of the candidate clusters from the cluster resource repository in response to a job running request of a job to be run. Among them, the resource attribute characteristics of the candidate clusters include the hardware performance and geographical location of the candidate clusters, the resource status characteristics of the candidate clusters include the amount of idle resources and queued resources of the candidate clusters, and the resource labels of the candidate clusters include specific job labels for running specific jobs. This step can facilitate calculating the schedulable scores of the candidate clusters based on the resource attribute characteristics, resource status characteristics, and resource labels of the candidate clusters.

[0028] Specifically, the above-mentioned candidate clusters may include one cluster or multiple candidate clusters.

[0029] Specifically, the above-mentioned multiple candidate clusters may include multiple clusters with different technical architectures, and may also include multiple clusters with the same technical architecture.

[0030] Specifically, the above-mentioned amount of queued resources can be understood as the amount of resources required for the queued jobs in the running waiting queue of the candidate cluster.

[0031] Optionally, when the above-mentioned candidate clusters include multiple candidate clusters, obtain the resource attribute characteristics, resource status characteristics, and resource labels of each candidate cluster from the cluster resource repository respectively.

[0032] Specifically, the above-mentioned cluster resource repository can be established in advance and updated in real time.

[0033] Optionally, before step 201, obtain the latest data of the resource attribute characteristics, resource status characteristics, and resource labels of the candidate clusters based on the resource information update period, and update the latest data of the resource attribute characteristics, resource status characteristics, and resource labels of the candidate clusters to the cluster resource repository.

[0034] Optionally, the process of obtaining the resource status characteristics of the candidate clusters includes: obtaining the original data of the amount of idle resources and queued resources of each candidate cluster, and converting the original data of the amount of idle resources and queued resources of each candidate cluster into data represented by the same resource amount unit to obtain the resource status characteristics of each candidate cluster.

[0035] It can be understood that the resource amount units of clusters with different technical architectures are different, so unified conversion is required for subsequent calculations.

[0036] Specifically, the above-mentioned resource update period can be set based on empirical data, or can be set based on the resource change law of the candidate clusters after statistically analyzing the resource change law of the candidate clusters.

[0037] It is understandable that the information of candidate clusters changes in real time under the influence of various factors such as workload, node failures, and dynamic resource adjustment. For example, for emergency situations, some cluster resources need to be freed up during certain activities. Therefore, updating the latest data of the resource attribute characteristics, resource status characteristics, and resource labels of candidate clusters to the cluster resource information library based on the resource information update cycle can facilitate more reasonable scheduling of cluster resources based on the latest cluster resource information.

[0038] Specifically, the hardware performance of the above candidate clusters can include computing performance, storage performance, network performance, power and heat dissipation performance, as well as reliability and availability. The above computing performance can specifically include whether a Graphics Processing Unit (GPU) is configured, and the above storage performance includes whether a Solid State Drive (SSD) is configured.

[0039] Step 202, calculate the schedulable score of the candidate cluster based on the job running characteristics of the job to be run, as well as the resource attribute characteristics, resource status characteristics, and resource labels of the candidate cluster. This step can facilitate scheduling the candidate cluster to run the job to be run based on the schedulable score.

[0040] Optionally, the job running characteristics of the above job to be run include the resource attributes required for running the job to be run, the resource status required for running, and the resource labels required for running.

[0041] Specifically, the job running characteristics of the above job to be run can also include the running priority of the job to be run, the estimated running time, and the dependency relationship with other jobs.

[0042] Specifically, the above job to be run can include one job to be run or multiple jobs to be run.

[0043] Specifically, when the above job to be run includes multiple jobs to be run, the schedulable scores of the candidate clusters corresponding to each job to be run can be calculated separately, or the schedulable scores of the candidate clusters corresponding to multiple jobs to be run can be calculated as a whole.

[0044] Optionally, when the above candidate clusters include multiple candidate clusters, for each candidate cluster, calculate the schedulable score of the current candidate cluster based on the resource attribute characteristics, resource status characteristics, and resource labels of the current candidate cluster, as well as the job running characteristics of the job to be run.

[0045] Specifically, the input features can be determined based on the job running characteristics of the job to be run, as well as the resource attribute characteristics, resource status characteristics, and resource labels of the candidate cluster, and the schedulable score of the above candidate cluster can be predicted and obtained based on the above input features through a pre-trained schedulable score prediction model.

[0046] Optionally, the process of calculating the schedulable score of the candidate cluster based on the job running characteristics of the job to be run, as well as the resource attribute characteristics, resource status characteristics, and resource labels of the candidate cluster includes: respectively determining whether the resource attribute characteristics, resource status characteristics, and resource labels of the candidate cluster match the required resource attributes, required resource status, and required resource labels for running the job to be run, and correspondingly determining the attribute matching score, status matching score, and label matching score based on the judgment results; then calculating the resource value score based on the attribute matching score, status matching score, and label matching score, and calculating the schedulable score based on the resource value score.

[0047] Optionally, the process of calculating the schedulable score of the candidate cluster based on the job running characteristics of the job to be run, as well as the resource attribute characteristics, resource status characteristics, and resource labels of the candidate cluster includes: based on the job running characteristics of the job to be run, as well as the resource attribute characteristics and resource status characteristics of the candidate cluster, predicting the probability of resource fragmentation in the candidate cluster through a pre-trained resource fragmentation prediction model and calculating the fragmentation optimization score based on the probability of resource fragmentation in the candidate cluster, and then calculating the schedulable score of the candidate cluster based on the fragmentation optimization score and the resource label.

[0048] Specifically, the network structure of the above resource fragmentation prediction model can be a multi-dimensional convolutional neural network.

[0049] Specifically, the schedulable score can also be calculated based on the resource value score and the fragmentation optimization score.

[0050] Optionally, the process of calculating the schedulable score of the candidate cluster based on the job running characteristics of the job to be run, as well as the resource attribute characteristics, resource status characteristics, and resource labels of the candidate cluster includes: calculating the cluster pressure score of the candidate cluster based on the resource status characteristics of the candidate cluster, and then calculating the schedulable score of the candidate cluster based on the cluster pressure score, resource attribute characteristics, and resource labels of the candidate cluster.

[0051] Step 203, scheduling the candidate cluster to run the job to be run based on the schedulable score. Based on Steps 201 and 202, this step can avoid the need to permanently run some resources when managing and allocating cluster resources with different technical architectures, reduce the performance loss caused by the cluster nesting scheme, and thus reasonably and optimally allocate the cluster resources with different technical architectures to achieve the efficient operation of the cluster system.

[0052] Optionally, the process of scheduling the candidate cluster to run the job to be run based on the schedulable score includes: determining the candidate cluster with the highest schedulable score as the target scheduling cluster, and adding the job to be run as a queued job of the target scheduling cluster to the running waiting queue of the target scheduling cluster.

[0053] Specifically, it is also possible to directly schedule the candidate cluster with the highest schedulable score to run the above-mentioned job to be run.

[0054] Optionally, the process of determining the candidate cluster with the highest schedulable score as the target scheduling cluster includes: determining one or more candidate clusters with the highest schedulable score as the target scheduling cluster.

[0055] The following further introduces the cluster resource scheduling method provided by the embodiments of the present invention. As Figure 3 shown, that is Figure 2 step 202 in [] can include the following steps:

[0056] Step 2021: respectively determine whether the resource attribute characteristics, resource status characteristics, and resource labels of the candidate cluster correspond to the resource attributes required for running the job to be run, the resource status required for running, and the resource labels required for running, and correspondingly determine the attribute matching score, status matching score, and label matching score based on the judgment results.

[0057] Optionally, the process of determining whether the resource attribute characteristics of the candidate cluster match the resource attributes required for running the job to be run includes: determining whether the hardware performance of the candidate cluster meets the hardware requirements during the running of the job to be run. For example, when the job to be run is an artificial intelligence training job, it is determined whether the candidate cluster is equipped with a graphics processor, and / or it is determined whether the candidate cluster is equipped with a solid-state drive. It can be understood that the process of artificial intelligence training is not only computationally intensive, data-intensive, has high real-time requirements, and many iterations, but also reads a large amount of data and needs to frequently perform random read and write of model parameters. The graphics processor can process multiple data simultaneously and perform parallel computing. Its powerful parallel computing ability can increase the speed of these matrix operations by several times or even dozens of times, greatly shortening the training time. The solid-state drive can quickly load a large amount of training data into the memory, providing data support for model training and avoiding the idle of computing resources caused by slow data reading. Therefore, when the job to be run is an artificial intelligence training job, determining whether the candidate cluster is equipped with a graphics processor and determining whether the candidate cluster is equipped with a solid-state drive can help accurately evaluate whether the hardware performance of the candidate cluster is suitable for the job to be run.

[0058] Optionally, the process of determining whether the resource attribute characteristics of the candidate cluster match the resource attributes required for the job to be run includes: determining whether the geographical location of the candidate cluster is within the location range determined based on the geographical location of the job to be run. It can be understood that a longer geographical distance will result in a longer network transmission distance, thereby generating a larger network latency. This will affect job task allocation, data transmission, and the communication efficiency between nodes. Therefore, determining whether the geographical location of the candidate cluster is within the location range determined based on the geographical location of the job to be run can facilitate accurately evaluating whether the candidate cluster is suitable for the job to be run based on the geographical location of the candidate cluster.

[0059] Optionally, the process of determining whether the resource status characteristics of the candidate cluster match the resource status required for the job to be run includes: determining whether the amount of idle resources in the candidate cluster is not less than the amount of resources required for the job to be run. Specifically, it can also be determined whether the amount of idle resources in the candidate cluster is not less than the sum of the resources required for the job to be run and queued jobs.

[0060] Optionally, the process of determining whether the resource labels of the candidate cluster match the resource labels required for the job to be run includes: determining whether the resource labels of the candidate cluster are the specific job labels corresponding to the job to be run, or determining whether the candidate cluster is not labeled with any specific job labels.

[0061] It can be understood that in practice, there will usually be situations where emergency or for certain activities, it is necessary to free up cluster resources for running specific jobs. This information can be used to label the corresponding cluster resources as tags. At this time, if the job to be run is this specific job, it is necessary to determine whether the resource labels of the candidate cluster are the specific job labels corresponding to the job to be run. If the job to be run is not this specific job, it is necessary to determine whether the candidate cluster is not labeled with any specific job labels.

[0062] Specifically, if the candidate cluster is labeled with specific job labels, but the job to be run is not this specific job, the label matching score can be determined as negative infinity.

[0063] Step 2022, calculate the resource value score based on the attribute matching score, status matching score, and label matching score.

[0064] Optionally, the process of calculating the resource value score based on the attribute matching score, status matching score, and label matching score includes:

[0065] Perform weighted summation on the above attribute matching score, status matching score, and label matching score, and then perform normalization processing on the corresponding sum value to obtain the above resource value score.

[0066] Step 2023: Based on the job running characteristics of the job to be run, as well as the resource attribute characteristics and resource status characteristics of the candidate clusters, predict the probability of resource fragmentation generated by the candidate clusters through a pre-trained resource fragmentation prediction model.

[0067] Specifically, the input characteristics of the resource fragmentation prediction model can be determined based on the job running characteristics of the job to be run, as well as the resource attribute characteristics and resource status characteristics of the candidate clusters. The probability of each candidate cluster generating resource fragmentation is output through the resource fragmentation prediction model to generate a fragmentation probability heat map.

[0068] Step 2024: Calculate the fragmentation optimization score based on the probability of the candidate clusters generating resource fragmentation.

[0069] Optionally, the process of calculating the fragmentation optimization score based on the probability of the candidate clusters generating resource fragmentation includes: for each candidate cluster, when the probability of the current candidate cluster generating resource fragmentation is relatively high, determine the fragmentation optimization score of the current candidate cluster as a relatively small value; when the probability of the current candidate cluster generating resource fragmentation is relatively low, determine the fragmentation optimization score of the current candidate cluster as a relatively large value.

[0070] Specifically, the reciprocal of the probability of the current candidate cluster generating resource fragmentation can be obtained to get the fragmentation optimization score of the current candidate cluster.

[0071] Step 2025: Calculate the cluster pressure score of the candidate clusters based on the resource status characteristics of the candidate clusters.

[0072] Optionally, when the candidate clusters include multiple candidate clusters, the process of calculating the cluster pressure score of the candidate clusters based on the resource status characteristics of the candidate clusters includes:

[0073] Calculate the difference between the free resource amount and the queued resource amount of each candidate cluster to obtain the available resource amount of each candidate cluster, and perform normalization processing on the available resource amount of each candidate cluster to obtain the cluster pressure score of each candidate cluster.

[0074] Step 2026: Perform weighted summation on the resource value score, the fragmentation optimization score, and the cluster pressure score to obtain the schedulable score.

[0075] Specifically, before performing weighted summation on the resource value score, the fragmentation optimization score, and the cluster pressure score, the weights of the resource value score, the fragmentation optimization score, and the cluster pressure score can be determined based on the actual application scenario, and then the resource value score, the fragmentation optimization score, and the cluster pressure score are weighted and summed based on the corresponding weights to obtain the above schedulable score.

[0076] In an alternative embodiment of the present invention, if a candidate cluster is labeled with a specific job tag, but the job to be run is not that specific job, the resource value score of the candidate cluster is assigned negative infinity, and the weight of the above resource value score can be set to 1.

[0077] By dynamically estimating the value, resource fragmentation, and pressure of the cluster, the embodiments of the present invention can consider from multiple dimensions whether a candidate cluster can be scheduled to run the job to be run, improve the credibility of the calculated schedulable score of the candidate cluster, and facilitate further rationalizing the cluster resource allocation scheme and further improving the utilization rate of cluster resources.

[0078] The following further introduces the cluster resource scheduling method provided by the embodiments of the present invention. As Figure 4 shown, it may include the following steps:

[0079] Step 401, in response to a job running request of the job to be run, obtain the resource attribute characteristics, resource status characteristics, and resource tags of the candidate clusters from the cluster resource information library.

[0080] Step 402, calculate the schedulable score of the candidate cluster based on the job running characteristics of the job to be run and the resource attribute characteristics, resource status characteristics, and resource tags of the candidate cluster.

[0081] Step 403, determine the candidate cluster with the highest schedulable score as the target scheduling cluster, and add the job to be run as a queued job of the target scheduling cluster to the running waiting queue of the target scheduling cluster.

[0082] Step 404, determine the candidate clusters with schedulable scores less than the schedulable score threshold as non-schedulable clusters, and add each queued job of the non-schedulable clusters to the running waiting queues of other candidate clusters in the order from front to back in the running waiting queue of the non-schedulable clusters.

[0083] Specifically, the above other candidate clusters can be understood as other candidate clusters except the non-schedulable clusters among the multiple candidate clusters.

[0084] In a specific example, the above schedulable score threshold can be set to 0, that is, determine the candidate clusters with schedulable scores less than 0 as non-schedulable clusters.

[0085] Optionally, the process of adding each queued job of the non-schedulable clusters to the running waiting queues of other candidate clusters includes: taking each queued job of the non-schedulable clusters as the job to be run, and re-executing steps 101 to 103 based on the other candidate clusters.

[0086] In the embodiments of the present invention, by performing job rollback and re-scheduling according to the operation conditions of resources and applications, the resources can be re-allocated in a timely manner according to the changes of resources, which is beneficial to ensuring the timely operation of jobs and avoiding the impact on normal jobs when the cluster resources change to unavailable.

[0087] Figure 5 It is a structural diagram of cluster resource scheduling provided by the embodiments of the present invention. This device is applicable to execute the cluster resource scheduling provided by the embodiments of the present invention. As Figure 5 shown, the device may specifically include:

[0088] A feature and label acquisition module 501, configured to obtain the resource attribute features, resource status features, and resource labels of candidate clusters from the cluster resource information library in response to a job running request of a job to be run. Among them, the resource attribute features of the candidate clusters include the hardware performance and geographical location of the candidate clusters, the resource status features of the candidate clusters include the amount of idle resources and queued resources of the candidate clusters, and the resource labels of the candidate clusters include specific job labels for running specific jobs. This module is beneficial to calculating the schedulable scores of candidate clusters based on the resource attribute features, resource status features, and resource labels of the candidate clusters.

[0089] Optionally, the candidate clusters include multiple candidate clusters. Specifically, the above-mentioned multiple candidate clusters may include clusters with different technical architectures.

[0090] A schedulable score calculation module 502, configured to calculate the schedulable scores of candidate clusters based on the job running characteristics of the job to be run and the resource attribute features, resource status features, and resource labels of the candidate clusters. This module is beneficial to scheduling candidate clusters to run the job to be run based on the schedulable scores.

[0091] Optionally, the job running characteristics of the job to be run include the resource attributes required for running, the resource status required for running, and the resource labels required for running.

[0092] Optionally, the above-mentioned schedulable score calculation module 502 can specifically be configured to respectively determine whether the resource attribute features, resource status features, and resource labels of the candidate clusters correspond to and match the resource attributes required for running, the resource status required for running, and the resource labels required for running of the job to be run, and determine the attribute matching score, status matching score, and label matching score based on the judgment results; and

[0093] calculate the resource value score based on the attribute matching score, status matching score, and label matching score, and calculate the schedulable score based on the resource value score.

[0094] Optionally, the above schedulable score calculation module 502 can specifically be used to predict the probability of resource fragmentation in a candidate cluster based on the job running characteristics of the job to be run, as well as the resource attribute characteristics and resource status characteristics of the candidate cluster; and

[0095] calculate a fragmentation optimization score based on the probability of resource fragmentation in the candidate cluster.

[0096] Optionally, the above schedulable score calculation module 502 can specifically be used to calculate a schedulable score based on the resource value score and the fragmentation optimization score.

[0097] Optionally, the above schedulable score calculation module 502 can specifically be used to calculate a cluster pressure score of the candidate cluster based on the resource status characteristics of the candidate cluster.

[0098] Optionally, the above schedulable score calculation module 502 can specifically be used to perform a weighted sum of the resource value score, the fragmentation optimization score, and the cluster pressure score to obtain a schedulable score.

[0099] The scheduling and running module 503 is used to schedule the candidate cluster to run the job to be run based on the schedulable score. Combining this module with modules 501 and 502 can enable the management and allocation of cluster resources with different technical architectures without permanently running some resources, reduce the performance loss caused by the cluster nesting solution, and thus reasonably and optimally allocate the cluster resources with different technical architectures to achieve the efficient operation of the cluster system.

[0100] Optionally, the above scheduling and running module 503 can specifically be used to determine the candidate cluster with the highest schedulable score as the target scheduling cluster, and add the job to be run as a queued job of the target scheduling cluster to the running waiting queue of the target scheduling cluster.

[0101] Optionally, the cluster resource scheduling device provided by the embodiments of the present invention further includes a cluster resource information library module, which is used to obtain the latest data of the resource attribute characteristics, resource status characteristics, and resource labels of the candidate cluster based on the resource information update period, and update the latest data of the resource attribute characteristics, resource status characteristics, and resource labels of the candidate cluster to the cluster resource information library.

[0102] Optionally, the cluster resource scheduling device provided by the embodiments of the present invention further includes a rollback module, which is used to determine the candidate cluster with a schedulable score less than the schedulable score threshold as a non-schedulable cluster, and add each queued job of the non-schedulable cluster to the running waiting queue of other candidate clusters in the order from front to back in the running waiting queue of the non-schedulable cluster.

[0103] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. For the specific working process of the above-described functional modules, reference can be made to the corresponding process in the foregoing method embodiments, and details are not described herein again.

[0104] An embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the cluster resource scheduling method provided in any of the foregoing embodiments.

[0105] An embodiment of the present invention further provides a computer-readable medium, on which a computer program is stored. When the program is executed by a processor, it implements the cluster resource scheduling method provided in any of the foregoing embodiments.

[0106] An embodiment of the present invention further provides a computer program product, including a computer program which, when executed by a processor, implements the cluster resource scheduling method as described in any of the embodiments of the present invention.

[0107] Next, refer to Figure 6 , which shows a schematic structural diagram of a computer system 600 of an electronic device suitable for implementing the embodiments of the present invention. Figure 6 The illustrated electronic device is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0108] As Figure 6 shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the system 600 are also stored. The CPU 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0109] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as required. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as required so that a computer program read from it is installed into the storage section 608 as required.

[0110] Specifically, according to the embodiments disclosed by the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed by the present invention include a computer program product which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by a central processing unit (CPU) 601, the above functions defined in the system of the present invention are executed.

[0111] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0112] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0113] The modules and / or units involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules and / or units can also be provided in a processor. For example, it can be described as: a processor includes a feature and label acquisition module, a schedulable score calculation module, and a scheduling operation module. Among them, the names of these modules do not constitute a limitation to the modules themselves in some cases.

[0114] As another aspect, the present invention also provides a computer-readable medium. The computer-readable medium can be included in the device described in the above embodiments; or it can exist separately without being assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the device, the device includes: obtaining the resource attribute features, resource status features, and resource labels of candidate clusters from the cluster resource information library in response to a job running request of a job to be run; calculating the schedulable scores of candidate clusters based on the job running features of the job to be run and the resource attribute features, resource status features, and resource labels of candidate clusters; and scheduling the candidate clusters to run the job to be run based on the schedulable scores. Among them, the resource attribute features of candidate clusters include the hardware performance and geographical location of candidate clusters, the resource status features of candidate clusters include the amount of idle resources and queued resources of candidate clusters, and the resource labels of candidate clusters include specific job labels for running specific jobs.

[0115] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for scheduling cluster resources, characterized in that, Including: Obtaining the resource attribute characteristics, resource status characteristics, and resource labels of candidate clusters from the cluster resource repository in response to a job running request of a job to be run; Calculating the schedulable scores of the candidate clusters based on the job running characteristics of the job to be run, as well as the resource attribute characteristics, resource status characteristics, and resource labels of the candidate clusters; And Scheduling the candidate clusters to run the job to be run based on the schedulable scores; Wherein, the resource attribute characteristics of the candidate clusters include the hardware performance and geographical location of the candidate clusters, the resource status characteristics of the candidate clusters include the amount of idle resources and queued resources of the candidate clusters, and the resource labels of the candidate clusters include specific job labels for running specific jobs.

2. The cluster resource scheduling method according to claim 1, wherein Before obtaining the resource attribute characteristics, resource status characteristics, and resource labels of candidate clusters from the cluster resource repository in response to a job running request of a job to be run, the method further includes: Obtaining the latest data of the resource attribute characteristics, the resource status characteristics, and the resource labels of the candidate clusters based on the resource information update period, and updating the latest data of the resource attribute characteristics, the resource status characteristics, and the resource labels of the candidate clusters to the cluster resource repository.

3. The cluster resource scheduling method according to claim 1, wherein The job running characteristics of the job to be run include the resource attributes required for running the job to be run, the resource status required for running, and the resource labels required for running; The calculating the schedulable scores of the candidate clusters based on the job running characteristics of the job to be run, as well as the resource attribute characteristics, resource status characteristics, and resource labels of the candidate clusters includes: Respectively determining whether the resource attribute characteristics, resource status characteristics, and resource labels of the candidate clusters correspond to and match the resource attributes required for running the job to be run, the resource status required for running, and the resource labels required for running, and correspondingly determining an attribute matching score, a status matching score, and a label matching score based on the judgment results; and Calculating a resource value score based on the attribute matching score, the status matching score, and the label matching score, and calculating the schedulable score based on the resource value score.

4. The cluster resource scheduling method according to claim 3, wherein Before calculating the schedulable score based on the resource value score, the method further includes: Predicting the probability of resource fragmentation of the candidate clusters through a pre-trained resource fragmentation prediction model based on the job running characteristics of the job to be run, as well as the resource attribute characteristics and resource status characteristics of the candidate clusters; and Calculating a fragmentation optimization score based on the probability of resource fragmentation of the candidate clusters; The calculating the schedulable score based on the resource value score includes: Calculating the schedulable score based on the resource value score and the fragmentation optimization score.

5. The cluster resource scheduling method according to claim 4, wherein Before calculating the schedulable score based on the resource value score and the fragmentation optimization score, the method further includes: Calculating the cluster pressure score of the candidate clusters based on the resource status characteristics of the candidate clusters; Calculating the schedulable score based on the resource value score and the fragmentation optimization score includes: Performing a weighted sum of the resource value score, the fragmentation optimization score, and the cluster pressure score to obtain the schedulable score.

6. The cluster resource scheduling method according to claim 1, wherein The candidate clusters include multiple candidate clusters. Scheduling the candidate clusters to run the to-be-run job based on the schedulable score includes: Determining the candidate cluster with the highest schedulable score as the target scheduling cluster, and adding the to-be-run job as a queued job of the target scheduling cluster to the running waiting queue of the target scheduling cluster.

7. The cluster resource scheduling method according to claim 6, wherein It further includes: Determining the candidate clusters with a schedulable score less than the schedulable score threshold as non-schedulable clusters, and adding each queued job of the non-schedulable clusters to the running waiting queues of other candidate clusters in the order from front to back in the running waiting queue of the non-schedulable clusters.

8. A cluster resource scheduling device, characterized in that, It includes: A feature and label acquisition module, configured to obtain the resource attribute features, resource status features, and resource labels of the candidate clusters from the cluster resource information library in response to a job running request of the to-be-run job; A schedulable score calculation module, configured to calculate the schedulable score of the candidate clusters based on the job running features of the to-be-run job and the resource attribute features, resource status features, and resource labels of the candidate clusters; And A scheduling and running module, configured to schedule the candidate clusters to run the to-be-run job based on the schedulable score; Wherein, the resource attribute features of the candidate clusters include the hardware performance and the geographical location of the candidate clusters, the resource status features of the candidate clusters include the amount of idle resources and the amount of queued resources of the candidate clusters, and the resource labels of the candidate clusters include specific job labels for running specific jobs.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the cluster resource scheduling method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the cluster resource scheduling method according to any one of claims 1 to 7.