Resource elastic scaling method and device and computing equipment
By receiving the hierarchical tags and cluster status information of the computing task, determining the priority of the target cluster, and performing resource elastic scaling processing, the problem of insufficient hierarchical management of computing tasks in the existing technology is solved, and efficient resource management of the computing cluster and rapid response to computing tasks is realized.
Patent Information
- Application Number
- CN202510444813.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
AI Technical Summary
The existing resource elastic scaling technology cannot achieve hierarchical management of computing tasks, and the lack of targeted resource elastic scaling mechanism for computing clusters, resulting in waste of resources and inefficient computing efficiency.
By receiving the hierarchical tags of the computing tasks, the target cluster priority is determined, and the resource elastic scaling processing is performed based on cluster status information such as the number of hoarding tasks and resource usage indicators, including expansion and shrinking, and different scaling strategies are adopted for computing clusters of different priorities.
The hierarchical scheduling of computing tasks is realized, the execution efficiency of computing tasks is improved, resource waste is avoided, burst traffic can be better dealt with, and the speed and efficiency of resource elastic scaling processing is improved.
Smart Images

Figure CN120336018A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and particularly to a method and apparatus for elastic resource scaling, a computing device, a computer storage medium, and a computer program product. Background Art
[0002] With the rapid development of artificial intelligence, edge computing, and the audio and video industries, computing tasks have become increasingly complex, and the requirements for the scale of distributed computing resources have also been growing. Elastic resource scaling refers to dynamically adjusting the scale of computing, storage, network, and other resources according to the real-time load of the system or a preset policy to match changes in business requirements, achieving a precise balance between resource supply and business load, and being able to effectively handle fluctuations in business traffic.
[0003] Currently, the commonly used resource elastic scaling technology is the HPA scaling method. However, in the process of implementing this application, the inventor found that this method has at least the following deficiencies: it is unable to implement hierarchical management of computing tasks and computing clusters, and lacks a targeted elastic resource scaling mechanism for computing cluster resources. Summary of the Invention
[0004] In view of the above problems, this application is proposed to provide a method and apparatus for elastic resource scaling, a computing device, a computer storage medium, and a computer program product that overcome the above problems or at least partially solve the above problems.
[0005] According to one aspect of this application, there is provided a method for elastic resource scaling, including:
[0006] Receiving a computing task issued by a requestor;
[0007] Determining a matching target cluster priority according to the hierarchical label of the computing task;
[0008] Detecting the cluster status information of the computing cluster with the target cluster priority; wherein, when the target cluster priority is the first priority, the cluster status information includes the number of hoarded tasks;
[0009] Performing elastic resource scaling processing on the computing cluster with the target cluster priority according to the cluster status information.
[0010] Optionally, when the target cluster priority is the first priority, the cluster status information further includes resource usage metrics;
[0011] After detecting the cluster status information of the computing cluster with the target cluster priority, the method further includes:
[0012] Judging the busy and idle status of the computing cluster with the first priority according to the resource usage metrics;
[0013] If the computing clusters with the first priority are all in a busy state, hoard computing tasks.
[0014] Optionally, when the target cluster priority is the second priority, the cluster status information includes resource usage metrics.
[0015] Optionally, in the case where the target cluster priority is the first priority, further including performing resource elastic scaling processing on the computing clusters with the target cluster priority according to the cluster status information:
[0016] If the number of hoarded tasks reaches the first threshold, perform an expansion process on the computing clusters with the first priority;
[0017] If the number of hoarded tasks does not reach the second threshold, perform a scaling-down process on the computing clusters with the first priority.
[0018] Optionally, in the case where the target cluster priority is the second priority, further including performing resource elastic scaling processing on the computing clusters with the target cluster priority according to the cluster status information:
[0019] If there are computing clusters among the computing clusters with the second priority whose resource usage metrics reach the third threshold, perform an expansion process on the computing clusters with the second priority;
[0020] For the computing clusters among the computing clusters with the second priority whose resource usage metrics do not reach the fourth threshold, perform a scaling-down process on them.
[0021] Optionally, the third threshold includes: a first sub-threshold and a second sub-threshold, the second sub-threshold is greater than the first sub-threshold, and further including performing an expansion process on the computing clusters with the second priority:
[0022] For the computing clusters among the computing clusters with the second priority whose resource usage metrics are between the first sub-threshold and the second sub-threshold, add new working nodes to them;
[0023] If there are computing clusters among the computing clusters with the second priority whose resource usage metrics reach the second sub-threshold, determine the computing clusters to be configured, and configure the cluster priority of the computing clusters to be configured as the second priority.
[0024] Optionally, further including determining the computing clusters to be configured:
[0025] According to the resource usage metrics, screen the computing clusters to be configured from the computing clusters with the first priority; or, create at least one computing cluster as the computing clusters to be configured.
[0026] Optionally, further including configuring the cluster priority of the computing clusters to be configured as the second priority:
[0027] Update the cluster label corresponding to the computing cluster to be configured to the cluster label corresponding to the computing cluster with the second priority;
[0028] Determining a matching target cluster priority according to the classification label of the computing task further includes:
[0029] Determine the cluster label that matches the classification label of the computing task;
[0030] Determine the target cluster priority according to the matching cluster label.
[0031] Optionally, the step of performing expansion processing on the computing cluster with the first priority further includes:
[0032] Add working nodes to one or more computing clusters in the computing cluster with the first priority.
[0033] Optionally, the shrinkage processing is specifically: take offline the working nodes to be taken offline in the computing cluster to be shrunk; the method further includes:
[0034] Screen the working nodes to be taken offline from the working nodes of the computing cluster to be shrunk according to the node status information and / or the node creation time.
[0035] Optionally, the method further includes:
[0036] Detect the number of working nodes included in the computing cluster with the second priority;
[0037] For the computing clusters in the computing cluster with the second priority whose number of included working nodes is lower than the number threshold, update their cluster priority to the first priority.
[0038] According to another aspect of the present application, there is provided a resource elastic scaling device, including:
[0039] A receiving module, adapted to receive a computing task issued by a requestor;
[0040] A scheduling module, adapted to determine a matching target cluster priority according to the classification label of the computing task;
[0041] A detection module, adapted to detect the cluster status information of the computing cluster with the target cluster priority; wherein, when the target cluster priority is the first priority, the cluster status information includes the number of hoarded tasks;
[0042] A processing module, adapted to perform resource elastic scaling processing on the computing cluster with the target cluster priority according to the cluster status information.
[0043] Optionally, when the target cluster priority is the first priority, the cluster status information further includes resource usage metrics;
[0044] The scheduling module is further adapted to: determine the busy or idle state of the computing clusters with the first priority according to the resource usage metrics; if all the computing clusters with the first priority are in a busy state, hoard computing tasks.
[0045] Optionally, when the target cluster priority is the second priority, the cluster status information includes resource usage metrics.
[0046] Optionally, in the case where the target cluster priority is the first priority, according to the cluster status information, the processing module is further adapted to:
[0047] If the number of hoarded tasks reaches the first threshold, perform an expansion process on the computing clusters with the first priority;
[0048] If the number of hoarded tasks does not reach the second threshold, perform a scaling-down process on the computing clusters with the first priority.
[0049] Optionally, in the case where the target cluster priority is the second priority, the processing module 540 is further adapted to:
[0050] If there is a computing cluster in the computing clusters with the second priority whose resource usage metrics reach the third threshold, perform an expansion process on the computing clusters with the second priority;
[0051] For the computing clusters in the computing clusters with the second priority whose resource usage metrics do not reach the fourth threshold, perform a scaling-down process on them.
[0052] Optionally, the third threshold includes: a first sub-threshold and a second sub-threshold, the second sub-threshold is greater than the first sub-threshold, and the processing module is further adapted to:
[0053] For the computing clusters in the computing clusters with the second priority whose resource usage metrics are between the first sub-threshold and the second sub-threshold, add new working nodes to them;
[0054] If there is a computing cluster in the computing clusters with the second priority whose resource usage metrics reach the second sub-threshold, determine the computing cluster to be configured, and configure the cluster priority of the computing cluster to be configured as the second priority.
[0055] Optionally, the processing module is further adapted to:
[0056] According to the resource usage metrics, screen the computing clusters to be configured from the computing clusters with the first priority; or, create at least one computing cluster as the computing cluster to be configured.
[0057] Optionally, the processing module is further adapted to:
[0058] Update the cluster label corresponding to the computing cluster to be configured to the cluster label corresponding to the computing clusters with the second priority;
[0059] The scheduling module is further adapted to:
[0060] Determine a cluster label that matches the hierarchical label of the computing task;
[0061] Determine the target cluster priority according to the matching cluster label.
[0062] Optionally, the processing module is further adapted to:
[0063] Add working nodes to one or more computing clusters in the computing cluster with the first priority.
[0064] Optionally, the scale-down process specifically is: take offline the working nodes to be scaled down in the computing cluster to be scaled down; the processing module is further adapted to:
[0065] Screen the working nodes to be taken offline from the working nodes of the computing cluster to be scaled down according to the node status information and / or the node creation time.
[0066] Optionally, the detection module is further adapted to:
[0067] Detect the number of working nodes included in the computing cluster with the second priority;
[0068] The processing module is further adapted to:
[0069] For a computing cluster in which the number of working nodes included in the computing cluster with the second priority is lower than the quantity threshold, update its cluster priority to the first priority.
[0070] According to another aspect of the present application, there is provided a computing device, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus;
[0071] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations corresponding to the above resource elastic scaling method.
[0072] According to still another aspect of the present application, there is provided a computer storage medium, in which at least one executable instruction is stored, and the executable instruction causes the processor to execute the operations corresponding to the above resource elastic scaling method.
[0073] According to yet another aspect of the present application, there is provided a computer program product, including at least one executable instruction, and the executable instruction causes the processor to execute the operations corresponding to the above resource elastic scaling method.
[0074] According to the resource elastic scaling method, device, computing device, computer storage medium and computer program product provided by the embodiments of the present application, a computing task sent by a requester is received; according to the classification label of the computing task, a matching target cluster priority is determined; the cluster status information of the computing cluster with the target cluster priority is detected; when the target cluster priority is the first priority, the cluster status information includes the number of hoarded tasks; according to the cluster status information, resource elastic scaling processing is performed on the computing cluster with the target cluster priority. By the above method, the computing tasks are executed by the target cluster with the target cluster priority, which can realize the hierarchical scheduling of the computing tasks and improve the execution efficiency of the computing tasks; different cluster status information can be used as the basis for triggering resource elastic scaling processing for computing clusters with different priorities, making the resource elastic scaling strategy of the computing cluster targeted. At the same time, the number of hoarded tasks is introduced as a judgment basis in the resource elastic scaling mechanism of the computing cluster with the first priority, avoiding the problem of resource waste in the computing cluster with a low priority caused by the method of relying on resource utilization rate for resource elastic scaling in the prior art, and can also improve the speed of resource elastic scaling processing to better cope with sudden traffic.
[0075] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically exemplified below. Brief Description of the Drawings
[0076] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0077] Figure 1 The flowchart of the resource elastic scaling method provided by an embodiment of the present application is shown;
[0078] Figure 2 The flowchart of the resource elastic scaling method provided by another embodiment of the present application is shown;
[0079] Figure 3 The schematic diagram of the system architecture provided by another embodiment of the present application is shown;
[0080] Figure 4 The schematic diagram of the system architecture provided by another embodiment of the present application is shown;
[0081] Figure 5The figure shows a schematic functional structure diagram of a resource elastic scaling device provided by an embodiment of the present application;
[0082] Figure 6 The figure shows a schematic structural diagram of a computing device provided by an embodiment of the present application. Detailed implementation manners
[0083] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be completely conveyed to those skilled in the art.
[0084] First, the noun terms related to one or more embodiments of the present application are explained.
[0085] K8S: The full name is Kubernetes, which is an open-source system for automatically deploying, scaling, and managing containerized applications.
[0086] HPA, the full name is Horizontal Pod Autoscaler, which is a mechanism for automatically horizontally scaling the number of K8S node replicas and is the core mechanism for dynamically adjusting the number of Pod replicas based on load metrics.
[0087] Cloud platform: A cloud computing platform based on Kubernetes, which is used for automatically deploying, scaling, and managing containerized applications.
[0088] vCore: Virtual core, which is a concept describing the virtualized processing ability of computing resources and is commonly used in cloud computing platforms or large-scale distributed computing frameworks.
[0089] Cluster: An open-source framework is adopted to provide a computing layer for parallel processing.
[0090] Worker node: It contains processes. After a computing task is sent to the worker node, the processes in the worker node execute it.
[0091] Distributed computing system: It is configured with several computing clusters, and several worker nodes are configured in the computing clusters. Among them, a computing task is submitted to the distributed computing system for execution, specifically, it is assigned to the worker nodes in the specified computing cluster in the distributed computing system for execution.
[0092] Status code: It is used to represent the processing result of the server for a request.
[0093] In the prior art, the resource utilization rate is used as the basis for determining whether to trigger resource elastic scaling. However, this method has limitations. First, there is a lack of targeted resource elastic scaling strategies. The resource utilization rate directly affects the computing cost. The higher the resource utilization rate, the more computing tasks can be processed with fewer computing resources. However, if the utilization rate is too high, the execution of computing tasks will slow down. Second, it will result in the resource utilization rate always being lower than the resource utilization rate threshold when triggering expansion, thus causing resource waste.
[0094] Figure 1 The flowchart of the resource elastic scaling method provided by an embodiment of the present application is shown. The method of the embodiment of the present application is executed by a scheduling service in a distributed computing system. The scheduling service is used to schedule computing tasks to specific computing clusters, such as Figure 1 As shown, the method includes the following steps:
[0095] Step S110, receive the computing task sent by the requester.
[0096] The requester sends the computing task to the scheduling service. The scheduling service schedules the computing task to a certain computing cluster in the distributed computing system. The working nodes in the computing cluster use resources to execute the computing task to obtain the expected output. For example, in the video service, the computing tasks involved in a video processing include: transcoding task, watermark task, subtitle task, noise reduction task, frame extraction task, super-resolution task, etc. The working node executing the transcoding task includes transcoding the source stream, and the working node executing the watermark task includes adding a watermark to the original video.
[0097] Among them, the type of computing task and the requester can be one-to-one, that is, one requester is used to send a single type of computing task; the type of computing task and the requester can also be many-to-one. For example, one requester is used to send multiple types of computing tasks under the same large classification. In specific implementation, multiple-priority computing clusters can be configured in the distributed computing system for each type or each large classification of computing tasks. Continuing with the above example, multiple-priority computing clusters are configured to execute transcoding tasks with different priorities, and multiple-priority computing clusters are respectively configured to execute picture-related computing tasks with different priorities (picture-related computing tasks include watermark tasks and subtitle tasks). The method of the embodiment of the present application is specifically a resource elastic scaling method for multiple-priority computing clusters corresponding to the received computing tasks.
[0098] Step S120, determine the matching target cluster priority according to the classification label of the computing task.
[0099] Among them, the hierarchical label of the computing task is used to indicate the priority of the computing task. The priority of the computing task is determined by evaluating the importance of the computing task, and then the hierarchical label corresponding to the priority is assigned to the computing task.
[0100] The priority of the computing task and the cluster priority corresponding to the computing cluster executing the computing task are in a one-to-one correspondence. When performing computing task scheduling, it is first necessary to determine the target cluster priority corresponding to the computing cluster executing the computing task according to the hierarchical label of the computing task. For example, for a high-priority computing task, the computing cluster executing the computing task is a high-priority computing cluster; for a medium-priority computing task, the computing cluster executing the computing task is a medium-priority computing cluster.
[0101] Specifically in implementation, the corresponding relationship between the hierarchical label and the cluster priority is pre-configured. Then, after receiving the computing task, the cluster priority corresponding to it is queried according to the hierarchical label of the computing task, and the queried cluster priority is the target cluster priority.
[0102] Step S130, detect the cluster status information of the computing cluster with the target cluster priority.
[0103] The distributed computing system includes at least two computing clusters with cluster priorities, and there is at least one computing cluster for each cluster priority. The cluster status information is used to indicate the computing task situation corresponding to the computing cluster, and can include the overall computing task situation of the computing clusters with the same cluster priority and the computing task situation of each computing cluster. The cluster status information specifically includes: the number of hoarded tasks, the relevant information of the computing tasks in execution, etc.
[0104] Step S140, perform resource elastic scaling processing on the computing cluster with the target cluster priority according to the cluster status information.
[0105] When the target cluster priority is the first priority, the cluster status information includes the number of hoarded tasks. The first priority is specifically at least one cluster priority lower than the highest cluster priority, preferably the lowest cluster priority. Correspondingly, the computing cluster with the first priority is used to execute non-highest-priority computing tasks.
[0106] The number of hoarded tasks refers to the total number of hoarded computing tasks waiting to be executed. There are many reasons for the computing tasks to be hoarded. The most important reason is that the resources of the computing cluster have been exhausted and no more resources can be allocated for the computing tasks, resulting in the computing tasks being hoarded.
[0107] After the scheduling service determines the target cluster priority corresponding to a computing task, if it is unable to determine the target computing cluster for executing the computing task, the computing task will be hoarded. Based on this, the number of hoarded tasks is for the computing clusters of the first priority as a whole, rather than the number of hoarded tasks of a specific computing cluster. The number of hoarded tasks of the computing clusters of the first priority is also the total number of computing tasks waiting to be executed by the computing clusters of the first priority.
[0108] In an alternative approach, if the target cluster priority corresponding to a computing task is the first priority, but none of the computing clusters of the first priority can execute the computing task, the computing task is added to the task queue.
[0109] Resource elastic scaling processing includes scaling-up processing and scaling-down processing. The purpose of scaling-up processing is to increase the resource scale of the computing clusters of the first priority, and the purpose of scaling-down processing is to reduce the resource scale of the computing clusters of the first priority. For the computing clusters of the first priority, it is determined whether to trigger scaling-up processing and scaling-down processing based on the number of hoarded tasks.
[0110] In addition, if it is determined according to the cluster status information that the computing clusters of the target cluster priority have idle resources, a target computing cluster is selected from the computing clusters of the target cluster priority, and the computing task is scheduled to the target computing cluster for execution.
[0111] In summary, according to the resource elastic scaling method provided in this embodiment, by receiving the computing tasks sent by the requestor, determining the matching target cluster priority according to the classification label of the computing tasks, and having the target cluster of the target cluster priority execute the computing tasks, hierarchical scheduling of the computing tasks can be achieved, and the execution efficiency of the computing tasks can be improved; by detecting the cluster status information of the computing clusters of the target cluster priority, when the target cluster priority is the first priority, the cluster status information includes the number of hoarded tasks, and according to the cluster status information, resource elastic scaling processing is performed on the computing clusters of the target cluster priority, which can use different cluster status information as the basis for triggering resource elastic scaling processing for computing clusters of different priorities, making the resource elastic scaling strategy of the computing clusters targeted. At the same time, the number of hoarded tasks is introduced as a judgment basis in the resource elastic scaling mechanism of the computing clusters of the first priority, avoiding the problem of resource waste in low-priority computing clusters caused by the method of relying on resource utilization rate for resource elastic scaling in the prior art, and also being able to improve the speed of resource elastic scaling processing to better handle sudden traffic.
[0112] Figure 2 Shows a flowchart of a resource elastic scaling method provided by another embodiment of the present application. As Figure 2 shown, the method includes the following steps:
[0113] Step S210: Receive the computing task sent by the requesting party.
[0114] The requesting party sends the computing task to the scheduling service. In the subsequent process, the scheduling service schedules the computing task to a certain computing cluster in the distributed computing system, and the worker nodes in the computing cluster utilize resources to execute the computing task to obtain the desired output.
[0115] Step S220: Determine the target cluster priority that matches according to the grading label of the computing task.
[0116] Among them, the grading label of the computing task is used to indicate the priority of the computing task. The priority of the computing task is determined by evaluating the importance of the computing task, and then the grading label corresponding to the priority is assigned to the computing task.
[0117] There is a one-to-one correspondence between the priority of the computing task and the cluster priority corresponding to the computing cluster that executes the computing task. When scheduling the computing task, it is first necessary to determine the target cluster priority corresponding to the computing cluster that executes the computing task according to the grading label of the computing task.
[0118] Specifically, determine the cluster label that matches the grading label of the computing task, and determine the target cluster priority according to the matching cluster label. The cluster label is used to indicate the priority of the computing cluster, and the cluster priority can be obtained by reading the cluster label of the computing cluster. The corresponding relationship between the grading label and the cluster label is pre-configured. Then, after receiving the computing task, query the cluster label that has a corresponding relationship with it according to the grading label of the computing task, and the cluster priority indicated by the queried cluster label is the target cluster priority.
[0119] Step S230: Detect the cluster status information of the computing cluster with the target cluster priority.
[0120] When the target cluster priority is the first priority, the cluster status information includes the number of hoarded tasks and the resource usage indicator; among them, the number of hoarded tasks is the total number of computing tasks that are hoarded and waiting for each computing cluster with a certain priority to execute, which is for the overall cluster priority; the resource usage indicator is used to indicate the resource usage situation of the computing cluster, and specifically can be the resource utilization rate, that is, the ratio of the allocated resource number to the total resource number, which is an indicator of a single computing cluster.
[0121] When the target cluster priority is the second priority, the cluster status information includes the resource usage indicator; among them, the second priority is specifically at least one cluster priority higher than the lowest cluster priority, preferably the highest cluster priority. This application does not make any limitations in this regard. Correspondingly, the computing cluster with the second priority is used to execute the computing tasks with non-lowest priorities.
[0122] When it is determined that the target cluster priority is the first priority, jump to execute step S240; when it is determined that the target cluster priority is the second priority, jump to execute step S260.
[0123] Step S240: Judge the busy and idle status of the computing clusters with the first priority according to the resource usage metrics; if all the computing clusters with the first priority are in a busy state, hoard the computing tasks.
[0124] Specifically, for the computing clusters with the first priority, the resource usage metric threshold for judging the busy and idle status can be set to a relatively high value. For example, when all the resources of the computing cluster are occupied, that is, the resource utilization rate is 100%, it is in a busy state; when there are unallocated resources, that is, the resource utilization rate does not reach 100%, it is in an idle state. In this way, the resource utilization rate of the computing clusters for processing low-priority computing tasks is improved.
[0125] If all the computing clusters with the first priority are in a busy state, then hoard the computing tasks. Further, return a specified status code to the requester, and the specified status code is used to indicate that the computing task has been hoarded. For each hoarded computing task, wait until there is an idle computing cluster in the computing clusters with the first priority and then execute them in sequence according to the order, and specifically determine the execution order according to the sequence of the issuance times of each computing task.
[0126] In addition, if there is an idle computing cluster among the computing clusters with the first priority, select a target computing cluster for executing the computing task from the idle computing clusters; specifically, if there is only one idle computing cluster, determine it as the target computing cluster; if there are multiple idle computing clusters, determine the computing cluster with the lowest resource usage metric among them as the target computing cluster for executing the computing task; then, schedule the computing task to the target computing cluster for the worker nodes in the target computing cluster to execute the computing task. Correspondingly, if there is an idle computing cluster among the computing clusters with the first priority, it means that there are no hoarded computing tasks at this time, and the number of hoarded tasks is zero, so there is no need to perform resource elastic scaling processing on the computing clusters with the first priority.
[0127] Step S250: Perform resource elastic scaling processing on the computing clusters with the first priority according to the number of hoarded tasks.
[0128] In an optional manner, hoard the computing tasks by adding the computing tasks to the task queue, that is, if the target cluster priority corresponding to the computing task is the first priority, but all the computing clusters with the first priority are in a busy state, then add the computing task to the task queue. The task queue can be set on the requester side or in the scheduling service, and this application does not make any limitations in this regard.
[0129] Among them, the resource elastic scaling process includes scaling-up processing. If the number of hoarded tasks reaches the first threshold, scaling-up processing is performed on the computing clusters with the first priority to increase the resource scale of the computing clusters with the first priority.
[0130] In an optional manner, the scaling-up processing includes the scaling-up method of adding working nodes to the computing cluster and the scaling-up method of increasing the computing cluster.
[0131] The specific method of adding working nodes to the computing cluster is as follows: For one or more computing clusters in the computing clusters with the first priority, new working nodes are added. Among them, if the required number is 1, that is, only one working node needs to be added, any one of the computing clusters is selected to add a working node; if the required number is greater than 1, that is, at least two working nodes need to be added, then multiple computing clusters are selected, and at least one working node is added to each of the computing clusters. The total number of newly added working nodes in the multiple computing clusters is equal to the required number. This method is more efficient than the method of adding the required number of working nodes in one computing cluster.
[0132] Specifically, the scheduling service creates working nodes through the cloud platform interface, and then connects the newly created working nodes to the master node of the computing cluster. Among them, the master node serves as the control center of the cluster, responsible for scheduling tasks, management, and maintaining the global state, etc. The working nodes are mainly responsible for executing actual computing tasks, receiving task instructions from the master node, and executing computing tasks.
[0133] The specific method of increasing the computing cluster is as follows: A new computing cluster is created and the newly created computing cluster is configured as a computing cluster with the first priority. That is, a new computing cluster is constructed and its cluster priority is updated to the first priority. This method will not affect the normal operation of other computing clusters; or, a computing cluster to be configured is screened from computing clusters with other priorities, and the computing cluster to be configured is configured as a computing cluster with the first priority. Specifically, the computing cluster with the lowest resource usage index among the computing clusters with other priorities is screened as the computing cluster to be configured. This method uses the resources of the computing clusters with other priorities to supplement the resources of the computing clusters with the first priority and will not incur the cost of creating a new computing cluster.
[0134] The resource elastic scaling process also includes the downscaling process. Specifically, if the number of hoarded tasks does not reach the second threshold, the downscaling process is performed on the computing clusters with the first priority. The second threshold is less than the first threshold, and the specific values of the first threshold and the second threshold can be flexibly set according to the actual business needs. The downscaling process specifically involves taking offline the worker nodes to be downscaled in the computing cluster to be downscaled (i.e., the computing cluster that needs to undergo the downscaling process). The scheduling service takes the offline worker nodes by calling the cloud platform. The downscaling process can specifically involve disconnecting the connection between the worker nodes to be downscaled and the master node of the computing cluster and reclaiming the resources of the worker nodes to be downscaled.
[0135] Step S260: Perform resource elastic scaling on the computing clusters with the second priority according to the resource usage metrics.
[0136] Specifically, if there is a computing cluster in the computing clusters with the second priority whose resource usage metric reaches the third threshold, the upscaling process is performed on the computing clusters with the second priority. The third threshold is less than 100%.
[0137] In an optional manner, the upscaling process includes the upscaling method of adding worker nodes to the computing cluster and the upscaling method of adding computing clusters. The triggering conditions for the two methods are different.
[0138] Specifically, the third threshold includes: a first sub-threshold and a second sub-threshold. The second sub-threshold is greater than the first sub-threshold, and both the first sub-threshold and the second sub-threshold are less than 100%. The upscaling method of adding worker nodes to the computing cluster is specifically: for the computing clusters in the computing clusters with the second priority whose resource usage metrics are between the first sub-threshold and the second sub-threshold, new worker nodes are added to them. That is, if the resource utilization rate is between the first sub-threshold and the second sub-threshold, new worker nodes are added to the computing cluster. The upscaling method of adding computing clusters is specifically: if there is a computing cluster in the computing clusters with the second priority whose resource usage metric reaches the second sub-threshold, the computing cluster to be configured is determined, and the cluster priority of the computing cluster to be configured is configured as the second priority. That is, once it is detected that there is a computing cluster in the computing clusters with the second priority whose resource usage metric reaches the second sub-threshold, the upscaling is achieved by adding computing clusters. Through the above method, a phased upscaling process is realized, which can ensure that the computing clusters with high priority can always reserve a certain amount of resources to ensure the execution efficiency of computing tasks, and can also avoid resource waste caused by excessive increase in resource scale.
[0139] In an optional manner, determining the computing cluster to be configured specifically includes: selecting the computing cluster to be configured from the computing clusters of the first priority according to the resource usage index, determining the computing cluster with the lowest resource usage index from the computing clusters of the first priority as the computing cluster to be configured, changing the cluster priority of the computing cluster to be configured from the first priority to the second priority, and after determining the computing cluster to be configured, first stop scheduling computing tasks to it, and at the same time, quickly exit the computing tasks being executed in the computing cluster to be configured, return a notification of computing task execution failure and computing task backlog to the requester, and then change the cluster priority. In the above manner, the low-priority computing cluster with the lowest resource usage is used to supplement the high-priority resource pool, and the cost of creating a new computing cluster will not be incurred.
[0140] In an optional manner, determining the computing cluster to be configured specifically includes: creating at least one computing cluster as the computing cluster to be configured, and configuring the cluster priority of the computing cluster to be configured as the second priority. This manner is to create a computing cluster and configure it as the second priority without affecting the normal operation of other computing clusters.
[0141] The step of configuring the cluster priority of the computing cluster to be configured as the second priority specifically includes: updating the cluster label corresponding to the computing cluster to be configured to the cluster label corresponding to the computing cluster of the second priority, where the cluster label is used to indicate the priority of the computing cluster.
[0142] In addition, if there is a computing cluster whose resource usage index does not reach the fourth threshold among the computing clusters of the second priority, the computing cluster whose resource usage index does not reach the fourth threshold among the computing clusters of the second priority is scaled down. The fourth threshold is less than the third threshold, and in particular, in the aforementioned staged expansion processing method, the fourth threshold is less than the first sub-threshold. If there is a computing cluster whose resource usage rate is less than the fourth threshold among the computing clusters of the second priority, it means that more resources of the computing cluster are idle, and the computing cluster is scaled down to improve the resource utilization rate of the computing cluster.
[0143] The specific process of downscaling is as follows: if the working nodes to be downgraded in the computing cluster to be downgraded are downgraded, then the working nodes to be downgraded need to be screened from the computing cluster to be downgraded (i.e., the computing cluster that needs to be downgraded), specifically: based on the node status information and / or the node creation time, the working nodes to be downgraded are screened from the working nodes of the computing cluster to be downgraded. Among them, the node status information is used to indicate the idle and busy status of the working node, which can be the resource utilization rate of the working node. When the resource scale of each working node in the computing cluster is consistent, the node status information can also be the number of computing tasks being executed.
[0144] Specifically, in the case where the node status information of each working node is inconsistent, select the most idle working node (i.e., the one with the lowest resource utilization rate or the fewest computing tasks in execution) as the working node to be taken offline; in the case where the node status information of each working node is consistent but the node creation times are inconsistent, select the working node with the earliest creation time as the working node to be taken offline; in the case where the node status information and the node creation times of each working node are both consistent, randomly select a working node to be taken offline from each working node.
[0145] After determining the working node to be taken offline, stop the computing task being executed by the working node to be taken offline, notify the requester corresponding to the computing task, and then disconnect the connection between the working node to be taken offline and the master node in the computing cluster.
[0146] In an optional manner, the method further includes: detecting the number of working nodes included in the computing clusters of the second priority; for the computing clusters in which the number of working nodes included in the computing clusters of the second priority is lower than the number threshold, update their cluster priority to the first priority. Specifically, after performing the scale-down processing on the computing clusters of the second priority, the number of working nodes included in each computing cluster of the second priority can be detected. If there is a computing cluster in which the number of working nodes included is lower than the number threshold, it indicates that relatively few computing tasks need to be executed by the computing clusters of the second priority. To avoid the problem of resource waste, the cluster priority of the computing cluster with fewer working nodes is updated from the second priority to the first priority to supplement the resource pool of the first priority, and the resource utilization rate of the distributed computing system and the execution efficiency of the computing tasks are guaranteed through dynamic resource adjustment.
[0147] In addition, if it is determined that the target cluster priority corresponding to the computing task is the second priority, then screen the target computing cluster from each computing cluster of the second priority, and schedule the computing task to the target computing cluster for execution. Specifically, screen out the computing cluster with the lowest resource utilization rate as the target computing cluster. Since the scale-out processing of the computing clusters of the second priority is performed when the resource utilization rate reaches a certain threshold, there are always idle resources in the computing clusters of the second priority. Therefore, for high-priority computing tasks, they can be directly scheduled to the computing cluster with the lowest resource utilization rate for execution, and a certain amount of resources can always be reserved for high-priority tasks to ensure the execution efficiency.
[0148] In summary, according to the resource elastic scaling method provided in this embodiment, by determining the target cluster priority that matches according to the hierarchical label of the computing task, and having the target cluster with the target cluster priority execute the computing task, hierarchical scheduling of the computing task can be achieved, and the execution efficiency of the computing task can be improved; judging whether to trigger resource elastic scaling processing and executing resource elastic scaling processing after receiving the computing task can effectively handle sudden traffic and improve the efficiency of resource elastic scaling; for low-priority computing clusters, the scaling logic uses the number of tasks hoarded when the resources are fully utilized as the judgment basis, and for high-priority computing clusters, the scaling logic uses the resource utilization rate as the judgment basis, realizing hierarchical scaling management of computing clusters, making the resource elastic scaling strategy of computing clusters have the characteristic of pertinence, which can improve the performance of computing cluster resource management, reserve a certain amount of resources for the computing cluster processing high-priority tasks to process new computing tasks to ensure the execution efficiency of high-priority tasks, and let the resources of the computing cluster processing low-priority tasks be fully utilized to improve the resource utilization rate; three different expansion processing methods are also provided, namely, the expansion method of adding working nodes, the expansion method of creating a new computing cluster, and the expansion method of converting the computing cluster priority, realizing phased expansion processing and intelligent expansion processing; the expansion method of adding working nodes can improve the expansion efficiency, the expansion method of creating a new computing cluster can avoid affecting the normal operation of other computing clusters, and the expansion method of converting the computing cluster priority can avoid the cost of creating a new computing cluster; in the scaling processing method of taking offline working nodes, the working nodes to be taken offline can be reasonably and intelligently selected to reduce other transactions and costs caused by taking offline the working nodes in the working process; for the computing cluster with the second priority, it is determined whether to perform priority downgrading according to the number of working nodes, which helps to achieve the balance of resource utilization rates among computing clusters with different priorities.
[0149] Figure 3 FIG. shows a schematic diagram of a system architecture provided by another embodiment of the present application, as Figure 3 shown. The main bodies involved in this system include: a scheduling service and a resource pool. The resource pool contains n computing clusters, each computing cluster has a cluster priority and contains several working nodes. The low-priority computing cluster is the computing cluster with the first priority, and the high-priority computing cluster is the computing cluster with the second priority. A request issues a computing task to the scheduling service. The scheduling service first selects a cluster, that is, filters out the target computing cluster with the target cluster priority for processing the computing task from the resource pool, and then the scheduling service submits the task, that is, submits the computing task to the selected target computing cluster.
[0150] Figure 4 FIG. shows a schematic diagram of a system architecture provided by another embodiment of the present application, as Figure 4As shown in the figure, the low-priority computing cluster contains x worker nodes, and the high-priority computing cluster contains m worker nodes. Resource scheduling for the cloud platform resource pool is achieved through virtual cores (vCores). Among them, for the low-priority computing cluster, resource elastic scaling is achieved according to the computing tasks hoarded in the task queue and the relevant thresholds of the hoarded task quantity; for the high-priority computing cluster, resource elastic scaling is achieved according to the resource utilization rate of each computing cluster and the relevant thresholds of the resource utilization rate.
[0151] For the low-priority computing cluster, if the number of hoarded tasks is greater than threshold 1, expansion is triggered, that is, the scheduling service creates worker nodes through the cloud platform interface and connects them to the master node of the computing cluster; if the number of hoarded tasks is less than threshold 2, contraction is triggered, that is, the scheduling service invokes the cloud platform interface to take offline the most idle or the earliest created worker node of the computing cluster.
[0152] For the high-priority computing cluster, if the resource utilization rate is greater than threshold 3, expansion is triggered, and the expansion method is to add worker nodes; if the resource utilization rate is less than threshold 4, contraction is triggered, and the contraction logic is to take offline worker nodes; if the resource utilization rate continues to increase and reaches threshold 5, the following logic is triggered: select the most idle low-priority computing cluster, stop scheduling, the computing tasks being executed in the low-priority computing cluster fail quickly and exit, and notify the requester, change the label of the low-priority computing cluster to high-priority, supplement the high-priority resource pool, so that high-priority computing tasks can be executed in a timely manner, and low-priority computing tasks are hoarded waiting for resources; in addition, after the high-priority computing cluster is contracted, if the number of worker nodes is lower than the threshold, its label is changed to low-priority to supplement the low-priority resource pool, and so on. Among them, the high-priority computing cluster is also configured with a task queue. Under normal circumstances, the high-priority computing cluster always reserves a certain amount of resources and will not cause hoarding of high-priority computing tasks, but in the case of sudden high traffic scenarios, it may cause hoarding of high-priority computing tasks.
[0153] Figure 5 The functional structure diagram of the resource elastic scaling device provided by the embodiment of the present application is shown.
[0154] As Figure 5 shown, the device includes:
[0155] A receiving module 510, adapted to receive computing tasks issued by the requester;
[0156] A scheduling module 520, adapted to determine the matching target cluster priority according to the classification label of the computing tasks;
[0157] A detection module 530, adapted to detect the cluster status information of the computing cluster with the target cluster priority; among them, when the target cluster priority is the first priority, the cluster status information includes the number of hoarded tasks;
[0158] The processing module 540 is adapted to perform resource elastic scaling processing on the computing clusters with the target cluster priority according to the cluster status information.
[0159] In an alternative manner, when the target cluster priority is the first priority, the cluster status information further includes resource usage metrics;
[0160] The scheduling module 520 is further adapted to: judge the busy or idle status of the computing clusters with the first priority according to the resource usage metrics; if all the computing clusters with the first priority are in a busy state, hoard computing tasks.
[0161] In an alternative manner, when the target cluster priority is the second priority, the cluster status information includes resource usage metrics.
[0162] In an alternative manner, when the target cluster priority is the first priority, according to the cluster status information, the processing module 540 is further adapted to:
[0163] If the number of hoarded tasks reaches the first threshold, perform an expansion process on the computing clusters with the first priority;
[0164] If the number of hoarded tasks does not reach the second threshold, perform a contraction process on the computing clusters with the first priority.
[0165] In an alternative manner, when the target cluster priority is the second priority, the processing module 540 is further adapted to:
[0166] If there are computing clusters in the second-priority computing clusters whose resource usage metrics reach the third threshold, perform an expansion process on the second-priority computing clusters;
[0167] For the computing clusters in the second-priority computing clusters whose resource usage metrics do not reach the fourth threshold, perform a contraction process on them.
[0168] In an alternative manner, the third threshold includes: a first sub-threshold and a second sub-threshold, the second sub-threshold is greater than the first sub-threshold, and the processing module 540 is further adapted to:
[0169] For the computing clusters in the second-priority computing clusters whose resource usage metrics are between the first sub-threshold and the second sub-threshold, add new working nodes to them;
[0170] If there are computing clusters in the second-priority computing clusters whose resource usage metrics reach the second sub-threshold, determine the computing clusters to be configured, and configure the cluster priority of the computing clusters to be configured as the second priority.
[0171] In an alternative manner, the processing module 540 is further adapted to:
[0172] Screen the computing clusters to be configured from the computing clusters of the first priority according to the resource usage metrics; or, create at least one computing cluster as the computing clusters to be configured.
[0173] In an alternative manner, the processing module 540 is further adapted to:
[0174] Update the cluster label corresponding to the computing clusters to be configured to the cluster label corresponding to the computing clusters of the second priority;
[0175] The scheduling module 520 is further adapted to:
[0176] Determine the cluster label that matches the hierarchical label of the computing task;
[0177] Determine the target cluster priority according to the matching cluster label.
[0178] In an alternative manner, the processing module 540 is further adapted to:
[0179] Add working nodes to one or more computing clusters in the computing clusters of the first priority.
[0180] In an alternative manner, the scale-down processing specifically is: take offline the working nodes to be taken offline in the computing clusters to be scaled down; the processing module 540 is further adapted to:
[0181] Screen the working nodes to be taken offline from the working nodes of the computing clusters to be scaled down according to the node status information and / or the node creation time.
[0182] In an alternative manner, the detection module 530 is further adapted to:
[0183] Detect the number of working nodes included in the computing clusters of the second priority;
[0184] The processing module 540 is further adapted to:
[0185] For the computing clusters in the computing clusters of the second priority where the number of included working nodes is lower than the number threshold, update their cluster priority to the first priority.
[0186] In summary, according to the resource elastic scaling device provided in this embodiment, it receives the computing tasks issued by the requester, determines the matching target cluster priority according to the classification tags of the computing tasks, and the target cluster with the target cluster priority executes the computing tasks, which can achieve hierarchical scheduling of the computing tasks and improve the execution efficiency of the computing tasks; it detects the cluster status information of the computing cluster with the target cluster priority. When the target cluster priority is the first priority, the cluster status information includes the number of hoarded tasks. According to the cluster status information, resource elastic scaling processing is performed on the computing cluster with the target cluster priority, which can use different cluster status information as the basis for triggering resource elastic scaling processing for computing clusters with different priorities, making the resource elastic scaling strategy of the computing cluster targeted. At the same time, the number of hoarded tasks is introduced as a judgment basis in the resource elastic scaling mechanism of the computing cluster with the first priority, avoiding the problem of resource waste in low-priority computing clusters caused by the method of resource elastic scaling relying on resource utilization rate in the prior art, and can also improve the speed of resource elastic scaling processing to better cope with sudden traffic.
[0187] An embodiment of the present application provides a non-volatile computer storage medium, and the computer storage medium stores at least one executable instruction or computer program, and the executable instruction or computer program can enable a processor to execute the operations corresponding to the resource elastic scaling method in any of the above method embodiments.
[0188] An embodiment of the present application provides a computer program product, and the computer program product includes at least one executable instruction or computer program, and the executable instruction or computer program can enable a processor to execute the operations corresponding to the resource elastic scaling method in any of the above method embodiments.
[0189] Figure 6 The structural schematic diagram of the computing device embodiment of the present application is shown, and the specific implementation of the computing device is not limited in the specific embodiments of the present application.
[0190] As Figure 6 shown, the computing device may include: a processor 602, a communications interface 604, a memory 606, and a communication bus 608.
[0191] Among them: the processor 602, the communications interface 604, and the memory 606 communicate with each other through the communication bus 608. The communications interface 604 is used to communicate with network elements of other devices such as clients or other servers. The processor 602 is used to execute the program 610, and specifically can execute the relevant steps in the resource elastic scaling method embodiment for the computing device described above.
[0192] Specifically, the program 610 may include program code that includes computer operation instructions.
[0193] The processor 602 may be a central processing unit (CPU), or a specific integrated circuit (ASIC) (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the computing device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0194] The memory 606 is used to store the program 610. The memory 606 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.
[0195] The program 610 may specifically be used to cause the processor 602 to execute the resource elastic scaling method in any of the above method embodiments. For the specific implementation of each step in the program 610, reference may be made to the corresponding steps and descriptions in the resource elastic scaling embodiments, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the devices and modules described above may refer to the corresponding process descriptions in the foregoing method embodiments, which will not be repeated here.
[0196] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems may also be used in conjunction with the teachings provided herein. The structure required to construct such a system will be apparent from the above description. In addition, the embodiments of the present application are not directed to any specific programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the descriptions made for specific languages above are to disclose the best mode of the present application.
[0197] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present application may be practiced without these specific details. In some instances, well-known methods, structures, and technologies have not been shown in detail so as not to obscure the understanding of this specification.
[0198] Similarly, it should be understood that, for the purpose of streamlining this application and assisting in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of this application, the various features of the embodiments of this application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed subject matter of this application requires more features than are expressly recited in each claim. Rather, as reflected by the claims, the inventive aspects lie in less than all of the features of the single embodiments disclosed previously. Thus, the claims following the detailed description hereby expressly incorporate the detailed description, where each claim itself serves as a separate embodiment of this application.
[0199] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.
[0200] In addition, those skilled in the art can understand that although some of the embodiments herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of this application and forms different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.
[0201] The various component embodiments of this application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components according to the embodiments of this application. This application can also be implemented as a device or apparatus program (e.g., a computer program and a computer program product) for executing a part or all of the methods described herein. Such a program implementing this application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0202] It should be noted that the above embodiments are illustrative of the present application rather than restrictive of the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.
Claims
1. A method for elastic scaling of resources, comprising: Receiving a computing task issued by a requester; Determining a matching target cluster priority according to the classification label of the computing task; Detecting the cluster status information of the computing cluster with the target cluster priority; wherein, when the target cluster priority is the first priority, the cluster status information includes the number of hoarded tasks; Performing elastic resource scaling processing on the computing cluster with the target cluster priority according to the cluster status information.
2. The method according to claim 1, wherein, When the target cluster priority is the first priority, the cluster status information further includes resource usage metrics; After detecting the cluster status information of the computing cluster with the target cluster priority, the method further includes: Judging the busy or idle state of the computing cluster with the first priority according to the resource usage metrics; If the computing clusters with the first priority are all in a busy state, hoard the computing tasks.
3. The method according to claim 1, wherein When the target cluster priority is the second priority, the cluster status information includes resource usage metrics.
4. The method according to claim 1 or 2, wherein In the case where the target cluster priority is the first priority, the performing elastic resource scaling processing on the computing cluster with the target cluster priority according to the cluster status information further includes: If the number of hoarded tasks reaches a first threshold, performing an expansion process on the computing cluster with the first priority; If the number of hoarded tasks does not reach a second threshold, performing a contraction process on the computing cluster with the first priority.
5. The method according to claim 3, wherein In the case where the target cluster priority is the second priority, the performing elastic resource scaling processing on the computing cluster with the target cluster priority according to the cluster status information further includes: If there is a computing cluster in the computing clusters with the second priority whose resource usage metrics reach a third threshold, performing an expansion process on the computing clusters with the second priority; For the computing clusters in the computing clusters with the second priority whose resource usage metrics do not reach a fourth threshold, performing a contraction process on them.
6. The method according to claim 5, wherein, The third threshold includes: a first sub-threshold and a second sub-threshold, and the second sub-threshold is greater than the first sub-threshold. The performing an expansion process on the computing clusters with the second priority further includes: For the computing clusters in the computing clusters with the second priority whose resource usage metrics are between the first sub-threshold and the second sub-threshold, adding new working nodes to them; If there is a computing cluster in the computing clusters with the second priority whose resource usage metrics reach the second sub-threshold, determining a computing cluster to be configured, and configuring the cluster priority of the computing cluster to be configured as the second priority.
7. The method according to claim 6, wherein, The determining the computing cluster to be configured further includes: Screening the computing cluster to be configured from the computing clusters with the first priority according to the resource usage metrics; or, creating at least one computing cluster as the computing cluster to be configured.
8. The method according to claim 6, wherein The configuring the cluster priority of the computing cluster to be configured as the second priority further includes: Updating the cluster label corresponding to the computing cluster to be configured to the cluster label corresponding to the computing clusters with the second priority; The determining a matching target cluster priority according to the classification label of the computing task further includes: Determine the cluster label that matches the classification label of the computing task; Determine the target cluster priority according to the matching cluster label.
9. The method according to claim 4, wherein The step of performing capacity expansion processing on the computing cluster with the first priority further includes: Add working nodes to one or more computing clusters in the computing cluster with the first priority.
10. The method according to claim 4 or 5, wherein The capacity reduction processing is specifically: take offline the working nodes to be taken offline in the computing cluster to be capacity-reduced; the method further includes: Screen the working nodes to be taken offline from the working nodes of the computing cluster to be capacity-reduced according to the node status information and / or the node creation time.
11. According to the method according to any one of claims 1-10, wherein, The method further includes: Detect the number of working nodes included in the computing cluster with the second priority; For the computing clusters in the computing cluster with the second priority whose number of included working nodes is lower than the number threshold, update their cluster priority to the first priority.
12. A resource elastic scaling device, comprising: A receiving module, adapted to receive a computing task issued by a requester; A scheduling module, adapted to determine a matching target cluster priority according to the classification label of the computing task; A detection module, adapted to detect the cluster status information of the computing cluster with the target cluster priority; wherein, when the target cluster priority is the first priority, the cluster status information includes the number of hoarded tasks; A processing module, adapted to perform resource elastic scaling processing on the computing cluster with the target cluster priority according to the cluster status information.
13. A computing device, comprising: A processor, a memory, a communication interface, and a communication bus, the processor, the memory, and the communication interface complete mutual communication through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the resource elastic scaling method according to any one of claims 1-11.
14. A computer storage medium, in which at least one executable instruction is stored, and the executable instruction causes a processor to perform the operations corresponding to the resource elastic scaling method according to any one of claims 1-11.
15. A computer program product, including at least one executable instruction, and the executable instruction causes a processor to perform the operations corresponding to the resource elastic scaling method according to any one of claims 1-11.