Service cluster management method and device and storage medium

By pooling the computing node resources of the service cluster into the resource pool, and dynamically allocating the computing nodes based on job requests, the resource waste caused by the heterogeneity of the computing node is solved, and efficient resource utilization and load balancing are achieved.

CN120416255APending Publication Date: 2025-08-01ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510573018.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

As the number of computing nodes increases, the heterogeneity between computing nodes leads to the problem of waste of resources and low utilization. In the existing technology, independent business clusters cannot be flexibly shared, resulting in resource segmentation and waste.

Method used

The computing node resource pool of multiple service clusters is used to filter the computing nodes based on the resource requirements of job requests to form a dynamic service cluster, realizing the decoupling and flexible allocation of resources.

Benefits of technology

It improves resource utilization, realizes dynamic allocation of isomorphic computing nodes, and improves efficient resource utilization and load balancing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416255A_ABST
    Figure CN120416255A_ABST
Patent Text Reader

Abstract

The invention provides a business cluster management method and device and a storage medium, and the method comprises the steps: pooling the resources of computing nodes contained in a plurality of business clusters to a business resource pool; the method comprises the steps that a job request of a user is received, the job request carries job parameters, and the job parameters are used for indicating resource requirements of the job request; based on a resource demand indicated by the operation parameter, screening out computing nodes meeting the resource demand from the service resource pool to form a dynamic service cluster corresponding to the user; and submitting the job request to the dynamic service cluster to enable the dynamic service cluster to complete corresponding jobs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of cloud technology, and in particular, to a method, device, and storage medium for managing a service cluster. Background Art

[0002] With the rapid growth of data scale, the number of computing nodes managed by a distributed data processing platform is also increasing rapidly. With the continuous increase in the number of computing nodes, there are significant differences in the software and hardware versions, the computer rooms where they are distributed, and the business domains among the computing nodes, which makes the computing nodes inherently heterogeneous. When users use a service cluster, they usually expect the computing nodes in the service cluster to be homogeneous and able to be used exclusively. However, this requirement results in resources being divided into multiple mutually isolated service clusters.

[0003] With the continuous development of business, the amount of resources managed by the platform continues to increase, and the number of independent service clusters is also increasing. However, the disadvantages of this way of managing resources with independent service clusters are becoming increasingly prominent. For example, a user's job is strictly bound to a service cluster, and only that user can use that service cluster for the job. Even if the service cluster is idle, it cannot be transferred to other users for temporary use; moreover, most of the time, each service cluster is not fully utilized, resulting in a great waste of resources. Summary of the Invention

[0004] In view of this, one or more embodiments of this specification provide the following technical solutions:

[0005] According to a first aspect of one or more embodiments of this specification, a method for managing a service cluster is proposed, in which the resources of the computing nodes included in multiple service clusters are pooled into a service resource pool; the method includes:

[0006] Receiving a job request from a user, the job request carrying job parameters for indicating the resource requirements of the job request;

[0007] Based on the resource requirements indicated by the job parameters, screening out computing nodes that meet the resource requirements from the service resource pool to form a dynamic service cluster corresponding to the user;

[0008] Submitting the job request to the dynamic service cluster so that the dynamic service cluster completes the corresponding job.

[0009] According to a second aspect of one or more embodiments of this specification, an electronic device is proposed, including: a processor; a memory for storing processor-executable instructions; wherein, the processor realizes the steps of the method as described in the first aspect above by running the executable instructions.

[0010] According to a third aspect of one or more embodiments of this specification, a computer-readable storage medium is provided, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method described in the first aspect above are implemented.

[0011] According to a fourth aspect of one or more embodiments of this specification, a computer program product is provided, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method described in the first aspect above are implemented.

[0012] In the above technical solutions, this specification pools the resources of the computing nodes included in multiple service clusters into a service resource pool, and based on the resource requirements of a job request, filters out the computing nodes that meet the resource requirements from the service resource pool to form a dynamic service cluster corresponding to the user, realizing the delivery of resources to the user in the form of resource quotas, and further realizing the decoupling of the job and the service cluster, enabling a service cluster to complete the jobs of multiple different users, thereby being able to utilize the resources in the service cluster more efficiently and improving the resource utilization rate. Description of the Drawings

[0013] Figure 1 is a flowchart of a method for managing a service cluster provided by an exemplary embodiment.

[0014] Figure 2 is a flowchart of a method for managing a service cluster based on a matching expression and pooling tags provided by an exemplary embodiment.

[0015] Figure 3 is a schematic structural diagram of a distributed data processing platform provided by an exemplary embodiment.

[0016] Figure 4 is a flowchart of scheduling a service cluster based on a matching expression provided by an exemplary embodiment.

[0017] Figure 5 is a schematic diagram of a service resource pool resource scheduling service provided by an exemplary embodiment.

[0018] Figure 6 is a schematic structural diagram of a device provided by an exemplary embodiment.

[0019] Figure 7 is a block diagram of a device for managing a service cluster provided by an exemplary embodiment. Detailed Embodiments

[0020] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0021] With the rapid growth of the data scale, the number of computing nodes managed by the distributed data processing platform is also increasing rapidly. As the number of computing nodes continues to increase, there are significant differences in the software and hardware versions, the computer rooms where they are distributed, and the business domains among the computing nodes, which makes the computing nodes inherently heterogeneous. When users use the business cluster, they usually expect the computing nodes in the business cluster to be homogeneous and can be used exclusively. However, this requirement results in resources being split into multiple mutually isolated business clusters.

[0022] With the continuous development of the business, the amount of resources managed by the platform continues to increase, and the number of independent business clusters is also increasing. However, the disadvantages of this way of managing resources with independent business clusters are becoming increasingly prominent. For example, the user's job is strictly bound to the business cluster, and only this user can use this business cluster for the job. Even if this business cluster is idle, it cannot be transferred to other users for temporary use; moreover, most of the time, each business cluster is not fully utilized, which greatly causes waste of resources.

[0023] Based on this, this specification provides a method for managing a business cluster, which can deliver resources to users in the form of resource quotas, realizes the decoupling of jobs and business clusters, can utilize the resources in the business cluster more efficiently, and improves the resource utilization rate.

[0024] In implementation, first, receive the user's job request, which carries job parameters used to indicate the resource requirements of the job request; then, based on the resource requirements indicated by the job parameters, screen out the computing nodes that meet the resource requirements from the business resource pool to form a dynamic business cluster corresponding to the user; finally, submit the job request to the dynamic business cluster so that the dynamic business cluster can complete the corresponding job.

[0025] In the above technical solution, by pooling the resources of the computing nodes included in multiple service clusters into a service resource pool, and screening out the computing nodes that meet the resource requirements from the service resource pool based on the resource requirements of the job request to form a dynamic service cluster corresponding to the user, the resources are delivered to the user in the form of resource quotas, thereby realizing the decoupling of the job and the service cluster, enabling a service cluster to complete the jobs of multiple different users, and thus being able to utilize the resources in the service cluster more efficiently and improving the resource utilization rate.

[0026] Next, a detailed introduction to the service cluster management method provided in this specification will be given:

[0027] Figure 1 FIG. is a flowchart of a service cluster management method provided by an exemplary embodiment. The execution subject of this method can be a distributed data processing platform or any device, and this specification does not limit the execution subject. It should be noted that in this embodiment, the resources of the computing nodes included in multiple service clusters are pooled into a service resource pool. As Figure 1 shown, the method includes:

[0028] Step S101, receive a job request from a user. The job request carries job parameters, and the job parameters are used to indicate the resource requirements of the job request.

[0029] Among them, the job request is a request for the service cluster to complete any job task. The job task can be a data processing task, a model training task, etc., and this specification does not limit the job task. In this specification, the job parameters are used to indicate the resource requirements of the job request, that is, how much resources are needed to complete the corresponding job. Exemplarily, the job parameters include the number of nodes, the memory size, the disk size, etc. This specification does not limit this.

[0030] Step S102, based on the resource requirements indicated by the job parameters, screen out the computing nodes that meet the resource requirements from the service resource pool to form a dynamic service cluster corresponding to the user.

[0031] Among them, the computing nodes that meet the resource requirements refer to: the resources of the computing nodes can meet the resource requirements. Screening out the computing nodes that meet the resource requirements from the service resource pool means: screening out multiple computing nodes from the service resource pool so that the resources of the multiple computing nodes can meet the resource requirements.

[0032] In this specification, the business cluster corresponding to a user is called a dynamic business cluster because the user is no longer bound to the business cluster. When the user submits a job request next time, the business resource pool will reassign a business cluster to the user. Therefore, the business cluster corresponding to the user is not fixed but dynamically changing. Hence, the business cluster assigned to the user is called a dynamic business cluster.

[0033] In one embodiment, the computing nodes within the same business cluster are all homogeneous computing nodes, and the computing nodes in different business clusters are homogeneous computing nodes or heterogeneous computing nodes. Therefore, the business resource pool contains both homogeneous and heterogeneous computing nodes. Since users usually expect the computing nodes in the business cluster to be homogeneous when using the business cluster, when screening out the computing nodes that meet the resource requirements from the business resource pool based on the resource requirements indicated by the job parameters, the screened nodes are homogeneous. That is to say, this dynamic business cluster is composed of homogeneous computing nodes in the same business cluster or different business clusters.

[0034] In one embodiment, the dynamic business cluster is composed of homogeneous computing nodes in different business clusters. Exemplarily, when none of the business clusters can meet the resource requirements indicated by the job parameters, homogeneous computing nodes in different business clusters can be selected to form the dynamic business cluster corresponding to the user. For example, the job parameters indicate that 100 computing nodes are needed. Business cluster 1 has 80 idle computing nodes, business cluster 2 has 30 idle computing nodes, and business cluster 3 has 70 idle computing nodes. Among them, the computing nodes in business cluster 1 and business cluster 2 are homogeneous computing nodes, and the computing nodes in business cluster 1 and business cluster 3 are heterogeneous computing nodes. Therefore, 100 idle computing nodes can be screened out from business cluster 1 and business cluster 2 and provided to the user. What the user perceives is that the platform provides a business cluster with 100 computing nodes.

[0035] In another embodiment, the dynamic business cluster is composed of homogeneous computing nodes in the same business cluster. Since the computing nodes within the same business cluster are all homogeneous computing nodes, when determining the computing nodes that meet the resource requirements in a business cluster, it can be ensured that the determined computing nodes are all homogeneous. Exemplarily, the job parameters indicate that 100 computing nodes are needed. Business cluster 1 has 500 idle computing nodes, business cluster 2 has 300 idle computing nodes, and business cluster 3 has 150 idle computing nodes. Therefore, 100 idle computing nodes can be determined from the 500 idle computing nodes in business cluster 1 and provided to the user. What the user perceives is that the platform provides a business cluster with 100 computing nodes.

[0036] To determine computing nodes that meet resource requirements in a service cluster, it is necessary to ensure that there are sufficient computing nodes in the service cluster. Therefore, before screening for computing nodes that meet resource requirements, it can be determined first whether the available resources of the service cluster meet the resource requirements. Optionally, the method further includes: after receiving a job request, traversing each service cluster in the service resource pool; if the available resources of any service cluster meet the resource requirements, then perform the step of screening out computing nodes that meet the resource requirements from the service resource pool based on the resource requirements indicated by the job parameters to form a dynamic service cluster corresponding to the user.

[0037] Exemplarily, after receiving a job request, first traverse each service cluster in the service resource pool; if the available resources of any service cluster meet the resource requirements, then based on the resource requirements indicated by the job parameters, screen out computing nodes that meet the resource requirements from one service cluster to form a dynamic service cluster corresponding to the user; if the available resources of no service cluster meet the resource requirements, then screen out homogeneous computing nodes that meet the resource requirements from different service clusters to form a dynamic service cluster corresponding to the user.

[0038] Step S103: Submit the job request to the dynamic service cluster so that the dynamic service cluster completes the corresponding job.

[0039] After selecting a dynamic service cluster through the service resource pool, the existing cluster submission job service can be used to submit a job on the dynamic service cluster; or it can be configured to select to submit a job on the service cluster, or to select to submit a job on the service resource pool, which is not limited in this specification.

[0040] In the above technical solution, this specification pools the resources of the computing nodes included in multiple service clusters into a service resource pool, and screens out computing nodes that meet the resource requirements from the service resource pool based on the resource requirements of the job request to form a dynamic service cluster corresponding to the user, realizing the delivery of resources to the user in the form of resource quotas, and further realizing the decoupling of jobs from service clusters, enabling a service cluster to complete jobs of multiple different users, thereby being able to utilize the resources in the service cluster more efficiently and improving the resource utilization rate.

[0041] Next, taking the dynamic service cluster being composed of computing nodes in the same service cluster as an example, the method provided in this specification is further described:

[0042] Figure 2 is a flowchart of a method for managing a service cluster based on a matching expression and a pooling label provided by an exemplary embodiment. The execution subject of this method can be a distributed data processing platform or any device, which is not limited in this specification. AsFigure 2 As shown in Figure 2 , the method includes:

[0043] Step S201: Receive a job request from a user. The job request carries job parameters, and the job parameters are used to indicate the resource requirements of the job request.

[0044] The above step S201 is the same as the above step S101, and reference can be made to the above step S101, which will not be elaborated here one by one.

[0045] Step S202: Based on the resource requirements indicated by the job parameters and the matching expression defined in the service resource pool, determine a target service cluster that meets the matching expression from the service resource pool. The matching expression defines the matching rule between the resource requirements and the available resources of the service cluster.

[0046] In this specification, the service resource pool defines a matching expression, and the matching expression defines the matching rule between the resource requirements and the available resources of the service cluster. Therefore, based on the resource requirements indicated by the job parameters and the matching expression, a target service cluster whose available resources meet the resource requirements can be screened out for the user.

[0047] Exemplarily, the matching expression includes key, value, and operator. Among them, key defines the name of the matching expression, value defines the value of the matching expression, and operator defines the operator used in the matching expression, such as in, notin, exist, not exist, lt, gt, etc.

[0048] In an illustrated embodiment, the matching expression includes a hard expression, where the hard expression is used to define the matching rule that must be satisfied. In one embodiment, based on the resource requirements indicated by the job parameters and the matching expression defined in the service resource pool, determining a target service cluster that meets the matching expression from the service resource pool includes: determining the service clusters in the service resource pool that meet the hard expression based on the resource requirements indicated by the job parameters and the hard expression defined in the service resource pool; if there is only one service cluster that meets the hard expression, then determine this service cluster as the target service cluster; if there are multiple service clusters that meet the hard expression, then randomly select one service cluster from the multiple service clusters to determine as the target service cluster, or select the service cluster with the lowest load from the multiple service clusters to determine as the target service cluster, or select the service cluster whose location in the computer room is closest to the data storage device from the multiple service clusters to determine as the target service cluster, or select the service cluster with the fastest computing speed from the multiple service clusters to determine as the target service cluster.

[0049] It should be noted that this specification only gives an exemplary description of the selection of the target business cluster. Of course, when there are multiple business clusters that meet the hard expression, one business cluster can also be selected from these multiple business clusters according to other criteria and determined as the target business cluster. This specification does not limit the selection criteria.

[0050] In another illustrated embodiment, the matching expression includes a soft expression. The soft expression is used to define a matching rule that does not have to be satisfied. Among them, based on the resource requirements indicated by the job parameters and the matching expression defined by the business resource pool, determining the target business cluster that meets the matching expression from the business resource pool includes: based on the resource requirements indicated by the job parameters and the soft expression defined by the business resource pool, determining the matching scores of each business cluster in the business resource pool, and determining the business cluster with the highest matching score as the target business cluster.

[0051] The degree of matching with the soft expression is different, and the corresponding matching scores are different. Among them, the higher the degree of matching with the soft expression, the higher the corresponding matching score; the lower the degree of matching with the soft expression, the lower the corresponding score. Exemplarily, the soft expression is used to represent location matching. If the data required by the job is stored in Shanghai, then the business cluster with the computer room in Shanghai has a higher matching score, and the business cluster with the computer room farther away from Shanghai has a lower matching score.

[0052] In one embodiment, when the business resource pool defines multiple soft expressions, the highest matching scores corresponding to each soft expression can be the same or different. Exemplarily, more important soft expressions can correspond to higher highest matching scores, and less important soft expressions can correspond to lower highest matching scores.

[0053] In another illustrated embodiment, the matching expression includes a hard expression and a soft expression. The hard expression is used to define a matching rule that must be satisfied, and the soft expression is used to define a matching rule that does not have to be satisfied; among them, based on the resource requirements indicated by the job parameters and the matching expression defined by the business resource pool, determining the target business cluster that meets the matching expression from the business resource pool includes: based on the resource requirements indicated by the job parameters and the hard expression defined by the business resource pool, determining the business clusters that meet the hard expression in the business resource pool as candidate business clusters; based on the resource requirements indicated by the job parameters and the soft expression defined by the business resource pool, determining the matching scores of each candidate business cluster, and determining the candidate business cluster with the highest matching score as the target business cluster.

[0054] Among them, the hard expression can be one expression or multiple expressions. This specification does not limit the number of hard expressions. Among them, the soft expression can be one expression or multiple expressions, and this specification does not limit the number of soft expressions.

[0055] In one embodiment, the matching expression defined by the service resource pool is modifiable. Relevant personnel can adjust the allocation of the service cluster by modifying the matching expression defined by the service resource pool and the matching score corresponding to the matching expression. Optionally, the soft expression defines the calculation of the matching score according to the load conditions of the candidate service clusters. That is to say, relevant personnel can define the soft matching expression of the service resource pool, so that the service cluster with a lower load has a higher corresponding matching score, so as to preferentially allocate jobs to the service cluster with a lower load, thereby achieving load balancing.

[0056] In an illustrated implementation manner, the service resource pool is used to manage each service cluster based on the pooling tags of each service cluster, and the pooling tag is used to describe the cluster information of the service cluster. Among them, based on the resource requirements indicated by the job parameters, computing nodes that meet the resource requirements are screened out from the service resource pool to form a dynamic service cluster corresponding to the user, including: determining, based on the resource requirements indicated by the job parameters and the pooling tags of each cluster, a target service cluster in the service resource pool whose pooling tag matches the resource requirements, and determining, from the target service cluster, computing nodes that meet the resource requirements to form a dynamic service cluster corresponding to the user.

[0057] Among them, the pooling tag includes a default tag and a custom tag. The default tag is used to represent the basic attributes of the service cluster, and the default tag is a non-modifiable tag. Exemplarily, the default tag includes the cluster name tag, cluster specification tag, cluster environment tag, computer room tag where the cluster is located, etc. of the service cluster.

[0058] The custom tag is a modifiable tag. Optionally, the custom tag is set by relevant personnel according to the actual situation or application requirements. Exemplarily, the custom tag includes the resource quota tag of the service cluster, and the resource quota tag is used to represent the total resource quota of the service cluster. For example, a service cluster includes a total of 1000 computing nodes, but relevant personnel only want to pool 500 computing nodes in this service cluster according to the actual situation or application requirements. Then, relevant personnel can modify the resource quota tag of the cluster in the custom tag from 1000 computing nodes to 500 computing nodes.

[0059] Exemplarily, the custom tag includes an online tag. When the custom tag of any business cluster includes the online tag, this business cluster can be used as a potential target business cluster. Subsequently, it is possible for the business resource pool to assign jobs to this business cluster. Exemplarily, the custom tag includes an offline tag. When the custom tag of any business cluster includes the offline tag, the business resource pool will not match this business cluster based on the resource requirements indicated by the job parameters and the matching expression defined by the business resource pool, and will not assign jobs to this business cluster. It should be noted that the custom tag can only include either the online tag or the offline tag, and cannot include both the online tag and the offline tag at the same time.

[0060] Step S203: Determine the computing nodes that meet the resource requirements from the target business cluster to form a dynamic business cluster corresponding to the user.

[0061] It should be noted that in the above step S203, determining the computing nodes that meet the resource requirements from the target business cluster to form a dynamic business cluster corresponding to the user does not actually build a new business cluster with some computing nodes in the target business cluster. Instead, it assigns the computing nodes that meet the resource requirements in the target business cluster to the user. From the user's perspective, the platform assigns a business cluster.

[0062] Exemplarily, the resource requirement indicates that 100 computing nodes are needed, and the target business cluster contains 1000 computing nodes. Just submit the job to this target business cluster, and this target business cluster can assign 100 computing nodes to complete this job. From the user's perspective, the platform assigns a business cluster with 100 computing nodes to complete this job.

[0063] Step S204: Submit the job request to the dynamic business cluster so that the dynamic business cluster can complete the corresponding job.

[0064] According to the description in the above step S203, submitting the job request to the dynamic business cluster actually means submitting the job request to the target business cluster where the dynamic business cluster is located.

[0065] According to the description in the above step S202, the pooling tag includes a default tag and a custom tag. The default tag is a tag that cannot be modified, and the custom tag is a tag that can be modified. The custom tag is set by relevant personnel according to the actual situation or application requirements. Therefore, in this specification, the following steps S205 and S206 are used as examples to exemplarily illustrate the usage and beneficial effects brought by the custom tag:

[0066] Step S205: In response to a first editing operation on the custom label of any service cluster, add a take-offline label to the custom label. After the any service cluster becomes idle, adjust the resource quota of the any service cluster. In response to a second editing operation on the custom label of the any service cluster, modify the take-offline label in the custom label to an online label, so as to enable the any service cluster to achieve seamless scaling for users.

[0067] Wherein, the first editing operation and the second editing operation can be any editing operation, and this specification does not limit the first editing operation and the second editing operation.

[0068] With the use of service clusters, there may be problems such as some service clusters being too large or some service clusters being too small. At this time, it is necessary to split or merge the service clusters. However, when splitting or merging service clusters, it is not desired to affect the normal use of users. Preferably, it is completed without the users' awareness. For this purpose, this specification provides the method described in step S205 above, that is, by modifying the custom label of the service cluster.

[0069] Exemplarily, with the use of service clusters, relevant personnel find that the specification of service cluster 1 is too large, with 1000 computing nodes, and it is often not fully utilized, resulting in waste of resources. Therefore, the relevant personnel want to split service cluster 1 into two service clusters, each containing 500 computing nodes. At this time, the relevant personnel can perform a first editing operation on the custom label of service cluster 1 to add a take-offline label to the custom label of service cluster 1. Although the relevant personnel add a take-offline label to the custom label of service cluster 1, it does not affect service cluster 1 from executing corresponding jobs. And because the custom label of service cluster 1 includes a take-offline label, the service resource pool will not continue to assign jobs to service cluster 1. Therefore, after service cluster 1 finishes executing the existing jobs, service cluster 1 will become completely idle. At this time, the relevant personnel can split service cluster 1 into service cluster 2 and service cluster 3, and add online labels to the custom labels of service cluster 2 and service cluster 3, so that the service resource pool can assign jobs to service cluster 2 and service cluster 3.

[0070] Step S206: In response to a second editing operation on the custom label of a newly added service cluster, add an online label to the custom label. In response to a first editing operation on the custom label of an original service cluster, add a take-offline label to the custom label, so as to achieve seamless job migration for users.

[0071] Considering reasons such as the possible differences in computing resource costs in different regions, it is necessary to migrate jobs from the business cluster in region A to the business cluster in another region B. At this time, relevant personnel can add an online label to the custom label of the business cluster in region B (i.e., the newly added business cluster), and add an offline label to the custom label of the business cluster in region A (i.e., the original business cluster). Therefore, the business resource pool will allocate new jobs to the business cluster in region B instead of the business cluster in region A. In this way, the user can migrate the job from region A to region B without being aware of it.

[0072] It should be noted that the above steps S205 to S206 are optional execution steps, and can be selected to execute or not execute according to the actual situation or application requirements. This specification does not limit this.

[0073] In the above technical solution, this specification realizes the delivery of resources to users in the form of resource quotas by pooling the resources of the computing nodes contained in multiple business clusters into a business resource pool, and screening out the computing nodes that meet the resource requirements from the business resource pool based on the resource requirements of the job request, thereby realizing the decoupling of the job and the business cluster, enabling a business cluster to complete the jobs of multiple different users, and thus being able to utilize the resources in the business cluster more efficiently and improving the resource utilization rate.

[0074] Moreover, the business resource pool manages the business clusters through the pooled labels of the business clusters. By editing the custom labels in the pooled labels, relevant personnel can not only achieve the load balancing of the business clusters, but also achieve the cluster scaling and job migration without the user being aware of it.

[0075] Next, this specification is based on Figure 1 and Figure 2 The embodiments shown are used to exemplarily illustrate the method provided in this specification:

[0076] As Figure 3 shown, the implementation architecture of the management method of this business cluster can include three modules from bottom to top, namely business cluster resources, unified pooled resource management, and business cluster resource scheduling. It should be noted that this specification only takes the implementation architecture including three modules as an example to exemplarily illustrate the implementation architecture. In another embodiment, the implementation architecture can include fewer or more modules, and this specification does not limit this.

[0077] Among them, the business cluster resources include multiple business clusters. Each business cluster is a homogeneous business cluster, and the homogeneous business clusters can be homogeneous or heterogeneous. Optionally, the business cluster can be built through Kubernetes (k8s, a distributed application orchestration platform based on containerization technology) or Apache Mesos (a cluster management system).

[0078] The business resource pool associates jobs and business clusters through the unified pooled resource description of the cluster (that is, the business resource pool associates jobs and business clusters through the pooled labels of the business clusters), as Figure 4 shown. The resource quantity requested will be configured in the job parameters, such as the number of nodes, memory size, disk size, etc. The pooled resource description of the business cluster includes the quotas of all types of resources (the quota includes the total quota of resources and the quota of used resources). At the same time, the pooled resource description of the business cluster also includes pooled labels, which include default labels and custom labels. The default label is a label that each business cluster has. The default label can include the cluster name label, cluster rule label, cluster environment label, and data center label where the cluster is located, etc. The custom label can set different labels and corresponding values for each business cluster according to requirements. The resource description of the business resource pool mainly includes resource quotas and matching expressions. The resource quota defines the quotas of different resources in the business resource pool, such as the number of nodes, CPU size, memory size, and disk size, etc. The matching expression includes a hard expression and a soft expression. Among them, the matching expression includes key, value, and operator. Among them, key defines the name of the matching expression, value defines the value of the matching expression, and operator defines the operator used in the matching expression. For example, in, not in, exist, not exist, lt, gt, etc.

[0079] The job request submitted by the user will be submitted to the unified pooled resource management, and the unified pooled resource management will send the job request to the business cluster resource scheduling. As Figure 5As shown, the verification unit for business cluster resource scheduling traverses each business cluster to verify whether there is any available resource in a business cluster that meets the resource requirements of the job request; if there is any available resource in a business cluster that meets the resource requirements of the job request, the routing unit is called. In the routing unit, first, a hard expression matching is performed, that is, the pooled label of the business cluster is matched with the hard expression defined in the business resource pool. If the pooled label of the business cluster can meet all the hard expressions, the business cluster is added to the candidate business cluster list; then, a soft expression matching is performed, that is, the candidate business clusters in the candidate business cluster list are matched with the soft expression defined in the business resource pool, and the weights or matching scores corresponding to all the matched soft expressions are summed to obtain the matching score of each candidate business cluster, and the candidate business cluster with the highest matching score is determined as the target business cluster. The business cluster resource scheduling returns the target business cluster to the unified pooled resource management. The unified pooled resource management calls the existing cluster submission job service to submit the job to the target business cluster, and the target business cluster will allocate the resources required by the job to execute the job to form a dynamic business cluster corresponding to the user. It should be noted that this specification only takes the call to the existing cluster submission job service as an example to illustrate the submission of the job exemplarily. In another embodiment, relevant personnel can configure the submission method of the job and choose to use business cluster submission or choose to use business resource pool submission, and this specification does not limit this.

[0080] Among them, the soft expression can be configured according to the requirements of the job. For example, if it is desired that the business cluster can achieve load balancing during the job, relevant personnel can define in the soft expression to calculate the matching score according to the load situation of the candidate business cluster. The lower the load of the business cluster, the higher the matching score of the business cluster; or, a higher weight or higher matching score is given to the soft expression that calculates the matching score according to the load situation of the candidate business cluster.

[0081] Since in the business cluster management method provided in this specification, the business resource pool manages the business cluster through the pooled label of the business cluster, and the pooled label contains a custom label, relevant personnel can achieve cluster scaling and job migration without user perception through the configuration of the custom label.

[0082] For example, with the use of business clusters, relevant personnel found that the specifications of business cluster 1 and business cluster 2 were too small to undertake certain jobs. Among them, business cluster 1 contains 100 computing nodes, business cluster 2 contains 120 computing nodes, and the computing nodes included in business cluster 1 and business cluster 2 are homogeneous computing nodes. Therefore, relevant personnel want to merge business cluster 1 and business cluster 2 into a new business cluster. At this time, relevant personnel can perform a first editing operation on the custom tags of business cluster 1 and business cluster 2 to add a take-off line tag to the custom tags of business cluster 1 and business cluster 2. Although relevant personnel have added a take-off line tag to the custom tags of business cluster 1 and business cluster 2, it will not affect the execution of corresponding jobs by business cluster 1 and business cluster 2. And because the custom tags of business cluster 1 and business cluster 2 include a take-off line tag, the business resource pool will not continue to assign jobs to business cluster 1 and business cluster 2. Therefore, after business cluster 1 and business cluster 2 have completed the existing jobs, business cluster 1 and business cluster 2 will be completely idle. At this time, relevant personnel can merge business cluster 1 and business cluster 2 into a new business cluster 3 and perform a second editing operation on business cluster 3 to add an on-line tag to the custom tag of business cluster 3. Subsequently, the business resource pool can assign jobs to business cluster 3. This process is imperceptible to users. Therefore, this method realizes the scaling of business clusters without the user's perception.

[0083] For another example, in some cases, it is necessary to migrate jobs from a business cluster in region A to a business cluster in another region B. At this time, relevant personnel can perform a second editing operation on the business cluster in region B to add an on-line tag to the custom tag of the business cluster in region B, and perform a first editing operation on the business cluster in region A to add a take-off line tag to the custom tag of the business cluster in region A. In this way, the business resource pool will assign new jobs to the business cluster in region B and no longer assign them to the business cluster in region A. In this way, it realizes the migration of jobs from region A to region B without the user's perception.

[0084] In the above technical solutions, this specification realizes the delivery of resources to users in the form of resource quotas by pooling the resources of the computing nodes contained in multiple business clusters into a business resource pool and screening out the computing nodes that meet the resource requirements from the business resource pool based on the resource requirements of job requests. Furthermore, it realizes the decoupling of jobs and business clusters, enabling a business cluster to complete the jobs of multiple different users, thereby being able to utilize the resources in the business cluster more efficiently and improving the utilization rate of resources.

[0085] Moreover, the business resource pool manages business clusters through the pooling tags of business clusters. Relevant personnel can not only achieve load balancing of business clusters but also achieve cluster scaling and job migration without user perception by editing the custom tags in the pooling tags.

[0086] Figure 6 is a schematic structural diagram of a device provided by an exemplary embodiment. Please refer to Figure 6 , at the hardware level, the device includes a processor 602, an internal bus 604, a network interface 606, a memory 608, and a non-volatile memory 610. Of course, it may also include other hardware required for other functions. One or more embodiments of this specification can be implemented in a software manner. For example, the processor 602 reads the corresponding computer program from the non-volatile memory 610 into the memory 608 and then runs it. Of course, in addition to the software implementation manner, one or more embodiments of this specification do not exclude other implementation manners, such as logical devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or a logical device.

[0087] Please refer to Figure 7 , the management device of the business cluster can be applied to a device such as Figure 6 shown to implement the technical solution of this specification. Among them, the management device of the business cluster may include:

[0088] A receiving unit 701, configured to receive a job request from a user, where the job request carries job parameters, and the job parameters are used to indicate the resource requirements of the job request;

[0089] A screening unit 702, configured to screen out computing nodes that meet the resource requirements from the business resource pool based on the resource requirements indicated by the job parameters to form a dynamic business cluster corresponding to the user;

[0090] A submitting unit 703, configured to submit the job request to the dynamic business cluster so that the dynamic business cluster completes the corresponding job.

[0091] In the above technical solution, this specification pools the resources of the computing nodes included in multiple business clusters into a business resource pool, and screens out computing nodes that meet the resource requirements from the business resource pool based on the resource requirements of the job request to form a dynamic business cluster corresponding to the user, thereby realizing the delivery of resources to the user in the form of resource quotas, and further realizing the decoupling of jobs and business clusters, enabling a business cluster to complete jobs of multiple different users, and thus being able to more efficiently utilize the resources in the business cluster and improving the resource utilization rate.

[0092] Optionally, the filtering unit 702 is configured to determine, based on the resource requirements indicated by the job parameters and the matching expression defined in the service resource pool, a target service cluster that meets the matching expression from the service resource pool, where the matching expression defines the matching rule between the resource requirements and the available resources of the service cluster; and determine computing nodes that meet the resource requirements from the target service cluster to form a dynamic service cluster corresponding to the user.

[0093] Optionally, the matching expression includes a hard expression and a soft expression. The hard expression is used to define the matching rules that must be satisfied, and the soft expression is used to define the matching rules that do not have to be satisfied.

[0094] The filtering unit 702 is configured to determine, based on the resource requirements indicated by the job parameters and the hard expression defined in the service resource pool, the service clusters in the service resource pool that meet the hard expression as candidate service clusters.

[0095] The filtering unit 702 is configured to determine the matching scores of each candidate service cluster based on the resource requirements indicated by the job parameters and the soft expression defined in the service resource pool, and determine the candidate service cluster with the highest matching score as the target service cluster.

[0096] Optionally, the soft expression defines the calculation of the matching score according to the load conditions of the candidate service clusters.

[0097] Optionally, after receiving the job request, the filtering unit 702 traverses each service cluster in the service resource pool. If there is any service cluster whose available resources meet the resource requirements, it performs the step of screening out computing nodes that meet the resource requirements from the service resource pool based on the resource requirements indicated by the job parameters to form a dynamic service cluster corresponding to the user.

[0098] Optionally, the service resource pool is used to manage each service cluster based on the pooling tags of each service cluster, and the pooling tags are used to describe the cluster information of the service cluster.

[0099] The filtering unit 702 is configured to determine, based on the resource requirements indicated by the job parameters and the pooling tags of each cluster, a target service cluster from the service resource pool whose pooling tags match the resource requirements, and determine computing nodes that meet the resource requirements from the target service cluster to form a dynamic service cluster corresponding to the user.

[0100] Optionally, the device further includes an editing unit, and the editing unit is configured to perform at least one of the following:

[0101] In response to a first editing operation on the custom tags of any service cluster, a offline tag is added to the custom tags. After the any service cluster becomes idle, the resource quota of the any service cluster is adjusted. In response to a second editing operation on the custom tags of the any cluster, the offline tag in the custom tags is modified to an online tag, so that the any service cluster can achieve seamless scaling for users.

[0102] In response to a second editing operation on the custom tags of a newly added service cluster, an online tag is added to the custom tags. In response to a first editing operation on the custom tags of an original service cluster, an offline tag is added to the custom tags to achieve seamless job migration for users.

[0103] Optionally, the computing nodes within the same service cluster are all homogeneous computing nodes, and the computing nodes in different service clusters are homogeneous computing nodes or heterogeneous computing nodes. The dynamic service cluster is composed of homogeneous computing nodes in the same service cluster or different service clusters.

[0104] Based on the same concept as the above method, this specification also provides an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor runs the executable instructions to implement the steps of the method as described in any of the above embodiments.

[0105] Based on the same concept as the above method, this specification also provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method as described in any of the above embodiments are implemented.

[0106] Based on the same concept as the above method, this specification also provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method as described in any of the above embodiments are implemented.

Claims

1. A management method for business clusters, where the resources of computing nodes in multiple business clusters are pooled into a business resource pool; the method includes: Receiving a job request from a user, the job request carrying job parameters, where the job parameters are used to indicate the resource requirements of the job request; Based on the resource requirements indicated by the job parameters, screening out computing nodes that meet the resource requirements from the business resource pool to form a dynamic business cluster corresponding to the user; Submitting the job request to the dynamic business cluster so that the dynamic business cluster completes the corresponding job.

2. The method according to claim 1, where the screening out computing nodes that meet the resource requirements from the business resource pool based on the resource requirements indicated by the job parameters to form a dynamic business cluster corresponding to the user includes: Based on the resource requirements indicated by the job parameters and a matching expression defined by the business resource pool, determining a target business cluster that meets the matching expression from the business resource pool, where the matching expression defines the matching rule between resource requirements and the available resources of the business cluster; Determining computing nodes that meet the resource requirements from the target business cluster to form a dynamic business cluster corresponding to the user.

3. The method according to claim 2, where the matching expression includes a hard expression and a soft expression, the hard expression is used to define a matching rule that must be met, and the soft expression is used to define a matching rule that does not have to be met; the determining a target business cluster that meets the matching expression from the business resource pool based on the resource requirements indicated by the job parameters and the matching expression defined by the business resource pool includes: Based on the resource requirements indicated by the job parameters and the hard expression defined by the business resource pool, determining business clusters in the business resource pool that meet the hard expression as candidate business clusters; Based on the resource requirements indicated by the job parameters and the soft expression defined by the business resource pool, determining the matching scores of each candidate business cluster, and determining the candidate business cluster with the highest matching score as the target business cluster.

4. The method according to claim 3, where the soft expression defines calculating the matching score according to the load condition of the candidate business cluster.

5. The method according to claim 1, the method further includes: After receiving the job request, traversing each business cluster in the business resource pool, if there is any business cluster whose available resources meet the resource requirements, then performing the step of screening out computing nodes that meet the resource requirements from the business resource pool based on the resource requirements indicated by the job parameters to form a dynamic business cluster corresponding to the user.

6. The method according to claim 1, wherein the service resource pool is used to manage each service cluster based on the pooled tags of each service cluster, and the pooled tags are used to describe the cluster information of the service cluster; the step of screening out computing nodes that meet the resource requirements from the service resource pool based on the resource requirements indicated by the job parameters to form the dynamic service cluster corresponding to the user includes: Based on the resource requirements indicated by the job parameters and the pooled tags of each cluster, determining a target service cluster in the service resource pool whose pooled tags match the resource requirements, and determining computing nodes that meet the resource requirements from the target service cluster to form the dynamic service cluster corresponding to the user.

7. The method according to claim 6, wherein the pooled tags include default tags and custom tags, the default tags are non-modifiable tags, and the custom tags are modifiable tags, and the method further includes at least one of the following: In response to a first editing operation on the custom tag of any service cluster, adding a take-offline tag to the custom tag, and after the any service cluster becomes idle, adjusting the resource quota of the any service cluster; in response to a second editing operation on the custom tag of the any cluster, modifying the take-offline tag in the custom tag to a take-online tag, so that the any service cluster realizes seamless scaling for users; In response to a second editing operation on the custom tag of a newly added service cluster, adding a take-online tag to the custom tag; in response to a first editing operation on the custom tag of an original service cluster, adding a take-offline tag to the custom tag, so as to realize seamless job migration for users.

8. The method according to claim 1, wherein the computing nodes within the same service cluster are all homogeneous computing nodes, and the computing nodes in different service clusters are homogeneous computing nodes or heterogeneous computing nodes, and the dynamic service cluster is composed of homogeneous computing nodes in the same service cluster or different service clusters.

9. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; wherein, the processor realizes the steps of the method according to any one of claims 1-8 by running the executable instructions.

10. A computer-readable storage medium, characterized in that, A computer instruction is stored thereon, and when the instruction is executed by the processor, the steps of the method according to any one of claims 1-8 are realized.