Method and apparatus for controlling cluster resources, and cloud computing system

The method for controlling cluster resources addresses the challenge of poor applicability by determining binding relationships, adding resources to application pools, and deploying data packets, thereby enhancing the efficiency and flexibility of resource expansion across different applications.

JP7684304B2Active Publication Date: 2025-05-27BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022534290
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-05
Filing Date
2020-09-24
Publication Date
2025-05-27
Estimated Expiration
2040-09-24

AI Technical Summary

Technical Problem

Existing methods for controlling cluster resources, particularly in cloud systems, face challenges in applying resource expansion across different business applications, leading to poor applicability.

Method used

A method for controlling cluster resources that involves determining the binding relationship between resources to be expanded and applications, adding initialized resources to the corresponding application's resource pool, generating and deploying data packets for execution, and managing resource expansion and contraction processes.

Benefits of technology

This approach improves the applicability of resource expansion by enabling seamless integration with various business applications, enhancing operational efficiency and flexibility in managing cluster resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007684304000001
    Figure 0007684304000001
  • Figure 0007684304000002
    Figure 0007684304000002
  • Figure 0007684304000003
    Figure 0007684304000003
Patent Text Reader

Abstract

The present disclosure relates to a cluster resource control method and apparatus, and a cloud computing system, and relates to the technical field of computers. The method includes the steps of: when a resource to be controlled is a resource to be extended, determining a binding relationship between the resource to be extended and an application, adding the resource to be extended after initialization to a resource pool of a corresponding application having a binding relationship therewith, generating a data packet to be executed of the application to be processed based on a deployment type of the application to be processed, and deploying the data packet to be executed on the resource to be extended in the resource pool of the application to be processed for execution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross-reference to Related Applications This application is based on and claims priority to Chinese Patent Application No. 201911232841.4, filed on December 5, 2019, the entire disclosure of which is incorporated herein by reference.

[0002] The present disclosure relates to the field of computer technology, and more particularly, to a method for controlling cluster resources, a control device for cluster resources, a cloud computing system, and a non-volatile computer-readable storage medium.

Background Art

[0003] With the continuous consumption of resources when using cloud systems, resource control (such as capacity expansion and contraction) needs to be implemented in various products in the cluster. Such resource control becomes an operation that needs to be regularly performed by operation and maintenance personnel.

[0004] In related technologies, an expansion method for the application of a specific service of a cluster in a specific scene has been developed.

Summary of the Invention

Means for Solving the Problems

[0005] According to some embodiments of the present disclosure, a method for controlling cluster resources is provided. The method includes: when the resource to be controlled is the resource to be expanded, determining a binding relationship between the resource to be expanded and an application; adding the initialized resource to be expanded to the resource pool of the corresponding application having a binding relationship with the resource to be expanded; generating a data packet scheduled for execution of the application to be processed according to the deployment type of the application to be processed; and deploying the data packet scheduled for execution for execution on the resource to be expanded in the resource pool of the application to be processed.

[0006] In some embodiments, the step of adding the to-be-expanded resource that has a binding relationship with the to-be-expanded resource to the resource pool of the corresponding application includes the step of transferring the related information of the to-be-expanded resource to the pre-script of the corresponding application, and the step of executing the pre-script to complete the initialization of the to-be-expanded resource.

[0007] In some embodiments, the step of generating the to-be-executed data packet of the to-be-processed application includes, when the deployment type is package deployment, determining that the to-be-expanded resource is a physical machine and generating the program package of the to-be-processed application as the to-be-executed data packet, and when the deployment type is image deployment, determining that the to-be-expanded resource is a container image, generating the program package of the to-be-processed application, and generating the to-be-executed data packet according to the program package of the to-be-processed application and the running image.

[0008] In some embodiments, the step of deploying the to-be-executed data packet for execution on the to-be-expanded resource in the resource pool of the to-be-processed application includes, when the to-be-expanded resource is a physical machine, sending the to-be-executed data packet to the physical machine for execution, and when the to-be-expanded resource is a container image, sending the to-be-executed data packet to an idle physical machine in the resource pool that has a binding relationship with the to-be-processed application for execution.

[0009] In some embodiments, the step of deploying an execution-scheduled data packet for execution on an extension-scheduled resource in a resource pool of a to-be-processed application includes obtaining related information of the extension-scheduled resource in the resource pool, and sending the related information of the extension-scheduled resource to a third-party program of the to-be-processed application through a deployment interface configured for the to-be-processed application, whereby the third-party program deploys the execution-scheduled data packet for execution on the extension-scheduled resource according to the deployment mode of the third-party program.

[0010] In some embodiments, the method further includes the step of executing a postscript of the to-be-processed application, and the postscript is used for at least one of returning an extension result to a management node of the cluster, creating a volume for an extended resource in the resource pool, or cleaning up garbage generated by the extension.

[0011] In some embodiments, the method further includes the step of establishing a Secure Shell (SSH) connection with the extension-scheduled resource to execute a related script of the corresponding application, and only one SSH connection is established with the extension-scheduled resource at a time.

[0012] In some embodiments, the method further includes the step of securing an SSH connection within a preset time period after the execution of the execution-scheduled data packet of the corresponding application is completed.

[0013] In some embodiments, when the control-scheduled resource is a shrink-scheduled resource, the method further includes determining whether there is important data in the shrink-scheduled resource and whether there is a service dependent on the shrink-scheduled resource, and when there is no important data and no service dependent on the shrink-scheduled resource, deleting the shrink-scheduled resource from the cluster.

[0014] In some embodiments, the obtained related information of the resources scheduled for reduction is transferred through the configured reduction interface, and the obtained related information is used when deleting the resources scheduled for reduction from the cluster.

[0015] In some embodiments, the control method further includes the step of obtaining the reduction result by polling the configured query interface.

[0016] In some embodiments, after the resources scheduled for reduction are deleted from the cluster, if the resources scheduled for reduction are physical machines, the method includes adding the resources scheduled for reduction to the resource pool, and if the resources scheduled for reduction are container images, destroying the container images and adding the resources scheduled for reduction to the resource pool.

[0017] In some embodiments, if the resources scheduled for reduction are physical machines, the method further includes the step of executing a post-script for reduction, and the post-script is used to start an installation process for performing an operating system reinstallation on the resources scheduled for reduction.

[0018] According to another embodiment of the present disclosure, a control device for cluster resources is provided. The device includes a determination unit configured to determine the binding relationship between the resources scheduled for expansion and an application when the resources scheduled for control are the resources scheduled for expansion, an addition unit configured to add the initialized resources scheduled for expansion to the resource pool of the application having a binding relationship with the resources scheduled for expansion, a generation unit configured to generate the execution-scheduled data packets of the application scheduled for processing according to the deployment type of the application scheduled for processing, and an execution unit configured to deploy the execution-scheduled data packets for execution on the resources scheduled for expansion in the resource pool of the application scheduled for processing.

[0019] In some embodiments, the additional unit transfers the related information of the resource to be extended to the pre-script of the corresponding application, and executes the pre-script to complete the initialization of the resource to be extended.

[0020] In some embodiments, when the deployment type is package deployment, it is determined that the resource to be extended is a physical machine, and the generation unit generates the program package of the application to be processed as an execution-scheduled data packet. When the deployment type is image deployment, it is determined that the resource to be extended is a container image, and the generation unit generates the program package of the application to be processed, and generates an execution-scheduled data packet according to the program package of the application to be processed and the running image.

[0021] In some embodiments, when the resource to be extended is a physical machine, the execution unit sends the execution-scheduled data packet to the physical machine for execution. When the resource to be extended is a container image, the execution unit sends the execution-scheduled data packet to an idle physical machine in the resource pool having a binding relationship with the application to be processed for execution.

[0022] In some embodiments, the execution unit obtains the related information of the resource to be extended in the resource pool, and sends the related information of the resource to be extended to the third-party program of the application to be processed through the deployment interface configured for the application to be processed. The related information is used for the deployment of the execution-scheduled data packet on the resource to be extended for execution by the third-party program according to the deployment mode of the third-party program.

[0023] In some embodiments, the execution unit executes the postscript of the application to be processed, and the postscript is used for at least one of returning the expansion result to the management node of the cluster, creating a volume for the expanded resources in the resource pool, or cleaning up the garbage generated by the expansion.

[0024] In some embodiments, the apparatus further comprises an establishment unit configured to establish an SSH connection with the resources scheduled for expansion, and the SSH connection is used to execute the relevant script of the corresponding application, and only one SSH connection is established with the resources scheduled for expansion at a time.

[0025] In some embodiments, the establishment unit secures the SSH connection within a preset time period after the execution of the execution-scheduled data packet of the corresponding application is completed.

[0026] In some embodiments, when the resources scheduled for control are the resources scheduled for contraction, the apparatus further comprises a contraction unit configured to determine whether there is important data in the resources scheduled for control and whether there is a service dependent on the resources scheduled for contraction, and the contraction unit deletes the resources scheduled for contraction from the cluster when there is no important data and no service dependent on the resources scheduled for contraction.

[0027] In some embodiments, the contraction unit transfers the obtained relevant information of the resources scheduled for contraction through the configured contraction interface, and the obtained relevant information is used when deleting the resources scheduled for contraction from the cluster.

[0028] In some embodiments, the contraction unit obtains the contraction result by polling the configured query interface.

[0029] In some embodiments, after the scaling-down unit deletes the resources scheduled for scaling down from the cluster, if the resources scheduled for scaling down are physical machines, the additional unit adds the resources scheduled for scaling down to the resource pool, and if the resources scheduled for scaling down are container images, the scaling-down unit destroys the container image and the additional unit adds the resources scheduled for scaling down to the resource pool.

[0030] In some embodiments, if the resources scheduled for scaling down are physical machines, the execution unit executes a post-script for scaling down, and the post-script is used to start an installation process for performing an operating system reinstallation on the resources scheduled for scaling down.

[0031] According to still other embodiments of the present disclosure, there is provided a control device for cluster resources including a memory and a processor coupled to the memory, and the processor is configured to implement a method for controlling cluster resources according to any of the above embodiments based on instructions stored in the memory.

[0032] According to a further embodiment of the present disclosure, there is provided a non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements a method for controlling cluster resources according to any of the above embodiments.

[0033] According to still further embodiments of the present disclosure, there is provided a cloud computing system including a control device for cluster resources, and the device is configured to implement a method for controlling cluster resources according to any of the above embodiments.

[0034] The accompanying drawings described herein are used to provide a further understanding of the present disclosure, form a part of this application, and the exemplary embodiments and their descriptions of the present disclosure are used to explain the present disclosure and do not constitute an undue limitation to the present disclosure.

Brief Description of the Drawings

[0035]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Embodiments for Carrying Out the Invention

[0036] Regarding the technical solutions in the embodiments of the present disclosure, they will be clearly and fully described below together with the accompanying drawings in the embodiments of the present disclosure. However, it is obvious that the described embodiments are only a part and not all of the embodiments of the present disclosure. The following description of at least one exemplary embodiment is inherently for illustrative purposes only and shall in no way act as any limitation to the present disclosure and its application or use. All other embodiments that can be derived by those skilled in the art from the embodiments in the present disclosure without creative work are within the protection scope of the present disclosure.

[0037] The relative order, numerical expressions, and numerical values of the elements and steps described in these embodiments do not limit the scope of the present disclosure unless otherwise specified. On the other hand, it should be understood that the sizes of various parts shown in the drawings are not drawn to actual scale for ease of explanation. Techniques, methods, and devices known to those skilled in the art may not be discussed in detail, but when appropriate, they should be regarded as part of the given detailed description. In all examples shown and discussed in this specification, any specific values should be regarded as for illustration only and not for limitation. Therefore, other examples of exemplary embodiments may have different values. Note that similar reference numerals and letters refer to similar items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further considered in subsequent drawings.

[0038] The inventors of the present disclosure have noticed that in the above-mentioned technical field, there is a problem that resource expansion cannot be applied to different business applications, thereby resulting in poor applicability.

[0039] In this regard, the present disclosure provides a technical solution for controlling cluster resources that can improve the applicability of resource expansion.

[0040] FIG. 1 shows a flowchart of some embodiments of a method for controlling cluster resources according to the present disclosure.

[0041] As shown in FIG. 1, the method includes a step S11 of determining a binding relationship, a step S12 of adding an extension-planned resource, a step S13 of generating a scheduled execution data packet, and a step S14 of expanding the scheduled execution data packet.

[0042] In step S11, when the scheduled control resource is an extension-planned resource, the binding relationship between the extension-planned resource and the corresponding application is determined. For example, the extension-planned resource may be a physical machine or a container image.

[0043] In step S12, the initialized extension-planned resource is added to the resource pool of the corresponding application according to the binding relationship.

[0044] In some embodiments, the related information of the extension-planned resource is transferred to the pre-script of the corresponding application, and the pre-script is executed to complete the initialization of the extension-planned resource.

[0045] In step S13, according to the deployment (provisioning) type of the application to be processed, a scheduled execution data packet of the application to be processed is generated.

[0046] In some embodiments, when the deployment type is package deployment, the program package of the application to be processed is generated as the scheduled execution data packet. In this case, the extension-planned resource is a physical machine.

[0047] In some embodiments, when the deployment type is image deployment, the program package of the application to be processed is generated, and the scheduled execution data packet is generated according to the program package and the running image of the application to be processed. In this case, the extension-planned resource is a container image.

[0048] In step S14, the execution-scheduled data packet is deployed for execution on the corresponding extended-scheduled resource in the resource pool of the processing-scheduled application.

[0049] In some embodiments, if the extended-scheduled resource is a physical machine, the execution-scheduled data packet is sent to the physical machine for execution, and if the extended-scheduled resource is a container image, the execution-scheduled data packet is sent to an idle physical machine in the corresponding resource pool (standby pool) for execution. For example, the standby pool is a standby physical machine resource pool in a cloud computing system.

[0050] In some embodiments, the related information of the extended-scheduled resource is obtained, and the related information of the extended-scheduled resource is sent to the third-party program of the processing-scheduled application through the deployment interface configured for the processing-scheduled application, and the third-party program deploys the execution-scheduled data packet for execution on the extended-scheduled resource according to the deployment mode of the third-party program.

[0051] In some embodiments, the postscript of the processing-scheduled application is executed, and the postscript is used for at least one of returning the extension result to the management node of the cluster, creating the corresponding volume for the corresponding extended resource, and cleaning the garbage generated by the extension.

[0052] In some embodiments, an SSH connection with each extended-scheduled resource is established to execute each related script of the corresponding application, and only one SSH connection is established with each extended-scheduled resource at a time.

[0053] In some embodiments, after the execution of the execution-scheduled data packet of the corresponding application is completed, the SSH connection is secured within a preset time period, and by doing so, when the execution-scheduled data packet of the corresponding application is executed next time, the SSH connection is reused.

[0054] FIG. 2 shows a flowchart of another embodiment of the method for controlling cluster resources according to the present disclosure.

[0055] As shown in FIG. 2, the method further includes step S21 of determining important data and dependent services, and step S22 of deleting the resources scheduled for reduction.

[0056] In step S21, when the resource scheduled for control is a resource scheduled for reduction, it is determined whether there is important data in the resources scheduled for reduction and whether there is a service dependent on the resources scheduled for reduction.

[0057] In step S22, when there is no important data and there is no service dependent on the resources scheduled for reduction, the resources scheduled for reduction are deleted from the cluster.

[0058] In some embodiments, the obtained related information of the resources scheduled for reduction may be transferred in the system through the configured reduction interface, and by doing so, the resources scheduled for reduction are deleted from the cluster, and the reduction result is obtained by polling the configured query interface.

[0059] In some embodiments, when the resource scheduled for reduction is a physical machine, the resource scheduled for reduction is added to a resource pool (standby pool), and when the resource scheduled for reduction is a container image, the container image is destroyed and the resource scheduled for reduction is added to the resource pool.

[0060] In some embodiments, when the resource scheduled for reduction is a physical machine, the post-script for reduction is executed. The post-script is used to start an installation process for performing an operating system reinstallation on the resource scheduled for reduction. For example, the reinstallation may be implemented by PXE (Preboot eXecution Environment).

[0061] In the above embodiment, the data packets of the application are deployed for execution on the corresponding resource scheduled for expansion according to the binding relationship between the resource scheduled for expansion and the application. In this way, the method may be applied to resource expansion of different business applications, thereby improving the applicability of resource expansion.

[0062] FIG. 3 shows a schematic diagram of some embodiments of a control device for cluster resources according to the present disclosure.

[0063] As shown in FIG. 3, a user can interact with a backend expansion (or reduction) controller (i.e., a control device) through a front-end web page. For example, the front-end web page may be implemented by an Nginx (Engine x) server, and the control device may be implemented in the Golang programming language.

[0064] The internal structure of the control device may be a unified resource control system that targets a composite native cloud cluster and supports multi-factor simultaneous expansion and reduction. Each business line script can be connected to the control device.

[0065] For example, the control device may include a plurality of modules such as a multi-factor expansion and reduction controller, a heterogeneous element resource management module, a simultaneous expansion and reduction executor, a unified element construction system, a multi-factor deployment system, and an element reduction recovery system.

[0066] In some embodiments, a multi-element expansion and contraction controller (e.g., including a judgment unit, an addition unit, etc.) is responsible for controlling the overall expansion and contraction process. For example, the multi-element expansion and contraction controller can coordinate the work of other modules, control and call each module used in the expansion and contraction, and provide the necessary data parameters.

[0067] In some embodiments, a heterogeneous element resource management module is responsible for managing all metadata required in the expansion and contraction process. For example, the metadata includes server (physical machine resource) management IP (Internet Protocol), server specified parameters, expansion and contraction application data, etc.

[0068] For example, the heterogeneous element resource management module can permanently store data using a relational database MySQL (My Structured Query Language).

[0069] For example, the heterogeneous element resource management module can be deployed independently and provide an external OpenAPI (Open Application Programming Interface) call interface.

[0070] In some embodiments, a simultaneous expansion and contraction executor (which may include an execution unit, an establishment unit, etc.) is responsible for executing scripts that need to be run by different resource servers at different stages.

[0071] For example, the lower layer of the simultaneous expansion and contraction executor can be a connection pool implemented based on the SSH protocol. The simultaneous expansion and contraction executor can be implemented by the Golang programming language to ensure high concurrency. The simultaneous expansion and contraction executor can establish an SSH connection between the control device and the resources to be expanded.

[0072] For example, the simultaneous expansion and contraction executor may execute a customized expansion and contraction process. For example, for related applications of IaaS (Infrastructure as a Service), a customized expansion and contraction process may be executed.

[0073] For example, the simultaneous expansion and contraction executor may execute a standardized expansion and contraction process. For example, for related applications of services other than big data, cloud storage, and IaaS, a customized expansion and contraction process may be executed.

[0074] For example, the simultaneous expansion and contraction executor may trigger a unified element construction system, a multi-element deployment system, and an element contraction recovery system. The unified element construction system, the multi-element deployment system, and the element contraction recovery system implement related expansion and contraction processes on a cluster of a cloud computing system.

[0075] In some embodiments, the unified element construction system (which may include a generation unit, for example) is responsible for compiling the source code online deployed by each resource server and preparing different compilation environments for different resource servers based on Docker containers. The unified element construction system can implement isolation of the compilation environment, and by doing so, the dependent services required by compilation do not interfere with each other, thus ensuring the smoothness of compilation.

[0076] In some embodiments, the multi-element deployment system (which may include an execution unit, for example) is responsible for online deploying the compiled program to a specified physical server or container.

[0077] For example, a multi-element deployment system may include package deployment and image deployment. Package deployment can deploy a compiled program online to a specified physical machine, and image deployment can deploy and start a container according to a running image packed during compilation and according to required resources (including the number of CPU cores, memory size, hard disk size, etc. required by the operation).

[0078] In some embodiments, an element reduction recovery system (which may include a reduction unit, an additional unit, etc.) is responsible for recovering reduced resources.

[0079] For example, the recovered container resources are put back into the resource pool, and the recovered physical machine is put into the standby pool. The element reduction recovery system destroys the container resources, reinstalls the recovered physical machine, and formats its data disk.

[0080] FIG. 4 shows a flowchart of still other embodiments of a method for controlling cluster resources according to the present disclosure.

[0081] As shown in FIG. 4, the method may include an expansion flow (the left flow in the figure) and a reduction flow (the right flow in the figure). The expansion flow may include a standardized expansion flow and a customized expansion flow, and the reduction flow may include a standardized reduction flow and a customized reduction flow.

[0082] For example, the processes of expanding and shrinking various resource nodes related to IaaS and cloud database products (such as MySQL resource nodes, mass storage resource nodes, etc.) all belong to customized expansion and shrinking flows, and the processes of expanding and shrinking various resource nodes related to cloud storage and data clouds (such as resource node ds2 - datanode and big data resource node datanode) all belong to standardized expansion and shrinking flows.

[0083] The expansion and shrinking flows of the above products need to pass through the following three common flows, namely, the flow of allocating usage (that is, binding the resources planned for expansion and shrinkage to the corresponding product applications), the flow of executing pre - scripts, and the flow of executing post - scripts. These common flows can be extracted to establish a resource control method with high applicability.

[0084] In some embodiments, the standardized expansion flow mainly includes installing the standby machine (resource) planned for expansion into the standby pool of the specified product application, compiling and building the program package or running image of the application, deploying the program package or image to the server in the standby pool, and polling the deployment result. For example, the expansion flow for data cloud products is mainly for the big data resource node datanode, and the standardized expansion flow may include the following steps.

[0085] In step 1, usage is allocated to the host planned for expansion. For example, the tag of the big data resource node datanode is allocated to the server (resource) planned for expansion, thereby realizing the binding between the resource planned for expansion and the application.

[0086] In step 2, a pre-script is executed. For example, a multi-element expansion and reduction controller can transfer the IP list of the servers to be expanded to a pre-set script (pre-script) as command line parameters by calling a pre-set script of the data cloud in the FTP (File Transfer Protocol) directory of the control machine. For example, the control machine may be a computer running the control method of the present disclosure in the cluster.

[0087] By executing the expansion script, the initialization of the servers to be expanded can be completed. The pre-set script is mainly responsible for the initialization of the expanded physical machines, for example, the installation of some basic software packages.

[0088] In step 3, the host to be expanded is installed. For example, the physical machine to be expanded is installed in the standby pool of the corresponding product line in the product service tree to facilitate the unified allocation of resources. For example, according to the relevant information of the product to be deployed in the cloud computing system, the resources to be expanded may be added to the resource pool of the product line to which the bound application belongs.

[0089] In some embodiments, the relevant information of the product to be deployed in the cloud computing system may be stored in a tree structure. For example, it may be stored according to a five-level tree structure of department, product line, product, system, and application.

[0090] In step 4, data packets are compiled and constructed. For example, according to the deployment type of the application, the program package of the application is compiled and constructed. If packet deployment is used by the big data resource node datanode, the corresponding program package is compiled as a data packet to be processed. For image deployment, the program package and the running image are packaged together to generate a data packet to be processed. The compiled and constructed program package is transferred to the bound extended machine in the standby pool, and the start script is executed.

[0091] In step 5, the extension result is polled. For example, the deployment may be polled periodically, and the deployment result may be recorded in the deployment unit of the CMDB (Configuration Management Database) of the information management module. The web front end can check the deployment result through the API interface and present the deployment result to the user.

[0092] In step 6, a postscript is executed to process the situation after extension. For example, the postscript can notify the management node in the cluster of the extended resource node, thereby completing the extension of the cluster.

[0093] In some embodiments, for the cloud storage extension flow, the extension of the resource node ds2 - datanode is mainly completed, and the flow steps are the same as those of the above - mentioned data cloud. However, the postscript of the cloud storage extension flow can implement the process of creating a volume.

[0094] In some embodiments, the standardization reduction flow may include business data check, dependent service check, deletion from the cluster, resource recovery, polling result, and execution of the postscript. For example, the standardization reduction flow may include the following steps.

[0095] In step 1, a business data check is performed on the resources scheduled for reduction to determine whether there is important data among the resources scheduled for reduction. If there is important data, the resources scheduled for reduction do not support reduction.

[0096] In step 2, a dependent service check is performed on the resources scheduled for reduction to determine whether there is a service that depends on the resources scheduled for reduction for operation. If there is a dependent service, the resources scheduled for reduction do not support reduction.

[0097] In step 3, the management node of the cluster is notified that the resources scheduled for reduction have been deleted from the cluster.

[0098] In step 4, the element reduction recovery system is called to recover the reduced resources.

[0099] In step 5, the reduction result is polled.

[0100] In step 6, a postscript is executed to handle the situation after reduction. For example, a PXE installation operation for resource recovery may be triggered.

[0101] In some embodiments, the customized extension flow includes installing the extension standby machine into the standby pool of the specified product application, and providing, by a third party, its own extension program (third-party program), where the extension program provides two standardized extension interfaces, namely, an extension trigger interface (deployment interface) and an extension result query interface (query interface), and a multi-factor extension and reduction controller may include calling the extension trigger interface to transfer the information of the extension standby machine in the form of parameters, and polling the extension result through the extension result query interface.

[0102] For example, for an IaaS product application, the customized extension flow can be implemented by the following steps.

[0103] In step 1, usage is allocated to the extension target host. For example, the tags of the resource nodes of the IaaS product application are allocated to the extension target server.

[0104] In step 2, the pre-script is executed. For example, the multi-factor extension and reduction controller calls a pre-set script in the data cloud in the FTP directory of the control machine, transfers the IP list of the extension machines to the pre-set script as command line parameters, and executes the extension script to complete the initialization of the machines. The pre-set script is mainly responsible for the initialization of the extended physical machine, for example, the installation of some basic software packages.

[0105] In step 3, the extension target host is installed. For example, the extension target physical machine is installed into the standby pool of the corresponding product line in the product service tree for unified acquisition by the third-party program.

[0106] In step 4, the host to be expanded is transferred. For example, a multi-factor expansion and contraction controller triggers the IaaS deployment interface, obtains the relevant information of the server to be expanded from the CMDB of the information management module, the relevant information is transferred to the IaaS product in the form of parameters, the IaaS product deployment program obtains the server to be expanded deployed from the standby pool, and the IaaS deployment service deploys the IaaS product online according to its own-owned business.

[0107] In step 5, the polling interface is called. For example, a multi-factor expansion and contraction controller periodically polls the expansion result of the IaaS, the deployment result is recorded in the deployment unit of the CMDB of the information management module, and the web front end can query the deployment result through the API interface and present the deployment result to the user.

[0108] In step 6, the postscript is executed. This step may be blank if the product does not have a postscript.

[0109] In some embodiments, the customized contraction flow may include business data check, dependent service check, calling the customized contraction interface, polling the contraction result, recovering the resources to the resource pool or standby pool, and executing the postscript. For example, the customized contraction flow may be realized by the following steps.

[0110] In step 1, a business data check is performed on the resources scheduled for contraction to determine whether there is important data in the resources scheduled for contraction. If there is important data, the resources scheduled for contraction do not support contraction.

[0111] In step 2, a dependency service check is performed on the resources scheduled for reduction to determine whether there is a service that depends on the resources scheduled for reduction for operation. If there is a dependent service, the resources scheduled for reduction do not support reduction.

[0112] In step 3, a customized reduction interface is called to transfer the relevant information of the resources scheduled for reduction.

[0113] In step 4, a customized reduction result interface (query interface) is polled to obtain the reduction result.

[0114] In step 5, the recovered resources are put into the resource pool or the standby pool.

[0115] FIG. 5 shows a schematic diagram of another embodiment of the control device for cluster resources according to the present disclosure.

[0116] As shown in FIG. 5, the heterogeneous element resource management module mainly consists of a service layer and a data layer, and provides metadata to other service modules and the web front-end UI (user interface) through the Http (Hypertext Transfer Protocol) API. A monitor can perform service monitoring on the heterogeneous element resource management module.

[0117] In some embodiments, the service layer may include an API server, a server information management unit, a container information management unit, a product information management unit, a compilation information management unit, and a deployment information management unit.

[0118] For example, the API server may be an Http server developed based on Golang. The API server provides an external Restful (Representational State Transfer) style interface and is responsible for securely and legally providing internal management data to users.

[0119] For example, the server information management unit is mainly responsible for managing all physical machine information in the cluster. The physical machine information may include the management node server, the resource node servers in use, and the standby machines for the resources planned to be expanded / scaled down. The server information management unit can manage all server information and provide the necessary information for subsequent deployment, configuration, and recovery of the servers.

[0120] For example, the container information management unit is responsible for managing the information of all containers in the cluster. The container information may include the number of CPU cores, memory size, hard disk size used when the container is running, and the physical machine, product application, etc. to which the container belongs.

[0121] For example, the product information management unit is responsible for managing the information of products deployed in the private cloud. The product information may adopt a tree structure. The tree structure can be divided according to a five-level structure of department, product line, product, system, and application. Before expansion and scaling down, it is necessary to allocate the physical machines planned to be expanded / scaled down to specific system applications to facilitate centralized management of the expansion and scaling down processes.

[0122] For example, the compilation information management unit is responsible for managing the compilation programs of each product application. The compilation information management unit may be developed based on Jenkins and support the automatic extraction of code from the source code repository (gitlab). The compilation information management unit can compile and build the binary program or image of the application (including the running image and program package) according to the configuration information.

[0123] For example, the deployment information management unit is responsible for deploying the compiled program package or packaged image to the specified physical machine or container and recording all deployment records.

[0124] The data layer may include a master storage node, a slave storage node, and memory. The master storage node stores various information sent by the service layer and synchronizes and backs up the information to the slave storage node. For example, the data in the master storage node may be regularly backed up to the memory so that backup data can be obtained when both the master storage node and the slave storage node are contaminated.

[0125] FIG. 6 shows a schematic diagram of still another embodiment of the cluster resource control device according to the present disclosure.

[0126] As shown in FIG. 6, the multi-factor expansion and contraction controller may be an SSH controller built by Golang. Each business line script may be connected to the internal structure of the multi-factor expansion and contraction controller.

[0127] In some embodiments, the multi-factor expansion and contraction controller requires the assistance of the simultaneous expansion and contraction executor when executing pre-scripts and post-scripts. For example, the lower layer of the simultaneous expansion and contraction executor is a Golang-based high-performance SSH simultaneous connection pool, which includes SSH connections to various resources scheduled for expansion.

[0128] For example, based on the connection pool, a series of operation interfaces for connecting physical machines, executing script commands, etc. can be encapsulated. The high concurrency characteristics of the Golang language combined with the above connection pool can guarantee large-scale simultaneous operations in servers and containers, thereby improving the efficiency of the execution of expansion and contraction.

[0129] For example, to ensure operation security, the connection pool only establishes one SSH connection for each resource scheduled for expansion or contraction (physical machine, container image), and periodically cleans up expired SSH connections.

[0130] FIG. 7 shows a schematic diagram of some embodiments of a method for controlling cluster resources according to the present disclosure.

[0131] As shown in FIG. 7, the method includes the steps of a user submitting relevant code to a git distributed version control system, and git triggering Jenkins to automate the server's compilation, construction, and online deployment flow. For example, triggering the compilation and construction flow may include manual triggering and automatic triggering after the user submits the code.

[0132] Jenkins performs compilation construction and compilation images through Docker and performs online deployment. For example, based on the type of application, program packages or program images can be generated. If it is packet deployment used by the big data resource node datanode, only program packages are generated by compilation construction, and the running images are not included. If it is the image deployment type, the program packages and running images may be packaged together into a program image.

[0133] FIG. 8 shows a schematic diagram of a further embodiment of a control device for cluster resources according to the present disclosure.

[0134] As shown in FIG. 8, after the application program is compiled and constructed, the multi-element expansion and contraction controller transfers the data packets compiled and constructed by the unified element construction system through the simultaneous expansion and contraction executor to the physical machine or idle machine to be expanded in the resource pool according to the type of the expanded application.

[0135] When performing expansion on the physical machine program package, the multi-element deployment system executes the start script, deploys the program package on the physical machine of the specific cloud cluster for operation, checks the running state. In the case of program image deployment, the multi-element deployment system starts the image and checks the running state of the container of the specific cloud cluster.

[0136] FIG. 9 shows a schematic diagram of a further embodiment of a control device for cluster resources according to the present disclosure.

[0137] As shown in FIG. 9, the element contraction recovery system performs container resource recovery and physical machine resource recovery according to the type of application deployment. The recovered ResourceIf it is a physical machine of the dedicated cloud cluster, the element reduction recovery system adds the physical machine to the standby pool, invokes PXE to reinstall the system, and the recovered Resource If it is a program image (container) of the dedicated cloud cluster, the element reduction recovery system destroys the container and recovers the resources to the resource pool.

[0138] In some embodiments, before reduction, the element reduction recovery system needs to check whether there is already business data in the reduced server or container. Moreover, the element reduction recovery system needs to check whether the reduced server or container has deployed a service that another application depends on. If there is business data or dependent services, the element reduction recovery system exits the reduction process.

[0139] In some embodiments, to clean up the services deployed on the physical machine, the element recovery system may just start the PXE installation process. After the operating system reinstallation is performed on the reduced physical machine and its data disk is formatted, the reduced resources are put into the standby pool, or after the container is destroyed, the reduced resources are put into the resource pool.

[0140] FIG. 10 shows a block diagram of some embodiments of a control device for cluster resources according to the present disclosure.

[0141] As shown in FIG. 10, the control device 10 for cluster resources includes a determination unit 101, an addition unit 102, a generation unit 103, and an execution unit 104.

[0142] When the resource to be controlled is a resource to be expanded, the determination unit 101 determines the binding relationship between the resource to be expanded and the corresponding application.

[0143] The additional unit 102 adds the to-be-expanded resources that have been initialized to the resource pool of the corresponding application according to the binding relationship.

[0144] In some embodiments, the additional unit 102 transfers the related information of the to-be-expanded resources to the pre-script of the corresponding application and executes the pre-script to complete the initialization of the to-be-expanded resources.

[0145] The generation unit 103 generates the to-be-executed data packets of the to-be-processed application according to the deployment type of the to-be-processed application.

[0146] In some embodiments, when the deployment type is package deployment, the to-be-expanded resource is a physical machine, and the generation unit 103 generates the program package of the to-be-processed application as the to-be-executed data packets. When the deployment type is image deployment, the to-be-expanded resource is a container image, and the generation unit 103 generates the program package of the to-be-processed application and generates the to-be-executed data packets according to the program package of the to-be-processed application and the running image.

[0147] The execution unit 104 deploys the to-be-executed data packets in the corresponding to-be-expanded resources in the resource pool of the to-be-processed application for execution.

[0148] In some embodiments, when the to-be-expanded resource is a physical machine, the execution unit 104 sends the to-be-executed data packets to the physical machine for execution. When the to-be-expanded resource is a container image, the execution unit 104 sends the to-be-executed data packets to the idle physical machine in the corresponding resource pool for execution.

[0149] In some embodiments, the execution unit 104 acquires the related information of the resources scheduled for expansion, and sends the related information of the resources scheduled for expansion to the third-party program of the application scheduled for processing through the deployment interface configured for the application scheduled for processing. By doing so, the third-party program deploys the execution-scheduled data packet on the resources scheduled for expansion according to the deployment mode of the third-party program.

[0150] In some embodiments, the execution unit 104 executes the postscript of the application scheduled for processing, and the postscript is used for at least one of returning the expansion result to the management node of the cluster, creating the corresponding volume for the corresponding expanded resource, or cleaning the garbage generated by the expansion.

[0151] In some embodiments, the control device 10 of the cluster resources further includes an establishment unit 105 configured to establish an SSH connection with the resources scheduled for expansion. The SSH connection is used to execute the related script of the corresponding application. The establishment unit 105 can establish only one SSH connection with the resources scheduled for expansion at a time.

[0152] In some embodiments, the establishment unit 105 secures the SSH connection within a preset time period after the execution of the execution-scheduled data packet of the corresponding application is completed.

[0153] In some embodiments, when the control resource scheduled for control is the resource scheduled for contraction, the control device 10 of the cluster resources further includes a contraction unit 106 configured to determine whether there is important data in the resources scheduled for contraction and whether there is a service dependent on the resources scheduled for contraction. When there is no important data and no service dependent on the resources scheduled for contraction, the contraction unit 106 deletes the resources scheduled for contraction from the cluster.

[0154] In some embodiments, the scaling unit transfers the acquired related information of the resources to be scaled through the configured scaling interface, and by doing so, the resources to be scaled are deleted from the cluster.

[0155] In some embodiments, the scaling unit 106 acquires the scaling result by polling the configured query interface.

[0156] In some embodiments, after the scaling unit 106 deletes the resources to be scaled from the cluster, if the resources to be scaled are a physical machine, the additional unit 102 adds the resources to be scaled to the resource pool, and if the resources to be scaled are a container image, the scaling unit 106 destroys the container image, and the additional unit adds the resources to be scaled to the resource pool.

[0157] In some embodiments, if the resources to be scaled are a physical machine, the execution unit 104 executes the post-script for scaling, and the post-script is used to start an installation process for performing an operating system reinstallation on the resources to be scaled.

[0158] In the above embodiments, the data packets of the application are deployed for execution on the resources to be extended according to the binding relationship between the resources to be extended and the application. In this way, the device may be applied to the resource extension of different business applications, and by doing so, the applicability of resource extension is improved.

[0159] FIG. 11 shows a block diagram of another embodiment of the cluster resource control device according to the present disclosure.

[0160] As shown in FIG. 11, the cluster resource control device 11 of the present embodiment includes a memory 111 and a processor 112 coupled to the memory 111. The processor 112 is configured to implement a method for controlling cluster resources in any of the embodiments of the present disclosure based on instructions stored in the memory 111.

[0161] The memory 111 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory stores, for example, an operating system, an application, a boot loader, a database, and other programs.

[0162] FIG. 12 shows a block diagram of still another embodiment of a cluster resource control device according to the present disclosure.

[0163] As shown in FIG. 12, the cluster resource control device 12 of the present embodiment includes a memory 1210 and a processor 1220 coupled to the memory 1210. The processor 1220 is configured to implement a method for controlling cluster resources in any of the above embodiments based on instructions stored in the memory 1210.

[0164] The memory 1210 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory stores, for example, an operating system, an application, a boot loader, and other programs.

[0165] The cluster resource control device 12 may further include an input / output interface 1230, a network interface 1240, a storage interface 1250, and the like. These interfaces 1230, 1240, 1250, as well as the memory 1210 and the processor 1220, may be connected, for example, through a bus 1260. The input / output interface 1230 provides a connection interface for input / output devices such as a display, a mouse, a keyboard, and a touch screen. The network interface 1240 provides a connection interface for various network devices. The storage interface 1250 provides a connection interface for external storage devices such as an SD card and a USB flash drive.

[0166] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure may take the form of an embodiment that is entirely a hardware embodiment, an embodiment that is entirely a software embodiment, or an embodiment that combines software and hardware aspects. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable non-transitory storage media embodying computer-usable program code.

[0167] FIG. 13 shows a block diagram of some embodiments of a cluster resource control system according to the present disclosure.

[0168] As shown in FIG. 13, the cloud computing system 13 includes a cluster resource control device 131, which is configured to implement the cluster resource control method in any of the above embodiments.

[0169] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure may take the form of an embodiment that is entirely in hardware, an embodiment that is entirely in software, or an embodiment that combines software and hardware aspects. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable non-transitory storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) embodying computer-usable program code.

[0170] As described in detail above according to the present disclosure. Some details well known in the art are not described in order to avoid obscuring the concept of the present disclosure. Those skilled in the art should now be fully able to understand how to implement the technical solutions disclosed herein in view of the above description.

[0171] The methods and systems of the present disclosure may be implemented in several ways. The methods and systems of the present disclosure may be implemented, for example, in any combination of software, hardware, firmware, or software, hardware, and firmware. The above order for the steps of the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above unless otherwise specified. Furthermore, in some embodiments, the present disclosure may be implemented as a program recorded on a recording medium, and the program includes machine-readable instructions for implementing the method according to the present disclosure. Therefore, the present disclosure also covers a recording medium storing a program for executing the method according to the present disclosure.

[0172] Although some specific embodiments of the present disclosure have been described in detail by way of example, those skilled in the art should understand that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. It should be understood by those skilled in the art that modifications may be made to the above embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Explanation of Reference Numerals

[0173] 10 Control device 11 Control device 12 Control device 13 Cloud computing system 101 Judgment unit 102 Addition unit 103 Generation unit 104 Execution unit 105 Establishment unit 106 Reduction unit 111 Memory 112 Processor 131 Control device 1210 Memory 1220 Processor 1230 Input / output interface, interface 1240 Network interface, interface 1250 Storage interface, interface 1260 Bus

Claims

A method for controlling cluster resources, executed by a cluster resource control device, the method comprising: When the resource to be controlled is a resource to be expanded, determining a binding relationship between the resource to be expanded and an application; Adding the initialized resource to be expanded to a resource pool of the corresponding application that has the binding relationship with the resource to be expanded; Generating an execution-scheduled data packet of the application to be processed according to the deployment type of the application; Deploying the execution-scheduled data packet for execution on the resource to be expanded in the resource pool of the application. **Claim 2** The step of adding the initialized resource to be expanded to a resource pool of the corresponding application that has the binding relationship with the resource to be expanded includes: Transferring related information of the resource to be expanded to a pre-script of the corresponding application; Executing the pre-script to complete the initialization of the resource to be expanded. **Claim 3** The step of generating an execution-scheduled data packet of the application to be processed includes: When the deployment type is package deployment, determining that the resource to be expanded is a physical machine, and generating the program package of the application to be processed as the execution-scheduled data packet; When the deployment type is image deployment, determining that the resource to be expanded is a container image, generating the program package of the application to be processed, and generating the execution-scheduled data packet according to the program package and the running image of the application to be processed. **Claim 4** The step of deploying the execution-scheduled data packet for execution on the resource to be expanded in the resource pool of the application to be processed includes: When the resource to be expanded is the physical machine, sending the execution-scheduled data packet to the physical machine for execution; When the resource to be expanded is the container image, the method includes sending the execution-scheduled data packet for execution to an idle physical machine in a resource pool having the binding relationship with the application to be processed, as described in claim 3.

5. The step of deploying the execution-scheduled data packet for execution on the resource to be expanded in the resource pool of the application includes: obtaining related information of the resource to be expanded in the resource pool; and sending the related information of the resource to be expanded to a third-party program of the application to be processed through a deployment interface configured for the application to be processed, where the related information is used for deployment of the execution-scheduled data packet on the resource to be expanded for execution by the third-party program according to the deployment mode of the third-party program, as described in claim 1.

6. The method further includes executing a postscript of the application to be processed, where the postscript is used to return the result of expansion to the management node of the cluster, create a volume for the expanded resource in the resource pool, or clean the garbage generated by the expansion, and is used for at least one of them, as described in claim 1.

7. The method further includes establishing a Secure Shell (SSH) connection with the resource to be expanded to execute the related script of the corresponding application, where only one SSH connection is established with the resource to be expanded at a time, as described in claim 1.

8. After the execution of the execution-scheduled data packet of the corresponding application is completed, the method further includes ensuring the SSH connection within a preset time period, as described in claim 7.

9. When the resource scheduled for control is a resource scheduled for reduction, the method includes determining whether there is important data in the resource scheduled for reduction and whether there is a service dependent on the resource scheduled for reduction. When there is no important data and there is no service depending on the resource scheduled for reduction, further including the step of deleting the resource scheduled for reduction from the cluster, the control method according to any one of claims 1 to 8.

10. The step of deleting the resource scheduled for reduction from the cluster is The step of transferring the obtained related information of the resource scheduled for reduction through the configured reduction interface, wherein the obtained related information includes the step used when deleting the resource scheduled for reduction from the cluster, The control method according to claim 9, further including the step of obtaining a reduction result by polling the configured query interface.

11. After the resource scheduled for reduction is deleted from the cluster, when the resource scheduled for reduction is a physical machine, the step of adding the resource scheduled for reduction to the resource pool, and When the resource scheduled for reduction is a container image, further including the step of destroying the container image and adding the resource scheduled for reduction to the resource pool, Optionally, the control method is When the resource scheduled for reduction is the physical machine, the step of executing a post-script for reduction, wherein the post-script is used to start an installation process for performing an operating system reinstallation on the resource scheduled for reduction, and the control method according to claim 9 further includes the step.

12. A control device for cluster resources, A determination unit configured to determine a binding relationship between the resource scheduled for expansion and an application when the resource scheduled for control is the resource scheduled for expansion, An addition unit configured to add the initialized resource scheduled for expansion to the resource pool of the corresponding application having the binding relationship with the resource scheduled for expansion, A generation unit configured to generate an execution-scheduled data packet of the application scheduled for processing according to the deployment type of the application scheduled for processing, A control device comprising an execution unit configured to deploy the execution-scheduled data packet for execution on the resource scheduled for expansion in the resource pool of the application scheduled for processing.

13. A memory, A control device for cluster resources, comprising a processor coupled to the memory, wherein the processor is configured to implement the control method for cluster resources according to any one of claims 1 to 11 based on instructions stored in the memory.

14. A cloud computing system comprising a control device for cluster resources, wherein the control device is configured to implement the control method for cluster resources according to any one of claims 1 to 11.

15. A non-volatile computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the control method for cluster resources according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • System, method and program for deploying application

    JP2011060035A

  • Method and cloud management node for automated application deployment - Patents.com

    JP2019503535A

  • Method and system for modeling and analyzing computing resource requirements of software applications in a shared and distributed computing environment

    US20140047119A1

  • Throughput maintenance support system, device, method, and program

    WO2011105001A1