Cluster management method, device and system and electronic equipment
By obtaining monitoring parameters and elastic scaling strategies, flexible scaling management of server clusters is achieved, and the problem of single scaling strategy in the existing technology is solved, improving scaling efficiency and system flexibility.
Patent Information
- Application Number
- CN202510573177.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the scaling strategy of the server cluster is single and cannot be effectively adjusted for different business scenarios, resulting in low scaling efficiency.
By obtaining the identification of monitoring parameters, the identification of the target node group and the elastic scaling policy, the parameters of the target node group are monitored, and the capacity is expanded or reduced when the conditions are met, and the user-defined configuration and editing policies are supported to achieve flexible scaling management.
It improves the scaling efficiency of server clusters, and can formulate and execute corresponding scaling strategies according to different business scenarios, ensures the flexibility and configurability of the system, and avoids resource waste and overload.
Smart Images

Figure CN120343024A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of servers, and particularly to a cluster management method, device, system and electronic device. Background Art
[0002] With the continuous development of technology, Internet technology is more and more widely used in various industries. For the business service industry, services can be provided externally based on a server cluster to meet different business requests of users. During the process of providing services by the business system, there may be a situation where the access traffic of the business system increases or decreases sharply within a certain period of time. To meet user needs and avoid waste of resources, it is necessary to perform scaling management on the business system.
[0003] In the related art, scaling decisions are made by obtaining resource metrics in the server cluster. For example, when the resource metrics are higher than a certain threshold, the resources in the server cluster are scaled out. When the resource metrics are lower than a certain threshold, the resources in the server cluster are scaled in. As a result, the scaling strategies in the related art are relatively single and lack scalability, and cannot execute corresponding scaling strategies for different business scenarios, resulting in low scaling efficiency in the related art. Summary of the Invention
[0004] This application provides a cluster management method, device, system and electronic device.
[0005] According to a first aspect of this application, a cluster management method is provided. The method includes:
[0006] Obtain the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy; wherein, the identifier of the monitoring parameter is used to indicate the parameter that the target node group needs to be monitored, and the elastic scaling policy is used to indicate that when the monitoring parameter of the target node group meets the first condition, scale out or scale in the target node group;
[0007] Based on the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy, monitor the monitoring parameter of the target node group to scale out or scale in the target node group.
[0008] The cluster management method provided by the embodiments of this application monitors the target node group through the obtained identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy, and scales out or scales in the target node group when the monitoring parameter obtained by monitoring the target node group meets the first condition. In this way, users can configure the node groups to be monitored and the elastic scaling policy as needed, and can formulate and execute corresponding scaling strategies for different business scenarios, thereby greatly improving the scaling efficiency.
[0009] Optionally, monitoring the monitoring parameters of the target node group to scale the target node group up or down specifically includes:
[0010] Monitoring the monitoring parameters of the target node group;
[0011] Determining a first node group in the target node group, where the monitoring parameters of the first node group meet the first condition of the elastic scaling policy;
[0012] Scaling the first node group up or down.
[0013] By monitoring the target node group, corresponding monitoring parameters are obtained. In this way, the first node group that meets the first condition of the elastic scaling policy can be determined from the target node group according to the monitoring parameters, and the first node group can be scaled up or down, thereby enabling accurate scaling up or down of the corresponding node group and improving the efficiency of scaling.
[0014] Optionally, the first condition includes an expansion condition and a contraction condition. Determining the first node group in the target node group and scaling the first node group up or down specifically includes:
[0015] If it is determined that the monitoring parameters of the first node group meet the expansion condition, then scale the first node group up;
[0016] Or, if it is determined that the monitoring parameters of the first node group meet the contraction condition, then scale the first node group down.
[0017] By determining whether the first node group meets the expansion condition or the contraction condition, it is convenient to scale up when the first node group meets the expansion condition and scale down when the first node group meets the contraction condition, thereby enabling the scaling needs of the first node group in different situations to be met.
[0018] Optionally, obtaining the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy specifically includes:
[0019] Obtaining an elastic scaling policy script file, where the elastic scaling policy script file includes the elastic scaling policy;
[0020] Obtaining a configuration file, where the configuration file includes the identifier of the monitoring parameter, the identifier of the target node group, and the storage address of the elastic scaling policy script file;
[0021] Determining the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy according to the configuration file.
[0022] By obtaining an elastic scaling policy script file including an elastic scaling policy, and obtaining a configuration file containing the identifier of the monitoring parameter, the identifier of the target node group, and the storage address of the elastic scaling policy script file, the corresponding identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy can be obtained according to the identifier of the monitoring parameter, the identifier of the target node group, and the storage address of the elastic scaling policy script file.
[0023] Optionally, after determining the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy according to the configuration file, the method further includes:
[0024] Receiving an editing operation of the user on the configuration file;
[0025] Wherein, the editing operation includes: modifying, adding, or deleting the identifier of the monitoring parameter, modifying, adding, or deleting the identifier of the target node group, and modifying, adding, or deleting the elastic scaling policy;
[0026] Updating at least one of the identifier of the monitoring parameter, the identifier of the target node group, or the elastic scaling policy corresponding to the edited configuration file.
[0027] In the embodiment of the present application, through the editing operation of the configuration file, the user is allowed to adjust the configuration file at any time in case of business changes, etc., without restarting or interrupting the service, which can ensure the continuity of the system permission. And corresponding scaling policies can be executed according to different business scenarios and system loads, etc., which can improve the flexibility and configurability of the system, and further improve the scaling processing efficiency of the server cluster.
[0028] Optionally, the elastic scaling policy includes a first monitoring threshold and a first duration; determining that the node group in the target node group whose monitoring parameter satisfies the first condition of the elastic scaling policy is the first node group includes:
[0029] When there is first monitoring data greater than the first monitoring threshold in the monitoring data where the target node group is located, and the continuous duration of the first monitoring data is greater than the first duration, taking the node group corresponding to the first monitoring data as the first node group;
[0030] When there is first monitoring data greater than the first monitoring threshold in the monitoring data, and the continuous duration of the first monitoring data is greater than the first duration, it indicates that there is a first node group with too high load in the server cluster, and it is necessary to perform capacity expansion processing on the first node group with too high load. In this way, the node group that needs to be expanded in the server cluster can be determined in time.
[0031] Alternatively, when there is second monitoring data less than a second monitoring threshold in the monitoring data where the target node group is located, and the duration of the first monitoring data is greater than a second duration, the node group corresponding to the second monitoring data is taken as the first node group; wherein, the first monitoring threshold is greater than the second monitoring threshold.
[0032] When there is second monitoring data less than a second monitoring threshold in the monitoring data where the target node group is located, and the duration of the first monitoring data is greater than a second duration, it indicates that there is an idle first node group in the server cluster, and there is a situation of resource waste. In this way, the first node group can be scaled down, which is convenient for improving the resource utilization rate of the node group.
[0033] Optionally, the method further includes:
[0034] When it is necessary to scale up the first node group, increase the number of virtual machines or physical machines in the target node group;
[0035] Or, when it is necessary to scale down the first node group, reduce the number of virtual machines or physical machines in the target node group.
[0036] By controlling the number of virtual machines or physical machines in the node group, the purpose of reasonably utilizing resources can be achieved, and problems such as resource waste caused by a large amount of idle resources and overload caused by resource tension can be avoided.
[0037] According to a second aspect of the present application, there is provided a cluster management device, the device includes:
[0038] A configuration module, configured to obtain an identifier of a monitoring parameter, an identifier of a target node group, and an elastic scaling policy; wherein, the identifier of the monitoring parameter is used to indicate the parameter that the target node group needs to be monitored, and the elastic scaling policy is used to indicate that when the monitoring parameter of the target node group meets a first condition, the target node group is scaled up or down;
[0039] A monitoring module, configured to monitor the monitoring parameter of the target node group based on the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy, so as to scale up or down the target node group.
[0040] According to a third aspect of the present application, there is provided a cluster management system, including: a configuration module, a monitoring module, a scaling module, and at least one node group; wherein, each node group includes at least one virtual machine, and the virtual machine is used to execute user tasks;
[0041] The configuration module is used to obtain the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy, and send the identifier of the monitoring parameter and the identifier of the target node group to the monitoring module; wherein, the identifier of the monitoring parameter is used to indicate the parameter that the target node group needs to be monitored, and the elastic scaling policy is used to indicate that when the monitoring parameter of the target node group meets the first condition, the target node group is scaled out or in.
[0042] The monitoring module is used to receive the identifier of the monitoring parameter and the identifier of the target node group, monitor the target node group based on the identifier of the monitoring parameter and the identifier of the target node group, and send the obtained monitoring parameter to the configuration module.
[0043] The configuration module is further used to receive the monitoring parameter, and send a scaling out / in instruction to the scaling out / in module based on the monitoring parameter and the elastic scaling policy.
[0044] The scaling out / in module is used to receive the scaling out / in instruction, and instruct the target node group to scale out or in based on the scaling out / in instruction.
[0045] According to the fourth aspect of the present application, an electronic device is provided. The electronic device includes: a memory and a processor, a computer program is stored on the memory, and when the processor executes the program, the method as described above is implemented.
[0046] According to the fifth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above method of the present application is implemented.
[0047] According to the sixth aspect of the present application, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the above method of the present application is implemented. Description of the Drawings
[0048] In the following description of the exemplary embodiments in conjunction with the drawings, more details, features, and advantages of the present application are disclosed. In the drawings:
[0049] Figure 1 It is a system architecture diagram provided for an exemplary embodiment of the present application;
[0050] Figure 2 It is a flowchart of a cluster management method provided for an exemplary embodiment of the present application;
[0051] Figure 3 It is a schematic block diagram of the functional modules of a cluster management device provided for an exemplary embodiment of the present application;
[0052] Figure 4Block diagram of an electronic device provided by an exemplary embodiment of the present application;
[0053] Figure 5 Block diagram of a computer system provided by an exemplary embodiment of the present application. Detailed implementation manners
[0054] Embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.
[0055] It should be understood that the various steps recited in the method embodiments of the present application can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this regard.
[0056] As used herein, the term "including" and its variants are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present application are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or the interdependence relationship.
[0057] It should be noted that the modifications of "one" and "multiple" mentioned in the present application are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".
[0058] The names of the messages or information exchanged between multiple devices in the embodiments of the present application are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0059] It can be understood that before using the technical solutions disclosed in the embodiments of the present application, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present application should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0060] For example, when a user's active request is received, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solution of this application according to the prompt message.
[0061] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device. It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation manner of this application, and other manners that comply with relevant laws and regulations can also be applied to the implementation manner of this application.
[0062] In the embodiments provided by this application, as Figure 1 shown, Figure 1 FIG. 10 is a system architecture diagram of a cluster management system provided by an embodiment of this application. The cluster management system 10 includes a terminal 11, a cluster management device 20, and a server cluster 30. Among them, the cluster management device 20 includes: a configuration module 21, a monitoring module 22, and a scaling module 23. Among them, the terminal 11 registers an elastic scaling policy and monitoring parameters with the configuration module 21, and obtains monitoring data through the monitoring module 22. The configuration module 21 sends a scaling instruction to the scaling module 23 according to the monitoring data and the elastic scaling policy, and the scaling module 23 performs scaling processing on the node group in the server cluster 30 according to the received scaling instruction.
[0063] The terminal 11 is configured to receive registration information sent by the user, where the registration information includes an elastic scaling policy and monitoring parameters defined by the user.
[0064] The configuration module 21 is configured to receive the registration information sent by the terminal 11, and generate a configuration file according to the elastic scaling policy, the identifier of the target node group, and the identifier of the monitoring parameters in the registration information. The configuration module 21 sends a monitoring request including the monitoring parameters to the monitoring module 22. Among them, the configuration file may specifically include the identifier of the monitoring parameters, the identifier of the target node group, and the storage address of the elastic scaling policy script file. The identifier of the monitoring parameters is used to indicate the parameters that need to be monitored for the target node group, and the elastic scaling policy is used to indicate that when the monitoring parameters of the target node group meet the first condition, the target node group is expanded or scaled down.
[0065] In the embodiment, the configuration file can be specifically registered in the configuration module 21 in the form of Yaml (yet Another Markup Language, still a markup language) configuration. This Yaml configuration will specify the location of the configuration file, the names of the configuration items in the configuration file, and the node group tags, so as to implement the elastic scaling policy and the association between each node group.
[0066] The monitoring module 22 is used to receive the monitoring requests sent by the configuration module 21, and according to the monitoring parameters included in the monitoring requests, monitor the resource utilization status of each node group in the server cluster 30, obtain the monitoring data of each node group, and send the monitoring data to the configuration module 21.
[0067] For example, the monitoring module 22 can periodically capture the status of the monitored node components according to the configuration file through the HTTP (hypertext transfer protocol) protocol to obtain the monitoring data.
[0068] Among them, the monitoring module 22 is an open-source monitoring and alarm system and a time series database, mainly used for the monitoring and alarm of the server cluster, and provides support for the scaling of the server cluster through functions such as data collection, storage, query, and alarm.
[0069] The configuration module 21 is also used to receive the monitoring data sent by the monitoring module 22, determine the node groups in the server cluster 30 that need to be scaled up or down based on the monitoring data, and send a scaling up or down instruction to the scaling module 23 by executing the configuration file. The scaling up or down instruction will carry the identifier of the node group. Among them, the configuration file can specifically be a script file.
[0070] Specifically, the configuration module 21 will determine the node groups in the server cluster 30 that need to be scaled up and the node groups that need to be scaled down according to the monitoring data. The configuration module 21 will send a scaling up instruction to the scaling module 23, and the scaling up instruction will carry the identifier of the node group that needs to be scaled up; the configuration module 21 will send a scaling down instruction to the scaling module 23, and the scaling up instruction will carry the identifier of the node group that needs to be scaled down.
[0071] The scaling module 23 is used to receive the scaling up or down instructions sent by the configuration module 21, such as a scaling up instruction or a scaling down instruction. And based on the received scaling up or down instruction, call the API (application programming interface) of the server cluster 30 to complete the scaling up or down processing of the node groups in the server cluster 30. The scaling module 23 is the interface layer in the system that interacts with the external cloud service or physical hardware environment, responsible for connecting to the private cloud or physical server provider, and realizing the management and operation of resources.
[0072] In an embodiment, a node group in the server cluster 30 may include multiple physical machines, such as physical machine 301 and physical machine 302, etc. The physical machine may specifically be a server. Each physical machine may include multiple virtual machines. When it is necessary to scale the physical machines in the node group up or down, the scaling instruction also carries the identifier of the node group and the identifiers of the physical machines in the node group, so that the number of virtual machines in the physical machines in the node group can be increased or decreased, realizing the scaling processing of the specific physical machines in the node group. Among them, the server cluster 30 may include multiple node groups, such as node group 1 and node group 2, etc. Each node group may include the same resource type, which is convenient for monitoring.
[0073] It should be noted that the configuration module 21 and the scaling module 23 may be deployed in Kubernetes, which is a portable and extensible open-source platform. Kubernetes aims to provide a flexible and extensible framework for managing containerized workloads and services.
[0074] In an embodiment, the configuration module 21 may include a decision-making module and a monitoring module. When there are multiple elastic scaling policies in the configuration file, the configuration module 21 may correspondingly include multiple groups of decision-making modules and monitoring modules. The monitoring module will be associated with the corresponding node groups in the server cluster 30. By monitoring each node group through the monitoring module, it can be realized that one set of elastic scaling policies corresponds to one node group, and different node groups can respectively correspond to different elastic scaling policies. When performing scaling processing on each node group, the efficiency of scaling can be improved.
[0075] In an embodiment, the configuration module 21 may be deployed in the Kubernetes system, bind the elastic scaling policy to the corresponding node group through the label selector of Kubernetes, query the monitoring data through the monitoring module, and call the elastic scaling policy through the decision-making module to obtain an action flag bit, such as 0, 1, or 2. Among them, the flag bit 0 represents no operation, the flag bit 1 represents scaling up, and the flag bit 2 represents scaling down. In this way, the scaling module 23 can be called through the flag bit to perform the corresponding scaling processing on the node groups in the server cluster 30.
[0076] In an embodiment, the elastic scaling policy included in the configuration file is an editable policy. By receiving the user's editing operation on the configuration file, at least one of the following operations on the elastic scaling policy in the configuration file can be realized: policy modification, policy addition, or policy deletion. In addition, the configuration file includes monitoring parameters, and the monitoring parameters are editable items, so that the user's editing operation on the configuration file can be received, such as monitoring parameter modification, monitoring parameter addition, or monitoring parameter deletion, etc.
[0077] In the embodiments of the present application, through editable monitoring parameters and editable elastic scaling policies, users are allowed to adjust the monitoring parameters and elastic scaling policies at any time when the business changes, etc., without restarting or interrupting the service, which can ensure the continuity of the system. And corresponding elastic scaling policies can be executed according to different business scenarios and system loads, etc., which can improve the flexibility and configurability of the system, and further improve the scaling processing efficiency of the server cluster.
[0078] Based on the above embodiments, the embodiments of the present application further provide a cluster management method, which can be applied to the device where the above configuration module 21 is located. The configuration module 21 can be located in a server, such as Figure 2 shown, and the method may include the following steps:
[0079] In step S210, obtain the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy.
[0080] Among them, the identifier of the monitoring parameter is used to indicate the parameter that the target node group needs to be monitored, and the elastic scaling policy is used to indicate that when the monitoring parameter of the target node group meets the first condition, the target node group is scaled out or in.
[0081] In the embodiment, the user can customize the elastic scaling policy in the configuration file. For example, multiple elastic scaling policies formulated by the user according to needs can be obtained, so that a configuration file including the multiple elastic scaling policies can be obtained. The configuration file may include the storage address of the elastic scaling policy script file, and the elastic scaling policy script file may include the elastic scaling policy. By executing the elastic scaling policy script file, the scaling out or in processing of each node group in the server cluster can be realized.
[0082] In step S220, based on the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy, monitor the monitoring parameter of the target node group to scale out or in the target node group.
[0083] In the embodiment, the configuration file includes monitoring parameters, such as the usage rate of GPU (graphics processing unit), the usage rate of CPU (central processing unit), and the usage rate of memory, etc.
[0084] The node groups in the embodiments include the same type of resources in the server cluster. For example, the same node group includes GPUs or CPUs, etc. This facilitates monitoring and obtaining monitoring data. For example, according to the resource types included in each node group, the corresponding monitoring data can be obtained through the corresponding configuration file. In this way, by obtaining the monitoring data of each node group in the server cluster, the monitoring data can be compared with the corresponding monitoring thresholds to determine the target node group that needs to be scaled in or out.
[0085] In the embodiments, by determining the target node group that needs to be scaled out or in, the corresponding elastic scaling policy can be executed to perform scaling out processing on the node group that needs to be scaled out and perform scaling in processing on the node group that needs to be scaled in.
[0086] The elastic scaling policy can specifically determine the specific scaling out method or scaling in method for the target node group. For example, when the target node needs to be scaled out, a step-by-step scaling out method or a direct scaling out method can be adopted. The step-by-step scaling out method can divide the scaling out into multiple steps. After each step of scaling out, the monitoring data of the target node group is obtained, and it is judged whether it meets the expectation. If it meets the expectation, the next step of scaling out is performed on the target node group; if it does not meet the expectation, the intensity of the next step of scaling out can be adjusted. The direct scaling out method can determine the resources that need to be scaled out according to the monitoring data. For example, the corresponding number of virtual machines in the physical machine is increased, and the resources that need to be scaled out are added to the target node group, and the monitoring data of the target node group is obtained. If the monitoring data meets the expectation, this scaling out processing is completed; otherwise, the node group is scaled out again until the expectation is met.
[0087] Similarly, when the target node needs to be scaled in, a step-by-step scaling in method or a direct scaling in method can be adopted. The step-by-step scaling in method can divide the scaling in into multiple steps. After each step of scaling in, the monitoring data of the target node group is obtained, and it is judged whether it meets the expectation. If it meets the expectation, the next step of scaling in is performed on the target node group; if it does not meet the expectation, the intensity of the next step of scaling in can be adjusted. The direct scaling in method can determine the resources that need to be scaled in according to the monitoring data. For example, the corresponding number of virtual machines in the physical machine is reduced, and the resources that need to be scaled in are reduced from the target node group, and the monitoring data of the target node group is obtained. If the monitoring data meets the expectation, this scaling in processing is completed; otherwise, the node group is scaled in again until the expectation is met.
[0088] The cluster management method provided by the embodiments of the present application monitors the target node group based on the identifiers of the monitored parameters, the identifier of the target node group, and the elastic scaling policy. When the monitored parameters obtained by monitoring the target node group meet the first condition, the target node group is scaled out or in. In this way, users can configure the node groups to be monitored and the elastic scaling policy as needed, formulate and execute corresponding scaling in and out policies for different business scenarios, and thus can greatly improve the efficiency of scaling in and out.
[0089] In the embodiment provided by the present application, step S220 above may specifically further include the following steps:
[0090] In step S221, monitor the monitored parameters of the target node group.
[0091] In step S222, determine the first node group in the target node group, and scale out or in the first node group. Among them, the monitored parameters of the first node group meet the first condition of the elastic scaling policy.
[0092] In the embodiment, by monitoring the target node group, corresponding monitored parameters are obtained. In this way, the first node group that meets the first condition of the elastic scaling policy can be determined from the target node group according to the monitored parameters, and the first node group is scaled out or in, and thus accurate scaling out or in of the corresponding node group can be realized, and the efficiency of scaling in and out is provided.
[0093] In the embodiment, users can customize the monitored parameters in the configuration file and edit the monitored parameters as needed. For example, the monitored parameters can be modified, new monitored parameters can be added to the configuration file, or relevant monitored parameters in the configuration file can be deleted, etc. The embodiment is not limited thereto.
[0094] The monitored parameters in the configuration file are for obtaining corresponding monitoring data, and the monitoring data can provide effective support for the adjustment of the elastic scaling policy. Therefore, in different business scenarios, users can customize the corresponding configuration file, and when the business scenario changes, etc., the monitored parameters in the configuration file can also be edited, and thus the required monitoring data can be obtained. In this way, based on the monitoring data, the target node group that needs to be scaled in and out can be determined in a timely manner, and the corresponding elastic scaling policy is executed on the target node group, so as to improve the efficiency of scaling in and out of the target node group.
[0095] Based on the above embodiments, in the embodiments provided in the present application, the above configuration file may include multiple elastic scaling policies. In this way, when receiving a policy selection operation of the user on the configuration file, the currently required elastic scaling policy can be determined from the multiple elastic scaling policies based on the policy selection operation. Alternatively, when receiving a policy adjustment operation of the user on the currently executed elastic scaling policy, the currently executed elastic scaling policy can be adjusted to the target elastic scaling policy that needs to be executed accordingly based on the policy adjustment operation.
[0096] In the embodiments, the user can pre-customize multiple elastic scaling policies in the configuration file. For different business scenarios, according to the user's selection operation, the elastic scaling policy that needs to be currently executed can be determined from the multiple elastic scaling policies in the configuration file to meet the needs of different scenarios. In addition, when the currently executed elastic scaling policy does not match the current business scenario, the user can switch the elastic scaling policy applicable to the current business scenario to meet the needs of the current business scenario, thereby improving the efficiency of scaling.
[0097] Based on the above embodiments, in another embodiment provided in the present application, the above first condition includes an expansion condition and a contraction condition. The step S222 may specifically further include: determining that the monitoring parameters of the first node group meet the expansion condition, and then expanding the first node group. Or, determining that the monitoring parameters of the first node group meet the contraction condition, and then contracting the first node group.
[0098] In the embodiments, by monitoring the target node group, the corresponding monitoring parameters are obtained. In this way, the first node group that meets the first condition of the elastic scaling policy can be determined from the target node group according to the monitoring parameters, and the first node group can be expanded or contracted, thereby accurately expanding or contracting the corresponding node group and improving the efficiency of scaling.
[0099] In the embodiments provided in the present application, the user can install a monitoring plugin in the Kubernetes system, such as the above monitoring module 22, and use the monitoring plugin to set the corresponding monitoring parameters, such as GPU usage rate, GPU video memory occupancy, or CPU usage rate, etc. In this way, the monitoring plugin can monitor each node group in the server cluster according to the monitoring parameters and obtain monitoring data. The monitoring plugin can send the obtained monitoring data items to the configuration module 21.
[0100] Since the monitoring parameters are editable items, the monitoring parameters can be customized according to the user's operation. According to the monitoring request containing the monitoring parameters received by the monitoring plugin, the monitoring plugin can obtain the corresponding monitoring data according to the monitoring parameters, thereby adapting to the needs of different business scenarios.
[0101] Based on the above embodiments, in another embodiment provided by the present application, the above step S210 may specifically further include the following steps:
[0102] In step S211, obtain an elastic scaling policy script file. Wherein, the elastic scaling policy script file includes an elastic scaling policy.
[0103] In the embodiment, since the elastic scaling policy script file includes an elastic scaling policy, by executing the elastic scaling policy script file, it is possible to scale out or scale in the target node group as needed.
[0104] In step S212, obtain a configuration file. Wherein, the configuration file includes the identifier of the monitoring parameter, the identifier of the target node group, and the storage address of the elastic scaling policy script file.
[0105] In step S213, determine the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy according to the configuration file.
[0106] In the embodiment, by obtaining an elastic scaling policy script file including an elastic scaling policy and obtaining a configuration file including the identifier of the monitoring parameter, the identifier of the target node group, and the storage address of the elastic scaling policy script file, it is possible to obtain the corresponding identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy according to the identifier of the monitoring parameter, the identifier of the target node group, and the storage address of the elastic scaling policy script file.
[0107] In the embodiment, the user can adaptively adjust the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy in the configuration file, and then can monitor and perform scaling processing on the corresponding node group in a targeted manner.
[0108] In the embodiment provided by the present application, it is also possible to obtain training samples to train a preset model so that the trained preset model is used to predict monitoring data. For example, it is possible to obtain the monitoring data of each node group in the server cluster in the past period of time, that is, historical monitoring data. In this way, training samples can be generated based on the historical monitoring data, and the preset model can be trained with the training samples to obtain a trained preset model. The preset model can be a neural network model, etc., and the embodiment is not limited thereto.
[0109] By obtaining the monitoring data of each node group in the current server cluster to the trained preset model, it is possible to predict the monitoring data of each node group in the server cluster in the future period of time to obtain preset monitoring data. In this way, it is possible to judge the change trend of the monitoring data of each node group in the server cluster based on the predicted monitoring data, and then it is possible to timely determine whether it is necessary to adjust the elastic scaling policy currently executed in the configuration file.
[0110] Based on the above embodiments, in another embodiment provided by the present application, the method may further include the following steps:
[0111] In step S230, receive an editing operation of the user on the configuration file.
[0112] Wherein, the editing operation includes at least one of the following: policy modification, policy addition, or policy deletion.
[0113] In step S240, update at least one of the identifier of the corresponding monitoring parameter, the identifier of the target node group, or the elastic scaling policy according to the edited configuration file.
[0114] In the embodiment, the user can edit the configuration file through the terminal. By receiving the editing operation of the user on the configuration file, one or more of policy modification, policy addition, or policy deletion of the elastic scaling policy can be realized. In this way, by receiving the editing operation of the user on the configuration file, the expansion of the elastic scaling policy can be realized, and then different service scenarios can be well dealt with, and the execution efficiency of the elastic scaling policy can be improved.
[0115] Based on the above embodiments, in another embodiment provided by the present application, when determining the first node group that needs to be scaled in or out in the target node group in the server cluster based on the monitoring data, when there is first monitoring data greater than the first monitoring threshold in the monitoring data where the target node group is located, and the continuous duration of the first monitoring data is greater than the first duration, the node group corresponding to the first monitoring data is used as the first node group. Or, when there is second monitoring data less than the second monitoring threshold in the monitoring data where the target node group is located, and the continuous duration of the second monitoring data is greater than the second duration, the node group corresponding to the second monitoring data is used as the first node group; wherein, the first monitoring threshold is greater than the second monitoring threshold.
[0116] In the embodiment, when there is first monitoring data greater than the first monitoring threshold in the monitoring data, and the continuous duration of the first monitoring data is greater than the first duration, it indicates that there is a first node group with too high load in the server cluster, and the first node group with too high load needs to be expanded. For example, when there is a target node group with the GPU usage rate continuously exceeding 80% and the continuous duration exceeding 5 minutes, the first node group is expanded. In this way, the node group that needs to be expanded in the server cluster can be determined in time.
[0117] In the embodiment, when there is second monitoring data less than the second monitoring threshold in the monitoring data, and the continuous duration of the first monitoring data is greater than the second duration, it indicates that there is a first node group that is idle in the server cluster, and there is a situation of resource waste. In this way, the first node group can be scaled down, which is convenient for improving the resource utilization rate of the node group.
[0118] In the embodiments provided by the present application, when it is necessary to expand the above-mentioned first node group, the number of virtual machines or physical machines in the first node group can be increased. When it is necessary to downsize the above-mentioned target node group, the number of virtual machines or physical machines in the first node group can be reduced. In this way, by controlling the number of virtual machines or physical machines in the node group, the purpose of reasonably utilizing resources can be achieved, and problems such as resource waste caused by a large amount of idle resources and overload caused by resource tension can be avoided.
[0119] In the embodiments provided by the present application, when the business requirements change, the user can adjust the monitoring parameters in the configuration file and the elastic scaling policy in the configuration file at any time without restarting or interrupting the service, thereby ensuring the continuity and stability of the system. Through the user-defined monitoring parameters and elastic scaling policy, the system can be flexibly adjusted according to the actual business requirements, and different node groups are supported to use different policies.
[0120] In the case of dividing each functional module according to the corresponding functions, the embodiments of the present application provide a cluster management device, and the cluster management device can be a server, a terminal, or a chip applied to a server. Figure 3 It is a schematic block diagram of the functional modules of the cluster management device provided by an exemplary embodiment of the present application. As Figure 3 shown, the cluster management device includes:
[0121] A configuration module 31, configured to obtain the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy; wherein, the identifier of the monitoring parameter is used to indicate the parameter that the target node group needs to be monitored, and the elastic scaling policy is used to indicate that when the monitoring parameter of the target node group meets the first condition, the target node group is expanded or downsized;
[0122] A monitoring module 32, configured to monitor the monitoring parameter of the target node group based on the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy;
[0123] An expansion and contraction module 33, configured to expand or downsize the target node group.
[0124] In another embodiment provided by the present application, the monitoring module 32 is specifically configured to:
[0125] Monitor the monitoring parameter of the target node group;
[0126] Determine the first node group in the target node group, and the monitoring parameter of the first node group meets the first condition of the elastic scaling policy;
[0127] Expand or downsize the first node group.
[0128] In another embodiment provided by the present application, the first condition includes an expansion condition and a contraction condition. In another embodiment provided by the present application, the monitoring module 32 is further specifically configured to:
[0129] If it is determined that the monitoring parameters of the first node group meet the expansion condition, expand the first node group;
[0130] Or, if it is determined that the monitoring parameters of the first node group meet the contraction condition, contract the first node group.
[0131] In another embodiment provided by the present application, the configuration module 31 is specifically configured to:
[0132] Obtain an elastic scaling policy script file, where the elastic scaling policy script file includes the elastic scaling policy;
[0133] Obtain a configuration file, where the configuration file includes the identifier of the monitoring parameter, the identifier of the target node group, and the storage address of the elastic scaling policy script file;
[0134] Determine the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy according to the configuration file.
[0135] In another embodiment provided by the present application, the device further includes:
[0136] An edit operation receiving module, configured to receive an edit operation of the user on the configuration file;
[0137] Wherein, the edit operation includes at least one of the following: modifying a monitoring item, adding a monitoring item, or deleting a monitoring item;
[0138] A configuration file update module, configured to update at least one of the identifier of the corresponding monitoring parameter, the identifier of the target node group, or the elastic scaling policy according to the edited configuration file.
[0139] In another embodiment provided by the present application, the elastic scaling policy includes a first monitoring threshold and a first duration. In another embodiment provided by the present application, the monitoring module 32 is further specifically configured to:
[0140] When there is first monitoring data greater than the first monitoring threshold in the monitoring data where the target node group is located, and the continuous duration of the first monitoring data is greater than the first duration, use the node group corresponding to the first monitoring data as the first node group;
[0141] Alternatively, when there is second monitoring data less than a second monitoring threshold in the monitoring data where the target node group is located, and the duration of the second monitoring data is greater than a second duration, the node group corresponding to the second monitoring data is taken as the first node group; wherein, the first monitoring threshold is greater than the second monitoring threshold.
[0142] In another embodiment provided by the present application, the device further includes a scaling module, configured to:
[0143] When it is necessary to perform an expansion process on the first node group, increase the number of virtual machines or physical machines in the target node group;
[0144] When it is necessary to perform a contraction process on the first node group, reduce the number of virtual machines or physical machines in the target node group.
[0145] The cluster management device provided by the embodiments of the present application monitors the target node group through the identifier of the obtained monitoring parameter, the identifier of the target node group, and the elastic scaling policy, and expands or contracts the target node group when the monitoring parameter obtained by monitoring the target node group meets the first condition. In this way, the user can configure the node group to be monitored and the elastic scaling policy as needed, and can formulate and execute corresponding scaling policies for different business scenarios, thereby greatly improving the efficiency of scaling.
[0146] Based on the above embodiments, the embodiments of the present application further provide a cluster management system. Referring to Figure 1 as shown, the system may include: a configuration module 21, a monitoring module 22, a scaling module 23, and at least one node group. Wherein, each of the node groups includes at least one virtual machine, and the virtual machine is used to execute user tasks;
[0147] The configuration module 21 is configured to obtain the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy, and send the identifier of the monitoring parameter and the identifier of the target node group to the monitoring module 22; wherein, the identifier of the monitoring parameter is used to indicate the parameter that the target node group needs to be monitored, and the elastic scaling policy is used to indicate that when the monitoring parameter of the target node group meets the first condition, expand or contract the target node group;
[0148] The monitoring module 22 is configured to receive the identifier of the monitoring parameter and the identifier of the target node group, monitor the target node group based on the identifier of the monitoring parameter and the identifier of the target node group, and send the obtained monitoring parameter to the configuration module 21;
[0149] The configuration module 21 is further configured to receive the monitoring parameter, and send a scaling instruction to the scaling module based on the monitoring parameter and the elastic scaling policy;
[0150] The scaling module 23 is configured to receive the scaling instruction and instruct the target node group to scale out or scale in based on the scaling instruction.
[0151] For specific details, please refer to the description of the above embodiments, which will not be elaborated here.
[0152] An embodiment of the present application further provides an electronic device. The electronic device may include the above-mentioned cluster management device 20, and the electronic device may include: at least one processor; a memory for storing executable instructions of the at least one processor; wherein, the at least one processor is configured to execute the instructions to implement the above-mentioned method disclosed in the embodiments of the present application.
[0153] Figure 4 It is a schematic structural diagram of an electronic device provided for an exemplary embodiment of the present application. As Figure 4 shown, the electronic device 400 includes at least one processor 401 and a memory 402 coupled to the processor 401. The processor 401 may execute the corresponding steps in the above-mentioned method disclosed in the embodiments of the present application.
[0154] The above-mentioned processor 401 may also be referred to as a central processing unit (CPU). It may be an integrated circuit chip with signal processing capabilities. Each step in the above-mentioned method disclosed in the embodiments of the present application may be completed by the integrated logic circuit in the hardware of the processor 401 or by the instructions in software form. The above-mentioned processor 401 may be a general-purpose processor, a digital signal processor (DSP), an ASIC (Application Specific Integrated Circuit), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in the memory 402, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, and other mature storage media in the art. The processor 401 reads the information in the memory 402 and combines its hardware to complete the steps of the above-mentioned method.
[0155] In addition, when various operations / processes according to the present application are implemented by software and / or firmware, a program constituting the software can be installed from a storage medium or a network into a computer system having a dedicated hardware structure, such as Figure 5 the computer system 500 shown. When various programs are installed in the computer system, it can execute various functions, including the functions described above and the like. Figure 5 The block diagram of the computer system provided for an exemplary embodiment of the present application.
[0156] The computer system 500 is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described herein and / or claimed.
[0157] As Figure 5 shown, the computer system 500 includes a computing unit 501, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the computer system 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0158] Multiple components in the computer system 500 are connected to the I / O interface 505, including: an input unit 506, an output unit 507, a storage unit 508, and a communication unit 509. The input unit 506 can be any type of device that can input information into the computer system 500. The input unit 506 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 507 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 508 can include but is not limited to magnetic disks, optical disks. The communication unit 509 allows the computer system 500 to exchange information / data with other devices through a network such as the Internet, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a BluetoothTM device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0159] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above. For example, in some embodiments, the above-described methods disclosed in the embodiments of the present application can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device via the ROM 502 and / or the communication unit 509. In some embodiments, the computing unit 501 can be configured to execute the above-described methods disclosed in the embodiments of the present application by any other suitable means (e.g., by means of firmware).
[0160] The embodiments of the present application also provide a computer-readable storage medium, wherein when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the above-described methods disclosed in the embodiments of the present application.
[0161] The computer-readable storage medium in the embodiments of the present application can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The above computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the above. More specifically, the above computer-readable storage medium can include electrical connections based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0162] The above computer-readable medium can be included in the above electronic device; or it can exist separately and not be assembled into the electronic device.
[0163] The embodiments of the present application also provide a computer program product, including a computer program, wherein when the computer program is executed by a processor, the above-described methods disclosed in the embodiments of the present application are implemented.
[0164] In an embodiment of the present application, computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The foregoing programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network (including a local area network (LAN) or a wide area network (WAN)), or may be connected to an external computer.
[0165] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions denoted by the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0166] The modules, components, or units described in the embodiments of the present application may be implemented in software or in hardware. Among them, the names of the modules, components, or units do not constitute a limitation to the modules, components, or units themselves in some cases.
[0167] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.
[0168] The above description is only some embodiments of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by the mutual replacement of the above features with the technical features (but not limited to) having similar functions disclosed in the present application.
[0169] Although some specific embodiments of the present application have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration and not for limiting the scope of the present application. Those skilled in the art should understand that the above embodiments can be modified without departing from the scope and spirit of the present application. The scope of the present application is defined by the appended claims.
Claims
1. A cluster management method, characterized in that, The method includes: Obtaining an identifier of a monitoring parameter, an identifier of a target node group, and an elastic scaling policy; wherein, the identifier of the monitoring parameter is used to indicate the parameter to be monitored for the target node group, and the elastic scaling policy is used to indicate that when the monitoring parameter of the target node group meets a first condition, scaling up or down the target node group; Based on the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy, monitoring the monitoring parameter of the target node group to scale up or down the target node group.
2. The method according to claim 1, characterized in that, The monitoring of the monitoring parameter of the target node group to scale up or down the target node group specifically includes: Monitoring the monitoring parameter of the target node group; Determining a first node group in the target node group, where the monitoring parameter of the first node group meets the first condition of the elastic scaling policy; Scaling up or down the first node group.
3. The method according to claim 2, wherein The first condition includes a scaling-up condition and a scaling-down condition. The determining of the first node group in the target node group and scaling up or down the first node group specifically includes: If it is determined that the monitoring parameter of the first node group meets the scaling-up condition, then scaling up the first node group; Or, if it is determined that the monitoring parameter of the first node group meets the scaling-down condition, then scaling down the first node group.
4. The method according to claim 1, wherein The obtaining of the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy specifically includes: Obtaining an elastic scaling policy script file, where the elastic scaling policy script file includes the elastic scaling policy; Obtaining a configuration file, where the configuration file includes the identifier of the monitoring parameter, the identifier of the target node group, and the storage address of the elastic scaling policy script file; Determining the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy according to the configuration file.
5. The method according to claim 4, wherein After determining the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy according to the configuration file, the method further includes: Receiving an editing operation of the user on the configuration file; Wherein, the editing operation includes: modifying, adding, or deleting the identifier of the monitoring parameter, modifying, adding, or deleting the identifier of the target node group, and modifying, adding, or deleting the elastic scaling policy; Updating at least one of the corresponding identifier of the monitoring parameter, the identifier of the target node group, or the elastic scaling policy according to the edited configuration file.
6. The method according to claim 2, characterized in that, The elastic scaling policy includes a first monitoring threshold and a first duration; the determining of the first node group in the target node group includes: When there is first monitoring data greater than the first monitoring threshold in the monitoring data where the target node group is located, and the continuous duration of the first monitoring data is greater than the first duration, using the node group corresponding to the first monitoring data as the first node group; Alternatively, when there is second monitoring data less than a second monitoring threshold in the monitoring data of the target node group, and the duration of the second monitoring data is greater than a second duration, the node group corresponding to the second monitoring data is taken as the first node group; wherein, the first monitoring threshold is greater than the second monitoring threshold.
7. The method according to claim 2, wherein The method further includes: When it is necessary to expand the first node group, increasing the number of virtual machines or physical machines in the target node group; Alternatively, when it is necessary to scale down the first node group, reducing the number of virtual machines or physical machines in the target node group.
8. A cluster management device, characterized in that, The apparatus includes: A configuration module, configured to obtain an identifier of a monitoring parameter, an identifier of a target node group, and an elastic scaling policy; wherein, the identifier of the monitoring parameter is used to indicate the parameter to be monitored for the target node group, and the elastic scaling policy is used to indicate that when the monitoring parameter of the target node group meets a first condition, the target node group is expanded or scaled down; A monitoring module, configured to monitor the monitoring parameter of the target node group based on the identifier of the monitoring parameter, the identifier of the target node group, and the elastic scaling policy, so as to expand or scale down the target node group.
9. A cluster management system, characterized in that, Includes: A configuration module, a monitoring module, a scaling module, and at least one node group; wherein, each node group includes at least one virtual machine, and the virtual machine is used to execute user tasks; The configuration module is configured to obtain an identifier of a monitoring parameter, an identifier of a target node group, and an elastic scaling policy, and send the identifier of the monitoring parameter and the identifier of the target node group to the monitoring module; wherein, the identifier of the monitoring parameter is used to indicate the parameter to be monitored for the target node group, and the elastic scaling policy is used to indicate that when the monitoring parameter of the target node group meets a first condition, the target node group is expanded or scaled down; The monitoring module is configured to receive the identifier of the monitoring parameter and the identifier of the target node group, monitor the target node group based on the identifier of the monitoring parameter and the identifier of the target node group, and send the obtained monitoring parameter to the configuration module; The configuration module is further configured to receive the monitoring parameter, and send a scaling instruction to the scaling module based on the monitoring parameter and the elastic scaling policy; The scaling module is configured to receive the scaling instruction, and instruct the target node group to expand or scale down based on the scaling instruction.
10. An electronic device, characterized in that, Includes: At least one processor; A memory for storing executable instructions of the at least one processor; Wherein, the at least one processor is configured to execute the instructions to implement the method according to any one of claims 1-7.