Method, device and equipment for expanding and shrinking capacity of working load, medium and program product

By automatically creating HPA resources and automatically adjusting workloads in Kubernetes clusters based on metrics, the problem of inefficiency of manual operations is solved, and more efficient resource utilization and accurate scaling is achieved.

CN119938217APending Publication Date: 2025-05-06CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411824207.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In Kubernetes clusters, manual operations increase or decrease Pod replicas to adjust workloads, especially when the cluster is large, inefficient and affect user experience.

Method used

Automatically create HPA resources by obtaining metrics associated with the target workload to achieve automatic scaling or shrinking of the workload.

Benefits of technology

It avoids the cumbersomeness of manual operations, saves time, improves efficiency, and users can customize indicators, improving the accuracy of capacity expansion or reduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938217A_ABST
    Figure CN119938217A_ABST
Patent Text Reader

Abstract

The invention discloses a capacity expansion and contraction method and device for workloads, equipment, a medium and a program product, and the method comprises the steps: obtaining at least one index related to a target workload in a cluster, and the at least one index comprises a first index set by a user and / or a second index corresponding to the cluster; based on the at least one index, creating an HPA resource corresponding to the target workload; and carrying out capacity expansion or capacity reduction on the target workload through the HPA resource and the at least one index. According to the method and the device, the corresponding HPA resource can be automatically created for the target workload, so that automatic capacity expansion or capacity reduction is realized, the complexity of manual operation is avoided, the time is saved, the efficiency is improved, a user can customize indexes, and the accuracy of capacity expansion or capacity reduction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to, but is not limited to, the field of computer technology, and in particular to a method, apparatus, device, medium, and program product for scaling workloads. Background Art

[0002] Kubernetes (k8s for short) is a popular container orchestration tool that provides excellent solutions in service deployment, service monitoring, application expansion and fault handling.

[0003] However, when the number of Pods of a workload in a k8s cluster needs to be adjusted, users need to manually enter instructions to increase or decrease Pod copies to improve the resource utilization of the entire cluster. When the k8s cluster is large, the number of workloads in the cluster will be relatively large. Scaling and shrinking in the above way is time-consuming and inefficient, affecting the user experience. Summary of the invention

[0004] In view of this, the present application provides a method, apparatus, device, medium and program product for scaling up or down of a workload, which can automatically create corresponding HPA resources for the target workload to achieve automatic scaling up or down, avoiding manual operation.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] In a first aspect, the present application provides a method for scaling a workload, comprising: obtaining at least one indicator associated with a target workload in a cluster, wherein the at least one indicator includes a first indicator set by a user, and / or a second indicator corresponding to the cluster; based on the at least one indicator, creating an HPA resource corresponding to the target workload; and scaling the target workload up or down through the HPA resources and at least one indicator.

[0007] In some embodiments, the target workload is scaled down through HPA resources and at least one indicator, including: determining the deletion cost of each Pod included in the target workload based on at least one indicator; determining the Pod replica to be deleted based on the deletion cost and the deletion policy corresponding to the target workload; and deleting the Pod replica to be deleted through HPA resources.

[0008] In some embodiments, at least one indicator includes a first indicator and a second indicator; based on the at least one indicator, determining the deletion cost of each Pod included in the target workload, including: determining a first weighted value based on a first sub-indicator associated with the second indicator in the first indicator and the second indicator; determining a second weighted value based on a second sub-indicator associated with each Pod in the first indicator and a parameter value corresponding to the second sub-indicator; determining the deletion cost of each Pod included in the target workload based on the first weighted value and the second weighted value.

[0009] In some embodiments, based on the deletion cost value and the deletion policy corresponding to the target workload, the Pod replica to be deleted is determined, including: obtaining the usage status of each Pod included in the target workload; when the usage status is a non-running state, determining the Pod replica corresponding to the non-running Pod as the Pod replica to be deleted; when the usage status is a running state, determining the Pod replica to be deleted according to the target parameters associated with the running Pod and the deletion cost value corresponding to the running Pod, wherein the target parameters include at least one of the readiness status corresponding to the running Pod, the number of Pod replicas in the node, the preparation time, the number of container restarts, and the creation time.

[0010] In some embodiments, a copy of the Pod to be deleted is determined based on target parameters associated with the running Pod and a deletion cost value corresponding to the running Pod, including: when the readiness status corresponding to the running Pod is not ready, determining that the copy corresponding to the first Pod is the copy of the Pod to be deleted, wherein the first Pod is the Pod with the smallest deletion cost value among the unready Pods.

[0011] In some embodiments, the method further includes: when the readiness status of the Pod in the running state is ready, and the deletion cost values ​​corresponding to the prepared Pods are the same, performing one of the following: when the number of Pod copies in the nodes where at least two Pods in the prepared Pods are located is different, determining that the copy corresponding to the second Pod is the Pod copy to be deleted, wherein the second Pod is the Pod with the largest number of Pod copies in the node where it is located among the prepared Pods; when the number of Pod copies in the node where each Pod in the prepared Pods is located is the same, and the preparation times corresponding to at least two Pods are different, determining that the copy corresponding to the third Pod is the Pod copy to be deleted, wherein, The third Pod is the Pod with the longest preparation time among the prepared Pods; when the number of Pod copies in the node where each Pod is located in the prepared Pods is the same and the preparation time corresponding to each Pod is the same, the copy corresponding to the fourth Pod is determined to be the Pod copy to be deleted, wherein the fourth Pod is the Pod with the largest number of container restarts among the prepared Pods; when the number of Pod copies in the node where each Pod is located in the prepared Pods is the same, the preparation time corresponding to each Pod is the same, and the number of container restarts for each Pod is the same, the copy corresponding to the fifth Pod is determined to be the Pod copy to be deleted, wherein the fifth Pod is the Pod with the latest creation time among the prepared Pods.

[0012] In some embodiments, obtaining at least one metric associated with a target workload in a cluster includes: filtering a namespace based on preset fields to obtain a creation event of the target workload in the cluster; and obtaining at least one metric through an interface of a target controller in the cluster in response to the creation event.

[0013] In some embodiments, the target workload is expanded through HPA resources and at least one indicator, including: determining the number N of Pod replicas to be added corresponding to the target workload through at least one indicator, where N is a positive integer; adding N Pod replicas through HPA resources and the number N of Pod replicas to be added.

[0014] In a second aspect, the present application provides a workload scaling device, comprising: an acquisition module for acquiring at least one indicator associated with a target workload in a cluster, wherein the at least one indicator includes a first indicator set by a user, and / or a second indicator corresponding to the cluster; a creation module for creating HPA resources corresponding to the target workload based on at least one indicator; and a scaling module for scaling the target workload through HPA resources and at least one indicator.

[0015] In some embodiments, the scaling module includes: a first determination submodule, used to determine the deletion cost of each Pod included in the target workload based on at least one indicator; a second determination submodule, used to determine the Pod replica to be deleted based on the deletion cost and the deletion policy corresponding to the target workload; a deletion submodule, used to delete the Pod replica to be deleted through HPA resources.

[0016] In some embodiments, at least one indicator includes a first indicator and a second indicator; the first determination submodule is further used to perform the following steps: determine a first weighted value based on the first sub-indicator and the second indicator associated with the second indicator in the first indicator; determine a second weighted value based on the second sub-indicator associated with each Pod in the first indicator and the parameter value corresponding to the second sub-indicator; determine the deletion cost of each Pod included in the target workload based on the first weighted value and the second weighted value.

[0017] In some embodiments, the second determination submodule includes: an acquisition unit, used to: acquire the usage status of each Pod included in the target workload; a first determination subunit, used to determine, when the usage status is a non-running state, that the Pod copy corresponding to the non-running Pod is the Pod copy to be deleted; a second determination subunit, used to determine, when the usage status is a running state, the Pod copy to be deleted based on the target parameters associated with the running Pod and the deletion cost value corresponding to the running Pod, wherein the target parameters include at least one of the readiness status corresponding to the running Pod, the number of Pod copies in the node, the preparation time, the number of container restarts, and the creation time.

[0018] In some embodiments, the second determination subunit is used to: when the readiness status corresponding to the Pod in the running state is not ready, determine that the copy corresponding to the first Pod is the Pod copy to be deleted, wherein the first Pod is the Pod with the smallest deletion cost among the unready Pods.

[0019] In some embodiments, the second determination subunit is further used to: when the readiness status of the Pod corresponding to the running state is prepared and the deletion cost values ​​corresponding to the prepared Pods are the same, perform one of the following: when the number of Pod copies in the nodes where at least two Pods in the prepared Pods are located is different, determine that the copy corresponding to the second Pod is the Pod copy to be deleted, wherein the second Pod is the Pod with the largest number of Pod copies in the node where it is located among the prepared Pods; when the number of Pod copies in the node where each Pod in the prepared Pods is located is the same, and the preparation times of at least two Pods are different, determine that the copy corresponding to the third Pod is the Pod copy to be deleted, wherein , the third Pod is the Pod with the longest preparation time among the prepared Pods; when the number of Pod copies in the node where each Pod in the prepared Pods is located is the same and the preparation time corresponding to each Pod is the same, the copy corresponding to the fourth Pod is determined to be the Pod copy to be deleted, wherein the fourth Pod is the Pod with the largest number of container restarts among the prepared Pods; when the number of Pod copies in the node where each Pod in the prepared Pods is located is the same, the preparation time corresponding to each Pod is the same, and the number of container restarts for each Pod is the same, the copy corresponding to the fifth Pod is determined to be the Pod copy to be deleted, wherein the fifth Pod is the Pod with the latest creation time among the prepared Pods.

[0020] In some embodiments, the acquisition module is used to: filter the namespace based on preset fields to obtain creation events of target workloads in the cluster; and obtain at least one indicator through an interface of a target controller in the cluster in response to the creation event.

[0021] In some embodiments, the scaling module is also used to perform the following steps: determine the number N of Pod replicas to be added corresponding to the target workload through at least one indicator, where N is a positive integer; add N Pod replicas through HPA resources and the number N of Pod replicas to be added.

[0022] In a third aspect, the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, and when the processor executes the computer program, some or all of the steps in the above method are implemented.

[0023] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements some or all of the steps in the above method when the computer program is executed by a processor.

[0024] In a fifth aspect, the present application provides a computer program product, including a computer program or instructions, which implement some or all of the steps in the above method when executed by a processor.

[0025] In a sixth aspect, the present application provides a computer program, comprising a computer-readable code. When the computer-readable code runs in an electronic device, a processor in the electronic device executes some or all of the steps for implementing the above method.

[0026] In the present application, by obtaining at least one indicator associated with the target workload in the cluster, namely: a first indicator set by the user, and / or a second indicator corresponding to the cluster, based on the at least one indicator, an HPA resource corresponding to the target workload is created, and the target workload is expanded or reduced through the HPA resources and at least one indicator. The corresponding HPA resources can be automatically created for the target workload to achieve automatic expansion or reduction, avoiding the tedious manual operation, saving time and improving efficiency, and users can customize indicators to improve the accuracy of expansion or reduction.

[0027] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to illustrate the technical solution of the present application.

[0029] Figure 1 is a structural schematic diagram of an electronic device provided in an embodiment of the present application;

[0030] Figure 2 It is a structural schematic diagram of a workload expansion and contraction device provided in an embodiment of the present application;

[0031] Figure 3 It is a structural schematic diagram of a workload expansion and contraction system provided in an embodiment of the present application;

[0032] Figure 4 It is a schematic diagram of an implementation process of the workload expansion and contraction method provided in an embodiment of the present application;

[0033] Figure 5 It is another implementation flow diagram of the workload expansion and contraction method provided in the embodiment of the present application;

[0034] Figure 6 This is another implementation flow diagram of the workload scaling method provided in the embodiment of the present application. DETAILED DESCRIPTION

[0035] The following is a description of the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. In the following description, reference is made to the drawings that form a part of the present application and show the specific aspects of the embodiments of the present application or the specific aspects of the embodiments of the present application in an illustrative manner. It should be understood that the embodiments of the present application can be used in other aspects and may include structural or logical changes not depicted in the drawings. Therefore, the following detailed description should not be understood in a restrictive sense, and the scope of the present application is defined by the appended claims. For example, it should be understood that the disclosure of the described method can be equally applicable to the corresponding device or system for performing the method, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units such as functional units to perform the one or more method steps described (for example, one unit performs one or more steps, or multiple units, each of which performs one or more of the multiple steps), even if such one or more units are not explicitly described or illustrated in the drawings. On the other hand, for example, if a specific device is described based on one or more units such as functional units, the corresponding method may include a step to perform the functionality of the one or more units (e.g., a step to perform the functionality of the one or more units, or multiple steps, each of which performs the functionality of one or more of the multiple units), even if such one or more steps are not explicitly described or illustrated in the drawings. Further, it should be understood that, unless otherwise explicitly stated, the features of the various exemplary embodiments and / or aspects described herein may be combined with each other.

[0036] As an open source container orchestration platform, k8s can manage containerized applications on multiple hosts in the cloud platform. Its goal is to make the deployment of containerized applications simple and efficient. It provides a mechanism for application deployment, planning, updating, and maintenance. When the number of Pods of a workload in the k8s cluster needs to be adjusted, you need to manually enter instructions to increase or decrease the number of Pod copies, thereby improving the resource utilization of the entire k8s cluster. If the k8s cluster is large, the number of workloads will be relatively large, and it will be time-consuming to enter instructions one by one to expand or shrink capacity.

[0037] To solve the above problems, k8s provides a horizontal pod autoscaler (HPA) method. HPA can automatically scale the number of Pods in the Replication Controller (RC), Deployment, and Replicaset based on the central processing unit (CPU) utilization or memory utilization. Horizontal scaling means that the response to the increase in workload is to deploy more Pods, which is different from vertical scaling. For k8s, vertical scaling means allocating more resources to the workload that is already running.

[0038] Here, RC is a resource object in k8s that is used to ensure that a specified number of Pod copies are running in the cluster. It is one of the core components of k8s and is used to achieve high availability and horizontal scaling of applications. Deployment is an advanced controller in k8s for defining and managing Pods. Replicaset is a resource object in k8s that is used to ensure that a specified number of Pod copies are running in the cluster. It is similar to ReplicationController, but provides a more powerful selector mechanism for managing Pod copies.

[0039] In k8s, the implementation of HPA is set as an intermittent control loop, and the control loop of the specified period (e.g., the default period is 15 seconds) is implemented by setting the period parameter (e.g., horizontal-Pod-autoscaler-sync-period). In each period, the controller manager of k8s queries the resource utilization according to the indicators specified in HPA. The specific process is as follows: the controller manager first determines the defined target resource (such as the scaleTargetRef field), then selects the selector according to the label of the target resource, such as (Pod.spec.selector), and then obtains the indicator from the resource indicator application programming interface (application programming interface, API) of each Pod targeted by HPA. If the utilization value is set, the scheduling controller in k8s will calculate the utilization value as a percentage of the equivalent resource request on the container in each Pod; if the target original value is set, the original indicator value is used directly. Finally, the controller manager obtains the utilization of all target Pods or the average of the original indicator values, and generates a ratio for scaling the required number of replicas.

[0040] Here, in the HPA of k8s, the scaleTargetRef field is used to specify the object that needs to be automatically scaled, which can be a controller object such as Deployment, ReplicaSet or StatefulSet. By specifying scaleTargetRef, HPA can monitor the resource utilization of the object and dynamically adjust the number of its replicas according to preset conditions to meet the load requirements. In k8s, Pod.spec.selector is the part used to define the label selector, which is used to identify which Pods belong to a specific controller (such as ReplicaSet, Deployment, etc.). This field is usually used in the configuration of controller objects (such as Deployment).

[0041] The specific implementation of HPA is to scale resources by creating HPA resources, but the HPA provided in k8s needs to create HPA resources manually, specifically: write a yaml file, specify the scaling object through the scaleTargetRef field in the yaml file, spec.minReplicas / maxReplicas define the minimum and maximum number of replicas, metrics.Resource specifies the indicator for resource monitoring, such as CPU utilization or memory utilization, and metrics.Pods is used to set Pod resources, such as packets-per-second, which is used to set the number of packets received per second. When the Pod's indicator value does not reach the specified target value, HPA will reduce the number of Pods for the target workload, and when the indicator value exceeds the specified target value, it will increase the number of Pods for the target workload.

[0042] However, the HPA mentioned above uses CPU utilization and memory utilization as metrics for elastic scaling, which requires more detailed settings in a real high-traffic environment and is more dependent on external resource indicators. The HPA scaling of k8s is difficult to solve this problem. The current HPA mechanism of k8s has the following limitations:

[0043] 1. Manual operation is required. When the cluster is large and the number of workloads is large, HPA resources need to be manually created for each workload, which is inefficient.

[0044] 2. HPA uses a simple indicator ratio to calculate the number of target replicas for scaling. This is only applicable to ideal scenarios where the number of application replicas and related indicators are strictly linearly correlated. However, in actual production, there is a complex relationship between various application indicators and the number of replicas. This method is difficult to obtain the number of replicas that meets the capacity requirements. In addition, HPA defines fewer indicators, and the method for selecting Pods to be deleted is relatively simple, which is difficult to meet the needs of workload scaling in various situations in actual scenarios.

[0045] Based on the above problems, the present application proposes a method for scaling workloads, by obtaining at least one indicator associated with the target workload in the cluster, namely: a first indicator set by the user, and / or a second indicator corresponding to the cluster, and based on at least one indicator, creating HPA resources corresponding to the target workload, and expanding or shrinking the target workload through HPA resources and at least one indicator. The corresponding HPA resources can be automatically created for the target workload to achieve automatic expansion or shrinking, avoiding the tedious manual operation, saving time and improving efficiency, and users can customize indicators to improve the accuracy of expansion or shrinking.

[0046] In some embodiments, the method described in the embodiments of the present application can be applied to the automatic expansion and contraction scenario of workload. In some embodiments, the method described in the embodiments of the present application can also be applied to application systems such as container services, container image services, and cloud native platforms. In some embodiments, the method described in the embodiments of the present application can also be applied to cloud computing, big data edge computing, education, medical care, social networking, entertainment, online courses, office and other fields.

[0047] The following describes an exemplary application of the electronic device provided in the embodiment of the present application. The electronic device provided in the embodiment of the present application can be a rechargeable device such as a laptop, a tablet computer, a desktop computer, a mobile device (e.g., a mobile phone, a wearable smart watch, a dedicated messaging device), but is not limited thereto. Alternatively, the electronic device can also be implemented as a server.

[0048] In some embodiments, the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms, but is not limited thereto. In some embodiments, the server and the electronic device may be directly or indirectly connected via wired or wireless communication, and this is not specifically limited in the embodiments of the present application.

[0049] See also Figure 1 , Figure 1 is a structural schematic diagram of an electronic device provided in an embodiment of the present application, Figure 1The electronic device 100 shown includes: at least one processor 110, a memory 150, at least one network interface 120 and a user interface 130. The various components in the electronic device 100 are coupled together via a bus system 140. It is understood that the bus system 140 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 140 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 140 is not described in detail. Figure 1 Various buses are labeled as bus system 140 .

[0050] The processor 110 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0051] The user interface 130 includes one or more output devices 131 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 130 also includes one or more input devices 132, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0052] The memory 150 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 150 may optionally include one or more storage devices that are physically remote from the processor 110.

[0053] The memory 150 includes a volatile memory or a nonvolatile memory, and may also include both volatile and nonvolatile memories. The nonvolatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 150 described in the embodiments of the present application is intended to include any suitable type of memory.

[0054] In some embodiments, the memory 150 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.

[0055] The operating system 151 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic businesses and process hardware-based tasks.

[0056] The network communication module 152 is used to reach other computing devices via one or more (wired or wireless) network interfaces 120. Exemplary network interfaces 120 include: Bluetooth, WiFi, and universal serial bus (USB).

[0057] The presentation module 153 is used to enable presentation of information (eg, a user interface for operating peripheral devices and displaying content and information) via one or more output devices 131 (eg, a display screen, a speaker, etc.) associated with the user interface 130 .

[0058] The input processing module 154 is configured to detect one or more user inputs or interactions from one of the one or more input devices 132 and to translate the detected inputs or interactions.

[0059] In some embodiments, the workload expansion and contraction method provided in the embodiments of the present application can be implemented in software and stored in the memory 150. Figure 2 , Figure 2 This is a structural diagram of a workload expansion and contraction device provided in an embodiment of the present application, which can be software in the form of programs and plug-ins, etc. The workload expansion and contraction device 155 includes the following software modules: an acquisition module 1551, a creation module 1552, and an expansion and contraction module 1553. These modules are logical, so they can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.

[0060] In other embodiments, the workload expansion and contraction device 155 provided in the embodiment of the present application can be implemented in hardware. As an example, the workload expansion and contraction device 155 provided in the embodiment of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the workload expansion and contraction method provided in the embodiment of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application specific integrated circuits (application specific integrated circuit, ASIC), DSP, programmable logic device (programmable logic device, PLD), complex programmable logic device (complex programmable logic device, CPLD), field programmable gate array (field-programmable gate array, FPGA) or other electronic components.

[0061] The following will explain the method for scaling workload provided by the embodiment of the present application in combination with the exemplary application and implementation of the electronic device provided by the embodiment of the present application.

[0062] In some embodiments, the workload scaling method provided in the embodiments of the present application can be applied to a workload scaling system, and the workload scaling system can be used to implement a method for automatic workload scaling.

[0063] See also Figure 3 , Figure 3 It is a structural diagram of the expansion and contraction system of the workload provided in the embodiment of the present application. The expansion and contraction system 300 of the workload can be deployed in a k8s cluster (k8s cluster), which may include: an automatic expansion and contraction controller (autoscale-controller) 301, a task processing component (jobHandler) 302, a work queue component (workqueue) 303, a task execution component (processHandle) 304, a k8s application programming interface server (k8sAPI Server) 305, an HPA controller (HPA-controller) 306, task 1 (job1), task 2 (job2), application 1 (app1) and application 2 (app2). The following describes each component included in the expansion and contraction system 300 of the workload.

[0064] 1. autoscale-controller 301

[0065] Based on the design concept of k8s controller, autoscale-controller 301 is designed, which can create corresponding HPA resources for the workload of Deployment, collect the configuration information of Deployment workload, initialize the deletion cost value of Pod in Pod annotation (such as controller.kubernetes.io / pod-deletion-cost value), and set its value to zero.

[0066] As a core component of the workload expansion and contraction method provided in the embodiment of the present application, the autoscale-controller 301 belongs to the control part and has the following functions: 1. Through the list-watch mechanism of the k8s API Server 305, it automatically monitors the creation of the Deployment workload in the k8s cluster, and obtains the first indicator set by the user and the second indicator corresponding to the k8s cluster (also called the HPA indicator) through its own interface, and can also obtain the information corresponding to each application; 2. Create tasks, specifically there are three types of tasks: the first is to create tasks (such as Figure 3 The first type of task is to create an HPA resource based on the acquired indicator. The second type is to update the task (such as Figure 3 The third type is a deletion task, that is, determining the Pod replica to be deleted. Then, the task to be executed is sent to jobHandler 302.

[0067] Here, list-watch is a resource monitoring mechanism provided by k8s API Server 305, which allows the client to obtain a list of resource objects (List) through k8s API Server 305, and monitor (Watch) changes in resource objects in real time by establishing a long connection. When a resource object is created, updated, or deleted, k8s API Server 305 will send event notifications to clients that have subscribed to the relevant resources to achieve real-time synchronization and response of resources.

[0068] In some embodiments, part of the code in the Deployment configuration file may be as follows:

[0069]

[0070]

[0071] It should be noted that the above codes are only used for illustrative purposes and are not specifically limited in the embodiments of the present application.

[0072] 2. jobHandler 302

[0073] The jobHandler 302 is mainly used to receive messages from the autoscale-controller 301, convert the received messages into a format that can be processed and recognized by the workqueue 303, and send the processed messages to the workqueue 303. There are three main types of messages received here: the first is the message of creating a task; the second is the message of updating a task; and the third is the message of deleting a task.

[0074] 3. workqueue 303

[0075] Workqueue 303 is a message queue mainly used to receive messages from jobHandle 302 and store the received messages persistently in a local cache. It can separate the delivery and processing of tasks and add identification information (such as identity (ID)) to each task to distinguish each task.

[0076] 4. processHandle 304

[0077] As the executor of the task, processHandle 304 receives and processes the message from workqueue 303, and starts an infinite loop function. The specific working process of the function is: obtain the message from workqueue 303 for processing and print the message. For example, according to the ID of the task, the detailed information of the task to be executed (such as the name of the workload, the namespace, the first indicator set by the user, etc.) is obtained from the local cache, and different task types are judged to execute the task. There are three types of tasks here. The first is a creation task. ProcessHandle 304 can create HPA resources corresponding to the workload based on the first indicator set by the user. If the user has not set the first indicator, HPA resources can be created according to the second indicator corresponding to the k8s cluster, namely, CPU utilization and memory utilization. The second is an update task. According to the obtained indicator value, the parameter value corresponding to the indicator is determined, and then the deletion cost value (pod-deletion-cost value) of each Pod is calculated, and the k8s API interface is called to update the annotation of the Pod in the workload. The third is a deletion task. When the user directly performs a scaling operation through the autoscale-controller 301 (for example, the user deletes a pod in the operation interface of the workload scaling system 300, or triggers the scaling button, etc.), the Pod copy to be deleted determined by the autoscale-controller 301 is received and deleted.

[0078] In some embodiments, part of the code for executing the task in processHandle 304 may be as follows:

[0079]

[0080] It should be noted that the above codes are only used for illustrative purposes and are not specifically limited in the embodiments of the present application.

[0081] 5. k8s API Server 305

[0082] The k8s API Server 305 provides a hypertext transfer protocol representational state transfer (HTTP REST) ​​interface for adding, deleting, modifying, checking, and monitoring various k8s resource objects (such as Pod, RC, Service, etc.). It is the data bus and data center of the entire workload expansion and contraction system 300. The creation and update of workloads are implemented by the k8s API Server 305, which can be sent to different applications (such as Figure 3 Application 1 and Application 2) shown in provide interfaces.

[0083] 6. HPA-controller 306

[0084] HPA-controller 306 is mainly used to call its interface to implement the creation of HPA resources, and it can also interact with various applications.

[0085] It should be noted that: Figure 3 The components included in the system shown in FIG. 1 are for illustrative purposes only and are not intended to be limiting.

[0086] The following describes the workload scaling method provided in the embodiment of the present application in combination with the above workload scaling system.

[0087] Figure 4 1 is a schematic diagram of an implementation flow of a method for scaling workloads provided in an embodiment of the present application. Figure 4 As shown, the workload scaling method of the embodiment of the present application may include steps S410 to S430.

[0088] In step S410, at least one indicator associated with a target workload in a cluster is obtained.

[0089] The at least one indicator includes a first indicator set by a user and / or a second indicator corresponding to the cluster.

[0090] Here, the cluster can be understood as a distributed system deployed in k8s, which can be called a k8s cluster. The cluster can be composed of a group of nodes, which can be physical servers or virtual machines. The cluster includes a collection of computing, storage and network resources. K8s can use these resources to run various container-based applications. Workload can be understood as various application programs and service instances running on the k8s cluster. It is the actual carrier for accessing services and the actual operating carrier for system applications such as node log collection and monitoring. The target workload can be understood as the workload that needs to be expanded or reduced in the Deployment, for example, the workload newly created in the Deployment. Expansion can be to increase the number of Pod copies, and reduction can be understood as reducing the number of Pod copies. The first indicator can be an indicator related to the expansion or reduction of the target workload set by the user, for example, the Pod runtime weight, whether the Pod runtime is positively correlated, the Pod storage usage weight, whether the Pod storage usage is positively correlated, etc., and the embodiment of the present disclosure does not specifically limit this. The second indicator can be CPU utilization and memory utilization. Pod can be understood as the smallest resource management component in k8s. It is the resource object that minimizes the running of containerized applications.

[0091] It is understandable that before expanding or shrinking the target workload, it is necessary to determine the number of Pod copies to be increased or decreased, and the number of Pod copies to be increased or decreased is related to the indicators associated with the target workload. Therefore, it is necessary to obtain at least one indicator associated with the target workload in the cluster, such as the first indicator and / or the second indicator. For example, at least one of the above indicators can be obtained through the interface of the autoscale-controller itself.

[0092] In some embodiments, the number of target workloads may be one or more, which is not specifically limited in the embodiments of the present application.

[0093] In one example, if the business in the cluster is a world wide web (web) business, since the database business focuses more on the number of connections, its corresponding expansion and contraction target may be the number of concurrent business processing. Based on this, these indicators related to the number of concurrent business processing can be used as the first indicator.

[0094] In some embodiments, the user may also set a template value of the first indicator for services of the same type, thereby saving time and avoiding the need to set the first indicator for each target workload.

[0095] In step S420, HPA resources corresponding to the target workload are created based on at least one indicator.

[0096] Here, HPA resources are used to achieve automatic expansion or reduction of the target workload.

[0097] It can be understood that after obtaining at least one indicator, the autoscale-controller can automatically create corresponding HPA resources for the target workload based on the at least one indicator.

[0098] In some embodiments, the autoscale-controller may also monitor the creation and update events of the target workload in the cluster through a list-watch mechanism, and based on the events, automatically create corresponding HPA resources for the target workload.

[0099] In step S430, the target workload is expanded or reduced in capacity through HPA resources and at least one indicator.

[0100] It can be understood that after the HPA resources are created, at least one indicator can be used to determine the number of Pod copies to be deleted or to be added. Then, the target workload can be automatically expanded or reduced through the HPA resources, the number of Pod copies to be deleted or to be added.

[0101] In some embodiments, in the case of capacity expansion, N Pod copies can be automatically added to the target load through HPA resources and the number of Pod copies that need to be added (N, N is a positive integer), thereby avoiding manual operations by users.

[0102] In some embodiments, in the case of scaling down, the Pod copies to be deleted can be automatically deleted through HPA resources and the Pod copies to be deleted, thereby avoiding manual operations by users.

[0103] In an embodiment of the present application, through the above steps S410 to S430, corresponding HPA resources can be automatically created for the target workload to achieve automatic expansion or reduction, avoiding the tedious manual operation, saving time and improving efficiency, and users can customize indicators to improve the accuracy of expansion or reduction.

[0104] In some embodiments, the above step S410 may include: filtering the namespace based on preset fields to obtain a creation event of the target workload in the cluster; and in response to the creation event, obtaining at least one indicator through an interface of the target controller in the cluster.

[0105] Here, the preset field may be an AutoScale field, and the target controller may be an autoscale-controller.

[0106] It can be understood that by filtering in the namespace through the preset field, the creation event of the target workload in the cluster associated with the preset field can be determined and the creation information of the creation event can be obtained. After monitoring the creation event, in response to the event, at least one indicator can be obtained through the interface of the autoscale-controller in the cluster. In this way, the monitoring scope of the autosace-controller can be reduced and its efficiency can be improved through the automatic expansion and contraction field, without affecting the business in other namespaces in the cluster.

[0107] In some embodiments, the original namespace object in k8s can be modified in advance to add an AutoScale field. For example, the method in the namespace object is modified. In the Create method and the Apply method, the AutoScale field is added when creating the namespace. After adding the AutoScale field, the List function can filter the namespace through AutoScale.

[0108] In some embodiments, the portion of code after adding the AutoScale field in the namespace may be as follows:

[0109] type Namespace struct{ / / Namespace structure

[0110] metav1.TypeMeta'json:",inline"'

[0111] metav1.ObjectMeta'json:"metadata,omitempty"protobuf:"bytes,1,opt,name=metadata"'

[0112] Spec NamespaceSpec'json:"spec,omitempty"protobuf:"bytes,2,opt,name=spec"' / / Namespace configuration specification

[0113] Status NamespaceStatus'json:"status,omitempty"protobuf:"bytes,3,opt,name=status"' / / Namespace status information

[0114] AutoScale bool'json:"autoScale,omitempty"protobuf:"byte,2,opt,name=autoscale"' / / Automatic expansion and contraction field Boolean type

[0115] }

[0116] In the above code, metav1.TypeMeta is a structure used to represent the metadata of an object; TypeMeta and ObjectMeta are two fields contained in metav1.TypeMeta, TypeMeta is used to store the type information of the object, and ObjectMeta is used to store the metadata information of the object; json:",inline" in metav1.TypeMeta means that this field is in JavaScript Object Notation (JavaScript Object Notation, JSON) will be inlined into the parent object during serialization instead of appearing as an independent JSON object; json:"xx,omitempty" means that during JSON serialization, if the xx field is empty, it is not included in the serialized JSON string. For example, json:"metadata,omitempty" means that during JSON serialization, if the metadata field is empty, the field is not included in the serialized JSON string; protobuf:"bytes,1,opt,name=xx" means that during protobuf serialization, the data type of the xx field is bytes, the tag is 1, which is optional, and the field name is xx. For example, protobuf:"bytes,1,opt,name=metadata" means that during protobuf serialization, the data type of the metadata field is bytes, the tag is 1, which is optional, and the field name is metadata.

[0117] It should be noted that the above codes are only used for illustrative purposes and are not specifically limited in the embodiments of the present application.

[0118] In some embodiments, the scaling down in the above step S430 may include: step 1, determining the deletion cost of each Pod included in the target workload based on at least one indicator; step 2, determining the Pod replica to be deleted based on the deletion cost and the deletion policy corresponding to the target workload; step 3, deleting the Pod replica to be deleted through HPA resources.

[0119] Among them, the deletion strategy can be understood as a strategy for how to screen the Pod copies to be deleted, which can be set by the autoscale-controller in the workload scaling system, or can be customized by the user, or can be determined based on the actual working conditions of the Pod. The embodiments of the present application do not make specific limitations on this.

[0120] It can be understood that based on at least one indicator associated with the target workload in the cluster, that is, the first indicator and / or the second indicator, the deletion cost of each Pod included in the target workload can be calculated by weighted summation or other algorithms. Then, according to the deletion cost of each Pod and the deletion policy corresponding to the target workload, the Pod copy to be deleted can be determined, and finally, the Pod copy to be deleted is automatically deleted through the created HPA resource. In this way, the Pod copy to be deleted can be automatically determined and deleted in the above manner. Compared with the manual operation required in the related technology, the above process is more efficient and fast, saves manpower, and can improve the user experience.

[0121] In some embodiments, at least one indicator includes a first indicator and a second indicator; the above step 1 may include: determining a first weighted value based on a first sub-indicator and a second indicator associated with the second indicator in the first indicator; determining a second weighted value based on a second sub-indicator associated with each Pod in the first indicator and a parameter value corresponding to the second sub-indicator; and determining a deletion cost of each Pod included in the target workload based on the first weighted value and the second weighted value.

[0122] Here, in the process of shrinking, the first indicator may include the first sub-indicator associated with the second indicator, the second sub-indicator associated with each Pod, and other third sub-indicators associated with shrinking set by the user, etc., which are not specifically limited in the embodiment of the present disclosure. The first sub-indicator may include indicators related to CPU utilization and memory utilization, such as: CPU utilization weight, whether CPU utilization is positively correlated, memory utilization weight and whether memory utilization is positively correlated, etc., which are not specifically limited in the embodiment of the present disclosure. The second sub-indicator may include Pod running time weight, whether Pod running time is positively correlated, Pod storage usage weight, whether Pod storage usage is positively correlated, Pod restart number weight and whether Pod restart number is positively correlated, etc., which are not specifically limited in the embodiment of the present disclosure. The above-mentioned positive correlation indicator can be used to determine whether the weighted values ​​of each indicator are summed or subtracted. If they are positively correlated, they are summed; if they are negatively correlated, they are subtracted. The third sub-indicator may include network traffic, shrinking operation execution cycle, etc., which are not specifically limited in the embodiment of the present disclosure.

[0123] It can be understood that in the case where at least one indicator includes a first indicator and a second indicator, the first weighted value is obtained by weighted summation according to the first sub-indicator and the second indicator associated with the second indicator in the first indicator, through the correspondence between the first sub-indicator and the second indicator. Then, the second weighted value is obtained by weighted summation according to the second sub-indicator associated with each Pod in the first indicator and the parameter value corresponding to the second sub-indicator. Finally, based on the first weighted value and the second weighted value, the two are summed to obtain the deletion cost of each Pod included in the target workload. In this way, the deletion cost of each Pod is determined by the above method, and the parameters based on it are more, which can meet the needs of user-defined control of workload scaling, reduce the tediousness of manual operations, improve the utilization of hardware resources, and better meet the scaling requirements in actual application scenarios.

[0124] In an example, if the first sub-indicator includes a CPU utilization weight (its value is recorded as a1), whether the CPU utilization is positively correlated (its value is positively correlated), a memory utilization weight (its value is recorded as a2), and whether the memory utilization is positively correlated (its value is positively correlated), and the second indicator includes CPU utilization (its value is recorded as b1) and memory utilization (its value is recorded as b2), then the first weighted value (recorded as cost1) can be obtained by the following formula (1):

[0125] cost1=a1*b1+a2*b2 (1)

[0126] In one example, if the second sub-indicator includes the Pod running time weight (its value is recorded as c1), whether the Pod running time is positively correlated (its value is positively correlated), the Pod storage usage weight (its value is recorded as c2), and whether the Pod storage usage is positively correlated (its value is positively correlated), the parameter corresponding to the Pod running time weight is the Pod running time, its value is recorded as d1, and the parameter corresponding to the Pod storage usage weight is the Pod storage usage, its value is recorded as d2, then the second weighted value (recorded as cost2) can be obtained by the following formula (2):

[0127] cost2=c1*d1+c2*d2 (2)

[0128] In some embodiments, the first indicator can be set as follows:

[0129]

[0130] Here, int represents integer type; bool represents Boolean type; string represents string type.

[0131] It should be noted that the above codes are only used for illustrative purposes and are not specifically limited in the embodiments of the present application.

[0132] In some embodiments, the above step 2 may include: step a, obtaining the usage status of each Pod included in the target workload; when the usage status is a non-running state, determining that the Pod copy corresponding to the non-running Pod is the Pod copy to be deleted; step b, when the usage status is a running state, determining the Pod copy to be deleted based on the target parameters associated with the running Pod and the deletion cost corresponding to the running Pod.

[0133] The target parameters include at least one of the readiness of the running Pod, the number of Pod replicas in the node, the preparation time, the number of container restarts, and the creation time. The non-running state may include: unscheduled state, pending state, unknown state, etc., which is not specifically limited in the embodiments of the present application.

[0134] It is understandable that when a Pod is in a running state, usually the copy of the Pod will not be determined as a Pod to be deleted, because it may be performing certain operations. Therefore, the usage state of each Pod is related to whether its copy can be determined as a Pod copy to be deleted. Based on this, the usage state of each Pod included in the target workload is obtained, and it is determined whether the usage state of each Pod is a non-running state. When the usage state is a non-running state, the Pod copy corresponding to the non-running Pod is determined as a Pod copy to be deleted. Specifically, the Pod copy corresponding to the unscheduled Pod can be first determined as a Pod copy to be deleted, and then the Pod copy corresponding to the Pod in the suspended state can be determined as a Pod copy to be deleted, and finally the Pod copy corresponding to the Pod in the unknown state can be determined as a Pod copy to be deleted.

[0135] In some embodiments, which Pod replicas corresponding to the Pods in a non-running state are determined as Pod replicas to be deleted may depend on the number of Pod replicas to be deleted, which is not specifically limited in the embodiments of the present application.

[0136] In some embodiments, for step b, when the usage status is Running, based on the deletion cost value corresponding to the Pod in the Running status and the target parameters associated with the Pod in the Running status, namely: the readiness status corresponding to the Pod in the Running status, the number of Pod copies in the node, the preparation time, the number of container restarts and the creation time, at least one of the actual working conditions corresponding to each Pod are compared to determine the Pod copies to be deleted so as not to affect the normal operation of the Pod.

[0137] In an embodiment of the present application, the Pod copies to be deleted are determined through the above steps a to b, which fully considers the usage status of each Pod, so that the determined Pod copies to be deleted are more in line with actual needs and will not affect the normal use of the Pod.

[0138] In some embodiments, the above step b may include: when the readiness status corresponding to the Pod in the running state is not ready, determining that the replica corresponding to the first Pod is a Pod replica to be deleted.

[0139] The first Pod is the Pod with the smallest deletion cost among the unready Pods. The smaller the deletion cost, the lower the actual impact of deleting the replica corresponding to the Pod on the workload.

[0140] It can be understood that when the readiness status of the Pod in the running state is unready, the deletion cost of each unready Pod is compared, and the copy corresponding to the Pod with the smallest deletion cost among the unready Pods is determined as the Pod copy to be deleted. In this way, the effect of minimizing the deletion cost and minimizing the loss can be achieved.

[0141] In some embodiments, the above step b may also include: when the readiness status of the Pod corresponding to the running state is prepared, and the deletion cost values ​​corresponding to each prepared Pod are the same, perform one of the following: when the number of Pod copies in the nodes where at least two Pods in the prepared Pods are located is different, determine that the copy corresponding to the second Pod is the Pod copy to be deleted; when the number of Pod copies in the node where each Pod in the prepared Pods is located is the same, and the preparation time corresponding to at least two Pods is different, determine that the copy corresponding to the third Pod is the Pod copy to be deleted; when the number of Pod copies in the node where each Pod in the prepared Pods is located is the same, and the preparation time corresponding to each Pod is the same, determine that the copy corresponding to the fourth Pod is the Pod copy to be deleted; when the number of Pod copies in the node where each Pod in the prepared Pods is located is the same, the preparation time corresponding to each Pod is the same, and the number of container restarts of each Pod is the same, determine that the copy corresponding to the fifth Pod is the Pod copy to be deleted.

[0142] Among them, the second Pod is the Pod with the largest number of Pod replicas in the node where it is located among the prepared Pods. The third Pod is the Pod with the longest preparation time among the prepared Pods. The fourth Pod is the Pod with the largest number of container restarts among the prepared Pods. The fifth Pod is the Pod with the latest creation time among the prepared Pods.

[0143] It can be understood that when the readiness status of the running Pod is ready, and the deletion cost values ​​corresponding to the prepared Pods are the same, it is necessary to determine the Pod copy to be deleted through at least one of the parameters such as the number of Pod copies in the node corresponding to the running Pod, the preparation time, the number of container restarts, and the creation time.

[0144] Specifically, when the number of Pod copies in the nodes where at least two Pods in the prepared Pods are located is different, the number of Pod copies in the nodes where the Pods are located can be compared, and the copy corresponding to the Pod with the largest number of Pod copies in the node where the prepared Pods are located is determined as the Pod copy to be deleted. Then, when the number of Pod copies in the nodes where each Pod in the prepared Pods is located is the same, and the preparation time corresponding to at least two Pods is different, the preparation time corresponding to each Pod can be compared, and the copy corresponding to the Pod with the longest preparation time in the prepared Pods is determined as the Pod copy to be deleted, because the deletion loss of this copy is relatively small compared to other copies. Then, when the number of Pod copies in the nodes where each Pod in the prepared Pods is located is the same, and the preparation time corresponding to each Pod is the same, the number of container restarts in each Pod is compared, and the copy corresponding to the Pod with the largest number of container restarts in the prepared Pod is determined as the Pod copy to be deleted, because the more restarts, the lower the usefulness of the Pod in the actual utilization process. Finally, when the number of Pod replicas in the nodes where each Pod is located is the same among the prepared Pods, the preparation time corresponding to each Pod is the same, and the number of container restarts for each Pod is the same, the creation time of each Pod is compared, and the replica corresponding to the Pod with the latest creation time among the prepared Pods is determined as the Pod replica to be deleted, because the later the Pod is created, the lower its usage frequency may be, and the loss of deleting its replica is relatively small compared to other replicas. In this way, determining the Pod replica to be deleted through the above process fully considers the usage and use value of each Pod, and the selected Pod replica to be deleted is more in line with actual needs.

[0145] In some embodiments, the capacity expansion in the above step S430 may include: determining the number N of Pod replicas to be added corresponding to the target workload through at least one indicator, where N is a positive integer; and adding N Pod replicas through HPA resources and the number N of Pod replicas to be added.

[0146] Here, at least one indicator can be understood as an indicator related to the expansion of workload, such as CPU utilization, memory utilization, etc., and the embodiments of the present application do not make specific limitations on this.

[0147] It can be understood that during the expansion process, the autoscale-controller can directly determine the number N of Pod replicas to be added corresponding to the target workload through at least one indicator, and then automatically add N Pod replicas for the target workload through HPA resources. In this way, expansion through the above method is simple, efficient, easy to implement, and does not require manual operation by the user.

[0148] Figure 5 FIG. 1 is another implementation flow diagram of the workload expansion and contraction method provided in the embodiment of the present application. Figure 5 As shown, the embodiment of the present application mainly introduces the process of reducing the workload, which may specifically include steps S510 to S550.

[0149] In step S510, at least one indicator associated with a target workload in a cluster is obtained.

[0150] In step S520, HPA resources corresponding to the target workload are created based on at least one indicator.

[0151] In step S530 , based on at least one indicator, a deletion cost value of each Pod included in the target workload is determined.

[0152] In step S540, the Pod replica to be deleted is determined based on the deletion cost value and the deletion policy corresponding to the target workload.

[0153] In step S550, the Pod replica to be deleted is deleted through the HPA resource.

[0154] Figure 6 FIG. 1 is another implementation flow diagram of the workload expansion and contraction method provided in the embodiment of the present application. Figure 6 As shown, the embodiment of the present application mainly introduces the workload expansion process, which may specifically include steps S610 to S640.

[0155] In step S610, at least one indicator associated with a target workload in a cluster is obtained.

[0156] In step S620, HPA resources corresponding to the target workload are created based on at least one indicator.

[0157] In step S630, the number N of Pod replicas to be added corresponding to the target workload is determined through at least one indicator.

[0158] In step S640, N Pod replicas are added using HPA resources and the number N of Pod replicas to be added.

[0159] Based on the same inventive concept, the embodiment of the present application also provides a workload expansion and contraction device, such as the workload expansion and contraction device 155 in the above embodiment. Figure 2 As shown, the workload scaling device 155 includes: an acquisition module 1551, used to acquire at least one indicator associated with the target workload in the cluster, wherein the at least one indicator includes a first indicator set by the user, and / or a second indicator corresponding to the cluster; a creation module 1552, used to create HPA resources corresponding to the target workload based on at least one indicator; a scaling module 1553, used to expand or shrink the target workload through HPA resources and at least one indicator.

[0160] In some embodiments, the scaling module 1553 includes: a first determination submodule, used to determine the deletion cost of each Pod included in the target workload based on at least one indicator; a second determination submodule, used to determine the Pod replica to be deleted based on the deletion cost and the deletion policy corresponding to the target workload; a deletion submodule, used to delete the Pod replica to be deleted through HPA resources.

[0161] In some embodiments, at least one indicator includes a first indicator and a second indicator; the first determination submodule is further used to perform the following steps: determine a first weighted value based on the first sub-indicator and the second indicator associated with the second indicator in the first indicator; determine a second weighted value based on the second sub-indicator associated with each Pod in the first indicator and the parameter value corresponding to the second sub-indicator; determine the deletion cost of each Pod included in the target workload based on the first weighted value and the second weighted value.

[0162] In some embodiments, the second determination submodule includes: an acquisition unit, used to: acquire the usage status of each Pod included in the target workload; a first determination subunit, used to determine, when the usage status is a non-running state, that the Pod copy corresponding to the non-running Pod is the Pod copy to be deleted; a second determination subunit, used to determine, when the usage status is a running state, the Pod copy to be deleted based on the target parameters associated with the running Pod and the deletion cost value corresponding to the running Pod, wherein the target parameters include at least one of the readiness status corresponding to the running Pod, the number of Pod copies in the node, the preparation time, the number of container restarts, and the creation time.

[0163] In some embodiments, the second determination subunit is used to: when the readiness status corresponding to the Pod in the running state is not ready, determine that the copy corresponding to the first Pod is the Pod copy to be deleted, wherein the first Pod is the Pod with the smallest deletion cost among the unready Pods.

[0164] In some embodiments, the second determination subunit is further used to: when the readiness status of the Pod corresponding to the running state is prepared and the deletion cost values ​​corresponding to the prepared Pods are the same, perform one of the following: when the number of Pod copies in the nodes where at least two Pods in the prepared Pods are located is different, determine that the copy corresponding to the second Pod is the Pod copy to be deleted, wherein the second Pod is the Pod with the largest number of Pod copies in the node where it is located among the prepared Pods; when the number of Pod copies in the node where each Pod in the prepared Pods is located is the same, and the preparation times of at least two Pods are different, determine that the copy corresponding to the third Pod is the Pod copy to be deleted, wherein , the third Pod is the Pod with the longest preparation time among the prepared Pods; when the number of Pod copies in the node where each Pod in the prepared Pods is located is the same and the preparation time corresponding to each Pod is the same, the copy corresponding to the fourth Pod is determined to be the Pod copy to be deleted, wherein the fourth Pod is the Pod with the largest number of container restarts among the prepared Pods; when the number of Pod copies in the node where each Pod in the prepared Pods is located is the same, the preparation time corresponding to each Pod is the same, and the number of container restarts for each Pod is the same, the copy corresponding to the fifth Pod is determined to be the Pod copy to be deleted, wherein the fifth Pod is the Pod with the latest creation time among the prepared Pods.

[0165] In some embodiments, the acquisition module 1551 is used to: filter the namespace based on preset fields to obtain creation events of target workloads in the cluster; and obtain at least one indicator through an interface of a target controller in the cluster in response to the creation event.

[0166] In some embodiments, the scaling module 1553 is also used to perform the following steps: determine the number N of Pod replicas to be added corresponding to the target workload through at least one indicator, where N is a positive integer; and add N Pod replicas through HPA resources and the number N of Pod replicas to be added.

[0167] The description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. In some embodiments, the functions or modules included in the device provided in the embodiment of the present application can be used to execute the method described in the above method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.

[0168] It should be noted that in the embodiment of the present application, if the above-mentioned workload expansion and contraction method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium, including a number of instructions to enable an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific hardware, software or firmware, or any combination of hardware, software, and firmware.

[0169] An embodiment of the present application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.

[0170] The embodiment of the present application provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, some or all of the steps in the above method are implemented. The computer-readable storage medium can be transient or non-transient.

[0171] An embodiment of the present application provides a computer program, including a computer-readable code. When the computer-readable code is run in an electronic device, a processor in the electronic device executes some or all of the steps for implementing the above method.

[0172] The present application embodiment provides a computer program product, including a computer program or an instruction, which implements some or all of the steps in the above method when the computer program or instruction is executed by a processor. The computer program product can be implemented specifically by hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium, and in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0173] It should be noted here that the description of the various embodiments above tends to emphasize the differences between the various embodiments, and the same or similar aspects can be referenced to each other. The description of the above device, storage medium, computer program and computer program product embodiments is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the embodiments of the device, storage medium, computer program and computer program product of this application, please refer to the description of the method embodiment of this application for understanding.

[0174] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the serial number of each step / process mentioned above does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. The serial numbers of the embodiments of the present application mentioned above are for description only and do not represent the advantages and disadvantages of the embodiments.

[0175] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0176] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0177] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0178] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0179] Those skilled in the art will appreciate that all or part of the steps of the above method embodiments may be implemented by hardware associated with program instructions, and the aforementioned program may be stored in a computer-readable storage medium, which, when executed, executes the steps of the above method embodiments; and the aforementioned storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROM), magnetic disks or optical disks.

[0180] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0181] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.

Claims

1. A method for scaling workload, characterized in that: include: Acquire at least one indicator associated with a target workload in a cluster, wherein the at least one indicator includes a first indicator set by a user and / or a second indicator corresponding to the cluster; Based on the at least one indicator, create Pod horizontal automatic scaling HPA resources corresponding to the target workload; The target workload is expanded or reduced in capacity through the HPA resources and the at least one indicator.

2. The method according to claim 1, characterized in that The scaling down the target workload by using the HPA resources and the at least one indicator includes: Determine, based on the at least one indicator, a deletion cost of each Pod included in the target workload; Determine the Pod replica to be deleted based on the deletion cost value and the deletion policy corresponding to the target workload; The Pod copy to be deleted is deleted through the HPA resource.

3. The method according to claim 2, characterized in that The at least one indicator includes a first indicator and a second indicator; The determining, based on the at least one indicator, a deletion cost value of each Pod included in the target workload includes: Determining a first weighted value according to a first sub-indicator associated with the second indicator in the first indicator and the second indicator; Determine a second weighted value according to a second sub-indicator associated with each Pod in the first indicator and a parameter value corresponding to the second sub-indicator; Based on the first weighted value and the second weighted value, a deletion cost value of each Pod included in the target workload is determined.

4. The method according to claim 2, characterized in that: The determining the Pod replica to be deleted based on the deletion cost value and the deletion policy corresponding to the target workload includes: Obtaining the usage status of each Pod included in the target workload; In the case where the usage state is a non-running state, determining that the Pod copy corresponding to the Pod in the non-running state is a Pod copy to be deleted; When the usage status is the running status, the Pod copy to be deleted is determined according to the target parameters associated with the Pod in the running status and the deletion cost value corresponding to the Pod in the running status, wherein the target parameters include at least one of the preparation status corresponding to the Pod in the running status, the number of Pod copies in the node, the preparation time, the number of container restart times, and the creation time.

5. The method according to claim 4, characterized in that The determining the Pod replica to be deleted according to the target parameter associated with the running Pod and the deletion cost value corresponding to the running Pod includes: When the readiness status corresponding to the Pod in the running state is not ready, it is determined that the replica corresponding to the first Pod is the Pod replica to be deleted, wherein the first Pod is the Pod with the smallest deletion cost among the unready Pods.

6. The method according to claim 4, characterized in that The method further comprises: When the readiness status of the Pod in the running state is Ready, and the deletion cost values ​​of the prepared Pods are the same, perform one of the following: In a case where the number of Pod replicas in the nodes where at least two of the prepared Pods are located is different, determining that the replica corresponding to the second Pod is the Pod replica to be deleted, wherein the second Pod is the Pod with the largest number of Pod replicas in the node where it is located among the prepared Pods; When the number of Pod replicas in the node where each Pod in the prepared Pod is located is the same and the preparation time corresponding to at least two Pods is different, determining that the replica corresponding to the third Pod is the Pod replica to be deleted, wherein the third Pod is the Pod with the longest preparation time among the prepared Pods; When the number of Pod replicas in the node where each Pod in the prepared Pod is located is the same and the preparation time corresponding to each Pod is the same, determine that the replica corresponding to the fourth Pod is the Pod replica to be deleted, wherein the fourth Pod is the Pod with the largest number of container restarts among the prepared Pods; When the number of Pod copies in the node where each Pod in the prepared Pods is located is the same, the preparation time corresponding to each Pod is the same, and the number of container restarts of each Pod is the same, determine that the copy corresponding to the fifth Pod is the Pod copy to be deleted, wherein the fifth Pod is the Pod with the latest creation time among the prepared Pods.

7. The method according to any one of claims 1 to 6, characterized in that: The obtaining of at least one indicator associated with the target workload in the cluster includes: Filter the namespace based on the preset fields to obtain creation events of the target workload in the cluster; In response to the creation event, the at least one indicator is obtained through an interface of a target controller in the cluster.

8. The method according to any one of claims 1 to 6, characterized in that: The scaling of the target workload by using the HPA resources and the at least one indicator includes: Determine, by means of the at least one indicator, the number N of Pod replicas to be added corresponding to the target workload, where N is a positive integer; N Pod replicas are added using the HPA resources and the number N of Pod replicas to be added.

9. A workload expansion and contraction device, characterized in that: include: An acquisition module, configured to acquire at least one indicator associated with a target workload in a cluster, wherein the at least one indicator includes a first indicator set by a user and / or a second indicator corresponding to the cluster; A creation module, configured to create Pod horizontal automatic scaling HPA resources corresponding to the target workload based on the at least one indicator; A capacity expansion and contraction module is used to expand or contract the target workload through the HPA resources and the at least one indicator.

10. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the steps in the method according to any one of claims 1 to 8 are implemented.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 8 are implemented.

12. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps in the method according to any one of claims 1 to 8 are implemented.