Method and equipment for expanding and shrinking capacity of application program and storage medium
By receiving the main configuration file and calling a custom algorithm model to determine the expected number of copies of the application, the problem of poor accuracy of the traditional scaling system is solved, and the stable operation and resource optimization of the application are achieved.
Patent Information
- Application Number
- CN202510410276.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional scaling systems have poor accuracy in scaling and scaling for applications and cannot meet the needs of complex environments, resulting in unstable application operation or waste of resources.
By receiving the main configuration file submitted by the target user, calling a custom algorithm model, determining the expected number of replicas of the target application based on the indicator constraints, and achieving scaling.
Improves the accuracy of scaling, ensures that applications operate stably at high loads, saves resources at low loads, and provides greater freedom and scalability.
Smart Images

Figure CN120335897A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the technical field of application management, and particularly to a method, device, and storage medium for scaling applications up and down. Background Art
[0002] Currently, application programs such as cloud-native ones often face unbalanced access traffic. In the face of traffic peaks, scaling up the application program can add resources to handle more traffic to ensure the stable operation of the application program. In the face of traffic valleys, scaling down the application program can reduce its resource occupancy to save costs. Therefore, it is particularly important to perform automatic scaling of application programs to elastically adjust resources.
[0003] Traditional scaling systems scale up and down any application program in a similar way. When the scaling system has difficulty meeting the requirements of the environment where some application programs are located, the accuracy of the scaling result will be poor, resulting in the inability of the application program to operate stably or not saving costs for customers. Therefore, a method for scaling application programs is needed that can improve the accuracy of scaling.
[0004] The content in the background art section is only information known to the inventor personally, and does not mean that the above information has entered the public domain before the filing date of this disclosure, nor does it mean that it can become the prior art of this disclosure. Summary of the Invention
[0005] The method, device, and storage medium for scaling application programs provided in this specification can improve the accuracy of scaling.
[0006] In a first aspect, this specification provides a method for scaling an application program, which is applied to a target scaling system. The method includes: receiving a main configuration file submitted by a target user for a target application, where the main configuration file includes the metric constraint conditions of the target application and a custom algorithm identifier; calling a custom algorithm model external to the target scaling system based on the algorithm identifier, so that the custom algorithm model determines the expected number of replicas of the target application under the constraint of the metric constraint conditions, and the custom algorithm model is a model generated for the target application; and changing the number of replicas of the target application to the expected number of replicas to achieve scaling of the target application.
[0007] In some embodiments, the metric constraint conditions include conditional metrics, and the custom algorithm model is pre-trained based on the historical values of the conditional metrics, the historical values of the operation metrics of the target application, and the corresponding historical number of replicas.
[0008] In some embodiments, the metric constraint condition further includes the target value of the conditional metric, and the desired number of replicas is determined by the following method: inputting the metric constraint condition into the custom algorithm model, where the custom algorithm model is used to obtain the current value of the operation metric and output the desired number of replicas based on the target value of the conditional metric and the current value of the operation metric.
[0009] In some embodiments, the metric constraint condition further includes the target value of the conditional metric, the main configuration file further includes prediction configuration information, and the custom algorithm model is used to predict the desired number of replicas corresponding to multiple future time points under the constraint of the metric constraint condition according to the prediction configuration information. For the desired number of replicas corresponding to each future time point, the custom algorithm model is used to obtain the predicted value of the operation metric at the future time point and determine the desired number of replicas corresponding to the future time point based on the predicted value of the operation metric and the target value of the conditional metric.
[0010] In some embodiments, changing the number of replicas of the target application to the desired number of replicas includes: for each future time point and its corresponding desired number of replicas, changing the number of replicas of the target application to the desired number of replicas before the future time point.
[0011] In some embodiments, it further includes: receiving the main configuration file through the management control module, creating an algorithm configuration file based on the main configuration file, where the algorithm configuration file includes a first declaration field for writing the metric constraint condition and the algorithm identifier; identifying the algorithm configuration file through the algorithm control module, calling the custom algorithm model based on the algorithm identifier, and writing the desired number of replicas determined by the custom algorithm model under the constraint of the metric constraint condition into the first status field in the algorithm configuration file; and identifying the desired number of replicas in the first status field through the management control module.
[0012] In some embodiments, the main configuration file further includes a custom batch gray-scale policy, and the batch gray-scale policy includes batch-changing the replicas of the target application. Changing the number of replicas of the target application to the desired number of replicas includes: obtaining the current number of replicas of the target application; determining the target number of replicas and its corresponding change method based on the current number of replicas and the desired number of replicas; and batch-changing the replicas according to the batch granularity and the change method until the changed quantity reaches the target number of replicas, and the number of replicas after the change is the desired number of replicas.
[0013] In some embodiments, the change method includes deleting replicas. Changing the replicas in batches according to the batch granularity and the change method includes: sequentially removing the traffic of each batch of replicas according to the batch granularity; and deleting each batch of replicas when the traffic removal of each batch of replicas is completed.
[0014] In some embodiments, for each batch of replicas, if the target application has a traffic anomaly after the traffic of the current batch of replicas is removed, traffic is mounted for the current batch of replicas.
[0015] In some embodiments, the change method includes deleting replicas. Changing the replicas in batches according to the batch granularity and the change method includes: sequentially controlling the process dormancy of each batch of replicas according to the batch granularity; and deleting each batch of replicas when the process dormancy of each batch of replicas is completed.
[0016] In some embodiments, for each batch of replicas, if the target application has an anomaly after the process dormancy of the current batch of replicas, the process of the current batch of replicas is awakened.
[0017] In some embodiments, changing the replicas in batches according to the batch granularity and the change method includes: sequentially deleting each batch of replicas according to the batch granularity; or sequentially creating each batch of replicas according to the batch granularity.
[0018] In some embodiments, it further includes: obtaining the current number of replicas of the target application through the management control module, and determining the target number of replicas and its corresponding change method based on the current number of replicas and the desired number of replicas; for each batch of replicas, creating a replica configuration file corresponding to the current batch of replicas through the management control module, where the replica configuration file includes a second declaration field for writing the batch granularity and the change method; identifying the replica configuration file through the execution control module, changing the current batch of replicas according to the batch granularity and the change method, and writing the change result of the current batch of replicas into the second status field in the replica configuration file, where the change result includes successful change or failed change; and identifying the change result in the second status field through the management control module.
[0019] In a second aspect, this specification also provides a computing device for scaling an application up or down, including: at least one storage medium storing a target scaling system for implementing scaling an application up or down; and at least one processor communicatively connected to the at least one storage medium, where when the computing device runs, the at least one processor reads the target scaling system and implements the method for scaling an application up or down according to any one of the first aspect.
[0020] In a third aspect, the present specification also provides a computer-readable non-transitory storage medium, wherein a target scaling system is stored in the computer-readable non-transitory storage medium, and when the target scaling system is executed by at least one processor, the method for scaling an application program according to any one of the first aspects is implemented.
[0021] As can be seen from the above technical solutions, in the method, device, and storage medium for scaling an application program provided in the present specification, a target user can access a custom algorithm model customized for a target application in a target scaling system and declare a corresponding algorithm identifier in a main configuration file. Therefore, the target scaling system can call an external algorithm to determine an expected number of replicas for the target application, thereby implementing the scaling of the target application. Through the method, device, and storage medium of the present specification, the custom algorithm model can be customized by the target user side according to the environment where the target application is located. The customized algorithm model can be strongly adapted to the environment where the target application is located. Therefore, the expected number of replicas can be accurately determined, the accuracy of scaling is improved, the stable operation of the application program is ensured when scaling up is required, and costs are saved for the target user when scaling down is required. Moreover, the target scaling system has scalability, providing greater freedom for users. Users can provide algorithm modules in any program form, not limited to the target scaling system.
[0022] Some other functions of the method, device, and storage medium for scaling an application program provided in the present specification will be partially listed in the following description. According to the description, the content introduced by the following numbers and examples will be obvious to those of ordinary skill in the art. The creative aspects of the method, device, and storage medium for scaling an application program provided in the present specification can be fully explained by practicing or using the methods, devices, and combinations described in the following detailed examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 Shows a schematic diagram of an environment for scaling an application program according to some embodiments of the present specification;
[0025] Figure 2 Shows a hardware structure diagram of a computing device according to some embodiments of the present specification;
[0026] Figure 3Schematic diagram showing a method for scaling an application according to some embodiments of this specification; and
[0027] Figure 4 Flowchart showing a method for scaling an application according to some embodiments of this specification. Detailed implementation
[0028] The following description provides specific application scenarios and requirements of this specification, aiming to enable those skilled in the art to manufacture and use the content in this specification. For those skilled in the art, various partial modifications to the disclosed embodiments are obvious, and without departing from the spirit and scope of this specification, the general principles defined here can be applied to other embodiments and applications. Therefore, this specification is not limited to the illustrated embodiments, but has the broadest scope consistent with the claims.
[0029] The terms used here are only for the purpose of describing specific example embodiments and are not restrictive. For example, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" used here may also include the plural forms. When used in this specification, the terms "comprising", "including" and / or "containing" mean that the associated integers, steps, operations, elements and / or components exist, but do not exclude the existence of one or more other features, integers, steps, operations, elements, components and / or groups, or the addition of other features, integers, steps, operations, elements, components and / or groups in the device / method.
[0030] Considering the following description, these features of this specification and other features, as well as the operations and functions of the related elements of the structure, and the economy of the combination and manufacture of the components can be significantly improved. Referring to the accompanying drawings, all of these form a part of this specification. However, it should be clearly understood that the drawings are only for the purpose of illustration and description and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.
[0031] The flowcharts used in this specification show the operations implemented by a device according to some embodiments in this specification. It should be clearly understood that the operations in the flowchart may not be implemented in sequence. On the contrary, the operations may be implemented in reverse order or simultaneously. In addition, one or more other operations may be added to the flowchart. One or more operations may be removed from the flowchart.
[0032] In this specification, "X includes at least one of A, B, or C" means that X includes at least A, or X includes at least B, or X includes at least C. That is to say, X can include only any one of A, B, and C, or can include any combination of A, B, and C and other possible content / elements. Any combination of A, B, and C can be A, B, C, AB, AC, BC, or ABC.
[0033] In this specification, unless explicitly stated, the associated relationships generated between structures can be direct or indirect. For example, when describing "A is connected to B", unless it is explicitly stated that A is directly connected to B, it should be understood that A can be directly connected to B or can be indirectly connected to B; for another example, when describing "A is above B", unless it is explicitly stated that A is directly above B (A and B are adjacent and A is above B), it should be understood that A can be directly above B or A can be indirectly above B (there are other elements between A and B and A is above B). And so on.
[0034] For convenience of description, some terms are explained as follows:
[0035] Kubernetes (abbreviated as K8s): An open-source platform for automating the deployment, scaling, and management of containerized applications. It helps manage the orchestration and operation and maintenance of containers in multiple environments and is the de facto standard for the current cloud-native application running environment. A K8s cluster consists of a group of physical machines or virtual machines on which the K8s platform runs. A K8s cluster usually includes a master node and several worker nodes. The master node is responsible for the management and control of the entire cluster, and the worker nodes are responsible for running containerized application programs.
[0036] Cloud-native application: An application designed for the cloud environment, usually using a microservices architecture and container technology to achieve high elasticity, fast deployment, and high availability.
[0037] Container: A lightweight, portable, self-contained software packaging technology that enables application programs to run on any operating system that supports container standards and behaves consistently. Containers can isolate the mutual influence between applications and ensure that they do not have problems due to different host environments.
[0038] Pod: The smallest deployable unit in K8s. A Pod represents an instance running on K8s. A Pod can contain one or more containers, and these containers share storage, network, and the specification of how to run. Each Pod has its own unique IP address and container environment, and different Pods can be independent of each other when running.
[0039] Replica: Through the controller in K8s, it ensures that a specified number of replicas of a certain Pod are always in the running state. For the same application, the number of replicas refers to the number of Pod instances of the application running simultaneously, that is, the number of replicas determines how many identical Pods of the application are running in the K8s cluster. In the embodiments of this specification, scaling the application up or down refers to adjusting the number of replicas of the application. When the application faces an increasing workload (such as traffic requests), the number of replicas can be increased for scaling up to share the traffic and improve the processing capacity and availability of the application. Conversely, when the workload of the application decreases, the number of replicas can be reduced for scaling down to save resources.
[0040] HPA (Horizontal Pod Autoscaler): A tool in K8s for automatically adjusting the number of replicas.
[0041] Application (abbreviated as APP): Abbreviated as an application, and its types can include many kinds. For example, a web application refers to an application accessed through a web browser, a mobile application refers to an application designed specifically for mobile devices, a cloud application refers to an application built based on cloud computing technology, and so on. Any type of application is within the scope of protection of the embodiments of this specification.
[0042] Before describing the specific embodiments of this specification in detail, a general introduction to the solution of this specification is given first:
[0043] Currently, the mainstream method is to use the software component HPA in K8s to scale the application up or down. HPA can calculate the number of replicas required for scaling the application up or down using its built-in algorithm. However, the built-in algorithm of HPA is relatively simple and can only be applied to application programs in some simple scenarios, such as application programs where the number of replicas is strictly linearly related to relevant metrics. In the complex production scenarios of enterprise customers, there are often intricate relationships between the relevant metrics of the application program and the number of replicas, and the built-in algorithm of HPA cannot calculate the correct number of replicas for it, resulting in HPA being unable to meet customer requirements. Moreover, some customers already have algorithm models specifically designed for the application program to calculate the number of replicas, but the scalability limitations of HPA prevent it from supporting these external algorithm models.
[0044] For the method of scaling the application up or down provided in this specification, the target user can provide their own algorithm model, and this algorithm model can be specifically designed for the target application, that is, the algorithm model is strongly adapted to the target application and can calculate the accurate number of replicas for the target application, thereby ensuring the accuracy of scaling up or down. Moreover, the target scaling system has scalability and can access any external custom algorithm model, improving the universality of the method of this specification.
[0045] The method for scaling application programs provided in this specification can be applied in various scenarios. For example, in the scenario of traffic fluctuations, many application programs will experience significant traffic fluctuations. For example, e-commerce websites during promotional periods, news websites during major events, or social platforms during specific time periods when user activities surge. Through the method of this specification, it can be ensured that the application program can handle more requests during peak periods and save resources during off-peak periods. Another example is the scenario of periodic workloads. Some application programs have obvious periodic characteristics, such as nightly batch jobs, regular data synchronization, etc. Through the method of this specification, the resource allocation can be automatically adjusted to adapt to different workload patterns. Another example is the scenario of geographical distribution and edge computing. For applications deployed across regions, there may be significant differences in the number of user accesses in different regions. Through the method of this specification, regional scaling can be performed based on location-related data to ensure that users in each region can obtain a good experience.
[0046] It should be noted that the above example scenarios are just a few usage scenarios provided in this specification. The method for scaling application programs provided in this specification can be applied not only to the above scenarios but also to other scenarios. Those skilled in the art should understand that the method for scaling application programs described in this specification being applied to other usage scenarios is also within the protection scope of this specification.
[0047] Figure 1 The figure shows a schematic diagram of an environment 001 for scaling an application program according to some embodiments of this specification. As Figure 1 shown, the environment 001 may include a target user 100, a client 200, a server 300, and a network 400.
[0048] The target user 100 can use the client 200 to submit a scaling request for the target application, such as submitting the main configuration file of the target application. The target user 100 and the target application may belong to the same enterprise. The developers of the enterprise can specifically generate an algorithm model for the target application and connect the algorithm model to the client 200. The connection to the client 200 means that the client 200 can call the algorithm model. The algorithm model can run on the client 200 or on other client devices or server devices. It should be noted that the user data obtained in this specification has been authorized by the user and does not involve user privacy.
[0049] The client 200 can display various interfaces for the target user 100 to view and trigger. In some embodiments, the target user 100 can write a main configuration file in YAML or JSON format and use a CMD (Command) tool, such as kubectl, to submit the main configuration file on the client 200. In some embodiments, the target user 100 can create and submit the main configuration file on the graphical user interface (UI) of the client 20. Resource objects, such as HPA resource objects and IHPA resource objects, can be created in the main configuration file. The target user 100 can declare some information for the created resource objects in the main configuration file, such as the target application, metric constraint conditions, specified frequency, custom algorithm model, custom batch grayscale policy, and so on.
[0050] The method for scaling an application can be executed on a computing device. In some embodiments, the computing device can be the client 200. At this time, the client 200 can store data or instructions for executing the method for scaling an application described in this specification and can execute or be used to execute the data or instructions. In some embodiments, the client 200 can include a hardware device with data information processing capabilities and the necessary programs for driving the hardware device to work. Such as Figure 1As shown, the client 200 can communicate with the server 300. In some embodiments, the server 300 can communicate with multiple clients 200. In some embodiments, the client 200 can interact with the server 300 through the network 400 to receive or send messages, etc. In some embodiments, the client 200 can include mobile devices, tablets, laptops, in-vehicle devices of motor vehicles or the like, vending machines, vending cabinets, or any combination thereof. In some embodiments, the mobile device can include smart home devices, smart mobile devices, virtual reality devices, augmented reality devices or the like, or any combination thereof. In some embodiments, the smart home device can include smart TVs, desktop computers, etc., or any combination. In some embodiments, the smart mobile device can include smartphones, personal digital assistants, gaming devices, navigation devices, etc., or any combination thereof. In some embodiments, the virtual reality device or augmented reality device may include virtual reality headsets, virtual reality glasses, virtual reality patches, augmented reality headsets, augmented reality glasses, augmented reality patches or the like, or any combination thereof. For example, the virtual reality device or the augmented reality device may include smart glasses, head-mounted displays, VR, etc. In some embodiments, the in-vehicle device in the motor vehicle can include in-vehicle computers, in-vehicle TVs, etc. In some embodiments, the client 200 can be a device with positioning technology for positioning the location of the client 200. In some embodiments, the client 200 can have one or more of the following functions: NFC (Near Field Communication), WIFI (Wireless Fidelity), 3G / 4G / 5G, POS (Point Of Sale) machine card swiping function, two-dimensional code scanning function, bar code scanning function, Bluetooth, infrared, SMS (Short Message Service), MMS (Multimedia Message Service).
[0051] In some embodiments, the client 200 may be installed with one or more applications (APPs). The APPs can provide the target user 100 with the ability and interface to interact with the outside world through the network 400. The APPs include, but are not limited to: web browser APP programs, search APP programs, chat APP programs, shopping APP programs, video APP programs, financial management APP programs, instant messaging tools, email clients, social platform software, and so on. In some embodiments, the client 200 may install a target application, and the client 200 may scale the target application up or down by the method of scaling the application program up or down. In some embodiments, the client 200 may install a scaling application, and the target user 100 may trigger a scaling request for the target application through the scaling application, such as submitting a main configuration file. The scaling application may execute the method of scaling the application program up or down in response to the scaling request.
[0052] In some embodiments, the computing device may be the server 300. At this time, the server 300 may store data or instructions for executing the method of scaling the application program up or down described in this specification, and may execute or be used to execute the data or instructions. In some embodiments, the server 300 may include a hardware device with data information processing functions and necessary programs for driving the hardware device to work. The server 300 may be communicatively connected to multiple clients 200 and receive data sent by the clients 200. The server 300 may be a server providing various services, such as a back-end server that supports the pages displayed by the target application and / or the scaling application on the client 200.
[0053] The network 400 is a medium for providing a communication connection between the client 200 and the server 300. The network 400 can facilitate the exchange of information or data. As Figure 1 shown, the client 200 and the server 300 can be connected to the network 400 and transmit information or data to each other through the network 400. In some embodiments, the network 400 can be any type of wired or wireless network, or a combination thereof. For example, the network 400 can include a cable network, a wired network, an optical fiber network, a telecommunication network, an intranet, the Internet, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a public switched telephone network (PSTN), a Bluetooth network, a ZigBee network, a near field communication (NFC) network, or a similar network. In some embodiments, the network 400 can include one or more network access points. For example, the network 400 can include wired or wireless network access points, such as base stations or Internet exchange points, through which one or more components of the client 200 and the server 300 can be connected to the network 400 to exchange data or information.
[0054] It should be understood that Figure 1 the numbers of the client 200, the server 300, and the network 400 in [[ ]] are merely illustrative. According to the implementation requirements, there can be any number of clients 200, servers 300, and networks 400.
[0055] It should be noted that the method for scaling the application can be executed entirely on the client 200, entirely on the server 300, or partially on the client 200 and partially on the server 300.
[0056] Figure 2 The hardware structure diagram of a computing device 600 provided according to some embodiments of the present specification is shown. The computing device 600 can execute the method for scaling the application described in the present specification. The method for scaling the application is introduced in other parts of the present specification. The computing device 600 can be a device of the client 200, a device of the server 300, or other computing devices, and can even be any combination of the above devices.
[0057] As Figure 2 shown, the computing device 600 can include at least one storage medium 630 and at least one processor 620. In some embodiments, the computing device 600 can further include a communication port 650 and an internal communication bus 610. At the same time, the computing device 600 can further include I / O components 660.
[0058] The internal communication bus 610 can connect different components, including the storage medium 630, the processor 620, and the communication port 650.
[0059] The I / O components 660 support input / output between the computing device 600 and other components.
[0060] The communication port 650 is used for data communication between the computing device 600 and the outside world. For example, the communication port 650 can be used for data communication between the computing device 600 and the network 400. The communication port 650 can be a wired communication port or a wireless communication port.
[0061] The storage medium 630 may include a data storage device. The data storage device may be a non-transitory storage medium or a transitory storage medium. For example, the data storage device may include one or more of a magnetic disk 632, a read-only storage medium (ROM) 634, or a random access storage medium (RAM) 636. The storage medium 630 may store at least one instruction set for implementing application scaling. The instructions are computer program codes, and the computer program codes may include programs, routines, objects, components, data structures, procedures, modules, etc. for executing the method for application scaling provided in this specification. The storage medium 630 may also store a software system for implementing the method for application scaling. At this time, the software system may be one or more instruction sets stored in the storage medium 630 that execute corresponding instructions, and are executed by the processor 620 in the computing device 600.
[0062] At least one processor 620 may be communicatively connected to at least one storage medium 630 and a communication port 650 through an internal communication bus 610. At least one processor 620 is configured to execute the above at least one instruction set. When the computing device 600 is running, at least one processor 620 may read the at least one instruction set, and according to the instructions of the at least one instruction set, execute the method for application scaling provided in this specification. The processor 620 may execute all steps included in the method for application scaling. The processor 620 may be in the form of one or more processors. In some embodiments, the processor 620 may include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application specific integrated circuit (ASIC), an application specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physics processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of executing one or more functions, etc., or any combination thereof. For illustrative purposes only, only one processor 620 is described in the computing device 600 in this specification. However, it should be noted that the computing device 600 in this specification may also include multiple processors. Therefore, the operations and / or method steps disclosed in this specification may be executed by one processor as described in this specification, or jointly executed by multiple processors. For example, if the processor 620 of the computing device 600 executes step A and step B in this specification, it should be understood that step A and step B may also be jointly or separately executed by two different processors 620 (e.g., the first processor executes step A, the second processor executes step B, or the first and second processors jointly execute steps A and B).
[0063] The target scaling system can be stored on the computing device 600. In some embodiments, the target scaling system can be a plug-in that can be added to an existing software system to expand or increase its functions without modifying the code of the original software system. In some embodiments, the target plug-in corresponding to the target scaling system can be a plug-in for K8s, capable of providing additional functions and services. The target plug-in can follow the protocol of K8s, i.e., the standards defined by K8s. The target plug-in can call the APIs provided by K8s to perform various operations. For example, the target user 100 submits the main configuration file through the standard K8s API, the target plug-in creates or deletes copies of the target application through the K8s API, and the target plug-in calls the network traffic control API to control the traffic of the target application, and so on. The target plug-in can also call the components running on the K8s nodes, such as calling the process control component (Agent) on the K8s node to control the wake-up and sleep of the application process. The target plug-in can also interface with the native components of K8s, such as controlling the addition and deletion of Pods to achieve scaling. In some embodiments, the target scaling system can also be an independent application program, which is not limited in the embodiments of this specification.
[0064] Figure 3 A schematic block diagram showing a method for scaling an application according to some embodiments of this specification is shown. The target scaling system can be referred to as the IHPA system. As Figure 3 shown, the IHPA system can include a management control module, an algorithm control module, a metric query component, and an execution control module.
[0065] The management control module (such as the IHPA Controller) can be the control plane module of the IHPA system and directly receive the main configuration file submitted by the target user 100. As Figure 3As shown, the target user 100 can submit the main configuration file through the CMD or UI method. In the main configuration file, an IHPA resource object can be created, and the declaration information including the resource object can be included. The declaration information can include, for example, the target application, metric constraint conditions, custom algorithm identifier, custom batch gray-scale policy, and so on. The management control module can send a determination instruction for the desired replica count to the algorithm control module. The determination instruction can be implemented through the declaration information of the resource object. For example, the management control module can convert the custom algorithm identifier and the relevant configuration parameters of the custom algorithm model into the declaration information of the algorithm resource object that the algorithm control module can recognize. In some embodiments, the management control module can create an algorithm configuration file based on the main configuration file. The algorithm configuration file contains a first declaration field for writing the metric constraint conditions and the algorithm identifier. The first declaration field belongs to the declaration information of the algorithm resource object. In some embodiments, the application portrait of the target application can include various features such as metrics for scaling, custom algorithm identifiers, and desired replica counts. Correspondingly, the determination instruction sent by the management control module can be a portrait task, which is used to instruct the algorithm control module to determine the desired replica count according to information such as the scaling metrics of the target application and the custom algorithm identifier, and receive the portrait result, that is, the desired replica count.
[0066] The algorithm control module (such as HPortrait Controller) can be a management component of the horizontal elastic scaling algorithm of the IHPA system, responsible for calling, running, and managing different algorithm models. The algorithm control module can receive the determination instruction. In some embodiments, the algorithm control module can recognize the algorithm configuration file, call a custom algorithm model outside the target scaling system based on the algorithm identifier, so that the custom algorithm model determines the desired replica count of the target application under the constraint of the metric constraint conditions. Thus, the algorithm control module can write the desired replica count into the first status field in the algorithm configuration file. The management control module can recognize the desired replica count in the first status field. The first status field belongs to the declaration information of the algorithm resource object. The custom algorithm model can run on the K8s platform or other big data or AI platforms. The hardware devices of different platforms can be the same or different. The algorithm control module can view the algorithm result (i.e., the desired replica count) generated by the custom algorithm model, and then write it into the status field.
[0067] In some embodiments, the algorithm control module may integrate built-in algorithms. The number of the built-in algorithms may be multiple. The target user 100 may declare the built-in algorithms to be used in the main configuration file, so that the algorithm controller may automatically run the declared built-in algorithms. The logic code of the built-in algorithms belongs to a part of the IHPA system code.
[0068] In some embodiments, the target user 100 may customize the algorithm control module. When the target user 100 declares to customize the algorithm control module in the main configuration file, the IHPA system may use the customized algorithm control module to manage various algorithms. This provides greater freedom for users, and users may provide algorithm control modules in any program form, not limited to the IHPA software system.
[0069] The metric query component (such as the Metrics Provider Server) may be the unified metric query component of the IHPA system, used to mask the differences between different underlying monitoring systems (such as the Monitor). The monitoring system may monitor the values of various metrics of the target application in real time. For the same metric, when different monitoring systems give different forms of the metric, the metric query component may mask these metric differences and provide a unified form, so as to provide a unified metric query service for externally running components (such as the customized algorithm model). When determining the desired number of replicas, the customized algorithm model may send a request to the metric query component to obtain the value of the required metric. Further, the metric query component requests the value of the operation metric from the monitoring system and returns it to the customized algorithm model. In some embodiments, when the customized algorithm model needs to predict the desired number of replicas at a future moment, the metric query component may provide it with the historical data of the required metric, so that it can make a prediction based on the historical data.
[0070] In some embodiments, when the algorithm control module determines the desired number of replicas using the built-in algorithm, it may directly request the value of the required metric from the monitoring system. The management and control module may also request the values of various metrics of the target application from the monitoring system in real time.
[0071] The execution control module (such as Replica Controller) can be the execution layer module of the IHPA system, responsible for controlling the number and status of replicas. When the management control module identifies the desired number of replicas from the algorithm resource object, it can send scaling instructions corresponding to each batch of replicas to the execution control module in combination with the batch gray-scale policy in the main configuration file. The scaling instructions can be implemented in the form of replica resource objects. In some embodiments, the management control module can obtain the current number of replicas of the target application and determine the target number of replicas and its corresponding change method based on the current number of replicas and the desired number of replicas. For each batch of replicas, the management control module can create a replica configuration file corresponding to the current batch of replicas, and the replica configuration file includes a second declaration field for writing the batch granularity and the change method. The second declaration field belongs to the replica resource object. The execution control module can identify the replica configuration file, change the current batch of replicas according to the batch granularity and the change method, and write the change result of the current batch of replicas into the second status field in the replica configuration file. The change result includes successful change or failed change. The management control module can identify the change result in the second status field. In some embodiments, the execution control module can call the replica control interface to perform the operation of changing the number of replicas, and control the traffic removal and mounting operations of the network traffic control system by calling the network traffic control interface. Among them, the replica control interface can be customized by the user to interface with the replicas of different application programs. The network traffic control interface can also be customized by the user to interface with different network traffic management systems. The execution control module can also call the process control component to perform the sleep or wake-up operation of the processes within the replicas.
[0072] In the embodiments of this specification, the IHPA system can be split into different extensible and replaceable modules according to the responsibilities of management, algorithm, and execution, and the modules can interact through configuration files that follow the same protocol, realizing the serial collaboration between the modules. Thus, these multiple modules can cooperate together to implement the complete scaling method. Moreover, customers can freely access customized algorithm models that are more suitable for their actual scenarios. At the same time, due to the extensibility characteristics of each module of the IHPA system, users can customize the required parts at the lowest cost and access the IHPA system, thereby solving problems such as effectiveness, stability, and heterogeneous adaptability that applications may encounter in complex production scenarios.
[0073] Figure 4 FIG. shows a flowchart of a method P100 for scaling an application program according to some embodiments of this specification. As described above, the computing device 600 can execute the method P100 for scaling an application program described in this specification. As Figure 4As shown, the method P100 may include:
[0074] S120: Receive a main configuration file submitted by a target user for a target application, where the main configuration file includes metric constraint conditions of the target application and a custom algorithm identifier.
[0075] As mentioned above, the target user 100 may submit the main configuration file via CMD or UI, and the management control module may receive this main configuration file. The main configuration file may declare various resource objects and configuration information, and it describes in a declarative manner the target state that the target user 100 hopes the target application to achieve.
[0076] The main configuration file may include an identifier of the target application to indicate that the object for which the computing device 600 scales in or out is the target application. The main configuration file may include metric constraint conditions, and the metric constraint conditions may be used as constraint conditions for the computing device 600 to calculate the desired number of replicas. The metric constraint conditions may include conditional metrics and their target values. The conditional metrics may be, for example, CPU utilization, memory utilization, QPS (Queries Per Second), RPM (Requests Per Minute), queue length, the number of completed tasks, the number of messages in the message queue, the number of database connections, and so on. When the target user 100 specifies a target value for a conditional metric, the computing device 600 may attempt to maintain the conditional metrics of each replica of the target application at the target value by increasing or decreasing the number of replicas, such as maintaining the average value of the conditional metrics of each replica at the target value. For example, the target value of the metric constraint condition for CPU utilization is 50%. In some embodiments, the number of metric constraint conditions may be multiple, such as the target value of CPU utilization is 50% and the target value of memory utilization is 60%.
[0077] As mentioned above, the target application corresponds to a custom algorithm model. The target user 100 may declare the algorithm identifier of the custom algorithm model in the main configuration file, such as the name of the algorithm model, or the name and IP address, to indicate that the computing device 600 calls the custom algorithm model according to this algorithm identifier. Of course, the target user 100 may also configure some parameters of the custom algorithm model in the main configuration file.
[0078] In some embodiments, the master configuration file may include the algorithm running frequency, that is, the target user 100 can specify how often the custom algorithm model runs. For example, it can run once a day, once a week, etc. Therefore, the custom algorithm model can run according to the algorithm running frequency. In some embodiments, the master configuration file may include prediction configuration information. When the custom algorithm model predicts the expected number of replicas at a future moment, it can make the prediction according to the prediction configuration information. For example, the prediction configuration information includes the prediction time range and the prediction frequency. The custom algorithm model can predict the expected number of replicas within the prediction time range according to the prediction frequency. For example, it can predict the expected number of replicas at multiple future moments within the next day at a frequency of once every 10 minutes. Of course, the prediction configuration information may also include other information, which is not limited in the embodiments of this specification.
[0079] In some embodiments, the master configuration file further includes a custom batch gray-scale policy. The batch gray-scale policy refers to changing the replicas of the target application in batches. The batch gray-scale policy may include the batch granularity, the number of batches, the waiting duration between batches, and / or the status of each batch of replicas, etc. The batch granularity among them refers to the number of replicas changed in each batch. The batch gray-scale policy may further include a rollback policy, which means restoring the number of replicas to the state of the previous change when the target application reaches the rollback condition. The rollback condition may be that the error rate of the target application reaches the target value, and this error rate is, for example, the ratio of the number of requests with processing failures to the total number of requests.
[0080] In some embodiments, the master configuration file may further include other information, such as the version of the API used, the resource type, and the resource name. For another example, the master configuration file may include information on whether to only display the expected number of replicas without performing changes. Through this information, the computing device 600 can use the execution control module provided by the target user 100 to perform changes. The target user 100 can also modify the master configuration file. For example, after the computing device 600 scales the target application up or down for a period of time, the target user 100 can add information to pause the scaling in the master configuration file to instruct the computing device 600 to pause the scaling of the target application.
[0081] S140: Invoke a custom algorithm model external to the target scaling system based on the algorithm identifier, so that the custom algorithm model determines the expected number of replicas of the target application under the constraint of the metric constraint condition, and the custom algorithm model is a model generated for the target application.
[0082] In some embodiments, the custom algorithm model is pre-trained based on the historical values of the conditional metrics, the historical values of the operation metrics of the target application, and the corresponding historical number of replicas. The custom algorithm model can be a trained machine learning model. The computing device 600 can collect the values of the conditional metrics, operation metrics, and their corresponding number of replicas of a large number of target applications at historical moments, and perform multiple rounds of training on the machine learning model based on the collected historical data. The custom algorithm model obtained through training can output the accurate number of replicas according to the input values of the conditional metrics and operation metrics. The custom algorithm model is a trained model that can describe the intricate non-linear relationship between the conditional metrics, operation metrics, and the number of replicas, solve complex problems that cannot be solved by simple linear algorithms, and effectively and accurately calculate the expected number of replicas in complex scenarios.
[0083] Among them, the conditional metrics and the operation metrics can be different. For example, the conditional metric is the CPU utilization rate, and the operation metrics are QPS and time. There is an association relationship between these metrics and the number of replicas. For example, the target application processes scheduled tasks every day, and the CPU utilization rate will increase when processing the scheduled tasks, but the QPS remains unchanged. At this time, the number of replicas needs to be increased. And the QPS of the target application may decrease at night, and the CPU utilization rate will also decrease. At this time, the number of replicas needs to be reduced. Therefore, the historical data of the CPU utilization rate, QPS, time, and the number of replicas can be used to train the custom algorithm model, so that the custom algorithm model can learn the association relationship between the CPU utilization rate, QPS, time, and the number of replicas, and thus calculate the corresponding expected number of replicas based on the values of the CPU utilization rate, QPS, and time. Another example is that the conditional metric is the CPU utilization rate, and the operation metrics are RPM and time, etc. The embodiments of this specification do not limit the type of operation metrics. The factors affecting the number of replicas are often multiple. Through the embodiments of this specification, the association relationship between multiple metrics and the number of replicas can be established, so as to obtain an accurate result of the number of replicas.
[0084] In some embodiments, the computing device 600 may determine a corresponding desired number of replicas for the current value of the operation metric. In some embodiments, the computing device 600 may input the metric constraint condition into the custom algorithm model, which is used to obtain the current value of the operation metric. For example, the current value of the operation metric may be requested through a metric query component. Furthermore, based on the target value of the conditional metric and the current value of the operation metric, the desired number of replicas is output. For example, the algorithm control module may obtain the metric constraint condition in the algorithm configuration file and input the metric constraint condition to the custom algorithm model when calling the custom algorithm model, so that the custom algorithm model outputs the desired number of replicas. For example, if the target value of the CPU utilization rate in the metric constraint condition is 50% and the current value of the operation metric QPS is 1000, the custom algorithm model may output the corresponding desired number of replicas. In the embodiments of this specification, the model can calculate the desired number of replicas based on the value of the operation metric under the constraint conditions (metric constraint conditions) desired by the user, which can not only meet the user's requirements but also output accurate desired number of replicas.
[0085] In some embodiments, the computing device 600 may determine a corresponding desired number of replicas for the predicted value of the operation metric. In some embodiments, the custom algorithm model is used to predict the desired number of replicas corresponding to multiple future time points under the constraint of the metric constraint condition according to the prediction configuration information. For each desired number of replicas corresponding to a future time point, the custom algorithm model is used to obtain the predicted value of the operation metric at the future time point and determine the desired number of replicas corresponding to the future time point based on the predicted value of the operation metric and the target value of the conditional metric.
[0086] Taking the prediction configuration information including the prediction time range and prediction frequency as an example. The algorithm control module may obtain the metric constraint condition and the prediction configuration information in the algorithm configuration file and transmit the metric constraint condition and the prediction configuration information to the custom algorithm model when calling the custom algorithm model. The custom algorithm model may predict the desired number of replicas within the prediction time range according to the prediction frequency. Specifically, for the desired number of replicas at each future time point, the custom algorithm model may obtain the predicted value of the operation metric (such as QPS) at the future time point. For example, the custom algorithm model may obtain the predicted value of the operation metric at the future time point from the operation metric prediction model (QPS prediction model), so as to calculate the desired number of replicas at the future time point. For example, the current time is 0:00, and a future time point is 0:10. The custom algorithm model obtains the predicted value of QPS at 0:10 as 10000, and thus calculates the desired number of replicas corresponding to 0:10 according to the predicted value of QPS 10000 and the target value of the CPU utilization rate of 50%.
[0087] In the embodiments of this specification, the computing device 600 may predict the predicted value of the operation metric, and determine the corresponding desired number of replicas based on the predicted value of the operation metric. That is, it may determine the number of replicas required for the target application for the upcoming situation, so as to facilitate early deployment to cope with the upcoming situation.
[0088] In some embodiments, the custom algorithm model may be pre-stored on the computing device 600. In some embodiments, the custom algorithm model may be stored on an external device. At this time, the computing device 600 may send a request for determining the desired number of replicas to the external device through an interface, so as to request the external device to determine the desired number of replicas based on the custom algorithm model, and receive the desired number of replicas returned by the external device.
[0089] S160: Change the number of replicas of the target application to the desired number of replicas to implement scaling of the target application.
[0090] In some embodiments, when the computing device 600 determines the corresponding desired number of replicas for the current value of the operation metric, it may obtain the current number of replicas of the target application, and when it determines that the current number of replicas is inconsistent with the desired number of replicas, change the current number of replicas to the desired number of replicas. For example, when the management control module obtains that the current number of replicas of the target application is 4 and determines that it is inconsistent with the desired number of replicas 10, it instructs the execution control module to increase the number of replicas from 4 to 10.
[0091] In some embodiments, when the computing device 600 determines the desired number of replicas corresponding to a future time based on the predicted value of the operation metric, it may make a change in advance. For example, for each future time and its corresponding desired number of replicas, the computing device 600 changes the number of replicas of the target application to the desired number of replicas before the future time. For example, for each future time and its corresponding desired number of replicas, the computing device 600 may determine the time consumption required for scaling the target application based on the desired number of replicas, subtract the time consumption from the future time to obtain the change time, and thus make a change at the change time. Therefore, before the future time arrives, the number of replicas of the target application has reached the desired number of replicas.
[0092] In the embodiments of this specification, the computing device 600 may perform scaling in advance. Scaling requires a certain amount of time. If scaling is performed when the target application already has load imbalance, then the target application will always be in a load imbalance state (such as a high load state) during the scaling period, making it unable to run stably or wasting resources. Therefore, performing scaling in advance can avoid the target application being in a load imbalance state, always ensure the stable operation of the application, and avoid resource waste.
[0093] As described above, the main configuration file may include a batch gray scale policy. Accordingly, the computing device may make changes according to the batch gray scale policy. In some embodiments, the computing device 600 may obtain the current number of copies of the target application, and determine the target number of copies and its corresponding change method based on the current number of copies and the desired number of copies. For example, if the current number of copies of the target application is 4 and the desired number of copies is 10, it means that 6 copies need to be created, that is, the target number of copies is 6, and the change method is to create copies. Another example, if the current number of copies of the target application is 10 and the desired number of copies is 3, it means that 7 copies need to be deleted, that is, the target number of copies is 7, and the change method is to delete copies. Further, the computing device 600 may change the copies in batches according to the batch granularity and the change method until the number of changes reaches the target number of copies, and the number of copies of the target application after the change is the desired number of copies. Among them, the batch granularity may be specified by the target user 100 in the batch gray scale policy, or may be automatically generated by the control control module according to the target number of copies. When the target number of copies is not a multiple of the batch granularity, the number of copies changed in the last batch may be less than the batch granularity, and the change is made according to the actual number of copies. For example, 6 copies need to be added, the batch granularity is 2, and 2 copies are added each time. Another example, 7 copies need to be deleted, the batch granularity is 2, 2 copies are deleted each time in the first 3 changes, and 1 copy is deleted in the 4th change.
[0094] The computing device 600 may change the copies by controlling the traffic of the copies. Specifically, the execution control module may call the network traffic control interface to dock with the network traffic control system of the target application, so as to control the removal and mounting of the traffic of the target application through the network traffic control system. The network traffic control systems of different applications may be different, and the network traffic control interfaces used to dock with different network traffic control systems may be different. Therefore, the developers of each application can set up dedicated network traffic control interfaces according to the characteristics of the network traffic control systems of their own applications, so that the execution control module can successfully dock with the network traffic control system through the network traffic control interface.
[0095] In some embodiments, when the change method is to delete replicas, the computing device 600 may sequentially remove the traffic of each batch of replicas according to the batch granularity, and delete each batch of replicas when the traffic removal of each batch of replicas is completed. For example, when removing the traffic of each batch of replicas, the management control module may send an instruction to remove the traffic to the execution control module. Further, the execution control module may select a batch of replicas from the current replicas and call the network traffic control interface to remove the traffic of this batch of replicas. For each batch of replicas whose traffic is removed, the execution control module may return the removal result to the management control module, that is, whether the removal is successful or failed. When the management control module determines that the removal result corresponding to each batch of replicas is successful, it may send an instruction to delete the replicas to the execution control module. Thus, the execution control module may call the replica control interface to delete each batch of replicas. In the embodiments of this specification, the traffic of the replicas is removed sequentially, and the replicas are deleted only when it is ensured that the application remains in a normal state throughout the traffic removal process, ensuring the operational stability of the application during the scaling process.
[0096] In some embodiments, for each batch of replicas, if the target application has a traffic anomaly after the traffic of the current batch of replicas is removed, traffic is mounted for the current batch of replicas. After the traffic of a certain batch of replicas is removed, this batch of replicas cannot receive traffic, and the target application may have a traffic anomaly. For example, the traffic suddenly increases, causing the replicas that have not had their traffic removed to be unable to carry the current traffic. At this time, the execution control module can quickly mount traffic for this batch of replicas to quickly restore it to the working state to handle the traffic anomaly. In the embodiments of this specification, the speed of mounting traffic is much faster than the speed of creating replicas. Therefore, in the event of a traffic anomaly, it can be quickly rolled back, that is, quickly restored to the original state to solve the anomaly problem, ensuring the operational stability of the application during the change process.
[0097] The computing device 600 can also change replicas by controlling the process state of the replicas. Specifically, a process control component may be installed on each node where a replica is located, and the execution control module may call the process control component corresponding to the replica to control whether the process of the replica is in a dormant state or a wake-up state.
[0098] In some embodiments, when the change method is to delete replicas, the computing device 600 may sequentially control the process suspension of each batch of replicas according to the batch granularity, and delete each batch of replicas when the process suspension of each batch of replicas is completed. For example, when controlling the process suspension of each batch of replicas, the management control module may send a process suspension instruction to the execution control module. Subsequently, the execution control module may select a batch of replicas from the current replicas and call the process control component corresponding to this batch of replicas to control the process suspension of this batch of replicas. For each batch of replicas whose process suspension is controlled, the execution control module may return a suspension result to the management control module, that is, whether the suspension is successful or failed. When the management control module determines that the suspension result corresponding to each batch of replicas is successful, it may send an instruction to delete the replicas to the execution control module. Thus, the execution control module may call the replica control interface to delete each batch of replicas. In the embodiments of this specification, the process suspension of the replicas is controlled sequentially first, and the replicas are deleted only when it is ensured that the application remains in a normal state throughout the process of removing traffic, ensuring the stable operation of the application during the scaling process.
[0099] In some embodiments, for each batch of replicas, if the target application has an abnormal situation after the process suspension of the current batch of replicas, the process of the current batch of replicas is awakened. This abnormal situation includes traffic abnormal situations and may also include other abnormal situations. If the target application has an abnormal situation after controlling the process suspension of a certain batch of replicas, at this time, the execution control module can quickly awaken the process of this batch of replicas to quickly restore it to the working state to handle the abnormal situation. In the embodiments of this specification, the speed of awakening the process is much faster than the speed of creating replicas. Therefore, when an abnormal situation occurs, it can be quickly rolled back, that is, quickly restored to the original state to solve the abnormal problem, ensuring the running stability of the application during the change process.
[0100] In some embodiments, when the change method is to create replicas, the computing device 600 may sequentially create each batch of replicas according to the batch granularity. Among them, when creating the current batch of replicas, the computing device 600 may awaken the process of the current batch of replicas and mount traffic to the current batch of replicas. If the current batch of replicas has a running abnormal situation, the traffic of the current batch of replicas is removed, or the traffic of the current batch of replicas is removed and the process is suspended until the current batch of replicas returns to normal. In the embodiments of this specification, for each batch of replicas created, the process of this batch of replicas is awakened and traffic is mounted. Abnormal situations may also occur when creating replicas. For example, replica creation errors, network errors, etc. can control the replicas to be temporarily in a non-working state by removing traffic, or removing traffic and suspending the process, and repair the abnormality. After successfully repairing the abnormal problem, the process is awakened and traffic is mounted, rather than directly deleting the replicas when a problem occurs and then creating replicas when normal. The method in this specification facilitates rapid expansion and improves the efficiency of expansion.
[0101] In some embodiments, the computing device 600 may also directly change the replicas. For example, the execution control module may sequentially delete each batch of replicas according to the batch granularity, or sequentially create each batch of replicas according to the batch granularity, thereby improving the efficiency of scaling up and down.
[0102] In summary, in the method, device, and storage medium for scaling up and down an application program provided in this specification, a target user can access a custom algorithm model customized for a target application in a target scaling up and down system and declare a corresponding algorithm identifier in a main configuration file. Therefore, the target scaling up and down system can call an external algorithm to determine the desired number of replicas for the target application, thereby implementing the scaling up and down of the target application. Through the method, device, and storage medium of this specification, the custom algorithm model can be customized by the target user side according to the environment where the target application is located. The customized algorithm model can be strongly adapted to the environment where the target application is located. Therefore, it can accurately determine the desired number of replicas, improve the accuracy of scaling up and down, ensure the stable operation of the application program when scaling up is required, and save costs for the target user when scaling down is required. Moreover, the target scaling up and down system has scalability, providing users with greater freedom. Users can provide algorithm modules in any program form, not limited to the target scaling up and down system.
[0103] On the other hand, this specification provides a non-transitory storage medium storing at least one set of executable instructions for performing scaling of an application. When the executable instructions are executed by a processor, the executable instructions direct the processor to perform the steps of the method P100 for scaling the application described in this specification. In some possible implementation manners, various aspects of this specification may also be implemented in the form of a program product, which includes program code. When the program product runs on a device for scaling an application, the program code is used to cause the device for scaling the application to perform the steps of the method P100 for scaling the application described in this specification. The program product for implementing the above method may use a portable compact disc read-only memory (CD-ROM) to include the program code and may run on the device for scaling the application. However, the program product of this specification is not limited to this. In this specification, the readable storage medium may be any tangible medium that contains or stores a program, and this program may be used by or in combination with an instruction execution device. The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer-readable storage medium may include a data signal propagated in a baseband or as a part of a carrier wave, which carries the readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium may also be any readable medium other than the readable storage medium, and this readable medium may send, propagate, or transmit a program used by or in combination with an instruction execution device, apparatus, or device. The program code included on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above. The program code for performing the operations of this specification may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages.The program code can be executed entirely on the device that scales the application, partially on the device that scales the application, executed as an independent software package, partially on the device that scales the application and partially on a remote computing device, or entirely on a remote computing device.
[0104] The specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require a particular order or a sequential order to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0105] In summary, after reading this detailed disclosure, those skilled in the art will understand that the foregoing detailed disclosure may be presented only by way of example and may not be restrictive. Although not explicitly stated herein, those skilled in the art can understand that this specification is intended to encompass various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be proposed by this specification and are within the spirit and scope of the exemplary embodiments of this specification.
[0106] In addition, certain terms in this specification have been used to describe the embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" mean that the specific features, structures, or characteristics described in connection with that embodiment may be included in at least one embodiment of this specification. Therefore, it should be emphasized and understood that two or more references to "an embodiment" or "one embodiment" or "alternative embodiments" in various parts of this specification do not necessarily all refer to the same embodiment. Additionally, the specific features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.
[0107] It should be understood that in the foregoing description of the embodiments of this specification, for the purpose of helping to understand a feature and for the purpose of simplifying this specification, this specification combines various features in a single embodiment, figure, or its description. However, this does not mean that the combination of these features is necessary. Those skilled in the art, when reading this specification, may very well mark out some of the devices as separate embodiments for understanding. That is to say, the embodiments in this specification can also be understood as the integration of multiple sub - embodiments. And it also holds when the content of each sub - embodiment contains fewer features than all the features of a single foregoing disclosed embodiment.
[0108] Each patent, patent application, published patent application, and other materials cited herein, such as articles, books, specifications, publications, documents, items, etc., except for any historical prosecution documents associated therewith, any identical ones that may be inconsistent or in conflict with this document, or any identical historical prosecution documents that may have a limiting effect on the broadest scope of the claims, may be incorporated herein by reference and used for all purposes now or hereafter associated with this document. In addition, in the event of any inconsistency or conflict between the description, definition, and / or use of terms associated with any of the included materials and the terms, descriptions, definitions, and / or uses associated with this document, the terms of this document shall prevail.
[0109] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can adopt alternative configurations based on the embodiments in this specification to implement the application in this specification. Therefore, the embodiments of this specification are not limited to the embodiments precisely described in the application.
Claims
1. A method for scaling an application, which is applied to a target scaling system, and the method includes: Receiving a main configuration file submitted by a target user for a target application, where the main configuration file includes metric constraint conditions of the target application and a custom algorithm identifier; Invoking a custom algorithm model external to the target scaling system based on the algorithm identifier, so that the custom algorithm model determines the desired number of replicas of the target application under the constraint of the metric constraint conditions, and the custom algorithm model is a model generated for the target application; And Changing the number of replicas of the target application to the desired number of replicas to achieve scaling of the target application.
2. The method according to claim 1, wherein, The metric constraint conditions include conditional metrics, and the custom algorithm model is pre-trained based on historical values of the conditional metrics, historical values of operation metrics of the target application, and corresponding historical numbers of replicas.
3. The method according to claim 2, wherein The metric constraint conditions further include target values of the conditional metrics, and the desired number of replicas is determined by the following method: Inputting the metric constraint conditions into the custom algorithm model, where the custom algorithm model is used to obtain the current value of the operation metric and output the desired number of replicas based on the target value of the conditional metric and the current value of the operation metric.
4. The method according to claim 2, wherein The metric constraint conditions further include target values of the conditional metrics, the main configuration file further includes prediction configuration information, and the custom algorithm model is used to predict the desired number of replicas corresponding to multiple future moments under the constraint of the metric constraint conditions according to the prediction configuration information, where, for the desired number of replicas corresponding to each future moment, the custom algorithm model is used to obtain the predicted value of the operation metric at the future moment and determine the desired number of replicas corresponding to the future moment based on the predicted value of the operation metric and the target value of the conditional metric.
5. The method according to claim 4, wherein, The changing the number of replicas of the target application to the desired number of replicas includes: For each future moment and its corresponding desired number of replicas, changing the number of replicas of the target application to the desired number of replicas before the future moment.
6. The method according to claim 1, further including: Receiving the main configuration file through a management control module and creating an algorithm configuration file based on the main configuration file, where the algorithm configuration file includes a first declaration field for writing the metric constraint conditions and the algorithm identifier; Identifying the algorithm configuration file through an algorithm control module, invoking the custom algorithm model based on the algorithm identifier, and writing the desired number of replicas determined by the custom algorithm model under the constraint of the metric constraint conditions into a first status field in the algorithm configuration file; And Identifying the desired number of replicas in the first status field through the management control module.
7. The method according to claim 1, wherein The main configuration file further includes a custom batch gray-scale policy, and the batch gray-scale policy includes changing the number of replicas of the target application in batches, The changing the number of replicas of the target application to the desired number of replicas includes: Obtaining the current number of replicas of the target application; Determine the target number of replicas and its corresponding change method based on the current number of replicas and the desired number of replicas; and Change the replicas in batches according to the batch granularity and the change method until the number of changed replicas reaches the target number of replicas, and the number of replicas after the change is the desired number of replicas.
8. The method according to claim 7, wherein, The change method includes deleting replicas,[ The changing of the replicas in batches according to the batch granularity and the change method includes: Remove the traffic of each batch of replicas in sequence according to the batch granularity; and Delete each batch of replicas when the traffic removal of each batch of replicas is completed.
9. The method according to claim 8, wherein, For each batch of replicas, if the target application has a traffic anomaly after the traffic of the current batch of replicas is removed, then mount traffic to the current batch of replicas.
10. The method according to claim 7, wherein the change method includes deleting replicas,[ The changing of the replicas in batches according to the batch granularity and the change method includes: Control the process of each batch of replicas to sleep in sequence according to the batch granularity; And Delete each batch of replicas when the process sleep of each batch of replicas is completed.
11. The method according to claim 10, wherein, For each batch of replicas, if the target application has an anomaly after the process of the current batch of replicas sleeps, then wake up the process of the current batch of replicas.
12. The method according to claim 7, wherein, The changing of the replicas in batches according to the batch granularity and the change method includes: Delete each batch of replicas in sequence according to the batch granularity; or Create each batch of replicas in sequence according to the batch granularity.
13. The method according to claim 7, further comprising: Obtain the current number of replicas of the target application through the management control module, and determine the target number of replicas and its corresponding change method based on the current number of replicas and the desired number of replicas; For each batch of replicas, create a replica configuration file corresponding to the current batch of replicas through the management control module, and the replica configuration file includes a second declaration field for writing the batch granularity and the change method; Identify the replica configuration file through the execution control module, change the current batch of replicas according to the batch granularity and the change method, and write the change result of the current batch of replicas into the second status field in the replica configuration file, and the change result includes successful change or failed change; And Identify the change result in the second status field through the management control module.
14. A computing device for scaling an application up or down, comprising: At least one storage medium storing a target scaling system for implementing scaling an application up or down; And At least one processor communicatively connected to the at least one storage medium, Wherein when the computing device for scaling an application up or down runs, the at least one processor reads the target scaling system and implements the method for scaling an application up or down according to any one of claims 1-13.
15. A computer-readable non-transitory storage medium, wherein, The computer-readable non-transitory storage medium stores a target scaling system, and when the target scaling system is executed by at least one processor, the method for scaling an application up or down according to any one of claims 1-13 is implemented.