Elastic scaling method and device based on multi-cluster multi-application bidirectional matching
By employing a bidirectional matching elastic scaling method across multiple clusters and applications, combined with score-based scheduling, the resource scheduling challenge in multi-cluster environments is solved, achieving efficient and automated resource management and scheduling, and improving the automation level of the container platform.
Patent Information
- Application Number
- CN202510822072.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-10-17
AI Technical Summary
In a multi-cluster environment, existing technologies struggle to effectively schedule resources, meet business service quality requirements, reduce cluster resource shortages and operating costs, improve resource utilization, and simultaneously reduce the frequency of manual maintenance interventions and enhance automation levels.
An elastic scaling method with bidirectional matching across multiple clusters and applications is adopted. By determining candidate schemes for applications and clusters, dynamic adaptive scheduling is performed based on the score value. Combined with indicators such as heterogeneous performance, scheduling cost, degree of dispersion and reliability, efficient resource allocation is achieved.
It enables highly available, efficient, and fair dynamic elastic scheduling of applications across clusters and architectures, simplifies the work of operations and maintenance personnel, and improves the automation level of the container platform.
Smart Images

Figure CN120803696A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of resource scheduling, in particular to an elastic scaling method and device based on multi-cluster and multi-application bidirectional matching. BACKGROUND
[0002] Kubernetes (K8S) as an open source and scalable container orchestration and management platform, is committed to simplifying, automating and expanding the deployment and management process of container applications in data centers. It has excellent disaster recovery and self-healing capabilities, clear automated container management logic, can realize fast version release, rollback and horizontal scaling operation, and has native service discovery and load balancing technology. This not only makes it the de facto standard of container platforms, but also promotes the vigorous development and landing of microservice architecture cloud native applications.
[0003] With the continuous expansion of business on the cloud, the scale of the container platform is expanding, and the number of container clusters that need to be managed is increasing. Under this background, how to reasonably schedule resources for various applications in a large number of cross-platform container clusters, not only to meet the quality of service requirements of business, but also to reduce the risk of insufficient cluster resources and operating costs, improve platform resource utilization, while reducing the frequency of manual operation and maintenance intervention and improving the level of automation, has become a complex decision-making problem. SUMMARY
[0004] The present application provides an elastic scaling method and device based on multi-cluster and multi-application bidirectional matching, to provide a dynamic adaptive scheduling and elastic scaling scheme, simplify the work of operation and maintenance personnel in application resource configuration adjustment and multi-cluster management, and improve the automation level of the container platform.
[0005] In a first aspect, the present application provides an elastic scaling method based on multi-cluster and multi-application bidirectional matching, the method comprising:
[0006] For each application, determine each first candidate scheme corresponding to the application; for each first candidate scheme, determine a first score value corresponding to the first candidate scheme according to the number of replicas of the application deployed in each cluster included in the first candidate scheme;
[0007] For each cluster, determine each second candidate scheme corresponding to the cluster; for each second candidate scheme, determine a second score value corresponding to the second candidate scheme according to the number of replicas of each application deployed in the cluster included in the second candidate scheme;
[0008] determine a target scheme with the highest score according to the first score values corresponding to each first candidate scheme and the second score values corresponding to each second candidate scheme; and perform deployment of application replicas of each application according to the number of replicas of each application on each cluster in the target scheme.
[0009] The technical scheme has the following advantages or beneficial effects:
[0010] In the present application, for each first candidate scheme of each application, a first score value corresponding to the first candidate scheme is determined according to the number of replicas of the application included in the first candidate scheme and deployed in each cluster. For each second candidate scheme of each cluster, a second score value corresponding to the second candidate scheme is determined according to the number of replicas of each application deployed in the cluster included in the second candidate scheme. Finally, a target scheme with the highest score is determined by combining the first score values corresponding to each first candidate scheme and the second score values corresponding to each second candidate scheme, and then deployment of application replicas of each application is performed according to the target scheme. Thus, a method for elastic scaling based on two-way matching of multiple clusters and multiple applications is realized, dynamic adaptive scheduling and elastic scaling are achieved, the work of application resource configuration adjustment and multi-cluster management of operation and maintenance personnel is simplified, and the automation level of the container platform is improved.
[0011] In a second aspect, the present application provides a device for elastic scaling based on two-way matching of multiple clusters and multiple applications, which comprises:
[0012] A first determination module is configured to determine, for each application, each first candidate scheme corresponding to the application; and determine, for each first candidate scheme, a first score value corresponding to the first candidate scheme according to the number of replicas of the application included in the first candidate scheme and deployed in each cluster.
[0013] A second determination module is configured to determine, for each cluster, each second candidate scheme corresponding to the cluster; and determine, for each second candidate scheme, a second score value corresponding to the second candidate scheme according to the number of replicas of each application deployed in the cluster included in the second candidate scheme.
[0014] An application deployment module is configured to determine a target scheme with the highest score according to the first score values corresponding to each first candidate scheme and the second score values corresponding to each second candidate scheme; and perform deployment of application replicas of each application according to the number of replicas of each application on each cluster in the target scheme.
[0015] In a third aspect, the present application provides an electronic device comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete communication with each other through the communication bus.
[0016] a memory for storing a computer program;
[0017] a processor for implementing the method when executing the program stored in the memory.
[0018] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method.
[0019] In a fifth aspect, the present application provides a computer program product, which comprises an executable program, and the executable program is executed by a processor to implement the method. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort based on these drawings.
[0021] Figure 1 a process diagram for determining the first score value corresponding to the first candidate scheme provided by the present application;
[0022] Figure 2 a process diagram for determining the first score value corresponding to the first candidate scheme provided by the present application;
[0023] Figure 3 a process diagram for determining the second score value corresponding to the second candidate scheme provided by the present application;
[0024] Figure 4 a process diagram for determining the first score value corresponding to the first candidate scheme provided by the present application;
[0025] Figure 5 a process diagram for determining the first score value corresponding to the first candidate scheme provided by the present application;
[0026] Figure 6 a process diagram for determining the first score value corresponding to the first candidate scheme provided by the present application;
[0027] Figure 7 a process diagram for determining the first score value corresponding to the first candidate scheme provided by the present application;
[0028] Figure 8 a process diagram for determining the first score value corresponding to the first candidate scheme provided by the present application; DETAILED DESCRIPTION
[0029] In order to make the objectives and the embodiments of the present application clearer, the following will clearly and completely describe the exemplary embodiments of the present application with reference to the accompanying drawings. Obviously, the described exemplary embodiments are only a part of the embodiments of the present application, but not all of them.
[0030] It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the subsequently described embodiments, and is not intended to limit the embodiments of the present application. Unless otherwise specified, these terms should be understood according to their ordinary and general meanings.
[0031] The terms "first", "second", "third", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar or identical objects or entities, and do not necessarily mean a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchanged under appropriate circumstances.
[0032] The terms "include" and "have" and any variations thereof are intended to cover but not exclusive inclusion, for example, a product or device including a series of components does not have to be limited to all the components clearly listed, but can include other components not clearly listed or inherent to these products or devices.
[0033] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware or / and software code capable of performing a function associated with the element.
[0034] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent substitutions for some or all of the technical features; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
[0035] For the convenience of explanation, the above description has been made in conjunction with specific embodiments. However, the above exemplary discussion is not intended to exhaust or limit the embodiments to the specific forms disclosed above. Various modifications and variations can be derived according to the above teachings. The selection and description of the above embodiments are for better explanation of the principles and practical applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific use considerations.
[0036] The English abbreviations involved in the present application are explained as follows:
[0037] Kubernetes (K8S): an open-source system for automatically deploying, scaling, and managing "containerized applications";
[0038] Horizontal Pod Autoscaling (HPA): a K8S native scaling method that automatically adjusts the number of application Pods within a specified upper and lower limit through certain metrics.
[0039] In related technologies, Kubernetes takes Pod as the smallest deployment unit of an application, and natively supports Horizontal Pod Autoscaling (HPA) and Vertical Pod Autoscaling (VPA) based on Pods, which to some extent meets the elastic scaling needs of applications and achieves good platform resource management. These two mechanisms adjust the number of Pod replicas and resource configurations by monitoring the resource occupation of Pods and running nodes, and according to the manually set adjustment criteria, realize the expansion and recovery of application resources, and meet the reasonable resource supply of applications in busy and idle periods.
[0040] In recent years, many automatic elastic scaling schemes have been proposed to improve the decision mechanism of HPA or VPA. These schemes use more comprehensive monitoring indicators, more sophisticated metric structures, and more complex decision models to achieve more accurate and effective elastic scaling of application load. Among them, AHP4HPA is representative, which uses a series of evaluation indicators and a hierarchical analysis method to sort the possible elastic scaling schemes in a period of time, thereby determining the application specification in real time.
[0041] The Karmada project is an open-source multi-cluster Kubernetes scheduler that aims to provide unified scheduling, deployment, and management services for applications and workloads across multiple Kubernetes clusters. Karmada supports resource-based multi-cluster application scheduling, which reasonably allocates workloads across clusters by considering the resource status of different clusters.
[0042] Although the HPA method supported by the Kubernetes community alleviates the contradiction between the relatively fixed resource capacity of the platform and the changing resource needs of the application to some extent, the rationality and timeliness of the scaling decision often face challenges. HPA relies on a series of parameters set by hand to control the scaling behavior, and reasonable HPA rules need to be determined based on specific application analysis and a large amount of observation experience, which not only increases the difficulty of HPA definition, consumes human and time costs, but also makes HPA prone to improper behavior.
[0043] VPA also relies on manual parameter tuning and is a passive scaling adjustment method with fixed rules. In addition, due to the limitations of Kubernetes features, when VPA adjusts the resource request and limit of a pod, it needs to be done by driving existing pods and rebuilding new pods, which leads to complex pod switching process during scaling, large time and resource overhead, and risk of business interruption.
[0044] The community's native auto-scaling method is limited to resource adjustment and scheduling of a single cluster and a single application, and each application is independently managed, which cannot evaluate the impact of application adjustment decisions on the cluster and even the entire platform from a higher dimension.
[0045] Some complex decision-making models proposed by the industry-academia-research community elevate the decision-making problem to the entire cluster level and make periodic cluster-level overall decisions. This method improves the accuracy of application scaling and scheduling, and to some extent avoids decision conflicts between applications, but in the common multi-cluster platform management scenario, it cannot provide an ideal application deployment solution for the entire platform, and the actual application value is limited.
[0046] The Karmada project is aimed at multi-cluster management scenarios, and can monitor multiple cluster resources and adjust the number of replicas of applications on each cluster using the FederatedHPA method on the managed clusters. Although it realizes cross-cluster application elasticity, it lacks comprehensive analysis in multiple application dimensions and belongs to the single-application multi-cluster control method, which is difficult to balance between applications.
[0047] The present application considers the above problems and proposes a dynamic adaptive scheduling and elastic scaling scheme with cross-cluster and multi-application linkage characteristics, which simplifies the work of operation and maintenance personnel in application resource configuration adjustment and multi-cluster management, and further improves the automation level of the container platform.
[0048] The present application proposes a two-way optimal matching mechanism suitable for multiple applications and multiple clusters: using the perception decision agent components of clusters and applications and the arbitration processing components on the platform cluster, high availability, high efficiency, and high fairness of dynamic elastic scheduling of multiple applications on multiple clusters are realized. Through multi-dimensional comprehensive analysis to determine the real-time scheduling scheme: a grid search resource load relationship analysis method based on online real-time performance history records, an application load prediction based on an autoregressive model, and an application scheduling scheme evaluation index system based on an energy consumption model.
[0049] The present application constructs an application scheduling and elastic scaling decision architecture in multiple application and multiple cluster dimensions, adopts a multiple-to-multiple two-way matching decision method of applications and clusters, and ensures that a large number of applications can achieve relatively ideal and fair resource allocation in a cross-cluster and cross-architecture environment.
[0050] By analyzing the application resource performance relationship, heterogeneous performance relationship, combining with request quantity prediction, energy consumption estimation and affinity state evaluation, the advantages and disadvantages of application proxy candidate decision are quantified, and the rationality, comprehensiveness and explainability of the decision system are improved from the application perspective.
[0051] Figure 1 The elastic scaling process based on the multi-cluster multi-application bidirectional matching provided in the application includes the following steps:
[0052] S101: For each application, determine the corresponding first candidate scheme of the application; for each first candidate scheme, determine the first score value corresponding to the first candidate scheme according to the number of replicas of the application deployed in each cluster included in the first candidate scheme;
[0053] S102: For each cluster, determine the corresponding second candidate scheme of the cluster; for each second candidate scheme, determine the second score value corresponding to the second candidate scheme according to the number of replicas of each application deployed in the cluster included in the second candidate scheme;
[0054] S103: Determine the target scheme with the highest score according to the first score value corresponding to each first candidate scheme and the second score value corresponding to each second candidate scheme; deploy the application replicas of each application according to the number of replicas of each application on each cluster in the target scheme.
[0055] The elastic scaling method based on the multi-cluster multi-application bidirectional matching provided in the application is applied to an electronic device, which can be a computer, a terminal, or the like, or a server.
[0056] In the application, for each application, first determine the corresponding first candidate scheme of the application. Among them, according to the pre-saved cluster architecture restrictions, disaster recovery requirements, SLA protocol requirements, etc. of the application, determine the corresponding first candidate scheme of the application. The first candidate scheme includes the number of replicas of the application deployed in each cluster. The first candidate scheme needs to meet the cluster architecture restrictions, disaster recovery requirements, and SLA protocol requirements of the application. The cluster architecture restrictions include Central Processing Unit (CPU), Graphics Processing Unit (GPU), etc. The disaster recovery requirements include cross-regional deployment of application replicas, cross-machine room deployment of application replicas, etc. The SLA protocol refers to the Service Level Agreement (SLA), which is a kind of agreement jointly defined and recognized by service providers and users or service providers within a certain cost range to ensure service performance and reliability.
[0057] For each application, determine each first candidate solution corresponding to the application; then for each first candidate solution corresponding to the application, determine a first score value corresponding to the first candidate solution according to the number of replicas of the application respectively deployed in each cluster included in the first candidate solution.
[0058] Optionally, determine a heterogeneous performance score value corresponding to the first candidate solution according to the sum value of the number of replicas of the application respectively deployed in each cluster included in the first candidate solution; determine a first scheduling score value corresponding to the first candidate solution according to the first distance value of the number of replicas of the application respectively deployed in each cluster included in the first candidate solution and the number of replicas of the application respectively deployed in each cluster at the previous moment; determine an application dispersion score value corresponding to the first candidate solution according to the entropy of the number of replicas of the application respectively deployed in each cluster included in the first candidate solution; determine a reliability score value corresponding to the first candidate solution by weighted summation according to the number of replicas of the application respectively deployed in each cluster included in the first candidate solution and the first weight value corresponding to each cluster respectively pre-saved.
[0059] In the present application, any one of the heterogeneous performance score value, the first scheduling score value, the application dispersion score value and the reliability score value can be taken as the first score value; or any two or more of the heterogeneous performance score value, the first scheduling score value, the application dispersion score value and the reliability score value are weighted and summed to obtain the first score value.
[0060] For each cluster, determine each second candidate solution corresponding to the cluster; wherein each second candidate solution corresponding to the cluster is determined according to the pre-saved cluster architecture restrictions, disaster recovery requirements, SLA agreement requirements, etc. of each application respectively. The second candidate solution includes the number of replicas of each application deployed in the cluster. The second candidate solution needs to meet the cluster architecture restrictions, disaster recovery requirements, SLA agreement requirements of the application.
[0061] For each cluster, determine each second candidate solution corresponding to the cluster, and for each second candidate solution corresponding to the cluster, determine a second score value corresponding to the second candidate solution according to the number of replicas of each application deployed in the cluster included in the second candidate solution.
[0062] Optionally, according to the number of replicas of each application deployed in the cluster included in the second candidate scheme and the second weight value corresponding to each application respectively, the priority score value corresponding to the second candidate scheme is determined by weighted summation; according to the number of replicas of each application deployed in the cluster included in the second candidate scheme and the resource consumption amount corresponding to the replicas of each application respectively, the total resource consumption amount corresponding to the second candidate scheme is determined; according to the ratio of the total resource consumption amount to the number of application types deployed in the cluster included in the second candidate scheme, the resource consumption score value corresponding to the second candidate scheme is determined; according to the second distance value of the number of replicas of each application deployed in the cluster included in the second candidate scheme and the number of replicas of each application deployed in the cluster at the last moment, the second scheduling score value corresponding to the second candidate scheme is determined.
[0063] In this application, any one of the priority score value, the resource consumption score value and the second scheduling score value can be used as the second score value; or any two or more of the priority score value, the resource consumption score value and the second scheduling score value are weighted and summed to obtain the second score value.
[0064] Finally, according to the first score value corresponding to each first candidate scheme and the second score value corresponding to each second candidate scheme, the target scheme with the highest score is determined; the target scheme includes the number of replicas of each application on each cluster. According to the target scheme, the deployment of application replicas of each application on each cluster is performed.
[0065] In this application, for each first candidate scheme of each application, the first score value corresponding to the first candidate scheme is determined according to the number of replicas of the application deployed in each cluster included in the first candidate scheme; for each second candidate scheme of each cluster, the second score value corresponding to the second candidate scheme is determined according to the number of replicas of each application deployed in the cluster included in the second candidate scheme; finally, the target scheme with the highest score is determined by combining the first score value corresponding to each first candidate scheme and the second score value corresponding to each second candidate scheme, and then the deployment of application replicas of each application is performed according to the target scheme. Thus, a multi-cluster multi-application bidirectional matching based elastic scaling method is realized, dynamic adaptive scheduling and elastic scaling are realized, the work of application resource configuration adjustment and multi-cluster management of operation and maintenance personnel is simplified, and the automation degree of the container platform is improved.
[0066] Figure 2 The process diagram provided in this application for determining the first score value corresponding to the first candidate scheme includes the following steps:
[0067] S201: Determine a heterogeneous performance score value corresponding to the first candidate scheme according to the sum of the number of replicas of the application included in the first candidate scheme deployed in each cluster.
[0068] The heterogeneous performance score value is represented as: is the number of replicas of the application i deployed in each cluster in the first candidate scheme k at the current time t. is the number of replicas of the application i deployed in the cluster j in the first candidate scheme k at the current time t.
[0069] S202: Determine a first scheduling score value corresponding to the first candidate scheme according to the first distance value between the number of replicas of the application included in the first candidate scheme deployed in each cluster and the number of replicas of the application respectively deployed in each cluster at the last time.
[0070] The first scheduling score value is represented as: is the number of replicas of the application i deployed in each cluster at the last time t-1.
[0071] S203: Determine an application dispersion score value corresponding to the first candidate scheme according to the entropy of the number of replicas of the application included in the first candidate scheme respectively deployed in each cluster.
[0072] The application dispersion score value is represented as: H represents the entropy operation.
[0073] S204: Determine a reliability score value corresponding to the first candidate scheme by weighted summation according to the number of replicas of the application included in the first candidate scheme respectively deployed in each cluster and the first weight value corresponding to each cluster respectively pre-saved.
[0074] The reliability score value is represented as: w j is the first weight value corresponding to the cluster j.
[0075] S205: Determine a first score value corresponding to the first candidate scheme according to the heterogeneous performance score value, the first scheduling score value, the application dispersion score value and the reliability score value corresponding to the first candidate scheme.
[0076] The first score value is represented as: α perf , α scale , α dis , α cls is the weight coefficient corresponding to each score value.
[0077] Figure 3 is the process schematic diagram provided by the present application for determining a second score value corresponding to a second candidate scheme, comprising the following steps:
[0078] S301: determining a priority score value corresponding to the second candidate scheme according to the number of replicas of each application deployed in the cluster included in the second candidate scheme and a second weight value corresponding to each application respectively and pre-stored;
[0079] The priority score value is expressed as: is the number of replicas of application i deployed in cluster j in the second candidate scheme m at the current time t. i is the second weight value corresponding to application i.
[0080] S302: determining a total resource consumption amount corresponding to the second candidate scheme according to the number of replicas of each application deployed in the cluster included in the second candidate scheme and a resource consumption amount corresponding to each replica of the application respectively and pre-determined; determining a resource consumption score value corresponding to the second candidate scheme according to a ratio of the total resource consumption amount to the number of application types deployed in the cluster included in the second candidate scheme;
[0081] The resource consumption score value is expressed as: wherein, Qi is the resource consumption amount of a single replica of application i, represents the number of application types deployed in the cluster included in the second candidate scheme m at the current time t.
[0082] S303: determining a second scheduling score value corresponding to the second candidate scheme according to the number of replicas of each application deployed in the cluster included in the second candidate scheme and a second distance value of the number of replicas of each application deployed in the cluster at the last time;
[0083] The second scheduling score value is expressed as wherein, is the number of replicas of each application deployed in cluster j included in the second candidate scheme m at the current time t, is the number of replicas of each application deployed in cluster j at the last time t-1.
[0084] S304: determining a second score value corresponding to the second candidate scheme according to the priority score value, the resource consumption score value and the second scheduling score value corresponding to the second candidate scheme.
[0085] The second score value is expressed as: α app , α res , α sche is a weight coefficient of each unilateral score value.
[0086] In this application, determining each first candidate solution corresponding to the application includes:
[0087] Get the target number of visits to the application at the current moment, and determine the target number of copies required when the target number of visits corresponds to the target number of visits and the number of copies required when the application is fully running on the jth cluster based on the pre-saved correspondence between the target number of visits to the application and the number of copies required when the application is fully running on the jth cluster.
[0088] Through the constraints Determine each first candidate solution corresponding to the application
[0089] in, is the kth first candidate solution for a single application i at time t; It represents the number of replicas that the k-th first candidate solution decides to deploy on the j-th cluster for application i at time t.
[0090] That is, according to the cluster architecture limitations, disaster recovery requirements, SLA agreement requirements, and constraints of the pre-saved application Determine each first candidate solution corresponding to the application.
[0091] In order to reduce energy consumption and improve the efficiency of elastic scaling based on multi-cluster multi-application bidirectional matching, in this application, after determining the first score corresponding to the first candidate solution and before determining the target solution with the highest score, the method further includes:
[0092] In descending order of the first scoring values, a first number of first candidate solutions with larger first scoring values and corresponding first scoring values are retained, and other first candidate solutions and corresponding first scoring values are deleted.
[0093] After determining the second score corresponding to the second candidate solution and before determining the target solution with the highest score, the method further includes:
[0094] In descending order of the second scoring values, a second number of second candidate solutions with larger second scoring values and corresponding second scoring values are retained, and other second candidate solutions and corresponding second scoring values are deleted.
[0095] That is to say, the target solution with the highest score is determined based on the first score values corresponding to the first number of first candidate solutions with larger first score values and the second score values corresponding to the second number of second candidate solutions with larger second score values; and the application copies of each application are deployed based on the number of copies of each application on each cluster in the target solution.
[0096] In the present application, the target scheme with the highest score is determined according to the first score value corresponding to each first candidate scheme and the second score value corresponding to each second candidate scheme, comprising:
[0097] According to the first score value corresponding to each first candidate scheme and the second score value corresponding to each second candidate scheme, the target scheme with the highest score is determined by a constraint condition and x ik +y jm ≤1.
[0098] Wherein, is the first score value corresponding to the kth first candidate scheme, is the second score value corresponding to the mth second candidate scheme, x ik ∈{0,1} represents whether the ith application adopts the kth first candidate scheme, y jm ∈{0,1} represents whether the jth cluster adopts the mth second candidate scheme, and a∈[0,1] is a preset preference coefficient.
[0099] The constraint condition represents that only one candidate scheme is selected for any application or cluster, and the kth first candidate scheme adopted by the ith application and the mth scheme adopted by the jth cluster cannot be satisfied at the same time, and the constraint condition x ik +y jm ≤1 must be met.
[0100] Figure 4 The present application provides a multi-cluster and multi-application bidirectional matching elastic scaling framework diagram, as shown in Figure 4 , which comprises a business cluster m, a control cluster. The business cluster m comprises a cluster controller and a cluster m agent (proxy), the control cluster comprises a platform controller, an arbitration decision maker and an application controller, and the application controller comprises an application n agent (proxy).
[0101] The present application considers a cloud container platform with a platform management cluster that can manage multiple business clusters, including monitoring the running state of the business cluster, issuing and adjusting the resource request, limit and replica number of the workload. There are multiple applications with different attributes on each business cluster, and the cluster architecture limit, priority, disaster recovery requirement, and SLA of different applications are different. When each application declares application attributes in the platform management plane, an application agent process is generated in the application controller, which independently evaluates the preferred deployment cluster of the application according to the pre-set evaluation index, and sorts different deployment strategies. Each cluster pulls the list of applications to be run from the platform side, and sorts the preferred applications. The application agent and the business cluster agent report the proposal to the platform layer arbitration decision maker, which decides the deployment and elastic scaling scheme of the application in the period, and feeds back the result to the application and cluster agent for state update, solves the many-to-many stable matching problem, and controls the cluster controller of the business cluster to run the application load according to the arbitration decision scheme.
[0102] Figure 5 The application agent processing flowchart provided by the present application includes: setting user SLA, priority, architecture limit; application resource performance characterization; application access pressure collection; determining whether the current deployment matches the application SLA, if yes, the application agent work cycle ends, if not, obtaining cluster information from the platform controller; filtering out clusters that do not match the resource demand or architecture; sorting the target deployment cluster using the decision evaluation model; reporting the preferred deployment scheme to the arbitrator; receiving the arbitration result and updating the current deployment scheme of the application; the application agent work cycle ends.
[0103] The working process of the application agent (agent) in a period is as shown in Figure 5 Before the application goes online, the SLA of the application is determined, and the application is stress tested in different running architectures and resource quota clusters through grid search, the performance limit under the given condition is recorded, and finally the load-resource relationship is characterized by linear fitting. The relationship can be expressed as a function:
[0104]
[0105] θ i (t) represents the access amount of the i-th application at time t, represents the number of replicas required if the application is completely run on the j-th cluster at this time. λ ij That is, the load-resource mapping of application i on the j-th cluster.
[0106] At time t, for a single application i, the k-th available decision (candidate scheme) can be expressed as:
[0107]
[0108] wherein represents the number of replicas of the application i decided by the kth decision at time t to deploy on the jth cluster. J is the number of all clusters that the application is adapted to. In order to meet the load-resource relationship, the decision satisfies the constraint:
[0109]
[0110] The application sets the application scheduling and scaling evaluation index to include:
[0111] Heterogeneous performance index: the running efficiency of the same application on different architecture and hardware level clusters is different. In the application performance description stage, the difference has been reflected in the load-resource mapping, and here the number of replicas of the decision can reflect the good or bad of the heterogeneous performance, that is, the less the number of replicas required, the better the heterogeneous performance.
[0112]
[0113] Scaling scheduling cost: the scaling scheduling behavior will generate running overhead in the platform and increase the instability risk of the platform. Setting this index aims to consider the prudence of the elastic scaling decision. For the replicas increased or decreased in the scheduling period, the cluster score list is adjusted by a penalty coefficient to reduce or increase the tendency to all clusters, so as to control the scaling cost. The scaling cost score is as follows, wherein is the current situation of the application i at the last time.
[0114]
[0115] Application dispersion level: in order to improve the high availability and disaster recovery capability of the application, and avoid the application replicas concentrated in a single or a small number of business clusters, it is necessary to evaluate the dispersion degree of the application deployment scheme. The dispersion degree of the replica of the decision scheme is evaluated by calculating the entropy of the decision.
[0116]
[0117] Reliability and energy consumption level of the cluster: according to the historical running level, scale and energy consumption level of the cluster, the cluster with high weight is artificially set, so that the application is preferentially applied to schedule to the business cluster with high service quality and better energy consumption level. Let the given cluster weight of the jth cluster be w j , then the decision score of this item is:
[0118]
[0119] The application agent scores the deployment and scaling scheme of the application by using the above index, and the score of the kth decision of the ith application at time t is:
[0120]
[0121] wherein a perf , a scale , a dis , a cls are the coefficients of each scoring item. The application agent sorts the decision scores and selects the top K schemes with the highest scores to report to the arbitration module. After the arbitration module responds to the comprehensive decision result, the application agent updates the latest deployment status of the application.
[0122] Figure 6 The cluster agent processing flowchart provided in the present application includes: cluster agent work begins; obtaining application information from the platform controller; filtering out clusters that do not match the resource needs or framework; sorting the application proposal by using application priority, working time, and resource demand; reporting the preferred deployment scheme to the arbitrator; receiving the arbitration result and updating the application deployment scheme; notifying the cluster management and control to complete the load expansion according to the request; and ending the cluster agent work cycle.
[0123] The working process of the cluster agent in a decision cycle is shown in Figure 6 . Its working mode is similar to that of the application agent. First, it pulls the application list carried by the platform, filters out the preferred applications that do not match the platform attributes, and then proposes deployment schemes for the adapted applications. For time t, for cluster j, the mth available decision can be represented as:
[0124]
[0125] wherein represents the number of replicas of application i decided to be deployed on the jth cluster by the mth decision at time t. I is the number of all applications adapted to the cluster.
[0126] The deployment schemes of the remaining applications on the cluster are sorted according to the following criteria.
[0127] Application priority: Different applications on the platform have different importance. The business cluster prioritizes the scheduling of high-importance clusters, so different weights can be set for the applications. Let the weight of application i be w i , then the application priority score of the decision scheme is:
[0128]
[0129] Resource consumption of a single copy: for the application with large resource request amount of a single copy, the number of copies deployed in a single cluster should be reduced as much as possible to promote the dispersion and fairness of the business on the business cluster. To evaluate this dimension, the average value of resource consumption of each application in the cluster is taken as the score part, and the punishment term is used to avoid providing too many copies for large resource request amount.
[0130]
[0131] wherein Qi is the resource consumption of a single copy of application i;
[0132]
[0133] Currently running applications on the cluster at the current time: for the existing running applications on the cluster, the resources such as data information and network link already exist in the cluster, and the cluster tends to ensure that the existing applications are not completely relocated, causing excessive deployment changes.
[0134]
[0135] The cluster agent scores the candidate applications according to the above evaluation indexes. For the mth scheme of the jth cluster at the tth time The score is:
[0136]
[0137] wherein α app , α res , and α sche are the coefficients of each unilateral scoring item. The cluster agent sorts the scores of each decision and selects the top M schemes with the highest scores to report to the arbitration module. After the arbitration module responds to the comprehensive decision result, the cluster agent adjusts the deployment of applications in the cluster according to the final scheduling scheme fed back by the arbitration module. The actual running state of the applications in the cluster is consistent with the expected state of the scheme.
[0138] After the arbitration module accepts the proposals of the application agent and the cluster agent, it comprehensively considers the schemes and scores of the two parties, and finally obtains the application deployment state on the entire platform, i.e., all business clusters controlled by the cluster management. The problem solved by the arbitration module can be expressed as:
[0139] Let x ik ∈{0,1} represent whether the ith application adopts the kth scheme, and y jm ∈{0,1} represent whether the jth cluster accepts the mth application absorption scheme. Then the objective function at the tth time is to solve:
[0140]
[0141] Wherein a is a preference coefficient in [0, 1], adjusting the optimization problem to consider the dimension of application or cluster. At the same time, there are constraint conditions:
[0142]
[0143] That is, it is required to select only one scheme for any application or cluster. At the same time, the arbiter needs to analyze the consistency constraint according to the scheme proposed by the application and the cluster. The scheme k proposed by the application i and the scheme m proposed by the cluster j are contradictory or cannot be satisfied at the same time:
[0144] x ik +y jm ≤1.
[0145] According to the above constraints and optimization objectives, the arbiter obtains the value of x ik , y jm , that is, determines the optimal application deployment scheme in the current period t. After the arbiter responds to the deployment decision, the cluster agent controls the cluster controller to change the deployment state of the application on the cluster according to the decision, and the application agent records the updated deployment state of the application.
[0146] The elastic scaling method based on two-way matching of multiple clusters and multiple applications provided in the application can be replaced by other algorithms and models in resource load analysis and decision model aspects. At this time, the information collection system of the application and the cluster, the decision mechanism and the entity framework will change, such as directly combining the evaluation indexes on both sides of the application and the cluster, and making a single decision. At this time, the candidate capacity of the decision is greatly expanded, and the establishment and calculation complexity of the evaluation model will be greatly improved.
[0147] Figure 7 The elastic scaling device structure diagram based on two-way matching of multiple clusters and multiple applications provided in the application, comprising:
[0148] The first determination module 11 is configured to determine each first candidate scheme corresponding to each application for each application; and for each first candidate scheme, determine a first score value corresponding to the first candidate scheme according to the number of replicas of the application in each cluster included in the first candidate scheme.
[0149] The second determination module 12 is configured to determine each second candidate scheme corresponding to each cluster for each cluster; and for each second candidate scheme, determine a second score value corresponding to the second candidate scheme according to the number of replicas of each application deployed in the cluster included in the second candidate scheme.
[0150] The application deployment module 13 is configured to determine a target scheme with the highest score according to the first score value corresponding to each first candidate scheme and the second score value corresponding to each second candidate scheme, and deploy application replicas of each application according to the number of replicas of each application on each cluster in the target scheme.
[0151] The first determining module 11 is specifically configured to determine a heterogeneous performance score value corresponding to the first candidate scheme according to the sum of the number of replicas of the applications respectively deployed in each cluster in the first candidate scheme, determine a first scheduling score value corresponding to the first candidate scheme according to the first distance value between the number of replicas of the applications respectively deployed in each cluster in the first candidate scheme and the number of replicas of the applications respectively deployed in each cluster at the previous moment, determine an application dispersion score value corresponding to the first candidate scheme according to the entropy of the number of replicas of the applications respectively deployed in each cluster in the first candidate scheme, determine a reliability score value corresponding to the first candidate scheme by weighted summation according to the number of replicas of the applications respectively deployed in each cluster in the first candidate scheme and the first weight value corresponding to each cluster, and determine the first score value corresponding to the first candidate scheme according to the heterogeneous performance score value, the first scheduling score value, the application dispersion score value and the reliability score value corresponding to the first candidate scheme.
[0152] The second determining module 12 is specifically configured to determine a priority score value corresponding to the second candidate scheme by weighted summation according to the number of replicas of each application deployed in each cluster in the second candidate scheme and the second weight value corresponding to each application, determine a total resource consumption amount corresponding to the second candidate scheme according to the number of replicas of each application deployed in each cluster in the second candidate scheme and the resource consumption amount corresponding to each replica of the application, determine a resource consumption score value corresponding to the second candidate scheme according to the ratio of the total resource consumption amount to the number of application types deployed in the cluster in the second candidate scheme, determine a second scheduling score value corresponding to the second candidate scheme according to the second distance value between the number of replicas of each application deployed in each cluster in the second candidate scheme and the number of replicas of each application deployed in each cluster at the previous moment, and determine the second score value corresponding to the second candidate scheme according to the priority score value, the resource consumption score value and the second scheduling score value corresponding to the second candidate scheme.
[0153] The first determination module 11 is specifically used to obtain the target access volume of the application at the current moment, and determine the target number of replicas required when the target access volume is fully running on the jth cluster based on the pre-saved correspondence between the access volume of the application and the number of replicas required when the application is fully running on the jth cluster.
[0154] Through the constraints Determine each first candidate solution corresponding to the application
[0155] in, is the kth first candidate solution for a single application i at time t; It represents the number of replicas that the k-th first candidate solution decides to deploy on the j-th cluster for application i at time t.
[0156] The first determining module 11 is further configured to retain a first number of first candidate solutions with larger first scores and their corresponding first scores in descending order of the first scores, and delete other first candidate solutions and their corresponding first scores.
[0157] The first determining module 12 is further configured to retain a second number of second candidate solutions with larger second scores and their corresponding second scores, and delete other second candidate solutions and their corresponding second scores, in descending order of the second scores.
[0158] The application deployment module 13 is specifically configured to: and x ik +y jm ≤1, determine the target solution with the highest score
[0159] in, is the first score value corresponding to the kth first candidate solution, is the second score value corresponding to the mth second candidate solution, x ik ∈{0,1} indicates whether the i-th application adopts the k-th first candidate solution, y jm ∈{0,1} indicates whether the jth cluster adopts the mth second candidate solution, and α∈[0,1] is the preset preference coefficient;
[0160] Constraints Indicates that only one candidate solution is selected for any application or cluster. When the kth and first candidate solutions adopted by application i and the mth solution adopted by cluster j cannot be satisfied at the same time, the constraint x must be satisfied. ik +yjm ≤1.
[0161] The application further provides an electronic device, such as Figure 8 as shown, comprising a processor 21, a communication interface 22, a memory 23 and a communication bus 24, wherein the processor 21, the communication interface 22 and the memory 23 complete communication with each other through the communication bus 24;
[0162] The memory 23 stores a computer program, and when the program is executed by the processor 21, the processor 21 executes any of the above method steps.
[0163] The communication bus mentioned in the above electronic device can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.
[0164] The communication interface 22 is used for communication between the above electronic device and other devices.
[0165] The memory can include a random access memory (RAM) and can also include a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located away from the aforementioned processor.
[0166] The above processor can be a general-purpose processor, including a central processing unit, a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc.
[0167] The application further provides a computer storage readable storage medium, the computer readable storage medium stores a computer program executable by an electronic device, and when the program runs on the electronic device, the electronic device executes the above any method step.
[0168] The application provides a computer program product, the computer program product includes an executable program, and the executable program is executed by the processor to realize the method.
[0169] While the preferred embodiments of the application have been described, additional variations and modifications can be made to the embodiments by those skilled in the art once they learn of the basic inventive concepts. Therefore, the appended claims are intended to encompass within their scope all such variations and modifications as are included within the spirit and scope of the application.
[0170] It is apparent that many modifications and variations of this application can be effected although only a few have been chosen for purposes of illustrative discussion. Thus, it is contemplated to cover the present application as broadly as possible. It is therefore desired that the appended claims shall cover all such modifications and variations as fall within the true scope of the application.
Claims
1. An elastic scaling method based on multi-cluster multi-application bidirectional matching, characterized in that: The method comprises: For each application, determining each first candidate solution corresponding to the application; for each first candidate solution, determining a first score corresponding to the first candidate solution based on the number of replicas of the application included in the first candidate solution deployed in each cluster; For each cluster, determining each second candidate solution corresponding to the cluster; for each second candidate solution, determining a second scoring value corresponding to the second candidate solution based on the number of replicas of each application deployed in the cluster included in the second candidate solution; Determine the target solution with the highest score based on the first score values corresponding to each first candidate solution and the second score values corresponding to each second candidate solution; and deploy application copies of each application based on the number of copies of each application on each cluster in the target solution.
2. The method according to claim 1, wherein Determining a first score corresponding to the first candidate solution according to the number of replicas of the application deployed in each cluster included in the first candidate solution includes: Determining a heterogeneous performance score corresponding to the first candidate solution according to a sum of the number of replicas of the application deployed in each cluster included in the first candidate solution; Determining a first scheduling score corresponding to the first candidate solution based on the number of replicas of the application deployed in each cluster included in the first candidate solution and a first distance value between the number of replicas of the application deployed in each cluster at a previous moment; Determining an application dispersion score corresponding to the first candidate solution according to the entropy of the number of replicas of the application deployed in each cluster included in the first candidate solution; Determining a reliability score corresponding to the first candidate solution by weighted summation based on the number of replicas of the application deployed in each cluster included in the first candidate solution and the pre-saved first weight values corresponding to each cluster; A first scoring value corresponding to the first candidate solution is determined according to the heterogeneous performance scoring value, the first scheduling scoring value, the application dispersion scoring value, and the reliability scoring value corresponding to the first candidate solution.
3. The method according to claim 1, wherein Determining, based on the number of replicas of each application deployed in the cluster included in the second candidate solution, a second scoring value corresponding to the second candidate solution includes: Determining a priority score corresponding to the second candidate solution by weighted summation based on the number of replicas of each application deployed in the cluster included in the second candidate solution and the pre-saved second weight values corresponding to each application; Determine the total resource consumption corresponding to the second candidate solution based on the number of replicas of each application deployed in the cluster included in the second candidate solution and the predetermined resource consumption corresponding to each replica of the application; determine the resource consumption score corresponding to the second candidate solution based on the ratio of the total resource consumption to the number of application types deployed in the cluster included in the second candidate solution; Determining a second scheduling score corresponding to the second candidate solution based on the number of replicas of each application deployed in the cluster included in the second candidate solution and a second distance value between the number of replicas of each application deployed in the cluster at a previous moment; A second scoring value corresponding to the second candidate solution is determined according to the priority scoring value, the resource consumption scoring value, and the second scheduling scoring value corresponding to the second candidate solution.
4. The method according to claim 1, wherein Determining each first candidate solution corresponding to the application includes: Get the target number of visits to the application at the current moment, and determine the target number of copies required when the target number of visits corresponds to the target number of visits and the number of copies required when the application is fully running on the jth cluster based on the pre-saved correspondence between the target number of visits to the application and the number of copies required when the application is fully running on the jth cluster. Through the constraints Determine each first candidate solution corresponding to the application in, is the kth first candidate solution for a single application i at time t; It represents the number of replicas that the k-th first candidate solution decides to deploy on the j-th cluster for application i at time t.
5. The method according to claim 1, wherein After determining the first score corresponding to the first candidate solution and before determining the target solution with the highest score, the method further includes: In descending order of the first scoring values, a first number of first candidate solutions with larger first scoring values and corresponding first scoring values are retained, and other first candidate solutions and corresponding first scoring values are deleted.
6. The method according to claim 1, wherein After determining the second score corresponding to the second candidate solution and before determining the target solution with the highest score, the method further includes: In descending order of the second scoring values, a second number of second candidate solutions with larger second scoring values and corresponding second scoring values are retained, and other second candidate solutions and corresponding second scoring values are deleted.
7. The method according to claim 1, wherein Determining the target solution with the highest score according to the first score values corresponding to the first candidate solutions and the second score values corresponding to the second candidate solutions includes: According to the first score value corresponding to each first candidate solution and the second score value corresponding to each second candidate solution, through the constraint condition and x ik +y jm ≤1, determine the target solution with the highest score in, is the first score value corresponding to the kth first candidate solution, is the second score value corresponding to the mth second candidate solution, x ik ∈{0,1} indicates whether the i-th application adopts the k-th first candidate solution, y jm ∈{0,1} indicates whether the jth cluster adopts the mth second candidate solution, and α∈[0,1] is the preset preference coefficient; Constraints Indicates that only one candidate solution is selected for any application or cluster. When the kth and first candidate solutions adopted by application i and the mth solution adopted by cluster j cannot be satisfied at the same time, the constraint x must be satisfied. ik +y jm ≤1.
8. An elastic scaling device based on multi-cluster multi-application bidirectional matching, characterized in that: The device comprises: A first determination module is configured to determine, for each application, each first candidate solution corresponding to the application; and for each first candidate solution, determine a first score corresponding to the first candidate solution based on the number of replicas of the application included in the first candidate solution deployed in each cluster; a second determination module configured to determine, for each cluster, each second candidate solution corresponding to the cluster; and, for each second candidate solution, determine a second scoring value corresponding to the second candidate solution based on the number of replicas of each application deployed in the cluster included in the second candidate solution; The application deployment module is used to determine the target solution with the highest score based on the first score values corresponding to each first candidate solution and the second score values corresponding to each second candidate solution; and deploy application copies of each application based on the number of copies of each application on each cluster in the target solution.
9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; A processor, configured to implement the method according to any one of claims 1 to 7 when executing a program stored in a memory.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.