A method, device and system for joint deployment and scaling of microservices

By fitting the microservice latency characteristic function and generating a tolerance scheme using a heuristic algorithm, the feasibility of microservices on a computing node cluster is evaluated. The optimal number and size of containers are iteratively calculated, solving the problem of excessive resource consumption in existing technologies and achieving efficient microservice deployment and scaling.

CN119621080BActive Publication Date: 2025-11-18BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411763511.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-11-18
Estimated Expiration
2044-12-03

AI Technical Summary

Technical Problem

Existing microservice deployment and scaling methods fail to effectively consider vertical scaling and node resource usage, resulting in insufficient accuracy in deployment decisions and unnecessary resource consumption.

Method used

By fitting the latency characteristic function of microservices and combining it with heuristic algorithms to generate tolerance schemes, the feasibility of microservices on computing node clusters is evaluated, and the optimal number and size of containers are iteratively calculated to minimize resource usage.

Benefits of technology

It improves the accuracy of microservice deployment decisions, reduces resource consumption, optimizes the impact of container configuration and node status on latency, and achieves a scaling solution that minimizes resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119621080B_ABST
    Figure CN119621080B_ABST
Patent Text Reader

Abstract

The application provides a micro-service joint deployment and expansion and contraction method, device and system, the method comprises the following steps: based on the historical data of the target application, fitting the relationship between the delay of each micro-service of the target application and the workload, the container size and the resource usage of the computing power node where the micro-service is located. Using heuristic algorithm, according to the resource usage and workload of the computing power node, the resource usage tolerance of each micro-service is calculated under the delay target limit, and the tolerance scheme is generated. The feasibility of successfully deploying the micro-service on the computing power node according to the tolerance scheme is evaluated. According to the feasible legal tolerance scheme, the number and size of the required containers for each micro-service are calculated to minimize resource usage, and the expansion and contraction scheme is generated. The tolerance and expansion and contraction scheme are iteratively updated using heuristic algorithm until convergence or the maximum number of iterations is reached. The optimal deployment scheme and container expansion and contraction scheme are selected and applied to the computing power node cluster. The application can meet the delay requirement of the application and reduce resource consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of microservice deployment technology, and in particular to a method, apparatus and system for joint deployment and scaling of microservices. Background Technology

[0002] Microservice architecture is a common software architecture in data centers. It decouples applications into independent, deployable microservices, each performing its own function and using isolated resources. This architecture allows system administrators to adjust microservice resources independently based on load, rather than the entire application. Container technology, due to its lightweight and isolation features, is often used in microservice architectures. Each microservice can run through one or more containers, and container resources such as CPU, memory, and network can be configured independently.

[0003] Service Level Objectives (SLOs) define key application attributes such as availability and latency, with latency SLOs often serving as the basis for algorithms managing microservice resources. To handle dynamic application loads, it's necessary to determine container deployment strategies, i.e., on which servers the containers run and how they are scaled up or down.

[0004] Existing microservice deployment and scaling methods typically rely on offline collection of historical data to assess microservice latency characteristics and online allocation of Service Level Losses (SLOs) and scaling decisions. However, these methods have limitations in modeling the relationship between latency and resource consumption. They only consider horizontal scaling (changing the number of containers) and neglect vertical scaling (changing the resource configuration of individual containers). They also ignore the impact of resource usage on microservice latency characteristics across different nodes, thus affecting the accuracy of deployment decisions and resulting in unnecessary resource consumption. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a method, apparatus and system for joint deployment and scaling of microservices, so as to improve the problem of insufficient accuracy of deployment decisions and resulting in more resource consumption in the prior art.

[0006] The first aspect of the present invention provides a method for joint deployment and scaling of microservices, the method comprising the following steps:

[0007] Based on historical data of the target application running in the computing node cluster using a microservice architecture, the relationship between the latency of each microservice of the target application and the workload, container size and resource usage of the computing node where each microservice is located is fitted to obtain the latency characteristic function of each microservice of the target application.

[0008] Based on the resource usage of each computing node in the current computing node cluster and the workload of the target application, a heuristic algorithm is used based on the latency feature function, with the user's latency target for the target application as the limit, to obtain the tolerance of each microservice of the target application to the resource usage of the computing node where it is located, and generate a tolerance scheme.

[0009] Assess the feasibility of deploying each microservice of the target application to the computing node cluster according to the tolerance scheme. If all microservices of the target application are successfully deployed on the computing node cluster, and the available resources of the deployed computing nodes meet the tolerance of each microservice for the resource usage of its respective computing node, then the tolerance scheme is feasible, and the tolerance scheme is recorded as a valid tolerance scheme, along with the corresponding deployment scheme. Otherwise, the tolerance scheme is not feasible, and a new tolerance scheme is obtained and evaluated.

[0010] Based on the latency feature function, the tolerance of each microservice to the resource usage of its computing node in the legal tolerance scheme, and the deployment scheme, the solution is to minimize the resource usage of the target application, calculate the number of containers and container size required to deploy each microservice, and generate a container scaling scheme.

[0011] Based on the container scaling scheme and the deployment scheme, the heuristic algorithm is used to generate a new legal tolerance scheme according to the latency feature function; with the goal of minimizing the resource usage of the target application, the legal tolerance scheme and the container scaling scheme are iteratively updated until convergence or a preset number of iterations is reached;

[0012] The optimal deployment scheme and the optimal container scaling scheme are selected and applied to the computing node cluster.

[0013] In some embodiments of the present invention, evaluating the feasibility of deploying the various microservices of the target application to the computing node cluster according to the tolerance scheme includes the following steps:

[0014] Initialize the number of requests each microservice can handle on the node;

[0015] All microservices are arranged into a first queue in ascending order based on each microservice's tolerance for resource usage on its respective computing node.

[0016] All computing nodes in the computing node cluster are arranged into a second queue in descending order based on the amount of available computing resources.

[0017] The microservice with the lowest tolerance is popped from the first queue in sorted order. A computing node with available resources that meets the microservice's tolerance for resource usage on its host computing node is searched from the second queue. If a computing node that meets the microservice's tolerance exists, the microservice is deployed to that computing node. The available resources of that computing node are updated, and the computing node is reinserted into the second queue in sorted order. The first queue continues to pop microservices in sorted order, searching for and deploying them to computing nodes in the second queue that meet their tolerance, until the first queue is empty. At this point, the tolerance is deemed feasible.

[0018] If no computing power node that meets the microservice tolerance can be found in the second queue, the computing power node with the largest available resources is obtained from the second queue, the microservice is partially deployed to the computing power node, the number of requests processed by the microservice on the computing power node is updated, and the microservice is reinserted into the first queue in sorted order; if the second queue is empty and no computing power node that meets the microservice tolerance can still be found, the tolerance scheme is determined to be infeasible.

[0019] In some embodiments of the present invention, the heuristic algorithm employs particle swarm optimization, ant colony optimization, genetic algorithm, or simulated annealing algorithm.

[0020] In some embodiments of the present invention, based on the latency characteristic function and the tolerance of each microservice to the resource usage of its computing node in the legal tolerance scheme, the goal is to minimize the resource usage of the target application, and the number and size of containers required to deploy each microservice are calculated. The process includes:

[0021] The resource usage of computing nodes in the latency characteristic function of microservices is fixed as the tolerance of each microservice to the resource usage of its computing node; a nonlinear constraint optimization algorithm is used to solve the problem with the goal of minimizing the resource usage of the target application, and the number and size of containers required to deploy each microservice are calculated.

[0022] In some embodiments of the present invention, the method for fitting the delay feature function of the target application includes: a neural network model, a support vector machine, or a decision tree model.

[0023] A second aspect of the present invention provides a microservice federated deployment and scaling system, the system comprising:

[0024] The latency fitting module is used to fit the relationship between the latency of each microservice of the target application and the workload, container size and resource usage of the computing nodes where each microservice is located, based on the historical data of the target application running in the computing node cluster using a microservice architecture, and to obtain the latency characteristic function of each microservice of the target application.

[0025] The tolerance generation module is used to generate a tolerance scheme by using a heuristic algorithm based on the resource usage of each computing node in the current computing node cluster and the workload of the target application, using the latency feature function and the user's latency target for the target application as a limit, to obtain the tolerance of each microservice of the target application for the resource usage of its respective computing node; based on the container scaling scheme and the deployment scheme, using the heuristic algorithm to generate a new valid tolerance scheme based on the latency feature function; iteratively updating the valid tolerance scheme and the container scaling scheme with the goal of minimizing the resource usage of the target application until convergence or reaching a preset number of iterations; and selecting the optimal deployment scheme and the optimal container scaling scheme and applying them to the computing node cluster.

[0026] The virtual deployment module is used to evaluate the feasibility of deploying each microservice of the target application to the computing node cluster according to the tolerance scheme. If all microservices of the target application are successfully deployed on the computing node cluster, and the available resources of the deployed computing nodes meet the tolerance of each microservice for the resource usage of its respective computing node, then the tolerance scheme is feasible, and the deployment scheme corresponding to the tolerance scheme is recorded; otherwise, the tolerance scheme is not feasible, and the next tolerance scheme is obtained and evaluated.

[0027] The optimal SLO allocation module is used to solve the problem based on the latency characteristic function, the tolerance of each microservice to the resource usage of its computing node in the legal tolerance scheme, and the deployment scheme, with the goal of minimizing the resource usage of the target application, to calculate the number of containers and container size required to deploy each microservice, and generate a container scaling scheme.

[0028] In some embodiments of the present invention, the system further includes:

[0029] The monitoring and data acquisition module is used to acquire historical data of the target application running in the computing node cluster using a microservice architecture.

[0030] The scaling deployment module is used to deploy the optimal deployment scheme and the optimal container scaling scheme to the computing node cluster.

[0031] A third aspect of the present invention provides a microservice federated deployment and scaling apparatus, including a processor, a memory, and a computer program stored in the memory, wherein the processor is configured to execute the computer program, and when the computer program is executed, the apparatus performs the steps of the method as described in any of the preceding claims.

[0032] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method as described in any of the preceding claims.

[0033] A fifth aspect of the present invention provides a computer program product comprising a computer program that, when executed by a processor, implements the steps of any of the methods described above.

[0034] The beneficial effects of the present invention are at least as follows:

[0035] This invention provides a method, apparatus, and system for joint deployment and scaling of microservices. The method includes: based on historical data of the target application's operation, fitting the relationship between the latency and workload, container size, and resource usage of the computing nodes where each microservice resides, to obtain the latency characteristic function of each microservice. Using a heuristic algorithm, under latency constraints, calculating the resource usage tolerance of each microservice based on computing node resource usage and workload, and generating a tolerance scheme. Evaluating the feasibility of successfully deploying the microservices on the computing nodes according to the tolerance scheme. Based on feasible and legal tolerance schemes, calculating the required number and size of containers for each microservice with the goal of minimizing resource usage, and generating a scaling scheme. Iteratively updating the tolerance and scaling scheme using a heuristic algorithm until convergence or the maximum number of iterations is reached. Selecting the optimal deployment scheme and container scaling scheme and applying them to the computing node cluster. This invention comprehensively considers the impact of container configuration, node status, and decision side effects on microservice latency, proposing a scaling and deployment decision that satisfies application latency with minimal resource consumption, and providing a fine-grained scheme for deploying microservices and scaling schemes.

[0036] Additional advantages, objects, and features of the invention will be set forth in part in the description which follows, and will also become apparent in part to those skilled in the art upon studying the description, or may be learned by practice of the invention. The objects and other advantages of the invention can be realized and obtained by means of the structures specifically pointed out in the description and drawings.

[0037] Those skilled in the art will understand that the objectives and advantages achievable with the present invention are not limited to those specifically described above, and that the above and other objectives achievable with the present invention will become clearer from the following detailed description. Attached Figure Description

[0038] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, are not intended to limit the scope of the invention. In the drawings:

[0039] Figure 1 This is a flowchart of a method for joint deployment and scaling of microservices in one embodiment of the present invention.

[0040] Figure 2 This is a data flow diagram of the SLO allocation technique in another embodiment of the present invention.

[0041] Figure 3 This is an example diagram of microservice invocation in another embodiment of the present invention.

[0042] Figure 4 This is a diagram illustrating the application of a microservice architecture in another embodiment of the present invention.

[0043] Figure 5 This is a diagram of a microservice joint deployment and scaling system architecture in another embodiment of the present invention. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and descriptions of this invention are used to explain the invention, but are not intended to limit the invention.

[0045] It should also be noted that, in order to avoid obscuring the invention with unnecessary details, only the structures and / or processing steps closely related to the solution according to the invention are shown in the accompanying drawings, while other details that are not closely related to the invention are omitted.

[0046] It should be emphasized that the term "including / comprises" as used herein refers to the presence of a feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, or components.

[0047] It should also be noted that, unless otherwise specified, the term "connection" in this article can refer not only to a direct connection, but also to an indirect connection involving an intermediary.

[0048] In the following description, embodiments of the invention will be illustrated with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar parts, or the same or similar steps.

[0049] Microservices are a software architectural style that designs an application as a series of small, independent services, each running in its own process and communicating with each other through lightweight communication mechanisms. Each microservice can be deployed, scaled, and maintained independently.

[0050] Containers are a virtualization technology used to isolate applications and their dependent runtime environments. Containers can package an application and its dependencies into a portable, self-contained unit, including application code, runtime environment, system tools, system libraries, etc. By packaging microservices into containers, rapid deployment, scaling, and migration of services can be achieved. Container technology allows microservices to more flexibly choose technology stacks and deployment environments, improving system flexibility and scalability.

[0051] One embodiment of the present invention provides a method, apparatus, and system for microservice deployment and scaling, wherein the method for joint deployment and scaling of microservices is as follows: Figure 1 As shown, the process includes the following steps S101 to S106:

[0052] Step S101: Based on the historical data of the target application running in the computing node cluster using a microservice architecture, fit the relationship between the latency of each microservice of the target application and the workload, container size and the resource usage of the computing node where each microservice is located, and obtain the latency characteristic function of each microservice of the target application.

[0053] Step S102: Based on the resource usage of each computing node in the current computing node cluster and the workload of the target application, a heuristic algorithm is used based on the latency feature function, with the user's latency target for the target application as the limit, to obtain the tolerance of each microservice of the target application for the resource usage of the computing node where it is located, and generate a tolerance scheme.

[0054] Step S103: Evaluate the feasibility of deploying each microservice of the target application to the computing node cluster according to the tolerance scheme. If all microservices of the target application are successfully deployed on the computing node cluster, and the available resources of the deployed computing nodes meet the tolerance of each microservice for the resource usage of its respective computing node, then the tolerance scheme is feasible, and the tolerance scheme is recorded as a valid tolerance scheme, along with the corresponding deployment scheme; otherwise, the tolerance scheme is not feasible, and a new tolerance scheme is obtained and evaluated.

[0055] Step S104: Based on the latency characteristic function, the tolerance of each microservice to the resource usage of its computing node and the deployment scheme in the legal tolerance scheme, the solution is obtained with the goal of minimizing the resource usage of the target application. The number of containers and the container size required to deploy each microservice are calculated, and a container scaling scheme is generated.

[0056] Step S105: Based on the container scaling and deployment schemes, use a heuristic algorithm to generate new legal tolerance schemes according to the latency characteristic function. With the goal of minimizing the resource usage of the target application, iteratively update the legal tolerance schemes and container scaling schemes until convergence or the preset number of iterations is reached.

[0057] Step S106: Select the optimal deployment scheme and the optimal container scaling scheme and apply them to the computing node cluster.

[0058] Among them, the user's latency target for the target application is the latency SLO (Service Level Objective) of the target application.

[0059] The workload, or the number of requests, is represented by RPS (Requests Per Second).

[0060] Among them, the fitting delay feature function can be used to fit the relationship between the delay and each variable using regression analysis or machine learning algorithms.

[0061] In some embodiments, the feasibility of deploying the various microservices of the target application to a cluster of computing nodes according to a tolerance scheme is evaluated, and the specific steps include:

[0062] Initialize the number of requests that each microservice can handle on the node.

[0063] All microservices are arranged into a first queue in ascending order based on each microservice's tolerance for resource usage on its respective computing node.

[0064] All computing nodes in the computing node cluster are arranged into a second queue in descending order based on the amount of available computing resources.

[0065] The microservice with the lowest tolerance is popped from the first queue in sorted order. Then, a computing node with available resources that meets the microservice's tolerance for resource usage on its host computing node is searched from the second queue. If a computing node that meets the microservice's tolerance exists, the microservice is deployed to that node. The available resources of that computing node are updated, and the node is reinserted into the second queue in sorted order. The first queue continues to pop microservices in sorted order, searching for and deploying them to computing nodes in the second queue that meet their tolerance, until the first queue is empty. At this point, the tolerance is considered feasible.

[0066] If no computing power node that meets the microservice tolerance can be found in the second queue, the computing power node with the largest available resources is obtained from the second queue, the microservice is partially deployed to the computing power node, the number of requests processed by the microservice on the computing power node is updated, and the microservice is reinserted into the first queue in sorted order; if the second queue is empty and no computing power node that meets the microservice tolerance can still be found, the tolerance scheme is determined to be infeasible.

[0067] In some embodiments, the heuristic algorithm employs particle swarm optimization, ant colony optimization, genetic algorithm, or simulated annealing.

[0068] In some embodiments, based on the latency characteristic function and the tolerance of each microservice to the resource usage of its host computing node in the legal tolerance scheme, the goal is to minimize the resource usage of the target application. This process calculates the number of containers and the container size required to deploy each microservice.

[0069] The resource usage of computing nodes in the latency characteristic function of microservices is fixed as the tolerance of each microservice to the resource usage of its respective computing node. A nonlinear constraint optimization algorithm is used to solve the problem by minimizing the resource usage of the target application, and the number and size of containers required to deploy each microservice are calculated.

[0070] In some embodiments, methods for fitting the delay feature function of the target application include: neural network models, support vector machines, or decision tree models.

[0071] The microservice joint deployment and scaling system proposed in this embodiment includes:

[0072] The latency fitting module is used to fit the relationship between the latency of each microservice of the target application and its workload, container size, and the resource usage of the computing nodes where each microservice is located, based on historical data of the target application running in the computing node cluster using a microservice architecture, and to obtain the latency characteristic function of each microservice of the target application.

[0073] The tolerance generation module is used to generate tolerance schemes based on the resource usage of each computing node in the current computing node cluster and the workload of the target application. It employs a heuristic algorithm based on a latency feature function, with the user's latency target for the target application as a limit, to obtain the tolerance of each microservice of the target application for the resource usage of its respective computing node. Based on the container scaling and deployment schemes, it uses a heuristic algorithm to generate new valid tolerance schemes based on the latency feature function. With the goal of minimizing the resource usage of the target application, it iteratively updates the valid tolerance schemes and container scaling schemes until convergence or a preset number of iterations is reached. Finally, it selects the optimal deployment scheme and the optimal container scaling scheme and applies them to the computing node cluster.

[0074] The virtual deployment module is used to evaluate the feasibility of deploying each microservice of the target application to the computing node cluster according to the tolerance scheme. If all microservices of the target application are successfully deployed on the computing node cluster, and the available resources of the deployed computing nodes meet the tolerance of each microservice for the use of resources of its computing node, then the tolerance scheme is feasible, and the deployment scheme corresponding to the tolerance scheme is recorded; otherwise, the tolerance scheme is not feasible, and the next tolerance scheme is obtained and evaluated.

[0075] The optimal SLO allocation module is used to solve the problem based on the latency characteristic function, the tolerance of each microservice to the resource usage of its computing node and the deployment scheme in the legal tolerance scheme, with the goal of minimizing the resource usage of the target application. It calculates the number of containers and the container size required to deploy each microservice and generates a container scaling scheme.

[0076] In some embodiments, the system further includes:

[0077] The monitoring and data collection module is used to acquire historical data of the target application running in the computing node cluster using a microservice architecture.

[0078] The scaling deployment module is used to deploy the optimal deployment scheme and the optimal container scaling scheme to the computing node cluster.

[0079] The monitoring and data collection module can collect historical data by collecting log records of the target application and calling the API interface of the target application.

[0080] The scaling and scaling deployment module can achieve automated deployment and scaling of containers through container orchestration tools (such as Kubernetes).

[0081] The microservice joint deployment and scaling device proposed in this embodiment includes a processor, a memory, and a computer program stored in the memory. The processor is used to execute the computer program, and when the computer program is executed, the device implements the steps of any of the methods described above.

[0082] Accordingly, this embodiment provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described above.

[0083] Accordingly, this embodiment provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above methods.

[0084] Another embodiment of the present invention provides a method for joint deployment and scaling of microservices. It proposes to meet the application SLO (Service Level Order) scaling and deployment decisions with minimal resource consumption by comprehensively considering the impact of container configuration, node status, and decision side effects on microservice latency, thereby overcoming the shortcomings of the prior art.

[0085] Existing SLO allocation techniques distribute the end-to-end SLO of a target application to multiple microservices that make up the application. Each microservice allocates resources according to its assigned sub-SLO. This technique typically requires offline collection of latency characteristics of each microservice based on historical data of the target application running in the cluster, expressed as a function of latency and other factors directly affecting latency, such as workload. Then, based on these latency characteristics and the current state of the cluster and the target application, the SLO allocation result is obtained online, along with the final scaling and deployment decisions. The entire process is as follows: Figure 2 As shown.

[0086] Specifically, during the online operation of existing SLO allocation algorithms, it is necessary to obtain the final application latency based on any SLO allocation scheme. This process requires analyzing microservice call relationships. For each microservice that makes up the application, its latency will be as close as possible to its allocated sub-SLO to reduce resource consumption. Therefore, the latency of each microservice after implementing a given SLO allocation scheme is considered as its allocated sub-SLO within that scheme. Then, the application latency is obtained based on the microservice latency. This process is typically accomplished by analyzing the application's microservice call graph. The microservice call graph usually appears in the form of a directed acyclic graph (DAG), and different methods may adjust its structure as needed. The structure used in this method is based on... Figure 3 For example, each node in the graph represents a microservice, and the edges represent the call relationships between microservices. As shown, microservice 1 only calls microservice 2, and microservice 2 calls the other 7 microservices. Microservice calls can be parallel or sequential. Parallel microservice calls use arrows of the same color. As shown, microservices 3 to 6 are called in parallel by microservice 2, and after execution, microservice 2 then sequentially calls microservices 9 to 11. Based on the call graph, the latency of the sequentially executed microservices is added together, and the latency of the parallelly executed microservices is taken as the maximum value to obtain the latency of the entire application. This process can be used to obtain the application latency based on a given SLO (Solution Time Limit) allocation scheme.

[0087] by Figure 4 Taking a microservice architecture as an example, this paper describes the methods and systems for joint deployment and scaling of microservices. In this embodiment, Figure 4 The application microservice architecture shown is deployed in a cluster with five identical nodes. The application has been running in the cluster for some time, and there is sufficient data to fit its latency characteristics. Because other applications are running in the cluster, each node has a different initial resource footprint.

[0088] The system's structural diagram is as follows: Figure 5 As shown, the steps of this method include:

[0089] 1. The Latency Profiling module collects historical runtime data of the target application's containers offline from the node cluster and fits a function of latency to load, container configuration, and the current resource usage of the node where the container resides.

[0090] 2. The Tolerance Generation module obtains the latency characteristic functions of each microservice in the target application from the Latency Profiling module;

[0091] 3. The Tolerance Generation module collects real-time data on the resource usage of each node in the cluster and the load of the target application. Then, based on the PSO (Particle Swarm Optimization) algorithm (or other heuristic algorithms, such as ant colony optimization, genetic algorithm, simulated annealing, etc.), it obtains the tolerance of each microservice of the target application for the maximum resource usage of its own node, that is, how much resource is occupied on this node at most (the tolerance of each microservice for the resources occupied by its own node).

[0092] 4. The Virtual Deployment module obtains a tolerance scheme from the Tolerance Generation module, and then attempts to deploy each microservice on the cluster according to this scheme to observe its feasibility. For example... Figure 5 The Virtual Deployment module demonstrates two tolerance schemes. In the first scheme, microservices 1, 2, 3, and 4 have high tolerance, while microservice 5 has low tolerance. Therefore, the first four microservices are run together on one node, while microservice 5 is run on another node. Since all nodes can deploy and meet their tolerances, this tolerance scheme is valid. In the second scheme, microservice 1 has extremely low tolerance, while the other microservices have high tolerance. Therefore, microservice 1 is deployed on one node, leaving it ample space. However, after deploying microservices 2, 3, and 4, it is found that there is not enough space to deploy microservice 5, therefore this tolerance scheme is invalid.

[0093] The specific process of deploying microservices using the Virtual Deployment module is as follows:

[0094] Step 1: For all microservices m∈[1,M], nodes s∈[1,N], where M is the number of microservices and N is the number of nodes, let λ m,s =0, Λ m =Λ, where λ m,s Let be the number of requests processed by microservice m on node s, and Λ be the current RPS of the application. The M microservices are then divided according to rr...m Sort the items in ascending order and place them in priority queue Q1, where rr m This represents the minimum remaining resources that microservice m can accept from the nodes where its instances reside. N nodes are then allocated according to Rem. s Sort the contents in descending order and place them in priority queue Q2, where Rem s This represents the currently available resources for this node. The actual resource usage (ru) for each microservice is obtained based on Λ. m =h m (Λ), h m (Λ) represents the actual resources consumed by microservice m when the application's RPS is Λ. Proceed to step 2;

[0095] Step 2: Pop rr from Q1 m The smallest microservice m, which satisfies Rem from Q2. s -ru m ≥rr m Rem in the node s Pop the smallest node s. Deploy the microservice m onto node s, and let λ m,s =Λ m Λ m =0. Let Rem s =Rem s -ru m ,ru m =0, insert node s again into Q2 in descending order, and proceed to step 3. If no such node exists, proceed to step 4;

[0096] Step 3: If Q1 is empty, the solution is valid and Virtual Deployment exits; otherwise, return to step 2.

[0097] Step 4: Retrieve Rem from Q2 s The largest node s will deploy the microservice m part to node s, and let λ m,s =(Rem s -rr m ) / (Rem s -ru m )*Λ m Λ m =Λ m -λ m,s . Let ru m =Rem s -rr m Rem s =0. Insert microservice m back into Q1 in ascending order, and return to step 3. If no such node exists, the scheme is invalid, and Virtual Deployment exits.

[0098] 5. If the tolerance scheme is valid, the deployment scheme is recorded and the application enters the Optimal SLO Assignment module; otherwise, it returns to the Tolerance Generation module to obtain the next scheme. After entering the tolerance scheme in the Optimal SLO Assignment module, the variable "resource usage of the node" in the application's latency characteristics is fixed as its tolerance. Then, the application solves the problem by using the SLSQP (Sequential Least Squares Quadratic Programming) algorithm or other algorithms for solving constrained nonlinear optimization problems, with the objective of minimizing the overall resource usage of the application, to obtain the number and size of each container.

[0099] By properly configuring the number and size of microservice containers, the system's latency can meet user needs while minimizing resource consumption and operational costs. The optimization process focuses on two aspects: first, finding the optimal container configuration to minimize resource usage; and second, ensuring that this configuration meets user latency requirements.

[0100] The process of solving the microservice latency optimization problem involves adjusting the number and size of containers to minimize resource costs. The specific process is as follows:

[0101] The previously fitted delay function was a function of load, container configuration, and the current resource usage of the node where the container resides, denoted as f. m,s (λ m ,θ m Rem s ′). Here λ m θ represents the load of each container in the microservice m. m This refers to the container size of the microservice m. Here, Rem... s ′ represents the resources of node s after adopting the resource allocation and deployment schemes, while Rem s These are resources on a node when the application is not deployed, therefore the two are not the same. Rem s ′ fixed as rr m , so f m,s That has nothing to do with s anymore, let's make it... r m It is the number of containers for microservice m, which can be used to... Turn it into a function that determines the number and size of microservice containers.

[0102] Then, the solution is obtained based on the microservice architecture in the embodiment:

[0103]

[0104] stg1(r1,θ1)+max(g2(r2,θ2)+g4(r4,θ4),g3(r3,θ3)+g5(r5,θ5))≤SLO;

[0105] SLO is the user-defined latency target.

[0106] 6. The deployment plan, horizontal and vertical scaling results, and final resource consumption are returned from the Optimal SLOAssignment module to the Tolerance Generation module. This module updates the state and generates a new tolerance plan based on the PSO algorithm with the aim of minimizing resource consumption. This guides convergence or reaching the maximum number of iterations.

[0107] 7. The Tolerance Generation module selects the globally optimal deployment and scaling scheme and applies it to the cluster.

[0108] In summary, this invention provides a method, apparatus, and system for joint deployment and scaling of microservices. The method includes: based on historical data of a target application, fitting the relationship between latency and workload, container size, and resource usage of the computing nodes where each microservice resides, to obtain latency characteristic functions for each microservice. Using a heuristic algorithm, under latency constraints, calculating the resource usage tolerance of each microservice based on computing node resource usage and workload, and generating a tolerance scheme. Evaluating the feasibility of successfully deploying microservices on computing nodes according to the tolerance scheme. Based on feasible and legal tolerance schemes, calculating the required number and size of containers for each microservice with the goal of minimizing resource usage, and generating a scaling scheme. Iteratively updating the tolerance and scaling scheme using a heuristic algorithm until convergence or the maximum number of iterations is reached. Selecting the optimal deployment scheme and container scaling scheme and applying them to the computing node cluster. This invention comprehensively considers the impact of container configuration, computing node status, and decisions on the resource usage status of computing nodes, iteratively generating the optimal container scaling scheme and microservice deployment decisions with the goal of minimizing application resource consumption.

[0109] Corresponding to the above method, the present invention also provides an apparatus comprising a computer device including a processor and a memory, the memory storing computer instructions, the processor executing the computer instructions stored in the memory, and the apparatus performing the steps of the method as described above when the computer instructions are executed by the processor.

[0110] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the aforementioned edge computing server deployment method. The computer-readable storage medium can be a tangible storage medium, such as random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, floppy disks, hard disks, removable storage disks, CD-ROMs, or any other form of storage medium known in the art.

[0111] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this invention are programs or code segments used to perform the desired tasks. The programs or code segments can be stored in a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried in a carrier wave.

[0112] It should be clarified that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of the present invention.

[0113] In this invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.

[0114] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations of the embodiments of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for joint deployment and scaling of microservices, the method comprising: receiving a request for a microservice; determining a target server for the microservice; and deploying the microservice to the target server. The method comprises the following steps: According to the historical data of the target application running in the computing node cluster adopting the micro-service architecture, the relationship between the delay of each micro-service of the target application and the workload, the container size and the resource usage of the computing node where each micro-service is located is fitted to obtain the delay characteristic function of each micro-service of the target application; According to the resource usage of each computing node in the computing node cluster and the workload of the target application, a heuristic algorithm is used to obtain the resource usage tolerance of each micro-service of the target application to the computing node where it is located based on the delay characteristic function and within the delay target of the target application by the user, and a tolerance scheme is generated; The feasibility of deploying each micro-service of the target application to the computing node cluster according to the tolerance scheme is evaluated; if all micro-services of the target application are successfully deployed on the computing node cluster and the available resources of the deployed computing node meet the resource usage tolerance of each micro-service to the computing node where it is located, the tolerance scheme is feasible, the tolerance scheme is recorded as a legal tolerance scheme, and the deployment scheme corresponding to the tolerance scheme is recorded; otherwise, the tolerance scheme is not feasible, the next tolerance scheme is obtained and evaluated again; Based on the delay characteristic function, the resource usage tolerance of each micro-service to the computing node where it is located in the legal tolerance scheme and the deployment scheme, the container quantity and the container size required for deploying each micro-service are calculated by solving the target of minimizing the resource usage of the target application, and a container expansion and contraction scheme is generated; According to the container expansion and contraction scheme and the deployment scheme, a new legal tolerance scheme is generated according to the delay characteristic function using the heuristic algorithm; the legal tolerance scheme and the container expansion and contraction scheme are iteratively updated until convergence or a preset iteration number is reached, with the target of minimizing the resource usage of the target application; The optimal deployment scheme and the optimal container expansion and contraction scheme are selected and applied to the computing node cluster. 2.The method of claim 1, wherein, The feasibility of deploying each micro-service of the target application to the computing node cluster according to the tolerance scheme is evaluated, and the specific steps comprise: Initialize the number of requests processed by each micro-service on the node; All micro-services are arranged in ascending order according to the resource usage tolerance of each micro-service to the computing node where it is located to form a first queue; All computing nodes in the computing node cluster are arranged in descending order according to the available resource quantity of the computing node to form a second queue; The micro-service with the smallest tolerance is popped out from the first queue according to the sorting, and a computing node with available resource quantity meeting the resource usage tolerance of the micro-service to the computing node where it is located is searched from the second queue; if there is a computing node meeting the tolerance of the micro-service, the micro-service is deployed to the computing node; the available resource quantity of the computing node is updated, and the computing node is inserted into the second queue again according to the sorting; the micro-service is continuously popped out from the first queue according to the sorting, and is deployed to the computing node in the second queue meeting its tolerance until the first queue is empty, and the tolerance is determined to be feasible; If no computing power node that meets the microservice tolerance can be found in the second queue, the computing power node with the largest available resources is obtained from the second queue, the microservice is partially deployed to the computing power node, the number of requests processed by the microservice on the computing power node is updated, and the microservice is reinserted into the first queue in sorted order; if the second queue is empty and no computing power node that meets the microservice tolerance can still be found, the tolerance scheme is determined to be infeasible. 3.The method of claim 1, wherein, The heuristic algorithm used is particle swarm optimization, ant colony optimization, genetic algorithm, or simulated annealing algorithm.

4. The method of claim 1, wherein, Based on the latency characteristic function and the tolerance of each microservice to the resource usage of its computing node in the legal tolerance scheme, the solution aims to minimize the resource usage of the target application, and calculates the number and size of containers required to deploy each microservice. The process includes: The resource usage of computing nodes in the latency characteristic function of microservices is fixed as the tolerance of each microservice to the resource usage of its computing node; a nonlinear constraint optimization algorithm is used to solve the problem with the goal of minimizing the resource usage of the target application, and the number and size of containers required to deploy each microservice are calculated.

5. The method of claim 1, wherein, Methods for fitting the delay feature function of the target application include: neural network models, support vector machines, or decision tree models.

6. A system for joint deployment and scaling of microservices, characterized in that, The system includes: The latency fitting module is used to fit the relationship between the latency of each microservice of the target application and the workload, container size and resource usage of the computing nodes where each microservice is located, based on the historical data of the target application running in the computing node cluster using a microservice architecture, and to obtain the latency characteristic function of each microservice of the target application. The tolerance generation module is used to generate a tolerance scheme by using a heuristic algorithm based on the resource usage of each computing node in the current computing node cluster and the workload of the target application, with the latency feature function as the limit and the user's latency target for the target application. The module generates a tolerance scheme based on this tolerance scheme. It then uses the heuristic algorithm to generate a new valid tolerance scheme based on the latency feature function, according to the container scaling and deployment schemes. Finally, it iteratively updates the valid tolerance scheme and the container scaling scheme with the goal of minimizing the resource usage of the target application until convergence or a preset number of iterations is reached. Finally, it selects the optimal deployment scheme and the optimal container scaling scheme and applies them to the computing node cluster. The virtual deployment module is used to evaluate the feasibility of deploying each microservice of the target application to the computing node cluster according to the tolerance scheme. If all microservices of the target application are successfully deployed on the computing node cluster, and the available resources of the deployed computing nodes meet the tolerance of each microservice for the resource usage of its respective computing node, then the tolerance scheme is feasible, and the deployment scheme corresponding to the tolerance scheme is recorded; otherwise, the tolerance scheme is not feasible, and the next tolerance scheme is obtained and evaluated. The optimal SLO allocation module is used to solve the problem based on the latency characteristic function, the tolerance of each microservice to the resource usage of its computing node in the legal tolerance scheme, and the deployment scheme, with the goal of minimizing the resource usage of the target application, to calculate the number of containers and container size required to deploy each microservice, and generate a container scaling scheme.

7. The system for joint deployment and scaling of microservices according to claim 6, wherein, The system also includes: The monitoring and data acquisition module is used to acquire historical data of the target application running in the computing node cluster using a microservice architecture. The scaling deployment module is used to deploy the optimal deployment scheme and the optimal container scaling scheme to the computing node cluster.

8. An apparatus for joint deployment and scaling of microservices, comprising a processor, a memory, and a computer program stored in the memory, characterized in that, The processor is used to execute the computer program, and when the computer program is executed, the device implements the steps of the method as described in any one of claims 1 to 5.

9. A computer-readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 5.

10. A computer program product comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Online application dynamic capacity expansion and shrinkage method based on micro-service call dependence perception

    CN112199150A

  • Micro-service dynamic scaling and migration method and device

    CN112988398A