A micro-service fault-tolerant scheduling method, system, device and storage medium

CN122593966BActive Publication Date: 2026-09-15NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611081652.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-21
Publication Date
2026-09-15
Estimated Expiration
2046-07-21

AI Technical Summary

Technical Problem

上述特性可能导致现有方法在缩短响应时间和降低资源消耗方面效果欠佳

Benefits of technology

1、设计了贝叶斯概率树(BPT)来建模调用图,从而更准确地反映请求的实际响应过程。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122593966B_ABST
    Figure CN122593966B_ABST
Patent Text Reader

Abstract

The application belongs to the field of micro-service scheduling, and particularly relates to a micro-service fault-tolerant scheduling method, system, device and storage medium based on Bayesian inference, which comprises the following steps: generating an application-level call graph according to historical micro-service call data, modeling the application-level call graph as a Bayesian probability tree; setting a resource pool; performing static micro-service instance deployment in the static resource pool according to the historical micro-service call data and the cold start time; pruning the Bayesian probability tree according to a micro-service request to obtain an actual call graph; calculating the minimum number of copies required for task execution; selecting the earliest task execution mode from the three modes of the deployed static micro-service instances in the static resource pool, the created dynamic instances and the newly created dynamic instances in the elastic resource pool as a target mode according to the minimum number of copies; and realizing micro-service scheduling of the micro-service request according to the target mode. The application can effectively reduce resource consumption and ensure system reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of microservice scheduling, specifically relating to a microservice fault-tolerant scheduling method, system, device, and storage medium. Background Technology

[0002] Microservice architecture is widely used in cloud applications due to its elasticity, scalability, and loose coupling. This architecture breaks down a monolithic application into multiple microservices, which improves development and maintenance efficiency and enhances system flexibility. For microservice-based online applications, when a user request is received, multiple microservices will work together to complete the response.

[0003] However, microservice-based online applications can experience failures, potentially triggering end-to-end request failures. This is especially true as the call graph grows, causing the end-to-end failure rate to spike dramatically, making it difficult to meet users' high reliability requirements. With the advancement of fault-tolerance technology research, master-slave replication technology has emerged. Its core concept is to improve system response reliability by increasing the number of replicas. For example, in active replication mode, a successful response is considered complete as long as at least one replica successfully completes the operation. However, too many replicas increase the system load, leading to longer response times and increased resource consumption. Therefore, how to minimize response time and optimize resource utilization while ensuring reliability has become a pressing problem to be solved.

[0004] Although fault-tolerant scheduling has been extensively studied, traditional methods such as RR, GSMS, RA-MOMA, and R_RIR typically model microservice workflows as directed acyclic graphs (DAGs). In this context, workflows represent the dependencies and execution order between tasks. However, microservice architectures contain complex call logic, the so-called call graph, which captures the intricate interactions between microservices. Microservice-based online application call graphs have three key characteristics that significantly differ from traditional DAGs and workflows: Incomplete information: The call timing and logic of microservices are determined internally by the application, and the scheduling system typically cannot obtain this information when a user request arrives, resulting in incomplete call graph information. However, many studies have emphasized ensuring the completeness of workflow information.

[0005] Diverse interactions: The call graph includes not only one-way calls (such as message queues that do not require a return result) but also bidirectional calls that require responses from downstream microservices. Furthermore, analysis of microservice tracing data shows that bidirectional calls account for 97.84% of all calls. This indicates a significant difference in interaction patterns between the call graph and the directed acyclic graph (DAG).

[0006] Dynamic structure: Calls between microservices are mainly determined by user requests, exhibiting significant dynamism. However, many studies treat workflows as static. Furthermore, analysis of microservice trace data shows that 73% of user requests generate call graphs smaller than the maximum call graph size, indicating that the system is highly dynamic.

[0007] Therefore, using a directed acyclic graph (DAG) to model the microservice call graph is not feasible. These characteristics may cause existing methods to be less effective in reducing response time and resource consumption. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to provide a microservice fault-tolerant scheduling method, system, device and storage medium that can effectively reduce resource consumption and ensure system reliability.

[0009] A microservice fault-tolerant scheduling method includes: Obtain historical microservice call data for the same application, and obtain a call graph for a single request based on the historical microservice call data; Set up a virtual super root that connects to the entry point of the call graph for all requests; Based on the virtual super root and historical microservice call data, an application-level call graph is generated, and the application-level call graph is modeled as a Bayesian probability tree. Set up a resource pool that includes a static resource pool and an elastic resource pool. The resource pool consists of a group of bare-metal nodes, each of which has a corresponding resource capacity. The static resource pool is used to deploy static microservice instances, and the elastic resource pool is used to create dynamic microservice instances on demand. Based on the historical microservice call data and cold start time, deploy static microservice instances in the static resource pool; Obtain the actual microservice request; The Bayesian probability tree is pruned based on the microservice request to obtain the actual call graph; Calculate the minimum number of replicas required to execute a task based on the reliability of each task in the actual call graph; Based on the minimum number of replicas for each task, the method that can execute the task earliest is selected from three options: static microservice instances already deployed in the static resource pool, dynamic microservice instances already created in the elastic resource pool, and newly created dynamic instances in the elastic resource pool. Among them, the deployed static microservice instances can be called directly, the dynamic microservice instances already created in the elastic resource pool have a buffer time to execute the task, and the newly created dynamic instances have a creation time. Based on the stated objective, microservice scheduling of microservice requests is implemented.

[0010] Optionally, the step of generating an application-level call graph based on the virtual superroot and historical microservice call data, and modeling the application-level call graph as a Bayesian probability tree, includes: Based on the virtual super root and historical microservice call data, vertex attributes and edge attributes are obtained. The vertex attributes include microservice type, microservice location, processing time, and partition attribute. The edge attributes include upstream or downstream vertex, edge type, probability of edge existence in a specific request, preparation time before call, and data transmission volume. Based on the vertex and edge attributes, an application-level call graph is generated, and the application-level call graph is modeled as a Bayesian probability tree.

[0011] Optionally, the step of deploying static microservice instances in the static resource pool based on the historical microservice call data and cold start time includes: Generate quantity-time curves for different microservice types based on the historical microservice call data; Initialize the number of static microservice instances, y. Based on the quantity-time curve, the number of intervals from when the number of static microservice instances drops below y to when it reaches or exceeds y, and the total duration during which the number of static instances is less than y at time t are obtained. The average call interval time is obtained based on the number of intervals and the total duration. If the average call interval is less than the interval parameter, the number of static instances is increased until the average call interval is greater than or equal to the interval parameter, which is obtained through cold start time and threshold parameter. If the average call interval is greater than or equal to the interval parameter, then the maximum number of static microservice instances deployed for different microservice types is obtained. Static microservice instances are deployed in the static resource pool based on the maximum number of static microservice instances deployed.

[0012] Optionally, calculating the minimum number of replicas required to execute a task based on the reliability of each task in the actual call graph includes: Based on the actual call graph, global parameters are obtained, including the number of tasks to be completed, the number of tasks that have been successfully scheduled, the overall request reliability required by the user, and the cumulative reliability of the tasks that have been successfully scheduled. Based on the actual microservice requests, a task list is obtained, and each task in the task list is sorted in ascending order according to its average processing time to obtain the sorting result. Select a task as the current task according to the sorting results, initialize the current number of replicas for the current task, and calculate the task success rate under the current number of replicas. Calculate the required reliability for the current task; If the task success rate is less than the required reliability, increase the current number of replicas until the task success rate is greater than or equal to the required reliability. If the task success rate is greater than or equal to the required reliability, then the number of replicas when the task success rate is greater than or equal to the required reliability is obtained as the minimum number of replicas for the current task. The sorted results are then iterated to obtain the minimum number of replicas required for each task to execute.

[0013] Optionally, the selection of the earliest task execution method from three options—static microservice instances already deployed in the static resource pool, dynamic microservice instances already created in the elastic resource pool, and newly created dynamic instances in the elastic resource pool—based on the minimum number of replicas for each task includes: Based on the minimum number of replicas for each task, search for all instances of the same type from the static microservice instances deployed in the static resource pool and the dynamic instances created in the elastic resource pool as a candidate instance list; Calculate the estimated start time for assigning tasks to candidate instances from the candidate instance list; Get the new estimated start time for creating a new dynamic instance in the Elastic Resource Pool; Compare the estimated start time of all candidate instances in the candidate instance list with the new estimated start time for creating a new instance, and find the minimum value as the minimum start time; Obtain the instance type corresponding to the minimum start time. The instance type includes static microservice instance deployment, created dynamic microservice instance, and newly created dynamic instance. Select the instance type corresponding to the minimum start time as the target method.

[0014] Optionally, the step of implementing microservice scheduling of microservice requests according to the target method includes: If the target is a deployed static microservice instance or a dynamically created instance in an elastic resource pool, then the replica will be scheduled to the deployed static microservice instance or the dynamically created instance in an elastic resource pool. If the target method is to create a new dynamic instance in the elastic resource pool, then a new microservice instance is created, and the replica is scheduled to the newly created microservice instance.

[0015] A microservice fault-tolerant scheduling system includes: The first acquisition module is used to acquire historical microservice call data of the same application and obtain the call graph of a single request based on the historical microservice call data. The first configuration module is used to configure the virtual super root that connects to the entry point of the call graph for all requests; The Bayesian probability tree creation module is used to generate an application-level call graph based on the virtual super root and historical microservice call data, and to model the application-level call graph as a Bayesian probability tree. The second setting module is used to set up a resource pool that includes a static resource pool and an elastic resource pool. The resource pool consists of a group of bare-metal nodes, each of which has a corresponding resource capacity. The static resource pool is used to deploy static microservice instances, and the elastic resource pool is used to create dynamic microservice instances on demand. The static deployment module is used to deploy static microservice instances in the static resource pool based on the historical microservice call data and cold start time. The second acquisition module is used to acquire the actual microservice requests; The pruning module is used to prune the Bayesian probability tree according to the microservice request to obtain the actual call graph; The minimum number of replicas module is used to calculate the minimum number of replicas required to execute a task based on the reliability of each task in the actual call graph. The selection module is used to select the earliest method to execute the task from three options: static microservice instances already deployed in the static resource pool, dynamic microservice instances already created in the elastic resource pool, and newly created dynamic instances in the elastic resource pool, based on the minimum number of replicas for each task. Among them, the deployed static microservice instances can be called directly, the dynamic microservice instances already created in the elastic resource pool have a buffer time to execute the task, and the newly created dynamic instances have a creation time. The scheduling module is used to schedule microservice requests according to the target method.

[0016] Optionally, the Bayesian probability tree creation module includes: The data acquisition unit obtains vertex attributes and edge attributes based on the virtual super root and historical microservice call data. The vertex attributes include microservice type, microservice location, processing time, and partition attribute. The edge attributes include upstream or downstream vertex, edge type, probability of edge existence in a specific request, pre-call preparation time, and data transmission volume. Create a unit to generate an application-level call graph based on the vertex and edge attributes, and model the application-level call graph as a Bayesian probability tree.

[0017] A terminal device includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor loads and executes the computer program, it employs a microservice fault-tolerant scheduling method.

[0018] A computer-readable storage medium storing a computer program, wherein the computer program employs a microservice fault-tolerant scheduling method when loaded and executed by a processor.

[0019] The beneficial effects of this invention are: 1. A Bayesian probability tree (BPT) was designed to model the call graph, thereby more accurately reflecting the actual response process of the request.

[0020] 2. Set up a resource pool containing a static resource pool and an elastic resource pool. The resource pool consists of a group of bare-metal nodes, each with corresponding resource capacity. The static resource pool is used to deploy static microservice instances, and the elastic resource pool is used to create dynamic microservice instances on demand. Based on the historical microservice call data and cold start time, deploy static microservice instances in the static resource pool; obtain the actual microservice requests; prune the Bayesian probability tree based on the microservice requests to obtain the actual call graph. The minimum number of replicas required to execute a task is calculated based on the reliability of each task in the actual call graph. Based on the minimum number of replicas for each task, the method that executes the task earliest is selected from three options: static microservice instances already deployed in the static resource pool, dynamic microservice instances already created in the elastic resource pool, and newly created dynamic instances in the elastic resource pool. A dual-pool scheduling framework is designed, in which static microservices are deployed using a call graph derived from historical data, while dynamic microservices utilize a real-time call graph, thereby effectively shortening the response time.

[0021] 3. The DRIFT algorithm is proposed. This algorithm is optimized for the dynamic and incomplete information characteristics of the call graph, effectively reducing resource consumption while meeting reliability requirements. Attached Figure Description

[0022] Figure 1 This is a flowchart illustrating a microservice fault-tolerant scheduling method according to the present invention. Figure 2 This is a diagram of the microservice-based online application system architecture of the present invention. Figure 3 This is the call diagram after modeling the present invention; Figure 4 This is a schematic diagram of the unidirectional and bidirectional edges of the present invention; Figure 5 This is a schematic diagram of the BPT model of the present invention. Figure 5 (a) and Figure 5 (b) is a diagram illustrating the merging of vertices to match edge type and vertex type. Figure 5 (c) and Figure 5 (d) A diagram illustrating the creation of a new branch when the edge type and vertex type do not match; Figure 6 This is a schematic diagram of the dual-pool scheduling framework of the present invention; Figure 7 This is a comparison of the response time and resource usage of the present invention, wherein, Figure 7 (a) Comparison of response times at a reliability of 0.99. Figure 7 (b) Comparison of resource usage at a reliability of 0.99. Figure 7 (c) Comparison of response times at a reliability of 0.9999. Figure 7 (d) Comparison of resource usage at a reliability of 0.9999; Figure 8 This is a schematic diagram of the response time distribution function (CDF) of the present invention, wherein... Figure 8 (a) is the response time distribution function plot at a reliability of 0.99. Figure 8 (b) is the response time distribution function graph at a reliability of 0.9999; Figure 9 This is a schematic diagram illustrating the average time consumption of four metrics for each request in this invention, wherein... Figure 9 (a) Average time elapsed when reliability is 0.99. Figure 9 (b) Average time consumed when reliability is 0.9999; Figure 10 This is a schematic diagram illustrating the changes in response time, resource usage, and average number of replicas (ANR) as a function of reliability in this invention. Figure 10 (a) is a schematic diagram of the response time. Figure 10 (b) is a diagram illustrating resource usage. Figure 10 (c) is a diagram illustrating the average number of copies; Figure 11 This is a schematic diagram illustrating the changes in response time, resource usage, and average replica count (ANR) as a function of dynamic ratios in this invention. Figure 11 (a) is a schematic diagram showing the change in response time as a function of the dynamic ratio. Figure 11 (b) is a schematic diagram showing how resource usage changes with dynamic ratios. Figure 11 (c) is a schematic diagram showing the change in the average number of copies as a function of the dynamic ratio; Figure 12 This is a schematic diagram illustrating how the response time, resource usage, and average number of replicas (ANR) of the present invention vary with the size of the call graph, wherein... Figure 12 (a) is a schematic diagram showing how the response time varies with the size of the call graph when the reliability is 0.99. Figure 12 (b) is a schematic diagram showing how resource usage changes with the size of the call graph when the reliability is 0.99. Figure 12 (c) is a schematic diagram showing how the average number of replicas varies with the size of the call graph when the reliability is 0.99. Figure 12(d) is a schematic diagram showing how the response time varies with the size of the call graph when the reliability is 0.9999. Figure 12 (e) is a schematic diagram showing how resource usage changes with the size of the call graph when the reliability is 0.9999. Figure 12 (f) is a schematic diagram showing how the average number of replicas varies with the size of the call graph when the reliability is 0.9999; Figure 13 The response time and resource usage of different pre-deployment methods of the present invention are shown below. Figure 13 (a) Comparison of response times when reliability is 0.99. Figure 13 (b) Comparison of resource usage when reliability is 0.99. Figure 13 (c) Comparison of response times when reliability is 0.9999. Figure 13 (d) Comparison of resource usage when reliability is 0.9999; Figure 14 This is a comparison chart of the average time consumption of various request metrics under the four pre-deployment methods of this invention, wherein... Figure 14 (a) represents the time consumption under different strategies when the reliability is 0.99. Figure 14 (b) Time consumption under different strategies when reliability is 0.9999; Figure 15 This invention relates to the impact of the threshold η on response time and resource usage, wherein... Figure 15 (a) Response time as η changes when reliability is 0.99. Figure 15 (b) When the reliability is 0.99, resource usage varies with η. Figure 15 (c) shows the response time as a function of η when the reliability is 0.9999. Figure 15 (d) is the resource usage as η changes when the reliability is 0.9999; Figure 16 This is a schematic diagram illustrating the response time and resource usage of different heuristic pre-deployment strategies of the present invention, wherein... Figure 16 (a) The ablation test results when the reliability is 0.99. Figure 16 (b) shows the ablation test results when the reliability is 0.9999. Detailed Implementation

[0023] A microservice fault-tolerant scheduling method, such as Figure 1 As shown, the present invention includes: S1. Obtain historical microservice call data for the same application, and obtain the call graph of a single request based on the historical microservice call data; Specifically, microservices are a software architecture style that breaks down a single application into a set of small, independent services. Each service is built around a specific business function, can be developed, deployed, and scaled independently, and collaborates through lightweight communication mechanisms such as HTTP / REST and gRPC. It is an evolution of the traditional monolithic architecture, aiming to improve the flexibility, maintainability, and scalability of the system.

[0024] In a microservice system, a "task" describes the execution load when a microservice is invoked. For online applications that require users to specify reliability requirements, when a user sends a request, the request must first be processed by the entry microservice. We call this load the entry task and place it in a task queue. Tasks in the queue are scheduled by the scheduling system and executed in the cloud resource pool. If a downstream microservice needs to be called during processing, the corresponding task will be put back into the task queue. Otherwise, after processing is complete, the system will decide whether to return the result to the upstream microservice based on the invocation pattern, such as... Figure 2 As shown.

[0025] Following common practices in fault tolerance research, we model the failures of microservices and bare-metal nodes as transient failures. These transient failures have a very short duration, and their distribution follows a Poisson process. We use... This represents the failure rate of a microservice, assumed to follow a Poisson distribution. The probability of a microservice successfully processing a task is given by... The following is an expression:

[0026] in, Indicates task Processing time, For the task The probability of successful processing For microservice failure rate.

[0027] lemma: Given two independent Poisson random variables, X ~ Poisson (parameters are...) ) and Y~Poisson (parameter is Their sum Z = X + Y is also a Poisson random variable with parameters. + That is, Z ~ Poisson.

[0028] Similarly, the failures of bare-metal nodes follow a Poisson distribution ( Assume that failures originating from microservices and bare-metal nodes are independent. Therefore, according to the lemma, when considering both types of failures, a single task... The probability of successful completion is determined by Given, as follows:

[0029] Where b is a bare-metal node.

[0030] Through active replication, the task It may be handled by different microservices Multiple copies are used to enhance the reliability of mission completion. The mission is considered successful if at least one copy completes successfully. Overall mission reliability... It is expressed as follows:

[0031] A request is considered successful only when all its constituent tasks complete successfully. The overall request reliability is calculated as follows:

[0032] Where G is the call graph, and |G| is the size of the graph.

[0033] The call graph describes how the various microservices call each other sequentially when processing a user request.

[0034] For example, when user A clicks "Buy" a book, the request first reaches the entry microservice (such as the gateway). The gateway first calls the user service to verify A's identity and address (a two-way call, requiring a return result). After successful verification, it calls the product service to confirm book inventory (a two-way call), then calls the order service to generate a new order (a two-way call). Next, the order service may asynchronously call the message queue service to notify the warehouse to prepare for shipment (a one-way call, sending a message is sufficient, no reply is needed). Simultaneously, the gateway calls the payment service to complete the deduction (a two-way call). After successful payment, the gateway returns the result to user A. The call relationships formed in this process constitute a call graph, containing a complex network of branches, loops (not shown in this example), and different types of calls (one-way / two-way), such as... Figure 3 and Figure 4 As shown.

[0035] S2. Set up a virtual super root that connects to the entry point of the call graph for all requests; Specifically, in order to construct a complete application-level call graph that reflects the commonalities and differences of all requests, a virtual super root is introduced, such as... Figure 5 As shown, the different entry vertices of all requests are connected. The request call graph is then merged in a top-down manner: if the edge type and vertex type match, the vertex is merged; if they do not match, a new branch is created at that point.

[0036] S3. Generate an application-level call graph based on the virtual super root and historical microservice call data, and model the application-level call graph as a Bayesian probability tree. Specifically, user requests first reach the entry microservice, which then issues a series of calls to other related microservices. These microservices are abstracted as vertices (regardless of their state), and dependencies are represented as directed edges. We distinguish between two edge types: unidirectional edges, representing downstream vertices (…). It will not return a result to the caller (e.g., a message queue). Upstream vertex In the direction The duration must be completed before sending the call. Preparation work. Subsequently, the call will be made with transmission time... Transmitted in a way .

[0037] Based on the virtual superroot and historical microservice call data, an application-level call graph is generated and modeled as a Bayesian probability tree, including: Based on the virtual super root and historical microservice call data, vertex attributes and edge attributes are obtained. Vertex attributes include microservice type, microservice location, processing time, and partition attribute. Edge attributes include upstream or downstream vertex, edge type, probability of edge existence in a specific request, preparation time before call, and data transfer volume. Based on vertex and edge attributes, an application-level call graph is generated, and the application-level call graph is modeled as a Bayesian probability tree.

[0038] Specifically, the call graph of a single request can be simplified into a tree structure. At the application level, requests within the same application may meet the following two conditions: 1. They originate from different entry microservices (entry nodes); 2. They share the same call prefix (edge). To construct a complete application-level call graph that simultaneously reflects both commonalities and differences, we adopt the following processing flow: A virtual superroot is introduced, connecting all the different entry vertices in the request-call graph. This prevents the application-level call graph from degenerating into a disconnected forest when requests have different entry points. Starting from the superroot node, we insert nodes into the request-call graph one by one from top to bottom. At each step, if the edge type and vertex type match, the vertices are merged (i.e., overlapped), such as... Figure 5 (a) and Figure 5 As shown in (b). If a mismatch occurs, a new branch is created at that point, and its descendants are no longer merged, as... Figure 5 (c) and Figure 5 As shown in (d), this method generates an application-level call graph that preserves the shared structure while making the branching points explicit, thus providing a more comprehensive view of the application.

[0039] Finally, we model the application-level call graph as a Bayesian probability tree. .in, This represents the set of all vertices (microservices or their tasks). Each vertex has four attributes: .in, This represents a vertex identifier used to locate its position in the call graph. Specify the vertex type (corresponding to the microservice type). Indicates processing time. Indicates a partition. This represents the set of all edges, which symbolize the dependencies between tasks, i.e., the calling relationships between microservices. Each edge contains seven attributes: .in, and These represent the upstream vertex and the downstream vertex, respectively. Indicates the type of edge (one-way edge and two-way edge). This indicates the probability that the edge exists in a specific request. This indicates the time interval between the start of execution of the upstream task and the sending of a request to the downstream task. Indicates from arrive The amount of data, similarly Indicates from arrive The amount of data. This related metadata enables fault-tolerant scheduling even when information is incomplete.

[0040] As can be seen from the above, the request-call graph has two significant characteristics: Partitioning: Only bidirectional edges force the caller to wait for the callee's result; unidirectional edges do not block the caller's response. Therefore, unidirectional edges are used as boundaries to partition the tree. Only vertices in the same partition as the entry vertex (partition 0) can directly influence it.

[0041] Microservice-based online applications run in an Infrastructure as a Service (IaaS) cloud environment, where bare-metal nodes are a core component of the infrastructure. The resource pool consists of a set of bare-metal nodes. Composition, in which For the first bare metal nodes Corresponding resource capacity >, among which CPU capacity For memory capacity, This refers to disk capacity. Microservices typically run in container environments, such as container orchestration platforms like Kubernetes. Given a microservice... Container resource requirements , representing their respective CPU, memory, and disk requirements. To ensure stable system operation, the total resource requirements (in sets) of all containers deployed on bare metal node b are... (This indicates that) the capacity limit of the node must not be exceeded, such as... Figure 5 As shown.

[0042]

[0043] In the call graph of a given request, the transfer time is... Calculations show that For the task and tasks The data size between. When and When processing on the same node, bandwidth It is considered to be infinitely large. Defined as The predecessor vertex, That is The successor vertex. The set of predecessor vertices is denoted as . , The set of successor vertices is denoted as .

[0044] container Medium task The earliest available time is the initialization end time. With container The maximum value between the latest completion times of the inner prerequisite tasks is expressed as:

[0045] in, Indicates the plan to schedule to the container The task set, Cold start time, For containers Prerequisites The latest completion time.

[0046] If container If it is running, there is no cold start time; if it is not running, a cold start time is required, as shown below:

[0047] Task The completion time is the sum of its start time and processing time. , is represented as:

[0048] in, For the task The start time, For the task The completion time.

[0049] The start time needs to consider two aspects: firstly, the availability time of the container; and secondly, the latest time of each previous task call in the call graph. For tasks with multiple replicas, the earliest call time among all replicas is considered the current task's call time. The call time, and ultimately, the start time is the maximum of these two times, expressed as:

[0050] in, For the task The start time, For the task container Available time, The time interval between the start of execution of the upstream task and the sending of a request to the downstream task. represent The father's task, A collection of copies of the parent quest. Father's Quest Dungeon To the mission Transmission time.

[0051] Task The return time is the maximum of its completion time and the maximum return time of subsequent tasks within the same partition.

[0052]

[0053] in, For the task Return time, For the task Completion time, Father's Quest Dungeon To the mission Transmission time, For the task The return time.

[0054] S4. Set up a resource pool that includes a static resource pool and an elastic resource pool. The resource pool consists of a group of bare-metal nodes, each of which has a corresponding resource capacity. The static resource pool is used to deploy static microservice instances, and the elastic resource pool is used to create dynamic microservice instances on demand. Specifically, for microservice-based online application systems, a scheduling model was built to optimize response time and resource consumption while ensuring the reliability of requests to the IaaS cloud platform. This model uses N to represent a set of mappings from containers to bare metal nodes, where each mapping... This indicates the container From the start date of the lease Until the end of the lease Deployed on bare metal nodes superior. and The calculation method is as follows:

[0055]

[0056] in, For return time, This is the start time of the task.

[0057] Total response time It is expressed as follows:

[0058] in, For partition 0, For the task A collection of copies, To carry out the mission Container.

[0059] Total resource usage It is expressed as follows:

[0060] in , and These represent the weights of CPU, memory, and disk resources, respectively, and can be configured according to user needs. For the task Required CPU capacity For the task Required memory capacity For the task Required disk capacity.

[0061] In microservice-based online applications, the response time and resource utilization optimization problem, under reliability guarantees, can be expressed by the following equation:

[0062] in, To call the graph The actual reliability, To meet the reliability requirements of users.

[0063] To address the aforementioned issues, a dual-pool scheduling framework was designed, and a scheduling algorithm called DRILF was proposed based on it. Given the stringent response time requirements of online applications, cold start time cannot be ignored. Therefore, the plan is to mitigate the impact of cold start on response time by pre-deploying microservices. To this end, a dual-pool scheduling framework was designed, which divides the cloud resource pool into a static resource pool and an elastic resource pool, and uses different deployers to deploy microservices, such as... Figure 6 As shown, the dual-pool scheduling framework consists of the following four components: 1. Call Graph Analyzer: Analyzes historical call graph data and constructs a call graph. In addition, it adjusts the call graph structure for specific requests in real time.

[0064] 2. Performance Evaluation: Includes evaluation of three metrics. (1) Response Time: Evaluate the completion time of different microservices, and the results are used to guide subsequent task scheduling and microservice deployment. (2) Resource Usage: Record and evaluate the resource usage of different bare metal nodes, and use it to guide microservice deployment.

[0065] Reliability: Calculate the number of replicas for each task based on the real-time call graph structure and reliability requirements.

[0066] 3. Task Scheduler: Schedules tasks to microservices based on the number of replicas and the status of existing microservices.

[0067] 4. Microservice Deployer: Includes two types of microservice deployment. (1) Static microservice: During the initialization phase, static microservices are deployed based on historical information in the static resource pool. (2) Elastic microservice: During the runtime phase, microservice deployment is adjusted based on the workload in the elastic resource pool.

[0068] To address the dynamic nature and incomplete information issues of online applications in microservice architectures, a fault-tolerant scheduling (drift) scheme driven by dynamic replica evaluation is proposed based on a dual-pool scheduling framework. This scheme achieves optimization through a dual mechanism: First, it shortens response time by accurately pre-selecting and deploying microservices. The system plots quantity change curves for various microservices based on historical data and calculates the average interval time under different quantities. By comparing this with a baseline time, the optimal quantity for each type of microservice can be determined. Second, the scheme effectively reduces resource consumption by dynamically calculating the number of replicas. Its core mechanism is dynamic adjustment based on the real-time call graph structure and reliability requirements. The system constructs a complete call graph based on historical data. When a task is invoked, the current call graph structure is pruned in real time to reduce its size. This mechanism reduces the reliability requirements of individual tasks, ultimately accurately determining the required number of replicas.

[0069] S5. Deploy static microservice instances in the static resource pool based on historical microservice call data and cold start time; Based on historical microservice call data and cold start time, deploy static microservice instances in the static resource pool, including: Generate quantity-time curves for different microservice types based on historical microservice call data; Initialize the number of static microservice instances, y. Based on the quantity-time curve, we can obtain the number of intervals from when the static microservice instance quantity demand is below y to when it reaches or exceeds y, and the total duration when the static instance quantity is less than y at time t. The average call interval time is obtained based on the number of intervals and the total duration. If the average call interval is less than the interval parameter, the number of static instances is increased until the average call interval is greater than or equal to the interval parameter. The interval parameter is obtained from the cold start time and the threshold parameter. If the average call interval is greater than or equal to the interval parameter, then the maximum number of static microservice instances deployed for different microservice types is obtained. Static microservice instances are deployed in the static resource pool based on the maximum number of static microservice instances deployed.

[0070] To mitigate the impact of frequent cold starts on response time, a subset of microservices requiring pre-deployment is identified through historical data analysis. Deploying too many static microservices can lead to excessive resource consumption, while deploying too few can cause response latency issues. Based on historical data, cold start time is used as the criterion to determine whether static microservices need to be deployed and their specific number. The core idea is: for a certain type of microservice, if the average call interval of n microservices in historical data is less than its cold start time multiplied by η, then the benefit of reducing latency is considered to outweigh the additional resource consumption, and n such static microservices can be deployed, where η is a threshold parameter.

[0071] First, a curve showing the relationship between the number of microservices and time needs to be obtained. Based on this curve, the average call interval time for different numbers of microservices is calculated. Simultaneously, the final deployment quantity is determined using cold start time × η, where η is a threshold reflecting the user's tolerance for idle time. The larger the value of η, the higher the tolerance for microservice idle time. Finally, microservices are deployed in descending order of quantity. Based on the call distribution of the deployed microservices, different types of microservices are deployed to different nodes. This deployment strategy not only improves system fault tolerance but also reduces latency caused by inter-node communication. On one hand, when deploying the first type of microservice, since its initial call volume is zero, a uniform distribution can be formed. This strategy effectively reduces the risk of system unavailability due to server failure, thereby enhancing system resilience. On the other hand, subsequent deployments will be adjusted according to the call frequency of each server, allowing high-frequency microservices to be deployed centrally, thereby reducing inter-node communication latency.

[0072] S6. Obtain the actual microservice request; S7. Prune the Bayesian probability tree according to the microservice requests to obtain the actual call graph; S8. For each task in the actual call graph, calculate the minimum number of replicas required to execute the task based on the reliability of each task. For each task in the actual call graph, calculate the minimum number of replicas required to execute the task based on the reliability of each task, including: Based on the actual call graph, global parameters are obtained. These global parameters include the number of tasks to be completed, the number of tasks that have been successfully scheduled, the overall request reliability required by the user, and the cumulative reliability of the tasks that have been successfully scheduled. The task list is obtained based on the actual microservice requests. Each task in the task list is sorted in ascending order according to its average processing time to obtain the sorting result. Select the task as the current task according to the sorting results, initialize the current number of replicas for the current task, and calculate the task success rate with the current number of replicas. Calculate the required reliability for the current task; If the task success rate is less than the required reliability, increase the current number of replicas until the task success rate is greater than or equal to the required reliability. If the task success rate is greater than or equal to the required reliability, then the number of replicas when the task success rate is greater than or equal to the required reliability is obtained as the minimum number of replicas for the current task. The sorted results are then iterated to obtain the minimum number of replicas required for each task to execute.

[0073] Specifically, to address the dynamic structure of the call graph, the number of replicas is dynamically evaluated in real time based on its configuration in order to optimize resources.

[0074] By reducing redundant replicas and optimizing resource utilization, a dual-core strategy is proposed: First, during the processing of each task, the corresponding call graph is pruned, thereby reducing the reliability requirements of individual tasks. Second, tasks are arranged in ascending order of processing time, allowing replicas of tasks with longer processing times to be computed later. This strategy effectively reduces the reliability requirements of high-processing-time tasks, ultimately reducing their number of replicas and lowering overall resource consumption. By focusing on these two dimensions, the proposed method significantly improves resource utilization efficiency.

[0075] First, update the structure of the call graph to obtain the total number of tasks. Number of scheduled tasks Overall reliability requirements and the reliability of scheduled tasks Then, the unscheduled tasks are sorted in ascending order based on their average processing time. Next, the tasks in the task list are traversed, and the number of replicas for each task is calculated, until the current task is the target task, at which point the loop exits. The required reliability of the current task is calculated using the following formula:

[0076] in, To ensure the required reliability for the current task, the algorithm should be used for the first time. and Initialized to 0 and 1. Average processing time is obtained from historical data.

[0077] In the DRIFT algorithm for microservices, there are two sub-algorithms: static microservice deployment and dynamic scheduling. The time complexity of static microservice deployment is O(n log n). Where I represents the number of microservice types, y represents the maximum number of individual microservice types (determined based on historical data), and |B| represents the number of bare-metal nodes. The time complexity of dynamic scheduling is O(n log n). Where |Ψ| represents the size of the unscheduled task list, p represents the number of replicas (usually close to 1), and q represents the number of microservices. Based on the above analysis, the total time complexity of the microservice DRIFT algorithm is O(n). + According to the analysis, 90% of the call graphs are smaller than 40, indicating that... The time complexity is usually small, so the time complexity of the DRIFT algorithm is manageable.

[0078] S9. Based on the minimum number of replicas for each task, select the method that can execute the task earliest from three options: static microservice instances already deployed in the static resource pool, dynamic microservice instances already created in the elastic resource pool, and newly created dynamic instances in the elastic resource pool. Among them, the deployed static microservice instances can be called directly, the dynamic microservice instances already created in the elastic resource pool have a buffer time to execute the task, and the newly created dynamic instances have a creation time. Based on the minimum number of replicas for each task, the target method is selected from three options: static microservice instances already deployed in the static resource pool, dynamic instances already created in the elastic resource pool, and newly created dynamic instances in the elastic resource pool. The method that executes the task earliest is chosen as the target method. Based on the minimum number of replicas for each task, search for all instances of the same type from the static microservice instances deployed in the static resource pool and the dynamic instances created in the elastic resource pool as a candidate instance list; Calculate the estimated start time for assigning tasks to candidate instances from the candidate instance list; Get the new estimated start time for creating a new dynamic instance in the Elastic Resource Pool; Compare the estimated start time of all candidate instances in the candidate instance list with the new estimated start time for creating a new instance, and find the minimum value as the minimum start time; Get the instance type corresponding to the minimum start time. The instance type includes static microservice instance deployment, existing dynamic instance, and newly created dynamic instance. Choose the instance type corresponding to the minimum start time as the target method.

[0079] Specifically, this application employs the DRIFT algorithm for scheduling, which includes two sub-algorithms: static instance deployment and dynamic scheduling. During system initialization, static microservices are deployed to a static resource pool. This step is performed only during system initialization. Subsequently, each task in the queue is scheduled sequentially. For each task, the number of replicas is obtained through the algorithm. Next, each task replica is scheduled in a greedy manner by searching for the earliest available microservice instance. It is worth noting that in online applications, it is impossible to accurately obtain specific information about the current call graph, especially processing time, which necessitates the evaluation of the completion time of each microservice instance based on historical data. The estimated completion time (EFT(c)) of microservice c is calculated as follows:

[0080] in, Indicates the current time. Indicates an ongoing task In microservice c, the estimated completion time, A, is the time excluding the current task. In addition, the sum of the execution times of other scheduled and running tasks is represented as follows:

[0081] in, express The estimated remaining execution time is expressed as follows:

[0082] in, This is the start time of the task. For task processing time, The container for handling the current task.

[0083] This represents the average execution time. Finally, the estimated start time of task k in microservice c is calculated and expressed as:

[0084] in, This indicates the upstream task. If the newly created elastic microservice instance starts earlier than an existing microservice instance, the new microservice instance will be deployed and the task will be scheduled to that microservice. Otherwise, the task will be scheduled to... Minimal microservices. When all elastic microservices are idle, the system will perform an instance deletion operation.

[0085] S10. Implement microservice scheduling for microservice requests based on the target method.

[0086] Based on the target approach, implement microservice scheduling for microservice requests, including: If the target is a deployed static microservice instance or a dynamically created instance in an elastic resource pool, the replica will be scheduled to the deployed static microservice instance or the dynamically created instance in an elastic resource pool. If the goal is to create a new dynamic instance in the elastic resource pool, then create a new microservice instance and schedule the replica to the newly created microservice instance.

[0087] Experimental setup: Using the SimPy toolkit developed in Python and combined with discrete event simulation theory, a simulation platform specifically designed for microservice architectures was constructed. This platform can systematically evaluate the performance differences between the proposed DRIFT and state-of-the-art methods. All experiments were conducted on a server equipped with an Intel Xeon Platinum 8352V processor (32 virtual CPUs) and 120GB of memory.

[0088] Dataset: All experiments in this study used real datasets. The dataset comes from the world microservice cluster dataset cluster-trace-microservice-v20221, which records the behavior of more than 20,000 microservices over 7 days. In the experiment, the dataset was divided into historical dataset and validation dataset: the first hour constituted the historical data and the second hour constituted the validation data.

[0089] Platform: In the simulated environment, the system consists of 100 bare-metal nodes, each configured with three resource types—CPU, memory, and disk. The maximum capacity for each resource type is 100 units per server. For each microservice, its resource requirements are randomly sampled from a uniform range of 1 to 5 units. The failure rate of the bare-metal nodes is distributed at 10 units per hour. -4 Up to 10 -3Within this timeframe, due to the complex causes and high frequency of microservice failures, the failure rate is uniformly distributed between 1 / 30 of an hour and 1 hour. The bandwidth between bare-metal nodes is set to 20 megabits per second. Microservice cold start time varies between 83 milliseconds and 1100 milliseconds, depending on the amount of code and initialization requirements. Values ​​uniformly distributed within this timeframe are used as baseline parameters for different microservices. It is worth noting that 0.99 and 0.9999 are widely used benchmarks for evaluating algorithm performance under different reliability constraints. In reliability testing, the required value range is gradually increased from 0.95 to 0.9999.

[0090] Baseline: Demonstrates the effectiveness of DRIFT. Comparative experiments were conducted on several cutting-edge fault-tolerant scheduling methods. (1) RR: Greedy scheduling iteration based on processor cumulative reliability. (2) R_RIR: Calculate the minimum redundancy of the workflow while meeting reliability requirements using the reliability increment ratio (RIR). (3) CGM: Reduce execution cost by minimizing redundancy using the geometric mean method. (4) QEFC: Effectively reduce execution cost while meeting workflow reliability requirements. (5) GSMS: Minimize the execution cost of microservice-based workflow applications while meeting deadline and reliability constraints. (6) A method combining GSMS with a high-frequency-first pre-deployment strategy (selecting the n microservices with the highest frequency). Since RR, RRIR, CGM, and QEFC lack microservice deployment strategies, they were adapted by integrating the microservice deployment strategy of GSMS. For algorithms that require input deadlines (such as GSMS), a deadline is assigned to each request. Furthermore, η is set to 0.6 by default. To ensure consistency in the experiments, the number of pre-deployed microservices in GSMS-T is set to be the same as that in DRIFT.

[0091] Metrics: DRIFT is evaluated from multiple perspectives. In addition to the main objectives—response time and resource usage—the following metrics are used in different experiments: Average Number of Replicas (ANR): The ANR is calculated as shown in the equation, where K represents the set of all tasks, and |·| represents the number of elements in the set. A collection of copies of the parent quest.

[0092]

[0093] Cold start time of the request: The total cold start time experienced by all tasks in the request within their respective microservices.

[0094] Request wait time: The total wait time for all tasks in the request, including queuing and downstream response wait.

[0095] Request execution time: The total processing time for all tasks in the request.

[0096] Request transmission time: The total time for data transmission between the two microservices in the request.

[0097] PersistentHitCount (PHC): The average number of microservice calls served by a persistent microservice per request.

[0098] To ensure the objectivity and accuracy of the experiment, ten tracking data points generated by the microservice tracing system were used for each experimental scenario.

[0099] Comparison with the latest methods: The systematic DRIFT approach is compared with cutting-edge techniques. Ten independent data trajectories were extracted from the microservice tracing dataset. Figure 7 The comparison results for two key metrics, response time and resource utilization, are presented. Data shows that DRIFT consistently minimizes response time at both 0.99 and 0.9999 reliability settings, outperforming other methods by 58.81–66.58%. Regarding resource consumption, DRIFT achieves a reduction of 6.71–20.06% at high reliability (0.9999). At lower reliability (0.99), while DRIFT's resource consumption is slightly higher than the non-pre-deployment baseline method (1.38–13.38%), it is still 17.51% lower than GSMS-T. This indicates that although pre-deployment introduces basic resource costs, DRIFT's strategy has significantly higher resource efficiency than the high-frequency-first strategy in GSMS-T, even when the benefits of reducing task replicas are not significant at low reliability levels.

[0100] To more intuitively illustrate the trend of response time, the cumulative distribution function (CDF) of the response time up to the 99th percentile (P99) was plotted, as shown below. Figure 8 As shown in the figure, the results show a high degree of consistency in trends across all reliability levels. DRIFT (red line) exhibits the steepest and leftmost curve, indicating that its latency is significantly lower than all benchmark methods. Notably, although GSMS-T outperforms non-pre-deployment methods such as RR with its static pre-deployment strategy, its efficiency still lags behind DRIFT's dynamic Bayesian method.

[0101] Because a proactive, static microservice architecture is used, cold starts are unnecessary. However, this also results in microservices remaining idle, leading to wasted resources.

[0102] In addition, the response time for each request was analyzed, and the cold start, waiting, execution, and transmission times were recorded. Figure 9The data shows the average values ​​of these four time types across all requests under different reliability requirements. DRIFT minimizes response latency by effectively eliminating cold start latency and its associated wait times. While GSMS-T offers a slight improvement in reducing cold start latency through static pre-deployment, it still lacks the precision of DRIFT. The core difference lies in DRIFT's Bayesian selectivity: unlike blind warm-up strategies that require massive over-provisioning of resources to achieve similar performance, DRIFT precisely targets high-probability services. This allows it to maintain extremely low latency while avoiding the enormous resource consumption associated with full pre-deployment.

[0103] The overall reduction in resource usage is primarily attributed to the advantage of the DRIFT algorithm using fewer replicas. Since the number of replicas directly impacts resource consumption, the average number of replicas for each task was statistically analyzed from a resource usage perspective (see Table 1). The results show that under high reliability requirements (0.9999), DRIFT is extremely efficient, reducing ANR by 33.12–29.45% compared to the baseline method. However, when reliability constraints are looser (0.99), the reduction in ANR is insufficient to fully offset the additional overhead from pre-deployed microservices, resulting in slightly higher resource consumption.

[0104] Table 1: Average Number of Replicas per Task (ANR)

[0105] Overall, DRIFT significantly reduces response time by pre-deploying microservices, but it increases resource consumption. Furthermore, by dynamically calculating reliability and reducing the number of replicas, DRIFT not only alleviates the resource pressure from pre-deployment but also slightly reduces overall resource usage.

[0106] Durability rating: To analyze the robustness of DRIFT performance, experiments were conducted on three dimensions: reliability, dynamic scaling of the call graph, and call graph size.

[0107] Reliability Evaluation: Used to analyze the reliability of each request under different reliability settings for DRIFT performance, with reliability values ​​set to 0.95, 0.99, 0.999, 0.9999, and 0.99999 respectively. All other settings remain at their default values.

[0108] The relationship between response time and resource usage as reliability changes was recorded. Figure 10 (a) shows that regardless of changes in reliability, DRIFT response time remains at a very low level, demonstrating a significant advantage (response time reduction of 43.09–67.22%). Regarding resource consumption, such as... Figure 10(b) DRIFT did not incur significant additional overhead; notably, at more stringent reliability settings (e.g., 0.99999), it even reduced resource consumption by 6.01–13.03%. This efficiency improvement stems from the trade-off between replica reduction and pre-deployment costs. At higher reliability levels, DRIFT's precise replica calculation significantly reduced the average replica count (ANR), effectively offsetting the overhead of pre-deployment. Conversely, at lower reliability levels, the reduction in ANR was too small to fully offset the pre-deployment costs. Specifically, Figure 10 (c) reveals that as reliability increases from 0.95 to 0.99999, the ANR reduction achieved by DRIFT increases dramatically from negligible 0.07–1.79% to significant 31.59–33.74%.

[0109] Regarding the impact of the dynamic ratio on drift performance, the dataset was divided into five trajectories based on the dynamic ratio D, with the following ranges: (0, 0.2], (0.2, 0.4], (0.4, 0.6], (0.6, 0.8], and (0.8, 1).

[0110] It records the correlation between response time and resource usage as a dynamic ratio, such as... Figure 11 As shown in the figure, experimental data shows that as the dynamic ratio increases, the drift algorithm's advantage in response time gradually strengthens, with its response time reduction compared to the baseline value reaching 65.83% to 66.88% within the (0.8, 1.0) range. Simultaneously, the algorithm's disadvantage in resource consumption is also alleviated, showing only a slight decrease of 9.28%-19.83% within the (0.8, 1.0) range. Overall, a higher dynamic ratio indicates a smaller call graph size, which leads to a gradual reduction in both response time and resource consumption. Furthermore, since the Drif mechanism is specifically designed to handle the dynamic characteristics of the call graph, its advantages are more significant at higher dynamic ratios. Figure 11 As shown in (c), this reduces the number of replicas. This mechanism, in turn, helps to narrow the resource consumption gap, such as... Figure 11 As shown in (b), requests with a dynamic ratio between 0.8 and 1 accounted for a high proportion of 86.39%. The Drift mechanism performs particularly well in handling these types of requests. The high dynamic ratio demonstrates its effectiveness.

[0111] Use the call graph size for evaluation: used to analyze the impact.

[0112] To investigate the impact of call graph size on system DRIFT performance, experiments were conducted on call graphs of different sizes. Since approximately 90% of request call graphs are smaller than 50, the dataset was divided into five size intervals: (0, 10], (10, 20], (20, 30], (30, 40], and (40, 50). Other parameters remained at their default settings. The patterns of response time and resource usage as a function of call graph size were recorded, such as... Figure 12 As shown.

[0113] The results show that the response advantage of DRIFT is most significant in smaller call graphs and gradually diminishes with increasing size. Within the (0, 10] range, the response time reduction peaks at 72.74–75.97% with 0.99 reliability and at 70.88–75.91% with 0.9999 reliability. Conversely, the improvement in resource efficiency increases with the size of the call graph. In the largest range (40, 50], resource savings peak at 8.18–29.36% (0.99 reliability) and 22.59–34.64% (0.9999 reliability), respectively. Similarly, the reduction in ANR also increases positively with call graph size, reaching 40.53–42.89% at high reliability.

[0114] According to statistics based on tracking data, 96.47% of requests call graph sizes are between 0 and 10. For this majority of requests, the DRIFT method is particularly effective, significantly reducing response time while maintaining low resource usage and ensuring reliability.

[0115] Operational analysis: This paper analyzes the problem from three dimensions: pre-deployment heuristic analysis, parameter analysis, and ablation analysis.

[0116] Heuristic pre-deployment analysis: DRIFT includes Heuristic methods, used for selecting static microservices, determine how many of each type of microservice should be static. Inaccurate type reservations or an excessive number of static microservices can lead to additional resource consumption. The following demonstrates three common strategies as benchmarks: Random strategy: Randomly select n microservices.

[0117] High-frequency priority strategy: Select the n microservices with the highest deployment frequency.

[0118] Proportional allocation strategy: Based on historical call frequency, select n microservices according to proportion.

[0119] This application (our strategy): For different microservice types, select n microservices by iteratively analyzing resource idleness and cold start latency.

[0120] For ease of comparison, in the experiment, the n value of the baseline group was set to be the same as that of DRIFT, that is, the same number of microservices were pre-deployed, and all other settings remained at their default values.

[0121] Response time and resource usage were recorded using different pre-deployment strategies, such as Figure 13 As shown in the diagram, the red dashed line represents the average performance of non-pre-deployment methods. DRIFT consistently outperforms all baseline strategies, reducing response time by 57.69-65.48% and saving resource consumption by 2.02-21.87%. Crucially, while the high-frequency-first and proportional allocation strategies only focus on call frequency, they fail to consider the overlap of requests over time. In contrast, DRIFT precisely identifies microservices that may become bottlenecks through interval analysis, thus avoiding both cold starts due to under-configuration and resource waste due to over-configuration.

[0122] In addition, the idle time of static microservices under different pre-deployment strategies was compared, such as Figure 14 As shown, the strategy presented in this application significantly outperforms baseline and other heuristic strategies. DRIFT's performance advantage primarily stems from its dramatic reduction in cold start latency (79.07%–87.03%) and corresponding waiting time (83.58%–89.61%). This demonstrates that our time-interval-based heuristic algorithm can effectively predict request spikes missed by frequency-based baseline methods, thereby ensuring that "hot" microservices are accurately ready when needed.

[0123] Table 2 shows the Persistent Hit Count (PHC) metric for each request. A higher PHC means that the warmed-up microservices are actively processing requests, rather than being idle. DRIFT achieved a PHC of 4.26–4.30, significantly outperforming the high-frequency-first (1.38) and proportional allocation (1.87) strategies. This stark contrast highlights the limitations of frequency-based approaches: high call frequency does not naturally imply a need for a large number of pre-deployed microservices. If the request lifecycle is short, even frequently called services may require only a very small number of persistent microservices. By calibrating the number of pre-deployed services based on actual concurrency requirements (considering frequency, execution duration, and arrival distribution), DRIFT maximizes resource utility.

[0124] Table 2: PHC index under different strategies

[0125] Parameter analysis: In the heuristic pre-deployment phase At the strategy level, there exists a threshold (η) that significantly affects the number of microservice deployments, thereby influencing response time and resource usage. This study focuses on experiments targeting this threshold, with η ranging from 0.1 to 1, increasing in increments of 0.1.

[0126] The response time and resource utilization changes were recorded as η changed, such as Figure 15 As shown, it is clear that as η increases, resource utilization increases almost linearly, while response time increases significantly.

[0127] The response time exhibits a clear inflection point at η=0.6, reaching a relatively ideal state at this point. Therefore, this study adopts η=0.6 as the default setting.

[0128] Ablation Analysis: To evaluate the impact of deployment and dynamic replicas on the system's DRIFT effect, ablation experiments were conducted. Experiments under each reliability level were divided into four groups, with options including pre-deployment and dynamic replicas. When pre-deployment was removed, the number of static microservices was set to zero; when dynamic replicas were omitted, the number of replicas was determined using the GSMS method. Each experiment used 10 data traces, with other parameters remaining at their default settings.

[0129] Experimental results are as follows Figure 16 As shown in the figure, the results reveal the different roles played by each component: First, pre-deployment is the primary acceleration engine, reducing response time by 50.21–50.76%. However, this speedup comes at the cost of resources, resulting in an increase in resource consumption of 8.13–24.28%. Second, dynamic replica generation acts as an efficiency stabilizer. It significantly reduces resource consumption by 22.21–30.17%, effectively offsetting the additional overhead of pre-deployment. Furthermore, when the two are integrated, it contributes an additional 14.21–14.39% to the response time reduction.

[0130] In summary, both mechanisms effectively reduce response time. While pre-deployment leads to a slight increase in resource consumption, dynamic replication further reduces resource usage, thus offsetting this impact.

[0131] Based on real-world microservice tracing data, we conducted extensive experiments. The results show that, under a typical reliability requirement of 0.9999, our method achieves a 58.81%–66.58% reduction in response time and a 6.71%–20.06% saving in resource usage compared to existing state-of-the-art methods.

[0132] A microservice fault-tolerant scheduling system includes: The first acquisition module is used to acquire historical microservice call data of the same application and obtain the call graph of a single request based on the historical microservice call data. The first configuration module is used to configure the virtual super root that connects to the entry point of the call graph for all requests; The Bayesian probability tree creation module is used to generate an application-level call graph based on the virtual super root and historical microservice call data, and to model the application-level call graph as a Bayesian probability tree. The second setting module is used to set up a resource pool that includes a static resource pool and an elastic resource pool. The resource pool consists of a group of bare-metal nodes, each of which has a corresponding resource capacity. The static resource pool is used to deploy static microservice instances, and the elastic resource pool is used to create dynamic microservice instances on demand. The static deployment module is used to deploy static microservice instances in the static resource pool based on the historical microservice call data and cold start time. The second acquisition module is used to acquire the actual microservice requests; The pruning module is used to prune the Bayesian probability tree according to the microservice request to obtain the actual call graph; The minimum number of replicas module is used to calculate the minimum number of replicas required to execute a task based on the reliability of each task in the actual call graph. The selection module is used to select the earliest method to execute the task from three options: static microservice instances already deployed in the static resource pool, dynamic microservice instances already created in the elastic resource pool, and newly created dynamic instances in the elastic resource pool, based on the minimum number of replicas for each task. Among them, the deployed static microservice instances can be called directly, the dynamic microservice instances already created in the elastic resource pool have a buffer time to execute the task, and the newly created dynamic instances have a creation time. The scheduling module is used to schedule microservice requests according to the target method.

[0133] Optionally, the Bayesian probability tree creation module includes: The data acquisition unit obtains vertex attributes and edge attributes based on the virtual super root and historical microservice call data. The vertex attributes include microservice type, microservice location, processing time, and partition attribute. The edge attributes include upstream or downstream vertex, edge type, probability of edge existence in a specific request, pre-call preparation time, and data transmission volume. Create a unit to generate an application-level call graph based on the vertex and edge attributes, and model the application-level call graph as a Bayesian probability tree.

[0134] This application also discloses a terminal device, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor loads and executes the computer program, a microservice fault-tolerant scheduling method is used.

[0135] The terminal device can be a computer device such as a desktop computer, a laptop computer, or a cloud server. The terminal device includes, but is not limited to, a processor and a memory. For example, the terminal device may also include input / output devices, network access devices, and buses.

[0136] The processor can be a central processing unit (CPU). Of course, depending on the actual use, it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc., and this application does not limit it.

[0137] The memory can be an internal storage unit of the terminal device, such as a hard disk or RAM of the terminal device, or an external storage device of the terminal device, such as a plug-in hard disk, smart memory card (SMC), secure digital card (SD), or flash memory card (FC) equipped on the terminal device. Furthermore, the memory can be a combination of internal storage units and external storage devices of the terminal device. The memory is used to store computer programs and other programs and data required by the terminal device. The memory can also be used to temporarily store data that has been output or will be output. This application does not limit this.

[0138] In this terminal device, a microservice fault-tolerant scheduling method from the above embodiments is stored in the terminal device's memory and loaded and executed on the terminal device's processor for convenient use.

[0139] This application also discloses a computer-readable storage medium, which stores a computer program, wherein when the computer program is executed by a processor, it employs a microservice fault-tolerant scheduling method as described in the above embodiments.

[0140] The computer program can be stored in a computer-readable medium. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or certain middleware. The computer-readable medium includes any entity or device capable of carrying computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the computer-readable medium includes, but is not limited to, the above-mentioned components.

[0141] In this embodiment, a microservice fault-tolerant scheduling method is stored in the computer-readable storage medium and loaded and executed on the processor to facilitate the storage and application of the method.

[0142] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of protection of this application is limited to these examples; under the concept of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of one or more embodiments of this application as described above, which are not provided in detail for the sake of brevity.

[0143] One or more embodiments in this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments in this application should be included within the protection scope of this application.

Claims

1. A microservice fault-tolerant scheduling method, characterized in that, include: Obtain historical microservice call data for the same application, and obtain a call graph for a single request based on the historical microservice call data; Set up a virtual super root that connects to the entry point of the call graph for all requests; Based on the virtual super root and historical microservice call data, an application-level call graph is generated, and the application-level call graph is modeled as a Bayesian probability tree. Set up a resource pool that includes a static resource pool and an elastic resource pool. The resource pool consists of a group of bare-metal nodes, each of which has a corresponding resource capacity. The static resource pool is used to deploy static microservice instances, and the elastic resource pool is used to create dynamic microservice instances on demand. Based on the historical microservice call data and cold start time, deploy static microservice instances in the static resource pool; Obtain the actual microservice request; The Bayesian probability tree is pruned based on the microservice request to obtain the actual call graph; Calculate the minimum number of replicas required to execute a task based on the reliability of each task in the actual call graph; Based on the minimum number of replicas for each task, the method that can execute the task earliest is selected from three options: static microservice instances already deployed in the static resource pool, dynamic microservice instances already created in the elastic resource pool, and newly created dynamic instances in the elastic resource pool. Among them, the deployed static microservice instances can be called directly, the dynamic microservice instances already created in the elastic resource pool have a buffer time to execute the task, and the newly created dynamic instances have a creation time. Based on the stated objective, implement microservice scheduling for microservice requests; The step of generating an application-level call graph based on the virtual super root and historical microservice call data, and modeling the application-level call graph as a Bayesian probability tree, includes: Based on the virtual super root and historical microservice call data, vertex attributes and edge attributes are obtained. The vertex attributes include microservice type, microservice location, processing time, and partition attribute. The edge attributes include upstream or downstream vertex, edge type, probability of edge existence in a specific request, preparation time before call, and data transmission volume. Based on the vertex and edge attributes, an application-level call graph is generated, and the application-level call graph is modeled as a Bayesian probability tree.

2. The microservice fault-tolerant scheduling method as described in claim 1, characterized in that, The step of deploying static microservice instances in the static resource pool based on the historical microservice call data and cold start time includes: Generate quantity-time curves for different microservice types based on the historical microservice call data; Initialize the number of static microservice instances, y. Based on the quantity-time curve, the number of intervals from when the number of static microservice instances drops below y to when it reaches or exceeds y, and the total duration during which the number of static instances is less than y at time t are obtained. The average call interval time is obtained based on the number of intervals and the total duration. If the average call interval is less than the interval parameter, the number of static instances is increased until the average call interval is greater than or equal to the interval parameter, which is obtained through cold start time and threshold parameter. If the average call interval is greater than or equal to the interval parameter, then the maximum number of static microservice instances deployed for different microservice types is obtained. Static microservice instances are deployed in the static resource pool based on the maximum number of static microservice instances deployed.

3. The microservice fault-tolerant scheduling method as described in claim 1, characterized in that, The calculation of the minimum number of replicas required to execute a task based on the reliability of each task in the actual call graph includes: Based on the actual call graph, global parameters are obtained, including the number of tasks to be completed, the number of tasks that have been successfully scheduled, the overall request reliability required by the user, and the cumulative reliability of the tasks that have been successfully scheduled. Based on the actual microservice requests, a task list is obtained, and each task in the task list is sorted in ascending order according to its average processing time to obtain the sorting result. Select a task as the current task according to the sorting results, initialize the current number of replicas for the current task, and calculate the task success rate under the current number of replicas. Calculate the required reliability for the current task; If the task success rate is less than the required reliability, increase the current number of replicas until the task success rate is greater than or equal to the required reliability. If the task success rate is greater than or equal to the required reliability, then the number of replicas when the task success rate is greater than or equal to the required reliability is obtained as the minimum number of replicas for the current task. The sorted results are then iterated to obtain the minimum number of replicas required for each task to execute.

4. The microservice fault-tolerant scheduling method as described in claim 1, characterized in that, The method for selecting the earliest task execution method from three options—static microservice instances already deployed in the static resource pool, dynamic microservice instances already created in the elastic resource pool, and newly created dynamic instances in the elastic resource pool—based on the minimum number of replicas for each task, includes: Based on the minimum number of replicas for each task, search for all instances of the same type from the static microservice instances deployed in the static resource pool and the dynamic instances created in the elastic resource pool as a candidate instance list; Calculate the estimated start time for assigning tasks to candidate instances from the candidate instance list; Get the new estimated start time for creating a new dynamic instance in the Elastic Resource Pool; Compare the estimated start time of all candidate instances in the candidate instance list with the new estimated start time for creating a new instance, and find the minimum value as the minimum start time; Obtain the instance type corresponding to the minimum start time. The instance type includes static microservice instance deployment, created dynamic microservice instance, and newly created dynamic instance. Select the instance type corresponding to the minimum start time as the target method.

5. The microservice fault-tolerant scheduling method as described in claim 1, characterized in that, The step of implementing microservice scheduling for microservice requests according to the target method includes: If the target is a deployed static microservice instance or a dynamically created instance in an elastic resource pool, then the replica will be scheduled to the deployed static microservice instance or the dynamically created instance in an elastic resource pool. If the target method is to create a new dynamic instance in the elastic resource pool, then a new microservice instance is created, and the replica is scheduled to the newly created microservice instance.

6. A microservice fault-tolerant scheduling system, characterized in that, include: The first acquisition module is used to acquire historical microservice call data of the same application and obtain the call graph of a single request based on the historical microservice call data. The first configuration module is used to configure the virtual super root that connects to the entry point of the call graph for all requests; The Bayesian probability tree creation module is used to generate an application-level call graph based on the virtual super root and historical microservice call data, and to model the application-level call graph as a Bayesian probability tree. The second setting module is used to set up a resource pool that includes a static resource pool and an elastic resource pool. The resource pool consists of a group of bare-metal nodes, each of which has a corresponding resource capacity. The static resource pool is used to deploy static microservice instances, and the elastic resource pool is used to create dynamic microservice instances on demand. The static deployment module is used to deploy static microservice instances in the static resource pool based on the historical microservice call data and cold start time. The second acquisition module is used to acquire the actual microservice requests; The pruning module is used to prune the Bayesian probability tree according to the microservice request to obtain the actual call graph; The minimum number of replicas module is used to calculate the minimum number of replicas required to execute a task based on the reliability of each task in the actual call graph. The selection module is used to select the earliest method to execute the task from three options: static microservice instances deployed in the static resource pool, dynamic microservice instances created in the elastic resource pool, and newly created dynamic instances in the elastic resource pool, based on the minimum number of replicas for each task. The scheduling module is used to schedule microservice requests according to the target method. The Bayesian probability tree creation module includes: The data acquisition unit obtains vertex attributes and edge attributes based on the virtual super root and historical microservice call data. The vertex attributes include microservice type, microservice location, processing time, and partition attribute. The edge attributes include upstream or downstream vertex, edge type, probability of edge existence in a specific request, preparation time before call, and data transmission volume. Create a unit to generate an application-level call graph based on the vertex and edge attributes, and model the application-level call graph as a Bayesian probability tree.

7. A terminal device, comprising a memory and a processor, characterized in that, The memory stores a computer program that can run on a processor, and when the processor loads and executes the computer program, it employs the method described in any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is loaded and executed by the processor, it employs the method described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and system for deploying micro-services in batches and storage medium

    CN121879786A

  • Real-time multi-modal audio anomaly detection system

    US12634322B1