Micro-service response delay prediction method, device, equipment, medium and product
By building a microservice call graph and using GNN to predict latency, the problems of resource management complexity and service quality violations in the microservice architecture are solved, and more efficient resource allocation and lower latency are achieved, reducing operational costs.
Patent Information
- Application Number
- CN202510180413.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-27
AI Technical Summary
Microservice architectures face resource management complexity and service quality violations in production environments, resulting in over-configuration of resources and tail delays, resulting in resource waste and performance degradation.
A microservice response delay prediction method is adopted to build a microservice call graph by obtaining the user's API call request, and a graph neural network (GNN) is used to calculate the prediction delay of each node, accumulate the node delay of the same microservice, and obtain the microservice response delay.
It significantly improves the accuracy of microservice response latency prediction, helps optimize resource allocation, reduce tail latency, improve resource utilization, and reduce operational costs.
Smart Images

Figure CN120045354A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of cloud service response, and particularly to a microservice response delay prediction method, device, equipment, medium and product. Background Art
[0002] In recent years, the design scale of software has become larger and larger, and business requirements have become more and more complex. Higher requirements are put forward for the performance, high throughput, high stability, high scalability and other characteristics of the system. In software development, the traditional monolithic architecture with strong coupling, cumbersome deployment and fault repair cannot meet the further expanding business needs, so the microservice architecture emerges as the times require. The microservice architecture is an important software development framework in the cloud native field. It constructs an application program as a collection of small, independent and self - contained services that can be independently deployed and managed. Compared with the monolithic architecture, the microservice framework has the advantage of flexible resource expansion. In an application program composed of multiple microservices, multiple microservices form a Microservice Call Graph. The user client can send an API call request to the application program through the API, and then call the microservice with a dedicated function to meet specific user requests.
[0003] However, although the microservice architecture supports agile development and resource flexibility, its modular design also brings complexity to resource management. Each microservice component needs to independently request resources to ensure service performance, resulting in complex resource management. The dependency relationships and complex topologies of microservices exacerbate the queuing effect, introducing service quality violation problems that are difficult to identify and correct in a timely manner. In the production environment, to ensure end - to - end service quality, the microservice architecture faces huge challenges because a single request may require hundreds of microservices to process. This leads to over - allocation of production cluster resources. For example, the CPU utilization rate may be as low as 20%, causing resource waste. The over - allocation of cloud resources is costly, and it is estimated that about $6.6 billion is wasted every year. Therefore, it is crucial to design a more efficient resource allocation framework to minimize tail latency and improve resource utilization. Summary of the Invention
[0004] The purpose of the present application is to provide a microservice response delay prediction method, device, equipment, medium and product, which can significantly improve the accuracy of delay prediction.
[0005] To achieve the above purpose, the present application provides the following solutions:
[0006] In the first aspect, the present application provides a microservice response delay prediction method, including:
[0007] Obtain an API call request of a user; the API call request includes multiple API requests.
[0008] For each API request, the API-based directed graph processing device obtains a microservice call graph corresponding to the API request; the microservice call graph is composed of a number of microservice call information with time dimension information; each microservice call information is represented by a node in the microservice call graph; each node is connected by a directed edge.
[0009] For the microservice call graph of each API request, the delay prediction device based on GNN determines the predicted delay of each node in the microservice call graph, and accumulates the node delays belonging to the same microservice to obtain the response delay of each microservice in the microservice call graph.
[0010] Predict the microservice response delays of all microservices according to the response delays of the microservices in each API request.
[0011] Optionally, for each type of API request, the API-based directed graph processing device obtains a microservice call graph corresponding to the API request, specifically including:
[0012] According to the API request, use the distributed tracing tool Jaeger in the API-based directed graph processing device to determine the actual call information in the program execution trace.
[0013] According to the actual call information in the program execution trace, use the microservice call graph construction algorithm to establish a microservice call graph corresponding to the microservice call timing information of the API request.
[0014] Optionally, according to the actual call information in the program execution trace, use the microservice call graph construction algorithm to establish a microservice call graph corresponding to the microservice call timing information of the API request, specifically including:
[0015] Traverse the span of each actual call information in the program execution trace; the span is the time spent by the actual call information in a microservice instance.
[0016] According to the out-degree number n of each span, and divide each span into 2*n + 2 nodes; where +2 represents the start and end nodes of the span.
[0017] Use directed edges to connect the nodes of the same span after division to obtain a topological sorting graph.
[0018] Traverse each span and check the start time and end time of each sub-call in the span.
[0019] If the start time of the sub-call matches the timestamp of a certain node in the span, add a directed edge from this node to the target node of the sub-call in the microservice call graph.
[0020] If the end time of the sub - call matches the timestamp of a certain node in the span, then a directed edge is added in the microservice call graph from the target node of the sub - call to this node.
[0021] Optionally, the GNN - based latency prediction device determines the predicted latency of each node in the microservice call graph, specifically including:
[0022] Based on the microservice call graph of microservice call timing information, an encoder is used to generate the feature vector of each node in the microservice call graph; the feature vector includes node name, resource quantity, resource utilization rate, node depth, concurrency, and the timing feature of the past response latency.
[0023] The feature vector of each node in the microservice call graph is combined with the microservice call graph, and the message aggregation ability of the graph neural network GNN is used to calculate the final feature representation of each node.
[0024] The final feature representation of each node is input into a fully - connected neural network to obtain the response latency of each microservice.
[0025] The response latencies of the nodes belonging to the same span are summed to obtain the response latency of each microservice.
[0026] Optionally, the update of the core formula of the graph neural network GNN is as follows:
[0027]
[0028] Among them, H (l) is the node feature matrix of the l - th layer, and its dimension is N×F (l) , where F (l) is the feature dimension of the l - th layer, W(v) is the weight matrix of the l - th layer, and its dimension is F (l) ×F (l+1) , σ is a non - linear activation function, and H (l+1) is the node feature matrix of the (l + 1) - th layer.
[0029] Optionally, the timing feature of the past response latency is obtained by inputting the past node latency information of the same type of API request into a long - short - term memory network.
[0030] In a second aspect, the present application provides a microservice response latency prediction device, including:
[0031] A call request acquisition module, configured to acquire the API call request of the user; the API call request includes multiple API requests.
[0032] The microservice call graph construction module is used to obtain the microservice call sequence information corresponding to the API request for each API request based on the directed graph processing device of the API; the microservice call graph is composed of several microservice call information with time dimension information; each microservice call information is represented by a node in the microservice call graph; each node is connected by a directed edge.
[0033] The response delay calculation module is used to determine the predicted delay of each node in the microservice call graph based on the delay prediction device of GNN for the microservice call graph of each API request, and accumulate the node delays belonging to the same microservice to obtain the response delay of each microservice in the microservice call graph.
[0034] The microservice response prediction module is used to predict the microservice response delays of all microservices according to the response delays of each microservice in each API request.
[0035] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement a microservice response delay prediction method according to any one of the above.
[0036] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements a microservice response delay prediction method according to any one of the above.
[0037] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements a microservice response delay prediction method according to any one of the above.
[0038] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application:
[0039] The present application provides a method, apparatus, device, medium and product for predicting microservice response latency. In this method, the first step is to obtain the API call requests of users, which provides the basic data for analysis. A microservice call graph for generating microservice call timing information is generated by using a directed graph processing apparatus based on APIs, which can clearly show the call relationship and time sequence between microservices. Each node in the microservice call graph represents a microservice call information and is connected by directed edges, which helps to understand the dependency relationship and execution order between services. A latency prediction apparatus based on a graph neural network (GNN) calculates the predicted latency of each node, and by accumulating the node latencies of the same microservice, the response latency of each microservice is obtained. Finally, based on the response latencies of the microservices in each API request, the response latencies of all microservices are predicted, thereby significantly improving the accuracy of microservice response latency prediction. Description of the Drawings
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0041] Figure 1 It is an application environment diagram of a method for predicting microservice response latency in an embodiment of the present application.
[0042] Figure 2 It is a schematic flowchart of a method for predicting microservice response latency provided in an embodiment of the present application.
[0043] Figure 3 It is a microservice call graph provided in an embodiment of the present application.
[0044] Figure 4 It is an actual microservice call graph of API1 provided in an embodiment of the present application.
[0045] Figure 5 It is a schematic diagram of the program execution trajectory of API1 provided in an embodiment of the present application.
[0046] Figure 6 It is a schematic diagram of the program execution trajectory of API1 after being segmented provided in an embodiment of the present application.
[0047] Figure 7 It is a microservice call graph corresponding to API1 provided in an embodiment of the present application.
[0048] Figure 8 It is a schematic diagram of a long short-term memory model provided in an embodiment of the present application.
[0049] Figure 9 A schematic diagram of an overall process provided in an embodiment of the present application.
[0050] Figure 10 A schematic design diagram of a microservice response delay prediction method provided in an embodiment of the present application.
[0051] Figure 11 A schematic diagram of functional modules of a microservice response delay prediction device provided in an embodiment of the present application.
[0052] Figure 12 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. Detailed implementation manners
[0053] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0054] In recent years, the design scale of software has become larger and larger, and business requirements have become more and more complex. Higher requirements are put forward for the performance, high throughput, high stability, high scalability, etc. of the system. In software development, the traditional monolithic architecture with strong coupling, cumbersome deployment and fault repair cannot meet the further expanding business needs, so the microservice architecture came into being. The microservice architecture is an important software development framework in the cloud native field. It constructs an application program as a collection of small, independent, self - contained services that can be independently deployed and managed.
[0055] The microservices architecture style is a way to develop a single application using a set of small services. Each service runs in its own process and uses lightweight mechanisms for communication, such as RPC or RESTful interfaces. These services are built based on business capabilities and can be independently deployed through an automated deployment mechanism. They can be implemented using different programming languages and different data storage technologies, with minimal centralized management. Now, a large number of companies in the world, such as Netflix, Twitter, Facebook, and Airbnb, etc., use this architecture to build cloud-based applications. Compared with the monolithic architecture, the microservices framework has the advantage of flexible resource scaling. In an application composed of multiple microservices, the multiple microservices form a Microservice CallGraph. The user client can send API call requests to the application through the API, and then call the microservices with dedicated functions to meet specific user requests. When the load of the application continues to increase, the service provider can locate and scale the individual microservices with heavier loads, rather than scaling the entire application. Therefore, the scalability of the microservices architecture has been greatly improved compared to the monolithic architecture. The independence of microservices enables development teams to develop and deploy in parallel, reducing the development cycle. Secondly, the loose coupling characteristic of microservices improves the maintainability of the system, and any change to one service will not affect the entire system. Moreover, microservices can choose the most suitable technology stack and database, making the technology selection more flexible and able to quickly adapt to changes in business requirements. In addition, microservices can enhance the fault tolerance of the system. The failure of a single service will not cause the entire application to crash, thus improving the availability of the system. Finally, the application of automated deployment and monitoring tools makes the operation and maintenance more efficient, able to respond to performance bottlenecks and failures in real time, further ensuring the stability and high throughput of the system. Through these advantages, the microservices architecture significantly improves the enterprise's ability to handle complex business requirements.
[0056] However, while this modular design of microservices architecture allows for agile development and resource flexibility, where each component can be developed and scaled independently, it complicates resource management. Application owners must request resources for each component (container or pod) to ensure that the application can serve traffic from users. In terms of resource management, since the complex topology of microservice dependencies exacerbates the queuing effect and introduces cascading quality of service (QoS) violations that are difficult to identify and correct in a timely manner. The microservices architecture poses a great challenge in providing end-to-end QoS guarantees in a production environment because a single user request needs to be processed by hundreds of fine-grained microservices. Today's production clusters usually over-provision resources to provide such guarantees, which easily leads to low overall resource utilization. For example, the CPU utilization in Alibaba's microservices cluster is as low as 20%. Over-provisioning is costly, and due to over-provisioning, approximately $6.6 billion is wasted in the cloud. Even a slight improvement in cloud resource allocation can save millions of dollars. An efficient resource allocation framework strictly minimizes the relevant resources in terms of tail latency. Therefore, it becomes crucial to design a more efficient resource scaling solution to meet QoS guarantees. In addition, an efficient resource scaling solution should proactively allocate resources for each microservice on an application according to the changes in the front-end workload. Such a measure is necessary to avoid the cascading effect, which severely degrades the performance of microservices during a traffic surge. The root cause of the cascading effect is the inevitable deployment overhead and the nature of the microservices architecture, i.e., requests are processed together with a series of microservices. When the front-end traffic surges, the deep microservices are not directly affected until the previous microservices handle the increased workload after deploying additional resources. Proactively allocating resources for the microservice chain is the key to avoiding the cascading effect. Otherwise, delays will accumulate when congestion is alleviated one by one from the first to the last microservice in the chain.
[0057] Therefore, the resource allocation framework for microservices should aim at two main goals: optimizing CPU resources according to the latency SLO (i.e., end-to-end tail latency), and proactively deploying CPU resources for each microservice according to the expected impact of the front-end workload. This is usually achieved with the help of resource estimation techniques to predict future utilization so that application owners can prepare and allocate such resources in advance to maintain QoS. The ability to achieve efficient proactive scaling highly depends on the accurate prediction of end-to-end latency, but this remains quite challenging. First, the component microservices in the same application can form complex interactions with each other. The response of one microservice will severely affect its downstream microservices, and thus affect the end-to-end latency of the entire trace. Therefore, failure to capture the complex dependencies between microservices will lead to inaccurate predictions. Second, tail latency is unstable due to its inherent randomness because it takes into account the worst-case scenario.
[0058] Regarding microservice response latency prediction, in previous work, people have dedicated themselves to capturing the complex dependencies between microservices and their operations to predict end-to-end latency. Specifically, these works mainly use deep neural networks to model dependencies based on the entire MCG. However, predictors based on MCG have several drawbacks: First, MCG does not have API awareness and cannot handle different APIs differently, resulting in a large amount of irrelevant information being included during prediction, which may lead to unsatisfactory results; Second, MCG is a static graph and cannot capture the temporal relationships between each call. That is, when a microservice calls multiple downstream microservices, MCG cannot show which microservice is called first, and it cannot distinguish whether these two microservices are called in parallel or sequentially. Finally, MCG cannot recognize runtime dynamics. Given the same API call, runtime dynamics may vary due to different input parameters. However, MCG cannot incorporate this information because it uses a single static graph to couple all runtime behaviors. All in all, designing a solution that can take into account the dependencies between microservices and can be API-aware and effectively predict microservice response latency is of great significance and has wide applications, which can significantly improve the performance of microservice systems.
[0059] The application scenarios of the microservice response latency prediction system include:
[0060] 1. Automated resource scheduling: In cloud or on-premises clusters, applications can automatically adjust resource allocation by predicting the response latency of microservices. For high-traffic applications, the system can proactively scale resources before a traffic surge to ensure the stable operation of the system during peak periods, thus avoiding service degradation or interruption caused by insufficient resources.
[0061] 2. Service Quality Assurance (QoS) optimization: In industries such as e-commerce and finance where high response speed is required, the microservice response latency prediction system can help ensure that the end-to-end latency of each user request meets the predefined Service Level Agreement (SLA). By estimating potential latencies, the system can take timely measures to avoid excessive latencies and enhance the user experience.
[0062] 3. Microservice performance tuning: In complex microservice architectures, the system can identify performance bottlenecks through response latency prediction and help development and operations teams optimize the interactions between microservices. For services with high latency, the system can recommend corresponding optimization measures, such as refactoring the service call path or processing multiple downstream services in parallel.
[0063] 4. Real-time Monitoring and Fault Warning: In a production environment, the microservice response latency prediction system can serve as a monitoring tool to help enterprises track the performance metrics of the system in real time. When the system predicts that a certain service is about to experience a latency anomaly, it can issue an early warning to prevent potential faults from further affecting other services and avoid a global performance collapse caused by the cascading effect.
[0064] 5. Cost Control and Resource Optimization: Enterprises can use the microservice response latency prediction system to conduct refined management of resource usage and avoid resource waste caused by over-allocation. When the system identifies that certain services will have low latency in the future, it can temporarily reduce the resource allocation for these services to lower the operating costs.
[0065] Currently, only a few works can take into account the dependencies between microservices and the importance of API awareness. Designing a more excellent microservice response prediction scheme is an inevitable trend.
[0066] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0067] A microservice response latency prediction method provided by an embodiment of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set separately, integrated on the server 104, placed on the cloud or other servers. The terminal 102 can send the API call request of the user to be processed to the server 104. After receiving the API call request of the user to be processed, for the API call request of the user to be processed, the server 104, for each API request, based on the directed graph processing device of the API, obtains the microservice call graph of the microservice call timing information corresponding to the API request; the microservice call graph is composed of several microservice call information with time dimension information; each microservice call information is represented by a node in the microservice call graph; each node is connected by a directed edge; for the microservice call graph of each API request, based on the GNN-based latency prediction device, determines the predicted latency of each node in the microservice call graph, and accumulates the node latencies belonging to the same microservice to obtain the response latency of each microservice in the microservice call graph; according to the response latencies of each microservice in each API request, predicts the microservice response latency of all microservices. The server 104 can feedback the microservice response latency of all microservices obtained to the terminal 102. In addition, in some embodiments, a microservice response latency prediction method can also be implemented separately by the server 104 or the terminal 102. For example, the terminal 102 can directly process the API call request of the user to be processed, or the server 104 can obtain the API call request of the user to be processed from the data storage system and process the API call request of the user to be processed.
[0068] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0069] In an exemplary embodiment, as Figure 2 shown, a microservice response latency prediction method is provided. This method is executed by a computer device, and can be specifically executed by a computer device such as a terminal or a server alone, or jointly executed by a terminal and a server. In the embodiments of the present application, taking this method applied to Figure 1 the server 104 in
[0070] Step 201, obtain the API call request of the user; the API call request includes multiple API requests.
[0071] Step 202, for each API request, based on the directed graph processing device of the API, obtain the microservice call graph of the microservice call timing information corresponding to the API request; the microservice call graph is composed of several microservice call information with time dimension information; each microservice call information is represented by a node in the microservice call graph; each node is connected by a directed edge.
[0072] Step 203, for the microservice call graph of each API request, based on the delay prediction device of GNN, determine the predicted delay of each node in the microservice call graph, and accumulate the node delays belonging to the same microservice to obtain the response delay of each microservice in the microservice call graph.
[0073] Step 204, predict the microservice response delay of all microservices according to the response delays of the microservices in each API request.
[0074] Among them, in some embodiments, when executing step 201, it can be specifically as follows:
[0075] Obtain the API call request of the user; the API call request includes multiple API requests
[0076] It is necessary to obtain the API call request of the user; in the API call request, there will usually be various types of API request information. For example: authentication information (such as API key or OAuth token), content type, cache control.
[0077] Among them, in some embodiments, when executing step 202, it can be specifically as follows:
[0078] According to the API request, use the distributed tracing tool Jaeger in the directed graph processing device of the API to determine the actual call information in the program execution trace.
[0079] According to the actual call information in the program execution trace, use the microservice call graph construction algorithm to establish the microservice call graph of the microservice call timing information corresponding to the API request.
[0080] Specifically, the API-based directed graph processing device is deployed in a computer. For each type of API request, the API-based directed graph processing device combines the request trace information collected by Jaeger with the local microservice call graph formed according to the microservices traversed by the request, and through analysis, obtains a more detailed microservice call graph that can reflect the microservice call timing information. In this embodiment, it is called Program Evaluation and Review Technique Graph (PERT Graph). As Figure 3 shown in the figure, this graph can reflect the microservice call information with time dimension information.
[0081] Specifically, in an application of a microservice architecture, its complete microservice call graph MCG is as Figure 3 shown. Among them, A, B, C, D, E, F, G, and H are all microservice nodes in the microservice architecture, and there may be dependencies between them. In an API request call, not all microservices will be accessed. If irrelevant microservices are introduced during analysis, irrelevant data will be introduced, resulting in a decrease in prediction accuracy. Here, in this embodiment, according to different APIs, a corresponding call graph that only contains the actually accessed microservices is generated for each API. In this embodiment, it is called the actual call graph.
[0082] The actual call graph is as Figure 4 shown. Among them, API1 is an API request of the microservice system. The microservices actually accessed by this request are only A, B, C, D, and E. Therefore, the generated actual call graph is as shown in the figure, which does not contain the unaccessed microservices. In this embodiment, the actual call graph of each API can be generated according to the prior knowledge of the microservice application. If this knowledge is lacking, test API requests can also be sent in the test environment, and the actual call information in the program execution trace (trace) obtained by the distributed link tracking tool Jaeger can be used. So far, this embodiment has successfully obtained the actual call graph.
[0083] Next, for each type of API call, in this embodiment, the trace (program execution trace) can be obtained through the distributed link tracking tool Jaeger deployed in advance in the server or container. The trace of API1 is as Figure 5 shown.
[0084] In the trace graph of API1, the horizontal axis represents time, and each one represents the time spent by a corresponding single request in a microservice instance. In this embodiment, it is called span. Figure 5Shows the spans experienced in five microservices after API1 arrives. A span represents a logical unit of work with microservice name, operation name, operation start time, duration, and parent span, and all this information can be obtained through trace. Subsequently, in this embodiment, based on the information in the trace, using the PERT Graph construction algorithm, a PERT Graph that can more accurately represent the microservice call situation is generated.
[0085] Among them, according to the actual call information in the program execution trace, using the microservice call graph construction algorithm to establish the microservice call graph of the microservice call timing information corresponding to the API request can be as follows:
[0086] Traverse the spans of each actual call information in the program execution trace; the span is the time spent by the actual call information in a microservice instance.
[0087] According to the out-degree number n of each span, and divide each span into 2*n + 2 nodes; where +2 represents the start and end nodes of the span.
[0088] Use directed edges to connect the nodes of the same span after division to obtain a topological sorting graph.
[0089] Traverse each span and check the start time and end time of each sub-call in the span.
[0090] If the start time of the sub-call matches the timestamp of a certain node in the span, add a directed edge from this node to the target node of the sub-call in the microservice call graph.
[0091] If the end time of the sub-call matches the timestamp of a certain node in the span, add a directed edge from the target node of the sub-call to this node in the microservice call graph.
[0092] Specifically, first, this embodiment traverses each span. According to the out-degree number n of each span, it is divided into 2*n + 2 nodes. Here, +2 represents the start and end nodes of the span, as Figure 7As shown. After that, in this embodiment, the nodes divided from the same span are connected by directed edges. Since there is a chronological relationship in the access order of these nodes, this embodiment can obtain a topological sorting graph. At this time, this embodiment has divided each span into several nodes connected by directed edges, obtaining an initial PERT Graph. Then, this embodiment will traverse each span again to check the start time and end time of each sub-call in the span. If the start time of the sub-call matches the timestamp of a certain node in the span, then a directed edge will be added in the PERT Graph from this node to the target node of the sub-call. If the end time of the sub-call matches the timestamp of a certain node in the span, then a directed edge will be added in the PERT Graph from the target node of the sub-call to this node, as Figure 6 shown. Thus, this embodiment obtains a PERT Graph that can reflect the call timing information. For each API, this embodiment will obtain a corresponding PERT Graph (microservice call graph).
[0093] In some embodiments, when performing step 203, it can be specifically as follows:
[0094] Based on the microservice call graph of microservice call timing information, use an encoder to generate feature vectors for each node in the microservice call graph; the feature vectors include node names, resource amounts, resource utilization rates, node depths, concurrency amounts, and chronological features of past response delays.
[0095] Combine the feature vectors of each node in the microservice call graph with the microservice call graph, and use the message aggregation ability of the graph neural network GNN to calculate the final feature representation of each node.
[0096] Input the final feature representation of each node into a fully connected neural network to obtain the response delays of each microservice.
[0097] Sum up the response delays of the nodes belonging to the same span to obtain the response delays of each microservice.
[0098] Specifically, the delay prediction device based on GNN is deployed in a computer, which receives the PERT Graph and the feature vectors of each node as inputs, uses the graph neural network GNN for calculation, and finally outputs the predicted delay of each node. Accumulate the node delays belonging to the same microservice together, and finally obtain the response delay of each microservice.
[0099] Graph Neural Networks (GNN for short) are a type of neural network specialized in processing graph-structured data. They can operate directly on graphs and capture complex relationships and patterns between nodes. The core advantage of GNNs lies in their ability to effectively process structured data, which makes them perform well when dealing with social networks, molecular structures, transportation networks, and various other graph-form data. The basic idea of GNNs is to represent the nodes in the graph structure as vectors and pass and update these vectors through neural network layers. In this process, the representation of each node not only contains its own feature information but also incorporates the information of its neighbor nodes. This information transfer is usually achieved through a message-passing mechanism, where each node aggregates information from its neighbors to update its feature representation.
[0100] The update of the core formula of the graph neural network is as follows:
[0101]
[0102] where H (l) is the node feature matrix of the l-th layer, with a dimension of N×F (l) , where F (l) is the feature dimension of the l-th layer. W(l) is the weight matrix of the l-th layer, with a dimension of F (l) ×F (l+1) . σ is a non-linear activation function, such as ReLU. H (l+1) is the node feature matrix of the (l + 1)-th layer.
[0103] Formula (1) shows that the new feature representation of each node is obtained by multiplying the feature representation of the current layer with the adjacency matrix and then applying the weight matrix and the non-linear activation function. In practice, the GNN model may contain multiple such graph convolutional layers, and the node feature representation will be updated in each layer. In addition, techniques such as skip connections and batch normalization may also be included to improve the performance and generalization ability of the model.
[0104] In the PERT Graph obtained in this embodiment, this embodiment needs to use an encoder to generate the feature vector (embedding) for each node. The feature vector includes the node name, resource quantity, resource utilization rate, node depth, concurrency, and temporal features of the past response latency. Among them, the resource utilization rate including cpu and memory utilization rate and the concurrency are obtained through Prometheus deployed on the server or container. Prometheus is a monitoring platform that collects metrics from monitored targets by scraping the metric HTTP endpoints on the targets.
[0105] To be able to effectively predict future delays, this embodiment also needs to utilize past delays to capture the trend of the delay time series. The temporal characteristics of past response delays are the feature representations obtained by inputting the past node delay information of the same type of API request into a Long Short-Term Memory (LSTM) network. LSTM (Long Short-Term Memory) is a special Recurrent Neural Network (RNN) architecture that excels in processing sequential data, especially long sequential data such as time series analysis and language modeling. The core of the LSTM lies in its unique gating mechanism, which enables it to effectively retain and transmit information between different time steps of the sequence, as Figure 8 shown. During the calculation process of the LSTM, the forget gate f t determines how much information to discard from the cell state, and the input gate i t and the candidate cell state together determine how much new information to update into the cell state. The output gate o t determines how much information to output for the hidden state h t , and the cell state c t and the hidden state h t are updated to new values for the current time step. These gating mechanisms enable the LSTM to effectively capture long-term dependence information while avoiding the problems of gradient vanishing or explosion.
[0106] After obtaining the features of each node to form a feature matrix, this embodiment can combine it with the PERT Graph and use the message aggregation ability of the GNN to calculate the final feature representation of each node. Finally, it is connected to a fully connected neural network to finally obtain the predicted delay of each node. Note that the delay obtained at this time is Figure 7 the delay of the nodes in, and to obtain the delay of each microservice, this embodiment needs to sum the delays of the nodes belonging to the same span, and finally obtain the delays of all microservices. The final process is as Figure 9 shown.
[0107] It can be seen that as Figure 10As shown in the figure, the present application designs a microservice latency prediction device with API awareness based on GNN in the context of a microservice architecture. This device is mainly deployed on a server or a container and mainly includes a directed graph processing device based on APIs and a latency prediction device based on GNN. The microservice architecture application is deployed on a server or a container, and Jaeger and Prometheus are deployed simultaneously to obtain monitoring data. The directed graph processing device based on APIs is used to process the data obtained through Jaeger and Prometheus, and after obtaining the request trace (tracking) and the call graph MCG, it generates a PERT Graph for each API. Jaeger is an end-to-end distributed tracing tool. Prometheus is an open-source system and service monitoring tool. The latency prediction device based on GNN is used to obtain the final predicted latency through node aggregation and a fully connected network using the PERT Graph and the node feature matrix obtained from Prometheus and processed by LSTM.
[0108] Based on the same inventive concept, an embodiment of the present application also provides a microservice response latency prediction device for implementing the microservice response latency prediction method involved above. The implementation solution provided by this device to solve the problem is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more of the following device embodiments can refer to the limitations on a microservice response latency prediction method in the above text and will not be elaborated here.
[0109] In an exemplary embodiment, as Figure 11 shown, a microservice response latency prediction device is provided, including:
[0110] A call request acquisition module 1101, configured to acquire an API call request of a user; the API call request includes multiple API requests.
[0111] A microservice call graph construction module 1102, configured to, for each API request, obtain a microservice call graph corresponding to the API request based on the directed graph processing device based on APIs; the microservice call graph is composed of several microservice call information with time dimension information; each microservice call information is represented by a node in the microservice call graph; each node is connected by a directed edge.
[0112] A response latency calculation module 1103, configured to, for the microservice call graph of each API request, determine the predicted latency of each node in the microservice call graph based on the latency prediction device based on GNN, and accumulate the node latencies belonging to the same microservice to obtain the response latency of each microservice in the microservice call graph.
[0113] The microservice response prediction module 1104 is configured to predict the microservice response delays of all microservices based on the response delays of each microservice in each API request.
[0114] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal, and its internal structure diagram may be as Figure 12 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store microservice response delay data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a microservice response delay prediction method.
[0115] Those skilled in the art can understand that Figure 12 the structure shown in
[0116] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0117] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0118] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0119] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0120] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0121] The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0122] In summary, this application has the following technical effects:
[0123] 1. The present application designs a directed graph processing device based on APIs. By generating customized PERT Graphs for each API request, the accuracy of the prediction model is significantly improved. Irrelevant microservices are excluded, reducing the interference of noisy data, thereby enhancing the efficiency and precision of the prediction model. Distributed tracing tools such as Jaeger are used to dynamically capture the call relationships between microservices, ensuring the real-time nature and accuracy of the graph.
[0124] 2. The present application designs a latency prediction device based on GNNs. Utilizing the powerful representation ability of GNNs, the present application can accurately capture the complex dependency relationships between microservice nodes. By considering the timing of API calls, the present application can more accurately simulate the interactions between microservices, thereby providing more refined latency predictions. An LSTM is integrated to analyze historical latency data, enhancing the model's ability to capture time series trends.
[0125] 3. By accurately predicting microservice latency, the present application creates good conditions for downstream tasks such as resource allocation and configuration optimization. By predicting potential latency, the system can take timely measures to avoid service quality (QoS) violations and enhance the user experience. Through refined resource management and cost control, the present application helps reduce resource waste and achieve more efficient operating costs.
[0126] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0127] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A microservice response delay prediction method, characterized in that: The microservice response delay prediction method comprises: Obtaining a user's API call request; the API call request includes multiple API requests; For each API request, a microservice call graph of microservice call timing information corresponding to the API request is obtained based on the API directed graph processing device; the microservice call graph is composed of a plurality of microservice call information with time dimension information; each microservice call information is represented by a node in the microservice call graph; each node is connected by a directed edge; For each microservice call graph of an API request, a GNN-based delay prediction device determines the predicted delay of each node in the microservice call graph, and accumulates the node delays belonging to the same microservice to obtain the response delay of each microservice in the microservice call graph; Predict the microservice response delay of all microservices based on the response delay of each microservice in each API request.
2. A microservice response delay prediction method according to claim 1, characterized in that: For each API request, the directed graph processing device based on the API obtains a microservice call graph of the microservice call timing information corresponding to the API request, specifically including: According to the API request, using Jaeger, a distributed link tracking tool in the directed graph processing device of the API, to determine the actual call information in the program execution trace; According to the actual call information in the program execution trajectory, a microservice call graph construction algorithm is used to establish a microservice call graph of the microservice call timing information corresponding to the API request.
3. A microservice response delay prediction method according to claim 2, characterized in that: According to the actual call information in the program execution trajectory, a microservice call graph construction algorithm is used to establish a microservice call graph of the microservice call timing information corresponding to the API request, specifically including: Traverse the span of each actual call information in the program execution trace; the span is the time the actual call information spends in a microservice instance; According to the number of out-degrees n of each span, each span is divided into 2*n+2 nodes; where +2 represents the start and end nodes of the span; Use directed edges to connect the nodes with the same span after division to obtain a topological sorting graph; Iterate over each span and check the start and end time of each subcall in the span; If the start time of the subcall matches the timestamp of a node in the span, a directed edge from that node to the subcall target node is added to the microservice call graph; If the end time of the subcall matches the timestamp of a node in the span, a directed edge from the subcall target node to that node is added to the microservice call graph.
4. A microservice response delay prediction method according to claim 1, characterized in that: The GNN-based delay prediction device determines the predicted delay of each node in the microservice call graph, including: Based on the microservice call graph of the microservice call timing information, an encoder is used to generate a feature vector of each node in the microservice call graph; the feature vector includes the timing characteristics of the node name, resource quantity, resource utilization, node depth, concurrency, and past response delay; Combine the feature vector of each node in the microservice call graph with the microservice call graph, and use the message aggregation capability of the graph neural network GNN to calculate the final feature representation of each node; The final feature representation of each node is input into the fully connected neural network to obtain the response delay of each microservice; The response delays of nodes belonging to the same span are summed to obtain the response delay of each microservice.
5. A microservice response delay prediction method according to claim 4, characterized in that: The core formula of the graph neural network GNN is updated as follows: Among them, H (l) is the node feature matrix of the lth layer, whose dimension is N×F (l) , where F (l) is the feature dimension of the lth layer, W(v) is the weight matrix of the lth layer, and its dimension is F (l) ×F (l+1) , σ is a nonlinear activation function, H (l+1) is the node feature matrix of the l+1th layer.
6. A microservice response delay prediction method according to claim 5, characterized in that: The temporal characteristics of the past response delay are obtained by inputting the past node delay information of the same API request into the long short-term memory network.
7. A microservice response delay prediction device, characterized in that: The microservice response delay prediction device comprises: A call request acquisition module is used to acquire a user's API call request; the API call request includes multiple API requests; A microservice call graph construction module is used to obtain a microservice call graph of microservice call timing information corresponding to each API request based on an API directed graph processing device; the microservice call graph is composed of a plurality of microservice call information with time dimension information; each microservice call information is represented by a node in the microservice call graph; and each node is connected by a directed edge; The response delay calculation module is used to determine the predicted delay of each node in the microservice call graph based on the GNN delay prediction device for each API request, and accumulate the node delays belonging to the same microservice to obtain the response delay of each microservice in the microservice call graph; The microservice response prediction module is used to predict the microservice response delay of all microservices based on the response delay of each microservice in each API request.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a microservice response delay prediction method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, a microservice response delay prediction method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, a microservice response delay prediction method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Data distribution method and device for distributed graph calculation
CN120295802A