Proactive anomaly detection

A neural network-based anomaly detection system for microservices addresses the inefficiencies in current monitoring by analyzing request context data, effectively predicting and preventing performance anomalies, thus enhancing system reliability and reducing downtime.

JP7705210B2Active Publication Date: 2025-07-09INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023532550
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-30
Filing Date
2021-10-21
Publication Date
2025-07-09
Estimated Expiration
2041-10-21

AI Technical Summary

Technical Problem

Current anomaly detection systems for microservices applications fail to effectively consider spatial and temporal dependencies between services, leading to inefficient monitoring and high false detection rates, which can result in SLA violations and downtime.

Method used

A neural network-based anomaly detection system that collects and analyzes request context data, including inter-request and intra-request factors, to predict performance anomalies and generate proactive warnings, supporting hybrid cloud deployments and various container orchestrators.

Benefits of technology

The system provides accurate proactive anomaly detection, reducing downtime by predicting potential issues and enabling timely resource adjustments, while enhancing observability and reducing false positives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007705210000001
    Figure 0007705210000001
  • Figure 0007705210000002
    Figure 0007705210000002
  • Figure 0007705210000003
    Figure 0007705210000003
Patent Text Reader

Abstract

A computer-implemented method, a computer program product, and a computer system are provided. For example, embodiments of the present invention may collect trace data and a specification of a series of requests for normal operation of a microservice application in response to receiving a request. Embodiments of the present invention may generate request contextual features from the collected trace data and specification. Embodiments of the present invention may train a neural network model based on the generated contextual features and use the trained neural network model to predict abnormal operation of the microservice application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to proactive anomaly detection, and more particularly to proactive anomaly detection for microservice applications using request context data and neural networks.

Background Art

[0002] A microservices architecture arranges an application as a collection of loosely coupled services. Microservices are not layers within a monolithic application (e.g., a web controller or backend-for-frontend). Thus, a microservices architecture is suitable for a continuous delivery software development process. Even when making changes to only a small part of an application, only one or a few services need to be rebuilt and redeployed.

[0003] Generally, a microservices architecture can be adopted for cloud-native applications, serverless computing, and applications using lightweight container deployment. In a monolithic approach, an application that supports three functions (such as a framework, a database, a message broker, etc.) needs to scale in the whole even if only one of these functions has resource constraints. In microservices, since only the microservices that support functions with resource constraints need to be scaled out, it provides the advantage of resource and cost optimization.

[0004] Machine learning (ML) is a scientific study of algorithms and statistical models used by computer systems to perform specific tasks by relying on patterns and inference instead of using explicit instructions. Machine learning is considered a subset of artificial intelligence. Machine learning algorithms build mathematical models based on sample data called training data and make predictions or decisions without being explicitly programmed to perform the task. Machine learning algorithms are used in various applications where it is difficult or impossible to develop traditional algorithms for effectively performing tasks such as email filtering and computer vision.

[0005] In machine learning, a hyperparameter is a setting that is external to the model and whose value cannot be estimated from the data. Hyperparameters are used in the process of assisting the estimation of model parameters. Hyperparameters are set before the learning (e.g., training) process begins. In contrast, the values of other parameters are derived by training. Different algorithms for model training require different hyperparameters, but some simple algorithms such as least squares regression may not require hyperparameters. Given a set of hyperparameters, the training algorithm learns the parameter values from the data. For example, the Least Absolute Shrinkage and Selection Operator (LASSO) is an algorithm that adds a regularization hyperparameter to least squares regression and needs to be set before the parameter is estimated by the training algorithm. Similar machine learning models may require different hyperparameters (e.g., different constraints, weights, or learning rates) to generalize different data patterns.

[0006] Deep learning is a field of machine learning based on a set of algorithms that model high-level abstractions in data by using model architectures that have complex structures, etc., or that are otherwise often composed of multiple non-linear transformations. Deep learning is part of a broader family of machine learning techniques based on learning the representation of data. Observations (e.g., images) can be represented in many ways, such as a vector of intensity values per pixel, or in more abstract ways such as a set of edges, regions of a particular shape, etc. In some representations, it is easier to learn the task from examples (e.g., face recognition or emotion recognition). Deep learning algorithms often use a cascade of many layers of non-linear processing units for feature extraction and transformation. Each successive layer uses the output from the previous layer as input. The algorithms can be supervised or unsupervised, and the applications include pattern analysis (unsupervised) and classification (supervised). Deep learning models include artificial neural networks (ANNs) inspired by information processing in biological systems and distributed communication nodes. ANNs differ in various ways from the biological brain.

[0007] A neural network (NN) is a computing system inspired by biological neural networks. An NN is not just a simple algorithm but a framework where many different machine learning algorithms work together to process complex data inputs. Such systems generally "learn" to perform tasks by considering examples without being programmed with task-specific rules. For example, in image recognition, an NN analyzes image examples correctly labeled as "cat" or "not a cat" to identify cats in other images and learns to identify images containing cats by using the results. The NN learns without any prior knowledge about cats, such as that cats have fur, tails, whiskers, and pointed ears. Instead, the NN automatically generates discriminative features from the training materials. An NN is based on a collection of connected units or nodes called artificial neurons, which loosely model the neurons of the biological brain. Each connection can transmit a signal from one artificial neuron to another, like a synapse in the biological brain. An artificial neuron that receives a signal can process the signal and transmit it to another artificial neuron.

[0008] In a typical NN implementation, the signals in the connections between artificial neurons are real numbers, and the output of each artificial neuron is calculated by some non-linear function of the sum of its inputs. The connections between artificial neurons are called "edges". Artificial neurons and edges usually have weights that are adjusted as learning progresses. The weights increase or decrease the strength of the signals at the connections. An artificial neuron may also have a threshold such that a signal is transmitted only when the aggregated signal exceeds that threshold. Generally, artificial neurons are grouped into layers. Different layers can perform different types of transformations on their inputs. Signals are sent from the first layer (input layer) to the last layer (output layer), possibly crossing layers multiple times in between. SUMMARY OF THE INVENTION

[0009] According to one aspect of the present invention, a computer-implemented method is provided. The method includes, in response to receiving a request, collecting trace data and specifications of a series of requests for the normal operation of a microservice application, generating request context features from the collected trace data and specifications, training a neural network model based on the generated context features, and predicting an abnormal operation of the microservice application using the trained neural network model.

[0010] Hereinafter, preferred embodiments of the present invention will be described by way of example only with reference to the following drawings.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Modes for Carrying Out the Invention

[0012] In embodiments of the present invention, microservices architectures are often used for applications deployed in hybrid cloud environments because loosely coupled components provide better scalability, flexibility, maintainability, and acceleration of developer productivity. Such applications are composed of many services, which in turn are replicated into multiple instances and run at different geographical locations. Over time, performance degradation due to anomalies may occur. Thus, embodiments of the present invention recognize that detecting anomalies in microservices applications is an important task that can enable taking certain actions that may help reduce downtime and loss of productivity. Current systems struggle to monitor microservices applications and optimize performance because observability is limited. Further, embodiments of the present invention recognize that typical approaches to anomaly detection currently lack the ability to consider spatial and temporal dependencies between services, which can lead to more false detections. Accordingly, embodiments of the present invention provide a solution for improving current anomaly detection systems and provide an efficient tool for technical service support personnel managing complex microservices applications. For example, embodiments of the present invention use neural networks to detect anomalies based on contextual data. In this aspect, embodiments of the present invention use a neural network approach to predict performance anomalies (e.g., service level agreement (SLA) violations) in applications that jointly consider dependencies available in request contextual data, as described in more detail hereinbelow. Embodiments of the present invention then generate notifications and can then correct detected anomalies before the user becomes aware.

[0013] FIG. 1 is a functional block diagram showing a computing environment generally designated as computing environment 100 according to an embodiment of the present invention. FIG. 1 provides only an illustration of one implementation and does not imply any limitation regarding the environments in which different embodiments may be implemented. Those skilled in the art can make many modifications to the illustrated environment without departing from the scope of the invention as recited in the claims.

[0014] Computing environment 100 includes client computing devices 102 and server computers 108 that are all interconnected via network 106. Client computing devices 102 and server computers 108 can be stand-alone computer devices, management servers, web servers, mobile computing devices, or any other electronic device or computing system capable of receiving, transmitting, and processing data. In other embodiments, client computing devices 102 and server computers 108 can represent a server computing system that utilizes multiple computers as a server system, such as a cloud computing environment. In another embodiment, client computing devices 102 and server computers 108 can be a laptop computer, tablet computer, netbook computer, personal computer (PC), desktop computer, personal digital assistant (PDA), smartphone, or any programmable electronic device capable of communicating with various components within computing environment 100 and other computing devices (not shown). In another embodiment, client computing devices 102 and server computers 108 each represent a computing system that utilizes clustered computers and components (e.g., database server computers, application server computers, etc.) that function as a single pool of seamless resources when accessed within computing environment 100. In some embodiments, client computing devices 102 and server computers 108 are a single device. Client computing devices 102 and server computers 108 may include internal and external hardware components capable of executing machine-readable program instructions, as depicted and described in further detail with respect to FIG. 6.

[0015] In this embodiment, the client computing device 102 is a user device related to the user and includes an application 104. The application 104 communicates with the server computer 108 to access the anomaly detector 110 (e.g., using TCP / IP) or to receive service requests and database information. The application 104 identifies the contextual features related to the received requests, generates or trains a neural network model, and further communicates with the anomaly detector 110 to predict future requests to be processed within the microservices application, as will be discussed in more detail with respect to FIGS. 2-5.

[0016] The network 106 can be, for example, a telecommunications network, a local area network (LAN), a wide area network (WAN) such as the Internet, or a combination of these three, and can include wired, wireless, or fiber optic connections. The network 106 can include one or more wired or wireless or both networks capable of receiving and transmitting data, voice, or video signals, or combinations thereof, including multimedia signals including voice, data, and video information. Generally, the network 106 can be any combination of connections and protocols that support communication between the client computing device 102 and the server computer 108 and other computing devices (not shown) within the computing environment 100.

[0017] Server computer 108 is a digital device that hosts anomaly detector 110 and database 112. In this embodiment, server computer 108 can reside in a cloud architecture (e.g., public, hybrid, or private). In this embodiment, anomaly detector 110 resides on server computer 108. In other embodiments, anomaly detector 110 can have an instance of a program (not shown) stored locally on client computer device 102. In other embodiments, anomaly detector 110 can be a stand-alone program or system that trains a multi-language neural network intent classifier. In still other embodiments, anomaly detector 110 can be stored on any number or computing devices.

[0018] Anomaly detector 110 enables proactive anomaly detection for microservices applications by using a neural network approach to consider the dependencies of request contextual data. The solution provided by anomaly detector 110 is independent of the deployment of the microservices application (e.g., private cloud, public cloud, or hybrid) and supports various container orchestrators (e.g., Kubenetes, OpenShift, etc.). Anomaly detector 110 provides a mechanism for hybrid data collection based on the operation of both applications and systems. In this embodiment, anomaly detector 110 can include one or more components that are described in more detail with respect to FIG. 2.

[0019] For example, the anomaly detector 110 can receive end-user requests for an application that includes N microservices. In each microservice instance, each collection agent (associated with the anomaly detector 110) extracts the trace data and specifications of each instance. Then, the collector agent of the anomaly detector 110 compiles the received information (the respective trace data and specifications) and normalizes the received information. From there, the collector agent can push the data to a queue for persistence. The feature extraction module (shown and described in FIG. 2) converts the raw data into request context features. Next, the anomaly detector 110 can use the formatted context features to build a neural network model and then use the built model to generate predictions. The anomaly detector 110 can then generate proactive alerts.

[0020] In this embodiment, in response to receiving a request for predicting an abnormal operation, the anomaly detector 110 can request additional information from each microservice. The additional information can include context features, that is, a hierarchical data structure representing the end-to-end details of the request. The context features can include one or more causally related services and call paths. The context features can further include the execution context in each service instance (e.g., CPU, accelerator, memory utilization, pod area, network traffic, I / O requests, etc.).

[0021] For example, a request for additional information (e.g., request specifications), a microservice path, and a function path. Examples of additional information can include a user name (anonymized ID) related to the user, a company name (anonymized ID), latency (e.g., 500 ms), region (e.g., Europe), browser type, device type, operating system, time (e.g., Friday, February 28, 2020, 2:55:02 PM GMT - 05:00).

[0022] Examples of paths for microservices can include a path from microservice A to microservice B. For example, the cluster ID, region (us), instance ID, duration (100ms), OS specifications (CPU, memory, disk, network) related to microservice A, and the respective cluster ID, region (us), instance ID, duration (400ms), OS specifications (CPU, memory, disk, network) of microservice B.

[0023] Examples of call paths (i.e., function paths) can include one or more functions. For example, functions 1 to 3: Function 1 includes a duration (40ms) and resource utilization (20%, 100MB), function 2 includes a duration (60ms) and resource utilization (20%, 100MB), and returns to function 1 which includes a duration (400ms) and resource utilization (20%, 100MB).

[0024] In this embodiment, the anomaly detector 110 provides hybrid data collection to request context features, that is, the request for context features can be sent to or collected from different sources. In this embodiment, the anomaly detector 110 includes a collection agent (shown and described in FIG. 2) deployed as a sidecar within each microservice instance (e.g., two containers of a single Kubernetes Pod), and can extract from two different sources: trace data from microservices such as Jaeger, and OpenTelemetry, and the runtime characteristics of the microservice (e.g., CPU, memory utilization, network, other co-located sidecars, Zabbix-Agent (e.g., CPU, Disk, memory, etc.), Istio's Envoy (e.g., network), etc.).

[0025] From these sources, the anomaly detector 110 can collect categorical data and numerical data. In this embodiment, the categorical data refers to requests and microservice instances extracted from either the request header or environment variables on the deployment host. In this embodiment, the numerical data refers to data reported by a distributed tracing library such as OpenTelemetry or Jaeger, which reports the time spent on each microservice and its important functions. Thus, the anomaly detector 110 can utilize numerical data reporting that reports, records, and obtains each system usage information with appropriate permissions. Therefore, by collecting contextual features from different sources, the anomaly detector 110 can enable a comprehensive view of processing requests across layers.

[0026] Next, the anomaly detector 110 can use the collected contextual features (i.e., additional information) to hierarchically handle the aforementioned request contextual features as input, and build and train a neural network model that can predict future requests to be processed within each microservice application.

[0027] In this way, the anomaly detector 110 can capture the inter-request factors and intra-request factors (using the constructed neural network model) and use the captured factors to predict future requests. In this embodiment, the inter-request factors describe the connections between characteristics in the request specifications (for example, a login request for a user ID from a specific region is likely to be followed by a get_request to a product catalog page from the user ID in the same region). In this embodiment, as the intra-request factors, considering the factors of individual requests, from the data of the causally related microservice paths and function paths, understand which services in the processing path play the most important role for future requests. By considering these two elements, the constructed neural network model can capture the correlation between each microservice and the last step. For example, the historical requests from a microservice can take two paths. The first path can utilize microservices A, B, and C with respective latencies of 40ms, 15ms, and 300ms. The second path can utilize microservices A, B, and D with latencies of 200ms, 40ms, and 1.2s. The constructed neural network can predict the path using microservices A, B, and D when the latency at microservice A is high. For example, the latency of microservice A can be 300ms and the latency of microservice B can be 50ms. In this example, the anomaly detector 110 can (using the constructed neural network) predict that the next request should be processed by microservice D with a latency of 2s instead of C with a latency of 100ms, and at time 2.35s, the anomaly detector 110 can send a warning (for example, 2.35s = 300ms(A) + 50ms(B) + 2s(D)). The trace path (A→B→D) is the prediction result of the neural network model and captures the correlation between the duration of A and the previous selection time. This is required (for the prediction) through the neural network model constructed and later shown and described with respect to FIGS. 3 and 4.Specifically, the LSTM model will be trained to learn the order relationships between microservices and predict which ones will be used next.

[0028] In this embodiment, the anomaly detector 110 can utilize a controller (shown and described in FIG. 2) to interpret the prediction sequence and determine whether an anomaly has occurred. In this embodiment, the controller weights key performance metrics (e.g., latency, throughput, failed RPC calls, etc.). In this embodiment, the key performance metrics can be determined or defined by the owner of the microservice application. The controller calculates statistical measures (e.g., deviation, percentile) and determines whether to issue a proactive warning. For example, the controller can calculate the deviation according to the following formula: deviation = |xi - average(X)|. In this embodiment, the larger the deviation, the more unstable the dataset indicating a specific anomaly becomes. In this embodiment, the percentile is defined such that a specific percentage of the scores is below that numerical value. For example, the 50th percentile of an ordered list of numerical values is its median.

[0029] In this embodiment, the anomaly detector 110 can generate a proactive warning in response to the predicted abnormal behavior. The generated proactive warning can include the reason the anomaly was predicted or the reason it was flagged or both. In this embodiment, the proactive warning can be generated by a component of the anomaly detector 110 (e.g., the controller shown and described in FIG. 2). In this embodiment, the controller can perform appropriate visualization, generation of proactive warnings, generation of root cause reports, provision of resource management functions, and system simulation.

[0030] For example, the anomaly detector 110 can generate visualizations of each component that processes end-user requests. The requests are sent to the following cloud infrastructure, which includes the following components: front-end services, router services, dispatcher services, adapter services, on-premises infrastructure (e.g., legacy code), consumers, back-end services, and private cloud software (SaaS) as a service that includes databases in two different locations (e.g., the United States and Europe). In this example, the anomaly detector 110 can generate visualizations of each component and function path of the request, and can generate one or more graphic icons to visually indicate that the detected root cause can be one of the services (e.g., the dispatcher). In this way, the anomaly detector 110 can generate a visualization of the end-to-end execution flow of an abnormal request and highlight the dispatcher server as the root cause.

[0031] In this embodiment, the root cause report includes the predicted abnormal service and possible reasons, and the generated proactive warnings including those reasons. Continuing with the above example, the root cause report can include an explanation of the abnormal behavior in the dispatcher and can generate a proactive warning that a long latency violating service quality assurance affects the end user.

[0032] In this embodiment, the anomaly detector 110 can provide a resource management function that warns the system administrator and takes appropriate actions. For example, if the predicted reason for the anomaly is due to insufficient computing resources such as CPU, low memory, and low network latency, the system administrator can provision more resources before affecting the application client.

[0033] In this embodiment, the anomaly detector 110 can also provide system simulation. For example, the prediction results include details of the end-to-end execution flow in each microservice, including CPU, memory, disk, and network usage. Such fine-grained and characteristic traces provide insights into the applications required on the underlying hardware system and are used as drivers for the system simulator to evaluate potential cloud system designs and learn about issues and trade-offs (e.g., local vs. remote, routing flow / traffic control, robust cores vs. weak cores, latency requirements, advantages of offloading, etc.). This process helps cloud system designers understand the interactions between various hardware components, such as storage, network, CPU, memory, and accelerators, that make up different applications. It also helps analyze the potential advantages and degradations of different hardware configurations and guides future cloud system design decisions.

[0034] In an end-to-end example, the system handled by the anomaly detector 110 can receive requests for processing. The requests can be sent to the following cloud infrastructure, which includes the following components: front-end service, router service, dispatcher service, adapter service, on-premises infrastructure (e.g., legacy code), consumer, back-end service, and private cloud software (SaaS) as a service including databases in two different locations (e.g., the United States and Europe).

[0035] In the first scenario, the request can be processed by the front-end service, sent to the router, transmitted to the adapter that returns to the consumer, and finally sent to the back-end component. In this scenario, the anomaly detector 110 can generate proactive warnings in response to predicting that either the dispatcher or the back-end service will experience long latencies that affect the end user and violate the SLA. By using the anomaly detector 110, abnormal behaviors in the dispatcher and the back-end service can be detected and appropriately attributed to the service instances causing the delays. In contrast, in the current system using a prediction model, due to the mixed logs collected from concurrent requests, low-accuracy results (e.g., low precision) are obtained. Embodiments of the present invention (e.g., the anomaly detector 110) differ from the current approach in that the request context data includes a trace that separates the logs into individual requests. For example, if the router service is processing 10 requests simultaneously, 4 of them will be routed to the dispatcher and the others will be routed to the back-end. In the current approach, it may only be possible to see the mixed log data interleaved for concurrent processing. Therefore, if one or more requests fail, it is difficult to identify which request failed. In contrast, the anomaly detector 110 provides trace data (i.e., request context data), and we can identify which request failed at which service.

[0036] In a second scenario using the above components, the anomaly detector 110 can predict that the backend service is experiencing a slow response from the database where it stores user information and generate a proactive warning to inform users of the delay in response for a specific set of users. In contrast, current systems have difficulty detecting problems with statistics on aggregated metrics. In some scenarios, the aggregated metrics can mislead the monitoring components. For example, even if the average latency is below a certain threshold, the system is not necessarily healthy. In this example, assume that 90% of the traffic is routed to the European Union (EU) database and 10% is routed to the United States (US) database. If the EU database is normal and there is an anomaly in the US database service, 90% of the requests will have normal latency, so the average latency will still appear normal. Instead, our model (e.g., anomaly detector 110) considers the latency of individual traces so that it can identify anomalies on the execution path to the US database.

[0037] In a third scenario using the above components, the anomaly detector 110 can predict that a job started by the dispatcher service cannot be completed due to performance degradation in legacy code and generate a warning about the delay in the backend that receives results from the consumer. In contrast, current systems have difficulty modeling asynchronous relationships using the metrics of producer and consumer logs. Current systems use log data for training machine learning models. As described above, the log data collected from individuals is interleaved in such a way that it is difficult to derive causal relationships. Instead, since the request context is built on top of the traces, the anomaly detector 110 can avoid this problem.

[0038] The anomaly detector 110 can further utilize the prediction results to perform root cause analysis, resource management, and system simulation. For example, the prediction results can be used to drive a system simulator to understand the potential advantages and degradations from various hardware configurations and to guide the design decisions of future cloud systems.

[0039] The database 112 can represent one or more databases or generally available databases that store the received information and provide authorized access to the anomaly detector 110. Generally, the database 112 can be implemented using any non-volatile storage medium known in the art. For example, the database 112 can be implemented using a tape library, an optical library, one or more independent hard disk drives, or multiple hard disk drives within a redundant array of independent disks (RAID). In this embodiment, the database 112 is stored on the server computer 108.

[0040] FIG. 2 is a diagram showing an example of a block diagram 200 of an anomaly detector for microservices according to an embodiment of the present invention.

[0041] This exemplary diagram shows one or more components of the anomaly detector 110. In some embodiments, the anomaly detector 110 can include one or more hosts having respective microservices and collection agents, but it should be understood that the anomaly detector 110 can access the microservices and collection agents across the cloud architecture.

[0042] In this example, the anomaly detector can include host 202A, hosts 202B to 202N. Each host can have respective microservices and collection agents (e.g., respective microservices 204A to N and collection agents 206A to N).

[0043] In this example, the anomaly detector 110 can receive the end-user request microservice 204A via the collection agent 206A. In this example, the collection agent 206 can receive requests from end-users and can also receive requests from one or more other components (e.g., other collocation sidecars, Zabbix-Agent (e.g., CPU, Disk, memory, etc.), Istio's Envoy (e.g., network), etc.).

[0044] The collection agent 206A is responsible for making collection requests and extracting the trace data and specifications of each instance. In this embodiment, each collection agent can interface with the collector module of the anomaly detector 110 (e.g., collector module 20 8 ). The collector module 20 8 is responsible for compiling the received information (the respective trace data and specifications). The collector module 20 8 can then normalize the data using the normalization module 210, that is, the normalization module 210 normalizes the data into a consistent format (e.g., JSON or a common data structure). The collector module 20 8 can then push the compiled information into a queue for persistence.

[0045] Next, the feature extraction module 21 2 can access the data in the queue and extract contextual features from the compiled data. In other words, the feature extraction module 21 2converts raw data into request context features. For example, the request context features (i.e., the request specifications) can include the following: username (anonymized ID), company name (anonymized ID), latency (500ms), region (Europe), browser (Firefox), device (iOS), operating system, time (e.g., Friday, February 28, 2020, 2:55:02 PM GMT-05:00), each microservice path (e.g., the path from microservice A to microservice B. For example, the cluster ID related to microservice A, region (us), instance ID, period (100ms), OS specifications (CPU, memory, disk, network), each cluster ID of microservice B, region (us), instance ID, period (400ms), OS specifications (CPU, memory, disk, network), and function path (e.g., functions 1 to 3: function 1 includes a period (40ms) and resource utilization (20%, 100MB), function 2 includes a period (60ms) and resource utilization (20%, 100MB), and returns to function 1 which includes a period (400ms) and resource utilization (20%, 100MB).).

[0046] The anomaly detector 110 can then use the formatted context features to build a neural network model using the neural network module 214 (shown and described in FIGS. 3 and 4). The controller module 216 can then generate predictions using the built neural network model and can perform appropriate visualization, proactive warnings, root cause report generation, resource management capabilities, and system simulation.

[0047] FIG. 3 is a diagram showing an example of a block diagram 300 for designing a neural network model according to an embodiment of the present invention.

[0048] Specifically, block diagram 300 depicts the design of a neural network (some hidden layers are omitted). The input is the requirement specifications of a series of requests. The input Si to the requirement embedding layer is the output of the microservice path neural network model shown and described in FIG. 4.

[0049] In this example, the anomaly detector 110 receives inputs 302A, 302B, through 302N (r1 specifications). For example, the request input, i.e., additional information, can include context hierarchical trace data collected during a specified time (e.g., time window, T). This request input can include the requirement specification, microservice path, and function path. Examples of additional information for the requirement specification can include the user name (anonymized ID) related to the user, company name (anonymized ID), latency (e.g., 500 ms), region (e.g., Europe), browser type, device type, operating system, time (e.g., Friday, February 28, 2020, 2:55:02 PM GMT - 05:00).

[0050] Examples of the microservice path can include the path from microservice A to microservice B. For example, the cluster ID, region (us), instance ID, period (100 ms), OS specifications (CPU, memory, disk, network) related to microservice A, and the respective cluster ID, region (us), instance ID, period (400 ms), OS specifications (CPU, memory, disk, network) of microservice B.

[0051] In the example of the call path (i.e., function path), it can include one or more functions. For example, functions 1 to 3: Function 1 includes a period (40 ms) and resource utilization (20%, 100 MB), function 2 includes a period (60 ms) and resource utilization (20%, 100 MB), and returns to function 1 which includes a period (400 ms) and resource utilization (20%, 100 MB).

[0052] Next, the received input performs embedded processing of the requirement specifications (e.g., R1 and A1, which are 304A~N and 306A~N respectively) at block 320. In this embodiment, "R1" is the embedding result of the string part in the requirement specification (e.g., username, browser type, etc.), and "A1" refers to the numerical part related to the requirement specification. In this embodiment, the anomaly detector 110 concatenates the embedding result with the numerical part of the requirement specification (e.g., latency, called A1~AN).

[0053] The anomaly detector can then combine the embedded requirement specifications with components B1 and S1, called 308A~N and 310A~N respectively. In this embodiment, B1~BN are the outputs with the requirement specifications embedded. In this embodiment, S1 is the output of the model described in FIG. 4. In this embodiment, S1 represents the modeled output of the end-to-end execution flow of a single request.

[0054] The processing for in-request embedding continues at block 330. The in-request factors include B1, S1, and C1. In this embodiment, B1, S1, and C1 are related to a single requirement specification. Similarly, B2, S2, and C1 are related to another requirement specification. C1 is an embedding layer (called 312A~N) for converting the combination of B1 and S1 into a vector.

[0055] The processing continues at blocks 340 and 350 (e.g., L STIt continues to add an inter - requirement factor including M340 and density 350. In block 340, the contextual features are supplied via a long short - term (LSTM) architecture used in the field of deep learning, D1 is added, and they are called 314A - N respectively. In this embodiment, D1 is a single unit of the LSTM model. Recall that C1, C2,... CN are the modeled outputs of individual requirements. The anomaly detector 110 uses the LSTM model to learn the intra - requirement relationships between requirements. In this embodiment, D1 - Dn are units of the LSTM model. Finally, at density 350, E1 called 316A - N is added. In this embodiment, E1 - EN are units of a tightly - coupled network, which reduces the dimensionality of the input to find their internal correlations. The resulting output is Y1, Y2 - Y N and are referred to as 318 A~N respectively.

[0056] FIG. 4 is an exemplary block diagram 400 of a neural network model that captures intra - requirement factors for individual requirements according to an embodiment of the present invention.

[0057] Inputs (e.g., F 1,1 , F 1,2 , F 2,1 and F B1 , called 402A, 402B, 402C, 402N respectively) are descriptions of functions in the requirement specifications of a series of requirements. The anomaly detector 110 receives the received inputs and performs an embedding of the requirement specifications (e.g., block 420). In this embodiment, G 1,1 , G 1,2 , G 2,1 and G B,1 are referred to as 404A, 404B, 404C - 404N, and H 1,1 , H 1,2 , H 2,1 , H B,1 are referred to as 406A, 406B, 406C, 406N respectively. G 1,1 , G 1,2 are embedding layers for the string parts in function F 1,1 . Similarly, G 2,1is the embedding unit for the string part in function F 2,1 H is the concatenation of the numerical parts of G 1,1 and F 1,1 collectively, 404A to N and 406A to N function in the same way as 304A to N and 306A to N described in FIG. 3 1,1

[0058] In this embodiment, in block 430, the embedded requirement specifications are supplied via long short-term memory (LSTM) and artificial recurrent neural network (RNN), and each K 1,1 K 1,2 K 2,1 and K B,1 (i.e., the units of the L ST M model are referred to as 408A, 408B, 408C, and 408N respectively) are added

[0059] The process continues to block 440 for microservice embedding, and M1, M2, and M B and O1, O2, and O B are added respectively. M1, M2, and M B are referred to as blocks 410A, 410B, and 410N and represent the output of the L ST M model of the B microservice (e.g., block 430), and O1, O2, and O B are referred to as blocks 412A, 412B, and 412N respectively and refer to the embedding of the specifications of the B microservice

[0060] The process continues to block 450, and the result of block 440 is supplied via another LTS layer, and P1, P2, and P B are added respectively. P1, P2, and P B are referred to as blocks 414A, 414B, and 414N respectively. In this embodiment, P1, P2, and P B are the units of the L ST M model of block 450

[0061] ​The result output of block 450 is supplied via block 460. Block 460 is a density layer that provides learned features from all combinations of features of the previous layer, and adds Q1, Q2, and QB, referred to as 416A, 416B, and 416N respectively.

[0062] In this embodiment, Z1, Z2, and Z N (respectively 418 A , 418 B and 418 N referred to as) are the result outputs of the workflow of block diagram 400. Collectively 418 A , 418 B and 418 N represent the modeled output of the end-to-end execution flow of a single request. 418 B and 418 N are referred to as S1 and are depicted incorporated into the model described in FIG. 3.

[0063] FIG. 5 is a flowchart 500 showing the operational steps for training an end-to-end speech, multi-language intent classifier according to an embodiment of the present invention.

[0064] In step 502, anomaly detector 110 receives information. In this embodiment, the received information can include an end-user request for an application that includes N microservices. For example, the end-user request can be a request triggered by a user's request for a front-end service. For example, when a user accesses a web page and presses a login button, a login request is generated for the application.

[0065] In this embodiment, anomaly detector 110 receives requests from client computing device 102. In other embodiments, anomaly detector 110 can receive information from one or more other components of computing environment 100.

[0066] In step 504, the anomaly detector 110 generates contextual information from the received information. In this embodiment, the anomaly detector 110 generates contextual information from the received request by requesting additional information and creating a hierarchical data structure representing the end-to-end details of the received request.

[0067] Specifically, the anomaly detector 110 can request additional information (e.g., request specifications) that may include a user name (anonymized ID), company name (anonymized ID), latency (e.g., 500 ms), region (e.g., Europe), browser type, device type, operating system, time (e.g., Friday, February 28, 2020, 2:55:02 PM GMT-05:00), microservice path, and function path related to the user.

[0068] The request for contextual features can be sent to the delta source or otherwise collected from the delta source. In this embodiment, the anomaly detector 110 includes a collection agent (shown and discussed in Figure 2) deployed within each microservice instance as a sidecar (e.g., two containers of a single Kubernetes Pod), and can derive from two different sources: trace data from microservices such as Jaeger, and OpenTelemetry) and the runtime characteristics of the microservice (e.g., CPU, memory utilization, network, other co-located sidecars, Zabbix-Agent (e.g., CPU, Disk, memory, etc.), Istio's Envoy (e.g., network), etc.).

[0069] From these sources, the anomaly detector 110 can collect categorical data and numerical data. In this embodiment, the categorical data refers to requests and microservice instances extracted from either the request header or environment variables on the deployment host. In this embodiment, the numerical data refers to data reported by a distributed tracing library such as OpenTelemetry or Jaeger about the time spent on each microservice and its important functions. In this way, the anomaly detector 110 can utilize numerical data reporting that reports, records, and obtains each system usage information with appropriate permissions. Therefore, by collecting contextual features from different sources, the anomaly detector 110 can enable a comprehensive view of processing requests across layers.

[0070] In step 506, the anomaly detector 110 trains a neural network based on the generated contextual information. In this embodiment, the anomaly detector 110 trains a neural network based on the generated contextual information including inter-request factors and intra-request factors. As described above, inter-request factors describe the connections between characteristics in the request specification (for example, a login request with a user ID from a specific region is likely to be followed by a get_request to a product catalog page from a user ID in the same region). In contrast, intra-request factors consider the factors of individual requests and understand which services in the processing path play the most important role for future requests from the data of causally related microservice paths and function paths. By considering these two elements, the constructed neural network model can capture the correlation between each microservice and the last step. In this way, the trained neural network can predict what the next series of requests and their contextual requests will be. And based on that prediction, the controller module determines whether there is an anomaly.

[0071] In step 508, the anomaly detector 110 predicts abnormal operations using the trained neural network model. For example, the anomaly detector 110 can predict anomalies such as SLA violations (e.g., the tail latency will increase in the next 10 minutes), the users who will be affected (e.g., a subset of users in the southern region of U), and the impact on a subset of requests (e.g., the acquisition of analysis results fails).

[0072] In step 510, the anomaly detector 110 takes appropriate actions based on the predicted abnormal operations. In this embodiment, the appropriate actions can be the generation of proactive warnings, the generation of root cause reports, the provision of resource management capabilities, and system simulation. For example, the anomaly detector 110 can then determine whether to send a proactive warning based on the prediction. In this embodiment, the anomaly detector 110 can automatically generate a proactive warning in response to predicting an anomaly. In another embodiment, the anomaly detector can generate a weighted score for the predicted anomaly and generate a proactive warning in response to the predicted anomaly meeting or exceeding a threshold for abnormal operations.

[0073] For example, proactive warnings can include predictions such as SLA violations (e.g., the tail latency will increase in the next 10 minutes), the users who will be affected (e.g., a subset of users in the southern region of U), and the impact on a subset of requests (e.g., the acquisition of analysis results fails).

[0074] Examples of root cause reports can include the identification of failed microservice instances and the reasons for the failures. For example, the database connection is slow, the computing resources are insufficient, etc.

[0075] In some embodiments, resource management can include recommended modifications. For example, the anomaly detector 110 can recommend provisioning microservice instances on nodes with greater capacity, increasing the network bandwidth between the backend and the database, adding nodes with more powerful CPUs, and so on.

[0076] FIG. 6 shows an exemplary FIG. 600 according to an embodiment of the present invention.

[0077] For example, FIG. 6 shows an overview of a sequence-to-sequence (seq2seq) model with encoder and decoder parts, and their inputs and outputs (representing the methodology described above). Both the encoder (e.g., block 602) and decoder (e.g., block 604) parts are RNN-based and can consume and return an output sequence corresponding to multiple time steps. The model takes inputs from the previous N values, and it returns the next N predictions. N is a hyperparameter, which is empirically set to 10 minutes in this figure. In the center of the figure, there is a hierarchical RNN-based anomaly detection neural network, which includes three main components: in-request factors, inter-request factors, and embeddings.

[0078] Specifically, the figure in FIG. 6 is an encoder-decoder architecture (commonly known as a seq2seq model). In this embodiment, X, X1, X2, ···, Xn represent the inputs to the model, which are the request contextual data of a series of requests. In this embodiment, Y, Y1, Y2, ··· Y n is the output of the model and is the predicted value of the model. The internal architecture of the model has been described in detail through FIGS. 3 and 4 and has been discussed previously.

[0079] FIGS. 7(A) and 7(B) are diagrams showing exemplary data collection code according to an embodiment of the present invention.

[0080] Specifically, FIG. 7(A) depicts exemplary data collection code 700 which is exemplary application code in each microservice.

[0081] Regarding FIG. 7(B), FIG. 7(B) depicts exemplary data collection code 750. Specifically, the exemplary data collection code 750 represents the code of the collection agent.

[0082] FIG. 8 is a block diagram of components of a computing system within the computing environment 100 of FIG. 1 according to an embodiment of the present invention. It should be understood that FIG. 8 provides only an example of one implementation and does not imply any limitation regarding the environment in which different embodiments may be implemented. Many modifications can be made to the depicted environment.

[0083] The programs described herein are identified based on the applications in which they are implemented in particular embodiments of the present invention. However, it should be understood that any specific program nomenclature herein is used merely for convenience, and thus the present invention should not be limited to use in any particular application that is identified by, or implied by, or both, such nomenclature.

[0084] The computer system 800 includes a communication fabric 802 that provides communication between a cache 816, a memory 806, a persistent storage 808, a communication unit 812, and an input / output (I / O) interface 814. The communication fabric 802 can be implemented in any architecture designed to pass data or control information or both between a processor (such as a microprocessor, a communication and network processor, etc.), the system memory, peripheral devices, and any other hardware components within the system. For example, the communication fabric 802 can be implemented using one or more buses or a crossbar switch.

[0085] Memory 806 and persistent storage 808 are computer-readable storage media. In this embodiment, memory 806 includes random access memory (RAM). Generally, memory 806 can include any suitable volatile or non-volatile computer-readable storage media. Cache 816 is fast memory and improves the performance of processor 804 by holding recently accessed data and data close to recently accessed data from memory 806.

[0086] Anomaly detector 110 (not shown) can be stored in persistent storage 808 and memory 806 for execution by one or more of respective computer processors 804 via cache 816. In one embodiment, persistent storage 808 includes a magnetic hard disk drive. Alternatively, or in addition to a magnetic hard disk drive, persistent storage 808 can include a solid state hard disk drive, a semiconductor memory device, read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, or any other computer-readable storage media capable of storing program instructions or digital information.

[0087] The media used by persistent storage 808 may be removable. For example, a removable hard drive may be used for persistent storage 808. Other examples include optical disks, magnetic disks, thumb drives, and smart cards, which are inserted into a drive for transfer to another computer-readable storage media that is also part of persistent storage 808.

[0088] In these examples, communication unit 812 enables communication with other data processing systems or devices. In these examples, communication unit 812 includes one or more network interface cards. Communication unit 812 may enable communication using either or both physical and wireless communication links. Anomaly detector 110 may be downloaded to persistent storage 808 via communication unit 812.

[0089] I / O interface 814 enables the input and output of data with other devices that may be connected to a client computing device or a server computer or both. For example, I / O interface 814 enables connection with an external device 820 such as a keyboard, keypad, touch screen, or other suitable input device or combinations thereof. Further, external device 820 may include a portable computer-readable storage medium such as, for example, a thumb drive, portable optical disk, portable magnetic disk, and memory card. Software and data (e.g., anomaly detector 110) used to implement embodiments of the present invention can be stored on such portable computer-readable storage media and loaded into persistent storage 808 via I / O interface 814. I / O interface 814 is also connected to a display 822.

[0090] Display 822 realizes a mechanism for displaying data to a user and can be, for example, a computer monitor.

[0091] The present invention can be a system, method, or computer program product or a combination thereof. The computer program product may include a computer-readable storage medium storing computer-readable program instructions for causing a processor to execute aspects of the present invention.

[0092] A computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. As an example, a computer-readable storage medium may be an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. As a more specific example of a computer-readable storage medium, there may be a portable computer diskette, a hard disk, a RAM, a ROM, an EPROM (or flash memory), an SRAM, a CD-ROM, a DVD, a memory stick, a floppy disk, a punched card, or a mechanically encoded device with instructions recorded in a raised structure in a groove, and suitable combinations thereof. A computer-readable storage device as used herein should not be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted via a wire.

[0093] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network or a combination thereof). The network is composed of copper wire transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers or a combination thereof. The network adapter card or network interface of each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in the computer-readable storage medium within each respective computing / processing device.

[0094] Computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or an object-oriented programming language such as Smalltalk, C++, or a conventional procedural programming language such as the "C" programming language and similar programming languages. The computer-readable program instructions may be executable entirely on the user's computer as a stand-alone software package, or partially on the user's computer. Alternatively, it may be executable partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may utilize the state information of the computer-readable program instructions to personalize and thereby execute the computer-readable program instructions for carrying out aspects of the present invention.

[0095] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0096] These computer-readable program instructions can be provided to a general-purpose computer, a processor of a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus generate means for implementing the functions / acts specified in one or more blocks of a flowchart, a block diagram, or both. These computer-readable program instructions may also be stored in a computer-readable storage medium that can be connected to a computer, a programmable data processing apparatus, or other devices that function in a particular manner or a combination thereof, such that the computer-readable program instructions stored therein constitute one of the manufactured articles that include instructions for implementing the aspects of the functions / acts specified in one or more blocks of a flowchart, a block diagram, or both.

[0097] Like the instructions for performing the functions / acts specified in one or more blocks of a flowchart, a block diagram, or both on a computer, other programmable apparatus, or other device, the computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to perform a series of operational steps on the computer, other programmable apparatus, or other device to generate a computer-implemented process.

[0098] The flowcharts and block diagrams in the figures illustrate the configuration, functionality, and operation of the implementation that can be executed by systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of an instruction, which constitutes one or more executable instructions for implementing the specified logical function. In some alternative embodiments, the functions shown in the blocks may occur in a different order than shown in the figures. For example, two blocks shown in succession may actually be executed substantially simultaneously, or the blocks may be executed in the reverse order depending on the related functions. It should also be noted that each block of the block diagram or flowchart diagram, or a combination of blocks of the block diagram or flowchart diagram, or both, can be implemented by a special-purpose hardware-based system that executes the specified function or operation, or a combination of special-purpose hardware and computer instructions.

[0099] The descriptions of the various embodiments of the present invention are presented for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. It will be apparent to those skilled in the art that many modifications and variations are possible without departing from the scope of the present invention. The terms used in this specification are selected to best explain the principles of the embodiments, the practical application to technologies found in the marketplace, or the technological improvements, or to enable those skilled in the art to understand the embodiments described in this specification.

[0100] (Additional comments or embodiments or both) Some embodiments of the present invention recognize the following facts, potential problems, or potential improvement areas, or combinations thereof, with respect to current state-of-the-art technology. Microservices architecture is attractive for applications deployed in hybrid cloud environments because loosely coupled components offer better scalability, flexibility, and acceleration of developer productivity. To avoid severe financial and business losses due to SLA violations, one of the most important tasks in the management of microservices applications is to effectively and efficiently detect and diagnose anomalies at specific time steps so that DevOps / SRE can take further actions to timely resolve root problems. However, existing approaches for issuing proactive warnings for detected anomalies are not yet effective for microservices applications because they do not consider spatial and temporal dependencies buried in multivariate time series data obtained from isolated services and end-user requests.

[0101] Some embodiments of the present invention can include one or more of the following features, characteristics, or advantages, or combinations thereof. The problem of tail latency is learned by the model and helps to predict before potential anomalies occur.

[0102] Embodiments of the present invention predict anomalies in microservices applications and identify root causes. Among existing studies on anomaly prediction, embodiments of the present invention are the first to perform a dual task for predicting request patterns and their paths (i.e., the services through which requests pass). Embodiments of the present invention design a collection agent for collecting data from the application deployment. This system supports the deployment of microservices applications in different environments, private, public, and hybrid.

[0103] In an embodiment of the present invention, the concept of request context features, which is a data structure including three levels of request information (request specifications, microservice paths, and function paths), is defined. The proposed features integrate two historical data, inter-request factors and intra-request factors, which affect the performance and processing path of received requests.

[0104] In an embodiment of the present invention, a hierarchical neural network model is designed to integrate training data of request context features. This model is based on a seq2seq architecture with heterogeneous data embedding and attention mechanisms, and can enhance the interpretability of results to a certain level.

[0105] There are two unique advantages of application-specific system trace information. Utilize timestamped system usage information to understand and predict system resource requirements, and further guide system administrators to reallocate resources to meet QoS requirements. In addition, the detailed and fine-grained system characteristics obtained from the application can understand various hardware impacts and trade-offs through system simulation, and utilize the lessons learned as input for future cloud system design.

[0106] Embodiments of the present invention enhance proactive warnings and anomaly diagnosis of microservice applications by analyzing the aforementioned dependencies available in request context data in horizontal and vertical directions using deep learning. The proposed approach addresses the following two specific questions: (1) Will a performance anomaly (e.g., SLA violation, increase in tail latency) occur at a specific time step elapsed from the current moment? and (2) If (1) is true, which microservice is most likely to cause the anomaly?. The first question is related to anomaly prediction, and the second question is related to the root cause of the predicted anomaly.

[0107] The issues of proactive warnings and anomaly diagnosis can be regarded as a pre - prediction task about how a series of microservices will cooperate to handle future requests. The technology we propose is a neural network approach that integrates the detailed characteristics of historical requests, including both its specifications and the trace information of each microservice instance along its path. This neural network model can predict whether anomalies (such as tail latency, SLA violations, etc.) will occur and what their root causes are. This solution is independent of the deployment of microservice applications (private cloud, public cloud, or hybrid) and supports various container orchestrators such as Kubernetes and OpenShift.

[0108] (Key idea) Key idea 1: Introduce the concept of request contextual characteristics. This is a hierarchical data structure representing the end - to - end details of a request, including causally related services and call paths, and the execution context (CPU, accelerator, memory usage, pod area, network track, IO requests, etc.) in each microservice. Request contextual characteristics are composed of information in three categories: request specifications, microservice paths, and function paths (see Section 6.2 for details). Each category contains heterogeneous forms of data such as scalars, vectors, and categorical. These collected feature points are provided as training data for the neural network.

[0109] Key Idea 2: Develop a method to collect data on the contextual features of requests from different sources (Section 6.1). The categorical data that describes requests and instances of microservices is extracted from either the request headers or the environment variables of the deployment host. The numerical data that reports the time spent on each microservice and its important functions is obtained from distributed tracing libraries such as OpenTelemetry or Jaeger, and the data that reports resource usage is recorded by obtaining information on system usage with appropriate permissions. As a result, the contextual features of requests can be comprehensively grasped across layers during the processing of requests.

[0110] Key Idea 3: Build a neural network model that predicts how future requests will be processed within a microservice application by hierarchically handling the aforementioned contextual features of requests as inputs. We consider that request processing prediction is a sequential problem with long-distance dependencies. That is, request processing in the near future depends on two groups of factors: inter-request factors and intra-request factors. Inter-request factors represent the connections between the characteristics included in the request specifications such as http methods, usernames, regions, etc. For example, a login request by a user ID from a certain region is likely to be followed by a request to obtain a product catalog page from the same region and the same user ID. Intra-request factors take into account the factors of individual requests. When processing requests, the microservices of the application cooperate by sending RPC calls to each other. Furthermore, since there are often many replicas for each microservice, not all instances will appear in the call path. An effective model should be able to understand which services in the processing path will play the most important role for future requests from the data of causal microservice paths and function paths. All of the above factors are captured during the training process by the proposed model.

[0111] Key Idea 4: During monitoring, the model generates the representation of the predicted requests one time step at a time, capturing dependencies between and within complex requests. A controller is created that interprets a series of pre-predictions to examine key performance indicators (e.g., latency), calculate statistical metrics (e.g., deviation values, percentiles), and determine whether to issue a warning. When the controller decides to issue a warning, the root cause analysis module interprets a series of representations supplemented with the current trend to identify the root cause (e.g., memory shortage in a specific microservice instance in a certain region, slow connection between a specific microservice instance and the backend storage).

[0112] (Motivating Example) As an example motivating this prediction problem, we will describe a microservices application consisting of four services. Each request must be processed by either A and B, and either C or D. In this particular scenario, there are two historical requests. The service paths are A→B→C and A→B→D. Considering only the order of these requests (i.e., the inter-request factors) to predict the next request and its path, the result would be A→B→C. A model learned from inter-request factors considers the order of requests as an important feature in the prediction process. Considering the effect of load balancing where C and D appear alternately in the historical data, this result is reasonable and the predicted total latency is <1 second. On the other hand, the model we proposed functions intelligently as it pays more attention to the latency along the service path. This may be due to an increase in the processing time at service instance A and a correlation between A and the selection of the last hop. Therefore, when the latency at A is high, service D is more likely to be selected, and the correct next request and its path A→B→D can be predicted well. The total predicted latency of the requests is 2.3 seconds, which is greater than the threshold (1.5 seconds), so a proactive warning will be sent to the SRE. To make a correct prediction, it is necessary to jointly consider the inter-request factors and intra-request factors in individual requests. These factors can be discovered from detailed information about the request path such as trace data, resource utilization, and specifications.

[0113] (Explanation) In this section, we introduce the methodological and technical details of the proposed approach to address the problems of proactive warning and anomaly diagnosis in microservices applications. In the first step, for both normal and abnormal operations, a series of request trace data and specifications are collected and prepared for feature extraction. In this second step, request context features are assembled from the collected data to generate a neural network model. In the third step, the pre-trained model is used to predict anomalies and serve to present a list of root causes.

[0114] As shown in FIG. 2, the high-level architecture of the proposed system consists of an application composed of N microservices with a uniquely designed collection agent, and a model creation and prediction pipeline. In the remainder of this section, the end-to-end process will be described in detail.

[0115] (Data Collection) First (as described in steps 502-504 of flowchart 500), the collection agent collects trace data from microservices located in the same place. The microservice and collection agent pair are executed in separate containers of a single Kubernetes pod. The microservice runs application code to process requests and passes them to downstream services. Additionally, the collection agent can aggregate important system information from sidecars such as the Zabbix agent or Istio's Envoy proxy.

[0116] The application code operating within the microservice uses a distributed tracing library such as Jaeger or OpenTelemetry to record the time spent on functions important to the business logic and sends the trace data to the collection agent via UDP packets. Note that in the proposed method, in the front-end service, it is necessary to capture the specification of the user request only once (see, for example, FIG. 7(A) described above). In addition to the trace information within the microservice, the collection agent also needs to obtain not only the static configuration of the microservice instance but also the dynamic resource utilization rate when receiving traces from the microservice (see, for example, FIG. 7(B) already described). Such data can be obtained from sidecars as described above. The collection agent batches this data and distributes it to a centralized collector.

[0117] Since the collector is implemented as a stateless server, it can be scaled out to a large number of replicas. The collector receives trace data and request specifications, normalizes them into a common representation, and pushes them into a queue. An example of a queue is Kafka. Kafka is open-source software that provides a high-throughput, low-latency platform for handling real-time data feeds (capable of writing up to one million records per second).

[0118] The anomaly detector can pull from the queue into a feature extraction module developed as a streaming-based job on top of the Flink framework. The feature extraction job is to transform the collected data into the form of request context features.

[0119] (Details of features) The collected features are summarized into three categories: request specifications, microservice paths, and function paths. Request specifications are static and contain self-descriptive information of the request, and the most important one is the end-to-end latency between a series of microservices that make up the application. The features of microservice paths and function paths are collected as causally related data to describe the request processing path. Figure 6 shows the hierarchical data structure collected at each step within the time window.

[0120] (Neural network model) The design of the neural network model is based on the seq2seq architecture. As described in Figure 6, the neural network model includes an encoder and a decoder part, and their inputs and outputs. Both the encoder and decoder parts are RNN-based and can consume and return an output sequence corresponding to multiple time steps. The model takes inputs from the previous N values and returns the next N predictions. N is a hyperparameter and is empirically set to 10 minutes in this figure. In the center of the figure, there is a hierarchical RNN-based anomaly detection neural network, which includes three main components (intra-request factors, inter-request factors, and embeddings). In the rest of this section, the details of the neural network will be described.

[0121] As described above, Figure 3 shows the design of the neural network. Regarding the intra-request factors, it combines the characteristics of a series of microservice paths and the corresponding requirement specifications. The characteristics of the microservice paths are detailed in Figure 4, which is also an RNN-based network. Regarding the inter-request factors, in order to train the inter-request patterns, the intra-request factors of a series of requests are fed into another RNN layer (e.g., LSTM). Throughout the network, different embedding layers (e.g., word2vec, ELMO) are applied to convert heterogeneous data into N-dimensional vectors (e.g., N = 300). The hierarchical requirement prediction neural network has the ability to learn the influence of inter-request patterns and intra-request patterns on the processing of future requests. As emphasized earlier, the embodiments of the present invention aim to predict the specifications of future requests and their paths through the microservice instances of the application.

[0122] (Monitoring and Insights) Our proactive anomaly detection problem involves two main tasks: predicting future requests with detailed service paths and anticipating SLA violations based on the predictions (step 508 in Figure 5). The first is performed by the prediction module (e.g., step 510 in Figure 5). During the monitoring stage, the system continuously collects request contextual data from the running applications and feeds them into the prediction module. These data are supplied to a neural network model fetched from storage. The output of the prediction module is a sequence of predicted requests that will occur within the next Wt seconds. For example, based on experience, Wt is set to 500ms so that the automatic resource partitioning software has the opportunity to act.

[0123] For the second task of determining proactive warnings, a controller that interprets the output from the prediction module is integrated. As shown in steps 510 of Figures 2 and 5, the controller has multiple functions. Regarding proactive warnings, it calculates the tail of the predicted latency. If the result is greater than a specific threshold, a proactive warning is issued. The details of the prediction results are further utilized for advanced missions such as root cause analysis, resource management, and system simulation.

[0124] System Simulation: The output in Figure 3 includes trace information of the detailed system of the on-the-fly application from the Zabbix agent (such as CPU, memory, disk, and network usage). As described in Figure 1, in system simulation, such fine-grained characteristic traces provide insights into the applications required on the underlying hardware system, and this can be used as a driver for the system simulator to evaluate potential cloud system designs and learn about issues and trade-offs. This process helps cloud system designers understand the interactions between various hardware components composed of various applications such as storage, network, CPU, memory, and accelerators. It also helps analyze the potential advantages and degradations of various hardware configurations and guides the design decisions of future cloud systems.

[0125] (Definition) The present invention: The term "the present invention" should not be construed as absolutely indicating that the subject matter described by this term is covered by the claims at the time of filing or by the claims that may ultimately be issued after patent examination. The term "the present invention" is used to help the reader get a general sense of which disclosures in this specification are potentially considered new, but the understanding indicated by the use of this term "the present invention" is provisional and tentative, indicating that it may change during the patent examination process due to the development of relevant information and the amendment of the claims.

[0126] Embodiment: Refer to the above definition of "the present invention". Similar considerations apply to the term "embodiment".

[0127] And / or: Inclusive disjunction; for example, A, B "and / or" C means that at least one of A or B or C is true and applicable.

[0128] Includes (inducing / include / includes): Unless otherwise specified, it means "including but not necessarily limited to".

[0129] User / subscriber: Includes but is not necessarily limited to the following: (i) a single human being, (ii) an artificial intelligence entity having sufficient intelligence to act as a user or subscriber, or (iii) a group of related users or subscribers, or a combination thereof.

[0130] Module / sub-module: Any set of hardware, firmware, or software, or a combination thereof, that operates functionally to perform a certain function, and the module may be (i) in a single local proximate location, (ii) widely distributed, (iii) in a single proximate location within a large software code, (iv) within a single software code, (v) within a single storage device, memory, or medium, (vi) mechanically connected, (vii) electrically connected, (viii) connected by data communication, and includes all such things.

[0131] Computer: Any device having significant data processing or machine-readable instruction reading capabilities or both, including but not limited to desktop computers, mainframe computers, laptop computers, field programmable gate array (FPGA)-based devices, smartphones, personal digital assistants (PDAs), body-mounted or plug-in computers, embedded device-style computers, application-specific integrated circuit (ASIC)-based devices, etc.

Claims

1. In response to receiving a request, collecting trace data recorded using a distributed tracing library when a microservice application processes the request, and the specification of the request; Generating request context features from the collected trace data and specifications; Training a neural network model based on the generated request context features; Predicting abnormal behavior of the microservice application using the trained neural network model; Including; Generating the request context features from the collected trace data and specifications includes integrating inter-request factors and intra-request factors related to the request, a computer-implemented method.

2. Further including generating a visualization related to the predicted abnormal behavior The computer-implemented method according to claim 1.

3. Further including generating a root cause report of the predicted abnormal behavior The computer-implemented method according to claim 1.

4. Further including providing a system simulation for the predicted abnormal behavior The computer-implemented method according to claim 1.

5. The trace data provides a hierarchical data structure that separates logs into individual requests, the computer-implemented method according to claim 1.

6. The neural network model is a recurrent neural network, the computer-implemented method according to claim 1.

7. The request context features include a data structure including information at three levels of the request: the specification of the request, the microservice path, and the function path, the computer-implemented method according to claim 1.

8. One or more computer-readable storage media, and program instructions stored on the one or more computer-readable storage media, the program instructions In response to receiving a request, program instructions for collecting trace data recorded using a distributed tracing library when a microservice application processes the request, and the specification of the request; Program instructions for generating request context features from the collected trace data and specifications; Program instructions for training a neural network model based on the generated request context features; Program instructions for predicting abnormal operations of the microservice application using the trained neural network model, and including The program instructions for generating the requirement context features from the collected trace data and specifications include A computer program product including program instructions for integrating between-requirement factors and within-requirement factors related to the requirement. **Claim 9** The program instructions stored in the one or more computer-readable storage media Further include program instructions for generating a visualization related to the predicted abnormal operation The computer program product according to claim 8. **Claim 10** The program instructions stored in the one or more computer-readable storage media Further include program instructions for generating a root cause report of the predicted abnormal operation The computer program product according to claim 8. **Claim 11** The program instructions stored in the one or more computer-readable storage media Further include program instructions for providing a system simulation for the predicted abnormal operation The computer program product according to claim 8. **Claim 12** The trace data provides a hierarchical data structure that separates logs into individual requests. The computer program product according to claim 8. **Claim 13** The neural network model is a recurrent neural network. The computer program product according to claim 8. **Claim 14** The requirement context features Include a data structure including information at three levels of requirements: requirement specifications, microservice paths, and function paths. The computer program product according to claim 8. **Claim 15** One or more computer processors, and One or more computer-readable storage media, and Program instructions stored in the one or more computer-readable storage media for execution by at least one of the one or more computer processors, the program instructions include In response to receiving a request, program instructions for collecting trace data recorded using a distributed tracing library when a microservice application processes the request, and the specifications of the request, and Program instructions for generating requirement context features from the collected trace data and specifications Program instructions for training a neural network model based on the generated requirement context features; Program instructions for predicting abnormal operations of the microservice application using the trained neural network model; comprising; The program instructions for generating the requirement context features from the collected trace data and specifications include program instructions for integrating inter-requirement factors and intra-requirement factors related to the requirements, a computer system. **Claim 16** The program instructions stored in the one or more computer-readable storage media further include program instructions for generating a visualization related to the predicted abnormal operation The computer system according to claim 15. **Claim 17** The program instructions stored in the one or more computer-readable storage media further include program instructions for generating a root cause report of the predicted abnormal operation The computer system according to claim 15. **Claim 18** The program instructions stored in the one or more computer-readable storage media further include program instructions for providing a system simulation for the predicted abnormal operation The computer system according to claim 15.

Citation Information

Patent Citations

  • System, Method and Apparatus for Determining Virtual Machine Performance

    US20140223427A1

  • Unsupervised learning to simplify distributed systems management

    US20200287923A1