Determining key logs for web applications

Through the cross-layer log analysis system, template mining models and knowledge graphs are used to solve the root cause determination problems of network application performance problems in data centers, and efficient and simplified troubleshooting is achieved.

CN120234207APending Publication Date: 2025-07-01JUNIPER NETWORKS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411717144.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-11-27
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In a data center environment, when investigating performance issues in network applications, independent management of each layer leads to inefficiency and challenging, making it difficult to efficiently determine the root cause of performance issues.

Method used

Through a cross-layer log analysis system, candidate logs are mapped to log templates, template mining models are used to generate and rank, key logs are selected to identify the root causes of performance problems, and combined with knowledge graphs and machine learning models to reduce analysis resource consumption.

Benefits of technology

Efficiently and thoroughly eliminate network application performance problems, reduce the amount of analytical log data, simplify troubleshooting, and improve the accuracy and efficiency of causal relationship analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234207A_ABST
    Figure CN120234207A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to determining key logs for web applications. Techniques are described for a computing system configured to obtain a plurality of candidate logs for computing a plurality of layers of an infrastructure. For each of the plurality of candidate logs, the computing system may map the candidate log to a log template of a plurality of log templates, where each log template to which the candidate log is mapped is a mapping log template. The computing system may rank the mapping log templates based on characteristics of the mapping log templates. The computing system may select one or more candidate logs as key logs based on the ranking. The computing system may output at least one of (1) an indication of a key log for determining potential root causes associated with performance issues of the web application, or (2) an indication of potential root causes associated with performance issues of the web application.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the benefit of U.S. Patent Application No. 18 / 400,481, filed Dec. 29, 2023, the entire content of which is incorporated herein by reference. Technical Field

[0003] The present disclosure relates to computing systems, and more particularly to managing web applications operating over a network. Background Art

[0004] In a typical data center environment, there are a large number of interconnected servers that provide computing and / or storage capabilities to run various applications. For example, a data center can include a facility that hosts applications and services for subscribers, i.e., customers of the data center. For example, a data center can host all infrastructure equipment, such as networking and storage systems, redundant power supplies, and environmental controls. In a typical data center, clusters of storage systems and application servers are interconnected via a high-speed switching fabric provided by one or more layers of physical network switches and routers. More complex data centers provide infrastructure worldwide to subscriber support equipment located in various physical hosting facilities.

[0005] Virtualized data centers are becoming the core foundation of modern information technology (IT) infrastructure. In particular, modern data centers have widely utilized virtualized environments in which virtual hosts (also referred to herein as virtual execution elements, such as virtual machines or containers) are deployed and executed on the underlying computing platforms of physical computing devices. Workloads can also include bare-metal processes.

[0006] Virtualization within a data center can provide several advantages. One advantage is that virtualization can provide a significant increase in efficiency. As underlying physical computing devices (i.e., servers) have become more powerful with the emergence of multi-core microprocessor architectures with a large number of cores per physical central processing unit (CPU), virtualization has become simpler and more efficient. A second advantage is that virtualization provides important control over the computing infrastructure. As physical computing resources become fungible resources, such as in cloud-based computing infrastructures, the provisioning and management of the computing infrastructure become simpler. Thus, in addition to the efficiency and increased return on investment (ROI) provided by virtualization, enterprise information technology (IT) staff generally prefer virtualized computing clusters in data centers because of their management advantages. Summary of the Invention

[0007] Generally, techniques for determining critical logs to rule out performance issues of a web application are described. The techniques include a unified framework for analyzing a cross-layer log collection from an application layer, a computing layer, and a network layer of a data center or other computing infrastructure to determine a reduced set of critical logs most relevant to ruling out performance issues of the web application. Examples of performance issues of a web application can include degradation of one or more services provided by the computing infrastructure. In some cases, critical logs can be used to identify possible causes of such performance issues.

[0008] In an example of the described techniques, an analysis system maps cross-layer candidate logs to log templates and determines critical logs from the candidate logs based on characteristics of the log templates. The analysis system utilizes a template mining model trained with historical cross-layer logs to learn to generate log templates to identify cross-layer log scenarios or patterns of cross-layer logs. The analysis system generates log templates for candidate logs based on the learned patterns of cross-layer logs to significantly reduce resources (e.g., memory, processing, etc.) used to determine critical logs. Candidate logs can include cross-layer logs collected during a time period surrounding the occurrence of a performance issue of the web application. Candidate logs can include a cross-layer log collection from nodes in each layer of the computing infrastructure that are identified as being associated with the performance issue of the web application. In some examples, the analysis system identifies candidate logs based on a knowledge graph specifying dependencies of nodes in each layer of the computing infrastructure.

[0009] The analysis system utilizes the template mining model to generate log templates. Log templates can be generated by providing the template mining model with historical log data and / or candidate logs. Log templates may include standard information items (e.g., keywords, identifiers, addresses, variables, etc.) included in cross-layer logs. Multiple candidate logs can be mapped to the same log template. The analysis system generates instances of log templates by mapping candidate logs to log templates. Instances of log templates can include the pattern or standard items of the log template that match the candidate logs, variable data from the candidate logs mapped to the log template, and the timestamp and source (e.g., layer of the computing infrastructure) of the mapped candidate logs.

[0010] The analysis system selects mapping log templates by ranking the mapping log templates according to the characteristics of each mapping log template in the mapping log template. The characteristics of the mapping log template can include keywords included in the log template, the number of instances of the log template (i.e., the number of candidate logs mapped to the log template for analysis runs), and / or whether the trained template mining model learned the log schema of the log template before generating the log template. In some examples, the analysis system assigns categories to the log templates based on these characteristics. The analysis system can select log templates based on the characteristics of each mapping log template in the mapping log template and the timestamps of the log template instances. In some examples, the analysis system selects log templates by calculating a critical template score based on the characteristics of the mapping log templates. The analysis system determines critical logs by selecting one or more instances of the log template from the selected log templates. The analysis system determines critical logs by extracting the corresponding candidate logs mapped to the selected instances of the log template. The analysis system outputs the critical logs or indications thereof to identify the root cause of the performance issues of the network application. In some examples, the analysis system can perform root cause analysis based on the critical logs and output an indication of the potential root cause associated with the performance issues of the network application.

[0011] In some examples, the analysis system considers various types of telemetry data, such as cross-layer logs, cross-layer metrics, and / or network traces. The analysis system determines structured log messages for historical telemetry data (e.g., cross-layer metrics and network traces), and the structured log messages can be provided as training logs, and the template mining model uses the training logs to learn the schema or pattern of various types of telemetry data. The analysis system collects various telemetry data, which may be considered abnormal compared to the baseline behavior. The analysis system includes the structured log messages for the collected telemetry data as candidate logs, and these candidate logs are mapped to the instances of the log templates generated by the trained template mining model. In this way, the analysis system outputs critical logs including various telemetry data (e.g., key performance indicators, network traces, etc.) to more accurately identify the root cause of the performance issues of the network application.

[0012] The techniques of the present disclosure can provide one or more technical advantages, which can enable one or more practical applications. For example, when the performance of a network application degrades, it may be necessary to investigate logs from the underlying infrastructure layer (e.g., compute layer logs, network layer logs, etc.) to determine the possible root cause of the performance issue. However, since each layer associated with the network application is developed and managed independently using different administrative domains and specific domain expertise, investigating each layer can be very inefficient and challenging. An analysis system configured to determine critical logs for each underlying infrastructure layer associated with a network application enables efficient, thorough, and simplified troubleshooting of performance issues of the network application. The critical logs determined by the analysis system significantly reduce the amount of log data that may need to be analyzed for causal analysis to determine the root cause of the performance issue associated with the network application. The analysis system can apply a dependency graph to reduce the number of cross-layer logs selected as candidate logs. The analysis system can generate log templates to categorize candidate logs based on patterns or log schemes, thereby reducing the number of candidate logs selected as critical logs. In this way, the analysis system can reduce the resource consumption and complexity associated with determining the root cause of a network application performance issue by identifying the most relevant cross-layer logs as critical logs for causal analysis for root cause determination.

[0013] In one example, a method includes obtaining, by a computing system, a plurality of candidate logs for a plurality of layers of a computing infrastructure. The method may further include, for each candidate log of the plurality of candidate logs, mapping, by the computing system, the candidate log to a log template of a plurality of logs, wherein each log template to which a candidate log is mapped is a mapped log template. The method may further include ranking, by the computing system, the mapped log templates based on characteristics of each mapped log template of the mapped log templates. The method may further include selecting, by the computing system, one or more candidate logs corresponding to the mapped log templates as critical logs based on the ranking of the mapped log templates. The method may further include outputting, by the computing system, at least one of the following: (1) an indication of the critical logs for determining a potential root cause associated with a performance issue of a network application, or (2) an indication of a potential root cause associated with a performance issue of a network application.

[0014] In another example, a computing system includes processing circuitry that can access a storage device, the processing circuitry being configured to obtain a plurality of candidate logs for a plurality of layers of a computing infrastructure. The processing circuitry can also be configured to, for each candidate log of the plurality of candidate logs: map the candidate log to a log template of a plurality of log templates, where each log template to which the candidate log is mapped is a mapped log template. The processing circuitry can also be configured to rank the mapped log templates based on characteristics of each mapped log template in the mapped log templates. The processing circuitry can also be configured to select, based on the ranking of the mapped log templates, one or more candidate logs corresponding to the mapped log templates as critical logs. The computing circuitry can also be configured to output at least one of the following: (1) an indication of the critical logs for determining a potential root cause associated with a performance issue of a network application, or (2) an indication of a potential root cause associated with a performance issue of the network application.

[0015] In another example, a computer-readable storage medium includes instructions that, when executed, cause processing circuitry to obtain a plurality of candidate logs for a plurality of layers of a computing infrastructure. The instructions can also cause the processing circuitry to, for each candidate log of the plurality of candidate logs: map the candidate log to a log template of a plurality of log templates, where each log template to which the candidate log is mapped is a mapped log template. The instructions can also cause the processing circuitry to rank the mapped log templates based on characteristics of each mapped log template in the mapped log templates. The instructions can also cause the processing circuitry to select, based on the ranking of the mapped log templates, one or more candidate logs corresponding to the mapped log templates as critical logs. The instructions can also cause the processing circuitry to output at least one of the following: (1) an indication of the critical logs for determining a potential root cause associated with a performance issue of a network application, or (2) an indication of a potential root cause associated with a performance issue of the network application.

[0016] Details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a block diagram of an example computing infrastructure that illustrates an example in which the techniques described herein can be implemented.

[0018] Figure 2 is a block diagram of an example computing system in accordance with the techniques described in the present disclosure.

[0019] Figure 3A is a block diagram of an example knowledge graph of a network layer node associated with a network application in accordance with one or more techniques of the present disclosure.

[0020] Figure 3B is a block diagram illustrating an example knowledge graph of network layer nodes associated with performance issues of a web application in accordance with one or more techniques of the present disclosure.

[0021] Figure 4A is a block diagram illustrating an example of an analysis system for determining critical logs in accordance with one or more techniques of the present disclosure.

[0022] Figure 4B is a block diagram illustrating an example of a log analysis engine for determining critical logs in accordance with one or more techniques of the present disclosure.

[0023] Figure 5 is a block diagram illustrating an example of an analysis system for determining critical logs in accordance with one or more techniques of the present disclosure.

[0024] Figure 6 is a flowchart illustrating an example operation for determining critical logs in accordance with one or more techniques of the present disclosure.

[0025] Like reference numerals refer to like elements throughout the description and the drawings. DETAILED DESCRIPTION

[0026] Figure 1 is a block diagram illustrating an example computing infrastructure 100 that can implement examples of the techniques described herein. Generally, a data center 101 provides an operating environment for applications and services of a customer site 104 (illustrated as “Customer 104”) having one or more customer networks coupled to the data center via a service provider network 106. For example, the data center 101 can host infrastructure equipment such as networking and storage systems, redundant power supplies, and environmental controls. The service provider network 106 is coupled to a public network 115, which can represent one or more networks managed by other providers and can thus form part of a large-scale public network infrastructure such as the Internet. For example, the public network 115 can represent a local area network (LAN), a wide area network (WAN), the Internet, a virtual LAN (VLAN), an enterprise LAN, a layer 3 virtual private network (VPN), an Internet protocol (IP) intranet operated by the service provider operating the service provider network 106, an enterprise IP network, or some combination thereof.

[0027] Although customer site 104 and public network 115 are primarily illustrated and described as edge networks of service provider network 106, in some examples, one or more of customer site 104 and public network 115 can be a tenant network within data center 10 or another data center. For example, data center 101 can host multiple tenants (customers), each tenant being associated with one or more virtual private networks (VPNs), and each VPN can be connected to one of the customer sites 104 within customer site 104.

[0028] Service provider network 106 provides packet-based connectivity to the attached customer sites 104, data center 101, and public network 115. Service provider network 106 can represent a network owned and operated by a service provider to interconnect multiple networks. Service provider network 106 can implement Multiprotocol Label Switching (MPLS) forwarding and in such instances can be referred to as an MPLS network or MPLS backbone. In some instances, service provider network 106 represents multiple interconnected autonomous systems providing services from one or more service providers, such as the Internet.

[0029] In some examples, data center 101 can represent one of many geographically distributed data centers in which the techniques and systems described herein can be implemented. As Figure 1 illustrated by the example of, data center 101 can be a facility that provides network services, cloud services, storage services, and / or application services to customers. Data center 101 can represent an on-premises data center, private cloud, public cloud, hybrid cloud, or other type of deployment. The customers of the service provider can be collective entities such as enterprises and governments or individuals. For example, a data center can host network services for several enterprises and end users. Other exemplary services can include data storage, virtual private networks, business engineering, file services, data mining, scientific computing, or supercomputing, etc. Although illustrated as a separate edge network of service provider network 106, elements of data center 101 (such as one or more physical network functions (PNFs) or virtualized network functions (VNFs)) can be included within the core of service provider network 106.

[0030] The switching fabric 121 may include interconnected top-of-rack (TOR) (or other "leaf") switches 16A through 16N (hereinafter "TOR switches 16"), which are coupled to the distribution layer of the rack (or "spine" or "core") switches 18A through 18N (hereinafter "rack switches 18"). The data center 101 may include a gateway 108. For example, the gateway 108 may include one or more non-edge switches, routers, hubs, gateways, security devices such as firewalls, intrusion detection and / or prevention devices, servers, computer terminals, laptops, printers, databases, wireless mobile devices such as cellular phones or personal digital assistants, wireless access points, bridges, cable modems, application accelerators, or other network devices. The data center 101 may also include one or more physical network functions (PNFs), such as physical firewalls, load balancers, routers, route reflectors, broadband network gateways (BNGs), evolved packet cores, or other cellular network elements or other PNFs.

[0031] The terms "packet flow", "traffic flow", or simply "flow" refer to a collection of packets that originate from a specific source device or endpoint and are sent to a specific destination device or endpoint. A single flow of packets can be identified by a 5-tuple: for example <source network address, destination network address, source port, destination port, protocol>. This 5-tuple typically identifies the packet flow corresponding to the received packets. An n-tuple refers to any n items drawn from the 5-tuple. For example, a 2-tuple for a packet can refer to a combination of <source network address, destination network address> or <source network address, source port> for the packet.

[0032] Any server in the data center 101 can be configured with a workload by virtualizing the resources of the server to provide isolation between one or more processes (applications) executing on the server. "Hypervisor-based" or "hardware-level" or "platform" virtualization refers to creating virtual machines that each include a guest operating system for executing one or more processes. Typically, the virtual machine provides the virtualized / guest operating system to execute applications in an isolated virtual environment. Since the virtual machines are virtualized from the physical hardware of the host server, the execution of applications is isolated from the host's hardware and other virtual machines. Each virtual machine can be configured with one or more virtual network interfaces to communicate on a corresponding virtual network.

[0033] "Container-based" or "operating system" virtualization refers to virtualizing an operating system to run multiple isolated systems on a single machine (virtual or physical). Such isolated systems represent containers, such as those provided by the open-source DOCKER container application or by CoreOS Rkt ("Rocket"). Like virtual machines, each container is virtualized and can be kept isolated from the host machine and other containers. However, unlike virtual machines, each container may omit a separate operating system and only provide an application suite and application-specific libraries. Generally, containers are executed by the host machine as isolated user-space instances and can share the operating system and common libraries with other containers executing on the host machine. Thus, compared to virtual machines, containers may require fewer processing capabilities, storage, and network resources. A group of one or more containers can be configured to share one or more virtual network interfaces for communication on a corresponding virtual network.

[0034] In some examples, containers are managed by their host kernel to allow for resource (CPU, memory, block I / O, network, etc.) limits and prioritization without the need to start any virtual machines. In some cases, namespace isolation functionality is used that allows for complete isolation of the application (e.g., a given container) view of the operating environment, including process trees, networking, user identifiers, and the mounted file system. In some examples, containers can be deployed according to Linux Containers (LXC), which is an operating system-level virtualization method for running multiple isolated Linux systems (containers) on a control host using a single Linux kernel. LXC is an operating system-level virtualization method for running multiple isolated Linux systems (containers) on a single control host (LXC host). LXC does not use virtual machines (although LXC can be hosted by a virtual machine). Instead, LXC uses a virtual environment with its own CPU, memory, block I / O, network, and / or other resource space. The LXC resource control mechanism is provided by namespaces and cgroups in the Linux kernel on the LXC host. Additional examples of containerization methods include OpenVZ, FreeBSD jails, AIX workload partitions, and Solaris containers. Thus, as used herein, the term "container" can cover not only LXC-type containers, but also any one or more of virtualization engines, virtual private servers, silos, or jails.

[0035] In Figure 1In the example, data center 101 includes storage and / or computing servers interconnected by one or more layers of physical network switches and routers, where computing nodes 110A through 110N (referred to herein as "computing nodes 110") are depicted as interconnected via TOR switches 16 and rack switches 18. Computing nodes 110 can be bare-metal machines and / or virtual machines within data center 101 and can also be referred to herein as "hosts" or "host devices". Computing nodes 110 can represent computing devices configured to operate according to the techniques described herein, such as x86 processor-based servers.

[0036] Computing nodes 110 can host virtual network endpoints for one or more virtual networks, which operate over the physical network provided by TOR switches 16 and rack switches 18. Although described primarily with respect to data center-based switching networks, other physical networks such as service provider network 7 can serve as the basis for one or more virtual networks.

[0037] Each of computing nodes 110 can host one or more workloads. The term "workload" encompasses virtual machines, containers, Kubernetes Pods, and other virtualized computing resources that provide at least partially isolated execution environments for applications. As Figure 1 shown, computing node 110A hosts a workload implementing service 122A. However, given the hardware resource limitations of computing nodes 110, computing nodes 110 can execute as many workloads as possible.

[0038] Computing infrastructure 100 implements an automation platform to automate the deployment, scaling, and operation of workloads on computing nodes 110 to provide a virtualized infrastructure for executing application workloads and services. In some examples, the platform can be a container orchestration platform that provides a container-centric infrastructure, which automates the deployment, scaling, and operation of containers to provide a container-centric infrastructure. In the context of virtualized computing infrastructure, "orchestration" generally refers to the provisioning, scheduling, and management of workloads and / or applications and services executing on such workloads to host servers available to the orchestration platform. Specifically, container orchestration enables container coordination and refers to, for example, the deployment, management, scaling, and configuration of containers to host servers via a container orchestration platform. Example instances of orchestration platforms include Kubernetes, Docker swarm, Mesos / Marathon, OpenShift, OpenStack, VMware, and Amazon ECS.

[0039] The components of the automation platform of the computing infrastructure 100 at least include computing nodes 110, network controllers 124, an orchestrator 130, and an analytics system 140. A cluster-based framework can be used to deploy workloads to a virtualized environment, where the cluster master node of the cluster manages the deployment and operation of containers to one or more cluster slave nodes of the cluster. The terms "master node" and "slave node" as used herein cover different orchestration platform terms for similar devices, which distinguish the main management components of the cluster and the main workload hosting devices of the cluster. For example, the Kubernetes platform uses the terms "cluster master" and "slave node", while the Docker Swarm platform refers to the cluster manager and cluster nodes.

[0040] Generally, the network controller 124 controls the network configuration of the data center 101 fabric to, for example, establish one or more virtual networks for packetized communication between virtual network endpoints. The network controller 124 provides a controller that is logically and in some cases physically centralized to facilitate the operation of one or more virtual networks within the data center 101. In some examples, the network controller 124 can operate in response to configuration inputs received from the orchestrator 130 and / or an administrator / operator. Additional information regarding example operations of an example network controller 124 operating in conjunction with other devices of the data center 101 or other software-defined networks can be found in International Application No. PCT / US2013 / 044378, filed on June 5, 2013, and titled "PHYSICAL PATH DETERMINATION FOR VIRTUAL NETWORK PACKET FLOWS"; and U.S. Patent Application No. 14 / 226,509, filed on March 26, 2014, and titled "TUNNELED PACKET AGGREGATION FOR VIRTUAL NETWORKS", each of which is incorporated herein by reference as if fully set forth herein.

[0041] The orchestrator 130 controls the deployment, scaling, and operation of workloads on a server cluster and provides a computing infrastructure, which can include a container-centric computing infrastructure. The orchestrator 130 and in some cases the network controller 124 can implement the respective cluster masters of one or more Kubernetes clusters. As an example, Kubernetes is a container management platform that provides portability across public and private clouds, each of which can provide virtualization infrastructure to the container management platform. The orchestrator 130 can represent any of the orchestration platforms listed above, such as Kubernetes.

[0042] In Kubernetes, by default, all workloads can communicate with all other workloads without using Network Address Translation (NAT). In some cases, the orchestrator 130 and the network controller 124 create a service virtual network and a workload virtual network shared by all namespaces, and allocate service and workload network addresses from them respectively. In some cases, all workloads in all namespaces spawned in a Kubernetes cluster may be able to communicate with each other, and the network addresses of all workloads can be allocated from the workload subnet specified by the orchestrator 130. When the user creates an isolated namespace for a workload, the orchestrator 130 and the network controller 124 can create a new workload virtual network and a new shared service virtual network for the new isolated namespace. Workloads in the isolated namespace spawned in the Kubernetes cluster draw network addresses from the new workload virtual network, and the corresponding services of such workloads draw network addresses from the new service virtual network.

[0043] As part of the process of creating a workload, the orchestrator 130 can request the network controller 124 to create corresponding virtual network interfaces for one or more virtual networks (indicated in the configuration data). For each virtual network to which the workload belongs, the workload can have a different virtual network interface. The network controller 124 processes the request to generate interface configuration data for the virtual network interfaces of the workload. The interface configuration data can include a container, pod, or other workload unique identifier and a list or other data structure specifying the network configuration data for each virtual network interface in the virtual network interfaces for configuring the virtual network interfaces. The network configuration data of the virtual network interfaces can include network names, assigned virtual network addresses, MAC addresses, and / or domain name server values.

[0044] Each of services 122A through 122N (collectively referred to as "services 122") is deployed using a workload. A service 122 may separately represent or include one or more containers deployed by a container orchestration system. One or more of the services 122 may together implement a network application, which includes a collection of one or more of the services 122. For example, a network application may include services 122A through 122N. Each of the services 122 may provide or implement one or more services, and in the case where a service 122 represents a pod or other container deployment, the one or more services are containerized services or "microservices". Compute node 110 may host services for multiple different network applications, with each network application having one or more services distributed thereamong. In some examples, the services of a network application are distributed across compute nodes managed by any combination of service providers, enterprises, or other entities. Such compute nodes may be located in multiple different data centers, on-premises or private, public, or hybrid clouds.

[0045] Orchestrator 130 may include a scheduler to schedule services 122 to compute node 110. Generally, orchestrator 130 may manage the placement of each of the services 122 within services 122 to compute node 110 according to a scheduling policy, the amount of resources requested by the service, and the available resources of compute node 110. When assigning a service 122 to compute node 110, the compute node resources considered by orchestrator 130 include CPU-related resources (such as cores, CPU / core utilization), memory-related resources (available main memory, e.g., 2GB), temporary storage devices, and user-defined extended resources. In Kubernetes, the scheduler is referred to as the kube scheduler.

[0046] Services 122 may have different performance requirements that need to be met within a highly dynamic application execution environment. In such an environment, application performance is a product of the dynamics of different resources such as worker node resources; network resources (e.g., bandwidth, latency, loss, jitter, firewall policies), network policies, and the communication graph between different services of a network application; and the performance of external services such as authentication and external cloud services.

[0047] As part of providing functionality to a network application, services 122 may communicate with each other. Each of the services 122 may provide functionality for one or more components of a network application. For example, service 122A may provide functionality for one part of a network application, while service 122N provides functionality for a different part of the network application.

[0048] Services 122 can communicate with each other using invocations such as remote procedure calls (RPCs). Services 122 can communicate with each other along an RPC chain to provide functionality for a network application. For example, service 122A can communicate with service 122N and send an RPC to service 122N as part of providing functionality for a network application.

[0049] Services 122 can call each other in the path of a service invocation. For example, service 122B can call service 122C, and then service 122C can call service 122F. As part of providing functionality for a network application, a series of services can call each other in sequence. In some cases, service 122A is an entry endpoint service, service 122N is a termination endpoint service, and one or more other services are called between service 122A and service 122N for an end-to-end call path of a network application.

[0050] Service requests that reach an entry point (also known as an endpoint) in a distributed system go through multiple "hops" via multiple microservice operations before being fully serviced. The lifespan of a request results in complex microservice interactions. These interactions are deeply nested, asynchronous, and invoke many other downstream operations. Due to this complexity, it can be difficult to identify which (if any) underlying services contribute to the overall end-to-end latency experienced by a top-level request.

[0051] Analysis system 140 can be executed on one or more devices of data center 101 as a network administrator application. However, analysis system 140 can be deployed separately from data center 101.

[0052] Analysis system 140 can be integrated as part of a telemetry system, a root cause determination system, or any system that a network administrator can implement to analyze the log data of computing infrastructure 100. In Figure 1 the example, analysis system 140 includes a log analysis engine 146, a log collector tool 144, and a log database 142.

[0053] According to the techniques described herein, the analysis system 140 determines a set of critical logs from cross-layer system logs of multiple layers of the computing infrastructure 100. The cross-layer system logs can include logs from different layers of the computing infrastructure 100, such as logs from the application layer (e.g., service 122), logs from the computing layer (e.g., computing node 110), and logs from the network layer (e.g., top-of-rack switch 18 and TOR switch 16). The analysis system 140 can output critical logs composed of cross-layer logs for root cause analysis of performance issues related to network applications. Network applications can include any combination of services or microservices in a data center that rely on network resources to perform specific functions, such as enabling communication, data sharing, and collaboration between network devices. In Figure 1 the example of, network applications can include any combination of services 122A through 122N or service 122. The log analysis engine 146 of the analysis system 140 can obtain cross-layer logs associated with service 122 to output critical logs using a machine learning model, such as model service 154. Model service 154 can include a machine learning model (e.g., template mining model) trained by model trainer 152 to identify patterns or schemes in cross-layer logs based on historical log data. Although illustrated as internal to the log analysis engine 146, model service 152 can include a template mining model trained offline on an external computing system or computing device.

[0054] Model trainer 152 can obtain historical log data for each layer of the computing infrastructure 100 to train the machine learning model of model service 154 to identify patterns in cross-layer logs. For example, model trainer 152 can obtain historical log data for the application layer, computing layer, and network layer of the computing infrastructure 100 to train the machine learning model into a trained template mining model that identifies schemes of keywords, functions, addresses, variables, identifiers, or other information included in different types of system logs.

[0055] Model trainer 152 can generate log templates based on historical log data. The log templates include common schemes output by the trained template mining model, such as an overview or representation of critical terms, functions, variables, identifiers, addresses, or other standard information included in system logs from multiple layers of a data center. Model trainer 152 can provide the multiple generated log templates to model service 154. Model service 154 can map candidate logs to the log templates to reduce the number of data items to be analyzed.

[0056] In some instances, model service 154 can execute the trained template mining model to generate log templates. In some instances, model service 154 can store log templates generated by the trained machine learning model.

[0057] The log analysis engine 146 can apply the model service 154 to map candidate logs to log templates. The log analysis engine 146 can obtain cross-layer system logs as candidate logs, and the candidate logs can be included in the critical log set. The model service 154 of the log analysis engine 146 can apply the trained template mining model to generate multiple log templates for the candidate logs. The model service 154 can generate an instance of the log template by mapping the candidate logs to the log template among the multiple log templates. The model service 154 can map multiple candidate logs to the log template based on whether the candidate logs match the patterns or scenarios included in the log template. Each log template to which one or more candidate logs are mapped is a mapped log template. The instance of the log template includes the scenario identified by the trained template mining model, as well as the timestamp of the candidate logs mapped to the log template and the source of the candidate logs (such as the application layer, the computing layer, the network layer, etc.). The log analysis engine 146 maps the candidate logs to the log template to reduce the number of data items analyzed to determine the critical log set. In some instances, the model trainer 152 can retrain the template mining model of the model service 154 with the candidate logs for future log template generation and critical log determination.

[0058] The log analysis engine 146 determines critical logs by selecting one or more mapped log templates. To determine critical logs from the mapped log templates, the log analysis engine 146 applies heuristics to the characteristics of the log templates. Example heuristics include considering temporal recency in view of various log template categories and critical template scores. Both of these heuristics involve the log analysis engine 146 selecting one or more mapped log templates based on the category assigned to each log template and other factors (such as the timestamp included in the instance of the log template or the keywords included in the mapped log template). These will be described in further detail below with respect to Figures 4A to 4B be described in further detail.

[0059] The log analysis engine 146 can determine which candidate logs are critical logs based on the characteristics of each mapped log template in the mapped log templates. The characteristics of the mapped log template can include the keywords included in the log template, the number of instances of the log template, and / or whether the log analysis engine learned the log scenario of the log template before generating the log template. The log analysis engine 146 determines critical logs based on the mapping of the candidate logs to the selected log templates. For example, the log analysis engine can extract the corresponding candidate logs from one or more instances of the selected log template as the candidate log set used in the root cause analysis of network application performance problems.

[0060] In some instances, the log analysis engine 146 can determine potential root causes of performance issues of a network application based on a set of critical logs and output an indication of the potential root causes. In some examples, the log analysis engine 146 can output an indication of the critical logs to an external computing system configured to perform causal analysis for determining the root cause of performance issues of the network application.

[0061] The log collector tool 144 can obtain cross-layer system logs associated with the computing infrastructure 100. For example, the log collector tool 144 can obtain logs from the application layer associated with the service 122A, the computing layer associated with the computing node 110A, and the network layer associated with the rack switch 18 and the TOR switch 16A.

[0062] The log collector tool 144 can include software tools (such as FluentBit and FluentD) configured to collect, filter, format, and annotate logs and metrics from multiple sources (such as the application layer, the computing layer, and the network layer). The log collector tool 144 can ingest and process system logs from each layer of the computing infrastructure 100. For example, the log collector tool 144 can add metadata (such as annotations) to each log to identify the time and source of the log data included in the log. The log collector tool 144 can append a corresponding timestamp of the log generation time before each collected log. The log collector tool 144 can attach each collected log to the corresponding source or layer where the log is collected. The log collector tool 144 can store the ingested and processed logs in the log database 142. The log collector tool 144 can store the logs on a specific index of the log database 142.

[0063] The log database 142 can include a database maintained by the analysis system 140 and stored to a storage medium. The analysis system 140 can store data regarding system logs, performance metrics, traces, etc. associated with the computing infrastructure 100. The analysis system 140 can store one or more maps or graphs of dependencies of layer nodes and network configurations in the log database 142. The log database 142 can include an OpenSearch database or other types of databases that may be capable of querying various types of data.

[0064] The log analysis engine 146 can obtain candidate logs from the log database 142. In some instances, the log analysis engine 146 can obtain candidate logs when generating each system log within a specific time period (e.g., a thirty - minute time period that results in performance issues with service 122A). In some instances, the log analysis engine 146 can use a knowledge graph that specifies particular nodes from each layer of the computing infrastructure 100 associated with a performance issue of a web application to obtain candidate logs as a reduced set of system logs generated over a period of time.

[0065] The techniques of the present disclosure can provide one or more technical advantages. For example, by identifying critical logs and thereby reducing the number of logs that must be analyzed, the log analysis engine 146 can help efficiently determine potential root causes of performance issues of a web application based on candidate logs. The log analysis engine 146 can identify candidate logs in a low - data manner from a large number of collected system logs from multiple layers of the computing infrastructure 100. The log analysis engine 146 can reduce the number of candidate logs to be analyzed by applying a knowledge graph that identifies nodes (e.g., any service 122 of the application layer, computing nodes 110 of the computing layer, and TOR switches 16 and rack switches 18 of the network layer) associated with a performance issue of a web application. The log analysis engine 146 can train and apply a machine - learning model to classify each candidate log into a log template, thereby categorizing the candidate logs based on patterns of a log schema. In this way, the log analysis engine 146 can select one or more candidate logs based on the ranking or scoring of the log templates rather than ranking or scoring each individual candidate log. By performing a causal analysis by the log analysis engine 146 or an external system to determine potential root causes of performance issues associated with a web application having the selected candidate logs that have been determined to be critical, the log analysis engine 146 can reduce the computational time required to effectively analyze candidate logs and reduce the amount of resources (e.g., processing power, memory storage, etc.) associated with determining the root cause of a performance issue of a web application.

[0066] Figure 2 is a block diagram illustrating an example computing system in accordance with the techniques described in the present disclosure. The computing system implements Figure 1Example instance of the analysis system 240. The computing system 202 can be implemented as any suitable computing system, such as one or more server computers, workstations, mainframes, appliances, cloud computing systems, and / or other computing systems that may be capable of performing the operations and / or functions described in accordance with one or more aspects of the present disclosure. In some examples, the computing system 202 represents a cloud computing system, a server farm, and / or a server cluster (or a portion thereof) that provides services to other devices or systems. In other examples, the computing system 202 may represent one or more virtualized computing instances (e.g., virtual machines, containers) of a cloud computing system, a server farm, a data center, and / or a server cluster or be implemented there through.

[0067] In Figure 2 the example, the computing system 202 may include one or more processors 213, a communication unit 215 (illustrated as COMM.215), one or more input devices 217, one or more output devices 218, and one or more storage devices of the storage system 205. The storage system 205 includes the analysis system 240 and the orchestrator 130. The communication channel 212 may interconnect one or more of the devices, modules, storage areas, or other components of the computing system 202 to (physically, communicatively, and / or operably) enable communication between the components. In some examples, the communication channel 212 may represent one or more of a system bus, a network connection, an interprocess communication data structure, or any other method for transferring data.

[0068] One or more of the (multiple) processors 213 may implement the functionality associated with the computing system 202 or with one or more of the modules illustrated herein and / or described below and / or execute instructions. One or more of the (multiple) processors 213 may be a processing circuitry that performs operations in accordance with one or more aspects of the present disclosure, may be a part of it, and / or may include the processing circuitry. Examples of the (multiple) processors 213 include microprocessors, application processors, display controllers, auxiliary processors, one or more sensor hubs, and any other hardware configured to function as a processor, a processing unit, or a processing device. The computing system 202 may use one or more processors 213 to perform operations in accordance with one or more aspects of the present disclosure using software, hardware, firmware, or a combination of hardware, software, and firmware residing in and / or executing at the computing system 202.

[0069] One or more communication units 215 of computing system 202 may communicate with devices external to computing system 202 by sending and / or receiving data, and in some aspects may operate as both an input device and an output device simultaneously. In some examples, communication unit 215 may communicate with other devices via a network. In other examples, communication unit 215 may send and / or receive radio signals over a radio network such as a cellular radio network. In other examples, communication unit 215 of computing system 202 may send and / or receive satellite signals over a satellite network. Examples of communication unit 215 include network interface cards (such as Ethernet cards), optical transceivers, radio frequency transceivers, GPS receivers, or any other type of device that can send and / or receive information. Such communication may conform to, implement, or comply with appropriate protocols, including Transmission Control Protocol / Internet Protocol (TCP / IP), Ethernet, or other technologies or protocols.

[0070] One or more input devices 217 may represent any input device of computing system 202 not separately described herein. Input device 217 may generate, receive, and / or process input. For example, one or more input devices 217 may generate or receive input from a network, a user input device, or any other type of device for detecting input from a human or a machine.

[0071] One or more output devices 218 may represent any output device of computing system 202 not separately described herein. Output device 218 may generate, present, and / or process output. For example, one or more output devices 218 may generate, present, and / or process output in any form. Output device 218 may include one or more USB interfaces, video and / or audio output interfaces, or any other type of device capable of generating tactile, audio, visual, video, electrical, or other output. Some devices may function as both input and output devices simultaneously. For example, a communication device may send data to other systems or devices via a network and may receive data from other systems or devices.

[0072] One or more storage devices of the storage system 205 within the computing system 202 may store information for processing during operation of the computing system 202. The storage system 205 may store program instructions and / or data associated with one or more modules described in accordance with one or more aspects of the present disclosure. One or more processors 213 and one or more storage devices may provide an operating environment or platform for such modules, which may be implemented as software, but in some examples may include any combination of hardware, firmware, and software. One or more processors 213 may execute instructions, and one or more storage devices of the storage system 105 may store instructions and / or data for one or more modules. The combination of the processor 213 and the storage system 205 may retrieve, store, and / or execute instructions and / or data for one or more applications, modules, or software. The processor 213 and / or the storage devices of the storage system 205 may also be operatively coupled to one or more other software and / or hardware components, including but not limited to one or more components of the components of the computing system 202 and / or one or more devices or systems illustrated as being connected to the computing system 202.

[0073] The processor 213 may execute the analytics system 240. The analytics system 240 may be an application, platform, or other form type of process configured to perform analytics. For example, the analytics system 240 may monitor the performance of various network applications and underlying services (such as Figure 1 the service 122 illustrated in

[0074] In accordance with the techniques described herein, as part of a monitoring service for network applications, the analytics system 240 may determine a critical log set for root cause analysis of performance issues for network applications. The log analysis engine 246 of the analytics system 240 may train a machine learning model to identify standard information patterns in cross-layer logs for log template generation. The model trainer 252 may train the machine learning model with historical log data 262. The historical log data 262 may include cross-layer system logs during normal system performance and / or candidate logs obtained during a previous critical log determination. The model trainer 252 may generate a log template using a template mining model. In some examples, the model trainer 252 may provide the log template to the model service 254 to map to candidate logs. Although illustrated as being within the computing system 202, the machine learning model may be trained offline at an external system.

[0075] The model service 254 of the log analysis engine 246 may include a machine learning model trained by the model trainer 252. The log template classifier 264 of the model service 254 may include an indication of the frequency or number of instances of a log template that have been mapped to candidate logs. The log template classifier 264 may include an indication of a log template instance that includes the pattern of the log template and the corresponding timestamp and source of the candidate logs mapped to the log template. The log template classifier 264 may store an indication of keywords and the importance of the keywords (e.g., the keyword "error" will receive a high importance value or the keyword "obtained" will receive a low importance value). The log template classifier 264 may store log templates generated during the training phase or during a previous critical log determination. The model service 254 may apply the indications of the stored log template classifier 264 to assign a category to a log template generated for candidate logs.

[0076] The anomaly detection engine 258 of the analysis system 240 may determine a performance issue of a network application. The anomaly detection engine 258 may determine a performance issue (e.g., a decrease in response time or latency) of a network application such as Figure 1 service 122). For example, the anomaly detection engine 258 may determine a performance issue of a network application based on performance data included in the performance metrics 238. The performance metrics 238 may include key performance indicators (KPIs) monitored for each layer of the computing infrastructure 100. For example, the performance metrics 238 may include application KPIs, computing KPIs, and network KPIs that are configured to measure latency, jitter, packet loss, throughput, network speed, bandwidth, network availability, packet replication, packet reordering, packet rate, interface flapping, processor utilization, memory utilization, user quality of experience, network congestion, round-trip time, network utilization, network error rate, network response time, etc. in each layer of the computing infrastructure 100. The anomaly detection engine 258 may determine a performance issue based on the thresholds of the KPIs included in the performance metrics 238.

[0077] In some instances, the anomaly detection engine 258 can determine performance issues that are not triggered by the KPI values included in the performance metrics 238. For example, the anomaly detection engine 258 can determine performance issues of a web application based on feedback from users or administrators. The anomaly detection engine 258 can generate an anomaly log based on the KPIs included in the performance metrics 238 to determine whether additional or different KPIs of a specific layer should be measured to detect performance issues. The anomaly detection engine 258 can generate an anomaly log by converting the time series data of the performance metrics 238 into certain message events. The anomaly detection engine 228 can provide the anomaly log to the log analysis engine 246. The log analysis engine 246 can apply the model trainer 252 to the historical log data 262 including historical anomaly logs to train the machine learning model of the model service 254, thereby generating a template for the anomaly log. The log analysis engine 246 can apply the model service 254 to determine whether an anomaly log is included as part of a critical log set. In this way, the log analysis engine 246 can output critical logs that concurrently consider different types of telemetry, such as logs, performance metrics, and traces, to help identify the potential causes of performance issues with higher accuracy.

[0078] The dependency mapping generator service (DMGS) 256 can include a knowledge graph that identifies the dependent nodes of each layer associated with a web application. The DMGS 256 can include an application tracing tool or toolkit (e.g., Jaegar or OpenTelemetry). The DMGS 256 can detect a web application including the service 122 and obtain call path information for a given time window. The DMGS 256 can use the call path information to determine the call paths between the services 122. The DMGS 256 can map or draw nodes from the application layer to other nodes of the application layer and / or nodes from the compute layer. The DMGS 256 can map or draw nodes from the compute layer to nodes of the application layer and / or the network layer. The DMGS 256 can map or draw nodes from the network layer to nodes of the application layer or other nodes of the network layer. The log analysis engine 246 can apply the knowledge graph included in the DMGS 256 to reduce the number of candidate logs from the perspective of the anomalous application layer based on the dependent nodes. For example, the log analysis engine 246 can determine candidate logs based on the nodes associated with one or more nodes (e.g., microservices) of the application that has experienced performance issues in the knowledge graph.

[0079] The log collector tool 244 may include software tools for collecting telemetry data from various sources across the computing infrastructure 100. The log collector tool 244 may collect telemetry data such as logs, traces, and performance metrics. The log collector tool 244 may store time series data associated with KPIs for each layer in the performance metrics 238. The log collector tool 244 may index the logs for each layer from the log database 242. For example, the log collection tool 244 may index the logs from the application layer in the application logs 232 of the log database 242. The log collection tool 244 may index the logs from the compute layer in the compute logs 234. The log collection tool 244 may index the logs from the network layer in the network logs 236.

[0080] The causality analysis service 266 may apply the critical logs determined by the log analysis engine 246 to determine potential root causes of performance issues. For example, the causality analysis service 266 may utilize the critical logs to implement a form of Granger causality analysis to determine potential root causes of performance issues in a network application. The causality analysis service 266 may efficiently determine potential root causes due to the critical logs, which include a reduced set of relevant logs on each layer associated with the network application experiencing performance issues.

[0081] The user interface (“UI”) 268 may generate a user interface including one or more visual elements. For example, the UI 268 may generate a user interface including one or more visual elements, each visual element associated with one or more services and network devices. In another example, the UI 268 may generate a user interface including a visual representation of the DAG of the network infrastructure. In yet another example, the UI 268 may generate a user interface that includes visual elements associated with an alert indicating that an underlying network device of a critical path is experiencing abnormal behavior.

[0082] The analytics system 240 may output the user interface via the output device 218. The analytics system 240 may output the user interface generated by the UI 268 for display to the user. For example, the UI 268 may generate a user interface that causes the output device 218 to display a user interface including a visual indicator of an alert regarding a device experiencing abnormal behavior. In another example, the UI 268 may generate a user interface that includes a visual representation of the DAG of the calls between services and the underlying network infrastructure. The UI 268 may generate a user interface that prompts the user to define a time period for candidate log determination, KPIs to monitor, knowledge graphs, or other aspects of the techniques described herein.

[0083] Figure 3Ais a block diagram illustrating an example knowledge graph of network layer nodes associated with a network application in accordance with one or more techniques of the present disclosure. For example purposes only, with respect to Figure 2 discussed Figure 3A . The analysis system 240 may maintain a knowledge graph that includes a mapping of cross-layer dependencies of application layer nodes, compute layer nodes, and network layer nodes in a data center environment (e.g., Figure 1 data center 101). In the Figure 3A example, the analysis system 240 may determine a knowledge graph (e.g., an application dependency graph) that includes service nodes 302A through 302G (collectively referred to herein as "service nodes 302" or "application nodes 302"), compute nodes 304A through 304H (collectively referred to herein as "compute nodes 304"), TOR switches 306A through 306D (collectively referred to herein as "TOR switches 306"), and spine switches 308A and 308B (collectively referred to herein as "spine switches 308").

[0084] The analysis system 240 may determine a knowledge graph having an application layer that includes service nodes 302, a compute layer that includes compute nodes 304, and a network layer that includes TOR switches 306 and spine switches 308. In some instances, service nodes 302 may correspond to multiple distributed services or network applications. In some examples, each service node of service nodes 302 may include multiple instances hosted on distributed compute nodes 304 in the compute layer of the knowledge graph. Each compute node 304 of compute nodes 304 may include a bare metal server, a computing device, a virtual machine, etc. Compute nodes 304 may be connected to a TOR switch of TOR switches 306 in the network layer. TOR switches 306 may be coupled to one or more spine switches of spine switches 308 in the network layer.

[0085] In accordance with the techniques described herein, the analysis system 240 may obtain candidate logs for each node of the knowledge graph. The analysis system 240 may obtain candidate logs of the nodes of the knowledge graph in response to determining a performance issue, anomaly, or unexpected behavior of one or more service nodes 302. In the Figure 3AIn the example, the analysis system 240 (or more specifically, the anomaly detection engine 258) can detect performance issues with the service nodes 302A, 302B, and 302C (e.g., unexpected behavior of the key performance indicator metrics of the service nodes 302A, 302B, 302C). The analysis system 240 can obtain cross-layer logs from each node for a period of time before detecting a performance issue associated with the service nodes 302A, 302B, 302C. For example, the analysis system 240 can determine a performance issue associated with the service nodes 302A, 302B, 302C at a time equal to "2:00". The analysis system 240 can obtain logs for each node with a timestamp indicating that the log was generated during the period between "1:30" and "2:00" (and including "1:30" to "2:00").

[0086] The analysis system 240 can apply the log collection tool 244 to obtain the system logs of each node of the knowledge graph. The analysis system 240 can apply the log collection tool 244 to obtain application logs for the service nodes 302 from the application layer log collector, the compute nodes 304 from the compute layer log collector, the TOR switches 306 from the network layer log collector, and the top-of-rack switches 308 from the network layer log collector.

[0087] The analysis system 240 can apply the log collection tool 244 to ingest and process the system logs obtained from the telemetry collector to cross-layer filter, format, and annotate the log data. For example, the analysis system 240 can process the application layer logs (e.g., logs for the service nodes 302) by adding or appending log data with an annotation such as "Application:Istio", indicating that the log corresponds to the application layer logs collected using the Istio service mesh. In another example, the analysis system 240 can use the application performance monitoring tool of the log collection tool 244 to collect logs and annotate the log data with "Application:APM:NewRelic" as the source of the log data. The analysis system 240 can annotate the network log data with source annotations such as "Network:Physical:Apstra" for the physical network managed by the Juniper Apstra fabric manager or "Network:Virtual:Contrail" for the virtual network monitored by Juniper Contrail. The analysis system 240 can append a corresponding timestamp corresponding to the log generation time before each log. The analysis system 240 can store the ingested cross-layer logs as candidate logs in the log database 242 for analysis by the log analysis engine 246 to determine the critical logs.

[0088] Figure 3BFIG. 0 is a block diagram illustrating an example knowledge graph of network layer nodes associated with performance issues of a web application in accordance with one or more techniques of the present disclosure. Analysis system 240 may prune or limit the knowledge graph used to obtain candidate logs to reduce the number of candidate logs stored in log database 242 and processed by log analysis engine 246. Analysis system 240 may apply anomaly detection engine 258 to determine performance issues (e.g., latency) associated with service nodes 302A, 302B, and 302C. Analysis system 240 may apply dependency mapping generator service 256 to prune the knowledge graph from the perspective of the anomalous application layer nodes (e.g., Figure 3A the knowledge graph) (e.g., generate a subgraph to include nodes associated with service nodes experiencing performance issues). In Figure 3B an example, analysis system 240 may prune Figure 3A the knowledge graph to include service nodes 302A, 302B, 302C, compute nodes 304A, 304B, 304C, TOR switches 306A, 306B, and top-of-rack switches 308A, 308B. Analysis system 240 may collect, filter, format, and annotate logs based on the nodes included in the Figure 3B pruned knowledge graph. Analysis system 240 may store the ingested cross-layer logs with a reduced number of nodes as candidate logs. In this way, log analysis engine 246 may analyze a smaller number of candidate logs compared to the number of candidate logs ingested according to Figure 3A the knowledge graph without affecting the accuracy of the resulting critical logs.

[0089] Figure 4A FIG. 14 is a block diagram illustrating an example of an analysis system for determining critical logs in accordance with one or more techniques of the present disclosure. Figure 4A It may be described with respect to Figure 2 merely for example purposes. In Figure 4A an example, application logs 432, compute logs 434, network logs 436, log analysis engine 446, and causal relationship analysis service 466 may include Figure 2 example implementations of application logs 232, compute logs 234, network logs 236, log analysis engine 246, and causal relationship analysis service 266 of

[0090] In some examples, the log analysis engine 446 can determine critical logs 470 in response to an indication of a performance degradation of a web application. The log analysis engine 446 can determine a subset of cross-application logs 432, compute logs 434, and network logs 436 as critical logs 470 based on the techniques described herein. The critical logs 470 can include the most relevant logs from each of the application logs 432, compute logs 434, and network logs 436 to efficiently and effectively troubleshoot performance issues of the web application. In some instances, the log analysis engine 446 can determine candidate logs from each of the application logs 432, compute logs 434, and network logs 436 as logs for each node from a knowledge graph or as a reduced set of logs for each node identified as being associated with a performance issue of the web application.

[0091] The log analysis engine 446 can apply a template mining model to generate log templates for candidate logs. In some instances, the log analysis engine 446 can generate log templates during the training of the template mining model. The log analysis engine 446 can map candidate logs to log templates generated by providing the candidate logs to the trained template mining model. The log analysis engine 446 maps candidate logs to log templates to reduce the number of logs to be searched, collected, processed, and analyzed by an order of magnitude. The log analysis engine 446 can assign different log template categories to the mapped log templates (e.g., the log templates that have been mapped to candidate logs). For example, the log analysis engine 446 can classify the mapped log templates into three different log template categories based on key keywords or the frequency or rarity of the log templates mapped to candidate logs (e.g., the number of instances of the log templates). The log analysis engine 446 can classify the mapped log templates into log template categories that can include a first category for well-known log templates, a second category for rarely occurring log templates, and a third category for log models that the trained template mining model cannot recognize. In some examples, the log analysis engine 446 can assign the mapped log template categories to the log templates in order. That is, the log analysis engine 446 can assign the first log template category to the log templates based on key keywords, then assign the second log template category to the remaining log templates based on the frequency with which the corresponding log templates have been mapped to candidate logs, and then assign the third log template category to the remaining log templates based on log templates that the trained template mining model cannot recognize (e.g., the trained template mining model has not learned the scheme or pattern of the log templates).

[0092] The log analysis engine 446 can classify mapped log templates by searching for log templates for key keywords. The log analysis engine 446 can automatically tag candidate logs associated with the mapped log templates for further analysis (e.g., classified as category 1) in response to determining that the mapped log template includes keywords strongly related to an error (e.g., a fault, an error, a crash, etc.). In some instances, the log analysis engine 446 can search for keywords within a sample log with unmasked log values. The log analysis engine 446 can search for corresponding sample logs for candidate logs, which can include key keywords masked by variable tokens. In response to determining that the log template mapped to one or more candidate logs includes key keywords, the log analysis engine 446 can assign the log template to a first category or classification of well-known log templates that historically indicate abnormal behavior.

[0093] The log analysis engine 446 can classify mapped log templates by searching for log templates that are rarely mapped in candidate logs. The log analysis engine 446 can classify mapped log templates by counting the number of candidate logs (e.g., the number of instances of a log template) mapped to each log template generated by a trained template mining model. The log analysis engine 446 can determine the frequency or rarity of a log template by counting the number of candidate logs (e.g., the time period corresponding to the timestamps of the candidate logs) mapped to certain log templates within an analysis window. In response to determining that a log template is rarely mapped to one or more candidate logs, the log analysis engine 446 can assign the mapped log template to a second category corresponding to rare network events.

[0094] The log analysis engine 446 can classify mapped log templates based on whether the log templates have been recognized by a trained template mining model. The trained template mining model can be trained to identify log patterns based on historical log data. The log analysis engine 446 can track log templates generated based on patterns of historical log data and candidate logs identified by the template mining model. The log analysis engine 446 can classify unrecognized mapped log templates generated for candidate logs because the historical training data used to train the template mining model includes log data of well-performing applications and infrastructure. In response to determining that a mapped log template has not been previously learned, the log analysis engine 446 can assign the log template to a third category of unrecognized log templates, which can be considered anomalies and may be related to abnormal system behavior. Once the log analysis engine 446 determines that a log template has been mapped to a candidate log, the log analysis engine 446 can track the mapped log template for future use and potential classification as a first or second category in subsequent analysis. In some examples, candidate logs with scenarios or patterns that cannot be recognized by the template mining model may themselves be log templates assigned to this category.

[0095] The log analysis engine 446 can select candidate logs as critical logs 470 at least in part based on the category assigned to the mapped log templates. The log analysis engine 446 can determine the critical logs 470 by ranking or scoring the mapped log templates based on multiple factors such as the category assigned to the mapped log templates, the timestamps of the corresponding candidate logs, and / or the keywords included in the candidate logs and the corresponding log templates.

[0096] In one example, the log analysis engine 446 can use the timestamps of the candidate logs and the categories assigned to the mapped log templates to determine the critical logs 470. The log analysis engine 446 can rank the mapped log templates chronologically with respect to the detection of performance issues of the network application. For example, the log analysis engine 446 can rank the mapped log templates based on the timestamps of the corresponding instances of the log templates or candidate logs, where the timestamps indicate times closest to the time when the performance issues of the network application are determined. The log analysis engine 446 can consistently select mapped log templates from different log template categories while considering the timestamps of the log template instances (e.g., discrete mapping of candidate logs to log templates). For each log template category, the log analysis engine 446 can give higher priority to log template instances with timestamps that indicate times closest to the time of the performance issue.

[0097] In this example, the log analysis engine 446 can determine candidate logs with timestamps for each mapped log template, where the timestamp indicates the time closest to the determination of the performance issue (referred to herein as the "latest timestamp"). For example, the log analysis engine 446 can map a log template to candidate logs to rule out performance issues of a web application at a time equal to "T". The log analysis engine 346 can map the first log template to a first candidate log with a timestamp of "T1" (e.g., 1 minute before time "T"), a second candidate log with a timestamp of "T2" (e.g., 2 minutes before time "T"), and a third candidate log with a timestamp of "T3" (e.g., 3 minutes before time "T"). The log analysis engine 446 can determine the first candidate log as the candidate log of the first log template with the latest timestamp. Similarly, the log analysis engine 446 can map the second log template to a fourth candidate log with a timestamp of "T4" (e.g., 20 seconds before time "T") and a fifth candidate log with a timestamp of "T5" (e.g., 30 seconds before time "T"). The log analysis engine 446 can determine the fourth candidate log as the candidate log of the second log template with the latest timestamp. Similarly, the log analysis engine 446 can map the third log template to a sixth candidate log with a timestamp of "T6" (e.g., 1 minute before time "T") and a seventh candidate log with a timestamp of "T7" (e.g., 2 minutes before time "T"). The log analysis engine 446 can determine the sixth candidate log as the candidate log of the third log template with the latest timestamp. Similarly, the log analysis engine 446 can map the fourth log template to an eighth candidate log with a timestamp of "T8" (e.g., 2 minutes before time "T"), a ninth candidate log with a timestamp of "T9" (e.g., 4 minutes before time "T"), and a tenth candidate log with a timestamp of "T10" (e.g., 6 minutes before time "T"). The log analysis engine 446 can determine the eighth candidate log as the candidate log of the fourth log template with the latest timestamp. Similarly, the log analysis engine 446 can map the fifth log template to an eleventh candidate log with a timestamp of "T11" (e.g., 1 minute before time "T"). The log analysis engine 446 can determine the eleventh candidate log as the candidate log of the fifth log template with the latest timestamp. Similarly, the log analysis engine 446 can map the sixth log template to a twelfth candidate log with a timestamp of "T12" (e.g., 1 minute before time "T") and a thirteenth candidate log with a timestamp of "T13" (e.g., 2 minutes before time "T"). The log analysis engine 446 can determine the twelfth candidate log as the candidate log of the sixth log template with the latest timestamp. Similarly, the log analysis engine 446 can map the seventh log template to a fourteenth candidate log with a timestamp of "T14" (e.g., 1 minute before time "T") and a fifteenth candidate log with a timestamp of "T15" (e.g., 2 minutes before time "T").The log analysis engine 446 can determine the fourteenth candidate log as the candidate log of the seventh log template with the latest timestamp. The log analysis engine 446 can determine the candidate log with the latest timestamp for each candidate log template in a similar manner.

[0098] Then, the log analysis engine 446 can rank the mapped log templates for each log template category based on the corresponding candidate logs with the latest timestamp. The log analysis engine 446 can group each mapped log template based on the log template category assigned to the log template. For example, the log analysis engine 446 can assign the log template category (e.g., category 1) corresponding to the well-known log template to the first log template and the second log template according to the above example. The log analysis engine 446 can group the first log template and the second log template and the corresponding candidate logs with the latest timestamp (e.g., the first candidate log and the fourth candidate log respectively). The log analysis engine 446 can assign the log template category (e.g., category 2) corresponding to the relatively rare log template to the fourth log template and the sixth log template according to the above example. The log analysis engine 446 can group the fourth log template and the sixth log template and the corresponding candidate logs with the latest timestamp (e.g., the eighth candidate log and the twelfth candidate log respectively). The log analysis engine 446 can assign the log template category (e.g., category 3) corresponding to the unrecognized log template to the third log template, the fifth log template, and the seventh log template according to the above example. The log analysis engine 446 can group the third log template, the fifth log template, and the seventh log template and the corresponding candidate logs with the latest timestamp (e.g., the sixth candidate log, the eleventh candidate log, and the fourteenth candidate log respectively).

[0099] The log analysis engine 446 can select a predefined number of mapped log templates based on the ranking of the log templates. The log analysis engine 446 can select the ranked log templates from each log template category in a round robin manner. In other words, the log analysis engine 446 selects an instance of the log template with the latest timestamp. According to the above example, the log analysis engine 446 can be configured to select three log templates in a round robin manner based on the log templates of the candidate logs with the latest timestamp. The log analysis engine 446 can select the first log template from category 1 (corresponding to the candidate log with the timestamp "T1"), the fourth log template from category 2 (corresponding to the candidate log with the timestamp "T8"), and the third log template from category 3 (corresponding to the candidate log with the timestamp "T6"). If the log analysis engine 446 is configured to select three log templates in a round robin manner, the log analysis engine 446 can select the first log template, the fourth log template, the third log template, and the second log template (corresponding to the candidate log with the timestamp "T4") from category 1, and the sixth log template (corresponding to the candidate log with the timestamp "T12") from category 2. The log analysis engine 446 can select a candidate log from each of the selected log templates by mapping the selected log templates back to the corresponding candidate logs. In some instances, the log analysis engine 446 can determine the critical log 470 based on the selected candidate logs with the latest timestamp.

[0100] In another example, the log analysis engine 446 can determine the critical log 470 based on the critical template score. The log analysis engine 446 can calculate the critical template score for each mapped log template. The log analysis engine 446 can calculate the critical template score based on multiple factors, such as the occurrence time of when a performance anomaly relative to the network application occurs, the rarity of the log template observed in the analysis or inference window, the log template category assigned to the mapped log template, and whether the mapped log template includes unimportant keywords that typically indicate less critical logs. The log analysis engine 446 can assign each factor value corresponding to the weight by which each factor should affect the calculated critical template score. In some examples, the log analysis engine 446 can multiply the weight values of each factor to generate a raw critical template score value between zero and negative infinity. The log analysis engine 446 can normalize the raw critical template score (e.g., using the SoftMax function) to generate a critical template score with a value between zero and one. The log analysis engine 446 can select multiple log templates based on the critical template scores of the log templates closest to one.

[0101] The log analysis engine 446 can calculate a critical template score based on factor weight values. The critical template score can include a template recency time weight, a template rarity weight, a template category weight, and a template unimportant keyword weight. The weight value of the template recency time weight can be determined based on the timestamps of the candidate logs mapped to the log template. Compared with a log template observed in the distant past relative to the application anomaly time, the log analysis engine 446 can give more weight to the log template observed closest to the time of the application anomaly occurrence. The template recency time weight value can be determined using the sigmoid function of the natural logarithm of the time difference between the candidate logs of the log template and the network application performance issue. By taking the negative sigmoid of the natural logarithm of the time recency, a template recency time weight value close to negative one indicates a log template with candidate logs that occurred in the distant past, while a template recency time weight value close to zero indicates a log template with candidate logs close to the application anomaly time. For example, the log analysis engine 446 can determine that the time recency of a first log template is 10, and the time recency of a second log template is 150. The log analysis engine 446 can determine that the template recency time weight value of the first log template is -0.90, and the template recency time weight value of the second log template is -0.993. In an example where the application anomaly time is unknown, the log analysis engine 446 can consider the middle time of the analysis window (e.g., the time period during which the candidate logs are collected) as the anomaly occurrence time.

[0102] The weight value of the template rarity weight can be determined based on the frequency of observing a specific log template in the analysis window. During the analysis window, a lower weight value is given to frequently occurring log templates than to rarely occurring log modules. The log analysis engine 446 determines the template rarity weight value by counting the number of specific log templates that appeared in the analysis window before the application anomaly was detected. The template rarity time weight value can be determined using the sigmoid function of the natural logarithm of the frequency (e.g., the number of count instances of the log template mapped to the candidate logs). By taking the negative sigmoid of the natural logarithm of the frequency, a template rarity time weight value close to negative one indicates a log template that was frequently observed during the analysis window, and a template rarity time weight value close to negative 0.5 indicates a log template that was rarely observed during the analysis window. For example, the log analysis engine 446 can determine that the occurrence count of a first log template is 50, and the occurrence count of a second log template is 2. The log analysis engine 446 can determine that the template rarity weight value of the first log template is -0.98, and the template rarity weight value of the second log template is -0.66.

[0103] The weight value of the template category weight can be determined based on the category or classification assigned to the log template. The log analysis engine 446 can assign a log template category to the log template mapped to the candidate log, as previously discussed. For example, the log analysis engine 446 can assign an unrecognized template category to the log template based on whether the trained template mining model can classify the log template during the analysis window. The log analysis engine 446 can determine that the template category weight value of the log template classified in the unrecognized template category is one. The log analysis engine 446 can assign a critical keyword category to the log template that contains specific predefined keywords (such as failure, error, crash, etc.). The log analysis engine 446 can determine that the template category weight value of the log template classified in the critical keyword category is greater than one. As previously discussed, the log analysis engine 446 can assign a rare template category to the log template based on the template rarity weight. The log analysis engine 446 can group the log templates based on the template rarity weight value of each log template. For example, the log analysis engine 446 can consider only the top 10% of the log templates with the lowest template rarity weight value as part of the rare template category. The log analysis engine 446 can determine that the template category weight of the top 10% of the log templates with the lowest template rarity weight value is 2.5. If the log template is rare, depending on how the rare template category is defined, greater weights are given to the log templates in the critical keyword category and the unrecognized template category so that they are selected as the critical log 470. In an example where a log template is assigned to multiple log template categories, the analysis engine 446 can obtain the total of the template category weights determined for each category by multiplying each weight value determined for each template category.

[0104] The weight value of the template unimportant keyword weight can be determined based on well-known keywords lacking relevance to critical log detection. The log analysis engine 446 can assign a template unimportant keyword weight to the log template that includes keywords lacking relevance (such as "GET" or "INFO"). The log analysis engine 446 can assign a template unimportant keyword weight of 7.5 to the log template that includes unimportant keywords to give a stronger penalty to the log module that contains keywords irrelevant to critical log determination.

[0105] The log analysis engine 446 can determine the raw critical template score based on the weight values determined for each factor. The log analysis engine 446 can determine the raw critical template score using the following equation:

[0106] Raw_critical_template_score = template_recency_time_weight * temp late_rarity_weight * template_category_weight * template_insignificant_k eyword_weight * -1

[0107] Applying the output of the above raw critical template score equation makes the raw critical template score between zero and negative infinity. The log analysis engine 446 can input the raw critical template score into a SoftMax function to transform the raw critical template score into a critical template score. The critical template score can represent the probability that a candidate log for a specific log template can be included in the critical log 470. The following table provides an example output of the weight values, raw template score values, and template score values for five different log templates mapped to candidate logs.

[0108]

[0109]

[0110] The log analysis engine 446 can select multiple log templates based on the critical template scores determined for each mapped log template. For example, according to the above table, the log analysis engine 446 can select the top three log templates of "Template 2", "Template 3", and "Template 5", where the critical template scores are closest to one. The log analysis engine 446 can select multiple candidate logs by mapping the selected templates back to the candidate logs with the latest timestamps. In other words, the log analysis engine 446 can select multiple instances of log templates based on the instances with the latest timestamps.

[0111] Figure 4B is a block diagram illustrating an example of a log analysis engine for determining critical logs using a model service according to one or more techniques of the present disclosure. In Figure 4B the example, the log analysis engine 446, historical log data 462, model trainer 452, model service 454, log collector tool 444, and log database 446 can respectively include Figure 2 examples of the log analysis engine 246, historical log data 262, model trainer 252, model service 254, log collector tool 244, and log database 242 of

[0112] The model trainer 452 can train a template mining model (e.g., model service 454) with historical log data 462. The model trainer 452 can utilize the historical log data 462 to train the template mining model of the model service 454, such as application layer logs, compute layer logs, and network layer logs collected during a training period. The model trainer 452 can train the template mining model of the model service 454 to observe or generate log templates for multiple cross-layer system logs. After the model trainer 452 trains the template mining model, the log analysis engine 446 can apply the model service 454 to generate log templates for candidate logs. In some examples, the model trainer 452 can utilize the template mining model to generate log templates and provide the log templates to the model service 454 for determining which of the multiple log templates will be the mapped log templates.

[0113] The log collector tool 444 can obtain and preprocess candidate logs from multiple layers of the network infrastructure. The log collector tool 444 can preprocess the candidate logs by including the source and timestamp in the metadata of the candidate logs. The log collector tool 444 can store and index the preprocessed candidate logs in the log database 444. The model service 454 can generate log templates for the candidate logs and map the candidate logs to the log templates to reduce the amount of data to be analyzed to determine critical logs. For example, the model service 454 can map multiple logs to the same log template based on structural similarity. The following example illustrates how the model service 454 maps candidate logs to log templates, where * represents variables, masked values:

[0114] Original log – sshd

[1234] : Accepted password for operation from 192.0.2.3 port 12345 ssh2

[0115] Corresponding template – sshd[*]: Accepted password for * from * port **

[0116] Application layer:

[0117] Original log:

[0118] Apr 25 15:14:32 2023-04-25T15:14:32.475051205Z http.req.id:9c5dace2-4dfe-4fd1-8c59-df0079f3496b,http.req.method:GET,http.req.path:cart,message:view user cart,session:0c0bbe20-8011-4781-afa0-80d205af8f9,severity:debug,timestamp:2023-04-25T15:14:32.475051205Z

[0119] Template:

[0120] <timestamp><*>"http.req.id:<*>http.req.method:GET,http.req.path:<*>message:<*><*><*>session:<*>severity:debug,timestamp:<*>APPLICATION

[0121] Network layer:

[0122] Original log:

[0123] Apr 25 14:14:05 2023-04-25T14:14:05.25Z alert_config_deviation_statustype:gauge,count:1,sum:1,min:1,max:1,latest:1,blueprint:CN2-Hetero,description:Telegraf collected metric,device:jfm-qnc-qfx5k-01,device_key:WS3120270462,device_name:jfm-qnc-qfx5k-01,entity.guid:Mzc0NjA2OXxFWFR8U0VSVklDRXw0NzI0MTU2MjgyODc2NjgzMT Yw,entity.name:aos-otel-collector,entity.type:SERVICE,host:045b990d5d7e, http.scheme: http, instrumentation.provider:opentelemetry, metricName: alert_config_deviation_status,net.host.name:10.213.15.123,net.host.port:9126,newrelic.source:api.metrics.otlp,otel.library.name:otel.library.version:,role:leaf,service.instance.id:10.213.15.123:9126,service.name:aos-otel-collector,severity:ALERT_CRITICAL\,timestamp:1682432045199

[0124] Template:

[0125] <timestamp><*>alert_config_deviation_status type:gauge,count:1,sum:<*>min:<*>max:<*>latest:<*>blueprint:CN2-Hetero,description:Telegraf collected metric,device:<*>device_key:<*>device_name:<*>entity.guid:<*>,entity.name:aos-otel-collector,entity.type: SERVICE, host: <*>, http.scheme: http,instrumentation.provider: opentelemetry, metricName:alert_config_deviation_status,net.host.name:<*>,net.host.port:<*>,newrelic.source:api.metrics.otlp,otel.library.name:otel.library.version,role:leaf,service.instance.id:<*>,service.name:aos-otel-collector,severity:ALERT_CRITICAL,timestamp:<*>DEVICE

[0126] The model trainer 452 trains a template mining model for determining log templates for a set of critical logs. For example, the model trainer 452 can implement Drain3 for the template generation algorithm. The model trainer 452 can use the Drain3 algorithm for template generation and customize the template generation to be optimized for various use cases. The model trainer 452 can use the Drain3 python library, which provides utility support for training and storing the template mining model. With the dynamic update template clustering system that adopts the Drain3 template mining model by the model trainer 452, the template mining model of the log analysis engine 446 can dynamically learn templates by masking variable tokens within common strings. Even after training, the template mining model can be retrained and will update its log templates to reflect new logs being ingested. The template mining model can be trained on application and infrastructure log data during normal execution and template the logs into clusters, saving the cluster sizes and template strings into a PostgreSQL database. The template mining model can optionally provide REGEX patterns and corresponding masks to be inserted into strings when found. This allows the template mining model to handle network-specific logs, including masks for timestamps, container IDs, HTTP response codes, etc.

[0127] The log collection tool 444 can ingest logs daily from the application layer, compute layer, and network layer. Conveniently, the Drain3 model may not require knowledge of the log structure and may also not require sourcing of the layers. The template mining model can be trained on logs from all layers of the system and can distinguish the layers from each other (given the different content of the logs). As a redundancy, the log collection tool 444 can append the name of the layer to the ingested logs so that they can always be traced back to their source.

[0128] Figure 5 is a block diagram illustrating an example of an analysis system for determining critical logs with performance metrics in accordance with one or more techniques of the present disclosure. In Figure 5 the example, the application logs 532, compute logs 534, network logs 536, dependency mapping generator service 556, log analysis engine 546, and causality analysis service 566 can include examples of the application logs 232, compute logs 234, network logs 236, dependency mapping generator service 256, log analysis engine 246, and causality analysis service 266.

[0129] The anomaly detection engine 558 can obtain cross-layer metric data from various layers of the data center. The cross-layer metric data can include key performance indicators, traces, or other types of metrics used to measure the performance of data center operations. The anomaly detection engine 558 can analyze the cross-layer metric data 538 for anomalies. For example, the anomaly detection engine 558 can independently analyze the application key performance indicator (KPI) 538A, the compute key performance indicator (KPI) 538B, and the network key performance indicator (KPI) 538C. The anomaly detection engine 558 can use machine learning-based methods to analyze the cross-layer metric data 538. The anomaly detection engine 558 can analyze the cross-layer performance metric data 538 of specific nodes identified by the dependency mapping generator service 556, which are determined to be affected by application anomalies. The anomaly detection engine 558 can generate an anomaly log 560 based on the content of the cross-layer metric data 538. The anomaly detection engine 558 can transform the values of the cross-layer performance metric data 538 (e.g., the values of the key performance indicators from each layer) to generate the anomaly log 560. The anomaly log 560 can include event messages associated with performance issues and the corresponding timestamps when the performance issues were observed.

[0130] The anomaly detection engine 558 can output the anomaly log 560. The anomaly log 560 can include a collection of structured log messages that capture specific metrics that are anomalous compared to the baseline behavior. The anomaly log 560 can include anomaly messages that specify the context of the specific KPI and layer, as well as the timestamp when the anomaly was observed. The following is an example schema for the anomaly log messages of the anomaly log 560:

[0131]

[0132] The following are examples of the anomaly log messages of the anomaly log 560:

[0133]

[0134]

[0135] The log analysis engine 546 can obtain candidate logs for ingestion or preprocessing of application logs 532, compute logs 534, and network logs 536. The log analysis engine 546 can reduce the number of candidate logs obtained based on nodes identified by the dependency mapping generator service 556, which specifies a knowledge graph of nodes associated with application anomalies. The dependency mapping generator service 556 can determine a subgraph from the perspective of an anomalous application based on the anomaly log 560. The log analysis engine 546 can obtain candidate logs from the nodes in the subgraph that are identified based on metric anomalies in the anomaly log 560. The log analysis engine 546 can determine critical logs 570 based on the candidate logs and the anomaly log 560. In some instances, the anomaly log can be included as part of the candidate logs. The following is an example of determining critical logs 570 based on the candidate logs and the anomaly log 560.

[0136] Application layer logs:

[0137] Sep 05 22:02:27{"log":"2023-09-05T22:02:27.165933914Z stderr F wget:can't connect to remote host(10.108.237.211):Connection refused","@timestamp":"2023-09-05T22:05:17.500937079+00:00","kub ernetes":{"pod_name":"busybox-76c87475f-xdbvn","namespace_name":"default","pod_id":"f84e78ed-69dd-4fec-8aa9-7e4a6fd7ed3d","host":"r5-u17-dell","container_name":"busybox","container_image":"docker.io / li brary / busybox:latest"}}pods

[0138] Network layer logs:

[0139] Sept 05 22:01:25"{"metric":"anomalous_interface_counters_rx_bps","timestamp":1693001574,

[0140] "Value":"644760""labels":{"entityLayer":"Network","entityType":"switch.interface"

[0141] "entityId":"dc1-must-esi-001-leaf1.xe-0 / 0 / 1","blueprint":"must_blueprint_dc1",

[0142] "cluster_name":"Apstra_Cluster1","anomalySource":"ad_service"}switches Sep 05 22:05:05 {"log":"Sep 5 22:05:05 10.6.1.442023-09-06T03:29:04.450527IST aos-server CEF:0|Apstra|AOS|4.1.2-269|101|Alert|10|msg={u'blueprint_label':u'must_blueprint_dc1',u'timestamp':1693951144450527,u'origin_name':u'XH3719090062::ae4',u'alert':{u'first_seen':1693951144450489,u'interface_link_status_mismatch_alert':{u'expected_ifstatus': 0, u'ifname':u'ae4', u'hostname':u'dc1-must-esi-001-leaf1',u'actual_ifstatus':1},u'raised':True,u'severity':3,u'id':u'a7f73d9b-9f03-452e-b001-3470c0c338ec'},u'origin_hostname':u'dc1-must-esi-001-leaf1','device_hostname':'dc1-must-esi-001-leaf1',u'origin_role':u'to_generic"}","@timestamp":"2023-09-05T22:05:05.495466407+00:00"}switches

[0143] The log analysis engine 546 can provide the critical log 5570 to the causal analysis service 566 to rule out performance issues of the web application. The causal analysis service 566 concurrently considers different types of telemetry, such as logs, metrics, traces, to help identify the cause of performance issues with higher accuracy.

[0144] Figure 6 is a flowchart illustrating an example operation for determining critical logs according to one or more techniques of the present disclosure. Figure 6 For example purposes only with respect to Figures 1 to 5 is discussed.

[0145] The analysis system 140 can obtain multiple candidate logs (602) for multiple layers of the computing infrastructure 100. The candidate logs can include system logs from the application layer, system logs from the computing layer, and system logs from the network layer. In some examples, the analysis system 140 can determine a subset of system logs from multiple layers based on a knowledge graph that identifies nodes in each layer affected by performance issues of the web application.

[0146] For each candidate log among the multiple candidate logs, the analysis system 140 can map the candidate log to a log template among multiple log templates, where each log template to which the candidate log is mapped is a mapped log template (604). In some examples, the log templates can be generated during the training of the template mining model (e.g., generated by the model trainer 152). In some instances, the model service 154 in the log analysis engine 154 can generate log templates.

[0147] The analysis system 140, or more specifically, the model service 154 can rank the mapped log templates based on the characteristics of each mapped log template in the mapped log templates (606). The model service 154 can rank the mapped log templates based on the characteristics of the mapped log templates, such as keywords included in the log template, the number of instances of the log template, and / or whether the log schema of the log template was learned by the trained template mining model before generating the log template. In some examples, the model service 154 can assign a category to the mapped log templates based on the characteristics of the mapped log templates. The model service 154 can consider the assigned category when ranking the mapped log templates.

[0148] The model service 154 may select one or more candidate logs corresponding to the mapped log template as critical logs (608) based on the ranking of the mapped log template. In some instances, the model service 154 may first select the mapped log template based on the ranking, and then select the corresponding candidate logs by mapping the candidate logs with the latest timestamp back to the mapped log template. In some examples, the model service 154 may extract candidate logs as critical logs according to an instance of the log template selected based on a ranking heuristic. The model service 154 may output at least one of an indication of the critical logs for determining a potential root cause associated with a performance issue of a network application or an indication of a potential root cause associated with a performance issue of a network application (610).

[0149] The techniques described herein may be implemented in hardware, software, firmware, or any combination thereof. The various features described as modules, units, or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices or other hardware devices. In some cases, the various features of an electronic circuit system may be implemented as one or more integrated circuit devices, such as an integrated circuit chip or chipset.

[0150] If implemented in hardware, the present disclosure may relate to an apparatus, such as a processor or an integrated circuit device, such as an integrated circuit chip or chipset. Alternatively or additionally, if implemented in software or firmware, the techniques may be at least partially implemented by a computer-readable data storage medium including instructions that, when executed, cause one or more processors to perform one or more of the methods described above. For example, the computer-readable data storage medium may store such instructions for execution by one or more processors.

[0151] The computer-readable medium may form part of a computer program product that may include packaging material. The computer-readable medium may include a computer data storage medium, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. In some examples, an article of manufacture may include one or more computer-readable storage media. The computer-readable storage media may be distributed among multiple packages, devices, or other components capable of being configured with computer instructions.

[0152] In some examples, the computer-readable storage medium may include a non-transitory medium. The term "non-transitory" may indicate that the storage medium is not embodied in a carrier wave or a propagated signal. In certain examples, the non-transitory storage medium may store data that may change over time (e.g., in RAM or a cache).

[0153] The code or instructions can be software and / or firmware executed by a processing circuitry that includes one or more processors, such as one or more digital signal processors (DSPs), general microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, the term "processor" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described in this disclosure may be provided within software modules or hardware modules.< / timestamp> < / timestamp>

Claims

1. A network application analysis method, comprising: obtaining, by a computing system, a plurality of candidate logs for a plurality of layers of a computing infrastructure; For each candidate log in the plurality of candidate logs, the computing system: Mapping the candidate log to a log template in a plurality of log templates, Each log template to which a candidate log is mapped is a mapped log template; The computing system ranks the mapping log templates based on a characteristic of each of the mapping log templates; selecting, by a computing system, one or more candidate logs corresponding to the mapping log template as key logs based on the ranking of the mapping log template; as well as The computing system outputs at least one of: (1) an indication of the critical log for determining a potential root cause associated with a performance problem of a network application, or (2) an indication of the potential root cause associated with the performance problem of the network application.

2. The network application analysis method according to claim 1, further comprising: training a template mining model to identify patterns of the plurality of candidate logs; as well as The plurality of log templates are generated by providing historical log data to the template mining model.

3. The network application analysis method according to claim 1, further comprising: storing, by the computing system, each candidate log of the plurality of candidate logs in a database having metadata identifying a tier of the plurality of tiers and a timestamp corresponding to the candidate log of the plurality of candidate logs; as well as The plurality of candidate logs are obtained from the database.

4. The network application analysis method according to claim 1, wherein the characteristics of each mapping log template in the mapping log template include keywords included in the mapping log template, the number of instances of the mapping log template, and whether the trained template mining model recognizes the log scheme of the log template.

5. The network application analysis method according to claim 1, further comprising: obtaining a knowledge graph identifying dependencies of nodes within each of the plurality of layers; determining one or more nodes affected by the performance issue based on the knowledge graph; as well as The plurality of candidate logs are obtained from the plurality of layers, wherein the plurality of candidate logs include system logs from the one or more nodes.

6. The network application analysis method according to claim 1, further comprising: obtaining a metric from each of the plurality of layers; Generate an exception log based on the indicator; One or more nodes are identified from each of the plurality of layers based on the anomaly log, wherein the plurality of candidate logs include system logs from the one or more nodes and the anomaly log.

7. The network application analysis method according to claim 6, wherein the indicators include values ​​of key performance indicators of the plurality of layers, and Wherein generating the one or more exception logs comprises converting the value of the key performance indicator into an event message.

8. The network application analysis method according to any one of claims 1 to 7, wherein selecting one or more candidate logs comprises: selecting one or more mapping log templates based on the ranking of the mapping log templates; For each of the selected one or more mapping log templates, determining an instance of the mapping log template having a latest timestamp; as well as The one or more candidate logs are selected by mapping the instance of the mapping log template back to corresponding candidate logs.

9. The network application analysis method according to any one of claims 1 to 7, wherein the multiple layers include an application layer, a computing layer, and a network layer, and The multiple candidate logs include application layer logs, computing layer logs and network layer logs.

10. The network application analysis method according to any one of claims 1 to 7, wherein ranking the mapping log templates comprises: assigning a category to each of the mapping log templates based on the characteristic of each of the mapping log templates; Determining a latest timestamp for each mapping log template based on timestamps of candidate logs mapped to the mapping log template; as well as One or more mapping log templates are selected based on a category assigned to each of the mapping log templates and each of the latest timestamps.

11. The network application analysis method according to any one of claims 1 to 7, wherein ranking the mapping log templates comprises: assigning a category to each of the mapping log templates based on the characteristic of each of the mapping log templates; For each mapping log template, determining a corresponding critical template score based on one or more of a time when the performance problem is determined, a number of instances of the mapping log template, the category assigned to the mapping log template, and a keyword included in the mapping log template; as well as One or more mapping log templates are selected based on the corresponding key template scores determined for each mapping log template.

12. An analysis system comprising processing circuitry capable of accessing a storage device, the processing circuitry being configured to: obtaining multiple candidate logs for multiple layers of the computing infrastructure; For each candidate log in the plurality of candidate logs: Mapping the candidate log to a log template in a plurality of log templates, Each log template to which a candidate log is mapped is a mapped log template; ranking the mapping log templates based on a characteristic of each of the mapping log templates; selecting one or more candidate logs corresponding to the mapping log template as key logs based on the ranking of the mapping log template; as well as Outputting at least one of: (1) an indication of the critical log for determining a potential root cause associated with a performance problem of a network application, or (2) an indication of the potential root cause associated with the performance problem of the network application.

13. The analysis system of claim 12, wherein the processing circuit system is further configured to: obtaining a knowledge graph identifying dependencies of nodes within each of the plurality of layers; determining one or more nodes affected by the performance issue based on the knowledge graph; as well as The plurality of candidate logs are obtained from the plurality of layers, wherein the plurality of candidate logs include system logs from the one or more nodes.

14. The analysis system of claim 12, wherein the processing circuit system is further configured to: obtaining a metric from each of the plurality of layers; Generate an exception log based on the indicator; identifying one or more nodes from each of the plurality of layers based on the anomaly log; The plurality of candidate logs are determined, wherein the plurality of candidate logs include system logs from the one or more nodes and the exception log.

15. The analysis system according to claim 12, wherein the characteristics of each of the mapping log templates include keywords included in the mapping log template, the number of instances of the mapping log template, and whether the trained template mining model identifies the log scheme of the log template.

16. The analysis system according to any one of claims 12 to 15, wherein in order to rank the mapping log templates, the processing is configured to: assigning a category to each of the mapping log templates based on the characteristic of each of the mapping log templates; Determining a latest timestamp for each mapping log template based on timestamps of candidate logs mapped to the mapping log template; as well as One or more mapping log templates are selected based on a category assigned to each of the mapping log templates and each of the latest timestamps.

17. The analysis system according to any one of claims 12 to 15, wherein in order to rank the mapping log templates, the processing is configured to: assigning a category to each of the mapping log templates based on the characteristic of each of the mapping log templates; For each mapping log template, determining a corresponding critical template score based on one or more of a time when the performance problem is determined, a number of instances of the mapping log template, the category assigned to the mapping log template, and a keyword included in the mapping log template; as well as One or more mapping log templates are selected based on the corresponding key template scores determined for each mapping log template.

18. A computer-readable storage medium encoded with instructions for causing one or more programmable processors to be configured to execute the method of any one of claims 1 to 11 or to be configured as the computer network system of any one of claims 12 to 17.

Citation Information

Patent Citations

  • Tunneled packet aggregation for virtual networks

    US9571394B1