INTELLIGENT LOG ANALYSES FOR LARGE-SCALE HIGH-PERFORMANCE COMPUTERS AND ARTIFICIAL INTELLIGENCE SYSTEMS

The described system addresses the challenge of complex event analysis in HPC and AI systems by transforming and classifying log data using a relationship hierarchy and decision tree, facilitating efficient anomaly detection and resolution.

DE102025106946A1Pending Publication Date: 2026-03-26HEWLETT PACKARD ENTERPRISE DEV LP
0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Large-scale high-performance computing (HPC) and artificial intelligence (AI) systems face challenges in efficiently identifying the root cause of anomalies due to event information being scattered across numerous subsystems in various formats, leading to complex tracing and time-consuming, computationally intensive analysis.

Method used

A system that extracts, filters, and formats logs from multiple subsystems, transforming them into events using a relationship hierarchy and decision tree classification, generating reports and visual representations to facilitate interactive user feedback for corrective actions.

Benefits of technology

Enables efficient identification of root causes by correlating events across subsystems, reducing analysis time and computational intensity, and allowing for quick resolution of anomalies through visualizations and recommended actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000018_0000
    Figure 00000018_0000
  • Figure 00000019_0000
    Figure 00000019_0000
  • Figure 00000020_0000
    Figure 00000020_0000
Patent Text Reader

Abstract

A system receives event information from components working together within the system. This information consists of a first set of events interpreted from log entries associated with the components, and a second set of events returned from queries for standard events. The system classifies the events interpreted from the log entries based on a hierarchy of the components. The system correlates two or more events based on their respective event classifications and a predetermined time window covering an event time associated with each event. The event time is derived from the log entries. The system generates a visual representation displaying the correlated events. In response to the visual representation indicating an anomaly, the system enables corrective actions to address the displayed anomaly.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Large systems such as high-performance computing (HPC) and artificial intelligence (AI) systems can comprise many subsystems, including storage infrastructure, network fabrics, host interfaces, centralized fabric managers (FMs), switches, and other controllers. Workloads in HPC and AI systems can be sensitive to events in these subsystems, impacting job performance. Anomaly detection and root cause analysis often involve extracting and analyzing event information from the subsystems. However, this event information can be scattered across numerous subsystems in various formats, such as host-level journal logs, FM console logs, and external system logs. Furthermore, relationships may exist between the different subsystems, potentially leading to complex tracing required for root cause analysis. BRIEF DESCRIPTION OF THE NUMBERS Fig. Figure 1A shows a system overview, including subsystems and protocols, of an environment that enables intelligent protocol analysis for large HPC and AI systems according to one aspect of the present application. Fig. Figure 1B shows an example component topology that enables intelligent protocol analysis for large HPC and AI systems according to one aspect of the present application. Fig. Figure 2 shows a high-level flow that facilitates intelligent protocol analysis for large HPC and AI systems in accordance with one aspect of the present application. Fig. Figure 3A shows a diagram of the conversion of log entries into a standard format according to one aspect of the present application. Fig. Figure 3B shows a diagram of the conversion of log entries into relevant events according to one aspect of the present application. Fig. 3C shows a decision tree used for event classification of log entries, according to one aspect of the present application. Fig. Figure 4 shows an environment with a protocol analysis system that communicates with multiple entities, enabling intelligent protocol analysis for large HPC and AI systems according to one aspect of the present application. Fig. Figure 5A shows an example screen with a visualization that shows anomalies of applications running on a host, a network interface controller (NIC) and hardware that exceed a certain threshold, in accordance with an aspect of the application at hand. Fig. Figure 5B shows an example screen with a visualization, including relevant time periods to be considered for correlations of events based on changes in power consumption, in accordance with one aspect of the present application. Fig. Figure 5C shows an example screen that represents a visualization, including events associated with anomalies of applications running on a host and hardware, in accordance with an aspect of the present application. Fig. Figure 5D shows an example screen that represents a visualization, including log extraction from a Fabric Manager and Fabric Controller agent, in accordance with one aspect of the present application. Fig. Figure 5E shows an example screen that represents a visualization, including anomalies of applications running on a host and events relating to a Fabric connection, in accordance with one aspect of the present application. Fig. Figure 5F shows an example display screen that provides a visualization, including network failure events and connection events, according to one aspect of the present application. Fig. 6A and Fig. Figure 6B shows flowcharts illustrating a procedure for facilitating intelligent protocol analysis for large HPC and AI systems in accordance with one aspect of the present application. Fig. Figure 7 shows a computer system that enables intelligent protocol analysis for large HPC and AI systems according to one aspect of the present application. Fig. Figure 8 shows a computer-readable medium that enables intelligent protocol analysis for large HPC and AI systems according to one aspect of the present application.

[0002] In the illustrations, identical numbers refer to the same elements of the illustration. DETAILED DESCRIPTION

[0003] Aspects of this application provide an intelligent analysis automation engine that: defines a hierarchy of relationships between the subsystems of an overall system; interprets log information from the subsystems into event information; and classifies these events to derive correlation information between them. The described aspects can also generate a report or visual representation of the correlations, enabling corrective action to be taken to address any identified anomalies.

[0004] Large systems (e.g., HPC and AI systems) can comprise many subsystems (e.g., storage infrastructure, network structure, host interfaces, centralized structure managers, switches, and other controllers). Workloads in such large systems can be sensitive to events in the subsystems, which can impact the performance of jobs running across them. Identifying relevant events and anomalies across the many subsystems and components can require extracting and analyzing event information distributed across numerous subsystems in various formats, such as host-level journal logs, fabric manager console logs, fabric controller agent console logs, external system logs, and so on. Furthermore, relationships may exist between the different subsystems, potentially leading to complex tracing for root cause analysis.

[0005] Extracting and analyzing event information distributed across numerous subsystems in various formats can be accomplished through custom-designed programs. However, such solutions can be time-consuming and computationally intensive. Furthermore, analyzing the relationships between subsystems can involve complex tasks. For example, the reliability service of a high-speed NIC might log events that are symptoms of a problem, rather than the problem itself. Reported timeouts can impact job performance caused by other factors, such as a network interface failure on another host or fabric link failures. Therefore, given the complexity of these tasks, analyzing the relationships between subsystems can be a limiting factor in efficiently identifying the root cause of various observed anomalous behaviors.

[0006] The described aspects address these limitations by providing a system that extracts, filters, and formats logs from multiple subsystems and then transforms the logs into events. The system can also process the events based on a relationship hierarchy (e.g., a decision tree, as shown below in relation to...). Fig. (Described in section 3C) classify and relate two or more events based on the classification and a specific time window associated with the respective events. The described aspects can also generate a report or visual representation of the correlations, which can lead to interactive user feedback that allows the user, for example, to take corrective action to resolve a displayed anomaly.

[0007] Fig. Figure 1A shows an environment 100, including subsystems and protocols, an environment that enables intelligent protocol analysis for large HPC and AI systems according to one aspect of the present application. The environment 100 can be a large HPC or AI system with multiple subsystems, each subsystem logging events in its own logs during operation. For example, an application 110 can log events in an application log 112. NIC controller agents 114 can log hardware events 116 in console logs and software events 118 in host logs. The host hardware 120 can include a central processing unit (CPU), a general-purpose processing unit (GPU), a peripheral component interconnect express (PCle) unit, a high-bandwidth memory (HBM) processor, and a dual in-line memory module (DIMM).Host hardware 120 can record hardware events 122 in console logs and software events 124 in job controller logs. A Fabric Manager (FM) 126 can log hardware events 118 in console logs of a Fabric Manager host and software events 130 in host logs of the Fabric Manager host. Domain Name Server (DNS) services 132 can log hardware events 134 in console logs and software events 136 in host logs. Chassis Managers (CMs) 138 can log events in chassis manager logs 140. Fabric Controller Agents (FCAs) 142 can log hardware events 144 in console logs of a switch and software events 146 in switch logs. Storage / Cluster Controller Agents 148 can log events in Storage / Cluster Logs 150. Rack Managers 152 can log events in Rack Manager Logs 154.The subsystems and protocols shown in Environment 100 are not exhaustive and are for illustrative purposes only. Other subsystems, components, units, and modules may create other protocols based on hardware, firmware, software, or a combination thereof.

[0008] Fig. Figure 1B shows an example of a component topology 160 that enables intelligent protocol analysis for large HPC and AI systems in accordance with one aspect of the present application. In the topology 160, a rack 162 can contain storage (or cluster) 164, a host 166, and chassis managers (CMs) 184. The host 166 can include a NIC 168, a CPU 174, a DIMM 176, an HBM 178, a GPU 180, and a resource allocation (and application launcher) service 182. The NIC 168 can interact based on the NIC controller software 170 and PCIe 172. The CPU 174 can also interact based on PCIe 172. CMs 184 can control or provide management services for switches 186. Fabric managers (FMs) 192 can also provide management services for and interact with switches 186. FM 192 can also interact with Fabric Controller Agents (FCAs) 188 and Domain Name Server / Network Time Protocol (DNS / NTP) 194.FCAs 188 can also interact with protocol agents 190. The organization of the elements (i.e., the subsystems) in the topology 160 is not limited and is for illustrative purposes only. Other topologies, elements (subsystems), and relationships between elements can be part of a network topology.

[0009] Fig. Figure 2 shows a high-level flow 200 that enables intelligent log analysis for large HPC and AI systems, as described in one aspect of the present application. During operation, the operations of modules 210, 212, and 214 can be performed by a log agent running in a specific component or subsystem (e.g., log agents 412, 432, 452, and 472, which are shown below). Fig. 4 are shown), while the operations of modules 216, 218, 220 and 222 can be performed by a central orchestrator (e.g. the Log Analytics Orchestrator 401, shown below in Fig. (as shown in Figure 4). A log extraction module 210 may contain a subsystem log agent that extracts various logs from components of the subsystem, such as host logs and console logs. A log filter module 212 may contain a log agent that eliminates noise in the extracted logs. A log transformation module 214 may contain a log agent that transforms the extracted and filtered log entries into event entries, as shown below with respect to the Fig. 3A and Fig. 3B described. An event classification module 216 can include a central orchestrator that classifies the events displayed in the converted event entries, as described below with respect to the Fig. 6A and Fig. 6B described. An event correlation module 218 can contain a central orchestrator that correlates the classified events based on a hierarchy of components, as described below in relation to Fig. 3C described. A Reporting Module 220 can include a central orchestrator that generates a report based on the correlated events, as described below with respect to the Fig. 4 and Fig. 5A-F described. A visual transformation module 222 can include a central orchestrator that generates a visual representation showing the correlated events, as below in relation to the Fig. 5A-F DESCRIBED. In addition, a user interaction module (not shown) may include a user's interactions with information generated by the Reporting module 220 or the Visual Transformation module 222, as described below in relation to the Fig. 5A-F DESCRIBED.

[0010] Fig. Figure 3A shows a diagram 300 illustrating the conversion of log entries into a standard format according to one aspect of the present application. Diagram 300 shows log entries 310, 320, and 330, all of which have the same standard format. For example, log entry 310 may contain event information such as: an entity 311 that corresponds to or is associated with an event; an event time 312 that indicates a time when the event occurred, such as a start time, an end time, or a time window; an event category 313 that indicates, for example, a severity level of the event; an event type 314 that indicates, for example, a software event, a hardware event, a processor event, a configuration event, or an error event; and event information 315 that provides a description of the event and other related information.Similarly, log entry 320 can contain: an entity 321; an event time 322; an event category 323; an event type 324; and event information 325. Furthermore, log entry 330 can contain: an entity 331; a time of the event 332; an event category 333; an event type 334; and event information 335.

[0011] A log agent running on a subsystem can create the formatted log entries of Diagram 300 based on the raw logs extracted from the various components of the subsystem. The log agent can further transform these formatted log entries, as described below with respect to Fig. 3B described.

[0012] Fig. Figure 3B shows a diagram 338 of the conversion of log entries into relevant events, in accordance with an aspect of the present application. Diagram 338 illustrates that log entries 360 can be converted into event entries 362 (as specified by 364) based on the event type (e.g., host, software, hardware, etc.) and the time format (e.g., a single time or a time window). For example, log entries occurring at time 340.A (or within a time window defined by time 340.A) may contain entries 342.1, 342.2, and 342.N.A (or within a time window defined by time 344.A) may contain entries 346.1, 346.2, and 346.N. and log entries that occur at time 348.A (or within a time window defined by time 348.A) may contain entries 350.1, 350.2 and 350.N.

[0013] A log agent running on a subsystem can convert log entries 360 into event entries 362, resulting in event entries grouped by similar corresponding times. For example, log entries 342.1-N, grouped at time 340.A, can be converted into events 352.1, 352.2, and 352.M, grouped at time 340.B. Log entries 346.1-N, grouped at time 344.A, can be converted into events 354.1, 354.2, and 354.M, grouped at time 344.B. Log entries 350.1-N, grouped at time 348.A, can be converted into events 356.1, 356.2, and 356.M, grouped at time 348.B. The log agent can convert a log entry into an event based on the event type information (e.g., as above regarding event type 314 of log entry 310 in Fig. 3A described).

[0014] Fig. Figure 3C shows a decision tree 368 that is used for event classification of log entries in accordance with one aspect of the present application. As above in relation to the event classification module 216 of Fig. As described in section 2, a central orchestrator can perform event classification after receiving the transformed event entries from the various subsystem log agents. An event classification 370 can refer to memory 371, host 374, or fabric 383. If the event is a memory event 371, the classification can be hardware 372 or software 373. If the event is a host event (374), the classification can be: hardware 375, which can be further classified as processor events (376), PCIe events (377), or DIMM / HBM events (378); or software 380, which can be further classified as associated with a memory leak (381) or software libraries (382). If the event is a Fabric event 383, the classification can be Hardware 384 or Software 391.Fabric event hardware 384 can be further classified as related to: a NIC 385, which can be further classified as related to hardware failure 386 or reliability service 387; or a switch 388, which can be further classified as related to hardware / application-specific integrated circuit (ASIC) failure 389 or fabric port failure 390. Fabric event software 391 can be further classified as related to a fabric manager (FM) 392 or a fabric controller agent (FCA) 394. FM 392 can be further classified as related to resources 393 or an invalid switch configuration 396. FCA 394 can be further classified as associated with resources 393, protocol agents 395, or an invalid switch configuration 396.

[0015] The organization and the elements that are in decision tree 368 of Fig. The examples shown in 3C are not limited and serve only for illustration. Other decision tree topologies and element relationships can also be used.

[0016] Fig. Figure 4 shows an environment 400 with a Log Analytics Orchestrator 401 that communicates with multiple units, enabling intelligent log analytics for large HPC and AI systems, as described in one aspect of the present application. In environment 400, the Log Analytics Orchestrator 401 (also referred to as "Orchestrator 401") can communicate with multiple units, including a Fabric Manager (FM) 410, a Switch 430, and Hosts 450 and 470. Each unit can contain its own log agent, which performs log extraction / collection of various logs generated and stored by that unit and also converts the logs into event entries. For example, FM 410 can contain a Log Agent 412, which includes a Log Extraction / Collection Module 414 that retrieves raw logs from, for example, a Fabric DB 427 or Host Logs 426.The Log Agent 412 may also contain a Log Transformation Module 416, which formats raw logs into log entries and event entries, as above in relation to the . Fig. 3A-B DESCRIBED. The Log Agent 412 can store the converted event entries, for example, in the Host Logs 426. The operations of the Log Extraction / Collection Module 414 can be assigned to Module 210 of Fig. 2 correspond, and the operations of the log-transform module 416 can be translated to the module 214 of Fig. 2 correspond. FM 410 can also include: a management level 420 with a health engine 421; a control level 422 with a routing engine 423; an operating system 424; and hardware 425.

[0017] Switch 430 can contain a log agent 432, which includes a log extraction / collection module 434 that obtains raw logs, for example, from an agent database 446 or host logs 445. Log agent 432 can also contain a log transformation module 436 that formats raw logs into log entries and event entries, as described above. Fig. 3A-B DESCRIBED. The log agent 432 can store the extracted / collected logs in the log event database 444 and store the converted event entries, for example, in the host logs 445. The intermediary 430 can also include: intermediary agents 440, platform services / software development kit (SDK) / drivers 441, an operating system 442, and hardware 443.

[0018] The host 450 can contain a log agent 452, which includes a log extraction / collection module 454 that retrieves raw logs, for example, from the host logs 465. The log agent 452 can also contain a log transformation module 456 that formats raw logs into log entries and event entries, as described above in relation to the Fig. DESCRIBED IN 3A-B. The log agent 452 can store the extracted / collected logs in the log event database 464 and store the transformed event entries, for example, in the host logs 465. Host 450 can also contain: host NIC agents 460; platform services / SDK / drivers 461; an operating system 462; and hardware 463. Similarly, host 470 can contain a log agent 472, which includes a log extraction / collection module 474 that retrieves raw logs, for example, from the host logs 485. Log agent 472 can also include a log transformation module 476 that formats raw logs into log entries and event entries, as described above in relation to the Fig. 3A-B DESCRIBED. The log agent 472 can store the extracted / collected logs in the log event database 484 and store the transformed event entries, for example, in the host logs 485. The host 470 can also include: host NIC agents 480; platform services / SDK / drivers 481; an operating system 482; and hardware 483.

[0019] The Orchestrator 401 can contain an Event Extraction / Collection module 404 and a Protocol Event Extraction module 405. The Event Extraction / Collection module 404 can query multiple entities for protocols, which may relate to standard events tracked by a given entity, for example, via a 490 communication from module 404 to FM 410. Although only the 490 communication with FM 410 is shown, module 404 can also query standard events from other locations. The Protocol Event Extraction module 405 can communicate with the protocol agents of the various entities to obtain the transformed event entries, e.g.,... B. via communication 491, 492, 493 and 494 with protocol agent 412 of FM 410, protocol agent 432 of Switch 430, protocol agent 452 of Host 450 and protocol agent 472 of Host 470.

[0020] After receiving both the events returned by queries for standard events (e.g., via 490) and the events interpreted from log entries associated with the entities or components (e.g., via 491-494), the orchestrator 401 can store the extracted data in one or more of relationship databases 406, time-series databases 407, or staging databases 408. In some aspects, the staging database 408 may contain the filtered, extracted, formatted, and transformed event entries, for example, those generated by the operations of module 214 in Fig. 2. Time-series DB 407 can contain log entries and event entries that have been grouped or clustered based on event type and time format (e.g., by a specific time or time window). The relationship DB 406 can contain information that relates two or more events based on their respective event classifications and event times. The system can also obtain component performance utilization from each component's management software over a period of time. Module 403 (or another module, not shown) can convert metrics related to application runtime and transaction results into time-series data and store this data, along with the obtained performance utilization, in DB 407.The system can use the data stored in one of the relationship DB 406, the time series 407 and the staging DB 408 to determine correlations and to identify relevant time periods with anomalous measurements or activities.

[0021] After classifying and correlating the events, a visualization and reporting module 402 of the Orchestrator 401 can generate reports and visualizations. An example of screen visualization is given below in relation to Fig. 5A-F. The processes of module 403 for classifying / correlating protocol events can be compared to modules 216 and 218 of Fig. 2 correspond, and the processes of the visualization and reporting module 402 can correspond to modules 222 and 220 respectively. Fig. 2 correspond.

[0022] The area around 400 of Fig. The four depicted units, components, and subsystems are not limited and are for illustrative purposes only. Other entities and relationships can also be used. For example, the functionality of Orchestrator 401 may reside in a single computer device, be accessible via a cloud computing environment, or be distributed across multiple virtual or physical network devices or nodes in a network environment. As another example, more or fewer elements or components may exist for each of the depicted units (FM 410, Switch 430, and Hosts 450 and 470).

[0023] Fig. Figure 5A shows an example display screen with a visualization, including anomalies of applications running on a host, network card, and hardware that exceed a certain threshold, according to one aspect of the application at hand. Graphs 500, 510, 520, and 530 in Fig. Figure 5A illustrates a representation of application measurements that depict anomalies in a sample used for correlating relevant transformed events from logs and standard events. A chart 500 shows measurements associated with a host event (e.g., Host Transaction A 502). A chart 510 shows measurements associated with a host event (e.g., Host Transaction B 512). A chart 520 shows measurements associated with a NIC event (e.g., NIC Event 522). A chart 530 shows measurements associated with a hardware event (e.g., Hardware Event 532). In charts 500, 510, 520, and 530, the x-axis displays time in ten-minute increments from 5:30 PM to 8:00 PM. In chart 500, the y-axis shows a time interval in seconds and minutes. In chart 510, the y-axis shows the time in milliseconds. In charts 520 and 530, the y-axis shows a number of errors (e.g.,(a number of errors at a specific point in time).

[0024] A user can view visualizations of event measurements from various entities based on the orchestrator's converted log entries. Visually reviewing the displayed information allows the user to quickly identify and resolve any issues.

[0025] In Chart 500, the partially shaded points represent a measurement of transaction A (504) at a specific time. Most measurements lie on the 0 ms line, indicating that most are below a certain expected threshold. However, Chart 500 also shows that transaction A lasts significantly longer than the threshold at 6:00 PM and between 6:45 PM and 6:50 PM.

[0026] In diagram 510, the partially shaded points correspond to a measurement of transaction B (514) taken at a specific time. Most measurements occur fairly evenly in the range between 1000 and 1800 milliseconds for the specified time period. No unusual or anomalous activity is apparent from diagram 510.

[0027] In diagram 520, the dots represent a count of the various NIC-related events (522). The partially hatched dots represent the power-on events (524), and the dots with bold outlines represent the flapping events (526). Diagram 520 shows that three NIC flapping events occurred between 6:45 PM and 6:50 PM, the same time period as the anomalous measurements of Host Transaction A (as shown in diagram 500). Consequently, a user can detect an anomaly in the Host Transaction A events and the NIC flapping, thus establishing a correlation between them. The user can then take corrective action to resolve the anomaly, such as rebooting or replacing the NIC.

[0028] In diagram 530, the dots represent a count of various hardware-related events (532). The partially shaded dots represent core failure events (534), the solid dots represent DIMM failure events (536), and the bold dots represent MCE failure events (538). Diagram 530 shows that two MCE failures occur between 5:45 PM and 5:50 PM, and diagram 510 shows that some anomalous Host Transaction B measurements also occur during the same time window. Consequently, a user can identify an anomaly and thus a correlation between Host Transaction B and the MCE failures detected in the hardware. The user can then take corrective action to resolve the anomaly, such as isolating or removing the node where the MCE failures were detected.

[0029] The system can also generate a report (not shown) that indicates the detected anomaly or correlation and suggests a corrective action for the user to take to resolve it. The report and visualization can include one or more interactive elements that facilitate viewing or manipulating the displayed information (whether in the report or visualization). These interactive elements might relate to: the detected anomaly; a recommended action indicating how to resolve the detected anomaly; or a configurable option indicating that the system should automatically perform the recommended action.In some aspects, the system can provide configurable or selectable default options at startup, relating to when to take a recommended option, what type of automatic action is approved by the user, how long an automatic action can be approved, etc.

[0030] Fig. Figure 5B shows an example screen with a visualization, including relevant time periods to consider for correlating events based on changes in power consumption, according to one aspect of the application at hand. Graph 550 shows power measurements (y-axis in megawatts (MW)) over time (x-axis of time in ten-minute increments from 5:30 PM to 8:00 PM). Graph 550 shows a distinct power spike at 6:06 PM. In conjunction with other visualizations, such as anomalies from applications running on a host, network card, or hardware, a user may determine that a specific event or events occurring around the same time may correlate with the power spike shown in Graph 550.Based on the visual representations generated by the system, the user can take corrective action to investigate or resolve the cause of the performance spike. For example, a user can identify patterns between the charts in [location / documentation]. Fig. 5A (i.e., 500, 510, 520 and 530) and diagram 550 of Fig. Observe 5B. A dip or spike in the performance curve that occurs during a similar timeframe to the application anomalies can identify a relevant period for further analysis. In graph 510, several anomalous measurements related to host transaction B 512 appear between 6:00 PM and 6:10 PM. During the same period, the performance curve in graph 550 shows a performance spike (between 6:00 PM and 6:10 PM). Based on the visual representations, the user (or the system) can correlate the events and take corrective action to further investigate the correlation between the anomalous measurements for host transaction B 512 and the performance curve in graph 550 that occurs during this relevant period, e.g., between 6:00 PM and 6:10 PM. The user can also further investigate other actions that may occur during this identified relevant period.

[0031] Fig. Figure 5C shows an example screen that provides a visualization, including events associated with anomalies of applications running on a host and hardware, consistent with an aspect of the application at hand. Graph 560 shows measurements associated with a host event (e.g., host transaction 562). In Graphs 560 and 565, the x-axis displays time in five-minute increments from 18:40 to 19:5. In Graph 560, the y-axis displays a time interval in milliseconds. In Graph 565, the y-axis indicates the number of errors (e.g., the number of errors at a given time).

[0032] In diagram 560, the partially shaded points correspond to a measurement of the host transaction (564) at a specific time. In diagram 560, transaction measurements lasting longer than 1000 milliseconds can be considered anomalies. For example, several anomalous measurements occur between 19:10 and 19:53. In diagram 565, the solid-colored points correspond to the DIMM error events (567). The same number of DIMM errors occur repeatedly throughout the measured period, including in groups of events that coincide with the anomalous events of host transaction 562, such as at 19:10 and 19:14, 19:45 and 19:26, 19:36 and 19:39, and 19:50 and 19:52. The DIMM errors (567) that occur consistently at a particular node can be correlated with the corresponding anomalous measurements for the host transaction (564). As a result, a user can take corrective action to resolve the anomaly, e.g.,Cancel the orders associated with the host transaction and take further action.

[0033] Fig. Figure 5D shows an example display screen that provides a visualization, including log extraction from a Fabric Manager and Fabric Controller agent, in accordance with one aspect of the present application. A graph 570 shows measurements associated with a Fabric Link event 571, and a graph 574 shows routing updates 575. In graphs 570 and 574, the x-axis shows time in five-minute increments from 18:40 to 19:55. In graph 570, the y-axis shows a number of Fabric Link events, and the solid dots represent link flaps or changes for a particular link (572). In graph 574, the y-axis shows routing updates, and the partially shaded dots indicate routing updates at a specified time (576). Based on Fig. 5D can establish a correlation between routing updates and fabric link changes during the time periods around 19:08 and 19:51. The routing updates can be observed as a result of fabric link changes, i.e., as correlated events. Therefore, times or time windows around these periods may be relevant for detecting anomalies or abnormal activity.

[0034] Fig. Figure 5E shows an example display screen with a visualization, including anomalies of applications running on a host and events relating to a Fabric connection, in accordance with one aspect of the present application. Fig. 5E shows an example of a protocol analysis of Fabric Controller agents, which is used in conjunction with the application protocols to correlate the behavior.

[0035] In charts 578 and 582, the x-axis shows time in ten-minute increments from 17:30 to 20:00. The data in charts 578 and 582 may be based on a sample high-performance benchmark run on thousands of nodes. Chart 578 shows measurements associated with transactions 579, where the y-axis indicates a time span in seconds, the partially shaded points show measurements for a swap transaction 580, and the solid-colored points show measurements for a broadcast transaction 581. Chart 582 shows measurements associated with a fabric link event 583, where the y-axis shows a number of fabric link events and the solid-colored points represent link flaps or changes for a particular link (584). Based on Fig. 5E can establish a correlation between certain transaction times and fabric events in the period around 6:41 PM. The fabric event (link flap or change 584) can lead to peak times for swap transactions (580) and broadcast transactions (581) in high-performance applications around 6:41 PM.

[0036] Fig. Figure 5F shows an example display screen that provides a visualization, including network dropout events and connection events, according to one aspect of the present application. In charts 586, 592, and 596, the x-axis displays time in 15-minute increments from 06:45 to 10:30. Local Connection_A and Local Connection_B can represent local connections in, for example, a dragonfly topology, while Global Connection_A can represent a global network connection in, for example, a dragonfly topology. The in Fig. The connections described in section 5F are used for illustrative purposes only. Other connections and network topologies can also be used. Fig. 5F shows an example of extracting standard state events in addition to log extraction to determine a relevant period (i.e., the specified time window) for detecting anomalous behavior.

[0037] Diagram 586 shows measurements related to network failure events 587, where: the y-axis indicates a number of failure events (e.g., a number of failed packets); the partially shaded points indicate failure events for a local fabric link (local link_A 588); the bold-bordered points indicate failure events for a local fabric link (local link_B 589); the solid-colored points indicate failure events for a global fabric link (global link_A 590); and the other points indicate failure events for other links (other links 591). Note that the other points, shown as other links, may represent separate local or global fabric links and are shown with the same label in Diagram 586 for illustrative purposes.Individual colors, labels, formatting, or other identifiers can be used to distinguish each of the other separate local or global Fabric links.

[0038] Diagram 592 shows measurements related to global link-flap events 593, where: the y-axis indicates a number of link flaps (e.g., at a specific time); the solid-colored dots represent link flaps for global Link_A 594; and other dots represent link flaps for other global links 595. Diagram 596 shows measurements related to local link-flap events 597, where: the y-axis indicates a number of link flaps (e.g., at a specific time); the solid-colored dots represent link flaps for local Link_A 598; and the dots with bold outlines represent link flaps for local Link_B 599.

[0039] Based on Fig. 5F allows for a correlation between network packet loss and link flaps at different levels (i.e., local link_A, local link_B, and global link_A). For example, at approximately 7:00 AM, a link flap for local link_A (as shown in diagram 596) can lead to the network packet loss shown for local link_A at the same time (as shown in diagram 586). Similarly, at approximately 8:17 AM, a link flap for local link_B (as shown in diagram 596) can lead to the network packet loss shown for local link_B at the same time (as shown in diagram 586). Furthermore, at approximately 9:38 AM, a link flap for global link_A (as shown in diagram 592) can lead to the network packet loss shown for global link_A at the same time (as shown in diagram 586).

[0040] Fig. 6A and Fig. Figure 6B shows flowcharts 600 and 630, illustrating a procedure that facilitates intelligent log analysis for large HPC and AI systems, according to one aspect of the present application. During operation, the system receives event information from components working together in the system. This information comprises a first set of events interpreted from log entries associated with the components and a second set of events returned from queries for standard events (Operation 602). Log agents running on various components can generate the log entries that represent the first set of events by extracting logs from one or more of the components in the system, as described above with respect to the log extraction / collection modules 414, 434, 454, and 474 of log agents 412, 432, 452, and 472. Fig. 4 or the protocol extraction module of Fig. 2 described. The log agents can remove noise from the extracted logs by filtering the extracted logs, and can also obtain reformatted log entries by reformatting the filtered logs. For example, log entries 310, 320, and 330 of Fig. 3A can be obtained after the filtering and reformatting described above (also as above with regard to the protocol filter module 212 of Fig. 2 described). The log agents can generate event information based on characteristics of the newly formatted log entries, as described above with respect to event information 315, 325 and 335 in Fig. 3A described.

[0041] The system classifies the events interpreted from the log entries based on a hierarchy of components (Operation 604). The Log Analytics Orchestrator 401 from Fig. 4 can, for example, use a decision tree, as described above in relation to Fig. 3C is shown.

[0042] The system correlates two or more events based on a respective event classification and a predetermined time window covering an event time associated with that event, with the event time being derived from the log entries (Operation 606). The predetermined time window can be determined from measurements of power consumption, application runtime, and transaction results of the components. For example, for a given time window, two or more events with a corresponding classification that occur during the same time window can be correlated, as above with respect to the NIC flapping errors (526) in NIC event 522 and the anomalous measurements (504) of host transaction A 502 in the visual representations of Fig. 5A described, as well as the above in Fig. 5B-5F.

[0043] The system stores information associated with the first and second event sets in entries within a data structure, where each entry specifies the determined event classification and any correlations to other events (operation 608). The system can store the information before classifying or correlating the events (as in operations 604 and 606, respectively). The information can be stored in a format similar to that described above for log entries 310, 320, and 330. Fig. The system is similar to that described in 3A. It can store these entries in a time-series database, for example, in the time-series database 407 of the Log Analytics Orchestrator 401 in Fig. 4.

[0044] The system determines whether to query the data structure directly or extract additional information (Decision 610). The system can make this decision based on a previously defined configuration that specifies whether additional information, such as performance metrics, should be used to determine the first predetermined period or identify the relevant period. If the system decides to query the data structure directly (Decision 610), it queries the data structure for events associated with a first predetermined period (Operation 612). The first predetermined period can be based on measurements related to power consumption, application runtime, and transaction results associated with the components.The system correlates the queried events by marking the respective entries for the queried events with the same correlation identification tag (operation 614). The system can also correlate the queried events by linking entries using pointers or other relational operations. This operation is labeled A in . Fig. Continued in 6B.

[0045] When the system decides to extract additional information (decision 610), it extracts performance and application metrics over a time window (operation 616), such as power utilization and application metrics associated with the components in the system during a specific time window. This can identify relevant periods with anomalous measurements, as described above in relation to Module 403 of Log Analytics Orchestrator 401. Fig. 4. The system identifies a relevant time period within the time window, for example, due to a performance drop, an increase in application runtime, or slower application measurements (e.g., slower than a predefined threshold) (Operation 618). The factors listed here as the basis for determining a relevant time period are for illustrative purposes only. Other factors can also be used. The system can use the identified relevant time period as the first predetermined time period, and the operation continues with Operation 610.

[0046] Fig. 6B shows a continuation of the processes from Fig. 6A following process 614. The system generates a visual representation that displays the correlated events (process 632). The visual representation can display the correlated events and the queried correlated events (from process 612). The visual representation can include charts that show a measurement (such as a time span or a number of errors) over a period of time, e.g., as in the charts of the Fig. 5A-5F. The system generates a report based on the correlated events (Operation 634). The system can display the report, and the report can contain one or more interactive elements that facilitate viewing or editing the displayed information, including, but not limited to: a detected anomaly; a recommended action indicating how to correct the detected anomaly; or a configurable option indicating that the computer should automatically perform the recommended action. The system can also perform an initial action based on the displayed report. The initial action can be a corrective action performed by a user connected to the system, or the initial action can be an action performed automatically by the system based on previously configured options for automatically accepting or executing recommended actions.

[0047] If the visual representation does not indicate an anomaly (Decision 636), the process returns. If the visual representation indicates an anomaly (Decision 636), the system allows corrective actions related to the displayed anomaly (Operation 638). For example, in response to charts 500 and 520, which indicate an anomaly based on the displayed measurements and correlated events, a user can take a corrective action to resolve the displayed anomaly, such as restarting a network card, removing a job or pausing a host transaction, removing or replacing a node or other hardware component, etc. In some respects, operations 616 and 618 can be performed by a user in response to the display of the generated visual representation or report.This means that by viewing the visual representation or report, the user can identify a relevant period within a specific timeframe based on the extracted and displayed performance and application metrics. The user (or the system) can then query the data structure for events within the identified relevant period and correlate the queried events (as described above in relation to operations 612 and 614 of [reference missing]). Fig. (described in 6A). Furthermore, based on the displayed report (generated in Operation 634 as described above), the user can take corrective action, such as based on a recommended action indicating the resolution of a detected anomaly. For example, the user can replace a network card that has been found to be associated with anomalous activity in a host transaction. The system can also perform other corrective actions, including inputting information about the correlated events into an external system to train a machine learning model. Anomalous activity or anomalies can be displayed in the visual representation if the measured values ​​for a corresponding event exceed a predefined benchmark or other threshold.

[0048] Fig. Figure 7 shows a computer system 700 that enables intelligent protocol analysis for large HPC and AI systems according to one aspect of the present application. The computer system 700 comprises a processor 702, a memory 704, and a storage device 706. The memory 704 can include volatile memory (e.g., random-access memory (RAM)) that serves as managed memory and can be used to store one or more memory pools. In addition, the computer system 700 can be connected to peripheral I / O user devices 710 (e.g., a display device 711, a keyboard 712, and a pointing device 713). The storage device 706 contains a non-transferable, computer-readable storage medium and stores an operating system 716, instructions 718, and data 730. The computer system 700 can have fewer or more units or instructions than those shown in Figure 706. Fig. 7 are included.

[0049] Instructions 718 may contain instructions which, when executed by computer system 700, can cause computer system 700 to perform the procedures and / or processes described in this disclosure. In particular, instructions 718 may contain instructions 720 to obtain event information from components operating together in a network environment, indicating a first set of events interpreted from the log entries associated with the components, and a second set of events returned from queries for standard events, as above with respect to operation 602 of Fig. 6A and protocol entries 310, 320 and 330 of Fig. 3A described.

[0050] Instructions 718 may contain instructions 722 to classify the events interpreted from the log entries based on a topology of components in the network environment, as above in relation to the event classification module 216 of Fig. 2, the decision tree 368 of Fig. 3C and Operation 604 of Fig. 6A described.

[0051] Instructions 718 can contain instructions 724 to correlate two or more events based on a respective event classification and a predetermined time window covering an event time associated with a respective event, where the event time is derived from the log entries and where the predetermined time window is determined from measurements relating to energy consumption, application runtime, and transaction results associated with the components, as above in relation to operation 604 of Fig. 6A and the diagrams of Fig. 5A-F.

[0052] Instructions 718 may include instructions 726 for generating a visual representation (and report) showing the correlated events, as above in relation to operations 632 / 634 and decision 636 of Fig. 6B and the diagrams of the FIGS are described. 5A-F.

[0053] Instructions 718 may include instructions 728 to enable corrective actions to resolve the displayed anomaly in response to the visual indication, as described above in relation to the operations in Fig. 6B described.

[0054] Instructions 718 can contain more instructions than those in Fig. The 7 shown contain instructions. For example, instructions 718 may contain instructions for performing the operations described above, specifically regarding: the high-level flow of Fig. 2; the capture, formatting, conversion and classification of log entries into Fig. 3A-C; the environment and communication of Fig. 4; the diagrams of Fig. 5A-F; those shown in the flowcharts of the Fig. 6A and Fig. The processes shown in 6B, as well as the instructions of the CRM 800 in Fig. 8.

[0055] Data 730 may contain all data required as input or generated as output by the methods, operations, communications, and / or processes described in this disclosure. In particular, Data 730 may store at least the following: event information; an entry; a first set of events interpreted from log entries; a second set of events returned from queries for standard events; a classification; an event classification; a correlation between two or more events; a time window; an event time; a visual representation; a report; an indicator of an anomaly; an indicator or identifier for hardware, software, or any other component in a system or associated with storage components, host components, or fabric components in the system; raw logs or log data; an extracted log; noise;a filtered log; a reformatted log entry; a log entry attribute; an entity or component identity; a time; an event category; an event type; an event description; a data structure; information; correlated events or correlated queried events; a report; an indicator or recommendation for action or corrective action; and an interactive element that facilitates viewing or manipulation of displayed information, including a detected anomaly, a recommended action, and a configurable option.

[0056] Fig. Figure 8 shows a computer-readable medium (CRM) 800 that enables intelligent log analysis for large HPC and AI systems according to one aspect of the present application. CRM 800 can be a non-transitory computer-readable medium or device that stores instructions which, when executed by a computer or processor, cause the computer or processor to perform a procedure. CRM 800 can store instructions 810 to obtain event information from components operating together in a system, indicating a first set of events interpreted from the log entries associated with the components, and a second set of events returned from queries for standard events, as above with respect to operation 602 of Fig. 6A and protocol entries 310, 320 and 330 of Fig. 3A described.

[0057] CRM 800 can store instructions 812 to classify the events interpreted from the log entries based on a hierarchy of components, as above in relation to the event classification module 216 of Fig. 2 and process 604 of Fig. 6A described.

[0058] CRM 800 can store instructions 814 to correlate two or more events based on a respective event classification and a predetermined time window covering an event time associated with a respective event, where the event time is derived from log entries and the predetermined time window is determined from measurements relating to energy consumption, application runtime, and transaction results associated with the components, as described above in relation to operation 604 of Fig. 6 and the diagrams of the Fig. 5A-C. CRM 800 can retrieve or extract the application host's power consumption and transaction metrics to identify deviations and anomalies, for example, in certain relevant time periods, as above in relation to operations 616 and 618 of Fig. 6A described.

[0059] CRM 800 can store instructions 816 to generate a visual representation or report showing the correlated events, as above in relation to operations 632 and 634 of Fig. 6B and the diagrams of the FIGS are described. 5A-F.

[0060] CRM 800 can store instructions 818 to enable corrective actions in response to the visual display or report indicating an anomaly, addressing the displayed anomaly, as described above in relation to the operations of Fig. 6B described.

[0061] CRM 800 can handle more instructions than those in Fig. The 8 shown are included. CRM 800 can, for example, store instructions for executing the operations described above, specifically regarding: the high-level flow of Fig. 2; the capture, formatting, transformation and classification of log entries into Fig. 3A-C; the environment and communication of Fig. 4; the diagrams of Fig. 5A-F; those shown in the flowcharts of the Fig. 6A and Fig. 6B depicted processes and instructions 718 of computer system 700 in Fig. 7.

[0062] The described aspects can thus enable improved anomaly detection in complex systems and enhanced root cause analysis. They can also facilitate more efficient identification of relationships between events in different subsystems and more efficient handling of various log formats and event types. Furthermore, these aspects can provide interactive user feedback for system optimization.

[0063] In general, the disclosed aspects provide a method, a computer system, and a computer-readable medium that enable intelligent log analysis for large HPC and AI systems. During operation, the system receives event information from components working together within the system. This information comprises a first set of events interpreted from log entries associated with the components and a second set of events returned from queries for standard events. The system classifies the events interpreted from the log entries based on a hierarchy of the components.The system correlates two or more events based on a specific event classification and a predetermined time window covering the event time associated with each event. The event time is derived from log entries, and the predetermined time window is determined from measurements related to power consumption, application runtime, and transaction results associated with the system components. The predetermined time window can also be determined based on the detection of errors and events in the system components. This determined time window can be used to search for fluctuations in application performance and performance anomalies. The system generates a visual representation displaying the correlated events. If the visual representation indicates an anomaly, the system enables corrective actions to address it.

[0064] In one variation of this aspect, the components comprise at least one of the following: hardware or software connected to storage components in the system; hardware or software connected to host components in the system, wherein the host components comprise one or more of the following: a graphics processing unit (GPU), high-bandwidth memory (HBM), a central processing unit (CPU) or core, CPU memory, and a PCIe (Peripheral Component Interconnect Express) component; or hardware or software connected to fabric components of the system, wherein the fabric components comprise one or more of the following: a network device, a switch, a switch agent, a centralized fabric manager, a fabric agent, and a network interface.

[0065] In another variation of this aspect, the system generates the log entries that display the first set of events by: extracting logs from one or more of the components in the system; removing noise in the extracted logs by filtering the extracted logs; obtaining reformatted log entries by reformatting the filtered logs; and generating event information based on properties of the reformatted log entries.

[0066] In another variant, the characteristics of the newly formatted log entries include at least one of the following: the identity of an entity or component associated with the log entry; a time associated with an event that generated the log entry; an event category; an event type; or a description of the event.

[0067] In another variant, the system stores information associated with the first and second sets of events in entries in a data structure and in a time series database, with each entry indicating the determined event classification and any correlations to other events.

[0068] In another variant, the system queries the data structure for events associated with a first predetermined time period. This first predetermined time period is based on at least one of the following: measurements related to energy consumption, application runtime, and transaction results associated with the components; or the detection of errors and events in the system's components. The system correlates the queried events by marking the respective entries for the queried events with the same correlation identification tag. The system then incorporates the correlated queried events into the generated visual representation.

[0069] In another variant, the system generates a report based on the correlated events and displays it. The system then performs an initial action based on the displayed report, which includes a corrective measure to resolve the identified anomaly.

[0070] In another variant, the displayed report contains one or more interactive elements that facilitate viewing or manipulating the displayed information, including at least one of the following: a detected anomaly; a recommended action indicating how to correct the detected anomaly; or a configurable option indicating that the computer should automatically perform the recommended action.

[0071] In another aspect, a computer system comprises a processor and a storage device that holds instructions. These instructions are used to retrieve event information from components operating together in a network environment. This information consists of a first set of events interpreted from log entries associated with the components, and a second set of events returned from queries for standard events. The instructions further classify the events interpreted from the log entries based on the topology of the components in the network environment. Finally, the instructions store the log entries in a time-series database.The instructions further serve to correlate two or more events based on a respective event classification and a predetermined time window covering an event time associated with each event. The event time is derived from the log entries, and the predetermined time window is determined from measurements relating to power consumption, application runtime, and transaction results associated with the components. The instructions also serve to generate a visual representation displaying the correlated events. Furthermore, the instructions serve to enable corrective actions in response to any anomalies displayed in the visual representation. The computer system may contain additional instructions to perform the operations described herein, including those relating to the high-level flow of [missing information]. Fig. 2; the collection, formatting, conversion and classification of log entries from Fig. 3A-C; the environment and communication of Fig. 4; the diagrams of Fig. 5A-F; those shown in the flowcharts of the Fig. 6A and Fig. The processes shown in 6B, as well as the instructions of the CRM 800 in Fig. 8.

[0072] In another aspect, a non-transferable, computer-readable storage medium (or CRM) stores instructions for retrieving event information from components working together in a system. This information comprises a first set of events interpreted from log entries associated with the components and a second set of events returned from queries for standard events. The instructions also serve to classify the events interpreted from the log entries based on a hierarchy of components.The instructions further serve to correlate two or more events based on a respective event classification and a predetermined time window covering an event time associated with each event. The event time is derived from log entries, and the predetermined time window is determined from measurements related to energy consumption, application runtime, and transaction results associated with the components. The instructions also generate a visual representation or report displaying the correlated events. Furthermore, in response to the visual representation or report displaying an anomaly, the instructions enable corrective actions to address the displayed anomaly. The CRM can also store instructions for executing the operations described above with respect to the high-level flow of [missing information]. Fig. 2; the collection, formatting, transformation, and classification of log entries into Fig. 3A-C; the environment and communication of Fig. 4; the diagrams of Fig. 5A-F; those shown in the flowcharts of Fig. 6A and Fig. 6B depicted processes; and the instructions 718 of computer system 700 in Fig. 7.

[0073] The foregoing description is intended to enable the person skilled in the art to produce and use the aspects and examples and is given in connection with a specific application and its requirements. Various modifications of the disclosed aspects are readily apparent to the person skilled in the art, and the general principles defined herein can be applied to other aspects and applications without departing from the spirit and scope of this disclosure. Therefore, the aspects described here are not limited to those shown but have the broadest possible scope consistent with the principles and features disclosed herein.

[0074] Furthermore, the foregoing descriptions of the aspects serve only for illustration and description. They do not claim to be exhaustive and do not limit the aspects described herein to the disclosed forms. Accordingly, many modifications and variations will be obvious to those skilled in the art. Moreover, the above disclosure is not intended to limit the aspects described herein. The scope of the aspects described herein is defined by the accompanying claims.

Claims

[1] A procedure comprising the following: Receiving event information from components working together in a system, specifying a first set of events interpreted from the log entries associated with the components, and a second set of events returned from queries for standard events; Classifying the events interpreted from the log entries based on a hierarchy of components; Correlating two or more events based on a respective event classification and a predefined time window that covers an event time assigned to a respective event, the event time derived from the log entries and the predetermined time window, which is determined from measurements of energy consumption, application runtime and the transaction results assigned to the components; Generating a visual representation that displays the correlated events; and In response to the visual display indicating an anomaly, enabling corrective action to rectify the indicated anomaly. [2] The method of claim 1, wherein the components comprise at least one of the following elements: Hardware or software that is assigned to the system's storage components; Hardware or software that is assigned to host components in the system, wherein the host components comprise one or more of a graphics processing unit (GPU), high-bandwidth memory (HBM), central processing unit (CPU) or core, CPU memory, and a PCIe (Peripheral Component Interconnect Express) component; or Hardware or software that is associated with fabric components of the system, wherein the fabric components include one or more of a network device, a switch, a switch agent, a centralized fabric manager, a fabric agent, and a network interface. [3] The method of claim 1, which further comprises generating the log entries indicating the first group of events by: Extracting logs from one or more components of the system; Removing noise from the extracted protocols by filtering the extracted protocols; Obtaining newly formatted log entries by reformatting the filtered logs; and Generating event information based on the characteristics of the newly formatted log entries. [4] Method according to claim 3, wherein the features of the newly formatted log entries include at least one of the following features: Identity of a unit or component associated with the log entry; a time associated with an event that generated the log entry; an event category; an event type; or a description of the event. [5] The method according to claim 1 further comprises: Storing information associated with the first and second sets of events in entries in a data structure and in a time series database, where a corresponding entry specifies the determined event classification and any correlations to other events. [6] The method of claim 5, which further comprises: Queries of the data structure for events that are associated with a first predetermined time period, where the first predetermined time period is based on at least one of the following: Measurements of energy consumption, application runtime, and transaction results assigned to the components; or Detecting errors and events in all components of the system; Correlate the queried events by marking the respective entries for the queried events with the same correlation identification tag; and Incorporating the queried correlated events into the generated visual representation. [7] The method of claim 1, which further comprises: Creating a report based on the correlated events; Displaying the report; and Taking initial action based on the displayed report, the first measure includes an appropriate corrective action to rectify the indicated anomaly. [8] Method according to claim 7, wherein the displayed report contains one or more interactive elements that facilitate viewing or manipulating the displayed information, including at least one of the following elements: a discovered anomaly; a recommendation for action to correct the identified anomaly; or a configurable option that indicates that the computer should automatically perform the recommended action. [9] A computer system comprising the following: a processor; and a storage device that stores instructions which, when executed by the processor, contain instructions to: to obtain event information from components that work together in a network environment, specifying a first set of events interpreted from the log entries associated with the components, and a second set of events returned from queries for standard events; to classify the events interpreted from the log entries based on a topology of the components in the network environment; to correlate two or more events based on a respective event classification and a predetermined time window that covers an event time assigned to a respective event, wherein the event time is derived from the log entries and wherein the predetermined time window is determined from measurements relating to energy consumption, application runtime and transaction results associated with the components; to create a visual representation that shows the correlated events; and to enable corrective action to be taken in response to the visual display indicating an anomaly. [10] Computer system according to claim 9, wherein the components comprise at least one of the following elements: Hardware or software that is assigned to storage components in the network environment; Hardware or software associated with host components in the network environment, wherein the host components comprise one or more components, namely a graphics processing unit (GPU), high-bandwidth memory (HBM), a central processing unit (CPU) or core, CPU memory, and a PCIe (Peripheral Component Interconnect Express) component; or Hardware or software associated with fabric components of the network environment, wherein the fabric components include one or more of the following: a network device, a switch, a switch agent, a centralized fabric manager that manages switches in the fabric, a fabric agent that operates on a switch, the fabric agent programming the switch and interacting with network protocol agents, and a network interface. [11] Computer system according to claim 9, wherein the instructions further comprise: Extracting logs from one or more components in the network environment; Removing noise from the extracted protocols by filtering the extracted protocols; Obtaining reformatted log entries by reformatting the filtered logs; and Generating event information based on the characteristics of the newly formatted log entries. [12] Computer system according to claim 11, wherein the features of the newly formatted log entries include at least one of the following features: Identity of a unit or component associated with the log entry; a time associated with an event that generated the log entry; an event category; an event type; or a description of the event. [13] Computer system according to claim 9, wherein the instructions further comprise: Storing information associated with the first and second sets of events in entries in a data structure and in a time series database, where a corresponding entry specifies the determined event classification and any correlations to other events. [14] Computer system according to claim 13, wherein the instructions further comprise: Queries of the data structure for events that are associated with a first predetermined time period, where the first predetermined time period is based on measurements relating to energy consumption, application runtime, and transaction results associated with the components; Correlate the queried events by marking the respective entries for the queried events with a suitable correlation tag; and Incorporating the queried correlated events into the generated visual representation. [15] Computer system according to claim 9, wherein the instructions further comprise: Creating a report based on the correlated events; Displaying the report; and Taking initial action based on the displayed report, the first measure includes an appropriate corrective action to rectify the indicated anomaly. [16] Computer system according to claim 15, wherein the displayed report contains one or more interactive elements that facilitate viewing or manipulating the displayed information, including at least one of the following elements: a discovered anomaly; a recommendation for action to correct the identified anomaly; or a configurable option that indicates that the computer should automatically perform the recommended action. [17] Computer system according to claim 15, wherein the instructions further comprise: in response to allowing corrective action to address the anomaly shown in the visual representation or taking the first action based on the displayed report: Receive updated event information from the components; Classify the updated events listed in the updated events information; two or more events correlate based on the updated events, a respective event classification, and the predetermined time window; to regenerate the visual representation that displays the correlated events; and in response to the newly generated visual representation, which indicates one or more other anomalies, allow further corrective actions affecting one or more other anomalies. [18] A non-transitory computer-readable medium that stores instructions to: to obtain event information from components working together in a system, specifying a first set of events interpreted from the log entries associated with the components, and a second set of events returned from queries for standard events; to classify the events interpreted from the log entries based on a hierarchy of components; to correlate two or more events based on a respective event classification and a predetermined time window covering an event time associated with a respective event, wherein the event time is derived from the log entries and the predetermined time window is determined from measurements relating to energy consumption, application runtime and transaction results associated with the components; to generate a visual representation or report that displays the correlated events; and to enable corrective action to be taken in response to the visual display or report indicating an anomaly. [19] Non-transient, computer-readable medium according to claim 18, wherein the instructions further generate the log entries indicating the first set of events by: Extract logs from one or more components of the system; Remove noise from the extracted logs by filtering the extracted logs; Newly formatted log entries are obtained by reformatting the filtered logs; and Generate event information based on the characteristics of the newly formatted log entries. [20] Non-transitory, computer-readable medium according to claim 18, wherein the instructions further comprise: Display the visual representation or report, wherein the displayed visual representation or report contains one or more interactive elements that facilitate viewing or manipulation of the displayed information, the displayed information contains at least one of the following elements: a discovered anomaly; a recommendation for action to correct the identified anomaly; or a configurable option that indicates that the computer should automatically perform the recommended action; and to allow the corrective measures to rectify the identified anomaly: Receive updated event information from the components; Classifying the updated events listed in the updated events information; Correlating two or more events based on the updated events, a respective event classification, and the predetermined time window; Regenerating the visual representation that displays the correlated events; and in response to the newly generated visual representation, which indicates one or more other anomalies, allow further corrective actions affecting one or more other anomalies.