Template-Ranked Critical Logs for Cross-Layer Network Troubleshooting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In data center environments, troubleshooting performance issues of network applications is inefficient due to the independent development and management of layers with different administrative domains, requiring extensive resource consumption and complexity in identifying root causes.
Innovation Solution
A unified framework that analyzes cross-layer logs using a template mining model to map and rank candidate logs, selecting critical logs based on properties and dependencies, reducing the number of logs needing analysis to pinpoint root causes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If all cross-layer logs are analyzed to identify root causes of performance issues, then the completeness of troubleshooting is improved, but the resource consumption and complexity increase significantly
Solution Approach 1:
The patent segments the log analysis process into multiple stages: first identifying candidate logs from specific layers (application, network, compute) associated with performance issues, then selecting critical logs from these candidates based on additional criteria. This segmentation divides the overwhelming task of analyzing all logs into manageable phases, reducing complexity while maintaining troubleshooting completeness.
Solution Approach 2:
The patent extracts and identifies a subset of candidate logs from the entire log set by focusing on logs from specific layers associated with performance issues. This extraction process separates the relevant logs from irrelevant ones, reducing the analysis scope to only those logs that could potentially contain root cause information, thereby reducing complexity without sacrificing reliability.
2Measurement precision
If all cross-layer logs are collected and analyzed, then the accuracy of root cause identification is improved, but the time and resources required for analysis increase
Solution Approach 1:
The patent performs preliminary actions by first identifying candidate logs from specific layers before conducting the actual root cause analysis. This preliminary filtering step prepares the data in advance by organizing and selecting only the relevant logs, so that when analysis is performed, it is done on a reduced, pre-processed set, saving time while maintaining accuracy.
Solution Approach 2:
The patent applies partial action by analyzing only a subset of logs (candidate logs and then critical logs) rather than all available logs. This partial analysis is sufficient to achieve accurate root cause identification because the selection process ensures that the most relevant logs are prioritized, making full log analysis unnecessary and thus reducing time loss.
3Ease of operation
If logs from multiple layers are analyzed independently, then the simplicity of management is improved, but the effectiveness of cross-layer troubleshooting deteriorates
Solution Approach 1:
The patent merges the analysis of logs from multiple layers (application, network, compute) into a unified cross-layer analysis framework. By combining these layers and identifying associations between them, the system maintains the simplicity of individual layer management while achieving effective cross-layer troubleshooting through the integrated candidate log identification process.
Data Source
AI summary
Techniques are described for a computing system configured to obtain a plurality of candidate logs for a plurality of layers of a computing infrastructure. The computing system may, for each candidate log of the plurality of candidate logs, map the candidate log to a log template of a plurality of log templates, wherein each log template to which a candidate log is mapped is a mapped log template. The computing system may rank mapped log templates based on properties of the mapped log templates. The computing system may select, based on the ranking, one or more candidate logs as critical logs. The computing system may output at least one of (1) an indication of the critical logs to determine a potential root cause associated with a performance issue of a network application or (2) an indication of the potential root cause associated with the performance issue of the network application.


