Adaptive Log Level Control for Remote Root Cause Diagnosis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing complexity of computer systems and the exponential growth of data collected for diagnosing anomalies lead to network bandwidth, storage, and computing power challenges, necessitating improved methods for remote management and support.
Innovation Solution
Implementing a machine learning model to autonomously control log data levels, adjusting them based on thresholds and probabilities to minimize data impact while ensuring sufficient diagnostic data is collected, particularly in network systems experiencing service level experience degradation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If log data level is increased to collect more diagnostic data, then diagnostic capability is improved, but network bandwidth consumption increases
Solution Approach 1:
The patent implements dynamic log data level adjustment where the logging system automatically modifies log data levels based on real-time system conditions, anomaly detection results, and diagnostic needs. This allows the system to collect detailed logs only when anomalies are detected or during critical diagnostic phases, while using minimal logging during normal operation, thereby resolving the contradiction between diagnostic capability and bandwidth consumption.
Solution Approach 2:
The system changes the parameter of log data level dynamically based on system state. By adjusting this parameter according to anomaly detection confidence levels, system performance metrics, and diagnostic priorities, the system optimizes the balance between collecting sufficient diagnostic information and minimizing network bandwidth usage.
2Measurement precision
If log data level is increased to ensure sufficient diagnostic data, then diagnostic accuracy is improved, but storage requirements increase
Solution Approach 1:
The patent implements dynamic log data level adjustment where the logging system automatically modifies log data levels based on real-time system conditions, anomaly detection results, and diagnostic needs. This allows the system to collect detailed logs only when anomalies are detected or during critical diagnostic phases, while using minimal logging during normal operation, thereby resolving the contradiction between diagnostic capability and bandwidth consumption.
Solution Approach 2:
The system changes the parameter of log data level dynamically based on system state. By adjusting this parameter according to anomaly detection confidence levels, system performance metrics, and diagnostic priorities, the system optimizes the balance between collecting sufficient diagnostic information and minimizing network bandwidth usage.
3Reliability
If log data level is increased to diagnose anomalies effectively, then system reliability is improved, but system performance overhead increases
Solution Approach 1:
The patent implements dynamic log data level adjustment where the logging system automatically modifies log data levels based on real-time system conditions, anomaly detection results, and diagnostic needs. This allows the system to collect detailed logs only when anomalies are detected or during critical diagnostic phases, while using minimal logging during normal operation, thereby resolving the contradiction between diagnostic capability and bandwidth consumption.
Solution Approach 2:
The system performs preliminary anomaly detection using lightweight monitoring mechanisms before initiating high-level logging. By detecting potential issues early through subtle performance deviations, resource usage patterns, or error trends, the system can activate detailed logging only when necessary, preventing performance overhead during normal operation while maintaining system reliability.
Data Source
AI summary
Disclosed are embodiments for improving remote diagnostics of a computer system. Some embodiments obtain operational parameter values and log data from a plurality of network devices, and provide the operational parameter values and log data to a machine learning model. The model is trained to identify a root cause of a degradation of the computer system based on the operational parameter values and log data, and to provide recommendations of log data level settings for the network devices. If the model identifies a root cause of the degradation with sufficient confidence, a remedial action is identified and applied to the computer system. If the confidence level is insufficient, log data level settings of the network devices are modified based on the recommendations of the model. This process may be performed iteratively until a root cause is identified with sufficient confidence.


