Unsupervised Auto Clustering for Network Log Diagnosis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Analyzing large volumes of system log data to diagnose network problems in online communications platforms is a time-consuming and difficult task for network administrators, as it typically involves manual processes and requires expertise in both standard queries and machine learning analysis.
Innovation Solution
A data processing system that includes a user interface for constructing queries and creating visualizations, utilizing machine learning algorithms to automatically identify clusters of data indicative of network issues, allowing administrators to refine query results and drill down into specific causes of performance problems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual processes are used to analyze system log data, then analysis accuracy can be maintained through expert judgment, but the time and effort required increases significantly
Solution Approach 1:
The patent introduces an automated analysis system that acts as an intermediary between the raw log data and the network administrator. This system includes components for automatically parsing logs, generating queries, analyzing results, and identifying root causes, thereby reducing the time required while maintaining diagnostic accuracy through structured automated processes.
Solution Approach 2:
The patent replaces the manual mechanical process of log analysis with an automated computational system. The system automatically performs data parsing, query generation, result analysis, and root cause identification, substituting human manual operations with automated algorithms and processing mechanisms.
2Productivity
If automated machine learning algorithms are used to analyze performance data, then analysis speed increases, but the complexity of the system increases
Solution Approach 1:
The patent divides the automated analysis system into distinct functional modules: a data parsing module for processing logs, a query generation module for creating analysis queries, a result analysis module for interpreting data, and a root cause identification module for determining problems. This segmentation manages complexity by organizing functions into separate, manageable components.
Solution Approach 2:
The automated analysis system is designed to handle multiple types of performance data and various diagnostic scenarios through a unified platform. The system can analyze different log formats, generate multiple query types, and identify various root causes, providing multi-functional capability that justifies the complexity through broad applicability.
3Measurement precision
If administrators manually formulate queries to analyze terabytes of log data, then query precision can be optimized for specific problems, but the effort and expertise required increases
Solution Approach 1:
The system performs self-service by automatically generating queries based on the performance data and diagnostic needs. The query generation module autonomously creates appropriate queries without requiring manual formulation by administrators, thereby maintaining query precision while significantly reducing the effort and expertise required for operation.
Solution Approach 2:
The system performs preliminary actions by automatically parsing logs, generating queries, and analyzing results before presenting findings to administrators. This preliminary automated processing reduces the manual effort required while maintaining precision through systematic pre-analysis.
Data Source
AI summary
Techniques performed by a data processing system for diagnosing problems with a communications platform include obtaining query parameters including an aggregation operator for invoking a machine learning algorithm configured to analyze performance data for the communications platform, automatically executing the query on the performance data to obtain query results by invoking the machine learning algorithm on the performance data to automatically identify a plurality of clusters of data indicative of a performance problem, and presenting a visualization of the query results. The visualization includes indicators identifying cluster properties for which the query results are further refinable and one or more second indicators identifying the second subset of the second cluster properties which are not relevant for further refining the first query results. The indicators are actuatable to automatically update and re-execute the first query based on the respective indicator that is actuated.


