Automated Log Clustering via Hash-Based Categorization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional log analytics tools face challenges in efficiently collecting and analyzing log records from large-scale computing systems due to their rudimentary abilities, inefficiencies in scaling, and the need for manual construction and maintenance of log parsers, which require significant time and resources.
Innovation Solution
An automated system that constructs a categorizer to parse and categorize machine-generated data records, using grammar rules to identify variable and non-variable components, allowing for automatic clustering and dynamic tracking of clusters with alert generation, and enabling efficient resource utilization and sharing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional log analytics tools are used to collect and analyze log records from large-scale computing systems, then basic log analysis capability is provided, but the system cannot efficiently scale when posed with massive systems involving large numbers of computing systems and applications
Solution Approach 1:
The patent merges log collection and analysis capabilities across multiple computing systems into a unified cloud-based platform. By consolidating parser construction, log collection, and analysis functions into a centralized system that serves multiple hosts, the solution achieves scalable efficiency without requiring redundant per-host resources.
Solution Approach 2:
The patent creates a universal log analytics platform that can handle diverse log formats from multiple computing systems and applications through a single system. The automated parser construction capability allows the platform to adapt to various log formats universally, eliminating the need for system-specific configuration and enabling scalable deployment across massive systems.
2Ease of operation
If conventional approaches are used where set-up and configuration activities are performed each time a new host is added or configured, then individual host log collection is achieved, but extensive redundant processing and resource usage occurs
Solution Approach 1:
The patent performs preliminary parser construction and configuration activities once, and then reuses these parsers across multiple hosts and computing systems. By pre-building the parsing capabilities and storing them centrally, the system eliminates the need to repeatedly perform configuration activities when new hosts are added, significantly reducing redundant processing and resource consumption.
Solution Approach 2:
The patent creates reusable parser templates and configuration artifacts that can be copied and applied across multiple computing systems. Instead of performing configuration activities from scratch for each host, the system copies and adapts proven parser configurations, eliminating redundant work while maintaining host-specific customization capabilities.
3Reliability
If on-premise solutions are used for log analytics, then local log processing is performed, but resource sharing and analysis component utilization is inadequate
Solution Approach 1:
The patent introduces a cloud-based intermediary platform that sits between local computing systems and log analysis operations. This intermediary collects logs from multiple hosts, performs centralized parsing and analysis, then returns results to local systems. This architecture maintains local processing reliability while enabling resource sharing and component utilization across the entire system through the cloud intermediary.
Solution Approach 2:
The patent transitions from a single-dimension local on-premise processing model to a multi-dimensional architecture that combines local log collection with centralized cloud-based analysis. By adding the cloud dimension, the system maintains local reliability for data collection while enabling resource sharing, scalability, and improved analysis capabilities through centralized components that serve multiple hosts simultaneously.
4Manufacturing precision
If log parsers are manually constructed by skilled personnel, then accurate log parsing is achieved, but significant time and resources are required to build and maintain parsers
Solution Approach 1:
The patent implements automated parser construction that enables the system to generate its own log parsers without requiring skilled personnel to manually build them. The automated system analyzes log formats, constructs appropriate parsers, and maintains them automatically, achieving accurate parsing while eliminating the significant time and resource investment previously required for manual parser construction and maintenance.
Solution Approach 2:
The patent transforms the parser construction process from a manual, skill-dependent activity into an automated parameter-driven process. By changing the parameters of the system from human expertise to automated algorithms that can learn and adapt to log formats, the solution maintains parsing accuracy while dramatically reducing the time and resources required to build and maintain parsers.
Data Source
AI summary
Some embodiments relate to assigning individual log messages to clusters. An initial cluster assignment may be performed by applying a hash function to one or more non-variable components of the message to generate an initial cluster identifier. Subsequently, clustering may be further refined (e.g., by determining whether to merge clusters based on similarity values). An interface can present a representative message of each cluster and indicate which portions of the message correspond to a variable component. Particular inputs detected at the input corresponding to one of these components can cause other values for the component to be presented. For a given cluster, timestamps of assigned messages can be used to generate a time series, which can facilitate grouping of clusters (with similar or complementary shapes) and/or triggering alerts (with a condition corresponding to a temporal aspect).


