Log Sampling in Distributed Analytics Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing complexity and scale of distributed computing systems lead to significant inefficiencies and overheads in managing and administering log/event messages, resulting in high computational, networking, and data-storage burdens, which can cause information loss and impair system management capabilities.
Innovation Solution
Implementing a method to sample log/event messages rather than processing and storing every message, using techniques like clustering and feature vector transformation to determine sampling rates, thereby reducing data-storage and bandwidth requirements while maintaining relevant information for downstream processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If every log/event message is processed and stored, then complete information is retained, but data-storage capacity and computational bandwidth overheads increase significantly
Solution Approach 1:
The patent extracts and retains only the most relevant log messages based on importance criteria, discarding redundant information. This selective extraction maintains essential system information while significantly reducing storage requirements by removing unnecessary data points.
Solution Approach 2:
The patent applies different retention qualities to different log messages based on their importance. Critical messages are retained with full detail while less important messages are discarded or summarized, creating a non-uniform retention strategy that optimizes storage efficiency.
2Reliability
If every log/event message is processed and stored, then comprehensive system monitoring is achieved, but computational bandwidth and networking bandwidth overheads increase
Solution Approach 1:
The system extracts only the most significant log messages for processing and transmission, filtering out redundant information. This reduces computational bandwidth requirements by focusing processing resources on high-value data while maintaining reliable system monitoring through selective message retention.
Solution Approach 2:
The patent applies partial processing to the log stream, handling only the most critical messages in detail while using sampling or aggregation for less important messages. This partial action approach maintains monitoring reliability for critical events while reducing overall computational overhead.
3Quantity of substance
If sampling is applied to log/messages, then data-storage and bandwidth overheads decrease, but information loss may occur
Solution Approach 1:
The patent implements differential sampling rates based on message importance. Critical messages are retained at 100% sampling rate while less important messages use lower sampling rates, ensuring that useful information is preserved while achieving storage reduction through targeted sampling of non-critical data.
Solution Approach 2:
The system dynamically adjusts sampling parameters based on message characteristics and system conditions. By changing sampling rates, retention criteria, and aggregation levels according to message importance, the system optimizes the balance between storage efficiency and information retention.
4Quantity of substance
If log/event message volume increases, then system complexity increases, but management and administration become less efficient
Solution Approach 1:
The patent extracts and removes redundant log messages that do not contribute to system management effectiveness. By filtering out duplicate and low-value messages, the system reduces log volume while maintaining management efficiency through retention of actionable information.
Solution Approach 2:
The system applies selective retention to log messages, maintaining full detail for critical management-related messages while using aggregation or summarization for routine messages. This partial retention strategy reduces overall log volume while preserving management efficiency by keeping essential information readily available.
Data Source
AI summary
The current document is directed to methods and systems that sample log/event messages for downstream processing by log/event-message systems incorporated within distributed computer facilities. The data-collection, data-storage, and data-querying functionalities of log/event-message systems provide a basis for distributed log-analytics systems which, in turn, provide a basis for automated and semi-automated system-administration-and-management systems. By sampling log/event-messages, rather than processing and storing every log/event-message generated within a distributed computer system, a log/event-message system significantly decreases data-storage-capacity, computational-bandwidth, and networking-bandwidth overheads involved in processing and retaining large numbers of log/event messages that do not provide sufficient useful information to justify these costs. Increase in efficiencies of log/event-message systems obtained by sampling translate directly into increases in bandwidths of distributed computer systems, in general, and to increases in time periods during which useful log/event messages can be stored.


