Token Permutation Clustering for IT Log Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The manual analysis of computer-generated data for troubleshooting IT issues is inefficient, as it involves scanning and repeatedly trying different drill-downs to identify the root cause, which is time-consuming and lacks automation.
Innovation Solution
A process for clustering computer-generated data entries based on token permutation groupings, where data entries are segmented into tokens, and unique identifiers are determined for each permutation grouping, allowing for efficient grouping and identification of clusters, thereby simplifying analysis and filtration.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual analysis of computer-generated data is performed, then expertise and detailed scanning can identify root causes, but the process is time-consuming and inefficient
Solution Approach 1:
The patent segments computer-generated data entries into tokens and creates permutation groupings of these tokens. By dividing the data into smaller manageable units (tokens) and organizing them into structured groups, the system enables efficient automated processing while maintaining the ability to accurately identify patterns and root causes, thus resolving the contradiction between analysis accuracy and time consumption
Solution Approach 2:
The patent creates permutation-based representations (copies) of the original data entries. Instead of manually scanning the entire original data, the system works with generated permutation groupings that capture the essential characteristics of the data, enabling rapid automated analysis that preserves the accuracy needed for root cause identification while dramatically reducing analysis time
2Productivity
If automated clustering is implemented, then analysis efficiency and speed improve, but data structure complexity increases
Solution Approach 1:
The patent applies segmentation by breaking down data entries into tokens and organizing them into permutation groupings. This structured segmentation creates a systematic framework that enables automated clustering algorithms to process data efficiently while maintaining organizational simplicity, thus improving productivity without excessively increasing structural complexity
Solution Approach 2:
The patent transforms the original data structure into a new representation based on token permutations and grouping identifiers. By changing the parameters of data organization from raw text to structured token groupings with unique identifiers, the system enables automated processing and clustering while maintaining a manageable data structure that balances complexity and efficiency
Data Source
AI summary
A computer-generated data entry is received. The computer-generated data entry is segmented into a set of tokens. A plurality of different token permutation groupings are determined. Each of the different token permutation groupings includes a different subset of tokens from the set of tokens of the computer-generated data entry. For the computer-generated data entry, a plurality of token permutation grouping identifiers associated with at least a portion of the plurality of different token permutation groupings is obtained. It is determined whether the computer-generated data entry belongs to any data entry cluster among a plurality of previously identified data entry clusters based on a search performed using the token permutation grouping identifiers of the computer-generated data entry.


