Data Summarization Linking via Affinity Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current tools lack the ability to efficiently search and analyze large quantities of raw machine-generated data from diverse sources, leading to challenges in identifying data subsets of interest due to the complexity and volume of data generated in IT environments.
Innovation Solution
The implementation of an event-based data intake and query system, such as the SPLUNKĀ® ENTERPRISE system, which uses a late-binding schema to process and store data, allowing for flexible extraction of information at search time and enabling the correlation of data across disparate sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If raw machine data from diverse sources is stored and analyzed, then greater flexibility and insights are obtained, but the complexity and difficulty of searching and analyzing the data increases
Solution Approach 1:
The patent introduces summarizations as intermediary objects that represent large sets of raw data in a condensed form. These summarizations contain extracted data elements and metadata that capture the essence of the underlying raw data without requiring analysts to process the entire raw data set. This intermediary layer simplifies the analysis process while preserving the ability to access detailed information when needed.
Solution Approach 2:
The patent segments large data sets into manageable summarizations that can be independently analyzed and correlated. Each summarization represents a specific subset of raw data with particular characteristics or time ranges, allowing analysts to work with divided, organized units rather than overwhelming monolithic data sets. This segmentation enables more efficient searching and analysis.
2Productivity
If summarizations of data sets are created to simplify analysis, then search efficiency improves, but information loss may occur during summarization
Solution Approach 1:
The patent performs preliminary extraction of key data elements and creation of summarizations before the actual analysis process. This advance preparation condenses large data sets into structured summaries with preserved metadata, enabling efficient searching and correlation without losing access to the underlying detailed information. The summarizations are pre-processed to contain the most relevant information for subsequent analysis.
Solution Approach 2:
The patent creates summarizations as simplified copies or representations of the original raw data sets. These summarizations contain extracted data elements that replicate the essential characteristics and relationships of the source data without containing all the raw detail. This copying approach enables efficient analysis while preserving the ability to reference original data when needed.
3Reliability
If affinities between summarizations are calculated to correlate data, then data correlation capability improves, but computational resources and time are consumed
Solution Approach 1:
The patent applies affinity calculations selectively to specific data elements and metadata within summarizations rather than performing comprehensive comparisons of entire data sets. By focusing computational effort on key extracted elements and metadata fields that are most indicative of relationships, the system achieves effective data correlation with reduced computational overhead compared to analyzing all raw data.
Data Source
AI summary
Systems and methods are disclosed involving user interface (UI) search tools for visualizing or summarizing a data set. A number of summarizations may be created that summarizes the data set in different ways. The summarizations may be linked, such that selecting a data element of a first summarization causes display of a second summarization. To assist in linking of summarizations, a user interface is further provided to display suggested linkings between summarizations based on affinities of the two summarizations. Affinities can reflect similarities in the data content of the two summarizations, such as an output of a first summarization being a valid input to the second summarization.


