Automatic Visualization Parameter Selection for IT Machine Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Analyzing and searching massive quantities of machine-generated data poses challenges due to the vast variety and format diversity, with existing tools often discarding non-preprocessed data and requiring pre-existing insights for visualization, limiting exploratory analysis and flexibility.
Innovation Solution
An event-based data intake and query system, such as the SPLUNKĀ® ENTERPRISE system, uses a late-binding schema to process and visualize data at search time, enabling flexible data analysis and visualization through a data analysis tool that automatically generates manipulable visualizations based on user selections and employs acceleration techniques like parallel processing and keyword indexing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If pre-specified data items are extracted and stored in a database during pre-processing, then data retrieval and analysis efficiency is improved, but data flexibility and exploratory analysis capability deteriorate because non-preprocessed data is discarded
Solution Approach 1:
The system performs preliminary indexing of all machine data at data intake time, creating a searchable structure without discarding any data. This preliminary action enables both efficient retrieval (by having data indexed) and flexible analysis (by retaining all original data), resolving the contradiction between productivity and adaptability
Solution Approach 2:
The system changes the parameter of data storage from selective extraction of pre-specified items to comprehensive retention of all machine data with automated field identification. This parameter change allows the system to maintain data flexibility while improving retrieval efficiency through intelligent field detection and indexing
2Adaptability or versatility
If all machine data is stored for later retrieval and analysis, then data flexibility and exploratory analysis capability are improved, but query processing time and system complexity worsen
Solution Approach 1:
The system performs preliminary field identification and data classification during data intake, automatically analyzing the structure and content of machine data before it needs to be queried. This preliminary action reduces query processing time by having fields pre-identified and indexed, while still maintaining flexibility to analyze all stored data
Solution Approach 2:
The system employs automated field identification that self-analyzes incoming machine data to determine field structures, data types, and relationships without requiring manual configuration. This self-service approach handles the complexity of analyzing all machine data automatically, reducing the burden on users and enabling flexible analysis of comprehensive data sets
3Ease of operation
If automated field identification and visualization generation are implemented, then ease of operation is improved, but device complexity increases
Solution Approach 1:
The system implements automated field identification that self-analyzes machine data structures and automatically generates appropriate visualizations without user intervention. This self-service capability simplifies operation for users while the system internally handles the complexity of data analysis and visualization selection
Solution Approach 2:
The system changes from manual visualization configuration to automated generation by detecting data parameters and automatically selecting appropriate visualization types. This parameter change improves ease of operation while the system manages the underlying complexity through intelligent algorithms
Data Source
AI summary
Embodiments are disclosed for a data analysis tool for facilitating iterative and exploratory analysis of large sets of data. In some embodiments a data analysis tool includes a graphical user interface through which an interactive set of field identifiers is displayed. Each of the listed field identifiers may reference fields associated with a set of events returned in response to a search query, the set of events including machine data produced by components within an information technology (IT) environment that reflects activity in the IT environment. In response to user selections of field identifiers included in the displayed set, a data analysis tool may cause display of manipulable visualizations based on values included in fields referenced by the selected field identifiers.


