Unstructured Data Reporting via Field Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modern data centers face challenges in processing and interpreting large volumes of machine-generated data due to its unstructured nature, leading to difficulties in indexing, searching, and presenting relevant information to users.
Innovation Solution
A reporting application that enables users to generate reports from unstructured data by identifying fields, selecting relevant data subsets, and creating visualizations and aggregates without requiring expertise in search processing languages, using a drag-and-drop interface and interactive graphical user interfaces.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If data is maintained in unstructured form to preserve more data for later use, then data availability is improved, but indexing and searching operations become difficult
Solution Approach 1:
The patent introduces an intermediary processing layer that automatically extracts structured fields from unstructured data during ingestion. This intermediary process creates a hybrid representation where the original unstructured data is preserved alongside extracted structured metadata, enabling both data availability and efficient searching without requiring the data to be fully structured beforehand
Solution Approach 2:
The system performs preliminary field extraction and indexing operations at data ingestion time rather than at query time. By pre-processing the unstructured data to identify and extract relevant fields, the system prepares the data for future searching operations while maintaining the original unstructured format for complete data availability
2Loss of information
If minimal processing is applied to preserve more data, then data completeness is improved, but data interpretability deteriorates
Solution Approach 1:
The patent segments the data representation into multiple layers: the original unstructured data layer for completeness, and an extracted structured fields layer for interpretability. Each layer serves its specific purpose, allowing users to access complete raw data when needed while also providing pre-processed structured data for easier analysis and reporting
Solution Approach 2:
Different parts of the data system are given different qualities - the raw data storage maintains unstructured format for completeness, while the indexing layer maintains structured extracted fields for interpretability. This local differentiation allows each component to optimize for its specific function without compromising the other
3Quantity of substance
If large volumes of machine-generated data are processed, then data coverage is improved, but processing complexity increases
Solution Approach 1:
The patent applies partial processing to large volumes of data by extracting only the most relevant fields needed for common search and analysis operations. Rather than fully processing every aspect of every data point, the system performs selective field extraction on a subset of data characteristics, reducing processing complexity while maintaining adequate data coverage for most use cases
4Loss of information
If search results are provided in large sets, then search completeness is improved, but user interpretation difficulty increases
Solution Approach 1:
The patent extracts and highlights key structured fields from within large sets of search results, presenting them in a simplified format that shows users the most relevant information upfront. This extraction approach allows users to quickly understand search results without having to manually parse through large volumes of unstructured data, while the complete results remain available for detailed analysis
Data Source
AI summary
The disclosure relates to certain system and method embodiments for generating reports from unstructured data. In one embodiment, a method can include identifying events matching criteria of an initial search query (each of the events including a portion of raw machine data that is associated with a time), identifying a set of fields, each field defined for one or more of the identified events, causing display of an interactive graphical user interface (GUI) that includes one or more interactive elements enabling a user to define a report for providing information relating to the matching events (each interactive element enabling processing or presentation of information in the matching events using one or more fields in the identified set of fields), receiving, via the GUI, a report definition indicating how to report information relating to the matching events, and generating, based on the report definition, a report including information relating to the matching events.


