Late-Binding Schema for Field Label-Value Pair Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modern data centers face challenges in processing and analyzing large volumes of heterogeneous, unstructured machine-generated data due to difficulties in applying semantic meaning and indexing, leading to inefficient data retrieval and interpretation.
Innovation Solution
The implementation of an event-based system with a late-binding schema that allows for flexible extraction rules and field label-value pairs, enabling the distinction between different data extractions and attributes, and facilitating the storage and processing of minimally processed data for later analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If data is maintained in unstructured form to preserve more data for later use, then data retention is improved, but indexing and searching operations become difficult
Solution Approach 1:
The patent segments unstructured data into structured field-label pairs during extraction, organizing data by semantic meaning while preserving the original unstructured content. This segmentation enables both data retention and efficient indexing by creating a structured representation that can be queried while maintaining access to the complete original data.
Solution Approach 2:
The patent introduces an intermediary extraction layer that sits between the unstructured data source and the storage system. This intermediary processes unstructured data into a semi-structured format with field labels, enabling efficient searching and indexing without losing the original unstructured data, thus acting as a mediator that resolves the contradiction between data retention and searchability.
2Quantity of substance
If data is pre-processed with extraction and storage of selected data, then storage space is saved in the short term, but data availability is reduced in the long term
Solution Approach 1:
The patent creates a copied structured representation of unstructured data through field-label pairs. Instead of storing only extracted data, the system extracts and structures selected data for efficient storage and retrieval while preserving the complete original unstructured data, effectively creating a copy that enables both space efficiency and data availability.
Solution Approach 2:
The patent performs preliminary extraction and structuring of data during the ingestion phase, organizing data into field-label pairs before storage. This preliminary action enables efficient future retrieval and processing without requiring re-processing of the original unstructured data, saving both storage space and ensuring long-term data availability.
3Loss of information
If a large amount of information is returned from data processing, then completeness of information is improved, but user interpretability deteriorates
Solution Approach 1:
The patent extracts and highlights specific field-label pairs from the complete processed data, separating the most relevant structured information from the full data set. This extraction presents users with organized, interpretable field-value pairs while the complete information remains available in the underlying unstructured data, resolving the contradiction between completeness and interpretability.
Data Source
AI summary
First one or more values are extracted from a plurality of events using a first extraction rule. The extracted first one or more values are assigned to a first field of the plurality of events as a first set of field-data item pairs and a field label is assigned to the first field. Second one or more values and a field label corresponding to the second one or more values are extracted from the plurality of the events using a second extraction rule, where the extracted field label corresponds to the assigned field label of the first field. The extracted second one or more values are assigned to a second field of the plurality of events as a second set of field-data item pairs, thereby distinguishing the extracted second one or more values from the extracted first one or more values.


