Late-Binding Schema for Field Label-Value Pair Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern data centers face challenges in processing and analyzing large volumes of heterogeneous, unstructured machine-generated data due to difficulties in applying semantic meaning and indexing, leading to inefficient data retrieval and interpretation.

Innovation Solution

The implementation of an event-based system with a late-binding schema that allows for flexible extraction rules and field label-value pairs, enabling the distinction between different data extractions and attributes, and facilitating the storage and processing of minimally processed data for later analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If data is maintained in unstructured form to preserve more data for later use, then data retention is improved, but indexing and searching operations become difficult

Engineering Contradiction:
Improvedata retentionVSAvoidindexing and searching difficulty
Core Design Contradiction:
Loss of informationVSDifficulty of detecting and measuring

Solution Approach 1:

The patent segments unstructured data into structured field-label pairs during extraction, organizing data by semantic meaning while preserving the original unstructured content. This segmentation enables both data retention and efficient indexing by creating a structured representation that can be queried while maintaining access to the complete original data.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary extraction layer that sits between the unstructured data source and the storage system. This intermediary processes unstructured data into a semi-structured format with field labels, enabling efficient searching and indexing without losing the original unstructured data, thus acting as a mediator that resolves the contradiction between data retention and searchability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If data is pre-processed with extraction and storage of selected data, then storage space is saved in the short term, but data availability is reduced in the long term

Engineering Contradiction:
Improvestorage spaceVSAvoiddata availability
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent creates a copied structured representation of unstructured data through field-label pairs. Instead of storing only extracted data, the system extracts and structures selected data for efficient storage and retrieval while preserving the complete original unstructured data, effectively creating a copy that enables both space efficiency and data availability.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary extraction and structuring of data during the ingestion phase, organizing data into field-label pairs before storage. This preliminary action enables efficient future retrieval and processing without requiring re-processing of the original unstructured data, saving both storage space and ensuring long-term data availability.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If a large amount of information is returned from data processing, then completeness of information is improved, but user interpretability deteriorates

Engineering Contradiction:
Improveinformation completenessVSAvoiduser interpretability
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent extracts and highlights specific field-label pairs from the complete processed data, separating the most relevant structured information from the full data set. This extraction presents users with organized, interpretable field-value pairs while the complete information remains available in the underlying unstructured data, resolving the contradiction between completeness and interpretability.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11907271B2Distinguishing between fields in field value extraction
Publication Date: 2024.02.20 CISCO TECHNOLOGY INC
  • US11907271B2 patent drawing
  • US11907271B2 patent drawing
  • US11907271B2 patent drawing

AI summary

First one or more values are extracted from a plurality of events using a first extraction rule. The extracted first one or more values are assigned to a first field of the plurality of events as a first set of field-data item pairs and a field label is assigned to the first field. Second one or more values and a field label corresponding to the second one or more values are extracted from the plurality of the events using a second extraction rule, where the extracted field label corresponds to the assigned field label of the first field. The extracted second one or more values are assigned to a second field of the plurality of events as a second set of field-data item pairs, thereby distinguishing the extracted second one or more values from the extracted first one or more values.