User-Defined Extraction Rules for Unstructured Data Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern data centers face challenges in processing and analyzing large volumes of heterogeneous, unstructured performance data due to difficulties in applying semantic meaning and indexing, leading to inefficiencies in data retrieval and analysis.

Innovation Solution

The implementation of an event-based system like SPLUNK ENTERPRISE, which uses a late-binding schema to extract field label-value pairs from events through user-defined extraction rules, allowing for flexible data processing and indexing at search time, enabling efficient retrieval and analysis of minimally processed data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If data is maintained in unstructured form to preserve more data, then data completeness is improved, but indexing and searching operations become difficult

Engineering Contradiction:
Improvedata completenessVSAvoidindexing and searching difficulty
Core Design Contradiction:
Loss of informationVSDifficulty of detecting and measuring

Solution Approach 1:

The patent applies dynamics by implementing a late-binding schema that allows the data structure to adapt dynamically at search time rather than being fixed during data collection. The system dynamically determines field extractions based on the specific search query, enabling the same data to be flexibly structured for different analytical needs without losing information.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies preliminary action by pre-processing data with minimal structuring during collection, while preparing extraction rules that will be applied later during search operations. This allows the system to preserve raw data integrity initially while having the capability to extract meaningful fields when needed.

Inventive Principle:
Principle #10Preliminary action

2Quantity of substance

If data is pre-processed with extraction and storage of selected fields, then storage space is reduced, but data availability for future use is compromised

Engineering Contradiction:
Improvestorage spaceVSAvoiddata availability
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent applies parameter changes by transforming the same raw data into different structured formats depending on the search parameters. The late-binding schema allows extraction rules to be applied dynamically with different parameters for different queries, enabling efficient storage while maintaining data availability for various future uses.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system dynamically adjusts the level of data structuring based on the specific search needs rather than applying a fixed pre-processing scheme. This allows the system to maintain compact storage while ensuring that required data fields are extracted and made available when specific analytical needs arise.

Inventive Principle:
Principle #15Dynamics

3Loss of information

If large volumes of data are processed to return comprehensive search results, then information completeness is improved, but user interpretation becomes difficult

Engineering Contradiction:
Improveinformation completenessVSAvoiduser interpretation ease
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent applies the extraction principle by selectively extracting only the relevant field label-value pairs that match the search criteria and user needs. Rather than returning all processed data, the system extracts and presents only the pertinent information, making comprehensive results more manageable and easier for users to interpret.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies partial action by performing field extractions only for the specific fields needed to answer the user's query rather than processing and presenting all possible data. This selective approach maintains information completeness for relevant aspects while reducing the overall volume to improve user interpretation ease.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11841908B1Extraction rule determination based on user-selected text
Publication Date: 2023.12.12 CISCO TECHNOLOGY INC
  • US11841908B1 patent drawing
  • US11841908B1 patent drawing
  • US11841908B1 patent drawing

AI summary

Based on a selection by a user of first one or more values of one or more events displayed in a graphical interface, an extraction rule is automatically determined that is capable of extracting a field label-value pair at least partially within at least the selected one or more values. An option is displayed that correspond to the determined extraction rule in the graphical interface. Based on the user selecting the option in the graphical interface, display is caused of second one or more values of one or more field label-value pairs extracted from the one or more events using the extraction rule. The one or more events may be displayed in a table format, and the first one or more value may be selected by the user selecting one or more cells, columns, or text portions in the table format.