Event-Based Data Intake System With Flexible Schema

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern data centers face challenges in analyzing and searching massive quantities of machine-generated data due to its unstructured nature and diverse formats, which complicates the application of semantic meaning and efficient processing.

Innovation Solution

An event-based data intake and query system with a flexible schema, allowing for late-binding schema application during search time, enables the collection, indexing, and retrieval of machine data from various sources, using extraction rules and configuration files to identify and extract relevant information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a rigid pre-defined schema is used for data collection and indexing, then data processing efficiency is improved, but adaptability to diverse data formats from multiple sources deteriorates

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidadaptability to diverse data formats
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent implements a dynamic schema system that evolves from rigid pre-defined schemas to flexible late-binding schemas. The schema is applied dynamically during search time rather than being fixed during data collection, allowing the system to adapt to diverse data formats while maintaining processing efficiency through indexed fields and extracted information.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If unstructured machine-generated data is collected without pre-defined formats, then adaptability to diverse sources is improved, but difficulty in applying semantic meaning and performing indexing operations increases

Engineering Contradiction:
Improveadaptability to diverse data sourcesVSAvoiddifficulty in applying semantic meaning
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent performs preliminary extraction of structured fields from unstructured data during the indexing phase. By extracting and organizing data into standardized fields (host, source, event type, etc.) before search operations, the system prepares semantic meaning in advance, making subsequent searching and analysis operations efficient despite the diverse origins of the data.

Inventive Principle:
Principle #10Preliminary action

3Speed

If data is indexed and searched using traditional methods, then processing speed is maintained, but ability to perform intelligent analysis and derive insights from large volumes of data deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidintelligent analysis capability
Core Design Contradiction:
SpeedVSProductivity

Solution Approach 1:

The patent introduces an intermediary search head component that coordinates between data sources, indexers, and users. The search head receives queries, distributes them to appropriate indexers, collects results, and presents them to users. This intermediary layer enables intelligent analysis by orchestrating complex search operations across multiple data sources while maintaining processing speed through distributed architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11693895B1Graphical user interface with chart for event inference into tasks
Publication Date: 2023.07.04 CISCO TECHNOLOGY INC
  • US11693895B1 patent drawing
  • US11693895B1 patent drawing
  • US11693895B1 patent drawing

AI summary

Machine data reflecting operation of a monitored system is ingested and made available for search by a data intake and query system (DIQS). Monitoring includes obtaining a subset of ordered events that are assigned to a task. In a graphical user interface on a display, a chart for the task is displayed. The chart includes an event identifier for each event of the subset of the ordered events, a confidence level value related to each event identifier of each event of the subset of ordered events, the confidence level value indicating the confidence level that the event is in the task. The chart further includes a time reference value identifying a time of each event.