Data Intake Query System Late-Binding Schema

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern data centers face challenges in analyzing and searching massive volumes of machine-generated data due to its unstructured nature and diverse formats, making it difficult to apply semantic meaning and efficiently process and index this data for effective retrieval and analysis.

Innovation Solution

A data intake and query system utilizing a flexible, late-binding schema that processes and stores machine data as events with timestamps, allowing for field-searchability and enabling users to define extraction rules at search time, thereby facilitating the extraction of semantically-related values from disparate data sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional indexing methods are used to process machine data, then data can be stored and retrieved, but the unstructured nature and diverse formats make it difficult to apply semantic meaning and efficiently process the data

Engineering Contradiction:
Improvedata processing efficiencyVSAvoiddata structure complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent transforms unstructured machine data into structured events by changing the data format parameters. Each event contains standardized fields (timestamp, host, source, type, token) that convert diverse unstructured data into a uniform structured format, enabling efficient processing while maintaining the ability to represent various data types through flexible field definitions

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary event structure that mediates between raw unstructured machine data and the indexing system. Events serve as intermediate representations that bridge the gap between diverse data sources and the search infrastructure, allowing semantic meaning to be applied during the intermediate processing stage rather than at the raw data or final storage stage

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If all machine data is collected and stored for analysis, then comprehensive operational intelligence can be obtained, but the volume of data makes indexing and searching operations challenging

Engineering Contradiction:
Improveinformation completenessVSAvoidindexing and searching efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent extracts only the essential semantic components from machine data during the event creation process. By identifying and extracting key fields (timestamp, host, source, type, token) from unstructured data, the system maintains information completeness for analysis while reducing the data volume that needs to be indexed and searched, improving processing efficiency

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments machine data into discrete events with standardized fields. This segmentation breaks down large volumes of unstructured data into manageable, uniformly structured units that can be efficiently indexed and searched, while the flexible field definitions ensure that semantically-related values from disparate sources are properly categorized and retrieved

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12039310B1Information technology networked entity monitoring with metric selection
Publication Date: 2024.07.16 CISCO TECHNOLOGY INC
  • US12039310B1 patent drawing
  • US12039310B1 patent drawing
  • US12039310B1 patent drawing

AI summary

Data intake and query system (DIQS) instances supporting applications including lower-tier, focused, work group oriented applications may be tailored to meet the specific needs of the users. Rather than offer pre-configured options, the DIQS-based application offers the user the ability to customize data collection before deploying the collectors for specified host entities within an IT environment. Once the user selects the metrics and/or log sources for data collection at a custom interface, the lower-tier DIQS generates custom script operable to establish collection of the source data having the selected metrics and events associated with selected log sources from the specified host entities. The user can display and analyze the collected data.