Dynamic Schema Generation for Cross-Source Data Field Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Analyzing and searching massive quantities of machine-generated data from diverse sources in data centers is challenging due to the vast types and formats of data, requiring efficient methods to extract and process data without pre-defined schemas, while conventional systems often discard minimally processed data, limiting flexibility and insights.

Innovation Solution

The implementation of an event-based data intake and query system, such as the SPLUNKĀ® ENTERPRISE system, which uses a late-binding schema to extract values from raw data at search time, allowing flexible data modeling and analysis across disparate data sources, enabling the identification of related data fields and patterns.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional systems use pre-defined schemas to process data, then data processing efficiency is improved, but flexibility and adaptability to diverse data formats deteriorate

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidflexibility to diverse data formats
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent applies dynamics by transitioning from static pre-defined schemas to dynamic schema generation. The system automatically generates schemas based on the actual data being processed, allowing the schema structure to adapt dynamically to different data formats and types. This enables the system to maintain high processing efficiency while simultaneously handling diverse and evolving data formats without requiring manual schema reconfiguration.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent utilizes parameter changes by transforming the schema from a fixed structural parameter to a variable parameter that can be automatically generated and modified. The system changes the schema parameters dynamically based on data characteristics, allowing the same processing system to efficiently handle multiple data formats by adjusting the schema parameters rather than requiring separate processing paths for each format.

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If minimal processing is applied to raw data, then data integrity and completeness are improved, but analysis capability and insight extraction deteriorate

Engineering Contradiction:
Improvedata completenessVSAvoidanalysis capability
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent applies preliminary action by performing automatic schema generation and data type inference before the actual analysis process. The system prepares the data structure in advance by automatically understanding and categorizing the raw data, which enables subsequent efficient analysis without requiring extensive manual preprocessing. This preliminary structuring maintains data completeness while enabling rapid analysis capability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies self-service by automatically generating schemas and inferring data types without human intervention. The data itself provides the information needed to create its own structure through automatic type inference mechanisms, allowing the system to maintain complete raw data while automatically extracting analytical value through self-organizing data structures.

Inventive Principle:
Principle #25Self-service

3Speed

If pre-defined schemas are used for data extraction, then data extraction speed is improved, but ability to identify related data across disparate sources deteriorates

Engineering Contradiction:
Improvedata extraction speedVSAvoidability to identify related data across sources
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The patent applies universality by creating a multi-functional schema generation system that can automatically adapt to various data sources and formats. The automatic schema generation mechanism serves multiple functions: it structures data for efficient extraction, identifies relationships across different source types, and maintains compatibility with diverse data formats. This universal approach enables fast extraction speed while simultaneously identifying related data across disparate sources through unified schema generation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11100172B2Providing similar field sets based on related source types
Publication Date: 2021.08.24 CISCO TECHNOLOGY INC
  • US11100172B2 patent drawing
  • US11100172B2 patent drawing
  • US11100172B2 patent drawing

AI summary

Embodiments of the present invention are directed to identifying and providing related data field sets. In one embodiment, a first portion of a graphical user interface (GUI) configured to receive a search query is displayed. The GUI enables user interaction to specify a source type in association with the search query. In accordance with a first source type specified in the search query, a first field set associated with the first source type is identified as related to a second field set associated with a second source type. A second portion of the GUI is displayed that includes a relationship indication that indicates the first field set associated with the first source type is related to the second field set associated with a second source type. Further, a third portion of the GUI is displayed that includes an explanation or recommendation associated with the relationship indication.