Dynamic Schema Generation for Cross-Source Data Field Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Analyzing and searching massive quantities of machine-generated data from diverse sources in data centers is challenging due to the vast types and formats of data, requiring efficient methods to extract and process data without pre-defined schemas, while conventional systems often discard minimally processed data, limiting flexibility and insights.
Innovation Solution
The implementation of an event-based data intake and query system, such as the SPLUNKĀ® ENTERPRISE system, which uses a late-binding schema to extract values from raw data at search time, allowing flexible data modeling and analysis across disparate data sources, enabling the identification of related data fields and patterns.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional systems use pre-defined schemas to process data, then data processing efficiency is improved, but flexibility and adaptability to diverse data formats deteriorate
Solution Approach 1:
The patent applies dynamics by transitioning from static pre-defined schemas to dynamic schema generation. The system automatically generates schemas based on the actual data being processed, allowing the schema structure to adapt dynamically to different data formats and types. This enables the system to maintain high processing efficiency while simultaneously handling diverse and evolving data formats without requiring manual schema reconfiguration.
Solution Approach 2:
The patent utilizes parameter changes by transforming the schema from a fixed structural parameter to a variable parameter that can be automatically generated and modified. The system changes the schema parameters dynamically based on data characteristics, allowing the same processing system to efficiently handle multiple data formats by adjusting the schema parameters rather than requiring separate processing paths for each format.
2Loss of information
If minimal processing is applied to raw data, then data integrity and completeness are improved, but analysis capability and insight extraction deteriorate
Solution Approach 1:
The patent applies preliminary action by performing automatic schema generation and data type inference before the actual analysis process. The system prepares the data structure in advance by automatically understanding and categorizing the raw data, which enables subsequent efficient analysis without requiring extensive manual preprocessing. This preliminary structuring maintains data completeness while enabling rapid analysis capability.
Solution Approach 2:
The system applies self-service by automatically generating schemas and inferring data types without human intervention. The data itself provides the information needed to create its own structure through automatic type inference mechanisms, allowing the system to maintain complete raw data while automatically extracting analytical value through self-organizing data structures.
3Speed
If pre-defined schemas are used for data extraction, then data extraction speed is improved, but ability to identify related data across disparate sources deteriorates
Solution Approach 1:
The patent applies universality by creating a multi-functional schema generation system that can automatically adapt to various data sources and formats. The automatic schema generation mechanism serves multiple functions: it structures data for efficient extraction, identifies relationships across different source types, and maintains compatibility with diverse data formats. This universal approach enables fast extraction speed while simultaneously identifying related data across disparate sources through unified schema generation.
Data Source
AI summary
Embodiments of the present invention are directed to identifying and providing related data field sets. In one embodiment, a first portion of a graphical user interface (GUI) configured to receive a search query is displayed. The GUI enables user interaction to specify a source type in association with the search query. In accordance with a first source type specified in the search query, a first field set associated with the first source type is identified as related to a second field set associated with a second source type. A second portion of the GUI is displayed that includes a relationship indication that indicates the first field set associated with the first source type is related to the second field set associated with a second source type. Further, a third portion of the GUI is displayed that includes an explanation or recommendation associated with the relationship indication.


