Query Execution at Remote Heterogeneous Data Stores

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data intake and query systems face challenges in seamlessly searching and analyzing large sets of diverse data from various sources, including structured, semi-structured, and unstructured data, due to limited scope and unidirectional processing flows that restrict the ability to route data to different destinations.

Innovation Solution

A data intake and query system that extends search and analytics capabilities by employing a search process master and query coordinators combined with a scalable network of distributed nodes, enabling the collection and processing of data from diverse data systems and providing search results across internal and external data sources, including MySQL, PostgreSQL, NoSQL data stores, cloud storage, and HDFS.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If data is stored in diverse external data systems with different data models and protocols, then data availability and search scope are improved, but system complexity and difficulty of integration increase

Engineering Contradiction:
Improvesearch scopeVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces external result providers (ERPs) as intermediary components that mediate between the search head and external data systems. These ERPs translate search requests into external system-specific queries and transform results back into a unified format, thereby extending search scope to diverse data sources while abstracting away the underlying system complexity and heterogeneity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The search head is designed with universal capabilities to coordinate searches across multiple external data systems with different data models and protocols. It implements a unified search interface that can adapt to various external systems through configurable ERPs, allowing a single system to handle diverse data sources without requiring separate specialized systems for each data type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If unidirectional processing flows are used in data intake systems, then system simplicity is maintained, but flexibility to route data to different destinations is reduced

Engineering Contradiction:
Improverouting flexibilityVSAvoidprocessing flow complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent implements dynamic, configurable data processing flows that can adapt to different routing requirements. The system allows configuration of multiple data intake sources and destinations with flexible routing rules, enabling the processing flow to dynamically direct data to appropriate destinations based on data type, source, and predefined conditions, thereby achieving routing flexibility without fixed unidirectional constraints.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If limited scope data processing is implemented, then system complexity is reduced, but ability to analyze diverse data types is constrained

Engineering Contradiction:
Improvedata type coverageVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the data processing architecture into distinct functional components: data intake sources, external result providers, search head, and data destinations. Each component handles specific data types and processing tasks independently, allowing the system to support diverse data types (structured, semi-structured, unstructured) through modular processing while managing complexity through clear separation of concerns.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11921672B2Query execution at a remote heterogeneous data store of a data fabric service
Publication Date: 2024.03.05 CISCO TECHNOLOGY INC
  • US11921672B2 patent drawing
  • US11921672B2 patent drawing
  • US11921672B2 patent drawing

AI summary

Systems and methods are described for executing a query of raw machine data that is stored at a remote data store that may store heterogeneous data. The system can determine the directories or file types that may store event data and may instruct one or more worker nodes to access files that may store events based on the determined directories of file types. Further, the system may exclude files at the remote data store that may not be identified as potentially storing events enabling a query that implicates a heterogeneous data store to be efficiently executed.