Unified Data Lake for Multi-Source Query Coordination

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data intake and query systems face challenges in seamlessly searching and analyzing diverse data types from various data sources, including enterprise systems and open source technologies, due to limited search and analytics capabilities that are often isolated to internal data stores, and lack the ability to route data to different destinations.

Innovation Solution

A data intake and query system that employs a search process master and query coordinators combined with a scalable network of distributed nodes to collect and process data from diverse data systems, extending search and analytics capabilities to include external data sources, common storage, and ingested data buffers, enabling scalable analytics across multiple data sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If data systems store and pre-process only specified data items based on anticipated analysis needs, then retrieval efficiency is improved, but data flexibility and analytical scope are reduced

Engineering Contradiction:
Improvedata retrieval speedVSAvoiddata analysis flexibility
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary data collection and storage actions by capturing all generated data from diverse sources before specific analysis needs are known. Data is stored in a unified data lake format that preserves raw information while enabling efficient subsequent retrieval and multiple analytical pathways without requiring pre-defined processing routes.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically adapts its data processing and retrieval operations based on actual query requirements. The unified data lake allows the system to transform and analyze data in multiple ways depending on the specific analytical needs, providing flexible query capabilities that can adjust to different analysis scenarios without requiring predetermined data processing paths.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If tools search data systems separately and collect results over a network, then data source coverage is improved, but search efficiency and user experience are reduced

Engineering Contradiction:
Improvedata source coverageVSAvoidsearch efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system merges multiple data sources from diverse systems into a single unified data lake, consolidating data from enterprise systems, open source technologies, and external sources into one accessible repository. This eliminates the need for separate tool searches across multiple systems while maintaining comprehensive data source coverage, significantly improving search efficiency and user experience.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The unified data lake serves as a universal data repository that can handle multiple types of data sources and support various analytical operations through a single interface. The system provides multi-functional capabilities including data collection, storage, processing, and analysis through one unified platform, replacing the need for multiple specialized tools while maintaining broad data source compatibility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If the system extends search capabilities to multiple data sources, then data accessibility is improved, but system complexity increases

Engineering Contradiction:
Improvedata accessibilityVSAvoidsystem architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system introduces a unified data lake as an intermediary layer between diverse data sources and analytical tools. This mediator consolidates connections to multiple external systems and internal data stores, providing a single standardized interface for data access. The data lake absorbs the complexity of integrating numerous data sources while presenting a simplified access model to users and applications, thereby improving data accessibility without exposing system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11232100B2Resource allocation for multiple datasets
Publication Date: 2022.01.25 CISCO TECHNOLOGY INC
  • US11232100B2 patent drawing
  • US11232100B2 patent drawing
  • US11232100B2 patent drawing

AI summary

Systems and methods are disclosed for processing queries against multiple dataset sources. One dataset source can include indexers that index and store data. The system can receive a query that identifies a set of data to be processed and a manner of processing the set of data. The set of data can include a first dataset that is accessible by one or more indexers and a second dataset that is accessible by one or more other dataset sources. A query coordinator can define a query processing scheme for obtaining and processing the set of data that includes a dynamic allocation of multiple layers of partitions. The partitions can operate on multiple worker nodes. The query can then be executed based on the query processing scheme.