Federated Data Lake Query Stitching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Complex network architectures, such as massively scalable data centers and hybrid cloud environments, pose challenges in guaranteeing end-to-end service legal agreements and troubleshooting due to increased latency and resource costs associated with centralized data lakes.

Innovation Solution

Implementing data stitching across federated data lakes, where individual data lakes are treated as a group, with queries broken down into operators for local data collection, and results stitched at a selected destination to minimize resource and computational costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If a centralized data lake architecture is used to analyze network traffic across multiple data centers, then comprehensive visibility into network performance is achieved, but resource costs and computational overhead increase significantly

Engineering Contradiction:
Improvenetwork performance visibilityVSAvoidresource cost
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The patent segments the centralized data lake architecture into multiple federated data lakes distributed across different data centers. Each data lake processes and stores data locally, eliminating the need to centralize all data in one location. This segmentation maintains comprehensive visibility through federated querying while reducing the resource costs associated with data movement and centralized processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a federated architecture dimension that allows querying across distributed data lakes without physically centralizing the data. This dimensional change enables comprehensive network performance visibility while maintaining data locality, thus reducing the resource costs associated with traditional centralized architectures.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If data is centralized in a single data lake for cross-data center analysis, then end-to-end service monitoring is improved, but latency increases due to data movement

Engineering Contradiction:
Improveend-to-end service monitoringVSAvoidlatency
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent segments the centralized data collection approach into distributed federated data lakes that remain at their source locations. This eliminates data movement across data centers while enabling end-to-end service monitoring through federated queries that virtualize access to distributed data, thereby reducing latency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a federated query processing layer as an intermediary that enables end-to-end service monitoring without physically moving data. This mediator layer coordinates queries across distributed data lakes, allowing comprehensive monitoring while keeping data local and minimizing latency.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Use of energy by moving object

If a federated data lake architecture is implemented with distributed data lakes, then resource costs are reduced, but system complexity increases

Engineering Contradiction:
Improveresource costVSAvoidsystem complexity
Core Design Contradiction:
Use of energy by moving objectVSDevice complexity

Solution Approach 1:

The patent implements a universal federated query processing layer that handles multiple data lakes with a single interface. This multi-functional layer abstracts the complexity of distributed data access, allowing resource-efficient federated queries while masking the underlying system complexity from users and applications.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The federated query processing layer acts as an intermediary that manages the complexity of distributed data lakes. It provides a unified interface for querying across multiple data centers, reducing the perceived system complexity while enabling resource-efficient distributed architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Use of energy by moving object

If data is distributed across multiple data lakes, then resource efficiency improves, but troubleshooting and analysis across domains become more difficult

Engineering Contradiction:
Improveresource efficiencyVSAvoidtroubleshooting ease
Core Design Contradiction:
Use of energy by moving objectVSEase of operation

Solution Approach 1:

The patent implements a universal federated query interface that simplifies troubleshooting across distributed data lakes. This multi-functional interface allows users to perform complex cross-domain analysis with simple queries, maintaining resource efficiency while improving ease of operation for troubleshooting and analysis.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11960508B2Data stitching across federated data lakes
Publication Date: 2024.04.16 CISCO TECHNOLOGY INC
  • US11960508B2 patent drawing
  • US11960508B2 patent drawing
  • US11960508B2 patent drawing

AI summary

In one embodiment, a device, in communication with a plurality of data lake sites, receives a federated data lake query. The device determines a plurality of data lake operator sets that each correspond to one of the plurality of data lake sites, wherein each of the plurality of data lake operator sets is used to establish a respective data pipeline for the federated data lake query. The device selects a particular data lake site of the plurality of data lake sites as a destination for data pipelines that are established for the federated data lake query. The device sends the plurality of data lake operator sets that each correspond to one of the plurality of data lake sites to cause the plurality of data lake sites to send query results to the particular data lake site using the data pipelines, wherein the particular data lake site stitches the query results.