Distributed Data Fabric for Scalable Multi-Source Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data intake and query systems face challenges in seamlessly searching and analyzing large sets of diverse data from various data sources, including structured, semi-structured, and unstructured data, due to limited scope and unidirectional processing flows that prevent comprehensive insights from being derived.

Innovation Solution

A data intake and query system with a network of distributed nodes and a search process master that extends search and analytics capabilities across diverse data systems, enabling scalable processing and visualization of data from internal and external sources, including MySQL, PostgreSQL, NoSQL databases, and cloud storage, through a data fabric platform.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If data is stored in a centralized data system for later retrieval and analysis, then data availability and flexibility are improved, but search and analysis efficiency deteriorate due to the vast amount of data

Engineering Contradiction:
Improvedata availabilityVSAvoidsearch and analysis efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent segments the centralized data system into multiple distributed worker nodes that independently store and process data subsets. Each worker node maintains local data storage and processing capabilities, allowing the system to handle vast amounts of data through parallel operations rather than centralized sequential processing, thus maintaining data availability while improving search and analysis efficiency

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimensional approach by distributing data across multiple spatial nodes (worker nodes) in the system architecture. This spatial distribution transforms the single-point centralized access model into a multi-point distributed access model, enabling simultaneous data retrieval and analysis operations across different nodes, thereby improving efficiency without sacrificing data availability

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of operation

If tools allow analysts to search data systems separately and collect results over a network, then data access flexibility is improved, but analysis comprehensiveness deteriorates due to piecemeal results

Engineering Contradiction:
Improvedata access flexibilityVSAvoidanalysis comprehensiveness
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent merges the separate search operations performed by analysts into a unified distributed query execution framework. The query coordinator consolidates multiple search requests and distributes them across worker nodes, then aggregates the results into a comprehensive unified result set, maintaining access flexibility while preventing information loss through systematic result integration

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces a query coordinator as an intermediary component that mediates between analysts and the distributed data storage system. This intermediary receives search requests, coordinates their execution across multiple worker nodes, and aggregates results, thereby maintaining ease of operation for analysts while ensuring comprehensive analysis through centralized result synthesis

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If a data fabric platform extends search capabilities across diverse data systems, then analysis versatility is improved, but system complexity increases

Engineering Contradiction:
Improveanalysis versatilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal data fabric platform that provides multi-functional capabilities across diverse data systems. The worker nodes are designed with universal interfaces and protocols that enable them to handle various data types and formats (structured, semi-structured, unstructured) from multiple sources, allowing the system to extend search capabilities across heterogeneous systems without proportionally increasing complexity through standardized multi-purpose components

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11580107B2Bucket data distribution for exporting data to worker nodes
Publication Date: 2023.02.14 CISCO TECHNOLOGY INC
  • US11580107B2 patent drawing
  • US11580107B2 patent drawing
  • US11580107B2 patent drawing

AI summary

Systems and methods are described for exporting bucket data from one or more buckets to one or more worker nodes. The system can identify data from different bucket data from buckets stored in a data intake and query system that is to be processed by one or more worker nodes. The system can allocate one or more execution resources, such as a processing pipeline, to process and export the bucket data from the buckets. The system can assign bucket data corresponding to individual buckets to the execution resource based on a bucket distribution policy. The indexer can export the bucket data to the worker nodes for further processing based on the bucket data-execution resource assignment.