Distributed Data Processing Platform Orchestrating Local Analytics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current distributed data processing frameworks face challenges in efficiently processing data across multiple geographic locations due to the need for a shared distributed file system, which is difficult to configure and maintain, and raises privacy concerns when data is copied to a centralized site for analysis.

Innovation Solution

The implementation of a distributed data processing platform that orchestrates the execution of applications across multiple processing nodes, allowing computations to be performed locally within each node's data zone, eliminating the need for data movement and enhancing privacy and security.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If data is copied from local sites to a centralized site for analysis, then centralized data processing can be performed, but it raises privacy concerns and increases data transmission time

Engineering Contradiction:
Improvedata processing capabilityVSAvoiddata transmission time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

Instead of copying data to a centralized site for processing, the patent inverts the approach by bringing the processing capability to the distributed nodes where data resides. The system orchestrates execution of data-intensive computing tasks across multiple geographically dispersed processing nodes, allowing computations to be performed locally without moving the actual data.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent segments the centralized processing function into distributed processing nodes across multiple geographic locations. Each node can independently execute computing tasks on local data, and the orchestration system coordinates these distributed operations to achieve the overall processing goal without requiring data centralization.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If a shared distributed file system is deployed across multiple geographic locations, then data can be accessed by all processing nodes, but it becomes difficult to configure and maintain

Engineering Contradiction:
Improvedata accessibilityVSAvoidsystem configuration and maintenance complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent extracts the file system dependency from the distributed processing model. Instead of requiring a shared distributed file system, the system allows processing nodes to access data through local file systems or other data access mechanisms at each node, eliminating the need for complex distributed file system infrastructure while maintaining data accessibility.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The orchestration system acts as an intermediary that manages data access between processing nodes and local data sources. It handles the complexity of coordinating data access across distributed nodes without requiring a shared file system, simplifying the overall system architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If data is copied to a centralized site for analysis, then comprehensive analytics can be performed, but it increases bandwidth consumption and energy usage

Engineering Contradiction:
Improveanalytics capabilityVSAvoidbandwidth and energy consumption
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent inverts the traditional data movement approach by keeping data stationary at distributed nodes and moving only the processing logic and results. This eliminates the need to transmit large volumes of data across the network, dramatically reducing bandwidth consumption and associated energy usage while maintaining comprehensive analytics capability.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent segments the analytics processing across multiple distributed nodes, allowing each node to perform local analytics on its data. The orchestration system coordinates these segmented processing tasks to achieve comprehensive analytics without requiring data centralization, thereby minimizing energy consumption.

Inventive Principle:
Principle #1Segmentation

4Adaptability or versatility

If a shared distributed file system is used, then all processing nodes can access the same data, but it raises security and governance challenges

Engineering Contradiction:
Improvedata sharing capabilityVSAvoidsecurity and governance
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments data access permissions and security policies at each distributed node. Each node maintains its own security context and can share data with other nodes according to node-specific policies, eliminating the need for a centralized security model while maintaining data sharing capability. This distributed security approach improves governance and reduces security risks.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10541938B1Integration of distributed data processing platform with one or more distinct supporting platforms
Publication Date: 2020.01.21 EMC IP HLDG CO LLC
  • US10541938B1 patent drawing
  • US10541938B1 patent drawing
  • US10541938B1 patent drawing

AI summary

An apparatus in one embodiment comprises at least one processing device having a processor coupled to a memory. The one or more processing devices are operative to configure a plurality of distributed processing nodes to communicate over a network, to obtain metadata characterizing data locally accessible in respective data zones of respective ones of the distributed processing nodes, and to populate catalog instances of a distributed catalog service for respective ones of the data zones utilizing the obtained metadata. Distributed data analytics are performed in the distributed processing nodes utilizing the populated catalog instances of the distributed catalog service and the locally accessible data of the respective data zones. The metadata characterizing the locally accessible data is illustratively obtained in a metadata repository from at least one of a master data management platform and a governance, risk and compliance platform.