Automated Dataset Placement via Global Name Repository

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data management systems in device ecosystems face challenges in efficiently managing and processing vast amounts of data produced at the edge, particularly in distributed edge systems, due to complexity, limited visibility, and organizational fragmentation, leading to unmanageability and difficulties in understanding data intent and purpose.

Innovation Solution

A method for distributed data management involving monitor agents that classify data intent using heuristic rules and machine learning classifiers, generating global names and metadata, and updating a global name repository to facilitate data placement and services management across the ecosystem.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If data is stored locally by edge devices, then data availability and processing speed are improved, but data management complexity and organizational fragmentation increase

Engineering Contradiction:
Improvedata processing speedVSAvoiddata management complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent introduces a data manager component that acts as an intermediary between edge devices and the cloud. This data manager centrally coordinates data placement decisions, manages dataset policies, and handles data lifecycle operations across distributed edge devices, thereby reducing the operational complexity at each individual device while maintaining fast local processing capabilities

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The data manager provides universal data management services across multiple edge devices and datasets. It handles diverse data types, applies unified placement policies, and manages various data operations (creation, access, deletion) through a single centralized system, reducing the need for device-specific management logic

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If distributed data management is implemented, then scalability is improved, but visibility into data intent and purpose decreases

Engineering Contradiction:
Improvesystem scalabilityVSAvoiddata intent visibility
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent attaches metadata and intent information to datasets at the time of their creation or placement. This preliminary action ensures that data intent and purpose are captured upfront and maintained throughout the data lifecycle, enabling the system to scale distributedly while preserving complete visibility into data meaning and purpose

Inventive Principle:
Principle #10Preliminary action

3Productivity

If automated data placement is performed, then data management efficiency is improved, but policy analysis complexity increases

Engineering Contradiction:
Improvedata management efficiencyVSAvoidpolicy analysis complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent uses data descriptors as standardized templates or copies that define data placement policies. Instead of complex ad-hoc decision logic, the system uses pre-defined policy templates that can be automatically applied to different datasets and edge devices, simplifying the policy analysis process while maintaining high automation efficiency

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11874848B2Automated dataset placement for application execution
Publication Date: 2024.01.16 EMC IP HLDG CO LLC
  • US11874848B2 patent drawing
  • US11874848B2 patent drawing
  • US11874848B2 patent drawing

AI summary

Techniques described herein relate to a method for distributed data management. The method may include obtaining data descriptors for an application executing on a data host, performing a dataset policy analysis using the data descriptors to determine a data placement for a dataset associated with the application using a global name repository, performing, based on the data policy analysis, the data placement, and based on the data placement, updating the global name repository.