Automated Dataset Placement via Global Name Repository
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data management systems in device ecosystems face challenges in efficiently managing and processing vast amounts of data produced at the edge, particularly in distributed edge systems, due to complexity, limited visibility, and organizational fragmentation, leading to unmanageability and difficulties in understanding data intent and purpose.
Innovation Solution
A method for distributed data management involving monitor agents that classify data intent using heuristic rules and machine learning classifiers, generating global names and metadata, and updating a global name repository to facilitate data placement and services management across the ecosystem.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If data is stored locally by edge devices, then data availability and processing speed are improved, but data management complexity and organizational fragmentation increase
Solution Approach 1:
The patent introduces a data manager component that acts as an intermediary between edge devices and the cloud. This data manager centrally coordinates data placement decisions, manages dataset policies, and handles data lifecycle operations across distributed edge devices, thereby reducing the operational complexity at each individual device while maintaining fast local processing capabilities
Solution Approach 2:
The data manager provides universal data management services across multiple edge devices and datasets. It handles diverse data types, applies unified placement policies, and manages various data operations (creation, access, deletion) through a single centralized system, reducing the need for device-specific management logic
2Adaptability or versatility
If distributed data management is implemented, then scalability is improved, but visibility into data intent and purpose decreases
Solution Approach 1:
The patent attaches metadata and intent information to datasets at the time of their creation or placement. This preliminary action ensures that data intent and purpose are captured upfront and maintained throughout the data lifecycle, enabling the system to scale distributedly while preserving complete visibility into data meaning and purpose
3Productivity
If automated data placement is performed, then data management efficiency is improved, but policy analysis complexity increases
Solution Approach 1:
The patent uses data descriptors as standardized templates or copies that define data placement policies. Instead of complex ad-hoc decision logic, the system uses pre-defined policy templates that can be automatically applied to different datasets and edge devices, simplifying the policy analysis process while maintaining high automation efficiency
Data Source
AI summary
Techniques described herein relate to a method for distributed data management. The method may include obtaining data descriptors for an application executing on a data host, performing a dataset policy analysis using the data descriptors to determine a data placement for a dataset associated with the application using a global name repository, performing, based on the data policy analysis, the data placement, and based on the data placement, updating the global name repository.


