Distributed Policy Agent for Edge Data Intent Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Managing vast amounts of data produced at the edge of device ecosystems is challenging due to complexity, scale, and limited visibility, with existing frameworks struggling to determine data intent and applicability, especially in distributed edge systems where computational resources are limited and network connectivity is not guaranteed.
Innovation Solution
Implementing monitor agents on computing devices that communicate with a global policy manager to classify data intent using heuristic rules and machine learning classifiers, generating global names and metadata that convey semantic meaning, and publishing this information to a global name repository for data management and service determination.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If distributed data management is implemented at edge devices, then data visibility and management capability are improved, but device complexity and computational resource requirements increase
Solution Approach 1:
The system segments data management functionality into distributed monitor agents deployed on individual edge devices. Each agent independently monitors and classifies data locally, while a central policy manager coordinates overall data governance. This segmentation enables improved data visibility across the ecosystem without overwhelming individual devices with full management complexity.
Solution Approach 2:
A global policy manager acts as an intermediary between distributed edge devices and the central data management system. The policy manager receives data information from monitor agents, applies classification rules and machine learning models, and returns pricing and management decisions. This intermediary layer reduces the computational burden on edge devices while maintaining centralized control and visibility.
2Measurement precision
If machine learning classifiers are used to determine data intent, then data classification accuracy is improved, but computational resources and processing time are consumed
Solution Approach 1:
The system performs preliminary data classification using heuristic rules and metadata extraction before applying more computationally intensive machine learning models. Monitor agents first gather data information and generate metadata, then the policy manager applies classification rules to determine intent. This preliminary action reduces the need for heavy computational resources by handling simple cases efficiently.
Solution Approach 2:
The system applies machine learning classifiers selectively rather than to all data uniformly. The policy manager determines when ML classification is necessary based on data characteristics and intent determination needs. This partial application of computational resources optimizes the balance between classification accuracy and energy consumption.
Data Source
AI summary
Techniques described herein relate to a method for distributed data management. The method may include making a first determination that data is written to a data structure of storage of a data host; obtaining, based on the first determination, data information associated with the data; making a second determination of intent corresponding to the data; generating a global name and metadata corresponding to the data, wherein the metadata comprises the intent; and publishing the global name and the metadata to a global name repository.


