Data Characterization Engine for Multicloud Intent Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Information processing systems, especially those in distributed multicloud edge platforms, face challenges in efficiently managing vast amounts of data generated by microservices, as existing methods largely ignore data characterization and rely on applications or storage services, leading to accessibility and cost issues due to nonlocality and untimely reachability assumptions.
Innovation Solution
Implementing a data characterization engine with a machine learning-based classification process that detects data intent, utilizing a multicloud edge platform to automatically select the most appropriate classifier for each application use case, enabling data visibility, access, movement, security, and orchestration decisions through a machine learning classification sub-system, feature extraction and selection sub-system, and parametric meta-learning decisioning sub-system.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If data characterization is left to applications or storage services, then application autonomy is maintained, but data management efficiency and accessibility deteriorate
Solution Approach 1:
The patent introduces a data characterization engine as an intermediary component between applications and storage services. This engine automatically detects data sources, classifies data by intent, and manages data characterization without requiring applications to perform these tasks themselves, thus maintaining application autonomy while improving data management efficiency
Solution Approach 2:
The data characterization engine operates autonomously to perform data detection, classification, and characterization tasks. It self-manages the complexity of data management by automatically selecting classifiers and determining data intent, freeing applications from these burdens while maintaining system efficiency
2Adaptability or versatility
If data is distributed across multiple cloud platforms, then processing capability and scalability are improved, but data accessibility and reachability worsen due to nonlocality
Solution Approach 1:
The data characterization engine provides universal data classification capabilities across multiple cloud platforms. By implementing a unified classification system that works consistently across different cloud environments, it enables data accessibility and reachability while maintaining the processing capability and scalability benefits of distributed architecture
Solution Approach 2:
The system changes the parameter of data characterization by introducing intent-based classification. This transformation allows data to be identified and accessed based on its purpose and meaning rather than its physical location, improving accessibility across distributed cloud platforms while preserving processing capabilities
3Device complexity
If traditional classification methods are used, then system complexity is reduced, but classification accuracy and data intent detection worsen
Solution Approach 1:
The patent replaces traditional rule-based or manual classification methods with machine learning-based classification. This substitution enables accurate detection of data intent and automatic selection of appropriate classifiers, significantly improving classification accuracy while the automated nature of the system manages the complexity rather than increasing it
4Speed
If data characterization is not performed, then processing speed is maintained, but data management costs and accessibility issues increase
Solution Approach 1:
The data characterization engine performs classification and characterization actions in advance, before data needs to be accessed or processed. By pre-tagging and organizing data based on intent, it enables faster subsequent access and reduces the need for expensive data egress operations, thus maintaining processing speed while reducing management costs
Data Source
AI summary
Data characterization techniques in an information processing system environment are disclosed. In one example, at least one processing device is configured to detect a source application associated with data obtained from execution of at least one of a plurality of applications in an information processing system, wherein the plurality of applications comprise services associated with multiple different policies. The processing device is further configured to classify the data to determine an intent associated with the data, wherein classifying comprises utilizing a machine learning classification process.


