Hybrid Cloud Data Clustering for Machine Learning Transfer
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current hybrid cloud systems face challenges in optimizing the transfer of data from private to public clouds for machine learning activities, balancing data security and performance, as existing methods do not effectively evaluate the resource allocation impact and information gain associated with data sensitivity and transfer costs.
Innovation Solution
The method involves clustering data elements based on attribute sensitivity, computing resource allocation impact, and information gain, and optimizing data transfer by determining the sensitivity values using natural language processing and machine learning algorithms like WORD2VEC, to identify the most beneficial data to transfer to public cloud resources for machine learning activities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is transferred from private cloud to public cloud for machine learning activities, then machine learning model performance is improved, but data security risks increase
Solution Approach 1:
The patent segments data into clusters based on sensitivity attributes, allowing different portions of data to be transferred to public cloud based on their security classification. This enables selective data transfer that improves machine learning performance while maintaining security control over sensitive information.
Solution Approach 2:
The patent applies different security policies and transfer decisions to different data elements based on their local sensitivity characteristics. Each data element is evaluated individually for its sensitivity attributes, allowing tailored decisions about which specific data to transfer to public cloud resources.
2Productivity
If more data is transferred to public cloud resources, then machine learning model performance improves, but transfer and processing costs increase
Solution Approach 1:
The patent changes the parameters used for data transfer decisions by evaluating information gain metrics and sensitivity attributes. This allows optimization of which data to transfer based on the actual value it provides to machine learning models, rather than transferring all data uniformly.
Solution Approach 2:
The patent transfers only the necessary portion of data to public cloud resources based on calculated information gain and sensitivity analysis. This partial action approach avoids the excessive cost of transferring all data while still achieving sufficient machine learning model performance.
3Reliability
If data sensitivity evaluation is performed for each data element, then data security is improved, but computational complexity increases
Solution Approach 1:
The patent segments data into clusters based on sensitivity attributes, reducing the computational burden by processing groups of similar data elements together rather than evaluating each element completely independently.
Solution Approach 2:
The system uses automated sensitivity evaluation and clustering algorithms that self-manage the complex computational tasks of data classification and security assessment, reducing the need for manual intervention in complex security evaluations.
Data Source
AI summary
Managing hybrid cloud resources by grouping at least a portion of the elements of a data set according to attribute sensitivity into a cluster of elements, computing a resource allocation impact of the cluster of elements, computing an information gain associated with the set of elements, and allocating cloud resources according to the resource allocation impact and information gain.


