Hybrid Cloud Data Clustering for Machine Learning Transfer

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current hybrid cloud systems face challenges in optimizing the transfer of data from private to public clouds for machine learning activities, balancing data security and performance, as existing methods do not effectively evaluate the resource allocation impact and information gain associated with data sensitivity and transfer costs.

Innovation Solution

The method involves clustering data elements based on attribute sensitivity, computing resource allocation impact, and information gain, and optimizing data transfer by determining the sensitivity values using natural language processing and machine learning algorithms like WORD2VEC, to identify the most beneficial data to transfer to public cloud resources for machine learning activities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If data is transferred from private cloud to public cloud for machine learning activities, then machine learning model performance is improved, but data security risks increase

Engineering Contradiction:
Improvemachine learning model performanceVSAvoiddata security
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments data into clusters based on sensitivity attributes, allowing different portions of data to be transferred to public cloud based on their security classification. This enables selective data transfer that improves machine learning performance while maintaining security control over sensitive information.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different security policies and transfer decisions to different data elements based on their local sensitivity characteristics. Each data element is evaluated individually for its sensitivity attributes, allowing tailored decisions about which specific data to transfer to public cloud resources.

Inventive Principle:
Principle #3Local quality

2Productivity

If more data is transferred to public cloud resources, then machine learning model performance improves, but transfer and processing costs increase

Engineering Contradiction:
Improvemachine learning model performanceVSAvoidtransfer and processing costs
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent changes the parameters used for data transfer decisions by evaluating information gain metrics and sensitivity attributes. This allows optimization of which data to transfer based on the actual value it provides to machine learning models, rather than transferring all data uniformly.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent transfers only the necessary portion of data to public cloud resources based on calculated information gain and sensitivity analysis. This partial action approach avoids the excessive cost of transferring all data while still achieving sufficient machine learning model performance.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If data sensitivity evaluation is performed for each data element, then data security is improved, but computational complexity increases

Engineering Contradiction:
Improvedata securityVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments data into clusters based on sensitivity attributes, reducing the computational burden by processing groups of similar data elements together rather than evaluating each element completely independently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system uses automated sensitivity evaluation and clustering algorithms that self-manage the complex computational tasks of data classification and security assessment, reducing the need for manual intervention in complex security evaluations.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11593013B2Management of data in a hybrid cloud for use in machine learning activities
Publication Date: 2023.02.28 KYNDRYL INC
  • US11593013B2 patent drawing
  • US11593013B2 patent drawing
  • US11593013B2 patent drawing

AI summary

Managing hybrid cloud resources by grouping at least a portion of the elements of a data set according to attribute sensitivity into a cluster of elements, computing a resource allocation impact of the cluster of elements, computing an information gain associated with the set of elements, and allocating cloud resources according to the resource allocation impact and information gain.