Cloud Bursting Data Copy and Node Scaling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data processing systems face performance deterioration and inability to cope with unexpected loads due to long data copy times during cloud bursting, especially when large amounts of data need to be copied from an on-premises environment to a public cloud, leading to communication bandwidth limitations.
Innovation Solution
A computer system and method that includes a management node, a processing cluster with scalable processing nodes, and a storage cluster, allowing for distributed data processing and storage. During cloud bursting, data is copied from a first data storage area to a second area in the public cloud while simultaneously increasing the number of processing nodes, thereby reducing copy waiting time and improving performance by parallel processing and caching frequently accessed data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data is copied from on-premises environment to public cloud during cloud bursting, then the system can handle unexpected loads, but the communication bandwidth limitation causes long copy time and performance deterioration
Solution Approach 1:
The system performs data copying in advance before the actual load occurs. When load increases are predicted or detected, the system proactively copies required data from on-premises storage to cloud storage beforehand, so that when the load actually arrives, the data is already available in the cloud, eliminating copy waiting time and performance deterioration.
Solution Approach 2:
The system dynamically adjusts the data copying strategy based on real-time load conditions. When load increases are detected, the system activates cloud bursting and initiates data copying. The copying process is dynamically managed based on available bandwidth, priority of data, and current system state, allowing flexible adaptation to changing conditions.
2Productivity
If cloud bursting is executed to handle unexpected load, then on-premises resources are not overloaded, but data must be copied from on-premises to cloud which takes time due to bandwidth limits
Solution Approach 1:
The system performs data copying in advance before the actual load occurs. When load increases are predicted or detected, the system proactively copies required data from on-premises storage to cloud storage beforehand, so that when the load actually arrives, the data is already available in the cloud, eliminating copy waiting time and performance deterioration.
Solution Approach 2:
The system ensures continuous data processing by maintaining data availability in both on-premises and cloud environments. During the transition to cloud bursting, the system continues processing operations without interruption by coordinating data access between both locations, ensuring uninterrupted productive action.
3Loss of time
If a part of data is copied to public cloud to reduce copy time, then some load can be handled, but the amount of data to be copied is still significant and copy waiting time remains long
Solution Approach 1:
The system applies different quality levels to different parts of the data based on their access patterns and importance. Frequently accessed or critical data is prioritized for copying to the cloud with higher bandwidth allocation, while less critical data remains on-premises. This selective approach ensures that the most important data is available in the cloud without copying the entire dataset.
Solution Approach 2:
The system changes parameters such as data selection criteria, copying priority, and bandwidth allocation based on current load conditions and data characteristics. When cloud bursting is activated, the system adjusts which data parameters are copied first, how much bandwidth is allocated to copying versus processing, and dynamically reconfigures the data distribution between on-premises and cloud environments.
Data Source
AI summary
An object of the present invention is to provide a computer system and a scale-out method of the computer system that can reduce the possibility of occurrence of performance deterioration in the case where cloud bursting is executed.In the case where scale-out is executed in a public cloud, a data processing system starts a data copy process of copying data stored in an on-premises first data storage area, to a second data storage area of a storage cluster in the public cloud via a first network. When starting the data copy process, the data processing system executes the scale-out by increasing the number of processing nodes while accessing the data stored in the first data storage area.


