Hadoop Storage Hierarchy for C3 Instance Bottlenecks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

New generation C3 instances in cloud computing offer improved performance but have limited local instance storage, making them unsuitable for processing large volumes of data in Hadoop clusters, and relying on additional storage solutions like EBS volumes increases costs and degrades performance due to bursty operations and potential outages.

Innovation Solution

Implementing a disk hierarchy that prioritizes local SSD storage and uses reserved EBS volumes only when local storage is insufficient, with delayed mounting of reserved volumes until disk utilization thresholds are met and immediate termination during downscaling, along with data balancing to minimize EBS usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If new generation C3 instances are used to improve compute performance, then processing speed and memory capacity are improved, but local instance storage capacity deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidlocal instance storage capacity
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The storage system is segmented into multiple layers: local instance storage (fast, small capacity) and EBS volumes (slower, large capacity). The system divides data storage responsibilities between these segments, with local storage handling frequently accessed data and EBS volumes handling bulk storage needs.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different storage locations are assigned different quality characteristics. Local instance storage provides high-speed access for critical data, while EBS volumes provide abundant capacity for less frequently accessed data. The system optimizes by placing the right data in the right storage location based on access patterns.

Inventive Principle:
Principle #3Local quality

2Quantity of substance

If EBS volumes are used to compensate for limited local storage, then storage capacity is improved, but system reliability deteriorates due to bursty operations and potential outages

Engineering Contradiction:
Improvestorage capacityVSAvoidsystem reliability
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The system prepares by pre-positioning data on local instance storage before it is needed, creating a buffer or cushion against EBS volume unavailability. This ensures that critical data is already cached locally and can be accessed even if EBS volumes experience outages or bursty behavior.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Solution Approach 2:

Local instance storage acts as an intermediary layer between the compute instance and EBS volumes. It mediates access to data, reducing direct dependency on EBS volumes and protecting the system from their reliability issues.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If EBS volumes are mounted immediately upon instance addition, then storage availability is improved, but additional storage costs increase

Engineering Contradiction:
Improvestorage availabilityVSAvoidadditional storage costs
Core Design Contradiction:
Ease of operationVSLoss of energy

Solution Approach 1:

The EBS volume mounting strategy is made dynamic rather than static. Volumes are mounted automatically when instances are added to the cluster, but the system continuously monitors local storage utilization and dynamically adjusts EBS mounting decisions based on current needs, optimizing the balance between availability and cost.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the parameter of EBS volume mounting from a fixed immediate action to a conditional action based on storage utilization thresholds. When local storage utilization exceeds predefined thresholds, EBS volumes are mounted; when thresholds are not exceeded, mounting is delayed, thereby reducing unnecessary storage costs.

Inventive Principle:
Principle #35Parameter changes

4Reliability

If reserved disk storage is mounted on all instances, then storage reliability is improved, but device complexity increases

Engineering Contradiction:
Improvestorage reliabilityVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

Instead of mounting EBS volumes on all instances (excessive action), the system applies partial action by mounting volumes only on instances where local storage utilization exceeds predefined thresholds. This reduces overall system complexity while maintaining reliability where needed.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10606478B2High performance hadoop with new generation instances
Publication Date: 2020.03.31 QUBOLE INC
  • US10606478B2 patent drawing
  • US10606478B2 patent drawing

AI summary

The present invention is generally directed to a distributed computing system comprising a plurality of computational clusters, each computational cluster comprising a plurality of compute optimized instances, each instance comprising local instance data storage and in communication with reserved disk storage, wherein processing hierarchy provides priority to local instance data storage before providing priority to reserved disk storage.