Warm-Tier Storage for Search Cost and Retention Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing costs of log analytics on large and growing data sets due to inefficient storage and retrieval methods are becoming untenable, forcing customers to reduce data retention periods, leading to missed insights.
Innovation Solution
A tiered storage structure utilizing hot and warm compute nodes, where frequently accessed data is stored locally by hot nodes and less-frequently accessed data is migrated to remote storage, with metadata management by warm nodes for efficient retrieval.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If all data is stored locally on compute nodes for fast access, then data retrieval speed is improved, but storage cost and computational resource consumption increase
Solution Approach 1:
The patent segments data into different tiers based on access frequency: hot data (frequently accessed) is stored locally on compute nodes, while cold data (rarely accessed) is stored remotely. This segmentation allows the system to maintain fast retrieval speeds for hot data while reducing overall storage resource consumption by moving cold data to cheaper remote storage.
Solution Approach 2:
The patent applies local quality by storing data locally only where it is most needed - on compute nodes for hot data that requires fast access. Less frequently accessed cold data is stored remotely, optimizing the balance between access speed and storage cost by making the storage location dependent on data access patterns.
2Loss of information
If data retention period is extended to maintain more historical data, then more insights are available, but storage cost increases
Solution Approach 1:
The patent segments data retention into two strategies: hot data is retained locally for extended periods to maintain access to frequently needed information, while cold data is retained remotely at lower cost. This allows the system to maintain comprehensive historical data for insights while controlling storage costs through tiered retention.
Solution Approach 2:
The patent changes the storage parameter (local vs. remote storage) based on data access patterns and retention needs. By dynamically adjusting where data is stored based on its hotness and retention requirements, the system optimizes the balance between maintaining historical data for insights and controlling storage expenditure.
3Productivity
If more compute nodes are deployed to handle larger data sets, then data processing capability is improved, but system cost increases
Solution Approach 1:
The patent segments data processing workload based on data location: hot data processing is handled by local compute nodes for fast operation, while cold data processing is handled by remote compute nodes. This segmentation allows the system to scale computational resources dynamically based on actual processing needs rather than provisioning for maximum capacity uniformly.
Solution Approach 2:
The patent applies partial action by deploying compute nodes selectively based on data hotness - local nodes for hot data processing and remote nodes for cold data processing. This partial deployment strategy reduces overall computational resource consumption while maintaining adequate processing capability for the most critical data access patterns.
Data Source
AI summary
Systems and techniques are described herein for tiered storage of customer data accessed by a search service of a computing resource service provider. In some aspects, customer data may be received by a search instance executed across a plurality of compute nodes and provisioned by a search service. The customer data may be indexed and the data and resulting index may be stored locally by a first pool of hot compute nodes of the search instance. The customer data and index may be migrated and stored remotely by a data storage service. Metadata associated with the customer data and/or index may be stored in a second pool of warm compute nodes of the search instance. The warm compute nodes, upon receiving a request to access the customer data, may identify a location of the customer data and retrieve the customer data from the data storage service according to the metadata.


