Cache Server Content Recovery to Limit SSD Wear
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Cache servers using nonvolatile memory face challenges in managing the wear caused by the finite number of program/erase cycles, leading to content read errors and potential data loss, necessitating a technology to suppress memory wear and ensure reliable content delivery.
Innovation Solution
The cache server employs a processor to determine content recovery based on delivery capability and writing cost, using error correction and selective recovery methods to minimize wear on nonvolatile memory by either writing recovered content to SSD or retaining it in main memory, depending on retention periods and wear levels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If content is recovered by rewriting to nonvolatile memory, then content availability is improved, but nonvolatile memory wear increases
Solution Approach 1:
The system changes the storage location parameter by selectively writing recovered content to either nonvolatile memory or main memory based on wear indicators and retention requirements. This parameter change allows the system to maintain content availability while controlling memory wear by adjusting where content is stored.
Solution Approach 2:
The system introduces main memory as an intermediary storage layer between the cache server and nonvolatile memory. When content needs recovery, it can be written to main memory first, which has higher write endurance, thereby reducing direct wear on the nonvolatile memory while still maintaining content availability for delivery.
2Reliability
If content is recovered and written to nonvolatile memory, then content delivery reliability is improved, but program/erase cycle consumption increases
Solution Approach 1:
The system monitors wear indicators and retention periods as parameters to decide whether to write recovered content to nonvolatile memory or main memory. By dynamically adjusting this parameter based on current memory health and content importance, the system optimizes the balance between delivery reliability and P/E cycle consumption.
Solution Approach 2:
Instead of always writing recovered content to nonvolatile memory, the system applies partial action by selectively writing only when necessary based on wear thresholds and retention requirements. This reduces unnecessary P/E cycles while maintaining adequate content delivery reliability.
3Reliability
If all read errors are treated as nonexistence, then content recovery operations increase, but this leads to excessive nonvolatile memory rewriting
Solution Approach 1:
The system changes the error handling parameter by introducing a verification mechanism that distinguishes between actual content nonexistence and read errors. By adjusting this parameter based on wear indicators and content criticality, the system avoids excessive recovery operations that would lead to unnecessary memory rewriting.
Solution Approach 2:
The system uses feedback from wear indicators and retention period analysis to determine whether to perform content recovery operations. This feedback mechanism prevents excessive recovery operations by only triggering them when truly necessary, thereby reducing unnecessary nonvolatile memory rewriting and extending memory endurance.
Data Source
AI summary
According to one embodiment, a processor of a cache server delivers content acquired from an origin server to a client, using a storage device as a cache of the contents. When an error occurs in reading the content from the storage device, the processor determines whether or not to recover the content based on a recovery amount of a delivery capability of the cache server as a result of recovering the content and a cost of writing data to the storage device associated with recovery of the content. When determined to recover the content, the processor selects a content recovery method based on a first remaining retention period during which the content should be retained and a second remaining retention period until the content is to be erased from the storage device.


