Two-Stage Hybrid Memory Buffer for High-Stream NVM Writes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional memory sub-systems with a single buffer, either external DRAM or internal SRAM, struggle to support a high number of streams at high performance, leading to increased costs, power consumption, and form factor issues when scaling to support multiple streams such as 32 to 1024 streams.
Innovation Solution
A two-stage hybrid memory buffer is introduced, comprising a host buffer component (e.g., external DRAM) and a staging buffer component (e.g., internal SRAM), where data from multiple streams is segregated and processed for error protection before being written to NVM memory, allowing for scalable support of multiple streams while maintaining performance and cost efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a single buffer (external DRAM or internal SRAM) is used in conventional memory sub-systems, then the system structure is simple, but the system cannot support a high number of streams at high performance, leading to increased costs, power consumption, and form factor issues when scaling to 32-1024 streams
Solution Approach 1:
The buffer is divided into two distinct stages: a first buffer (host buffer) that receives data from the host system and a second buffer (staging buffer) that prepares data for NVM programming. This segmentation allows each buffer to be optimized for its specific function, enabling support for multiple streams (32-1024) while maintaining manageable complexity through clear functional separation.
Solution Approach 2:
The patent introduces a temporal dimension to the buffer architecture by implementing a two-stage processing flow. Data flows through the first buffer, undergoes processing (error protection, aggregation), then moves to the second buffer before NVM programming. This dimensional expansion in the data flow process enables high-performance multi-stream support without proportionally increasing overall system complexity.
2Productivity
If internal SRAM is used as the buffer, then the bandwidth and performance are high, but the cost and die area become prohibitive when scaling to support multiple streams
Solution Approach 1:
The patent merges two different memory technologies into a hybrid buffer system: SRAM (providing high bandwidth) and DRAM or other non-volatile memory (providing large capacity at lower cost). The first buffer can use DRAM or other memory for capacity, while the second buffer uses SRAM for high-speed processing. This combination allows the system to achieve both high bandwidth and large capacity without the prohibitive cost of using only SRAM for all buffer needs.
Solution Approach 2:
Different parts of the buffer system are assigned different memory technologies based on their specific functional requirements. The staging buffer (second buffer) that requires high bandwidth for NVM programming uses SRAM, while the host buffer (first buffer) that requires large capacity for multi-stream data aggregation uses DRAM or other cost-effective memory. This local optimization of memory quality matches performance characteristics to functional needs.
3Quantity of substance
If external DRAM is used as the buffer, then the cost is lower and capacity is larger, but the bandwidth and performance are significantly reduced compared to SRAM
Solution Approach 1:
The buffer is segmented into two functional stages with different bandwidth requirements. The first buffer (host buffer) using DRAM handles data aggregation from multiple host streams, where large capacity is more critical than maximum bandwidth. The second buffer (staging buffer) using SRAM handles the final programming preparation to NVM, where high bandwidth is essential. This segmentation allows DRAM to provide cost-effective large capacity where it matters most, while SRAM provides high bandwidth where it is most needed.
Solution Approach 2:
The first buffer (host buffer) using DRAM acts as an intermediary that aggregates data from multiple host streams before passing it to the second buffer. This intermediary role allows the system to leverage DRAM's large capacity and low cost for the bulk data aggregation phase, while the SRAM-based second buffer handles the high-bandwidth final processing, thus overcoming DRAM's bandwidth limitations for the critical NVM programming path.
4Adaptability or versatility
If the buffer size is increased to support all streams simultaneously, then all streams can be supported at high performance, but the cost, power consumption, and form factor become prohibitive
Solution Approach 1:
The buffer system is segmented into two stages with distinct functions and optimized capacities. The first buffer aggregates data from multiple streams (32-1024) over time, requiring large capacity but not maximum simultaneous bandwidth. The second buffer processes data for NVM programming, requiring high bandwidth but smaller capacity. This segmentation allows support for many streams without requiring a single enormous high-speed buffer, thus reducing power consumption while maintaining multi-stream capability.
Solution Approach 2:
The first buffer performs preliminary data aggregation and error protection coding for multiple streams before the data is transferred to the second buffer for final NVM programming preparation. This preliminary action allows data from multiple streams to be collected and pre-processed over time in the first buffer, reducing the need for the second buffer to handle all streams simultaneously at full bandwidth, thus lowering overall power consumption while supporting high stream counts.
Data Source
AI summary
Described herein is a system comprising one or more external dynamic random access memory (DRAM) devices having a first programming unit buffer and a second programming unit buffer, an internal static RAM (SRAM) device, one or more non-volatile memory (NVM) devices, and a processing device, operatively coupled with the one or more external DRAM devices, the internal SRAM device, and the one or more NVM devices. The processing device transfers first write data from the first programming unit buffer to the internal SRAM device responsive to the first write data satisfying a programming unit (PU) threshold, the PU threshold pertaining to a PU of the one or more NVM devices. The processing device also writes the first write data from the internal SRAM device as a first programming unit to the one or more NVM devices, and transfers a second write data from the second programming unit buffer to the internal SRAM device responsive to the second write data satisfying the PU threshold. The processing device further writes the second write data from the internal SRAM device as a second programming unit to the one or more NVM devices.


