Integrated GPU Storage for Big Data Analytics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current big data analytics systems face performance bottlenecks due to I/O limitations, as they rely on conventional storage systems that introduce latencies and bus congestion, limiting the ability to process large volumes of unstructured or semi-structured data efficiently, even with high-end GPGPU expansion cards.
Innovation Solution
An integrated storage/processing system with a non-volatile memory array directly coupled to a GPU, allowing for low-latency data access without relying on host system memory, using PCIe-based interfaces and peer-to-peer data transfers to bypass traditional storage bottlenecks, enabling faster throughput and parallel processing capabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional storage systems (hard disk drives or NAND flash-based SSDs) are used to store large data sets before loading into system memory, then data storage capacity is sufficient, but I/O latency and bus congestion increase significantly
Solution Approach 1:
The patent merges the storage device and processing device into a single integrated unit, where the non-volatile memory array is directly coupled to the GPU through a memory controller on the same expansion card. This eliminates the need for data to travel through the host system's PCIe root complex and system memory, thereby reducing I/O latency and bus congestion while maintaining large storage capacity.
2Ease of operation
If data are transferred through multiple hops (storage array → system memory → DMA channel → local frame buffer), then data can be accessed by processing units, but the number of protocol conversions and latency increases
Solution Approach 1:
The patent extracts the memory controller function from the host system and integrates it directly onto the expansion card with the GPU. This allows the GPU to access the non-volatile memory array directly without going through the host system's memory subsystem, eliminating multiple protocol conversions and reducing transfer latency while maintaining data accessibility.
3Power
If high-end GPGPU expansion cards with large local frame buffers are used, then computational processing capability is enhanced, but the on-board volatile memory capacity remains limited to 6 GB
Solution Approach 1:
The patent implements a nested memory hierarchy where a large non-volatile memory array (providing terabyte-scale capacity) is nested within the expansion card that also contains the GPU and smaller volatile frame buffer. This allows the system to maintain the 6 GB volatile buffer for active processing while providing access to much larger容量的非易失性存储,effectively combining the benefits of both small fast memory and large slow memory.
4Adaptability or versatility
If conventional PCIe root complex architecture is used for data transfer, then system compatibility is maintained, but bus congestion and protocol conversions occur
Solution Approach 1:
The patent introduces a dedicated memory controller as an intermediary between the GPU and the non-volatile memory array, eliminating the need for data to pass through the PCIe root complex. This intermediary handles all memory access operations locally on the expansion card, maintaining system compatibility through standard PCIe interfaces while dramatically improving data transfer throughput by avoiding bus congestion and protocol conversions.
Data Source
AI summary
Architectures and methods for performing big data analytics by providing an integrated storage/processing system containing non-volatile memory devices that form a large, non-volatile memory array and a graphics processing unit (GPU) configured for general purpose (GPGPU) computing. The non-volatile memory array is directly functionally coupled (local) with the GPU and optionally mounted on the same board (on-board) as the GPU.


