Storage-Server NDP Architecture for Data Movement Bottlenecks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modern data centers suffer from performance and energy efficiency imbalances due to latency bottlenecks in data processing and movement, exacerbated by big-data analytics and compute-intensive applications, leading to underutilization of components and increased energy consumption.
Innovation Solution
Implementing a server system architecture with near-data processing (NDP) engines, such as CPUs and FPGAs, interposed between processors and storage devices to minimize data movement and reduce load on processors, using a communication fabric and switch layer to optimize data processing near the data location.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is processed through traditional server architecture with centralized processing, then system simplicity is maintained, but performance bottleneck and energy efficiency degrade due to excessive data movement and processor load
Solution Approach 1:
The patent divides the centralized processing function into distributed NDP engines positioned at storage devices. Each storage device is paired with its own NDP engine, segmenting the processing workload across multiple locations rather than concentrating it at a single centralized processor. This segmentation reduces data movement distance and processing bottlenecks while maintaining architectural modularity.
Solution Approach 2:
The patent introduces NDP engines as intermediary components between storage devices and the host processor. These intermediaries handle data processing tasks locally at storage locations, reducing the burden on the central processor and minimizing data movement through the system. The NDP engines act as mediators that offload processing functions from the centralized architecture.
2Loss of energy
If data movement is reduced through near-data processing, then energy efficiency improves, but system complexity increases due to additional processing components
Solution Approach 1:
The patent implements self-service processing where storage devices equipped with NDP engines process data locally without requiring constant intervention from the host processor. The NDP engines autonomously handle data processing tasks, reducing energy consumption by eliminating unnecessary data transfers and processor wake-ups. This self-service capability allows the system to maintain low power states while still performing required processing functions.
3Use of energy by stationary object
If processor load is reduced by offloading to NDP engines, then energy efficiency improves, but data processing latency may increase due to additional processing stages
Solution Approach 1:
The patent applies preliminary action by having NDP engines perform data processing tasks locally at storage locations before data needs to be accessed by the host processor. This advance processing reduces the amount of data that needs to be transferred and processed later, effectively reducing overall latency despite the additional processing stage. The NDP engines prepare data in advance, so when the host processor needs it, the data is already ready in an optimized state.
Data Source
AI summary
A server system includes a first plurality of mass-storage devices, a central processing unit (CPU), and at least one near data processing (NDP) engine. The CPU is coupled to the first plurality of the mass-storage devices, such as solid-state drive (SSD) devices, and the at least one NDP engine is associated with a second plurality of the mass-storage devices and interposed between the CPU and the second plurality of the mass-storage devices associated with the NDP engine. The second plurality of the mass-storage devices is less than or equal to the first plurality of the mass-storage devices. A number of NDP engines may be based on a minimum bandwidth of a bandwidth associated with the CPU, a bandwidth associated with a network, a bandwidth associated with the communication fabric and a bandwidth associated with all NDP engines divided by a bandwidth associated with a single NDP engine.


