Coherent Host-Managed Device Memory Pre-Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional coherent shared memory systems are limited by bandwidth and latency issues in chip-to-chip interconnects and external buses, leading to increased data movement and power consumption during pre-processing operations for computationally intensive workloads like AI and ML, which require extensive data decoding and transformation.
Innovation Solution
The implementation of storage devices with data-processing engines that perform pre- and post-processing operations on data within coherent host-managed device memory, reducing data movement by processing data locally before transmission to host processors or accelerators via cache-coherent interconnects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If data is processed externally by host processors or accelerators, then computational capability is improved, but data movement and power consumption increase
Solution Approach 1:
The patent merges storage and processing functions by integrating data-processing engines directly into the storage device. This allows pre-processing operations to be performed on data while it resides in device memory, combining what were previously separate storage and computation operations into a unified system that reduces data movement and power consumption.
Solution Approach 2:
The storage device with integrated data-processing engines acts as an intermediary between device memory and host processors. Instead of moving all data to external processors for computation, the storage device performs preliminary processing operations locally, reducing the computational burden on host processors while minimizing data movement across the system.
2Ease of operation
If data is moved across chip-to-chip interconnects and external buses, then data accessibility is improved, but bandwidth consumption and latency increase
Solution Approach 1:
The system performs preliminary data processing operations within the storage device before data needs to be accessed by host processors. By pre-processing data while it remains in device memory, the system prepares data in advance, reducing the time and bandwidth required for subsequent data retrieval and transmission operations.
3Adaptability or versatility
If pre-processing operations are performed externally, then data transformation capability is improved, but data movement increases
Solution Approach 1:
The patent merges storage and processing functions by integrating data-processing engines directly into the storage device. This allows pre-processing operations to be performed on data while it resides in device memory, combining what were previously separate storage and computation operations into a unified system that reduces data movement and power consumption.
Data Source
AI summary
The disclosed computer-implemented method may include receiving, from a host via a cache-coherent interconnect, a request to access an address of a coherent memory space of the host. When the request is to write data, the computer-implemented method may include (1) performing, after receiving the data, a post-processing operation on the data to generate post-processed data and (2) writing the post-processed data to a physical address of a device-attached physical memory mapped to the address. When the request is to read data, the computer-implemented method may include (1) reading the data from the physical address of a device-attached physical memory mapped to the address, (2) performing, before responding to the request, a pre-processing operation on the data to generate pre-processed data, and (3) returning the pre-processed data to the external host via the cache-coherent interconnect. Various other methods, systems, and computer-readable media are also disclosed.


