Coherent Host-Managed Device Memory Pre-Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional coherent shared memory systems are limited by bandwidth and latency issues in chip-to-chip interconnects and external buses, leading to increased data movement and power consumption during pre-processing operations for computationally intensive workloads like AI and ML, which require extensive data decoding and transformation.

Innovation Solution

The implementation of storage devices with data-processing engines that perform pre- and post-processing operations on data within coherent host-managed device memory, reducing data movement by processing data locally before transmission to host processors or accelerators via cache-coherent interconnects.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If data is processed externally by host processors or accelerators, then computational capability is improved, but data movement and power consumption increase

Engineering Contradiction:
Improvecomputational capabilityVSAvoidpower consumption
Core Design Contradiction:
Extent of automationVSLoss of energy

Solution Approach 1:

The patent merges storage and processing functions by integrating data-processing engines directly into the storage device. This allows pre-processing operations to be performed on data while it resides in device memory, combining what were previously separate storage and computation operations into a unified system that reduces data movement and power consumption.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The storage device with integrated data-processing engines acts as an intermediary between device memory and host processors. Instead of moving all data to external processors for computation, the storage device performs preliminary processing operations locally, reducing the computational burden on host processors while minimizing data movement across the system.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If data is moved across chip-to-chip interconnects and external buses, then data accessibility is improved, but bandwidth consumption and latency increase

Engineering Contradiction:
Improvedata accessibilityVSAvoidlatency
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system performs preliminary data processing operations within the storage device before data needs to be accessed by host processors. By pre-processing data while it remains in device memory, the system prepares data in advance, reducing the time and bandwidth required for subsequent data retrieval and transmission operations.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If pre-processing operations are performed externally, then data transformation capability is improved, but data movement increases

Engineering Contradiction:
Improvedata transformation capabilityVSAvoiddata movement
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent merges storage and processing functions by integrating data-processing engines directly into the storage device. This allows pre-processing operations to be performed on data while it resides in device memory, combining what were previously separate storage and computation operations into a unified system that reduces data movement and power consumption.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11775434B2Systems and methods for pre-processing and post-processing coherent host-managed device memory
Publication Date: 2023.10.03 META PLATFORMS INC
  • US11775434B2 patent drawing
  • US11775434B2 patent drawing
  • US11775434B2 patent drawing

AI summary

The disclosed computer-implemented method may include receiving, from a host via a cache-coherent interconnect, a request to access an address of a coherent memory space of the host. When the request is to write data, the computer-implemented method may include (1) performing, after receiving the data, a post-processing operation on the data to generate post-processed data and (2) writing the post-processed data to a physical address of a device-attached physical memory mapped to the address. When the request is to read data, the computer-implemented method may include (1) reading the data from the physical address of a device-attached physical memory mapped to the address, (2) performing, before responding to the request, a pre-processing operation on the data to generate pre-processed data, and (3) returning the pre-processed data to the external host via the cache-coherent interconnect. Various other methods, systems, and computer-readable media are also disclosed.