Memory Sub-System with Dual Buses for In-Memory ML Throughput

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional memory sub-systems experience latency issues during machine learning operations due to the need for external buses and interfaces to transmit data, models, and intermediate results between memory components and machine learning processors, which hinders performance.

Innovation Solution

Integrating internal logic within memory components to perform machine learning operations, eliminating the need for external machine learning processors and reducing data transmission latency by using internal buses for data and model processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If data is transmitted between memory components and machine learning processors using external buses, then the system can perform machine learning operations, but data transmission latency increases and performance decreases

Engineering Contradiction:
Improvedata transmission speedVSAvoiddata transmission latency
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The patent combines the machine learning processor with the memory sub-system into a single integrated unit. The machine learning processor is coupled to the memory components through internal interconnects rather than external buses, merging previously separate functions into one unified system. This integration eliminates the need for external data transmission and reduces latency.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent extracts the machine learning processing function from external processors and places it directly within the memory sub-system. By taking out the processing function and embedding it in the memory architecture, the system eliminates external bus dependencies for data transmission between storage and processing units.

Inventive Principle:
Principle #2Taking out (Extraction)

2Device complexity

If a single bus is used to transmit both machine learning data and host data, then device complexity is reduced, but data transmission efficiency decreases

Engineering Contradiction:
Improvebus structure complexityVSAvoiddata transmission efficiency
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent segments the data transmission paths by providing separate interfaces: a first interface for machine learning operations and a second interface for host data operations. This segmentation allows parallel processing of different data types without interference, improving overall system productivity while maintaining manageable complexity through structured separation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The memory sub-system is designed with multi-functionality to handle both machine learning workloads and traditional host data operations simultaneously. By incorporating dedicated interfaces for different operation types within a single integrated system, the design achieves universal functionality without requiring completely separate systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11676010B2Memory sub-system with a bus to transmit data for a machine learning operation and another bus to transmit host data
Publication Date: 2023.06.13 MICRON TECHNOLOGY INC
  • US11676010B2 patent drawing
  • US11676010B2 patent drawing
  • US11676010B2 patent drawing

AI summary

A system includes a memory component to store host data from a host system and to store a machine learning model and input data. A controller includes an in-memory logic to perform a machine learning operation by applying the machine learning model to the input data to generate an output data. A bus can receive additional host data from the host system and provide the additional host data to the memory component. An additional bus can receive machine learning data from the host system and provide the machine learning data to the in-memory logic that is to perform the machine learning operation.