Memory Sub-System with Internal ML Logic for Low-Latency Inference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional memory sub-systems experience latency issues due to the need for external buses or interfaces to transmit data, machine learning models, and intermediate data between memory components and machine learning processors, which hampers the performance of machine learning operations.

Innovation Solution

Implementing internal logic within memory components to perform machine learning operations, eliminating the need for external machine learning processors and reducing data transmission latency by using digital logic or resistor arrays integrated within the memory components.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If external buses and interfaces are used to transmit data between memory components and machine learning processors, then connectivity and data transmission capability are improved, but latency increases and performance deteriorates

Engineering Contradiction:
Improvedata transmission capabilityVSAvoiddata transmission latency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent merges the machine learning processor functionality directly into the memory component by integrating a neural network engine within the memory device. This integration eliminates the need for separate external machines to perform machine learning operations, thereby reducing data transmission latency while maintaining connectivity through internal buses within the memory component itself.

Inventive Principle:
Principle #5Merging (Combining)

2Productivity

If separate external machines are used for machine learning operations, then processing capability is improved, but system complexity and data transmission requirements increase

Engineering Contradiction:
Improvemachine learning processing capabilityVSAvoidsystem architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent combines the machine learning processing capability directly within the memory component by integrating a neural network engine. This integration consolidates what were previously separate external machines into a single unified device, reducing system architecture complexity while maintaining machine learning processing capability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The memory component is designed to perform multiple functions: it serves as both data storage and machine learning processing unit. The neural network engine within the memory device can execute machine learning operations directly on stored data, making the memory component a universal device that handles both storage and processing tasks.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of time

If internal logic is implemented within memory components to perform machine learning operations, then data transmission latency is reduced, but manufacturing complexity increases

Engineering Contradiction:
Improvedata transmission timeVSAvoidmemory component manufacturing complexity
Core Design Contradiction:
Loss of timeVSEase of manufacture

Solution Approach 1:

The patent integrates a neural network engine directly into the memory component structure, merging storage and processing functions. This integration reduces data transmission time by eliminating external communication requirements while managing manufacturing complexity through standardized integration processes.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11694076B2Memory sub-system with internal logic to perform a machine learning operation
Publication Date: 2023.07.04 MICRON TECHNOLOGY INC
  • US11694076B2 patent drawing
  • US11694076B2 patent drawing
  • US11694076B2 patent drawing

AI summary

A memory component can include memory cells where a first region of the memory cells is to store a machine learning model and a second region of the memory cells is to store input data and output data of a machine learning operation. A controller can be coupled to the memory component with one more internal buses to perform the machine learning operation by applying the machine learning model to the input data to generate the output data.