Memory Sub-System with Internal ML Logic for Low-Latency Inference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional memory sub-systems experience latency issues due to the need for external buses or interfaces to transmit data, machine learning models, and intermediate data between memory components and machine learning processors, which hampers the performance of machine learning operations.
Innovation Solution
Implementing internal logic within memory components to perform machine learning operations, eliminating the need for external machine learning processors and reducing data transmission latency by using digital logic or resistor arrays integrated within the memory components.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If external buses and interfaces are used to transmit data between memory components and machine learning processors, then connectivity and data transmission capability are improved, but latency increases and performance deteriorates
Solution Approach 1:
The patent merges the machine learning processor functionality directly into the memory component by integrating a neural network engine within the memory device. This integration eliminates the need for separate external machines to perform machine learning operations, thereby reducing data transmission latency while maintaining connectivity through internal buses within the memory component itself.
2Productivity
If separate external machines are used for machine learning operations, then processing capability is improved, but system complexity and data transmission requirements increase
Solution Approach 1:
The patent combines the machine learning processing capability directly within the memory component by integrating a neural network engine. This integration consolidates what were previously separate external machines into a single unified device, reducing system architecture complexity while maintaining machine learning processing capability.
Solution Approach 2:
The memory component is designed to perform multiple functions: it serves as both data storage and machine learning processing unit. The neural network engine within the memory device can execute machine learning operations directly on stored data, making the memory component a universal device that handles both storage and processing tasks.
3Loss of time
If internal logic is implemented within memory components to perform machine learning operations, then data transmission latency is reduced, but manufacturing complexity increases
Solution Approach 1:
The patent integrates a neural network engine directly into the memory component structure, merging storage and processing functions. This integration reduces data transmission time by eliminating external communication requirements while managing manufacturing complexity through standardized integration processes.
Data Source
AI summary
A memory component can include memory cells where a first region of the memory cells is to store a machine learning model and a second region of the memory cells is to store input data and output data of a machine learning operation. A controller can be coupled to the memory component with one more internal buses to perform the machine learning operation by applying the machine learning model to the input data to generate the output data.


