In-Memory Neural Network Processing via Hierarchical Sub-Instructions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current memory devices face challenges in efficiently processing neural network operations due to limitations in power consumption and processing performance, particularly in devices like computers, smartphones, and wearables, where specialized hardware accelerators are needed for imaging and computer vision applications.
Innovation Solution
A memory device with a device controller and processing engine that generates sub-processing instructions from host processing instructions, allowing for operations to be performed near or in memory, utilizing a hierarchical structure of processing engines to optimize data processing and reduce latency and bandwidth usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If specialized hardware accelerators are added to improve processing performance, then neural network processing efficiency is improved, but device complexity increases
Solution Approach 1:
The patent combines memory functionality with processing capabilities by integrating processing engines directly into the memory device. This merging allows the memory device to perform neural network operations (such as matrix multiplication and data processing) without requiring separate dedicated hardware accelerators, thereby improving processing efficiency while avoiding the complexity increase that would result from adding separate accelerator components.
Solution Approach 2:
The memory device is designed with multi-functionality, serving both as data storage and as a processing unit. The processing engines within the memory device can execute various neural network operations, making the memory device universal rather than specialized. This eliminates the need for separate hardware accelerators while maintaining high processing efficiency for neural network workloads.
2Speed
If data processing is performed near or in memory to reduce latency, then processing speed is improved, but device complexity increases
Solution Approach 1:
The patent merges storage and processing functions within the same memory device, enabling data processing to occur near or in memory without requiring additional separate processing components. This integration reduces data transfer latency while avoiding the complexity increase that would result from adding separate processing units, as the processing capability is embedded within the memory structure itself.
3Productivity
If hierarchical processing structure is implemented to optimize data processing, then processing efficiency is improved, but device complexity increases
Solution Approach 1:
The patent implements a hierarchical processing structure that segments neural network operations into different levels of processing engines within the memory device. This segmentation allows complex operations to be divided into manageable tasks handled by specialized processing units at different hierarchy levels, improving processing efficiency through parallelization and task distribution while keeping the overall device complexity manageable through structured organization.
Solution Approach 2:
The hierarchical processing structure introduces an additional organizational dimension to the memory device architecture. By arranging processing engines in hierarchical levels rather than a flat structure, the system achieves improved processing efficiency through layered task execution while managing complexity through clear structural separation and defined interfaces between hierarchy levels.
Data Source
AI summary
A memory device includes: a device controller configured to generate a sub-processing instruction based on a host processing instruction received from a host; and a processing engine configured to perform an operation based on the generated sub-processing instruction.


