Neural Network Processing Circuit in Memory Device
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current memory devices lack the capability to perform neural network processing efficiently, as they are not designed to handle the complex operations required for deep learning techniques like Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs within their internal processing circuits.
Innovation Solution
A memory device is designed with a neural network processing circuit that includes multiple cell array regions, a computation processing block, a data operation block, and an operation control block, enabling parallel processing of neural network operations across multiple layers, with the ability to store and retrieve weight and computation information, and perform neural network computations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a memory device is designed with traditional memory architecture, then it can store data reliably, but it cannot perform neural network processing operations
Solution Approach 1:
The memory device is designed to perform both traditional memory storage functions and neural network processing operations. The computation processing block enables the memory device to execute neural network computations (multiplication and accumulation operations) directly within the memory array, allowing the same hardware to serve dual purposes: data storage and data processing.
Solution Approach 2:
The patent combines memory storage functionality with computation processing functionality into a single integrated device. The memory cell array regions are coupled with computation processing blocks that can perform neural network operations, merging what were traditionally separate components (memory and processor) into one unified system.
2Productivity
If neural network processing is performed using external processors, then processing capability is sufficient, but data transfer time and energy consumption increase
Solution Approach 1:
The memory device is divided into N cell array regions that can independently store and process data. Each region can be configured to store different types of data (input data, weight information, computation information, or computation-completed data), enabling parallel processing operations and reducing the time needed to move data between storage and processing units.
Solution Approach 2:
The patent introduces a new operational dimension by enabling in-memory computation. Instead of the traditional von Neumann architecture where data must be transferred between separate memory and processor units, the computation processing block allows neural network operations to be performed directly within the memory array, adding a spatial dimension to data processing efficiency.
3Productivity
If parallel processing is implemented across multiple memory regions, then processing throughput increases, but control complexity increases
Solution Approach 1:
The operation control block serves as an intermediary that manages the complex coordination between N cell array regions and M computation processing blocks. It receives commands and addresses, decodes them appropriately, and generates the necessary control signals to coordinate parallel operations across multiple memory regions, thereby managing control complexity centrally rather than distributing it throughout the system.
Solution Approach 2:
The device includes buffer regions that can pre-load and store data (input data, weight information, computation information) before processing operations begin. This preliminary action allows data to be staged and organized in advance, reducing the complexity of real-time data management during parallel processing operations.
Data Source
AI summary
A memory device comprising: N cell array regions, a computation processing block suitable for generating computation-completion data by performing a network-level operation on input data, the network-level operation indicating an operation of repeating a layer-level operation M times in a loop, the layer-level operation indicating an operation of performing N neural network computations in parallel, a data operation block suitable for storing the input data and (M*N) pieces of neural network processing information in the N cell array regions, and outputting the computation-completion data through the data transfer buffer, and an operation control block suitable for controlling the computation processing block and the data operation block.


