Tiled Neural Processor With Memristor Crossbars for Von Neumann Bottlenecks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional computer architectures are not optimal for implementing neural nets due to the von Neumann bottleneck and lack of general-purpose neural processor designs, which are necessary for emerging applications like image classification and malware analysis.
Innovation Solution
A computer processor with a tiled array architecture and a compact, power-efficient comparator design, utilizing memristor technology for dynamic weight programming and on-chip learning, allowing for flexible and efficient implementation of neural networks across various array sizes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional von Neumann architecture is used, then general-purpose computing is achieved, but performance and energy efficiency deteriorate due to the von Neumann bottleneck
Solution Approach 1:
The patent merges memory and processing functions into a unified processor-in-memory architecture. The neural network weights are stored in memory cells that directly perform multiply-accumulate operations, eliminating the need to move data between separate memory and processor units. This integration resolves the von Neumann bottleneck while maintaining adaptability for various neural network workloads.
Solution Approach 2:
The processor is designed with a configurable array of processing elements that can be dynamically allocated to handle different neural network architectures and workloads. The same hardware infrastructure supports various network sizes and configurations, providing general-purpose capability without requiring application-specific customization for each task.
2Adaptability or versatility
If traditional von Neumann architecture is used, then general-purpose computing is achieved, but energy consumption increases due to frequent data movement
Solution Approach 1:
By combining memory storage and processing operations into a single integrated structure, the patent eliminates energy-intensive data transfers between separate memory and processor components. The processing elements are embedded within the memory array, allowing computations to be performed directly where data is stored, thereby dramatically reducing energy consumption while supporting general-purpose neural network workloads.
3Productivity
If fixed-functionality processors are used, then performance for specific applications is improved, but adaptability to different neural network sizes and applications deteriorates
Solution Approach 1:
The processor employs dynamically configurable processing elements that can be activated or deactivated based on the specific neural network workload. The array structure allows flexible allocation of computational resources, enabling the system to optimize performance for different network sizes and types while maintaining the same physical hardware infrastructure.
Solution Approach 2:
The processor is designed as a universal platform that can accommodate various neural network architectures through configurable processing elements. The same hardware can be adapted to handle different network sizes, layers, and operation types, providing both high performance and broad adaptability without requiring multiple specialized processors.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables efficient and flexible implementation of neural networks, overcoming the limitations of traditional architectures by providing a general-purpose neural processor capable of handling diverse neural network sizes and applications with reduced power consumption and increased performance.
Implementation Method 1
a programmable resistor connecting the voltage signal to the first bit line
Data Source
AI summary
A computer processor includes an on-chip network and a plurality of tiles. Each tile includes an input circuit to receive a voltage signal from the network, and a crossbar array, including at least one neuron. The neuron includes first and second bit lines, a programmable resistor connecting the voltage signal to the first bit line, and a comparator to receive inputs from the two bit lines and to output a voltage, when a bypass condition is not active. Each tile includes a programming circuit to set a resistance value of the resistor, a pass-through circuit to provide the voltage signal to an input circuit of a first additional tile, when a pass-through condition is active, a bypass circuit to provide values of the bit lines to a second additional tile, when the bypass condition is active; and at least one output circuit to provide an output signal to the network.


