Computational Memory Array for Neural Network Power Efficiency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Deep learning and neural network computations require significant power due to data movement between memory and processing elements, leading to inefficiencies such as high power consumption, increased complexity, and larger chip area requirements, particularly in battery-powered devices.

Innovation Solution

A processing device with an array of interconnected processing elements that allows direct communication among neighbors, enabling parallel operations and efficient data movement through interconnections, reducing the need for long data paths and optimizing memory access.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Power

If data is moved between memory and processing elements in traditional computer architecture, then computations can be performed, but power consumption increases significantly

Engineering Contradiction:
Improvepower consumptionVSAvoidcomputational efficiency
Core Design Contradiction:
PowerVSProductivity

Solution Approach 1:

The patent merges memory and processing elements into a unified neuromorphic architecture where synaptic weights are stored locally at the processing elements. This eliminates the need for frequent data movement between separate memory and processing units, thereby reducing power consumption while maintaining computational efficiency through parallel neural network operations.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transitions from traditional von Neumann architecture to a neuromorphic architecture that adds a spatial dimension to computation by organizing processing elements in arrays with direct interconnections. This dimensional change enables parallel processing of neural network operations without requiring data to traverse long physical paths, thus reducing power consumption while enhancing computational throughput.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If traditional computer architecture is used for deep learning computations, then general-purpose processing is achieved, but processing time increases

Engineering Contradiction:
Improveprocessing speedVSAvoidarchitecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the processing system into multiple independent neuromorphic processing elements arranged in arrays. Each element can perform neural network operations independently and in parallel, significantly reducing processing time for deep learning computations. The segmented architecture maintains manageable complexity through standardized interconnections between elements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The neuromorphic processing elements are designed with universal functionality to perform various neural network operations including matrix multiplications, activations, and pooling operations. This multi-functionality allows a single array architecture to handle different deep learning workloads without requiring specialized hardware for each operation, thus improving processing speed while controlling complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If more processing elements are added to handle large-scale neural networks, then computational capacity increases, but chip area requirements increase

Engineering Contradiction:
Improvecomputational capacityVSAvoidchip area
Core Design Contradiction:
ProductivityVSArea of stationary object

Solution Approach 1:

The patent implements a hierarchical nested architecture where processing elements are organized in arrays within banks, which are further organized into larger processing systems. This nested structure allows efficient utilization of chip area by packing processing elements densely with minimal interconnection overhead, enabling high computational capacity without linearly increasing chip area requirements.

Inventive Principle:
Principle #7Nested doll (Nesting)

Solution Approach 2:

The patent utilizes two-dimensional arrays of processing elements with direct interconnections between neighboring elements, reducing the physical distance data must travel compared to linear or hierarchical interconnection schemes. This dimensional organization increases computational capacity while minimizing chip area by exploiting spatial efficiency in the array layout.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11256503B2Computational memory
Publication Date: 2022.02.22 AT-MEMORY COMPUTING LP
  • US11256503B2 patent drawing
  • US11256503B2 patent drawing
  • US11256503B2 patent drawing

AI summary

A processing device includes an array of processing elements, each processing element including an arithmetic logic unit to perform an operation. The processing device further includes interconnections among the array of processing elements to provide direct communication among neighboring processing elements of the array of processing elements. A processing element of the array of processing elements may be connected to a first neighbor processing element that is immediately adjacent the processing element. The processing element may be further connected to a second neighbor processing element that is immediately adjacent the first neighbor processing element. A processing element of the array of processing elements may be connected to a neighbor processing element via an input selector to selectively take output of the neighbor processing element as input to the processing element. A computing device may include such processing devices in an arrangement of banks.