In-Memory Binary Convolution via Differential Memory Array

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The von-Neumann bottleneck in computing systems, where the data transfer rate between the CPU and memory has not kept pace with increasing processor speeds and memory density, leading to latency and inefficiency in processing, particularly in deep binary neural networks.

Innovation Solution

A circuit and method for in-memory binary convolution using a non-volatile memory structure, specifically a differential memory array with enhanced decoder and analog-to-digital converter, enabling binary neural network computations directly within the memory array by accumulating currents through bits, thereby overcoming the bottleneck.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If in-memory computing is implemented to overcome the von-Neumann bottleneck, then processing speed and latency are improved, but device complexity increases

Engineering Contradiction:
Improveprocessing speedVSAvoiddevice complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent merges memory storage and computation functions into a single integrated structure. The differential memory array simultaneously stores binary data and performs convolution operations, eliminating the separation between CPU and memory that causes the von-Neumann bottleneck. This integration allows data to be processed in-place without transfer overhead.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The memory array is designed to serve multiple functions: it acts as both storage medium and computational engine. The same physical structure that holds binary data also performs the convolution operation through controlled current accumulation, making the system multi-functional and reducing overall system complexity despite the added computational capability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Loss of time

If computations are performed within the memory array, then data transfer time is reduced, but manufacturing complexity increases

Engineering Contradiction:
Improvedata transfer timeVSAvoidmanufacturing ease
Core Design Contradiction:
Loss of timeVSEase of manufacture

Solution Approach 1:

The memory array performs computations using its own stored data and internal circuitry without requiring external computational resources. The convolution operation is executed by activating word lines and bit lines that already exist in the memory structure, allowing the memory to serve itself for both storage and processing functions.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the operational parameters of the memory array by applying specific voltages to word lines and bit lines to enable convolution operations. By controlling the activation state of memory cells through voltage parameters, the system transitions from pure storage mode to computational mode without requiring structural modifications.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If binary neural network operations are accelerated through in-memory computing, then productivity increases, but energy consumption increases

Engineering Contradiction:
Improvecomputation throughputVSAvoidenergy consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent replaces traditional electronic switching and sequential processing with direct physical current accumulation in the memory array. Instead of moving data through multiple processing stages, the system uses the physical property of current flow and accumulation to perform convolution operations, substituting mechanical/electronic processing with a more efficient physical phenomenon.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach reduces latency, enhances processing speeds, and provides an on-demand accelerator for deep binary neural networks by performing computations within the memory itself, improving performance without altering memory density or regular memory operations.

Implementation Method 1

performing a binary convolution of the first input word operand and the second input word operand in the differential memory array circuit by accumulating a summation of currents through a plurality of bits in the differential memory array circuit

Methodology Applied
Scientific EffectCurrent accumulation: Conduction (electrical)

Data Source

PatentUS10997498B2Apparatus and method for in-memory binary convolution for accelerating deep binary neural networks based on a non-volatile memory structure
Publication Date: 2021.05.04 GLOBALFOUNDRIES US INC
  • US10997498B2 patent drawing
  • US10997498B2 patent drawing
  • US10997498B2 patent drawing

AI summary

The present disclosure relates to a structure including a differential memory array circuit which is configured to perform a binary convolution of two input word operands by accumulating a summation of currents through a plurality of bits which are each arranged between a wordline and a sourceline in a horizontal direction and bitlines in a vertical direction.