Computational Memory with Sorting Network ALU

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The von Neumann bottleneck in traditional computer architecture limits computational speed due to congestion between processors and memory, as it separates semiconductors into logic and memory chips, leading to inefficient data movement and processing.

Innovation Solution

The development of a computational memory, referred to as Superstrider, which integrates computation and memory into a single component with an arithmetic and logic unit operating as a sorting network, allowing for efficient data sorting and merging within the memory bank, thereby alleviating the von Neumann bottleneck.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional von Neumann architecture separates processors and memory into distinct chips, then device complexity is reduced and manufacturing is easier, but computational speed and processing efficiency deteriorate due to data movement congestion

Engineering Contradiction:
Improvecomputational speedVSAvoidarchitecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent merges processors and memory into a single integrated device, combining computation and storage functions in one chip. This eliminates the physical separation between CPU and memory, allowing data to be processed in-place without being moved across chip boundaries, thereby resolving the von Neumann bottleneck while maintaining manageable device complexity through unified architecture design

Inventive Principle:
Principle #5Merging (Combining)

2Use of energy by moving object

If data is frequently moved between memory and processor in traditional architecture, then data access flexibility is maintained, but energy consumption increases due to repeated data transfer operations

Engineering Contradiction:
Improveenergy efficiencyVSAvoiddata access time
Core Design Contradiction:
Use of energy by moving objectVSLoss of time

Solution Approach 1:

The patent extracts the data movement step from the computation process by enabling in-memory processing. Data remains stationary in the memory array while processing operations are performed directly on the stored data, eliminating the energy-consuming data transfer operations between separate memory and processor components while maintaining fast access times

Inventive Principle:
Principle #2Taking out (Extraction)

3Speed

If computation and memory are integrated in a single chip, then data movement latency is reduced and processing speed increases, but manufacturing precision requirements and device complexity increase

Engineering Contradiction:
Improveprocessing speedVSAvoidfabrication precision
Core Design Contradiction:
SpeedVSManufacturing precision

Solution Approach 1:

The patent implements a universal memory cell design that can perform both storage and computation functions. The same memory infrastructure serves dual purposes: storing data and executing processing operations on that data. This multi-functionality reduces the need for separate dedicated processing circuits, thereby lowering manufacturing precision requirements despite the integrated architecture

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10936230B2Computational processor-in-memory with enhanced strided memory access
Publication Date: 2021.03.02 NATIONAL TECHNOLOGY & ENGINEERING SOLUTIONS OF SANDIA LLC
  • US10936230B2 patent drawing
  • US10936230B2 patent drawing
  • US10936230B2 patent drawing

AI summary

A computational memory for a computer. The memory includes a memory bank having a selected-row buffer and being configured to store records up to a number, K. The memory also includes an accumulator connected to the memory bank, the accumulator configured to store up to K records. The memory also includes an arithmetic and logic unit (ALU) connected to the accumulator and to the selected row buffer of the memory bank, the ALU having an indirect network of 2K ports for reading and writing records in the memory bank and the accumulator, and the ALU further physically configured to operate as a sorting network. The memory also includes a controller connected to the memory bank, the ALU, and the accumulator, the controller being hardware configured to direct operation of the ALU.