Content-Addressable Processing Engine for PIM Speedup

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing processing-in-memory (PIM) architectures face challenges in overcoming the von Neumann bottleneck and require expensive memory technologies or restrictive programming languages to leverage content-addressable memories for parallel processing.

Innovation Solution

A CMOS-based content-addressable processing engine (CAPE) that integrates computation and storage, utilizing dense 6T SRAM arrays and a general-purpose microarchitecture to perform associative computing with standard RISC-V instructions, enabling efficient vector operations and high programmability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If content-addressable memories are used for parallel processing, then processing speed is improved, but manufacturing cost increases due to expensive memory technology requirements

Engineering Contradiction:
Improveprocessing speedVSAvoidmanufacturing cost
Core Design Contradiction:
ProductivityVSEase of manufacture

Solution Approach 1:

The patent changes the fundamental parameter of memory technology from emerging/expensive types to standard CMOS SRAM, enabling content-addressable parallel processing at conventional manufacturing costs while maintaining high processing speeds through architectural optimization

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The invention replaces expensive, specialized memory technology with inexpensive, widely-manufacturable CMOS SRAM cells, making the system economically viable for mass production while achieving comparable or superior performance through clever use of standard components

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

2Productivity

If content-addressable memories are used for parallel processing, then processing speed is improved, but programming complexity increases due to restrictive programming language requirements

Engineering Contradiction:
Improveprocessing speedVSAvoidprogramming complexity
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent makes the content-addressable processing engine universally programmable using standard RISC-V instruction sets, allowing the same hardware to handle both traditional sequential operations and parallel vector operations without requiring specialized programming languages or compilers

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

Instead of requiring specialized programming languages to access parallel processing capabilities, the invention inverts the approach by making standard RISC-V instructions capable of triggering parallel execution, thus simplifying the programming model while maintaining high productivity

Inventive Principle:
Principle #13The other way round (Inversion)

3Productivity

If computation and storage logic are combined into a single component, then data access efficiency is improved, but device complexity increases

Engineering Contradiction:
Improvedata access efficiencyVSAvoiddevice complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the processing engine into distinct functional units (control logic, compute units, storage arrays) that can operate independently yet cooperatively, managing complexity through modular design while achieving efficient data access by keeping computation and storage in close proximity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention merges computation and storage logic into a unified content-addressable processing engine where SRAM arrays serve both as storage and as the basis for computational operations, eliminating the von Neumann bottleneck while managing complexity through shared hardware resources

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12001841B2Content-addressable processing engine
Publication Date: 2024.06.04 CORNELL UNIVERSITY
  • US12001841B2 patent drawing
  • US12001841B2 patent drawing
  • US12001841B2 patent drawing

AI summary

A content-addressable processing engine, also referred to herein as CAPE, is provided. Processing-in-memory (PIM) architectures attempt to overcome the von Neumann bottleneck by combining computation and storage logic into a single component. CAPE provides a general-purpose PIM microarchitecture that provides acceleration of vector operations while being programmable with standard reduced instruction set computing (RISC) instructions, such as RISC-V instructions with standard vector extensions. CAPE can be implemented as a standalone core that specializes in associative computing, and that can be integrated in a tiled multicore chip alongside other types of compute engines. Certain embodiments of CAPE achieve average speedups of 14× (up to 254×) over an area-equivalent out-of-order processor core tile with three levels of caches across a diverse set of representative applications.