ECC-Enabled PIM MAC Architecture for AI Latency Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Processing-in-memory (PIM) devices face challenges in efficiently performing multiplication and accumulation operations due to limitations in data communication between memory and processor, which hinders the performance of artificial intelligence applications, especially in deep learning tasks where increased computation demands are exponential.

Innovation Solution

A PIM device is designed with an error correction code (ECC) logic circuit and a multiplication and accumulation (MAC) operator that generates and processes write and read data, parity, and converted data to perform MAC operations directly within the memory, enhancing data processing speed and reducing errors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If a general hardware system with separate memory and processor is used, then device complexity is reduced, but data processing speed and AI performance are degraded due to limitations in data communication between memory and processor

Engineering Contradiction:
Improvedata processing speedVSAvoiddevice complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent merges the processor and memory into a single integrated PIM device, allowing arithmetic operations to be performed directly within the memory structure. This integration eliminates the need for data communication between separate memory and processor components, thereby improving data processing speed while accepting increased device complexity as a necessary trade-off for achieving high-performance AI computing

Inventive Principle:
Principle #5Merging (Combining)

2Productivity

If the number of layers in neural network is increased to improve AI performance, then computational capability is enhanced, but the amount of computation required increases exponentially

Engineering Contradiction:
Improvecomputational capabilityVSAvoidcomputation time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent implements preliminary action by performing multiplication and accumulation operations directly within the memory structure before data needs to be processed further. The PIM device pre-computes intermediate results using integrated MAC operators, reducing the amount of data that needs to be transferred and processed subsequently, thereby decreasing overall computation time for deep neural networks with increased layers

Inventive Principle:
Principle #10Preliminary action

3Ease of manufacture

If separate memory and processor systems are used, then ease of manufacture is improved, but performance of artificial intelligence is degraded due to limitation of data communication

Engineering Contradiction:
Improveease of manufactureVSAvoidAI performance
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent combines memory and processing functions into a single PIM device that can be manufactured as an integrated semiconductor component. While this integration increases manufacturing complexity compared to separate components, the patent addresses this by using standard semiconductor fabrication processes and modular design approaches, achieving a balance between ease of manufacture and high AI performance through direct in-memory computation

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11996157B2Processing-in-memory (PIM) devices
Publication Date: 2024.05.28 SK HYNIX INC
  • US11996157B2 patent drawing
  • US11996157B2 patent drawing
  • US11996157B2 patent drawing

AI summary

A processing-in-memory (PIM) device includes an ECC logic circuit configured to generate first write data, first write parity, second write data, and second write parity from first write input data and second write input data when a write operation in an operation mode is performed, and generate first converted data and second converted data from first read data, first read parity, second read data, and second read parity when a read operation in the operation mode is performed; and a MAC operator configured to perform a MAC arithmetic operation for the first converted data and the second converted data to generate MAC operation result data.