Processing-in-Memory MAC Arrays for Neural Network Data Transfer

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The separation of processor and memory in traditional hardware systems limits data communication, degrading the performance of artificial intelligence due to increased computational demands in neural networks, particularly in deep learning applications.

Innovation Solution

Integration of a processing-in-memory (PIM) device with MAC operators and memory banks, allowing for deterministic arithmetic operations within a semiconductor chip, facilitating direct data processing and reducing latency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If processor and memory are separated in traditional hardware systems, then device complexity is reduced and manufacturing is easier, but data communication between memory and processor is limited, degrading artificial intelligence performance

Engineering Contradiction:
Improvedata processing speedVSAvoiddevice complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent merges the processor and memory into a single integrated device, specifically combining MAC operators (processing units) with memory banks within the same semiconductor chip. This integration eliminates the need for separate processor and memory components, thereby improving data processing speed by enabling direct computation within the memory structure without external data transfer delays.

Inventive Principle:
Principle #5Merging (Combining)

2Productivity

If the number of layers in neural network is increased to improve artificial intelligence performance, then computational capability is enhanced, but the amount of computations required increases exponentially

Engineering Contradiction:
Improvecomputational capabilityVSAvoidcomputation time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent segments the computational function into distributed MAC operators that are embedded within multiple memory banks. Each memory bank contains one or more MAC operators that can perform computations locally on data stored in that bank. This segmentation allows parallel processing across multiple banks, significantly reducing the time required for exponential computations in deep neural networks.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimensional aspect to memory architecture by adding computational capabilities (MAC operators) to the traditional storage dimension. This transforms the memory structure from a passive storage medium into an active computing resource, enabling in-memory computation that reduces the time complexity of neural network operations.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If data communication between memory and processor is increased to support deep learning, then artificial intelligence performance is improved, but the limitation of data communication bandwidth becomes a bottleneck

Engineering Contradiction:
Improveartificial intelligence performanceVSAvoidenergy loss
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

By merging MAC operators with memory banks into an integrated structure, the patent eliminates the need for frequent data transfer between separate processor and memory components. Computations are performed directly within the memory banks using stored data, thereby reducing energy loss associated with data communication and improving AI performance.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12353986B2Processing-in-memory (PIM) device, controller for controlling the PIM device, and PIM system including the PIM device and the controller
Publication Date: 2025.07.08 SK HYNIX INC
  • US12353986B2 patent drawing
  • US12353986B2 patent drawing
  • US12353986B2 patent drawing

AI summary

A processing-in-memory (PIM) device includes a plurality of multiplication/accumulation (MAC) operators and a plurality of memory banks. The MAC operators are included in each of a plurality of channels. Each of the plurality of MAC operators performs a MAC arithmetic operation using weight data of a weight matrix. The memory banks are included in each of the plurality of channels and are configured to transmit the weight data of the weight matrix to the plurality of MAC operators. The weight data arrayed in one row of the weight matrix are stored into one row of each of the plurality of memory banks.