PIM Accelerator Register Control for In-Memory AI Operations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing accelerators, such as GPUs and NPUs, face inefficiencies in performing predetermined tasks like AI and machine learning operations due to differences in instruction set architectures and data access limitations, particularly in processing-in-memory (PIM) operations.

Innovation Solution

An electronic device with a PIM operation register and PIM processing elements (PEs) that load operation type information, control PIM operations in memory banks, and manage data access to optimize PIM operations, including parallel and serial execution based on operation characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If data is stored in separate memory from processor (traditional architecture), then memory capacity and processor functionality are improved, but data access time and processing efficiency deteriorate

Engineering Contradiction:
Improveprocessing efficiencyVSAvoiddata access time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent merges the memory storage function and processing function into a single integrated structure. PIM processing elements are embedded within memory banks, allowing data to be processed in-place without being transferred to separate processors. This eliminates the time loss associated with data transfer between memory and processor while maintaining both storage capacity and processing capability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The memory is divided into multiple memory banks, each containing PIM processing elements. This segmentation allows different portions of data to be processed simultaneously in different banks, improving overall processing efficiency while reducing access time through parallel operation.

Inventive Principle:
Principle #1Segmentation

2Productivity

If general-purpose processors are used for AI and machine learning tasks, then device versatility is improved, but processing speed and efficiency for specific tasks deteriorate

Engineering Contradiction:
Improveprocessing speedVSAvoidinstruction set compatibility
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The PIM processing elements are specifically optimized for AI and machine learning operations, with specialized hardware components like multiply-accumulate units tailored for neural network computations. This local optimization delivers high processing speed for specific tasks while the host processor maintains overall system versatility.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent introduces an intermediary layer between the host processor and memory that handles PIM operation execution. This intermediary translates host processor instructions into PIM-specific operations, enabling efficient AI/ML processing while maintaining compatibility with the host processor's instruction set architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If data is rearranged in memory according to PIM operation characteristics, then PIM operation efficiency is improved, but data access complexity and overhead increase

Engineering Contradiction:
ImprovePIM operation efficiencyVSAvoiddata arrangement complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

Data is rearranged into PIM-optimized formats during initial loading into memory banks, before PIM operations are executed. This preliminary arrangement eliminates the need for complex data reorganization during PIM operation execution, reducing overhead while maintaining high efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The PIM processing elements automatically detect and adapt to the data arrangement format stored in memory banks. This self-service capability reduces the complexity of data management by eliminating the need for complex control logic to manage data reorganization, as the PIM elements handle format adaptation autonomously.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12554632B2Electronic device and method with hardware acceleration
Publication Date: 2026.02.17 SAMSUNG ELECTRONICS CO LTD
  • US12554632B2 patent drawing
  • US12554632B2 patent drawing
  • US12554632B2 patent drawing

AI summary

An electronic device and method are provided. The method includes, in response to a result of a performed determination, between a processing-in-memory (PIM) operation and a non-PIM operation, being that a target operation included in an instruction set is the PIM operation, loading information regarding a PIM operation type of the target operation to a PIM operation register of an accelerator of the electronic device based on a performed determination that input data, for use in the target operation, is arranged in a memory of the accelerator according to a characteristic of the target operation, and controlling, using a PIM operation execution instruction, one or more PIM processing elements (PEs) of the memory to perform the target operation using the arranged input data and the information regarding the operation type loaded in the PIM operation register.