PIM Controller Queue Scheduling for Deterministic AI Memory Computing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Processing-in-memory (PIM) systems face challenges in efficiently managing data communication between separate processor and memory units, leading to degraded performance in artificial intelligence applications, particularly in deep learning tasks due to the exponential increase in computation requirements and limitations in data communication.

Innovation Solution

A PIM controller is configured with read/arithmetic queue logic, write queue logic, and scheduling logic circuits to manage and prioritize queues, enabling efficient data processing and arithmetic operations within a PIM device that integrates processor and memory functions, optimizing the output sequence and enabling deterministic arithmetic operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If a general hardware system with separate memory and processor is used, then the system structure is simple and easy to manufacture, but the performance of artificial intelligence is degraded due to limitation of data communication between memory and processor

Engineering Contradiction:
Improvesystem structureVSAvoidartificial intelligence performance
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent merges the processor and memory into a single integrated PIM device, where the processor core is directly coupled to the memory array. This integration eliminates the communication bottleneck between separate memory and processor units, enabling high-speed data access for AI computations while maintaining manufacturing feasibility through standardized semiconductor fabrication processes.

Inventive Principle:
Principle #5Merging (Combining)

2Productivity

If the number of layers in neural network is increased to improve AI performance, then the computation capability is enhanced, but the amount of computation required increases exponentially

Engineering Contradiction:
ImproveAI performanceVSAvoidcomputation requirement
Core Design Contradiction:
ProductivityVSPower

Solution Approach 1:

The PIM device performs preliminary data processing and arithmetic operations directly within the memory array before data needs to be transferred to external processors. By pre-computing intermediate results and maintaining data in high-speed memory during computation, the system reduces the exponential growth of computation requirements for deep neural networks with increased layers.

Inventive Principle:
Principle #10Preliminary action

3Speed

If PIM device directly performs arithmetic operations internally, then data processing speed is improved, but the complexity of queue management and scheduling increases

Engineering Contradiction:
Improvedata processing speedVSAvoidqueue management complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent segments the queue management system into distinct functional units: a command queue for storing arithmetic operation instructions, a data queue for buffering operands, and a result queue for storing computation outputs. The scheduling logic is divided into separate modules that independently manage each queue type, reducing overall system complexity while enabling high-speed parallel arithmetic operations within the PIM device.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11720354B2Processing-in-memory (PIM) system and operating methods of the PIM system
Publication Date: 2023.08.08 SK HYNIX INC
  • US11720354B2 patent drawing
  • US11720354B2 patent drawing
  • US11720354B2 patent drawing

AI summary

A processing-in-memory (PIM) controller includes a read/arithmetic queue logic circuit, a write queue logic circuit, and a scheduling logic circuit. The read/arithmetic queue logic circuit stores read queues and arithmetic queues, generates an arithmetic mode signal when an arithmetic queue exists in the read/arithmetic queue logic circuit, and outputs the arithmetic queue in response to an arithmetic mode enablement signal. The write queue logic circuit stores write queues, generates an arithmetic write signal when an arithmetic write queue exists in the write queue logic circuit, and outputs the write queue in response to an arithmetic write enablement signal. The scheduling logic circuit transmits the arithmetic mode enablement signal to the read/arithmetic queue logic circuit in response to the arithmetic mode signal and transmits the arithmetic write enablement signal to the write queue logic circuit in response to the arithmetic write signal.