Hardware Balancer Messaging for Scalable Compute-Near-Memory Links

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing computing architectures face challenges with high latency and energy consumption due to significant data movement between processors and memory, constraining performance and capacity, particularly in systems with large numbers of compute elements.

Innovation Solution

Implementing compute-near-memory (CNM) systems with hybrid threading processors and custom compute fabrics, utilizing network structures for efficient communication between hardware compute elements, and employing a balancer element to manage communication requests, reducing the need for maintaining state data and minimizing network structure capacity constraints.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If data is moved between processors and memory in conventional architectures, then data access is enabled, but latency and energy consumption increase significantly

Engineering Contradiction:
Improvedata access speedVSAvoiddata movement latency
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The patent merges processing elements and memory elements into integrated compute-near-memory units, where processing logic is embedded directly within or adjacent to memory structures. This eliminates the need for separate data movement between distinct processor and memory components, thereby reducing latency and improving data access speed simultaneously.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces local buffers and intermediate storage structures within the memory device that act as mediators between the processing elements and main memory. These intermediaries enable data to be processed locally without requiring frequent trips to external memory, reducing data movement latency while maintaining efficient data access.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If more compute elements are added to increase processing capacity, then computational power increases, but network structure capacity constraints and deadlocks occur

Engineering Contradiction:
Improvecomputational capacityVSAvoidnetwork structure complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the computing system into multiple independent compute-near-memory units that can operate autonomously. Each unit has its own integrated processing and memory resources, reducing the need for complex inter-unit communication and minimizing network structure complexity even as the number of compute elements scales.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from a centralized processing architecture to a distributed, multi-dimensional array of compute-near-memory units. This spatial reorganization allows data and processing to occur in parallel across multiple dimensions, increasing computational capacity while reducing the burden on any single network path.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Reliability

If state data is maintained for communication requests, then request tracking is enabled, but network structure capacity is constrained

Engineering Contradiction:
Improverequest tracking accuracyVSAvoidnetwork structure capacity
Core Design Contradiction:
ReliabilityVSVolume of stationary object

Solution Approach 1:

The patent implements self-service mechanisms where compute-near-memory units autonomously manage their own communication requests and state tracking. Each unit maintains minimal local state information and uses event-driven protocols to coordinate with others, eliminating the need for centralized state management and freeing up network structure capacity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent employs event-driven communication where state information is temporarily maintained only when needed for specific operations and then discarded once the operation completes. This on-demand state management reduces the overall volume of stationary data required in the network structure while maintaining reliable request tracking.

Inventive Principle:
Principle #34Discarding and recovering

Data Source

PatentUS20260072752A1Methods and systems for communications between hardware components
Publication Date: 2026.03.12 MICRON TECHNOLOGY INC
  • US20260072752A1 patent drawing
  • US20260072752A1 patent drawing
  • US20260072752A1 patent drawing

AI summary

Various examples are directed to an arrangement comprising a first hardware compute element and a hardware balancer element. The first hardware compute element may send a first request message to a hardware balancer element. The first request message may describe a processing task. The hardware balancer element may send a second request message towards a second hardware compute element for executing the processing task and send to the first compute element a first reply message in reply to the first request message. After sending the first reply message, the hardware balancer element may receive a first completion request message indicating that the processing task is assigned and send, to the first hardware computing element, a second completion request message, the second completion request message indicating that the processing task is assigned.