Hardware Balancer Messaging for Scalable Compute-Near-Memory Links
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computing architectures face challenges with high latency and energy consumption due to significant data movement between processors and memory, constraining performance and capacity, particularly in systems with large numbers of compute elements.
Innovation Solution
Implementing compute-near-memory (CNM) systems with hybrid threading processors and custom compute fabrics, utilizing network structures for efficient communication between hardware compute elements, and employing a balancer element to manage communication requests, reducing the need for maintaining state data and minimizing network structure capacity constraints.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If data is moved between processors and memory in conventional architectures, then data access is enabled, but latency and energy consumption increase significantly
Solution Approach 1:
The patent merges processing elements and memory elements into integrated compute-near-memory units, where processing logic is embedded directly within or adjacent to memory structures. This eliminates the need for separate data movement between distinct processor and memory components, thereby reducing latency and improving data access speed simultaneously.
Solution Approach 2:
The patent introduces local buffers and intermediate storage structures within the memory device that act as mediators between the processing elements and main memory. These intermediaries enable data to be processed locally without requiring frequent trips to external memory, reducing data movement latency while maintaining efficient data access.
2Productivity
If more compute elements are added to increase processing capacity, then computational power increases, but network structure capacity constraints and deadlocks occur
Solution Approach 1:
The patent segments the computing system into multiple independent compute-near-memory units that can operate autonomously. Each unit has its own integrated processing and memory resources, reducing the need for complex inter-unit communication and minimizing network structure complexity even as the number of compute elements scales.
Solution Approach 2:
The patent transitions from a centralized processing architecture to a distributed, multi-dimensional array of compute-near-memory units. This spatial reorganization allows data and processing to occur in parallel across multiple dimensions, increasing computational capacity while reducing the burden on any single network path.
3Reliability
If state data is maintained for communication requests, then request tracking is enabled, but network structure capacity is constrained
Solution Approach 1:
The patent implements self-service mechanisms where compute-near-memory units autonomously manage their own communication requests and state tracking. Each unit maintains minimal local state information and uses event-driven protocols to coordinate with others, eliminating the need for centralized state management and freeing up network structure capacity.
Solution Approach 2:
The patent employs event-driven communication where state information is temporarily maintained only when needed for specific operations and then discarded once the operation completes. This on-demand state management reduces the overall volume of stationary data required in the network structure while maintaining reliable request tracking.
Data Source
AI summary
Various examples are directed to an arrangement comprising a first hardware compute element and a hardware balancer element. The first hardware compute element may send a first request message to a hardware balancer element. The first request message may describe a processing task. The hardware balancer element may send a second request message towards a second hardware compute element for executing the processing task and send to the first compute element a first reply message in reply to the first request message. After sending the first reply message, the hardware balancer element may receive a first completion request message indicating that the processing task is assigned and send, to the first hardware computing element, a second completion request message, the second completion request message indicating that the processing task is assigned.


