Reconfigurable Compute Fabric Loop Data I/O

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer architectures face limitations in performance and efficiency due to the time and energy required for data movement between processors and memory, constraining compute systems beyond the capabilities of transistor scaling.

Innovation Solution

The implementation of memory-centric compute topologies with compute-near-memory (CNM) systems, which integrate processors with memory or data storage components, utilizing hybrid threading processors and fabrics to facilitate low-latency operations through custom compute fabrics and reconfigurable compute fabrics that support synchronous flows and nested loops.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If data is moved between processors and memory using conventional architectures, then data access is possible, but time and energy consumption increase significantly

Engineering Contradiction:
Improvedata access timeVSAvoidenergy consumption
Core Design Contradiction:
Loss of timeVSUse of energy by moving object

Solution Approach 1:

The patent merges memory and compute resources by integrating compute units directly within the memory device, creating a memory-compute device where processing elements are coupled to memory arrays. This integration eliminates the need for data to be moved between separate processor and memory components, thereby reducing both access time and energy consumption associated with data transfer operations.

Inventive Principle:
Principle #5Merging (Combining)

2Productivity

If compute resources are integrated with memory, then compute efficiency improves, but device complexity increases

Engineering Contradiction:
Improvecompute efficiencyVSAvoiddevice complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the memory-compute device into distinct functional components including memory arrays, compute units, and interconnect structures. Each segment performs a specific function, allowing for modular design and independent optimization. This segmentation manages device complexity by organizing complex functionality into manageable, specialized units that can be independently designed and controlled.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements reconfigurable compute fabrics that can dynamically change their interconnect topology and data flow paths based on the computational task being performed. This dynamic reconfiguration allows the same hardware structure to adapt to different algorithms and workloads, improving compute efficiency while managing complexity through software-controlled flexibility rather than hardwired dedicated structures.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If reconfigurable compute fabrics are used, then adaptability to different workloads improves, but control complexity increases

Engineering Contradiction:
Improveworkload adaptabilityVSAvoidcontrol complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal compute fabric architecture that can perform multiple functions through a single reconfigurable structure. The same interconnect network and compute units can be configured to execute different types of operations (e.g., neural network inference, matrix operations, data processing) by changing configuration parameters rather than requiring separate dedicated hardware for each workload type. This universality improves adaptability while controlling complexity by eliminating the need for multiple specialized subsystems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11709796B2Data input/output operations during loop execution in a reconfigurable compute fabric
Publication Date: 2023.07.25 MICRON TECHNOLOGY INC
  • US11709796B2 patent drawing
  • US11709796B2 patent drawing
  • US11709796B2 patent drawing

AI summary

Various examples are directed to systems and methods in which a first flow controller of a first synchronous flow may receive an instruction to execute a first loop using the first synchronous flow. The first flow controller may determine a first iteration index for a first iteration of the first loop. The first flow controller may send, to a first compute element of the first synchronous flow, a first synchronous message to initiate a first synchronous flow thread for executing the first iteration of the first loop. The first synchronous message may comprise the iteration index. The first compute element may execute an input/output operation at a first location of a first compute element memory indicated by the first iteration index.