Execution Mask Generator for Multi-threaded Reconfigurable Fabric

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing computing systems face limitations in computation processing speed, energy efficiency, and heat dissipation, particularly in handling compute-intensive tasks like Fast Fourier Transforms and finite impulse response filters, which are essential for applications such as artificial intelligence, 5G base stations, and machine learning, requiring a high-performance and energy-efficient computing architecture capable of dynamic self-configuration.

Innovation Solution

A multi-threaded, coarse-grained configurable computing architecture with dynamic self-scheduling and self-reconfiguration capabilities, featuring a configurable computation circuit, synchronous network inputs and outputs, and a configuration memory for data path configuration, enabling conditional branching, backpressure control, and efficient loop execution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If existing computing systems are used for compute-intensive tasks, then basic computation processing is achieved, but computation processing speed and energy efficiency are limited

Engineering Contradiction:
Improvecomputation processing speedVSAvoidenergy efficiency
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent implements dynamic reconfiguration of the computing fabric, allowing the system to adapt its architecture runtime based on computational workload characteristics. Configuration memory and control circuits enable runtime modification of data paths, compute element connectivity, and resource allocation, transforming the static architecture into a dynamic one that optimizes performance and energy efficiency for different computational patterns

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The computing fabric is divided into multiple independent compute elements (tiles) that can be individually configured and activated. Each tile contains configurable logic units, arithmetic units, and local memory, allowing the system to segment computational tasks across multiple independent processing units, thereby improving throughput and energy efficiency by activating only the necessary segments for each task

Inventive Principle:
Principle #1Segmentation

2Productivity

If computing systems are designed for high performance, then computation speed improves, but heat dissipation increases

Engineering Contradiction:
Improvecomputation speedVSAvoidheat dissipation
Core Design Contradiction:
ProductivityVSTemperature

Solution Approach 1:

The patent implements fine-grained control over which compute elements and data paths are active at any given time. Configuration memory stores runtime configuration data that enables selective activation of specific tiles, logic units, and arithmetic units based on the computational requirements of the current task, ensuring that only necessary components consume power and generate heat

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The control circuitry automatically manages the activation and configuration of compute elements based on workload characteristics, eliminating the need for manual intervention. The system self-adjusts its operational state to optimize performance while minimizing energy consumption and heat generation by activating only the minimum necessary computational resources

Inventive Principle:
Principle #25Self-service

3Productivity

If computing systems are designed for specialized applications, then performance for specific tasks improves, but adaptability to other applications deteriorates

Engineering Contradiction:
Improveperformance for specific tasksVSAvoidadaptability to various applications
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal computing fabric where each compute element can be dynamically reconfigured to perform different computational functions. The configurable logic units, arithmetic units, and interconnection networks can be programmed via configuration memory to implement various algorithms and computational patterns, allowing the same hardware to efficiently execute diverse applications from neural network inference to signal processing to graph analytics

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system employs runtime reconfiguration capabilities that allow the architecture to dynamically adapt to different computational workloads. Configuration data stored in configuration memory enables the system to transform its data paths, activate different compute elements, and modify interconnections based on the specific requirements of each application, providing both specialization for current tasks and adaptability for future workloads

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12106099B2Execution or write mask generation for data selection in a multi-threaded, self-scheduling reconfigurable computing fabric
Publication Date: 2024.10.01 MICRON TECHNOLOGY INC
  • US12106099B2 patent drawing
  • US12106099B2 patent drawing
  • US12106099B2 patent drawing

AI summary

Representative apparatus, method, and system embodiments are disclosed for configurable computing. A representative system includes an asynchronous packet network having a plurality of data transmission lines forming a data path transmitting operand data; a synchronous mesh communication network; a plurality of configurable circuits arranged in an array, each configurable circuit of the plurality of configurable circuits coupled to the asynchronous packet network and to the synchronous mesh communication network, each configurable circuit of the plurality of configurable circuits adapted to perform a plurality of computations; each configurable circuit of the plurality of configurable circuits comprising: a memory storing operand data; and an execution or write mask generator adapted to generate an execution mask or a write mask identifying valid bits or bytes transmitted on the data path or stored in the memory for a current or next computation.