Execution Mask Generator for Multi-threaded Reconfigurable Fabric
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computing systems face limitations in computation processing speed, energy efficiency, and heat dissipation, particularly in handling compute-intensive tasks like Fast Fourier Transforms and finite impulse response filters, which are essential for applications such as artificial intelligence, 5G base stations, and machine learning, requiring a high-performance and energy-efficient computing architecture capable of dynamic self-configuration.
Innovation Solution
A multi-threaded, coarse-grained configurable computing architecture with dynamic self-scheduling and self-reconfiguration capabilities, featuring a configurable computation circuit, synchronous network inputs and outputs, and a configuration memory for data path configuration, enabling conditional branching, backpressure control, and efficient loop execution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If existing computing systems are used for compute-intensive tasks, then basic computation processing is achieved, but computation processing speed and energy efficiency are limited
Solution Approach 1:
The patent implements dynamic reconfiguration of the computing fabric, allowing the system to adapt its architecture runtime based on computational workload characteristics. Configuration memory and control circuits enable runtime modification of data paths, compute element connectivity, and resource allocation, transforming the static architecture into a dynamic one that optimizes performance and energy efficiency for different computational patterns
Solution Approach 2:
The computing fabric is divided into multiple independent compute elements (tiles) that can be individually configured and activated. Each tile contains configurable logic units, arithmetic units, and local memory, allowing the system to segment computational tasks across multiple independent processing units, thereby improving throughput and energy efficiency by activating only the necessary segments for each task
2Productivity
If computing systems are designed for high performance, then computation speed improves, but heat dissipation increases
Solution Approach 1:
The patent implements fine-grained control over which compute elements and data paths are active at any given time. Configuration memory stores runtime configuration data that enables selective activation of specific tiles, logic units, and arithmetic units based on the computational requirements of the current task, ensuring that only necessary components consume power and generate heat
Solution Approach 2:
The control circuitry automatically manages the activation and configuration of compute elements based on workload characteristics, eliminating the need for manual intervention. The system self-adjusts its operational state to optimize performance while minimizing energy consumption and heat generation by activating only the minimum necessary computational resources
3Productivity
If computing systems are designed for specialized applications, then performance for specific tasks improves, but adaptability to other applications deteriorates
Solution Approach 1:
The patent implements a universal computing fabric where each compute element can be dynamically reconfigured to perform different computational functions. The configurable logic units, arithmetic units, and interconnection networks can be programmed via configuration memory to implement various algorithms and computational patterns, allowing the same hardware to efficiently execute diverse applications from neural network inference to signal processing to graph analytics
Solution Approach 2:
The system employs runtime reconfiguration capabilities that allow the architecture to dynamically adapt to different computational workloads. Configuration data stored in configuration memory enables the system to transform its data paths, activate different compute elements, and modify interconnections based on the specific requirements of each application, providing both specialization for current tasks and adaptability for future workloads
Data Source
AI summary
Representative apparatus, method, and system embodiments are disclosed for configurable computing. A representative system includes an asynchronous packet network having a plurality of data transmission lines forming a data path transmitting operand data; a synchronous mesh communication network; a plurality of configurable circuits arranged in an array, each configurable circuit of the plurality of configurable circuits coupled to the asynchronous packet network and to the synchronous mesh communication network, each configurable circuit of the plurality of configurable circuits adapted to perform a plurality of computations; each configurable circuit of the plurality of configurable circuits comprising: a memory storing operand data; and an execution or write mask generator adapted to generate an execution mask or a write mask identifying valid bits or bytes transmitted on the data path or stored in the memory for a current or next computation.


