Loop Thread Execution Control in Self-Scheduling Computing Fabric
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computing systems face limitations in performance and energy efficiency for compute-intensive tasks such as Fast Fourier Transforms and finite impulse response filters, particularly in applications like artificial intelligence and 5G technologies, where dynamic reconfiguration and self-scheduling capabilities are needed.
Innovation Solution
A multi-threaded, coarse-grained configurable computing architecture with dynamic self-configuration and self-reconfiguration capabilities, including conditional branching, backpressure control, and ordered thread execution, utilizing a reenter queue and thread identifiers for advanced loop execution, and incorporating a configuration memory with instruction and instruction index memories to manage data path configurations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If existing computing systems are used for compute-intensive tasks, then basic computation can be performed, but performance and energy efficiency are insufficient
Solution Approach 1:
The patent implements dynamic reconfiguration of the computing fabric, allowing the system to adapt its architecture at runtime based on workload requirements. The reconfigurable logic elements can be dynamically programmed to optimize for specific compute-intensive tasks, improving both performance and energy efficiency by eliminating static architectural limitations
Solution Approach 2:
The computing fabric is divided into multiple independent logic elements that can be individually configured and orchestrated. This segmentation allows parallel execution of multiple computation threads across different logic elements, significantly improving computational throughput and energy efficiency for tasks like FFTs and FIR filters
2Productivity
If multi-threaded execution is implemented, then computational throughput is improved, but thread scheduling complexity increases
Solution Approach 1:
The system implements self-scheduling capability where the computing fabric autonomously manages thread execution without requiring external control. The logic elements can independently select and execute computation threads based on available resources and workload characteristics, reducing scheduling complexity while maintaining high throughput
Solution Approach 2:
The patent pre-configures multiple computation threads in the instruction memory before execution. This preliminary preparation allows the self-scheduling mechanism to efficiently select and execute pre-compiled threads without complex runtime compilation or scheduling decisions, simplifying the scheduling process
3Adaptability or versatility
If dynamic reconfiguration capability is added, then adaptability to different applications is improved, but system complexity increases
Solution Approach 1:
The patent implements a universal reconfiguration mechanism where a single control interface and instruction set can program all logic elements in the fabric. This universal approach allows the same system to be configured for different applications (FFT, FIR, machine learning, etc.) without requiring application-specific control logic, reducing overall system complexity while maintaining high adaptability
Data Source
AI summary
Representative apparatus, method, and system embodiments are disclosed for configurable computing. A representative system includes an interconnection network; a processor; and a plurality of configurable circuit clusters. Each configurable circuit cluster includes a plurality of configurable circuits arranged in an array; a synchronous network coupled to each configurable circuit of the array; and an asynchronous packet network coupled to each configurable circuit of the array. A representative configurable circuit includes a configurable computation circuit and a configuration memory having a first, instruction memory storing a plurality of data path configuration instructions to configure a data path of the configurable computation circuit; and a second, instruction and instruction index memory storing a plurality of spoke instructions and data path configuration instruction indices for selection of a master synchronous input, a current data path configuration instruction, and a next data path configuration instruction for a next configurable computation circuit.


