Reconfigurable Multi-Processing Array for Thread Parallelism
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing coarse-grained reconfigurable architectures (CGRAs) face challenges in programming complexity, limiting their widespread adoption despite offering high parallelism and efficiency, as they require innovative architectures and design methodologies to meet the demands of modern embedded systems needing high performance, low energy consumption, flexibility, and rapid design-to-market capabilities.
Innovation Solution
A signal processing device with multiple functional units capable of word- or subword-level operations and dynamically switchable interconnect arrangements, allowing for simultaneous processing of multiple threads and flexible configuration between single and multi-threading modes, supported by control modules and multiplexing/demultiplexing circuits that can change settings per clock cycle.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If coarse-grained reconfigurable architectures are used to achieve high parallelism and efficiency, then processing performance is improved, but programming complexity increases
Solution Approach 1:
The functional units are divided into multiple processing units (PUs), each with its own control module. This segmentation allows the system to handle multiple threads simultaneously while keeping each control module relatively simple, resolving the contradiction between high parallelism and programming complexity.
Solution Approach 2:
The interconnect arrangements are made dynamically switchable between single-threading and multi-threading modes. This dynamic reconfigurability allows the system to adapt its complexity level based on the application requirements, achieving high performance when needed while simplifying programming for less demanding tasks.
2Adaptability or versatility
If multiple interconnect arrangements are provided for dynamic switching between single and multi-threading modes, then flexibility is improved, but device complexity increases
Solution Approach 1:
The routing resources are designed to support multiple interconnect arrangements that can be dynamically switched. The same physical routing infrastructure serves both single-threading and multi-threading modes, achieving versatility without proportionally increasing complexity.
Solution Approach 2:
Multiple interconnect arrangements are merged into a unified routing resource structure. By combining the single-threading and multi-threading interconnect paths into a shared infrastructure with dynamic switching capability, the system achieves high flexibility while avoiding the complexity of completely separate interconnect systems.
3Productivity
If functional units are grouped into processing units with predetermined topology, then processing efficiency is improved, but reconfigurability decreases
Solution Approach 1:
The processing unit topology is dynamically reconfigurable between predetermined groupings for efficient processing and other configurations for adaptability. This dynamic switching capability allows the system to optimize for processing efficiency when needed while maintaining reconfigurability for different application requirements.
Solution Approach 2:
Processing units are pre-configured with predetermined topologies that are optimized for specific processing tasks. These preliminary configurations enable high processing efficiency for common workloads, while the ability to dynamically reconfigure allows adaptation to less common requirements.
Data Source
AI summary
A signal processing device is adapted for simultaneous processing of at least two process threads in a multi-processing manner. The device comprises a plurality of functional units capable of executing word- or subword-level operations on data. The device further comprises means for interconnecting the plurality of functional units, the means for interconnecting supporting a plurality of dynamically switchable interconnect arrangements, and at least one of the interconnect arrangements interconnects the plurality of functional units into at least two non-overlapping processing units each with a pre-determined topology. The device further comprises at least two control modules each assigned to one of the processing units.


