DSP Recirculation Path for Low Latency Compute
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current digital signal processor (DSP) architectures face challenges in reducing latency and meeting the performance requirements of high-intensity applications, such as base stations in wireless systems, due to limitations in computation speed and power consumption, often necessitating the use of expensive and inflexible combinations with ASICs and FPGAs.
Innovation Solution
A DSP architecture is designed with a compute array featuring a serial connection of compute engines, allowing data and instructions to recirculate directly between the initial and final compute engines within a single clock cycle, reducing latency and optimizing performance through a recirculation path that bypasses intermediate engines.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data and instructions circulate through all compute engines in a DSP, then all compute engines can process the data, but latency increases due to the time required for data to traverse each engine
Solution Approach 1:
The compute array is segmented into multiple compute engines (e.g., eight compute engines) that can independently process data. This segmentation allows the system to balance between throughput (all engines working) and latency (direct recirculation path bypassing intermediate engines).
Solution Approach 2:
A recirculation path acts as an intermediary mechanism that enables data to return from the final compute engine to the initial compute engine. This recirculation path can operate in two modes: through all compute engines for high throughput, or directly for low latency, allowing the system to switch between operational modes as needed.
2Speed
If DSP performance is increased to meet high-intensity application requirements, then computation speed improves, but power dissipation increases
Solution Approach 1:
The DSP architecture implements dynamic operation modes where compute engines can be selectively activated or deactivated based on application requirements. This allows the system to adjust computation speed and power consumption dynamically, enabling high-performance operations when needed while conserving power during normal operations.
Solution Approach 2:
The system can change operational parameters such as the number of active compute engines, clock frequency, and data flow patterns to optimize the balance between computation speed and power dissipation. The recirculation mechanism allows flexible parameter adjustment without hardware changes.
3Productivity
If multiple computation blocks are used to increase DSP performance, then computation capability improves, but latency increases as instructions circulate among computation blocks
Solution Approach 1:
The computation blocks are segmented into discrete compute engines within a compute array, each capable of independent operation. This segmentation enables selective activation of engines based on task requirements, reducing unnecessary circulation latency while maintaining high computation capability through parallel processing of different data streams.
Solution Approach 2:
The recirculation path enables continuous data flow through the compute engines without interruption or idle circulation. Data can continuously recirculate from the final engine back to the initial engine, maintaining productive computation across all engines while avoiding wasted time in non-productive data movement.
Data Source
AI summary
A processor includes a compute array comprising a first plurality of compute engines serially connected along a data flow path such that data flows between successive compute engines at successive times. The first plurality of compute engines includes an initial compute engine and a final compute engine. The data flow path includes a recirculation path connecting the final compute engine to the initial compute engine with no compute engine therebetween.


