Programmable DSP Accelerator With Coherent Cache for Memory Stalls
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing accelerator architectures face inefficiencies in memory access, leading to stalled task execution and missed instructions due to high data volumes, particularly in digital signal processors (DSPs).
Innovation Solution
A configurable DSP with embedded Arithmetic Logic Unit (ALU), data cache, and alignment units is introduced, integrated into a vector Single Instruction, Multiple Data (SIMD) architecture, which includes a hardware cache coherent memory to enhance data movement and reduce external memory access attempts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is accessed from external memory in high-volume processing, then data volume handling capability is improved, but memory access efficiency deteriorates causing stalls
Solution Approach 1:
The patent divides the memory system into multiple segments: external memory for bulk storage and on-chip block memory for frequent access. The DSP block is segmented into functional units (multipliers, adders, ALU) that can be independently configured. This segmentation allows high-volume data to be stored externally while maintaining efficient access to frequently used data through local block memory.
Solution Approach 2:
The patent introduces block memory as an intermediary between external memory and the DSP computational units. This intermediary buffer reduces the number of external memory access attempts by keeping frequently accessed data locally, thereby preventing stalls while still enabling high-volume data processing.
2Productivity
If multiple DSPs operate in parallel, then processing throughput is improved, but data movement complexity increases
Solution Approach 1:
The patent merges multiple DSP blocks into a unified accelerator structure with shared block memory and interconnect resources. This merging allows parallel operation of multiple DSPs while reducing overall data movement complexity through shared infrastructure rather than requiring separate memory systems for each DSP.
Solution Approach 2:
The block memory and interconnect structures are designed as universal resources that can serve multiple DSP blocks simultaneously. This multi-functionality enables efficient data movement in parallel architectures without requiring dedicated data paths for each DSP, thereby reducing complexity while maintaining high throughput.
3Speed
If dedicated hardware accelerators are used, then computational speed is improved, but flexibility and reconfigurability deteriorate
Solution Approach 1:
The patent implements dynamically reconfigurable DSP blocks where the internal architecture (number of multipliers, adders, ALU configurations) can be changed at runtime. This dynamic reconfigurability allows the same hardware to adapt to different computational tasks while maintaining high-speed performance, resolving the contradiction between speed and flexibility.
Solution Approach 2:
The patent enables parameter changes in the DSP block configuration, such as adjusting the number of parallel multipliers, adder configurations, and data path widths. These parameter changes allow the hardware accelerator to be optimized for different algorithms and data types while maintaining dedicated hardware performance, thus achieving both speed and adaptability.
Data Source
AI summary
An accelerated processor structure on a programmable integrated circuit device includes a processor and a plurality of configurable digital signal processors (DSPs). Each configurable DSP includes a circuit block, which in turn includes a plurality of multipliers. The accelerated processor structure further includes a first bus to transfer data from the processor to the configurable DSPs, and a second bus to transfer data from the configurable DSPs to the processor.


