Cascaded DSP Register Blocks for Concurrent Arithmetic Throughput
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing demand for real-time, concurrent arithmetic operations in digital signal processing (DSP) applications exceeds the capabilities of traditional DSP chips and Programmable Logic Devices (PLDs), as the latter face bottlenecks due to the limitations of multiple DSP elements programmed in the PLD fabric.
Innovation Solution
A digital signal processing circuit with cascaded DSP elements and a programmable logic device (PLD) architecture that includes a top DSPE with input register blocks coupled to a multiplier and arithmetic logic units (ALUs), allowing for improved arithmetic operations through enhanced interconnectivity and pipelining within the FPGA fabric.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If multiple DSP elements are configured in the programmable logic of a PLD to allow concurrent DSP operations, then the capability for real-time arithmetic operations is improved, but the bottleneck becomes the fabric of the PLD, limiting further performance improvement
Solution Approach 1:
The PLD is segmented into multiple dedicated DSP blocks, each containing complete DSP elements with multipliers, adders, and register blocks. This segmentation allows each DSP block to operate independently with dedicated resources, eliminating the fabric bottleneck that occurs when multiple DSP elements share common interconnect resources in traditional PLD architectures.
Solution Approach 2:
Dedicated interconnect structures are introduced as intermediaries between DSP blocks and the PLD fabric. These specialized interconnects provide high-speed data transfer paths specifically optimized for DSP operations, separating the critical data movement paths from the general-purpose fabric and eliminating the bottleneck effect.
2Adaptability or versatility
If a single DSP microprocessor is used in traditional DSP chips, then the device is programmable with a native instruction set, but the capability to handle massive concurrent arithmetic operations is insufficient
Solution Approach 1:
Multiple DSP elements are merged into a single integrated DSP block structure, combining multiple multipliers, adders, and register blocks into one cohesive unit. This merging allows the system to perform multiple arithmetic operations simultaneously within a single programmable block, achieving both high concurrency and programmability through a unified instruction set architecture.
Solution Approach 2:
Each DSP block is designed as a universal unit capable of performing multiple types of arithmetic operations (multiplication, addition, subtraction, accumulation) through a single programmable interface. The multi-functional DSP blocks can be configured via the native instruction set to handle various DSP algorithms, eliminating the need for multiple specialized processors while maintaining programmability.
3Productivity
If more DSP microprocessors are added to perform DSP applications in parallel, then the arithmetic operations per second increase, but the same disadvantages of general-purpose microprocessors persist, making them not ideally suited for numerically-intensive requirements
Solution Approach 1:
Instead of using multiple general-purpose microprocessors, the invention creates dedicated DSP blocks that copy and integrate the essential arithmetic computation resources (multipliers, adders, registers) directly into the PLD fabric. This copying of critical computational resources eliminates the overhead and inefficiency of general-purpose processors while maintaining the ability to perform numerically-intensive operations with high efficiency.
Data Source
AI summary
An integrated circuit that includes a digital signal processing element (DSPE) having a first and a second register block coupled to a first arithmetic logic unit (ALU) circuit; a middle DSPE adjacent to the top DSPE having a third and a fourth register block coupled to a second ALU circuit, where the third register block is coupled to the first register block, and the fourth register block register block is coupled to the second register block; and a bottom DSPE adjacent to the middle DSPE having a fifth and a sixth register block coupled to a third ALU circuit, where the fifth register block is coupled to the third register block and the sixth register block register block is coupled to the fourth register block.


