Parallel Pipelined Stream Processor Switching Noise Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing stream processing systems face performance limitations due to inefficient control logic optimization, which leads to significant switching noise and resource bottlenecks, particularly in parallel pipelined hardware implementations like FPGAs, where control logic fan-out and signal propagation issues hinder optimal performance.
Innovation Solution
The method involves high-level synthesis to optimize control logic concurrently with data path scheduling, using minimum-cut partitioning and phase transition registers to minimize control logic fan-out and align data across different clock phases, thereby reducing switching noise and optimizing hardware resource utilization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If control logic is optimized at the RTL stage, then the design complexity is reduced, but the optimization scope is limited and cannot achieve global optimization
Solution Approach 1:
The patent moves the optimization process from the traditional RTL dimension to the high-level synthesis dimension, adding a new dimension to the design flow. This allows control logic optimization to occur earlier in the design process, before the RTL code is generated, thereby expanding the optimization scope to include the entire system rather than just the control logic portion.
Solution Approach 2:
The patent performs control logic optimization as a preliminary action during the high-level synthesis stage, before the RTL implementation is finalized. By optimizing control logic early in the design process, the system achieves global optimization of both data path and control logic together, rather than limiting optimization to the later RTL stage.
2Productivity
If parallel pipelined hardware is used to increase computing power, then calculation speed is improved, but switching noise increases due to simultaneous logic switching
Solution Approach 1:
The patent segments the parallel pipelined hardware into multiple clock phases, dividing the simultaneous switching events into sequential phases. This segmentation of the clocking scheme allows the system to maintain high computational throughput while reducing peak switching noise by distributing logic switching across different time phases.
Solution Approach 2:
The patent implements periodic action by using multi-phase clocking, where different groups of logic elements are activated in periodic phases rather than simultaneously. This periodic activation pattern maintains the overall calculation speed while reducing instantaneous switching noise through phased operation.
3Productivity
If more transistors are accommodated per unit area by reducing transistor size, then processing power is increased, but switching noise and signal integrity issues worsen
Solution Approach 1:
The patent segments the dense transistor array into multiple clock phases, where different spatial regions are activated in different phases. This spatial and temporal segmentation allows high-density transistor placement while reducing switching noise by preventing simultaneous switching across the entire chip.
Solution Approach 2:
The patent introduces dynamic clock phase assignment, where the activation phase of different logic elements can be dynamically adjusted. This dynamic approach allows the system to optimize for both high processing power through dense transistor usage and low switching noise through phased activation patterns.
Data Source
AI summary
A method of configuring a hardware design for a pipelined parallel stream processor includes obtaining a scheduled graph representing a processing operation in the time domain as a function of clock cycles. The graph includes a data path to be implemented in hardware as part of the stream processor, an input, an output, and parallel branches to enable data values to be streamed therethrough from the input to the output as a function of increasing clock cycle. The data path is partitioned into a plurality of discrete regions, each region operating on a different clock phase and having discrete control logic elements. Phase transition registers to align data separated by a boundary between regions having different clock phases are introduced into the data path at the boundary. The graph and control logic elements define a hardware design for the pipelined parallel stream processor.


