Fully CMOS MUX Slices for 200G+ Serializer Timing Balance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing serializer designs in high-speed wireline systems face challenges in achieving low power consumption and efficient performance due to improper architecture selection, particularly at data rates over 200 gigabits/second, where generating and distributing half-rate clocks with stringent timing requirements lead to poor power efficiency and performance.
Innovation Solution
The implementation of a multiplexer system with Q-mux and I-mux slices, utilizing interconnection, inverters, and buffers to balance clock and data signal propagation delays, and incorporating digital-to-analog converters to support high-speed data transmission with reduced power consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If half-rate clock generation and distribution is used in serializer design, then data rate over 200 gigabits/second can be supported, but power consumption increases and timing requirements become stringent
Solution Approach 1:
The multiplexer is divided into multiple independent stages (first stage with multiple input branches, second stage combining outputs), allowing each stage to operate at relaxed clock rates while achieving full-rate output through cascaded operation. This segmentation eliminates the need for a single high-speed half-rate clock throughout the entire circuit.
Solution Approach 2:
The patent transitions from a single-stage time-multiplexing approach to a multi-stage architecture that adds spatial dimension (multiple parallel branches in first stage, multiple parallel paths in second stage). This dimensional expansion allows lower clock frequencies to achieve the same effective data rate through parallel processing.
2Device complexity
If single stage multiplexer with shared output node is used, then device complexity is reduced, but large self-loading makes it difficult to obtain full-rate symbol with low power consumption
Solution Approach 1:
The single shared output node is segmented into multiple output nodes distributed across different stages. The first stage has multiple output branches that feed into the second stage, distributing the loading across multiple nodes rather than concentrating it all in one node. This reduces the self-loading effect at each individual node.
Solution Approach 2:
The patent adds a temporal and spatial dimension to the output structure by using multiple stages with multiple output nodes per stage. Instead of one output node handling all signals simultaneously, the output function is distributed across multiple nodes operating at different times and locations in the circuit.
3Loss of time
If quarter-rate architecture is used for MUX implementation, then timing requirements are relaxed, but large self-loading due to shared output node persists
Solution Approach 1:
The quarter-rate architecture is implemented with segmented stages where each stage operates independently at the relaxed quarter-rate clock. The first stage processes four input branches in parallel, and the second stage combines their outputs, allowing the relaxed timing to be maintained throughout while achieving full-rate output through the cascaded structure.
Data Source
Figure 1
Figure 2A
Figure 2B
AI summary
A module including a first slice and a second slice. The first slice and the second slice receive data from a plurality of inputs. A first stage of the first slice selects a first subset of the inputs in synchronization with an edge of a first clock. In synchronization with a second clock, a second stage of the first slice selects an input from the first subset. A first stage of the second slice selects a second subset of the inputs in synchronization with an edge of the second clock. In synchronization with the first clock, a second stage of the second slice selects an input from the second subset.