Multi-level pipeline phase rotator and application thereof in clock and data recovery loop

Through the combination of pipeline operation, dual-phase rotator and FIR filter, the problems of high-power consumption and large-area in high-speed communication are solved, and the CDR loop with low power consumption and low-area is realized, which is suitable for data center communication of 200-Gb/s and above.

CN120303879APending Publication Date: 2025-07-11西纳公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380082693.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-01
Filing Date
2023-11-29
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

While pursuing faster speed, the optical and electrical communication hardware of existing data centers faces the problems of high power consumption and large-area consumption. Especially in the 200-Gb/s CDR that meets the strict jitter requirements, the LCVCO implementation method consumes too much power.

Method used

A method of pipeline operation combined with a dual-phase rotator (PR) is used in combination with a finite impulse response (FIR) filter for clock and data recovery (CDR) loops, reducing power and area consumption, and reducing phase noise and jitter through staging phase selection.

Benefits of technology

It realizes a low power consumption and low area CDR loop, can meet the jitter requirements of 200-Gb/s and above, reduces power consumption and hardware area, and is suitable for high-speed wired transceivers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303879A_ABST
    Figure CN120303879A_ABST
Patent Text Reader

Abstract

Aspects of the subject disclosure may include, for example, implementing: a first stage including a first number of phase rotators in parallel, the first number of phase rotators generating respective clock phases offset by a fixed amount; and a second stage including a second number of phase rotators, the second number of phase rotators receiving outputs of the first number of phase rotators from the first stage, the second stage outputting a first weighted sum of respective clock phases generated by the second number of phase rotators. The subject disclosure also includes that the second number of phase rotators is less than the first number of phase rotators, and the total number of bits dedicated to phase selection is divided over the first stage and the second stage. Other embodiments are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the priority of U.S. Patent Application No. 18 / 060,787, filed on December 1, 2022. All parts of the above application are incorporated herein by reference in their entirety. Technical Field

[0003] The present disclosure relates to a method and apparatus for clock and data alignment that reduces power consumption. Background Art

[0004] The demand for greater bandwidth in data centers continues to increase, forcing the need for faster optical communication hardware and electrical communication hardware. Despite the need for faster communication, the power that hardware can consume is limited, driven by capacity and environmental issues. Existing data centers are equipped to handle a limited amount of power from the power grid, and current estimates suggest that data centers will consume 8% of the world's total electricity by 2030. Brief Description of the Drawings

[0005] Reference will now be made to the drawings, which are not necessarily to scale, and in which:

[0006] Figure 1 An exemplary non-limiting embodiment of a serializer / deserializer (SerDes) in accordance with various aspects described herein is shown.

[0007] Figure 2 An exemplary non-limiting embodiment of a SerDes using a digital-to-analog converter (DAC) and an analog-to-digital converter (ADC) in accordance with various aspects described herein is shown.

[0008] Figure 3 An exemplary non-limiting embodiment of a standard high-speed DAC and a serializer in accordance with various aspects described herein is shown.

[0009] Figure 4 An exemplary non-limiting embodiment of a standard high-speed ADC and a deserializer in accordance with various aspects described herein is shown.

[0010] Figure 5 An exemplary non-limiting embodiment of a clock and data recovery loop (a) phase rotation-based clock and data recovery (PR-based CDR), (b) LC voltage-controlled oscillator-based clock and data recovery (LCVCO-based CDR) in accordance with various aspects described herein is shown.

[0011] Figure 6 An exemplary non-limiting embodiment of phase interpolation using a phase rotator in accordance with various aspects described herein is shown.

[0012] Figure 7 An exemplary non - limiting embodiment of the constellation of an octagonal quadrature phase rotator in accordance with various aspects described herein is shown.

[0013] Figure 8 An exemplary non - limiting embodiment of the integrated nonlinearity of a 6b quadrature and octagonal phase rotator in accordance with various aspects described herein is shown.

[0014] Figure 9 An exemplary non - limiting embodiment of a dual - phase rotator in accordance with various aspects described herein is shown.

[0015] Figure 10 An exemplary non - limiting embodiment of the constellation of a dual - phase rotator in accordance with various aspects described herein is shown.

[0016] Figure 11 An exemplary non - limiting embodiment of the integrated nonlinearity of a dual - phase rotator in accordance with various aspects described herein is shown.

[0017] Figure 12 An exemplary non - limiting embodiment of a two - stage tournament - type pipelined phase rotator in accordance with various aspects described herein is shown.

[0018] Figure 13 An exemplary non - limiting embodiment of a two - stage pipelined phase rotator in accordance with various aspects described herein is shown.

[0019] Figure 14 An exemplary non - limiting embodiment of a two - stage pipelined dual - phase rotator in accordance with various aspects described herein is shown.

[0020] Figure 15 An exemplary non - limiting embodiment of a three - stage pipelined phase rotator in accordance with various aspects described herein is shown.

[0021] Figure 16 Another exemplary non - limiting embodiment of a two - stage pipelined dual - phase rotator in accordance with various aspects described herein is shown.

[0022] Figure 17 An exemplary non - limiting embodiment of a finite impulse response filter in accordance with various aspects described herein is shown.

[0023] Figure 18 An exemplary non - limiting embodiment of the 27 - tap Kaiser - Bessel finite impulse response (FIR) frequency response in accordance with various aspects described herein is shown, where f 3dB = 1 GHz, f sample(采样) = 17 GHz.

[0024] Figure 19 An exemplary non - limiting embodiment of an analog FIR filter using voltage - mode summation in accordance with various aspects described herein is shown.

[0025] Figure 20 An exemplary non - limiting embodiment of the layout of a two - stage pipelined dual - phase rotator in accordance with various aspects described herein is shown.

[0026] Figure 21 An exemplary non - limiting embodiment of the phase - noise plot of an FIR filter connected to the output of a PR in accordance with various aspects described herein is shown. Detailed Description

[0027] The present disclosure describes illustrative embodiments of a method for pipelining a phase rotator, which results in power and area savings. Other embodiments are described in the present subject matter disclosure.

[0028] One or more aspects of the present subject matter disclosure include the combination of pipelining operation with a dual - phase rotator (PR) to form a low - power, low - jitter, and low - area clock and data recovery (CDR) loop.

[0029] One or more aspects of the present subject matter disclosure include the use of a finite impulse response (FIR) filter in a CDR loop that helps suppress phase noise and reduce jitter.

[0030] The techniques herein include a method for pipelining a PR, which results in power and area savings. The combination of pipelining operation and dual PRs results in a low - power, low - jitter, and low - area CDR loop. The use of an FIR filter in the CDR loop helps suppress phase noise and reduce jitter. The combination of these features enables the use of a PR in a 200 - Gb / s CDR, where strict jitter requirements have led to the LC voltage - controlled oscillator (LCVCO) being the only proven option. These features also result in the lowest possible area and power for the CDR, both of which are highly valuable in high - speed wired transceivers.

[0031] The method for pipelining a PR presented herein results in lower power and area consumption than traditional PRs or league - type PRs. For example, the reduction factors for power and area of a two - stage pipelined PR with equal bits in the first and second stages implemented with the method disclosed herein are (3)(2 - N / 2). If more bits are used in the second stage or if additional stages are added to the PR, the power and area reduction factors are further improved.

[0032] The pipelined operation PR method disclosed herein is combined with the concept of dual PR to further improve PR non-linearity, i.e., integral non-linearity (INL), which directly translates to jitter during quasi-synchronous operation. Since power and area are very precious in high-speed wired transceivers, traditional dual PR is not common in serializer-deserializer (SerDes). Here, the low power and area consumption of pipelined PR enables the use of dual PR.

[0033] An FIR filter is added to the CDR loop to suppress phase noise and filter jitter. This strategy has been used in fractional-N frequency synthesizers to suppress ΣΔ noise, but has never been used in CDRs to suppress noise caused by quasi-synchronous operation.

[0034] 200-Gb / s+ SerDes has strict jitter requirements, which are only met in LCVCO-based CDRs. However, such implementations suffer from high power and area consumption. The pipelined operation PR method described here, combined with the use of dual PR and FIR filters in CDRs, shows a way to use PR in 200-Gb / s+ SerDes.

[0035] The CDR loop is implemented in 3nm FinFET CMOS to test the concepts described herein. The dual-pipelined PR achieves 390-fs peak-to-peak INL while consuming 10.6 mA from a 0.7V supply. During quasi-synchronous operation, the PR phase noise is measured and integrated from DC to Nyquist, resulting in 255-fs RMS jitter. Adding a 27-tap Kaiser-Bessel FIR filter at the PR output further reduces the PR RMS jitter to 207-fs. Once shaped by the complete CDR loop, the jitter specification of this PR will enable a low-power and area CDR, and in turn a 200-Gb / s+ wired transceiver that meets the <75fs RMS random jitter standard.

[0036] The PR concepts described herein are not limited to use within a CDR loop. The same concepts can be used elsewhere where PR exists to achieve low jitter, power, and area. For example, these PR concepts can be used for internal clock generation in analog-to-digital converters (DACs) or digital-to-analog converters (ADCs).

[0037] To limit the total power consumed in a data center, key hardware, i.e., ADCs, DACs, and SerDes, must increase their power only at the same rate as their speed. For example, a very short reach (VSR) SerDes expected to operate at 224 gigabits per second (Gb / s) is expected to consume a total of 448 mW, corresponding to a power efficiency of 2 picojoules per bit (pJ / b).

[0038] SerDes( Figure 1 ) consists of two main blocks: a transmitter (Tx) and a receiver (Rx). The main responsibility of the Tx is to serialize many low-speed data paths into a single high-speed data path. Conversely, the Rx deserializes the high-speed data path into many low-speed data paths. As the transmission speed increases, SerDes becomes increasingly dependent on high-speed medium-resolution DACs and ADCs to perform their serialization and deserialization.

[0039] High-speed DACs and ADCs( Figure 2 ) can be further divided into two parts: a data path and a clock path. This main body disclosure addresses the clock path that consumes most of the power budget of the DAC and ADC.

[0040] The basic purpose of a DAC is to receive an N-bit binary bus and convert it into a single analog signal. Modern DACs also perform serialization through cascaded multiplexers (MUXes) that use progressively higher-speed clocks as their select bits to combine several low-speed data paths.

[0041] The basic purpose of an ADC is to receive a single analog signal and convert it into an N-bit binary bus. Modern ADCs use a time-interleaved architecture where a sampling front end (SFE) first deserializes the data into lower-speed paths before the parallel sub-ADCs perform the actual data conversion, each sub-ADC operating at operation. Fs is the total sampling rate of the ADC. Rank 1 and Rank 2 are integers representing the number of low-speed data paths after the first and second levels of interleaving, respectively.

[0042] See Figure 3 and Figure 4 , modern DACs and ADCs have sampling rates in the range of 100 to 200 gigasamples per second (GS / s) and may require multi-phase clocks that operate at any frequency from to . As an example, 112-GS / s DACs and ADCs are required to perform 224-Gb / s PAM4 encoded data transmission. A common approach for DACs is to utilize 16:8, 8:4, and 4:1 MUX stages. This requires a four-phase clock at , an eight-phase clock at , and a 16-phase clock at . For ADCs, it is common that Rank 1 = 8 and Rank 2 = 12, requiring an eight-phase clock at and a 96-phase clock at .

[0043] Another important aspect of SerDes is clock - to - data alignment to ensure that sampling occurs at the optimal point. A CDR loop ( Figure 5 ) is used for this alignment. The basic operation of the CDR is as follows. A phase detector (PD) is used to recover the data and compare it with its sampling clock. The PD outputs pulses that are equivalent to the phase mismatch between the data and the sampling clock. These pulses are then filtered and used to drive a PR or an LCVCO.

[0044] Among these two strategies, LCVCO - based CDR is less common due to its high power and area consumption. However, CDRs have started using LCVCOs to meet strict jitter requirements. For example, the target for a 200 - Gb / s SerDes implementation is <75 fs, rms random jitter. This offset is mainly due to the difficulty of designing a PR that can meet such jitter requirements. However, the present subject disclosure presents new concepts that enable the implementation of PR - based CDRs to be possible at 200 - Gb / s and above.

[0045] Most simply, a PR ( Figure 6 ) takes the weighted sum of two input clocks CK in1 and CK in2 with phases θ1 and θ2 respectively, to produce an output CK out with phase θ out . Treating the clocks as phasors, the output is related to the inputs by CK out =(1 - α)(cosθ1 + jsinθ1)+α(cosθ2 + jsinθ2). The weights of the two clocks are arranged such that α is between 0 and 1, so that as the weight of CK in1 increases, the weight of CK in2 decreases by the same amount. This results in θ out being closer to θ1 when α is low and θ out being closer to θ2 when α is high.

[0046] When CK in1 and CK in2 are separated by 90°, the output clock simplifies to CK with output phase out =(1 - α)+jα. One way to define the PR characteristics is their INL, which is a measure of how much the output phase deviates from the ideal output phase. The ideal output phase of an orthogonal PR is given by θ out,ideal =(α)(90°), so the INL can be defined as:

[0047] The INL may be more useful when defined in seconds. Given the input clock period of T CK and the N bits dedicated to phase selection, where α can increase from 0 to 1 in steps of 90° / 2 N From this formula, it is clear that to improve INL, either the number of bits dedicated to phase selection must be increased, or the input clock period must be decreased. For each additional bit added, the phase rotator power consumption doubles. The phase rotator power consumption also increases linearly with the input clock frequency. Additionally, there are diminishing returns when increasing the PR bits. Beyond 8 or 9 bits for phase selection, the PR peak-to-peak INL stops improving in practical implementations.

[0048] Another way to improve PR linearity is to make CK in1 and CK in2 separated by 45°. In this case, the output clock simplifies to where the output phase Now the INL becomes and

[0049] Although the INL is improved, the octagonal PR implementation requires 8 input phases, which adds significant complexity and power earlier in the CDR.

[0050] A comparison of the phase rotator constellations of the ideal octagonal and orthogonal implementations can be seen in Figure 7 , Figure 8 which compares the theoretical INL of 6b octagonal and orthogonal PRs. During quasi-synchronous operation, the worst-case jitter of the CDR occurs, where there is a frequency mismatch between the input data and the sampling clock in quasi-synchronous operation. This causes the PR to spin and not lock to a single phase. The peak-to-peak INL measured in seconds is converted to random jitter during quasi-synchronous operation. Therefore, finding new low-power methods to improve INL is crucial for PRs to be used in the CDR of 200-Gb / s links. The INL formula gives a local maximum at α = 0.25 and a local minimum at α = 0.75. It has been shown that one strategy to improve PR linearity is to use a dual PR( Figure 9 ). The idea is to add the outputs of two PRs, where α PR2 = α PR1 + 0.5. When the INL of PR2 is maximum, the INL of PR1 is minimum and vice versa, thus resulting in approximate cancellation of the INL. The dual PR constellation can be seen in Figure 10 .

[0051] Figure 11 compares the theoretical INL of traditional 6b orthogonal PR with dual 6b orthogonal PR.

[0052] Process, voltage, and temperature (PVT) variations cause the INL of each PR to lose some of its correlation, so the result is not zero non-linearity, but the INL is significantly reduced. The disadvantage of this concept is that the PR power and area consumption are doubled.

[0053] The present disclosure provides a PR concept that can be used to reduce power and area consumption using pipelined operations. The concept behind PR pipelined operations is to split the bits dedicated to phase selection into multiple stages. The previously used PR pipelined operations were done in a league manner, where additional PRs were used in each stage, and MUXing was used after each stage to decide which phase to forward ( Figure 12 ). Although each driver stage in the league-type PR uses less power, the total number of driver stages increases, resulting in no power or area savings. This PR approach is more similar to multi-phase generation, so improved linearity is achieved (similar to going from orthogonal PR to octagonal PR), but the power consumption is equal to or greater than that of traditional PR.

[0054] The PR pipelined operation method disclosed herein uses the same number of PRs in the first stage as the total number of stages, but in each subsequent stage, the total number of PRs decreases. For example, consider a two-stage PR separated by a dashed line using the pipelined operation method disclosed herein ( Figure 13 ). The first stage depicted by reference numeral 101 uses two parallel PRs to generate two clock phases that are offset by a single least significant bit (LSB). The second stage depicted by reference numeral 102 then uses a single PR to interpolate between these two phases entering the second stage.

[0055] For comparison, consider a traditional PR with N bits dedicated to phase selection. Each controllable driver stage requires 2 N cell devices, so a total of (2)(2 N ) cell devices are used in the PR. In the pipelined operation method disclosed herein, bits are used in the first stage and bits are used in the second stage. The same resolution is achieved, but (4)(2 N / 2 ) cell devices are required in the first stage and (2)(2 N / 2 ) devices are required in the second stage. Therefore, the total number of cell devices is now (6)(2 N / 2 ), resulting in a significant power and area reduction factor of (3)(2 -N / 2)。If more bits are used in the second stage than in the first stage, the power and area reduction factor can be further improved. For example, in a non-limiting embodiment, 4 bits can be used in the first stage, while 5 bits can be used in the second stage. As more stages are added to the PR pipelining operation method disclosed herein, the total number of cell devices continues to decrease, and the power and area reduction factor is further improved.

[0056] Figure 14 An example of a two-stage pipelined dual-phase rotator is shown. The first stage, represented by reference numeral 103, utilizes two (dual) parallel PRs 111. The second stage, depicted by reference numeral 104, uses two parallel PRs. The summing node is represented by reference numeral 110. This enables the use of dual PRs without significant power or area penalty. Combining these concepts results in a highly linear low-power PR that can be used in future generations of SerDes.

[0057] In Figure 15 In another non-limiting embodiment shown, a three-stage pipelined phase rotator can be implemented. The first stage, represented by reference numeral 105, utilizes two (dual) parallel PRs 111. The second stage, depicted by reference numeral 106, uses two parallel PRs. The third stage, depicted by reference numeral 107, uses a single PR. Any value of α1 can exist compared to α2 (and α3 to α4). Preferably these codes are offset from each other by 1 LSB, but this is not required. They can be any offset while still achieving many benefits. For example, assume CKin1 is 0°, and Ckin2 is 90°, and 3 bits are dedicated to stage 1. If α1 and α2 are offset by 1 LSB, then stage 2 interpolates between phases offset by 90° / 2 3 = 11.25°. If α1 and α2 are offset by 2 LSBs, then stage 2 interpolates between phases offset by 22.5°.

[0058] In another non-limiting embodiment, the position of the summing node 110 in the two-stage pipelined phase rotator can be changed, for example see Figure 16 , and there is still some pipelining operation occurring afterwards. Note the rearrangement of the alpha (α) values among the stages. The benefit of this embodiment is that 2 groups of unit cells are discarded from stage 2, thus saving an additional 2×2 N unit cells in the overall design, where N is the number of bits dedicated to stage 2.

[0059] Another embodiment of the present disclosure includes improving PR linearity by adding an FIR filter to the output of the PR. FIR filters have been used in the feedback path of fractional-N frequency synthesizers to suppress ΣΔ noise. The present disclosure proposes using an FIR filter in the CDR. The FIR filter can beFigure 17 As seen in, in a basic FIR filter, all α coefficients are equal, resulting in a stopband attenuation of 20log sample at f / 2 of 20log 10 N and a 3-dB BW proportional to f sample / N, where N is the number of FIR taps and f sample = 1 / T D , where T D is the z -1 delay.

[0060] In an analog implementation, increasing the FIR taps to more than a few dozen taps can be difficult and error-prone, but this still provides sufficient stopband attenuation and bandwidth to significantly reduce the phase noise introduced by PR. Other types of FIR filters, such as Kaiser-Bessel FIR, can be used to improve the stopband attenuation and 3-dB bandwidth without increasing the number of taps. In this implementation, the first zero-order modified Bessel function is used to calculate the coefficients. In Figure 18 an example of the frequency response of a 27-tap Kaiser-Bessel FIR can be seen, where f sample = 17 GHz. The stopband attenuation is around 40 dB and the 3-dB bandwidth is around 1 GHz. An example of an analog FIR filter implementation can be seen in Figure 19 . The summation is done in the voltage domain and the resistor values are modified to set the tap coefficients.

[0061] To test the concepts disclosed herein, an 11-bit 17-GHz version of the phase rotator from Figure 14 was implemented in 3-nm FinFET CMOS and the phase rotator was simulated using Cadence Spectre. The layout of the phase rotator can be seen in Figure 20 . The first stage is represented by reference numeral 108 and the second stage is represented by reference numeral 109. The two-stage pipelined dual-phase rotator has an effective area of 83 μm × 94 μm.

[0062] Simulations of quasi-synchronous operation with PR were performed. The resulting phase noise plots can be seen in Figure 21 . The RMS jitter was measured to be 255 fs integrated from DC to f CK / 2. This number is also prior art. In another simulation where PR was in static operation, the measured jitter was 55 fs. Then, another simulation was performed using the FIR filter after PR. The resulting phase noise can also be seen in Figure 21 , and the improvement in phase noise using the FIR filter is clear. This improvement is again shown by measuring the RMS jitter that dropped from 255 fs to 207 fs.

[0063] Unless the context clearly indicates otherwise, the use of the terms "first", "second", "third", etc. in the claims is for clarity only and does not otherwise indicate or imply any order.

[0064] In addition, the words "example" and "exemplary" are used herein to mean serving as an instance or illustration. Any embodiment or design described herein as "example" or "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs. Instead, the use of the word example or exemplary is intended to present concepts in a concrete manner. As used in this application, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X employs A or B" is intended to mean any natural inclusive arrangement. That is, if X uses A; X uses B; or X employs both A and B, then "X uses A or B" is satisfied in any of the foregoing instances. In addition, unless otherwise specified or clear from the context with respect to the singular form, the articles "a" and "an" used in this application and the appended claims shall generally be construed to mean "one or more".

[0065] As used herein, the term "processor" can refer to substantially any computing processing unit or device, including but not limited to a single-core processor; a single-processor with software multithreading execution capabilities; a multi-core processor; a multi-core processor with software multithreading execution capabilities; a multi-core processor with hardware multithreading technology; a parallel platform; and a parallel platform with distributed shared memory. In addition, a processor can refer to an integrated circuit, an application specific integrated circuit (ASIC), a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic controller (PLC), a complex programmable logic device (CPLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A processor can employ a nanoscale architecture, such as but not limited to transistors, switches, and gates based on molecules and quantum dots, in order to optimize space usage or enhance the performance of a user device. A processor can also be implemented as a combination of computing processing units.

[0066] The foregoing description includes only examples of various embodiments. Of course, it is not possible to describe every conceivable combination of components or methods in order to describe these examples, but one of ordinary skill in the art will recognize that many further combinations and permutations of this embodiment are possible. Accordingly, the embodiments disclosed and / or claimed herein are intended to cover all such changes, modifications, and variations that fall within the spirit and scope of the appended claims. In addition, insofar as the term "comprising" is used in the detailed description or the claims, this term is intended to be inclusive in a manner similar to the way the term "including" is interpreted when used as a transitional word in the claims.

[0067] Although specific embodiments have been illustrated and described herein, it should be understood that any arrangement that achieves the same or similar purpose may be substituted for the embodiments described or shown by the subject disclosure. The subject disclosure is intended to cover any and all modifications or variations of various embodiments. Combinations of the above-described embodiments and other embodiments not specifically described herein may be used in the subject disclosure. For example, one or more features from one or more embodiments may be combined with one or more features of one or more other embodiments. In one or more embodiments, features stated positively may also be stated negatively and excluded from the embodiment with or without being replaced by another structural and / or functional feature. The steps or functions described with respect to the embodiments of the subject disclosure may be performed in any order. The steps or functions described with respect to the embodiments of the subject disclosure may be performed alone or in combination with other steps or functions of the subject disclosure, and with other steps from other embodiments or steps not described in the subject disclosure. In addition, more or fewer than all of the features described for the embodiments may also be utilized.

[0068] The embodiments disclosed herein have been described with reference to the accompanying drawings. Similarly, for purposes of explanation, specific numbers, materials, and configurations have been set forth in order to provide a thorough understanding. However, the embodiments may be practiced without these specific details.

Claims

1. A multi - stage pipelined phase rotator, comprising: A first stage, including a first number of phase rotators in parallel, the first number of phase rotators generating corresponding clock phases offset by a fixed amount; A second stage, including a second number of phase rotators, the second number of phase rotators receiving the outputs of the first number of phase rotators from the first stage, and the second stage outputting a first weighted sum of the corresponding clock phases generated by the second number of phase rotators; Wherein the second number of phase rotators is less than the first number of phase rotators, and Wherein the total number of bits dedicated to phase selection is divided between the first stage and the second stage.

2. The multi - stage pipelined phase rotator according to claim 1, wherein the first stage has a first number of bits based on the total number of bits dedicated to phase selection, wherein the second stage has a second number of bits based on the total number of bits dedicated to phase selection, wherein each phase rotator in the first stage includes a first number of phase interpolation unit cells determined according to the first number of bits dedicated to phase selection, and wherein each phase rotator in the second stage includes a second number of phase interpolation unit cells determined according to the second number of bits dedicated to phase selection.

3. The multi - stage pipelined phase rotator according to claim 2, wherein the first stage and the second stage further include a first plurality of summing nodes and a second plurality of summing nodes respectively, for receiving the outputs of the phase interpolation unit cells of the corresponding phase rotators.

4. The multi-stage pipelined phase rotator according to claim 1 further includes a third stage, the third stage includes a third number of phase rotators, the third number of phase rotators receive the output of the second number of phase rotators from the second stage, and the third stage outputs a second weighted sum of the corresponding clock phases generated by the third number of phase rotators, wherein, The third number of phase rotators is less than the second number of phase rotators.

5. The multi - stage pipelined phase rotator according to claim 2, wherein The second number of bits dedicated to phase selection used in the second stage is greater than the first number of bits dedicated to phase selection used in the first stage.

6. The multi - stage pipelined phase rotator according to claim 2, wherein The first number of phase interpolation unit cells and the second number of phase interpolation unit cells include inverters or common - source amplifiers with controllable variable drive strength.

7. The multi - stage pipelined phase rotator according to claim 3, further comprising an intermediate stage between the first stage and the second stage, the intermediate stage including a third plurality of summing nodes for receiving the outputs of the first plurality of summing nodes.

8. The multi - stage pipelined phase rotator according to claim 1, wherein the total number of bits dedicated to phase selection on the first stage and the second stage is 9.

9. A phase - rotator - based clock and data recovery (CDR) loop, comprising: A data sampler for sampling input data; A phase detector for receiving the sampled data from the data sampler, detecting whether the data is sampled at the optimum point, and passing the detection result to a loop filter; A loop filter for receiving the detection result from the phase detector and outputting a code based on the detection result to indicate to a phase rotator (PR) how many codes it must rotate to sample the data at the optimum point; A phase rotator for receiving the code from the loop filter and instructing the data sampler to sample the data, wherein A finite impulse response (FIR) filter is placed between the phase detector and the data sampler.

10. The CDR loop according to claim 9, wherein, The FIR filter includes a Kaiser-Bessel FIR or an analog FIR using voltage-mode summation.

11. The CDR loop according to claim 9, wherein, The PR includes: A first stage including a first number of phase rotators in parallel, the first number of phase rotators generating respective clock phases offset by a fixed amount; A second stage including a second number of phase rotators, the second number of phase rotators receiving the outputs of the first number of phase rotators from the first stage, the second stage outputting a first weighted sum of the respective clock phases generated by the second number of phase rotators; wherein the second number of phase rotators is less than the first number of phase rotators, and wherein the total number of bits dedicated to phase selection is divided between the first stage and the second stage.

12. The CDR loop according to claim 11, wherein, The first stage has a first number of bits based on the total number of bits dedicated to phase selection, wherein the second stage has a second number of bits based on the total number of bits dedicated to phase selection, wherein each phase rotator in the first stage includes a first number of phase interpolator unit cells determined according to the first number of bits dedicated to phase selection, and wherein each phase rotator in the second stage includes a second number of phase interpolator unit cells determined according to the second number of bits dedicated to phase selection.

13. The CDR loop according to claim 12, wherein, The first stage and the second stage further include a first plurality of summing nodes and a second plurality of summing nodes respectively for receiving the outputs of the phase interpolator unit cells of the respective phase rotators.

14. The CDR loop according to claim 11, wherein the second number of bits dedicated to phase selection used in the second stage is greater than the first number of bits dedicated to phase selection used in the first stage, and wherein the first number of phase interpolator unit cells and the second number of phase interpolator unit cells include inverters or common-source amplifiers having controllable variable drive strengths.

15. The CDR loop according to claim 9, wherein the data sampler is part of an analog-to-digital converter (ADC) and is implemented together with the phase rotators in an analog macro, and the phase detector and the loop filter are implemented in a digital signal processing (DSP) engine.