Dynamic timing loop gain for compensating phase interpolator nonlinearity

By introducing a digital timing recovery loop into digital communication, and using phase interpolation and dynamic loop gain to compensate for the nonlinearity of the phase interpolator, the instability problem of the clock recovery circuit is solved, and the accuracy and stability of data recovery are improved, especially under high symbol rate and complex channel conditions.

CN120729348BActive Publication Date: 2026-01-13CREDO TECHNOLOGY GROUP LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202411634957.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2024-04-22
Filing Date
2024-11-15
Publication Date
2026-01-13
Estimated Expiration
2044-11-15

AI Technical Summary

Technical Problem

In digital communication, the nonlinearity of the phase interpolator causes oscillations and instability in the clock recovery circuit, especially at high clock frequencies, where performance degrades and it becomes difficult to accurately recover data in the presence of inter-symbol interference (ISI) and additive noise.

Method used

The digital timing recovery loop in the integrated circuit transceiver is used to compensate for the nonlinearity of the phase interpolation through phase interpolation and dynamic loop gain. It includes a phase interpolator, sampling element, timing error estimator and feedback circuit. The gain of the phase interpolator is dynamically adjusted by the scaling element and lookup table in the feedback circuit to minimize the timing error.

Benefits of technology

It effectively compensates for the nonlinearity of the phase interpolator, improves the stability of the clock recovery circuit and the accuracy of data recovery, and enhances the performance of the receiver, especially under high symbol rate and complex channel conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120729348B_ABST
    Figure CN120729348B_ABST
Patent Text Reader

Abstract

The present disclosure relates to dynamic timing loop gain for compensating for phase interpolator nonlinearity. An integrated circuit transceiver having a digital timing recovery loop with phase interpolation can incorporate dynamic loop gain to compensate for nonlinearity of the phase interpolation. An illustrative integrated receiver circuit includes a phase interpolator, a sampling element, a timing error estimator, and a feedback circuit. The phase interpolator provides a sampling signal by applying a phase shift to a clock signal in response to a phase control signal. The sampling element produces a digital receive signal by sampling an analog receive signal in accordance with the sampling signal; the timing error estimator produces a timing error signal indicative of an estimated timing error of the sampling signal relative to the analog receive signal. The feedback circuit derives the phase control signal from the timing error signal using a scaling element configured to scale the estimated timing error by a scaling factor that depends on the phase control signal.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Digital communication occurs between a transmitting device and a receiving device over an intermediate communication medium (e.g., a fiber optic cable or insulated copper wire) having one or more designated communication channels (e.g., carrier wavelengths or frequency bands). Each transmitting device typically transmits symbols at a fixed symbol rate, while each receiving device detects a potentially corrupted sequence of symbols and attempts to reconstruct the transmitted data.

[0002] A “symbol” is a state or condition of a channel that persists for a fixed period of time, referred to as a “symbol interval.” For example, a symbol can be a voltage or current level, an optical power level, a phase value, or a particular frequency or wavelength. A change from one channel state to another is referred to as a symbol transition. Each symbol can represent (i.e., encode) one or more binary bits of data. Alternatively, data can be represented by a symbol transition or by a sequence of two or more symbols. The simplest digital communication link uses only one bit per symbol; a binary “0” is represented by one symbol (e.g., a voltage or current signal in a first range), while a binary “1” is represented by another symbol (e.g., a voltage or current signal in a second range).

[0003] Channel imperfections give rise to dispersion, which can cause each symbol to disturb its neighbors, resulting in inter-symbol interference (ISI). As the symbol rate increases, ISI can make it difficult for a receiving device to determine which symbols were transmitted in each interval, especially when such ISI is combined with additive noise.

[0004] The disclosure literature discloses many equalization and demodulation techniques for recovering digital data from a degraded received signal even in the presence of ISI. A key to such techniques is determining the correct sampling timing, as sampling timing directly impacts the signal-to-noise ratio of the discrete samples. There are many strategies for detecting and tracking the optimal sampling time, with varying degrees of tradeoff between simplicity and performance. Notable examples can be found, for example, in U.S. Patent 7,058,150 “High-Speed Serial Data Transceiver and Related Methods”; and U.S. Patent 10,892,763 “Second-order Clock Recovery Using Three Feedback Paths,” both of which are incorporated herein by reference. These examples are particularly notable for their use of a sampling clock phase interpolator.

[0005] Nonlinearity is a potential issue with the use of phase interpolators, and it becomes more challenging at higher clock frequencies. The authors have found that such nonlinearity can lead to oscillations and instability in the clock recovery circuit, resulting in a corresponding degradation of receiver performance. SUMMARY

[0006] Accordingly, disclosed herein are integrated circuit transceivers, receivers, and methods having digital timing recovery loops with phase interpolation and dynamic loop gain to compensate for nonlinearity of the phase interpolation. An illustrative integrated receiver circuit includes a phase interpolator, a sampling element, a timing error estimator, and a feedback circuit. The phase interpolator is configured to provide a sampling signal by applying a phase shift to a clock signal in response to a phase control signal. The sampling element is configured to produce a digital receive signal by sampling an analog receive signal in accordance with the sampling signal. The timing error estimator is configured to produce a timing error signal indicative of an estimated timing error of the sampling signal relative to the analog receive signal. The feedback circuit is configured to derive the phase control signal from the timing error signal, the feedback circuit including a scaling element configured to scale the estimated timing error by a scaling factor that depends on the phase control signal.

[0007] An illustrative clock recovery method includes providing a phase control signal to a phase interpolator to derive a sampling signal from a clock signal; sampling an analog receive signal in accordance with the sampling signal to obtain a digital receive signal; producing a timing error signal indicative of an estimated timing error of the sampling signal relative to the analog receive signal; and deriving the phase control signal from the estimated timing error, the deriving including scaling the estimated timing error by a scaling factor that depends on the phase control signal.

[0008] The foregoing circuits can be embodied as a semiconductor IP core that resides on a non-transitory information storage medium. The core can represent a circuit schematic using, for example, a hardware description language, or can represent semiconductor fabrication process mask patterns using, for example, GDSII or OASIS languages.

[0009] Each of the above can be implemented alone or in combination, and can be implemented with any one or more of the following features in any suitable combination: 1. a demodulator that extracts a transmitted symbol stream from a digital receive signal. 2. a histogram circuit configured to determine a relative probability of each phase control signal value. 3. an inversion element configured to use an inverse of the relative probability to determine a scaling factor. 4. the scaling factor is stored in a lookup table configured to receive the phase control signal. 5. the scaling factor is derived from contents of a lookup table configured to receive the phase control signal. 6. the scaling element is part of a first feedback path in a feedback circuit configured to minimize a phase component of an estimated timing error. 7. the feedback circuit includes a second feedback path configured to minimize a frequency component of the estimated timing error. 8. a fractional frequency phase-locked loop configured to generate the clock signal. 9. the scaling element is part of a first feedback path in a feedback circuit configured to minimize a phase component of an estimated timing error. 10. the feedback circuit includes an additional feedback path configured to derive a frequency division ratio error from the estimated timing error. BRIEF DESCRIPTION OF DRAWINGS

[0010] Figure 1 An illustrative network is shown.

[0011] Figure 2 is a block diagram of an illustrative switch.

[0012] Figures 3-6 An illustrative digital communication receiver with different clock recovery circuit configurations is shown.

[0013] Figure 7 is a block diagram of an illustrative non-linear compensation circuit.

[0014] Figure 8 is a block diagram of an illustrative decision feedback equalizer ("DFE").

[0015] Figure 9 is a block diagram of an illustrative parallelized DFE. DETAILED DESCRIPTION

[0016] While specific embodiments are given in the drawings and above description, it is understood that they are not limiting of the present disclosure. On the contrary, they are provided as examples of implementations and equivalents and alternatives to claim the full scope of the disclosure.

[0017] For the purposes of context, Figure 1An illustrative network such as can be found in a data processing center is shown, in which multiple server racks 102-106 each contain multiple servers 110 and at least one“top of rack” (TOR) switch 112. The TOR switches 112 are connected to aggregator switches 114 for interconnection and connection to regional networks and the Internet. (As used herein, the term“switch” includes not only traditional network switches, but also routers, bridges, hubs, and other devices that forward network communication packets between ports.) Each of the servers 110 is connected to the TOR switch 112 by a network cable 120, which can carry signals at high symbol rates.

[0018] Figure 2 An illustrative switch 112 is shown having an application specific integrated circuit (ASIC) 202 that implements the packet switching functionality of the port connectors 204 coupled to line cards or“pluggable modules” 206. The pluggable modules 206 are coupled between the port connectors 204 and cable connectors 208 to improve communication performance through equalization and optional format conversion (e.g., between electrical and optical signals). The pluggable modules 206 can conform to any of a variety of pluggable module standards including SFP, SFP-DD, QSFP, QSFP-DD, and OSFP. Alternatively, the cable itself can have a connector that conforms to a pluggable module standard and contains pluggable module circuitry.

[0019] The pluggable modules 206 can each include a retimer chip 210 and a microcontroller chip 212 that controls the operation of the retimer chip 210 according to firmware and parameters that can be stored in non-volatile memory 214. The operating mode and parameters of the pluggable retimer modules 206 can be set via a two-wire bus (such as I2C or MDIO) that connects the microcontroller chip 212 to a host device (e.g., the switch 112). The microcontroller chip 212 responds to queries and commands received via the two-wire bus and responsively retrieves information from and saves information to the control registers 218 of the retimer chip 210.

[0020] The retimer chip 210 includes a host-side transceiver 220 coupled to a line- side transceiver 222 by a first-in-first-out (FIFO) buffer 224. Although only a single lane is shown in the figure, the transceivers can support multiple lanes of communication via multiple corresponding optical or electrical conductors. A controller 226 coordinates the operation of the transceivers according to control register contents, and can provide multiple communication phases according to a communication standard such as the Fibre Channel standard published by the American National Standards Institute Accredited Standards Committee INCITS, which provides phases for link speed negotiation (LSN), equalizer training, and normal operation.

[0021] The receiver portion of each transceiver can employ any of a number of equalization and demodulation techniques disclosed in the public literature for recovering digital data from a degraded received signal even in the presence of ISI. As previously noted, a key to such techniques is the determination of the correct sampling timing, since sampling timing directly impacts the signal-to-noise ratio of the discrete samples.

[0022] Figures 3-6 Various clock recovery methods that can be implemented by the illustrative integrated receiver circuits are shown. The illustrated receivers each employ a phase interpolator as part of the clock recovery circuit, and further employ dynamic gain in the digital timing loop to compensate for potential non-linearities of the phase interpolator.

[0023] Figure 3 The receiver of FIG. 3 includes an analog-to-digital converter 304 or other sampling element that samples the analog received signal 302 at sampling instants corresponding to transitions in the sampled signal 305, thereby providing a digital received signal to a demodulator 306. The demodulator 306 can apply equalization and symbol detection using, for example, matched filters, decision feedback equalizers, maximum likelihood sequence estimators, or other suitable techniques, for extracting the stream of digital symbols conveyed by the analog received signal. The resulting detected symbol stream 308 can be provided as a parallelized symbol stream for handling by "on-chip" circuitry (e.g., FIFO buffering, error correction, and retransmission).

[0024] The demodulator 306 includes some form of timing error estimator for generating an estimated timing error signal 310. Any suitable design can be used for the timing error estimator, including, for example, a bang-bang phase detector or a proportional phase detector. Suitable timing error estimators are set forth in commonly owned U.S. Patent 10,447,509, “Precompensator-based quantization for clock recovery,” which is incorporated herein by reference in its entirety. Other suitable timing error estimators can be found in the public literature, including, for example, Mueller, “Timing Recovery in Digital Synchronous Data Receivers,” IEEE Transactions on Communications, May 1976, Vol. 24, No. 5, and Musa, “High-speed Baud-Rate Clock Recovery,” University of Toronto thesis, 2008.

[0025] In Figure 3 , a feedback circuit derives a phase control signal from the timing error signal 310 to control the phase interpolator 320 in a manner that statistically minimizes the timing error signal 310. A frequency error accumulator 312 integrates the timing error signal after it has been scaled by a frequency coefficient (K F ) to obtain a frequency offset signal. A multiplier or other scaling element 314 scales the phase error with a dynamic phase coefficient (K P ). An adder 316 adds the scaled phase error to the frequency offset signal. A filter 318 operates on the output of the adder 316 to obtain the control signal for the phase interpolator 320. In Figure 3 , the filter 318 is shown as an accumulator, but other filter implementations would also be suitable.

[0026] The phase interpolator 320 also receives a clock signal from a phase-locked loop (PLL) 322. The phase control signal causes the phase interpolator 320 to generate the sampling signal by adjusting the phase of the clock signal in a manner that minimizes the expected value of the timing error signal. In other words, the control signal compensates for both the frequency offset component and the phase component in the estimated timing error of the clock signal relative to the analog received signal 302, thereby phase aligning the sampling signal 305 with the data symbols in the analog received signal 302. Various suitable phase interpolator implementations can be found in the public literature. See, for example, U.S. Patent 7,058,150, "High-Speed Serial Data Transceiver and Related Methods," by Buchwald et al.

[0027] The clock signal generated by the PLL 322 is a multiplied version of the reference clock signal from a reference oscillator 324. A voltage-controlled oscillator (VCO) 326 supplies the clock signal to both the phase interpolator 320 and a counter 328, which divides the frequency of the clock signal by a constant modulus N. The counter supplies the divided clock signal to a phase frequency detector (PFD) 330. The PFD 330 can use a charge pump (CP) as part of determining which input (i.e., the divided clock signal or the reference clock signal) has an earlier or more frequent transition than the other. A low-pass filter 332 filters the output of the PFD 330 to provide a control voltage for the VCO 326. The filter coefficients are selected so that the divided clock becomes phase aligned with the reference oscillator.

[0028] It should be noted that for at least some contemplated uses, the reference clock used by the receiver will typically drift relative to the reference clock used by the transmitter, and can differ by hundreds of ppm. In Figure 3 In embodiments of the '150 patent, the resulting frequency offset between the clock signal output of the PLL and the analog data signal can require continuous phase rotation by the phase interpolator 320 to correct. This mode of operation places stringent requirements on the linearity of the phase interpolator 320 over its entire tuning range, as the interpolator will repeatedly cycle through each of the phase interpolations during continuous rotation.

[0029] The presence of phase interpolation nonlinearity can be visualized as a dependence of the feedback loop gain on the current phase interpolation setting of the phase interpolator, resulting in a clock recovery circuit that is more or less sensitive to timing errors depending on the control signal value. In the presence of continuous phase rotation, this sensitivity can be observed as a change in probability of each phase interpolation setting. (Continuous phase rotation can be introduced by, for example, adding a small offset bias to the phase error signal 310, adjusting the frequency offset stored by the frequency accumulator 312, having the adder 316 introduce an additional offset, or adjusting the PLL to introduce an actual frequency offset). Those phase interpolation settings with reduced sensitivity to phase error will exhibit higher probability relative to those with enhanced sensitivity and thus accelerated response to small phase errors. Once the relative probabilities of each phase control signal value are determined, the feedback loop gain can be dynamically adjusted to compensate for the phase interpolator nonlinearity. In Figure 3 , the dynamic adjustment is provided by a lookup table 340 that provides a phase coefficient K P depending on the control signal value of the phase interpolator.

[0030] Figure 3 The receiver in further includes a histogram circuit 342 for determining the relative probabilities of the phase interpolation settings when continuous phase rotation is introduced. The histogram circuit 342 can intermittently sample the values of the control signal over a sufficiently long window of time to ensure that the histogram statistics are representative of the relative probabilities. A histogram count is obtained for each possible value of the control signal. An inversion element 344 converts each histogram count to a corresponding phase coefficient K P or, optionally, to a normalized scaling factor for the phase coefficient K P , storing the result in a location of the lookup table 340 associated with that value of the control signal. A firmware programmed microcontroller can perform the function of the inversion element 344.

[0031] Figure 4 A receiver module is provided that embodies an alternative clock recovery circuit configuration. The receiver module retains the analog-to-digital converter 304 for sampling the analog receive signal 302 and providing a digital receive signal to the demodulator 306. As previously described, the demodulator includes a timing error estimator that generates the timing error signal 310, and a feedback circuit having a feedback path with a dynamic phase coefficient (K P ) accumulator 314 and a filter. However, Figure 3 The frequency offset accumulator 312 in the embodiment is replaced by another feedback path that couples the timing error signal 310 to a fractional frequency division phase locked loop 422 to correct for frequency offset separately from the phase interpolator 320. This feedback path includes a frequency division ratio scaling coefficient (K D) and a divide ratio error accumulator 412 that supplies a divide ratio control signal to a fractional divide phase locked loop 422.

[0032] The fractional divide phase locked loop 422 is used in place of the original phase locked loop 322 to provide finer granularity frequency control of the clock signal supplied to the phase interpolator 322. The divide ratio control signal adjusts the frequency offset of the clock signal relative to the data in the analog receive signal 302, essentially reducing the rate of phase rotation required from the phase interpolator 320.

[0033] Figure 3 And Figure 4 A comparison of the phase locked loop 322 and the fractional divide phase locked loop 422 shows that both employ a PFD / CP 330 (to compare the divided clock signal to a reference clock), a low pass filter 332 (to filter the error to reduce noise), and a voltage controlled oscillator 326 (to supply the output clock signal). The fractional divide phase locked loop 422 uses a multi-modulus divider 428 to divide the output clock signal instead of using the fixed modulus divider 328, dividing by N or N+1 depending on whether the modulus select signal is asserted at the end of the count period (or at the beginning of the count period or at any point during the count period in alternative embodiments). A delta-sigma modulator (DSM) 429 converts the divide ratio control signal to pulses of the modulus select signal. The pulse density controls which fractional value between N and N+1 is implemented by the divider, enabling very fine control of the clock frequency supplied to the interpolator 320.

[0034] We now turn to Figure 5 which shows a receiver including Figure 3 the feedback path in the embodiment and Figure 4 the feedback path in the embodiment. To ensure that the divide ratio error accumulator 412 operates in concert with the frequency offset accumulator 512, the frequency offset error accumulator 312 can be modified to include a leak coefficient K L . In the modified accumulator 512, the frequency offset signal is multiplied by (1 - K L ) in each integration period. The leak coefficient (K L) indicates gradual memory loss, although it enables the feedback path to provide fast response, it causes the frequency offset signal to tend toward zero on a longer time scale. The divide ratio error accumulator 412, in combination with the low pass filter 332 of the phase locked loop 422, operates on a longer time scale to overcome the memory loss of the modified accumulator 512. Under steady state or slowly changing conditions, the frequency offset correction is provided by the divide ratio error accumulator 412, thereby reducing the rate of continuous phase rotation that the phase interpolator 320 might otherwise need to provide. Under conditions where the frequency offset is changing more rapidly, the modified frequency offset accumulator 512 provides more transient correction.

[0035] Figure 6 Yet another feedback circuit configuration is shown, in which the modified frequency offset accumulator 512 is further modified to be a second order filter 612, and the adder 616 combines the output of the second order filter 612 with the output of the phase error filter 318 to form the phase interpolator control signal. The divide ratio error accumulator 412 remains unchanged from the embodiment of Figure 5 This implementation can enable different filters and accumulators to be driven at different clock frequencies, which can be potentially advantageous for some applications.

[0036] Figure 7 A more specific illustrative implementation of the non-linear compensation circuit (elements 340-344) is shown. In Figure 7 The histogram circuit 342 contains a memory 702 that can be reset to a clear state by a CLR signal from the controller 710. When an address is supplied to the memory 702, the memory responsively provides read data (RD) to the adder 704 representing the contents of the selected address location, which the adder increments by one. A saturation element 706 prevents the contents from rolling over a maximum value that the memory 702 can store, for example, detecting whether the incremented value is all zeros, in which case the saturation element can set the memory location contents to all ones. The memory 702 can be configured to accept the value from the saturation element 706 as write data (WD) for storage at the address location.

[0037] The multiplexer 708 supplies the address to the memory 702. When the histogram circuit 342 is collecting statistics, the multiplexer 708 is set to supply the phase interpolator phase setting control signal value as the address to the memory 702, enabling the memory 702 to increment the contents of the location corresponding to the current value of the control signal. Once the collection window is closed, the multiplexer 708 is set to forward any address supplied by the controller 710, enabling the controller to access the histogram counts for processing or storage elsewhere.

[0038] The controller 710 can supply a clock signal to the memory 702 during the collection window. The memory clock signal can be derived from the sampling clock signal, with the counter 712 asserting an enable signal EN for the length of the collection window (e.g., 2 N N sample clock signal counts), where N is large enough to provide reliable statistical data collection, without being so large that any histogram count saturates. Some implementations can implement adjustments to N, such as decreasing N if saturation is detected, or increasing N if any histogram count is below a predetermined threshold at the end of the collection window. When the collection window is closed, i.e., when the counter 712 reaches 2 N N, the gate 714 blocks the memory clock signal. To reduce correlation, the frequency divider 716 can reduce the sampling clock frequency by a factor of 2 L N, so that the phase interpolator control signal values are sampled apart, such as once every 2 L N symbol intervals.

[0039] Once the collection window is closed, the controller 710 can systematically retrieve the histogram counts and store the corresponding reciprocal (multiplicative inverse) values provided by the inversion element 344 in the lookup table 340. The illustrated lookup table 340 includes a memory 720 that receives an address from a multiplexer 728. The multiplexer 728 is initially set to provide an address location selected by the controller 710, which can further provide a read / write signal 724 to store the inverse histogram count or other multiplicative K P factor. Once the lookup table 720 has been populated, the multiplexer 728 is set to provide the current value of the phase interpolator control signal PHS as the address, causing the lookup table to provide the corresponding K P factor. The multiplier 726 can determine the product of the selected K P factor and the nominal K P coefficient value, which product is used as the dynamic phase coefficient K P_dyn .

[0040] In a contemplated variation, the memory 702 can be employed simultaneously for both the collection of histogram counts and as a lookup table during normal operation of the receiver. Once the histogram count collection process is complete, the incrementing circuitry is disabled. The controller 710 can store the K P factors derived from the histogram counts, or alternatively, the K P factors can be derived from the histogram counts on an as-needed basis in accordance with the employment of the inversion element.

[0041] Figure 8 and Figure 9 An illustrative receiver embodiment is shown to provide additional details of implementations of the demodulator 306, as well as insight into adjusting clock recovery circuitry for parallelism.

[0042] Figure 8 An illustrative digital receiver is shown that includes a continuous-time linear equalizer ("CTLE") 801 to attenuate out-of-band noise and optionally provide some spectral shaping to improve the response to high frequency components of the received signal. An ADC 304 is provided to digitize the received signal, and a digital filter (also known as a feed-forward equalizer or "FFE") 802 performs further equalization to further shape the overall channel response of the system and minimize the impact of pre-cursor ISI on the current symbol. As part of the shaping of the overall channel response, the FFE 802 can also be designed to shorten the channel response of the filtered signal while minimizing any accompanying noise enhancement.

[0043] A summer 803 subtracts an optional feedback signal from the output of the FFE 802 to minimize the impact of post-cursor ISI on the current symbol, resulting in an equalized signal that is coupled to a decision element ("slicer") 804. The decision element 804 includes one or more comparators that compare the equalized signal to corresponding decision thresholds to determine which constellation symbol the value of the signal most closely corresponds to for each symbol interval. Here, the equalized signal can also be referred to as the "combined signal" herein.

[0044] The decision element 804 accordingly produces a sequence of symbol decisions (denoted as A k , where k is a time index). In certain contemplated embodiments, the signal constellation is a bipolar (non-return-to-zero) constellation representing -1 and +1, requiring the use of only one comparator with a decision threshold of zero. In certain other contemplated embodiments, the signal constellation is PAM4 (-3, -1, +1, +3), requiring the use of three comparators with decision thresholds of -2, 0, and +2, respectively. (For generality, the units used to express the symbols and thresholds are omitted, but can be assumed to be volts for purposes of explanation. In practice, a scaling factor will be employed.)

[0045] A feedback filter ("FBF") 805 uses a series of delay elements (e.g., latches, flip-flops, or registers) that store recent output symbol decisions (A k-1 , A k-2 ,..., A k-N , where N is the number of filter coefficients f i ) to derive a feedback signal. Each stored symbol is multiplied by a corresponding filter coefficient f i , and the products are combined to obtain the feedback signal.

[0046] In addition, we note here that the receiver also includes a filter coefficient adaptation unit, but such considerations are addressed in the literature and are well known to those skilled in the art. However, we note here that at least some contemplated embodiments include one or more additional comparators in decision element 804 to be used to compare the combined signal to one or more of the symbol values, thereby providing an error signal that can be used for timing recovery and / or coefficient adaptation.

[0047] As the symbol rate increases to the gigahertz range, it becomes increasingly difficult for the ADC 304 and demodulator 306 components to perform the operations they require entirely within each symbol interval, at which point it becomes advantageous to parallelize their operations. Parallelization generally involves the use of multiple components that share the workload by taking turns and thereby provide more time for each of the individual components to complete its operations. Such parallel components are driven by a set of staggered clock signals. For example, four-fold parallelization employs a set of four clock signals, each having a frequency of one quarter of the symbol rate, such that in the set of staggered clock signals, each symbol interval contains only one upward transition. While four-fold parallelization is used here for purposes of discussion, the actual degree of parallelization can be higher, e.g., 8-fold, 16-fold, 32-fold, or 64-fold. Moreover, the degree of parallelization is not limited to powers of two.

[0048] Figure 9 An illustrative receiver is shown with a parallelized equalizer implementation, including an optional feedback filter for DFE. As with the implementation of Figure 8 As with the implementation of FIG. 8, the CTLE 801 filters the channel signal to provide a received signal that is provided in parallel to an array of analog-to-digital converters (ADC0-ADC3). Each of the ADC elements is provided with one of the corresponding ones of the staggered clock signals. The clock signals have different phases, causing the ADC elements to take turns sampling and digitizing the received signal, such that at any given time, only one of the ADC element outputs is transitioning.

[0049] An array of FFEs (FFE0-FFE3), each forming a weighted sum of the ADC element outputs. The weighted sums employ filter coefficients that are cyclically shifted relative to one another. FFE0 operates on the held signals from ADC3 (an element that operates before CLK0), ADC0 (an element that responds to CLK0), and ADC1 (an element that operates after CLK0), such that during the assertion of CLK2, the weighted sum produced by FFE0 is the same as the weighted sum produced by FFE 802 Figure 8) correspond. FFE1 operates on the hold signals from ADC0 (an element operating before CLK1), ADC1 (an element responsive to CLK1), and ADC2 (an element operating after CLK1), so that during the assertion of CLK3, the weighted sum corresponds to the output of FFE 802. And the operation of the remaining FFEs in the array follow the same pattern with the relevant phase shift. In practice, the number of filter taps can be smaller, or the number of elements in the array can be larger, in order to provide a longer effective output window.

[0050] As with the receiver of Figure 8 , the adder can combine the output of each FFE with the feedback signal to provide an equalized signal for the corresponding decision element. Figure 9 An array of decision elements (slicer 0 through slicer 3) is shown, each operating on an equalized signal derived from the output of a corresponding FFE. As with the decision elements of Figure 8 , the illustrated decision elements employ comparators to determine which symbol the equalized signal most likely represents. Decisions are made when the corresponding FFE output is valid (e.g., slicer 0 operates when CLK2 is asserted, slicer 1 operates when CLK3 is asserted, etc.). Preferably, the decisions are provided in parallel on an output bus to enable a lower clock rate for subsequent operations.

[0051] An array of feedback filters (FBF0-FBF3) operates on the preceding symbol decisions to provide feedback signals for the adders. As with the FFEs, the inputs to the FBFs are circularly shifted, and valid outputs are provided only when the inputs correspond to the contents of FBF 805( Figure 8 ) and are consistent with the time window of the corresponding FFE. In practice, the number of feedback filter taps can be fewer than shown, or the number of array elements can be larger, in order to provide a longer valid output window.

[0052] As with the decision elements of Figure 8 , the decision elements in Figure 9 may each employ additional comparators to provide timing recovery information, coefficient training information, and / or pre-computations to unfold one or more taps of the feedback filter. In Figure 9 , the digital timing circuitry is also parallelized, with timing error estimator 910 accepting symbol decisions and equalized signals in parallel to determine timing error signal 310( Figure 3) of the parallelized version of the equalizer 900. The set of timing loop filters 912 implement the feedback circuitry discussed previously to provide control signals for the phase interpolator 920 and a divide ratio control signal for the PLL. The phase interpolator 920 operates similarly to the phase interpolator 320 to convert the PLL clock signal into a set of interleaved clock signals having uniformly spaced phases and symbol-aligned transitions. A set of delay lines (DLO-DL3) are provided for fine tuning the individual clock phases relative to one another as needed, e.g., to compensate for different propagation delays of the individual ADC elements.

[0053] The delay lines can be individually adjusted by a clock skew adjustment circuit 944 based on parameters from a controller 942. The controller 942 can optimize the clock skew adjustment settings based on reliability indicators from a monitoring circuit. In Figure 9 In one implementation, the monitoring circuit is a margin calculator 940 that computes the minimum difference between the equalized signal and the decision threshold (or equivalently, the maximum error between the equalized signal and the nominal symbol values). Clock skew adjustment is described in greater detail in commonly owned U.S. Application 16 / 836,553 (now U.S. Patent No. 10,992,501), “Eye Monitor for Parallelized Digital Equalizers,” filed March 31, 2020, which is incorporated herein by reference in its entirety.

[0054] Most integrated circuit devices that will employ receiver and clock recovery circuitry have become so complex that it is impractical for an electronic device designer to design the integrated circuit device from scratch. Instead, the electronic device designer relies on pre-defined modular units of integrated circuit layout design that are arranged and interfaced as needed to implement the various functions of the desired device. Each modular unit has a defined interface and behavior that has been verified by its creator. Although it can take a significant amount of time and investment to create each modular unit, the availability of modular unit re-use and further development greatly reduces product cycle time and enables better products. The pre-defined units can be organized hierarchically, with a given unit containing one or more lower-level units, and in turn being contained in a higher-level unit. Many organizations have libraries of such pre-defined modular units for sale or license, including, for example, embedded processors, memories, interfaces for different bus standards, power converters, frequency multipliers, sensor transducer interfaces, etc. The pre-defined modular units are also referred to as cells, blocks, cores, and macros, with these terms having different meanings and variations (“intellectual property (IP) cores,” “soft macros”), but are often used interchangeably.

[0055] Modular units can be expressed in different ways, for example, in the form of hardware description language (HDL) files, or in the form of fully routed designs that can be directly printed to create a series of fabrication process masks. Fully routed design files are typically process-specific, meaning that additional design work is usually required to port a modular unit to a different process or foundry. Modular units in HDL form require subsequent synthesis, placement, and routing steps to implement, but they are process-independent, meaning that different foundries can apply their preferred automatic synthesis, placement, and routing processes to implement the unit using a wide range of fabrication processes. Due to their higher level of representation, HDL units can be more amenable to modification and use of variable design parameters, while fully routed units can provide better predictability in area requirements, reliability, and performance. While there is no fixed rule, digital module designs are more commonly specified in HDL form, while analog and mixed-signal units are more commonly specified in lower-level physical descriptions. In any case, such semiconductor IP cores can be retained in a design database, which resides on a non-transitory information storage medium. Once a device has been fully designed, commercially available software can convert the semiconductor intellectual property core and other integrated circuit components into semiconductor mask patterns, which are also stored on a non-transitory information storage medium. Thereafter, these patterns can be transferred to various processing units on the appropriate assembly line of an integrated circuit foundry.

[0056] Numerous alternative forms, equivalents, and modifications will become apparent to those skilled in the art once the above disclosure is fully appreciated. For example, the above description focuses on the use of an integral-based accumulator, but other recursive filter or moving average filter implementations that provide a low-pass filter response can also be employed. It is intended that the claims be interpreted to embrace all such alterations, equivalents, and modifications as falling within the scope of the appended claims.

Claims

1. An integrated receiver circuit, comprising: a phase interpolator configured to provide a sampling signal by applying a phase shift to a clock signal in response to a phase control signal; a sampling element configured to produce a digital receive signal by sampling an analog receive signal in accordance with the sampling signal; a timing error estimator configured to produce a timing error signal indicative of an estimated timing error of the sampling signal relative to the analog receive signal; and a feedback circuit configured to derive the phase control signal from the timing error signal, the feedback circuit comprising a scaling element configured to scale the estimated timing error by a scaling factor dependent on the phase control signal.

2. The integrated receiver circuit of claim 1, further comprising a demodulator to derive a digital symbol stream from the digital receive signal.

3. The integrated receiver circuit of claim 1, further comprising a histogram circuit configured to determine a relative probability of each phase control signal value.

4. The integrated receiver circuit of claim 3, further comprising an inversion element configured to determine the scaling factor using an inverse of the relative probability.

5. The integrated receiver circuit of claim 1, wherein the scaling factor is stored in a lookup table configured to receive the phase control signal.

6. The integrated receiver circuit of claim 1, wherein the scaling factor is derived from contents of a lookup table configured to receive the phase control signal.

7. The integrated receiver circuit of claim 1, wherein scaling element is part of a first feedback path in the feedback circuit, the first feedback path configured to minimize a phase component of the estimated timing error, the feedback circuit further comprising a second feedback path configured to minimize a frequency component of the estimated timing error.

8. The integrated receiver circuit of claim 1, further comprising a fractional- division phase-locked loop configured to generate the clock signal, wherein the scaling element is part of a first feedback path in the feedback circuit, the first feedback path configured to minimize a phase component of the estimated timing error, the feedback circuit further comprising an additional feedback path configured to derive a division ratio error from the estimated timing error.

9. A method of clock recovery, the method comprising, in an integrated receiver circuit: providing a phase control signal to a phase interpolator to derive a sampling signal from a clock signal; sampling an analog receive signal in accordance with the sampling signal to obtain a digital receive signal; producing a timing error signal indicative of an estimated timing error of the sampling signal relative to the analog receive signal; and ​ ​ deriving the phase control signal from the estimated timing error, the deriving including scaling the estimated timing error by a scaling factor that depends on the phase control signal.

10. The clock recovery method of claim 9, further comprising demodulating the digital receive signal to extract a stream of digital symbols.

11. The clock recovery method of claim 9, further comprising determining a relative probability of each phase control signal value.

12. The clock recovery method of claim 11, further comprising using an inverse of the relative probability to determine the scaling factor.

13. The clock recovery method of claim 9, wherein the phase control signal is used to retrieve the scaling factor from a lookup table.

14. The clock recovery method of claim 9, wherein the phase control signal is used to retrieve a relative probability from a lookup table, and wherein the method further comprises deriving the scaling factor from the relative probability.

15. The clock recovery method of claim 9, wherein the scaling is performed in a first feedback path to minimize a phase component of the estimated timing error, and wherein the deriving the phase control signal employs a second feedback path to minimize a frequency component of the estimated timing error.

16. The clock recovery method of claim 9, wherein the scaling is performed in a first feedback path to minimize a phase component of the estimated timing error, and wherein the method further comprises providing a fractional division ratio error to a fractional division phase-locked loop, the fractional division phase-locked loop providing the clock signal, the fractional division ratio error being derived from the estimated timing error.

17. A non-transitory information storage medium having a semiconductor intellectual property core for generating circuitry, the circuitry comprising: a phase interpolator configured to provide a sample signal by applying a phase shift to a clock signal in response to a phase control signal; a sampling element configured to produce a digital receive signal by sampling an analog receive signal according to the sample signal; a timing error estimator configured to produce a timing error signal indicative of an estimated timing error of the sample signal relative to the analog receive signal; and a feedback circuit configured to derive the phase control signal from the timing error signal, the feedback circuit including a scaling element configured to scale the estimated timing error by a scaling factor that depends on the phase control signal.

18. The non-transitory information storage medium of claim 17, wherein the circuitry further comprises a histogram circuit configured to determine a relative probability of each phase control signal value.

19. The non-transitory information storage medium of claim 18, wherein the circuitry further comprises an inversion element configured to use an inverse of the relative probability to determine the scaling factor.

20. The non-transitory information storage medium of claim 17, wherein the circuitry further comprises a lookup table configured to retrieve the scaling factor in response to the phase control signal.

Citation Information

Patent Citations

  • Precompensator-based quantization for clock recovery

    US10447509B1

  • Second-order clock recovery using three feedback paths

    US10892763B1

  • Eye monitor for parallelized digital equalizers

    US10992501B1

  • High-speed serial data transceiver and related methods

    US7058150B2

  • Spur and quantization noise cancellation for PLLS with non-linear phase detection

    CN113328742A