Digital signal processing method and system based on FPGA
By digitizing and analyzing the frequency of analog high-speed signals and dynamically adjusting the CDR loop parameters, FPGA-based fast locking and high-precision analysis are achieved, solving the problem of the trade-off between speed and accuracy in existing CDR designs and improving the versatility and efficiency of the test equipment.
Patent Information
- Application Number
- CN202511388952.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-01-09
AI Technical Summary
Existing FPGA-based CDR designs cannot simultaneously balance loop locking speed and measurement accuracy, nor can they be dynamically adjusted for different application scenarios, thus limiting the versatility and efficiency of test and measurement equipment.
A dynamic adaptive digital signal processing architecture is adopted. By digitizing the analog high-speed signal and performing feedforward frequency analysis, the frequency offset value is obtained. The CDR loop is preset and high bandwidth lock is started. The phase error signal is monitored and the proportional and integral coefficients of the loop filter are smoothly switched after the lock is achieved. The switch is made from high bandwidth mode to low bandwidth mode to achieve high-precision clock recovery and jitter analysis.
It achieves the integration of fast locking and high-precision measurement in a single operation, solving the problem of the trade-off between locking speed and measurement accuracy in traditional CDR designs, and improving testing efficiency and adaptability.
Smart Images

Figure CN121308932A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital signal processing, and more specifically, to a digital signal processing method and system based on FPGA. Background Technology
[0002] With the ever-increasing data transmission rates in modern communications, data centers, and semiconductor testing, real-time and accurate processing and analysis of high-speed digital signals have become crucial. Field-Programmable Gate Arrays (FPGAs), with their parallel processing capabilities, high throughput, and reconfigurable flexibility, have become an ideal platform for building such high-performance digital signal processing systems, especially in implementing critical Clock Data Recovery (CDR) functionality. The core objective of CDR technology is to extract the embedded clock from the serial data stream and use this clock to make accurate decisions about the data, which is fundamental to ensuring signal integrity.
[0003] However, current FPGA-based CDR designs generally face an inherent and irreconcilable contradiction: the trade-off between loop locking speed and measurement accuracy. Traditional CDR designs typically employ a phase-locked loop (PLL) structure with a fixed loop bandwidth, where the proportional (Kp) and integral (Ki) coefficients of the loop filter are fixed during design. If a high loop bandwidth is chosen to pursue fast locking, the system's ability to suppress jitter will deteriorate, resulting in insufficient accuracy of the recovered clock, failing to meet the stringent requirements of precise jitter component decomposition (such as random jitter and deterministic jitter) in the R&D verification phase. Conversely, if a low loop bandwidth is chosen to achieve high-precision measurement, although it can effectively filter out jitter and obtain a high-quality recovered clock, its locking process will become extremely slow, unable to meet the extreme pursuit of test efficiency in scenarios such as automated test equipment (ATE), and even difficult to quickly reconstruct the connection when the link state changes (such as hot-plugging).
[0004] This singular, static design paradigm makes existing CDR solutions unable to simultaneously address the diverse needs of different application scenarios. Therefore, current technology lacks an intelligent CDR architecture capable of dynamically adjusting its characteristics according to the working stage, and cannot seamlessly transition from a fast locking mode to a high-precision analysis mode within a single operation. This severely limits the versatility and efficiency of test and measurement equipment, necessitating an innovative technical solution to address this challenge. Summary of the Invention
[0005] To address the aforementioned problems in the prior art, according to one aspect of this application, an FPGA-based digital signal processing method is provided, comprising:
[0006] The acquired analog high-speed signal is digitized and subjected to feedforward frequency analysis to obtain the frequency offset value;
[0007] Based on the frequency offset value, the real-time digital signal stream is subjected to CDR loop preset and high-bandwidth lock-in start to obtain the phase error signal;
[0008] Hardware verification of the locked state is performed on the continuously monitored phase error signal to obtain a lock achievement flag;
[0009] In response to the lock attainment flag being high, the proportional coefficient and integral coefficient of the high-level triggered loop filter are smoothly switched from high-bandwidth mode to low-bandwidth mode based on a multi-step interpolation state machine to obtain a high-precision recovery clock and time interval error measurement stream;
[0010] Based on high-precision recovery of clock and time interval error measurement stream, data deserialization and jitter analysis are performed on real-time digital signal stream to obtain deserialized data output and jitter analysis results.
[0011] According to another aspect of this application, an FPGA-based digital signal processing system is provided, comprising:
[0012] The signal digitization frequency analysis module is used to perform signal digitization and feedforward frequency analysis on the acquired analog high-speed signals to obtain the frequency offset value;
[0013] The phase error signal generation module is used to perform CDR loop preset and high bandwidth lock-on activation on the real-time digital signal stream based on the frequency offset value to obtain the phase error signal.
[0014] The hardware verification module is used to perform hardware verification of the lock state of the continuously monitored phase error signal to obtain a lock achievement flag.
[0015] The mode switching module is used to smoothly switch the proportional coefficient and integral coefficient of the high-level triggered loop filter from high-bandwidth mode to low-bandwidth mode based on a multi-step interpolation state machine in response to the lock achievement flag being high-level, so as to obtain a high-precision recovery clock and time interval error measurement stream.
[0016] The deserialization jitter analysis module is used to perform data deserialization and jitter analysis on real-time digital signal streams based on high-precision recovery of clock and time interval error measurement streams, so as to obtain the deserialized data output and jitter analysis results.
[0017] Compared with existing technologies, this application provides an FPGA-based digital signal processing method and system that addresses the challenge of balancing locking speed and measurement accuracy in traditional CDR design by proposing a dynamically adaptive digital signal processing architecture. This architecture introduces a phased, variable bandwidth CDR strategy. Initially, signal frequency offset is rapidly estimated through feedforward frequency analysis, and a high-bandwidth loop is initiated based on this, achieving extremely fast locking and meeting the efficiency requirements of production testing and other scenarios. Once the system confirms stable locking through hardware logic, a multi-step interpolation state machine is triggered. This state machine smoothly and seamlessly switches the coefficients of the loop filter from the high-bandwidth mode to the low-bandwidth mode, avoiding the risk of lock loss during the switching process. Finally, in the stable low-bandwidth mode, the system can utilize a high-precision recovery clock for accurate jitter decomposition and data analysis. This integrates both rapid acquisition and high-precision measurement modes in a single operation, overcoming the limitations of static design in existing technologies. Attached Figure Description
[0018] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0019] Figure 1 This is a flowchart of an FPGA-based digital signal processing method according to an embodiment of this application.
[0020] Figure 2 This is a schematic diagram of the data flow of an FPGA-based digital signal processing method according to an embodiment of this application.
[0021] Figure 3 This is a flowchart of step 2 in the FPGA-based digital signal processing method according to an embodiment of this application.
[0022] Figure 4 This is a flowchart of step 4 in the FPGA-based digital signal processing method according to an embodiment of this application.
[0023] Figure 5 This is a block diagram of an FPGA-based digital signal processing system according to an embodiment of this application. Detailed Implementation
[0024] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0025] To address the challenge of balancing locking speed and measurement accuracy in traditional CDR design, this application proposes a digital signal processing method based on FPGA. Figure 1 This is a flowchart of an FPGA-based digital signal processing method according to an embodiment of this application. Figure 2 This is a schematic diagram of the data flow in an FPGA-based digital signal processing method according to an embodiment of this application. Figure 1 and Figure 2 As shown, the FPGA-based digital signal processing method according to an embodiment of this application includes: Step 1, performing signal digitization and feedforward frequency analysis on the acquired analog high-speed signal to obtain a frequency offset value; Step 2, based on the frequency offset value, performing CDR loop preset and high-bandwidth lock-in activation on the real-time digitized signal stream to obtain a phase error signal; Step 3, performing hardware verification of the lock-in state on the continuously monitored phase error signal to obtain a lock-in achievement flag; Step 4, in response to the lock-in achievement flag being high, smoothly switching the proportional coefficient and integral coefficient of the high-level triggered loop filter from high-bandwidth mode to low-bandwidth mode based on a multi-step interpolation state machine to obtain a high-precision recovered clock and time interval error measurement stream; Step 5, based on the high-precision recovered clock and time interval error measurement stream, performing data deserialization and jitter analysis on the real-time digitized signal stream to obtain deserialized data output and jitter analysis results.
[0026] In step 1, the acquired analog high-speed signal is digitized and subjected to feedforward frequency analysis to obtain the frequency offset value. It should be understood that in high-speed digital signal processing, the actual data rate of the signal being processed often deviates from its nominal rate due to factors such as inherent deviations of the transmitting crystal oscillator and temperature drift. Traditional clock data recovery loops typically start at the ideal nominal frequency and slowly search and pull towards the actual signal frequency through a feedback control mechanism. This frequency acquisition process constitutes the main part of the lock-in delay, which is unacceptable in production testing with extremely high efficiency requirements or in online monitoring scenarios requiring rapid response to link changes. Therefore, this application pre-executes digitization and feedforward frequency analysis of the acquired signal before formally starting the phase-locked loop (PLL) to directly and quickly calculate the difference between the actual signal frequency and the theoretical nominal frequency—the frequency offset value—through a single open-loop measurement. Using this offset value to pre-compensate the core oscillator of the loop allows its initial operating frequency to be directly set near the actual signal frequency, thereby greatly shortening or even eliminating the time-consuming frequency search phase and laying the foundation for subsequent rapid phase locking.
[0027] In one possible implementation, step 1, performing signal digitization and feedforward frequency analysis on the acquired analog high-speed signal to obtain a frequency offset value, includes: step 11, performing digital sampling and snapshot buffering on the acquired analog high-speed signal to obtain a time-domain sampled data block; step 12, performing a fast Fourier transform on the time-domain sampled data block to obtain a spectral amplitude vector; and step 13, performing spectral peak search and frequency offset calculation on the spectral amplitude vector to obtain the frequency offset value.
[0028] In the above implementation, step 1 is as follows: First, step 11 is executed. The analog high-speed signal refers to a continuously changing voltage waveform without digital decision, such as a non-return-to-zero (NRZ) serial data stream with a nominal data rate of 10 gigabits per second, whose voltage rapidly switches between two logic levels (high and low), but the signal itself is continuous in the time domain. This analog high-speed signal is first fed into the input of a high-speed analog-to-digital converter (ADC). The key parameters of this ADC are pre-configured according to the characteristics of the signal under test. To ensure distortion-free capture of the entire spectral information of the signal, its sampling rate must be much higher than the Nyquist frequency of the signal. For example, for a 10 gigabits per second NRZ signal with a fundamental frequency of 5 gigahertz, a high-speed ADC with a sampling rate of 40 gigabits per second (GSa / s) can be configured, and its quantization bit width can be set to 10 bits to ensure sufficient dynamic range and quantization accuracy. Inside the field-programmable gate array (FPGA), a data acquisition controller is used to manage the entire sampling and buffering process. The controller is implemented in the FPGA using a hardware description language. Its core consists of a state machine and an address counter. When an externally or internally generated snapshot trigger signal becomes active, the controller's state machine transitions from an idle state to a capture state. In the capture state, the analog-to-digital converter, driven by the sampling clock, continuously quantizes and samples the input high-speed analog signal, generating a continuous 10-bit wide parallel digital stream. The data capture controller continuously captures a preset number of samples from this digital stream and writes them sequentially to a pre-designated memory cell. This memory cell is a dual-port block random access memory (BRAM) within the FPGA, with its depth N preset to a specific value, an integer power of 2, to accommodate subsequent fast Fourier transform processing; for example, it can be set to 4096. The address counter in the data capture controller starts from zero and increments with each sampling clock cycle, providing an address for each write operation. When the address counter reaches 4095 and the last write operation is completed, the data capture controller stops writing and switches the state machine back to the idle state. Simultaneously, it generates a buffer full flag to notify subsequent processing modules that the data is ready. At this point, all 4096 10-bit sampling points stored in the block random access memory together constitute the time-domain sampling data block. It is worth mentioning that the sampling rate is set based on the bandwidth characteristics of the signal under test, the quantization bit width is set based on the required measurement accuracy, and the storage depth is set based on the desired frequency resolution and hardware resource limitations.
[0029] Next, step 12 is executed. Once the buffer full flag is valid, a control logic is triggered to sequentially read the time-domain sampled data blocks from memory and stream them into a dedicated Fast Fourier Transform (FFT) processing engine. The FFT processing engine is implemented in the FPGA by instantiating a pre-designed IP hard core. The parameters of this IP core are precisely configured during the design phase to match the data processing requirements. For example, its transform point number N is set to 4096 to correspond to the length of the time-domain sampled data block; the input data bit width is set to 10 bits; and its internal architecture is configured in pipelined mode to support high-throughput real-time data stream processing. After receiving all 4096 time-domain sample points, the FFT engine performs an N-point Fast Fourier Transform operation. Essentially, this operation replaces the signal's basis from the time basis with a set of orthogonal sine and cosine function bases, and the result is a complex spectral vector of length 4096. Each element in this vector is a complex number, consisting of a real part and an imaginary part, which together encode the signal amplitude and phase information at the corresponding frequency band. This complex spectrum vector, output by the FFT engine, is then streamed into an amplitude calculation unit. This unit can be efficiently implemented in the FPGA using a coordinate rotation digital computer IP core, specifically designed for performing trigonometric functions, vector rotation, and modulus calculations. The amplitude calculation unit performs a modulo operation on each complex element in the input complex spectrum vector, which has the form a + jb, i.e., calculates its distance to the origin, using the formula (a...jb ... 2 +b 2 The arithmetic square root of the signal is used. This process converts the amplitude and phase information of each complex point into a single real value, representing the energy intensity of the signal at that frequency. After this unit, the complex vector of length 4096 is converted into a real amplitude vector of length 4096. Finally, data simplification is performed using the conjugate symmetry of the real signal spectrum. Since the input time-domain sampled data block is a real sequence, its Fourier transform result has symmetry, that is, the spectrum from the Nyquist frequency to the sampling frequency is a mirror image of the first half of the spectrum. Therefore, to save resources and time in subsequent processing, only the first half of the spectrum needs to be retained. A simple control logic is used to truncate the elements from index 0 (DC component) to index N / 2 (Nyquist frequency) in the real amplitude vector. For the example of N = 4096, this means retaining 2049 amplitude values from index 0 to 2048. This truncated real vector of length 2049 is the final output spectrum amplitude vector.
[0030] Then, in step 132, in one possible implementation, step 13, performing spectral peak search and frequency offset calculation on the spectral amplitude vector to obtain the frequency offset value, includes: step 131, inputting the spectral amplitude vector into a peak search engine to obtain the peak amplitude and its corresponding peak index; step 132, inputting the peak index and FFT parameters into a frequency calculation unit to obtain the actual fundamental frequency of the signal; step 133, inputting the nominal data rate into a nominal frequency calculation unit to obtain the theoretical fundamental frequency, wherein the theoretical fundamental frequency is half of the nominal data rate; step 134, inputting the actual fundamental frequency and the theoretical fundamental frequency into a difference calculation unit to obtain the frequency offset value.
[0031] Specifically, step 131 is executed first. The peak search engine is a dedicated circuit implemented in hardware logic within the FPGA. Its input is a spectral amplitude vector. The core architecture of this engine includes a comparator, two registers (one for storing the current maximum amplitude value, and the other for storing the corresponding index), and a control state machine. When the spectral amplitude vector is fed in as streaming data, the state machine initiates the search process. In the first clock cycle, the first element of the vector, with index 0, and its index are loaded into the maximum value register and the index register. Starting from the second clock cycle, each subsequent vector element is compared with the value in the current maximum value register. If the value of the new element is larger, the comparator outputs a high level, triggering a register update, storing the value of the new element and its index into the maximum value register and the index register, respectively. This process continues until the entire vector of length 2049 has been traversed. After traversal, the maximum value register stores the peak amplitude, and the index register stores its corresponding peak index. For example, after the search, the obtained peak index might be 513. Next, step 132 is executed. This frequency calculation unit is an arithmetic logic module implemented in the FPGA. It receives the peak index 513 obtained in the previous step, along with preset FFT parameters, namely the sampling frequency (Fs = 40 GHz) and the number of FFT points (N = 4096). This unit first performs a preliminary calculation based on the formula: Frequency = Peak Index × (Fs / N), obtaining a discrete frequency estimate. To achieve higher frequency measurement accuracy, this unit also integrates an interpolation calculation submodule, employing a parabolic interpolation algorithm. This submodule reads the peak index k = 513 and its two adjacent indices, k-1 = 512 and k+1 = 514, corresponding to the spectral amplitude values, denoted as y(k-1), y(k), and y(k+1). Using these three points, a parabola can be fitted, and the precise position of the parabola's vertex can be calculated. The offset δ of the vertex position relative to the center point k can be calculated using the formula δ = 0.5 × (y(k-1) - y(k+1)) / (y(k-1) - 2y(k) + y(k+1)). This yields a more precise fractional index value, namely, the precise peak index = k + δ. For example, if the calculated δ is 0.2, then the precise peak index is 513.2. Finally, the frequency calculation unit uses this precise peak index to calculate the actual fundamental frequency of the signal: actual fundamental frequency = precise peak index × (Fs / N) = 513.2 × (40GHz / 4096) ≈ 5.01171875 gigahertz. Simultaneously, step 133 is executed. The nominal frequency calculation unit receives a preset nominal data rate determined according to the communication protocol standard; for a 10 gigabits per second non-return-to-zero (NRZ) signal in this example, this value is 10 gigabits per second.Since the theoretical fundamental frequency of the NRZ signal is half its data rate, this unit performs a division by 2 operation, which is implemented in hardware as a right shift operation, to obtain the theoretical fundamental frequency. The calculation result is: Theoretical fundamental frequency = 10Gbps / 2 = 5 GHz. Finally, in step 134, the difference calculation unit is a hardware subtractor. It receives the actual fundamental frequency of the signal (5.01171875 GHz) from the frequency calculation unit and the theoretical fundamental frequency (5 GHz) from the nominal frequency calculation unit. By performing the subtraction operation, the final output is obtained: Frequency offset value = 5.01171875 GHz - 5.0 GHz = +0.01171875 GHz = +11.71875 MHz.
[0032] In step 2, based on the frequency offset value, the real-time digitized signal stream is preset with a CDR loop and started with high-bandwidth locking to obtain the phase error signal. Correspondingly, after completing the feedforward frequency analysis and obtaining the accurate frequency offset value, the Clock Data Recovery (CDR) loop is about to enter the actual locking phase. When a traditional CDR loop starts, its internal oscillator begins operating from a fixed nominal frequency, requiring a lengthy frequency search and pulling process to synchronize with the input signal, which severely impacts locking efficiency. Simultaneously, the loop's dynamic characteristics (such as locking speed and jitter suppression capability) are determined by its loop bandwidth. To achieve optimal performance in different application scenarios, different loop bandwidths need to be used in the initial locking phase and the stable tracking phase. Therefore, CDR loop preset and high-bandwidth locking start based on the frequency offset value can utilize known frequency deviation information to directly preset the loop oscillator's operating frequency to a position close to the actual signal frequency, while simultaneously configuring the loop in high-bandwidth mode, thereby skipping the time-consuming frequency search phase and achieving rapid phase acquisition and locking.
[0033] In one possible implementation, Figure 3 This is a flowchart of step 2 in the FPGA-based digital signal processing method according to an embodiment of this application. Figure 3 As shown, step 2, based on the frequency offset value, performs CDR loop presetting and high-bandwidth lock-in activation on the real-time digital signal stream to obtain a phase error signal, including: step 21, performing NCO frequency pre-compensation based on the frequency offset value to obtain an initial NCO clock; step 22, loading a high-bandwidth coefficient onto the loop filter to obtain a configured loop filter; step 23, performing phase error detection based on the pre-compensated clock on the real-time digital signal stream based on the initial NCO clock to obtain an original phase error stream; step 24, performing high-bandwidth filtering and NCO phase fine-tuning on the original phase error stream based on the configured loop filter to obtain the phase error signal and the fine-tuned NCO clock.
[0034] In the above implementation, step 2 is as follows: First, step 21 is executed. In one possible implementation, step 21, performing NCO frequency pre-compensation based on the frequency offset value to obtain the initial NCO clock, includes: step 211, inputting the frequency offset value into the control word conversion unit to obtain an offset control word; step 212, adding the offset control word to the nominal frequency control word to obtain a compensated frequency control word; step 213, loading the compensated frequency control word into the NCO's frequency control register to obtain the initial NCO clock.
[0035] Specifically, step 211 is executed first. The control word conversion unit is an arithmetic logic module implemented in the FPGA using a hardware description language. Its core function is to convert a frequency value expressed in physical units (Hertz) into a unitless digital frequency control word that can be directly used by the numerically controlled oscillator (NCO). The conversion ratio is determined by the parameters of the NCO itself, and the specific formula is: Frequency control word = (Input frequency / NCO reference clock frequency) × 2^W. Here, the NCO reference clock frequency is the clock frequency that drives the internal phase accumulator of the NCO, provided by a stable clock source inside the FPGA, for example, it can be set to 500 MHz. W is the bit width of the NCO frequency control register, which determines the resolution of the frequency setting, and is set to a relatively large value, such as 32 bits. In other words, the setting of the NCO reference clock frequency and the bit width of the frequency control register is determined based on a comprehensive trade-off between the FPGA's hardware resource capabilities, the maximum operating frequency, and the resolution and range of the output clock frequency. Based on these preset parameters, the control word conversion unit performs the calculation: Offset control word = (+11.71875MHz / 500MHz) × 2^32 ≈ +100663296. This calculated positive integer +100663296 is the offset control word, which precisely represents the portion of the input signal frequency that exceeds the nominal frequency. Next, step 212 is executed. The nominal frequency control word is a pre-calculated constant stored in the FPGA's internal register, representing the NCO control word corresponding to the theoretical fundamental frequency of the signal under ideal conditions. Its calculation method is the same as the offset control word, using the theoretical fundamental frequency (e.g., 5 GHz) as input: Nominal frequency control word = (5 GHz / 500MHz) × 2^32 = 10 × 4294967296 = 42949672960. A dedicated hardware adder is used to add the offset control word (+100663296) obtained in the previous step to the nominal frequency control word (42949672960). The result of the addition operation is: compensated frequency control word = 42949672960 + 100663296 = 43050336256. This final compensated frequency control word precisely corresponds to the actual fundamental frequency of the signal (5.01171875 GHz) at the digital level. Finally, step 213 is executed. The NCO is the core clock generation component in the CDR loop, which contains a phase accumulator and a phase-to-amplitude converter. The phase accumulator adds the value in the frequency control register to its own phase value in each reference clock cycle. Therefore, the value of the frequency control register directly and linearly determines the frequency of the NCO output signal. The compensated frequency control word (43050336256) is written into the NCO's 32-bit frequency control register in one go using a control signal.Once loaded, the NCO immediately begins operating according to this new frequency instruction, and its output digital clock signal sequence is precisely set to a frequency of 5.01171875 gigahertz. This pre-compensated NCO output clock, which is already very close to the actual signal frequency before the loop is closed, is the initial NCO clock.
[0036] Step 22 is executed in parallel. The core component of this processing flow is the loop filter, which is implemented in the FPGA as a digital proportional-integral (PI) controller. The hardware architecture of this PI controller mainly includes two parallel arithmetic paths: a proportional path and an integral path, and a final adder. The proportional path contains a multiplier to multiply the input phase error by a scaling factor (Kp). The integral path contains an adder and an accumulator consisting of a register, and a multiplier to accumulate the phase error and multiply it by the integral factor (Ki). The outputs of the two paths are summed at the final adder to generate the total output of the filter. The values of the coefficients Kp and Ki directly determine the bandwidth and damping characteristics of the loop filter. They are stored in a dedicated coefficient register inside the FPGA and can be dynamically modified at runtime. To achieve switching between different modes, at least two sets of PI coefficient values are pre-stored in the FPGA's memory units (such as block random access memory or distributed registers). One set is for high-bandwidth coefficients, and the other set is for low-bandwidth coefficients. These coefficient values are determined based on the loop theory model, after simulation and optimization for specific performance objectives (such as fastest lock-in time or best jitter suppression). For example, the high-bandwidth coefficients can be set to Kp_high = 0.125 and Ki_high = 0.015625. These values are represented as fixed-point numbers or binary decimals to facilitate efficient multiplication operations in hardware. For example, 0.125 can be represented as binary 0.001, implemented in hardware as a 3-bit arithmetic right shift. When the entire CDR process starts, a top-level control state machine enters a lock-in start state. In this state, the control state machine issues a load instruction. This instruction drives a multiplexer or a bus controller to select and read the pre-stored high-bandwidth coefficients (Kp_high and Ki_high) from the memory. Subsequently, these two coefficient values are written to the corresponding Kp coefficient register and Ki coefficient register inside the loop filter, respectively. This write operation is a one-time operation, instantly setting the loop filter's behavior mode to high-bandwidth mode. Once loading is complete, the loop filter is configured and becomes a configured loop filter. In this state, any subsequent input of raw phase error to the filter will be processed according to the high bandwidth coefficient. This means the filter will respond very quickly and strongly to changes in phase error, driving the NCO to rapidly adjust its phase to minimize phase error and ultimately achieve fast locking. This configured loop filter will maintain this high bandwidth state until subsequent steps detect that locking has been achieved and trigger a mode switch.
[0037] Step 23 is then executed. A component of this process is a phase detector (PD) implemented in an FPGA. This phase detector receives two synchronous inputs: a continuous, real-time digitized signal stream from a high-speed analog-to-digital converter (ADC), such as a 10-bit wide digital sequence; and an initial NCO clock from step 21, with a pre-compensated frequency of approximately 5.01171875 GHz. This embodiment employs an Alexander-type phase detector, also known as a Baud-Rate phase detector, because it performs a phase assessment once per data bit cycle (UI). The hardware architecture of this Alexander phase detector primarily consists of a set of sampling registers and a combinational logic circuit. It utilizes the initial NCO clock to drive three key sampling moments: a data sampling point at the center of the data eye, an edge sampling point at the data transition edge, and a previous data sampling point referencing the center of the previous data eye. Specifically, at the start of one cycle of the initial NCO clock, a delay phase-locked loop (PLL) or fine delay line generates three sampling clocks: one aligned with the NCO clock for data sampling, one delayed by half a clock cycle relative to the NCO clock for edge sampling, and the other delayed by a full clock cycle for the previous data sampling. Within each NCO clock cycle, the phase detector performs the following operations: First, using these three sampling clocks, it samples the real-time digitized signal stream three times, obtaining three sample values, denoted as the current data sample, the current edge sample, and the previous data sample, respectively. These sample values are latched into their respective registers. Subsequently, a combinational logic circuit evaluates these three latched sample values to determine the phase relationship. This logic first determines whether a data transition has occurred within the current cycle by comparing the current data sample value with the previous data sample value. If the previous data sample is not equal to the previous data sample, it indicates a valid data transition edge exists. Phase detection is only meaningful when a valid transition is detected. In the case of a transition, the logic circuit further compares the edge sample value with the data value before the transition. If the edge sample value is the same as the data value before the transition, it indicates that the data transition has not yet been completed at the edge sampling moment, meaning the data transition edge occurred after the edge sampling point of the NCO clock, i.e., the data is later than the clock. Conversely, if the edge sample value is equal to the data value after the transition, it indicates that the data transition edge occurred before the edge sampling point, i.e., the data is earlier than the clock. Based on this logical judgment, the phase detector outputs a digital code in each evaluation cycle. For example, an output of +1 (binary representation 01) represents data earlier than the clock, an output of -1 (binary representation 11) represents data later than the clock, and an output of 0 (binary representation 00) represents no data transition detected. This continuously output discrete numerical sequence consisting of +1, -1, and 0 is the original phase error stream.
[0038] Finally, step 24 is executed. The component of this process is the loop filter configured in high-bandwidth mode in step 22. This filter receives the raw phase error stream consisting of +1, -1, and 0 from step 23. The filter's hardware architecture is a digital proportional-integral (PI) controller. It contains two parallel processing paths: a proportional path and an integral path. In the proportional path, the input raw phase error value, for example +1, is fed into a multiplier. This multiplier is connected to a register storing a high scaling factor Kp_high, such as 0.125. The multiplier calculates the output of the scaling term, i.e., +1 × 0.125 = +0.125. This result represents the direct response to the current instantaneous phase error. In the integral path, the input raw phase error value +1 is first fed into an adder, which is connected to a register acting as an accumulator, forming an integrator. If the accumulated value in the integrator register in the previous cycle was -0.5, then in this cycle, the new error value +1 will be added, making the updated value of the register -0.5 + 1 = +0.5. This accumulated value +0.5 is then fed into another multiplier and multiplied by a register storing a high integration coefficient Ki_high, such as 0.015625, to obtain the integral term output as +0.5 × 0.015625 = +0.0078125. This multiplication result represents the correction for the cumulative effect of historical phase errors, including the current error, and is mainly used to eliminate long-term static phase deviations. The outputs of the two paths (proportional term +0.125 and integral term +0.0078125) are fed into a final adder for summation. The summation result, +0.125 + 0.0078125 = +0.1328125, is the phase error signal. It is a filtered, smoothed digital control word representing the required phase adjustment. Because the loop filter operates in high-bandwidth mode, this phase error signal can very quickly reflect changes in the original phase error stream, thus achieving a fast loop response. Finally, this generated phase error signal is directly fed into the phase control port of the numerically controlled oscillator (NCO). The NCO's internal phase accumulator accumulates its frequency control word in each reference clock cycle, and also adds this phase error signal (after appropriate scaling). This is equivalent to performing real-time, fine-tuned phase adjustment based on the NCO's original frequency. If the phase error signal is positive, the NCO's phase advances faster; if it is negative, it advances slower. In this way, the NCO's output clock phase is continuously fine-tuned to continuously reduce the phase difference with the input data stream. After this closed-loop fine-tuning process, the NCO's output clock is the fine-tuned NCO clock, achieving tight phase synchronization with the input data stream, indicating that the loop has entered a locked state.
[0039] In step 3, hardware verification of the locked state is performed on the continuously monitored phase error signal to obtain a lock achievement indicator. It is understood that during the high-bandwidth lock-up startup phase of the Clock Data Recovery (CDR) loop, the loop converges rapidly to reduce the initial phase error. However, the mere fact that the loop has started operating does not immediately indicate that it has reached a stable and accurate locked state. The phase error signal, as a key indicator reflecting the instantaneous state of the loop, directly reveals whether the loop has truly converged and stably tracked the input signal. If subsequent operations (such as switching to a low-bandwidth mode) are performed before the error has been sufficiently reduced, it may lead to loop lockout or performance degradation. Therefore, this application performs hardware verification of the locked state on the continuously monitored phase error signal, that is, by analyzing the statistical characteristics of the signal in real time, it objectively and automatically determines whether the loop has entered and maintained a stable locked state that meets the accuracy requirements, thereby providing a reliable hardware trigger signal for subsequent mode switching decisions.
[0040] In one possible implementation, step 3 is as follows: using a lock detector module implemented in an FPGA. The hardware architecture of the lock detector mainly consists of three functional units: a sliding window statistics calculation unit, a threshold comparator, and a continuous pass / fail counter and its associated control logic.
[0041] First, the input phase error signal stream is fed into the sliding window statistics calculation unit. This unit calculates the statistics of the phase error signal within a configurable sliding time window. In this embodiment, the root mean square (RMS) value of the phase error signal within the window is chosen as the statistical indicator, which comprehensively reflects the fluctuation intensity of the phase error within the window period. The calculation unit internally includes a squarer, an accumulator, a window shift register, and a divider. For each newly arrived phase error value, the calculation process is as follows: the value is first fed into the squarer to calculate its square. Simultaneously, the value is stored in a shift register (window memory) of length M, and the oldest value is popped out. The accumulator continuously maintains the sum of squares of all M error values within the current window. Whenever a new value enters the window, the accumulator executes: New sum of squares = Old sum of squares - Square of the popped-out old value + Square of the newly entered value. Then, the divider calculates the RMS value of the error within the current window: RMS value = sqrt((New sum of squares) / M). The preset window length M is a result of a trade-off between desired statistical stability and hardware resource consumption. For example, it can be set to a power of 2, such as the number of sampling points corresponding to 1024 UIs, to facilitate hardware division operations. For instance, if each UI corresponds to one phase error sampling point, then M can be set to 1024, representing the evaluation of phase error fluctuations over 1024 UI periods.
[0042] Next, the calculated root mean square (RMS) value is fed into a threshold comparator. This comparator compares the RMS value with a preset lockout threshold. This lockout threshold is preset to a small value, such as 0.05 UI, based on the desired lockout accuracy and loop stability performance requirements. The lockout threshold is stored in the FPGA's configuration register. If the current RMS value is less than or equal to the lockout threshold, the comparator outputs a high level (logic 1), indicating that the phase error fluctuation within the current window meets the lockout accuracy requirements; otherwise, it outputs a low level (logic 0).
[0043] Finally, the comparator's output is fed into a persistent valid counter and its control logic. The goal of this logic is to verify the persistence of the lock-in condition met by the phase error signal. The counter is preset with a count value N, pre-set according to the lock-in stability requirements, for example, 128. Whenever the comparator outputs a high level (condition met), the counter increments by 1 in each evaluation cycle (each UI or each sampling cycle); when the comparator outputs a low level (condition not met), the counter is immediately reset to zero. A state machine monitors the counter's value. When the counter's value reaches or exceeds the preset persistent count threshold N (e.g., 128), the state machine determines that the lock-in state has been stable for a sufficiently long time, thus setting the final lock-in achievement flag high (logic 1). This flag will remain high until cleared by a subsequent operation (e.g., a loop reset). For example, the lock-in achievement flag will only be set high if the root mean square value of the phase error signal is below 0.05UI for 128 consecutive UI cycles.
[0044] In step 4, in response to the lock achievement flag being high, the proportional and integral coefficients of the high-level triggered loop filter are smoothly switched from high-bandwidth mode to low-bandwidth mode based on a multi-step interpolation state machine to obtain a high-precision recovered clock and time interval error measurement stream. It should be understood that after the clock data recovery (CDR) loop successfully achieves fast lock-up, its objective shifts from rapid acquisition to accurate tracking and high-precision measurement. While the high-bandwidth mode used in the initial stage is beneficial for fast lock-up, its ability to suppress high-frequency jitter is poor, leading to significant jitter in the recovered clock and large noise in the phase error signal output by the loop filter, making it unsuitable as a basis for accurate time interval error (TIE) measurement. If the loop bandwidth is suddenly switched from high to low at this point, the step change in coefficients will cause severe oscillations in the loop's transient response, potentially leading to temporary loss of lock-up and disrupting the established stable state. Therefore, in order to respond to the signal that the lock is achieved, the characteristics of the loop filter are smoothly transitioned from the high-bandwidth mode to the low-bandwidth mode in a gradual and perturbation-free manner. This allows the CDR loop to be seamlessly switched to the high-precision tracking and measurement mode without compromising loop stability, ultimately resulting in a low-jitter, high-precision recovery clock and a clean time interval error measurement stream that can be used for jitter analysis. This application requires a smooth switching of the multi-step interpolation state machine.
[0045] In one possible implementation, Figure 4 This is a flowchart of step 4 in the FPGA-based digital signal processing method according to an embodiment of this application. Figure 4 As shown, step 4, in response to the lock-up achievement flag being high, smoothly switches the proportional coefficient and integral coefficient of the high-level triggered loop filter from high-bandwidth mode to low-bandwidth mode based on a multi-step interpolation state machine to obtain a high-precision recovery clock and time interval error measurement stream, includes: step 41, in response to detecting the lock-up achievement flag being high, extracting high proportional coefficient, high integral coefficient, low proportional coefficient, and low integral coefficient from high-bandwidth coefficient and low-bandwidth coefficient; step 42, performing incremental calculation on the high proportional coefficient, high integral coefficient, low proportional coefficient, and low integral coefficient to obtain coefficient value increments composed of proportional increments and integral increments; step 43, the multi-step interpolation state machine uses the coefficient value increments to perform step-by-step interpolation of the loop filter coefficients so that the proportional coefficient and integral coefficient of the loop filter are smoothly switched to low-bandwidth mode; step 44, inputting the original phase error stream into the loop filter switched to low-bandwidth mode to obtain the time interval error measurement stream; step 45, the NCO performs phase adjustment based on the time interval error measurement stream to obtain the high-precision recovery clock.
[0046] In the above implementation, step 4 is as follows: First, step 41 is executed. This is performed by a top-level control logic unit implemented in the FPGA. The hardware architecture of this unit includes an edge detector and a set of data loading logic. The edge detector continuously monitors the latch achievement flag signal line from step 3. When the signal line transitions from a low level (logic 0) to a high level (logic 1), the edge detector generates a single-cycle pulse as a trigger signal. This trigger signal activates the data loading logic. This logic is connected to a preset coefficient storage area, which can be a read-only memory (ROM) inside the FPGA or a set of configurable registers. At least two sets of loop filter coefficients are stored in this storage area. These coefficients are pre-simulated and optimized based on the loop theory model for different performance targets. For example, the high bandwidth coefficient is set to {Kp_high = 0.125, Ki_high = 0.015625}, and the low bandwidth coefficient is set to {Kp_low = 0.0078125, Ki_low = 0.000244140625}. Upon receiving the trigger pulse, the data loading logic simultaneously reads these four fixed-point values from this memory area and writes them respectively into four dedicated working registers via the data bus. These four working registers are named: High Proportional Coefficient Register, High Integral Coefficient Register, Low Proportional Coefficient Register, and Low Integral Coefficient Register.
[0047] Next, step 42 is executed. This is performed by the Arithmetic Logic Unit (ALU), which is activated after step 41. Its hardware architecture contains two parallel computation paths: one for calculating the proportional increment and the other for calculating the integral increment. Each path consists of a subtractor and a divider (or their equivalent shift and multiplication logic). The ALU first reads the input values from four working registers: high proportional coefficient Kp_high = 0.125, high integral coefficient Ki_high = 0.015625, low proportional coefficient Kp_low = 0.0078125, and low integral coefficient Ki_low = 0.000244140625. Simultaneously, it reads a preset interpolation step number S from a configuration register. This S value determines the smoothness and duration of the switching process and is a key design parameter that can be set according to the required smoothness. For example, setting S = 256 achieves a balance between smoothness and switching speed. In the proportional increment calculation path, the subtractor first calculates the total difference in proportional coefficients: ΔKp_total = Kp_low - Kp_high = 0.0078125 - 0.125 = -0.1171875. Then, the divider divides this total difference by the number of interpolation steps S: proportional increment ΔKp = ΔKp_total / S = -0.1171875 / 256 ≈ -0.00045776. In the integral increment calculation path, the operation is completely similar. The subtractor calculates the total difference in integral coefficients: ΔKi_total = Ki_low - Ki_high = 0.000244140625 - 0.015625 = -0.015380859375. Then, the divider calculates the integral increment ΔKi = ΔKi_total / S = -0.015380859375 / 256 ≈ -0.00006008. Since implementing a floating-point divider in an FPGA consumes significant resources, the division operation is converted to a division by a power of 2 (i.e., an arithmetic right shift) or multiplied by the reciprocal of S using a fixed-point multiplier. The calculated proportional increment ΔKp and integral increment ΔKi, these two fixed-point values, together constitute the coefficient increment. They are then written into two dedicated increment registers.
[0048] Then, step 43 is executed. The multi-step interpolation state machine mainly includes a main controller, a step counter, and two parallel coefficient update arithmetic units (one for proportional coefficients and one for integral coefficients). Upon receiving a lock attainment flag, the state machine is activated from the idle state and enters the interpolation running state. At the initial moment of entering the interpolation running state, the two coefficient update arithmetic units load the high proportional coefficient Kp_high = 0.125 and the high integral coefficient Ki_high = 0.015625 from the working registers of step 41 as the initial values of their internal current coefficient value registers. Simultaneously, the step counter is cleared. In the interpolation running state, the state machine drives an interpolation update according to a preset rhythm, such as every unit interval UI or a fixed number of clock cycles. The counting range of the step counter is set from 0 to S-1, where S is the preset total number of interpolation steps, such as S = 256. In the update cycle when the counter count is i (i ranges from 0 to 255), the multi-step interpolation state machine performs the following operations: First, the two coefficient update arithmetic units read the proportional increment ΔKp≈-0.00045776 and the integral increment ΔKi≈-0.00006008 from the increment register of step 42, respectively. Next, each arithmetic unit contains an adder. The proportional coefficient update unit calculates: Kp_current(i) = Kp_high + i × ΔKp. The integral coefficient update unit calculates: Ki_current(i) = Ki_high + i × ΔKi. A more efficient implementation is to accumulate the increment value into the current coefficient value register in each cycle: Kp_current(i) = Kp_current(i-1) + ΔKp. Finally, the two calculated instantaneous coefficient values Kp_current(i) and Ki_current(i) are immediately written into the Kp and Ki coefficient registers inside the loop filter that is actually working in the CDR loop. This write operation changes the dynamic characteristics of the loop filter in real time. This stepwise interpolation process continues for S (256) update cycles. As the step counter i increases from 0 to 255, the coefficients loaded into the loop filter transition linearly and smoothly from high bandwidth values to low bandwidth values. When the step counter completes the 255th update, the state machine enters the interpolation complete state. In this state, to ensure accuracy, it directly reads the final low proportional coefficient Kp_low and low integral coefficient Ki_low from the register in step 41 and writes them fixedly into the coefficient register of the loop filter. At this point, the coefficients of the loop filter have been completely and smoothly switched to the low bandwidth mode and are operating stably in this mode.
[0049] Specifically, in the clock data recovery loop, while performing linear incremental calculations and correspondingly updating the loop filter coefficients step-by-step via linear interpolation is simple and effective, it does not fully consider the preservation of the loop filter's key dynamic characteristics in control theory. In other words, the update of the loop filter coefficients should be performed in the parameter space along a path that maintains the relative constancy of key system dynamic characteristics (such as damping ratio), which is usually not a simple linear path. Consider the proportional coefficient K of the loop filter in a second-order phase-locked loop model. p and integral coefficient K i Together, they determine two important system dynamic characteristics: the natural frequency ω n Damping ratio δ:
[0050]
[0051] Among them, K pd It is the phase detector gain, K vco This is the NCO gain, which can be considered a constant during switching. Natural frequency ω n The loop bandwidth (BW≈2δω) is determined. n ), that is, loop response speed, the goal is to improve from high ω n (High bandwidth) Smooth transition to low ω n (Low bandwidth). The damping ratio δ determines the system's response to a step input. A ratio less than one indicates underdamping, resulting in a fast response but with overshoot and oscillations; a ratio equal to one indicates critical damping, with the fastest response and no overshoot; and a ratio greater than one indicates overdamping, resulting in a slow response and no overshoot. When the proportional and integral coefficients change linearly, the damping ratio, a key parameter determining loop stability, exhibits a complex nonlinear relationship, as shown in the second-order phase-locked loop model. Therefore, when the proportional coefficient K... p and integral coefficient K i When the damping ratio δ changes linearly, its path is a complex nonlinear relationship. This means that during switching, the damping ratio δ may fluctuate drastically, even entering a deep underdamped region (δ << 1), causing instantaneous oscillations in the loop. Even without lockout, this leaves observable, severe fluctuations in phase error, affecting the smoothness of the transition. Therefore, by interpolating the loop filter coefficients based on a constant damping ratio, this unstable linear update path can be completely eliminated. Instead, a precise nonlinear update strategy based on control theory is adopted, forcing the update trajectories of the proportional and integral coefficients to strictly follow a specific path that keeps the damping ratio at an ideal constant value throughout the switching process. This fundamentally eliminates instantaneous overshoot and oscillations caused by drastic fluctuations in the damping ratio, achieving a truly smooth, stable, and predictable bandwidth switching process.
[0052] In a possible preferred embodiment, step 43, where the multi-step interpolation state machine uses the coefficient value increments to perform step-by-step interpolation of the loop filter coefficients so that the proportional coefficients and integral coefficients of the loop filter are smoothly switched to a low-bandwidth mode, includes:
[0053] Based on the target damping ratio and loop gain, the constant damping ratio constraint constant is determined. It should be understood that in the second-order phase-locked loop model, the damping ratio δ is determined by… The formula determines this. To maintain δ at a constant value during the switching process, a K needs to be found. p and K i The mathematical relationship that is always satisfied between them. By transforming this formula, we can obtain... In this relationship, the target damping ratio δ is a pre-set ideal value, for example, set to 0.707 near the critical damping to obtain the fastest overshoot-free response. Loop gain K pd (phase detector gain) and K vco The gain of the numerically controlled oscillator is an inherent property of the hardware, determined during the design phase, and can be considered a constant during switching. Therefore, all terms on the right-hand side of the equation... Together they form a constant composite constant, denoted as C. This reduces the time required to update two variables (K) independently. p and K i The complex problem can be simplified into a constraint problem: only one variable K needs to be determined. i The update path, another variable K p It can be passed The relationship is uniquely determined. This paves the way for subsequent iterative calculations, ensuring that as long as this constraint is followed, the damping ratio will remain constant, thus guaranteeing the stability of the switching process.
[0054] Based on high integral coefficients, low proportional coefficients, low integral coefficients, and the number of steps, the single-step increment in the logarithmic field is calculated. Correspondingly, the integral coefficients K are directly... i Linear interpolation results in a constant rate of change, which can seem harsh in control systems. However, by performing linear interpolation in the logarithmic domain and mapping it back to the linear domain, K... i The changes in K will exhibit exponential characteristics. This means that K i The rate of change is proportional to its current value, especially in the initial switching phase (K). i When the value is large, it changes rapidly, while when it approaches the target value (K)... i When the value is small, the change slows down. This exponential change, which is initially fast and then slows down, is inherently smoother than a linear change, better meeting the smooth transition requirements of physical systems. In practical implementation, it is based on a high integral coefficient K. ih Low integral coefficient K il And the preset total number of interpolation steps N, through formula Ki (Δ)=[log(K il )-log(K ih The constant increment that needs to be changed at each step in the logarithmic field is calculated using ] / N. Where K ih and K il The number of steps, N, is determined by the design requirements of high and low bandwidth, and is a configurable parameter, for example, set to 256, to strike a balance between switching smoothness and switching time. This provides a constant step increment used in the logarithmic domain for subsequent iterations. By planning in the more stable space of the logarithmic domain, the final K is ensured. i The update trajectory in the linear domain is a smooth exponential curve, providing an ideal driving force for achieving perturbation-free switching.
[0055] At each step of the switching process, perform the following operations:
[0056] The integral coefficient of the current step is calculated by performing linear interpolation in the number domain and then exponential operation; the proportional coefficient of the current step is calculated based on the constant damping ratio constraint constant and the integral coefficient of the current step; the integral coefficient of the current step and the current step are then written into the corresponding coefficient register inside the loop filter. That is, in each interpolation period t, first, linear interpolation in the number domain is performed and then exponential operation is performed, i.e., log[K... i [(t)]=logK i (start)+tlogK i (Δ), K i (start) = log(K) ih K is obtained through antilogarithmic (exponential) operations. i linear value Calculate the integral coefficient K for the current step. i (t). This step is to implement K. i The key to exponential smooth changes. Next, based on the constant damping ratio constraint constant C determined in the first step and the newly obtained K... i (t), through Calculate the scaling factor K for the current step. p (t). This step is crucial for enforcing the constant damping ratio constraint, ensuring K... p The update path strictly matches K i The change is adjusted to maintain a constant δ. Finally, the calculated instantaneous coefficient value K is... p (t) and K i(t) is written into the coefficient register inside the loop filter. In other words, it converts the theoretically calculated value into the dynamic characteristics of the physical loop in real time. As the number of steps t increases from 1 to N, the loop coefficients are updated precisely, step by step, along a preset nonlinear trajectory. Because the damping ratio is precisely maintained at an ideal non-oscillating constant value (e.g., 0.707) throughout the switching process, the loop's response characteristics remain predictable and stable. This fundamentally eliminates instantaneous overshoot and oscillations caused by drastic fluctuations in the damping ratio, resulting in an extremely smooth transition curve reflected in the phase error signal. Simultaneously, even if there is external noise or signal jitter interference during the switching process, the loop can suppress it with its constant and excellent dynamic characteristics, without being amplified by interference due to its own position within an unstable damping ratio range. Ultimately, this achieves a fully controllable and highly reliable smooth bandwidth switching. After the multi-step interpolation state machine completes the final iterative calculation (i.e., when step t equals N), the integral coefficient Ki(t) and the proportional coefficient Kp(t) have precisely reached the preset target values of low integral coefficient Kil and low proportional coefficient Kpl. At this point, the coefficients of the loop filter are fixed at these low values and no longer change. According to the theory of second-order phase-locked loops, the lower integral coefficient Ki and proportional coefficient Kp directly lead to a reduction in the natural frequency of the loop, thereby significantly narrowing the loop bandwidth. Therefore, through this smooth, step-by-step interpolation process, the loop finally operates stably in the low-bandwidth mode determined by the target low coefficient values, realizing the state transition from fast locking to high-precision tracking.
[0057] During and after the smooth switching of the loop filter coefficients, step 44 is executed. After step 43, the proportional coefficient (Kp) and integral coefficient (Ki) registers inside the loop filter have been fixed to the coefficient values of the low-bandwidth mode. The raw phase error stream from the phase detector in step 23, i.e., a three-valued digital sequence consisting of +1, -1, and 0, is continuously fed into the input port of this loop filter configured for low-bandwidth mode. The loop filter processes each input error value. Its internal proportional path multiplies the input error value, for example -1, by the low proportional coefficient Kp_low to obtain the proportional term output -0.0078125. Its integral path accumulates the input error value -1 into the integrator and multiplies the accumulation result by the low integral coefficient Ki_low to obtain the integral term output. Since the values of Kp_low and Ki_low are very small, the filter's response to a single input error value becomes very weak and slow. It mainly responds to the long-term average trend of the error, i.e., the low-frequency phase drift. The outputs of the two paths are summed to obtain the final output of the loop filter. This output signal stream, due to low-pass filtering, has its high-frequency components significantly attenuated, resulting in a very smooth signal. This smooth, precise digital sequence reflects the long-term phase deviation of the data signal relative to the ideal clock, and its physical meaning aligns with the definition of Time Interval Error (TIE). Therefore, this output signal stream is the Time Interval Error Measurement Stream; it is a continuous, high-precision digital stream that can continue to be used as a closed-loop control signal and sent to the NCO.
[0058] Finally, step 45 is executed. The numerically controlled oscillator (NCO) mainly consists of a phase accumulator and a phase-to-amplitude converter, which can be a lookup table or CORDIC module. The output frequency of the NCO is determined by the value of its frequency control register, while its output phase can be finely adjusted in real time through a dedicated phase control port. The time interval error measurement stream from step 44, which is a smooth digital sequence representing the required phase correction amount, is continuously fed into the phase control port of the NCO. Inside the NCO, this input control value is added to the value of the phase accumulator. Specifically, in each reference clock cycle of the NCO, the phase accumulator performs the following operation: New phase value = Old phase value + Frequency control word + Phase adjustment value. Here, the phase adjustment value is the input time interval error measurement stream, and the frequency control word is the compensated frequency control word from step 212. For example, at a certain moment, the value of the time interval error measurement stream may be a small positive number, such as +0.002UI, indicating that the NCO clock is slightly lagging behind the data signal. This positive value is added to the phase accumulator, causing the NCO's phase to advance slightly faster than normal, thus catching up with the data signal's phase. Conversely, if the input value is negative, such as -0.0015UI, the NCO's phase advance will be slightly slower to correct for lead. Since the input time interval error measurement stream has been strongly smoothed by a low-bandwidth filter, it does not contain drastic high-frequency fluctuations. Therefore, the phase adjustment applied to the NCO is also smooth and fine. This fine phase adjustment based on the smooth control signal allows the NCO to stably track low-frequency jitter and long-term phase drift of the input signal, while effectively suppressing phase disturbances caused by noise and high-frequency jitter. Finally, the NCO's phase-to-amplitude converter generates a digitized sine or square wave sequence based on this steadily changing phase accumulator value. This output digital clock sequence has extremely low jitter and its phase remains precisely synchronized with the input data stream over a long period, thus constituting a high-precision recovered clock.
[0059] In step 5, based on the high-precision recovered clock and the time interval error (TIE) measurement stream, data deserialization and jitter analysis are performed on the real-time digital signal stream to obtain the deserialized data output and jitter analysis results. Correspondingly, after the Clock Data Recovery (CDR) loop successfully generates a high-quality recovered clock and an accurate time interval error (TIE) measurement stream, the ultimate goal of the entire process is to utilize these intermediate results to accomplish two core tasks: first, to recover the original data; and second, to quantify signal quality. The high-precision recovered clock provides an ideal time reference for error-free data decision-making, while the time interval error measurement stream contains all the information about signal jitter. Therefore, performing data deserialization and jitter analysis based on these two key signal streams can transform the results of the CDR process into final application value. On the one hand, by sampling the signal at the optimal time, the original serial data stream can be accurately recovered; on the other hand, by performing in-depth statistical analysis on the TIE stream, the jitter characteristics of the signal can be quantitatively evaluated, providing key quantitative indicators for link performance diagnosis and optimization.
[0060] In one possible implementation, step 5, based on the high-precision recovery clock and time interval error measurement stream, performs data deserialization and jitter analysis on the real-time digital signal stream to obtain the deserialized data output and jitter analysis results, including: step 51, using the optimal sampling phase of the high-precision recovery clock to sample the real-time digital signal stream to obtain the deserialized data output; step 52, inputting the time interval error measurement stream into the real-time jitter analysis engine to obtain the jitter analysis results.
[0061] In the above implementation, step 5 is as follows: First, step 51 is executed. This is performed by the data sampling decision module. This module mainly consists of a finely adjustable digital delay line and a binary decision unit, such as a comparator. Before operation begins, the optimal sampling phase offset value needs to be set. This value represents the time delay of the ideal sampling point relative to the recovery clock reference edge (e.g., the rising edge). Theoretically, the optimal sampling point is located at the center of the data eye diagram, i.e., the position farthest from the data transition edge, typically half a unit interval (0.5UI) after the clock reference edge. This value can be determined through link characteristic simulation or pre-calibration testing and written into a configuration register of the FPGA. For example, if a UI is 100 picoseconds, the offset value can be set to 50 picoseconds. After processing begins, the high-precision recovery clock is first fed into the finely adjustable digital delay line. This delay line can apply a precise time delay to the input clock signal according to the optimal sampling phase offset value in the configuration register. For example, it will delay each valid edge of the recovery clock by 50 picoseconds. After processing by the delay line, the output is a sampling clock with a precisely adjusted phase. The effective edge of this sampling clock is now precisely aligned with the center of the eye diagram for each data bit cycle. This phase-adjusted sampling clock is then used as the trigger signal for a binary decision unit. The decision unit is a digital comparator that samples the real-time digitized signal stream once at the effective edge (e.g., the rising edge) of each sampling clock cycle. That is, it latches the amplitude of the 10-bit wide digital signal stream at that moment. The decision unit then compares this latched sample value with a preset decision threshold. This threshold represents the level boundary between logic 0 and logic 1, typically 0 for differential signals and the midpoint of the signal swing for single-ended signals. This threshold is also preset and stored in a configuration register. Based on the comparison result, the decision unit outputs a single-bit binary value: a logic 1 if the sample value is greater than the decision threshold, and a logic 0 if the sample value is less than or equal to the decision threshold. This single-bit binary sequence of 0s and 1s output each sampling clock cycle ultimately yields the deserialized data output.
[0062] In parallel, step 52 is performed. In one possible implementation, step 52, inputting the time interval error measurement stream into a real-time jitter analysis engine to obtain the jitter analysis result, includes: step 521, constructing a real-time histogram of the time interval error measurement stream based on histogram configuration parameters to obtain a TIE histogram vector; step 522, extracting dual Dirac model parameters from the TIE histogram vector; and step 523, inputting the dual Dirac model parameters into a dual Dirac jitter model to obtain a jitter analysis result consisting of deterministic jitter values, random jitter values, and total jitter values.
[0063] Specifically, step 521 is executed first. This step is performed by the histogram construction unit. This unit receives the Time Interval Error Measurement (TIE) stream from step 44, which is a continuous, low-bandwidth filtered digital sequence where each value represents the phase deviation at a given moment, expressed as a fraction of a unit interval (UI). Simultaneously, the unit reads a set of preset histogram configuration parameters from the configuration register. These parameters are pre-set for statistical analysis and primarily include: the histogram quantization range defines the interval of the TIE of interest, for example, set to -0.5UI to +0.5UI, covering a complete bit period; the number of quantization bins determines the histogram resolution, for example, set to 1024 bins. This means the entire quantization range is divided into 1024 equally wide intervals, each with a width of (0.5 - (-0.5)) / 1024 ≈ 0.0009765625UI. The core hardware architecture of the histogram construction unit is a dual-port block random access memory (BRAM) and address calculation logic. The depth of the BRAM is consistent with the number of quantization bins (1024), and its bit width needs to be large enough to avoid count overflow (e.g., 32 bits). The address calculation logic is responsible for mapping the input TIE values to addresses in the BRAM. For each input TIE value, it performs a mapping calculation once. For example, a time interval error of +0.125UI corresponds to the BRAM address floor(((+0.125-(-0.5)) / (1.0))*1024)=floor(0.625*1024)=640. In each clock cycle, this unit performs a read-modify-write operation on a new TIE value: it reads the current count value from the BRAM according to the calculated address, increments the count value by 1 using an adder, and then writes the new result back to the same address in the BRAM. This process continues, and as a large number of TIE samples accumulate, the 1024 count values stored in the BRAM form a vector that depicts the probability density distribution of the TIE values, i.e., the IE histogram vector. Next, step 522 is executed. Once the histogram has accumulated a sufficient number of samples (e.g., reaching a preset threshold for the total number of samples), the dual Dirac model parameter extraction unit is activated. It reads the entire TIE histogram vector (i.e., the 1024 count values in the BRAM) at once and analyzes it to fit an industry-standard dual Dirac jitter model. This model decomposes the total jitter into deterministic jitter (DJ) and random jitter (RJ). The unit's hardware logic is designed to execute an algorithmic flow. First, it traverses the entire histogram vector using a peak search algorithm to locate the two peaks with the highest probability density. These two peaks correspond to the average positions of the edges during the data transitions from 0 to 1 and from 1 to 0.The distance between the center locations of the two peaks (i.e., the corresponding BRAM addresses or converted TIE values) is calculated as an estimate of the deterministic jitter peak-to-peak value (DJ_pp). For example, if the two peaks are located at addresses 200 and 824 respectively, then DJ_pp ≈ (824-200)*(1.0 / 1024) = 0.609UI. Then, the unit performs Gaussian curve fitting on the data on both sides of each peak. The standard deviation (σ) of the Gaussian distribution that best matches the peak shape can be estimated using a hardware-implemented least squares method or a lookup table method. Since random jitter (RJ) is modeled as a Gaussian distribution, the two standard deviations (σ1 and σ2) extracted from the peak shape are averaged or root mean squared to obtain a unified random jitter standard deviation RJ_rms. These two key model parameters, DJ_pp and RJ_rms, are the dual Dirac model parameters. Finally, step 523 is executed. The jitter calculation unit receives the dual Dirac model parameters extracted in the previous step as input. This unit mainly consists of multipliers and adders, used to perform the calculation of the standard jitter model. It reads a Q factor related to the target bit error rate (BER) from the configuration register. The Q factor is a statistical constant; for example, for the commonly used 10^-12 BER in the communications field, the Q factor is preset to 14.069. The calculation process is as follows: The deterministic jitter value (DJ) is directly taken as its peak-to-peak value, i.e., DJ = DJ_pp. The random jitter value (RJ) is extrapolated to the peak-to-peak value at a specific BER by the standard deviation (rms value) of the random jitter using the Q factor. The calculation formula is: RJ_pp = Q × RJ_rms = 14.069 × RJ_rms. The total jitter value (TJ) at a specific BER is a linear superposition of deterministic jitter and random jitter. The calculation formula is: TJ@10^-12=DJ_pp+RJ_pp, where TJ@10^-12 is the value at a bit error rate of 10. -12 Under the given conditions, the total jitter (TJ) measured is DJ_pp + RJ_pp. Finally, the unit outputs three digitized jitter analysis results: deterministic jitter value, random jitter value, and total jitter value. These results can be displayed, recorded, or used for further link margin analysis, completing the full conversion from raw phase error to the final quantified jitter metric.
[0064] In summary, the FPGA-based digital signal processing method based on the embodiments of this application has been clarified. To address the challenge of balancing locking speed and measurement accuracy in traditional CDR design, this application proposes a dynamically adaptive digital signal processing architecture. This involves introducing a staged, variable bandwidth CDR strategy. Initially, signal frequency offset is rapidly estimated through feedforward frequency analysis, and a high-bandwidth loop is initiated based on this, achieving extremely fast locking and meeting the efficiency requirements of production testing and other scenarios. Once the system confirms stable locking through hardware logic, a multi-step interpolation state machine is triggered. This state machine smoothly and seamlessly switches the coefficients of the loop filter from the high-bandwidth mode to the low-bandwidth mode, avoiding the risk of lock loss during the switching process. Finally, in the stable low-bandwidth mode, the system can utilize a high-precision recovery clock for accurate jitter decomposition and data analysis. This integrates both rapid acquisition and high-precision measurement modes in a single operation, overcoming the limitations of static design in existing technologies.
[0065] Figure 5 This is a block diagram of an FPGA-based digital signal processing system according to an embodiment of this application. Figure 5 As shown, the FPGA-based digital signal processing system 100 according to an embodiment of this application includes: a signal digitization frequency analysis module 110, used to perform signal digitization and feedforward frequency analysis on the acquired analog high-speed signal to obtain a frequency offset value; a phase error signal generation module 120, used to perform CDR loop preset and high bandwidth lock-in activation on the real-time digitized signal stream based on the frequency offset value to obtain a phase error signal; a hardware verification module 130, used to perform hardware verification of the lock-in state of the continuously monitored phase error signal to obtain a lock-in achievement flag; a mode switching module 140, used to smoothly switch the proportional coefficient and integral coefficient of the high-level triggered loop filter from a high bandwidth mode to a low bandwidth mode based on a multi-step interpolation state machine in response to the lock-in achievement flag being high level to obtain a high-precision recovered clock and time interval error measurement stream; and a deserialization jitter analysis module 150, used to perform data deserialization and jitter analysis on the real-time digitized signal stream based on the high-precision recovered clock and time interval error measurement stream to obtain deserialized data output and jitter analysis results.
[0066] Here, those skilled in the art will understand that the specific operations of each step in the above-described FPGA-based digital signal processing system have been referenced above. Figures 1 to 4 The FPGA-based digital signal processing method has been described in detail, and therefore, its repeated description will be omitted.
Claims
1. A digital signal processing method based on FPGA, characterized in that, include: The acquired analog high-speed signal is digitized and subjected to feedforward frequency analysis to obtain the frequency offset value; Based on the frequency offset value, the real-time digital signal stream is subjected to CDR loop preset and high-bandwidth lock-in start to obtain the phase error signal; Hardware verification of the locked state is performed on the continuously monitored phase error signal to obtain a lock achievement flag; In response to the lock attainment flag being high, the proportional coefficient and integral coefficient of the high-level triggered loop filter are smoothly switched from high-bandwidth mode to low-bandwidth mode based on a multi-step interpolation state machine to obtain a high-precision recovery clock and time interval error measurement stream; Based on high-precision recovery of clock and time interval error measurement stream, data deserialization and jitter analysis are performed on real-time digital signal stream to obtain deserialized data output and jitter analysis results.
2. The FPGA-based digital signal processing method according to claim 1, characterized in that, The acquired analog high-speed signal is digitized and subjected to feedforward frequency analysis to obtain the frequency offset value, including: The acquired analog high-speed signal is digitally sampled and snapshot buffered to obtain time-domain sampled data blocks; Perform a Fast Fourier Transform on the time-domain sampled data block to obtain the spectral amplitude vector; The frequency offset value is obtained by performing spectral peak search and frequency offset calculation on the spectral amplitude vector.
3. The FPGA-based digital signal processing method according to claim 2, characterized in that, Performing spectral peak search and frequency offset calculation on the spectral amplitude vector to obtain the frequency offset value includes: The spectral amplitude vector is input into the peak search engine to obtain the peak amplitude and its corresponding peak index; The peak index and FFT parameters are input into the frequency calculation unit to obtain the actual fundamental frequency of the signal. The nominal data rate is input into the nominal frequency calculation unit to obtain the theoretical fundamental frequency, which is half of the nominal data rate; The actual fundamental frequency and the theoretical fundamental frequency of the signal are input into the difference calculation unit to obtain the frequency offset value.
4. The FPGA-based digital signal processing method according to claim 1, characterized in that, Based on the frequency offset value, CDR loop preset and high-bandwidth locked start are performed on the real-time digital signal stream to obtain the phase error signal, including: Based on the frequency offset value, NCO frequency pre-compensation is performed to obtain the initial NCO clock; A high bandwidth coefficient is applied to the loop filter to obtain the configured loop filter; Based on the initial NCO clock, phase error detection based on the pre-compensated clock is performed on the real-time digital signal stream to obtain the original phase error stream; Based on the configured loop filter, the original phase error stream is subjected to high-bandwidth filtering and NCO phase fine-tuning to obtain the phase error signal and the fine-tuned NCO clock.
5. The FPGA-based digital signal processing method according to claim 4, characterized in that, Based on the frequency offset value, NCO frequency pre-compensation is performed to obtain the initial NCO clock, including: The frequency offset value is input into the control word conversion unit to obtain the offset control word; The offset control word is added to the nominal frequency control word to obtain the compensated frequency control word; The compensated frequency control word is loaded into the NCO's frequency control register to obtain the initial NCO clock.
6. The FPGA-based digital signal processing method according to claim 1, characterized in that, In response to the lock attainment flag being high, the proportional and integral coefficients of the high-level triggered loop filter are smoothly switched from high-bandwidth mode to low-bandwidth mode based on a multi-step interpolation state machine to obtain a high-precision recovery clock and time interval error measurement stream, including: In response to detecting that the lock achievement flag is high, high proportional coefficient, high integral coefficient, low proportional coefficient, and low integral coefficient are extracted from high bandwidth coefficient and low bandwidth coefficient; Incremental calculations are performed on high proportional coefficients, high integral coefficients, low proportional coefficients, and low integral coefficients to obtain coefficient value increments composed of proportional increments and integral increments; The multi-step interpolation state machine uses the coefficient value increment to interpolate the loop filter coefficients step by step so that the proportional coefficient and integral coefficient of the loop filter are smoothly switched to low bandwidth mode. The original phase error stream input is switched to a loop filter in low-bandwidth mode to obtain the time interval error measurement stream; The NCO performs phase adjustment based on the time interval error measurement stream to obtain the high-precision recovery clock.
7. The FPGA-based digital signal processing method according to claim 1, characterized in that, Based on high-precision clock and time interval error measurement streams, data deserialization and jitter analysis are performed on real-time digital signal streams to obtain deserialized data output and jitter analysis results, including: The real-time digital signal stream is sampled using the optimal sampling phase of the high-precision recovery clock to obtain the deserialized data output. The time interval error measurement stream is input into the real-time jitter analysis engine to obtain the jitter analysis results.
8. The FPGA-based digital signal processing method according to claim 7, characterized in that, The time interval error measurement stream is input into the real-time jitter analysis engine to obtain the jitter analysis results, including: Based on histogram configuration parameters, a real-time histogram is constructed on the time interval error measurement stream to obtain the TIE histogram vector. Extract the parameters of the dual Dirac model from the TIE histogram vector; The parameters of the dual Dirac model are input into the dual Dirac jitter model to obtain jitter analysis results consisting of deterministic jitter values, random jitter values, and total jitter values.
9. A digital signal processing system based on FPGA, characterized in that, include: The signal digitization frequency analysis module is used to perform signal digitization and feedforward frequency analysis on the acquired analog high-speed signals to obtain the frequency offset value; The phase error signal generation module is used to perform CDR loop preset and high bandwidth lock-on activation on the real-time digital signal stream based on the frequency offset value to obtain the phase error signal. The hardware verification module is used to perform hardware verification of the lock state of the continuously monitored phase error signal to obtain a lock achievement flag. The mode switching module is used to smoothly switch the proportional coefficient and integral coefficient of the high-level triggered loop filter from high-bandwidth mode to low-bandwidth mode based on a multi-step interpolation state machine in response to the lock achievement flag being high-level, so as to obtain a high-precision recovery clock and time interval error measurement stream. The deserialization jitter analysis module is used to perform data deserialization and jitter analysis on real-time digital signal streams based on high-precision recovery of clock and time interval error measurement streams, so as to obtain the deserialized data output and jitter analysis results.
Citation Information
Cited By
Real-time vector signal analyzer based on FPGA and signal processing method
CN122640050A