Low-delay sequence detection device for cpo photoelectric signal reception

By using an adaptive module and a partitioned look-ahead parallel detector, combined with partitioning and branching metric calculations, the high latency and high complexity of PAM4 signal detection in CPO photoelectric signal reception are solved, achieving low latency and low power consumption detection results.

CN121841465BActive Publication Date: 2026-05-19NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NAT UNIV OF DEFENSE TECH
Filing Date
2026-03-12
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In existing technologies for CPO optoelectronic signal reception, PAM4 signal detection suffers from high latency, complexity, and power consumption, which limits its application in high-speed serial interfaces.

Method used

An adaptive module and multiple partitioned look-ahead parallel detectors are employed, including a partition and branch metric calculation module, a grid merging module, and a decision module. By adaptively adjusting the tap coefficients and ISI coefficients, combined with partitioned look-ahead parallel detection technology, the complexity and latency of PAM4 signal sequence detection are reduced.

Benefits of technology

This study reduces the complexity, power consumption, and delay of PAM4 signal sequence detection in CPO photoelectric signal reception, thereby improving detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121841465B_ABST
    Figure CN121841465B_ABST
Patent Text Reader

Abstract

The application discloses a low-delay sequence detection device for CPO photoelectric signal reception, comprising an adaptive module and a plurality of partitioned forward-looking parallel detectors, the partitioned forward-looking parallel detectors comprising partitioned and branch metric calculation modules, trellis merging modules and decision modules connected in sequence, the partitioned and branch metric calculation modules being used for partitioning and branch metric calculation of PAM4 signals based on tap coefficients and ISI coefficients provided by the adaptive module for input signals, the trellis merging modules being used for trellis signal merging calculation of path metrics of synchronization blocks and backtracking blocks according to obtained branch metrics, and the decision modules being used for decision of input signals to obtain binary PAM4 signals according to obtained path metrics, and the adaptive module being used for adaptive update of tap coefficients and ISI coefficients. The application aims at CPO photoelectric signal reception while reducing the complexity, power consumption and delay of PAM4 signal sequence detection technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of PAM4 signal detection technology in co-packaged optics (CPO), and specifically to a low-delay sequence detection device for CPO photoelectric signal reception. Background Technology

[0002] Currently, the data deluge driven by artificial intelligence and high-performance computing is propelling data rates towards 1.6T / 3.2T, posing bottlenecks for traditional pluggable optical modules in terms of power consumption, density, and signal integrity. Against this backdrop, co-packaged optics (CPO) technology has emerged. Its core advantage lies in co-packaging the optical engine with the computing chip, significantly shortening the electrical interconnect distance, thereby substantially reducing system power consumption (estimated at 30%-50%) and improving port density and transmission performance.

[0003] In CPO technology, the electrical interface employs parallel high-speed channels with extremely low voltage swing (e.g., 224G PAM4) to achieve ultra-high bandwidth density. This makes the signal integrity issues it faces different from those of traditional pluggable modules' board-level electrical channels. The focus of signal integrity issues therefore shifts from board-level link attenuation to package-level interconnect optimization: firstly, the increased inter-symbol crosstalk due to link insertion loss under ultra-wideband conditions; secondly, near-end crosstalk between extremely high-density parallel channels; and thirdly, power integrity and simultaneous switching noise arising from the shared power supply system between the high-speed SerDes and the optical engine. Therefore, CPO technology presents new challenges for PAM4 signal detection. Traditional maximum likelihood sequence detection (MLSD) has demonstrated its advantage in efficiently utilizing inter-signal correlations in four-level pulse amplitude modulation (PAM4) signal detection, thus providing better bit error rate performance in the presence of inter-symbol interference (ISI). However, MLSD also suffers from significant resource consumption, high power consumption, and large detection delays, which limit its widespread application in high-speed PAM4 receivers. While low-complexity sequence detection techniques for PAM4 signals exist, reducing the number of surviving paths to decrease the computational complexity of maximum likelihood sequence detection, thus laying the foundation for reducing resource consumption and power consumption, the high latency problem of maximum likelihood sequence detection remains unresolved. As signal rates increase, the latency issue of maximum likelihood sequence detection algorithms becomes particularly prominent, limiting their application in high-speed serial interfaces. Current research has explored solutions to the high latency problem of traditional maximum likelihood sequence detection. By optimizing the mesh structure of maximum likelihood sequence detection and implementing parallel processing of the Add-Compare-Select (ACS) operation, the sequence computation latency of maximum likelihood sequence detection can be effectively shortened. Although the above methods reduce detection latency, they are mainly geared towards non-return-to-zero (NRZ) signals. When applied to PAM4 signals, they increase complexity and power consumption. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a low-latency sequence detection device for CPO photoelectric signal reception, which addresses the above-mentioned problems of the prior art. The present invention aims to reduce the complexity, power consumption and latency of PAM4 signal sequence detection technology for CPO photoelectric signal reception.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0006] A low-latency sequence detection device for CPO photoelectric signal reception includes an adaptive module and multiple partitioned look-ahead parallel detectors. Each partitioned look-ahead parallel detector includes a partitioning and branching metric calculation module, a grid merging module, and a decision module connected in sequence. The partitioning and branching metric calculation module is used to process the input signal... The PAM4 signal partitioning and branching metrics are calculated based on the tap coefficients and ISI coefficients provided by the adaptive module; the grid merging module is used to calculate the path metrics of synchronization blocks and backtracking blocks by merging grid signals according to the obtained branch metrics; the decision module is used to evaluate the input signal based on the obtained path metrics. The decision is made to obtain a binary PAM4 signal. The adaptive module is used to determine the binary PAM4 signal output by the decision module. The generation partition and branch metric calculation module is used to perform the tap coefficients and ISI coefficients required for signal detection.

[0007] Optionally, the input signal of the partition and branch metric calculation module The data consists of 32 channels of 8-bit data. The partitioning and branching metric calculation module calculates 4 sets of branch metrics, and the mesh merging module calculates 8 path metrics for the synchronization block and backtracking block. The final decision module outputs a binary PAM4 signal. It is a 64-bit signal.

[0008] Optionally, the partition and branch metric calculation module includes: an expected calculation unit, two partition units, and an adder, wherein the expected calculation unit is used to calculate according to the following formula. 16 expected levels at the time :

[0009] ;

[0010] in, and They are respectively Time and The desired level at time t, each desired level contains four combinations from 0 to 3, resulting in a total of 16 possible combinations of desired levels. The ISI coefficient;

[0011] And the input signal is based on the following formula. Calculate the expected input :

[0012] ;

[0013] in, This is the tap factor;

[0014] The partitioning unit is used to determine the expected input. 8 expected levels and the partition number of the input signal id Select two expected levels for partitioning. Used for branch metrics Calculation;

[0015] The adder is used to calculate the branch metric according to the following formula. :

[0016] ;

[0017] in, for Branching metric at time , for Expected level at time, for Expected input at any given time.

[0018] Optionally, the partitioning unit includes:

[0019] Comparator circuit, used to compare the expected input The decision levels dLev0=-1 and dLev1=1 of the partition are compared to determine whether the signal is numbered 0 or 1 in the partition. The PAM4 signal with constellation set {-3,-1,1,3} is divided into two partitions numbered 0 and 1, where partition 0 contains {-3,1} and partition 1 contains {-1,3}.

[0020] The selection circuit is used to first select from 8 expected levels for each partition through four selectors (MUX) based on the comparison result obtained from the comparison circuit. Four expected levels are selected from the input signal for subsequent calculations, and then the partition numbering is used as the basis for the calculations. id By using two selectors (MUX) to remove two expected levels based on the partitioning results at the next time step, two expected levels are ultimately selected for each partition. Used for branch metrics The calculation.

[0021] Optionally, the mesh merging module includes a five-layer improved addition-comparison-selection operation module, and the partitioning and branching metric calculation module outputs 32 sets of branching metrics. bm and the corresponding partition number of the input signal The 16 improved add-compare-select operation modules sent to the first level each measure the two sets of branches according to the partition look-ahead algorithm. bm and the corresponding partition number of the input signal Merge into a set of branch metrics bm and the corresponding partition number of the input signal Then it is fed into the next level of 8 improved addition-comparison-selection operation modules, and finally a set of branch metrics is obtained through the fifth level of 1 improved addition-comparison-selection operation module. bm and the corresponding partition number of the input signal .

[0022] Optionally, the improved add-compare-select operation module includes:

[0023] Two 2-bit carry-lookahead adders (CLAs) are used to implement 2-bit branch addition operations, performing addition operations on the four branches of the input;

[0024] An improved two-bit carry-lookahead adder MCLA is used to implement a comparison operation with two-bit branches. The comparison operation is performed on the addition results of two two-bit carry-lookahead adders CLA to obtain the comparison result.

[0025] Two selectors MUX, one of which is used to select the smaller branch metric from the addition results of the two two-bit carry-lookahead adders CLA based on the comparison result of the improved two-bit carry-lookahead adder MCLA. bm The output, another selector MUX, is used to partition the corresponding two input signals in four branches based on the comparison result of the improved two-bit carry-lookahead adder MCLA. Choose a smaller branch metric bm Corresponding partition number Output.

[0026] Optionally, the two-bit carry-lookahead adder (CLA) is a logic gate device.

[0027] Optionally, the logical expression of the improved two-bit carry-lookahead adder MCLA is:

[0028] ;

[0029] in, and For the input operands, for The result obtained through an inverter, and the carry output flag of the improved two-bit carry-lookahead adder MCLA. As the selection input for the two selectors MUX.

[0030] Optionally, the decision module includes:

[0031] The add-compare-select operation module is used to add the path metric output by the mesh merging module. The path metrics of the synchronization block and the backtracking block are added separately. After performing an add-compare-select operation on the current path metric and the synchronization block and the backtracking block, the final four path metrics are output. The synchronization block and the backtracking block are the merged path metrics cached in the previous and next time steps, respectively.

[0032] The decision circuit is used to output the final four path metrics from the add-compare-select operation module. Make a decision to find the path with the minimum metric. Corresponding partition number Output;

[0033] The selector MUX is used to determine the corresponding partition number output by the decision circuit. Output the corresponding partition numbers for the four paths. The corresponding path is output and converted into a binary PAM4 signal.

[0034] Optionally, the adaptive module includes:

[0035] The lookup table module is used to output the decision results in each cycle. Eight sets of 2-bit results are selected, and the input 2-bit detection results are converted into 8-bit signed numbers using a lookup table. Then follow the formula below:

[0036]

[0037] Generate an estimated signal, where For the first One expected level, For the first -1 desired level, The first ISI parameter is then set as a post-processor, and the updated tap coefficients and ISI coefficients are generated by combining them with the expected input signal according to the following formula:

[0038] ;

[0039] ;

[0040] in, The updated tap coefficients, For the input tap coefficient, These are the step parameters used to adjust the convergence speed. For input signal, The updated ISI coefficients, These are the input ISI coefficients.

[0041] Compared with existing technologies, the present invention mainly achieves the following beneficial effects: The present invention includes an adaptive module and multiple partitioned look-ahead parallel detectors. The partitioned look-ahead parallel detectors include a partition and branch metric calculation module, a grid merging module, and a decision module connected in sequence. On the one hand, the partitioned look-ahead parallel detectors combine the principles of partitioning and grid merging, reducing the number of surviving paths while further reducing the sequence detection latency, and propose a hybrid expected value calculation method to reduce the computational load. On the other hand, the partitioned and branch metric calculation modules are used to calculate the input signal... The PAM4 signal partitioning and branching metrics are calculated based on the tap coefficients and ISI coefficients provided by the adaptive module. The adaptive module is used to calculate the binary PAM4 signal output by the decision module. The generation of partition and branch metric calculation modules for the tap coefficients and ISI coefficients required for signal detection is achieved through adaptive mixing parameters, which dynamically adjust the expected value calculation under the mixing mode. This allows the present invention to simultaneously reduce the complexity, power consumption, and latency of PAM4 signal sequence detection technology for CPO photoelectric signal reception. Attached Figure Description

[0042] Figure 1 This is a schematic diagram of the low-latency sequence detection device in an embodiment of the present invention.

[0043] Figure 2 This is a schematic diagram of the partition and branch metric calculation module in an embodiment of the present invention.

[0044] Figure 3 This is a schematic diagram of the partition unit in an embodiment of the present invention.

[0045] Figure 4 This is a schematic diagram of the grid merging module in an embodiment of the present invention.

[0046] Figure 5 This is a schematic diagram of the improved addition-comparison-selection operation module (MACS) in an embodiment of the present invention.

[0047] Figure 6 This is a schematic diagram illustrating the principle of the improved add-compare-select operation module (MACS) and its two-bit carry-lookahead adder CLA (2-bit CLA) and improved two-bit carry-lookahead adder MCLA (2-bit MCLA) in the embodiments of the present invention.

[0048] Figure 7 This is a schematic diagram of the decision module in an embodiment of the present invention.

[0049] Figure 8 This is a schematic diagram of the adaptive module in an embodiment of the present invention.

[0050] Figure 9 The following are block diagrams of the mesh structure of MLSD and the structure of ACS used for comparison in the embodiments of the present invention. Figure 9 (a) shows the mesh structure of MLSD. Figure 9 (b) is a structural block diagram of ACS.

[0051] Figure 10 The circuit diagram shown is of a conventional 2-bit ACS circuit used as a comparison in this embodiment of the invention.

[0052] Figure 11 As included in the embodiments of the present invention to 3-state NRZ signal grid diagram at time 1.

[0053] Figure 12 This is the MLSD process for the lookahead partitioning of 2-state PAM4 signals in an embodiment of the present invention. Figure 12 (a) is a 2-state PAM4 signal. Figure 12 (b) For partitioned calculations k At the specified time point, a partitioning operation is performed, reducing the number of branches from 16 to 8. Figure 12 (c) is for partitioned calculations. k+1 Continuously perform partitioning operations, reducing the previous grid to 4 branches and the next grid to 8 branches. Figure 12 (d) To perform a partitioning operation at time k+2, the subsequent grid is reduced to 4 branches. Figure 12 (e) is the grid merging operation, which selects paths from k to k+2 that pass through different k+1 and performs an addition-comparison-selection operation. Figure 12 (f) represents the four branches that survived the grid merging.

[0054] Figure 13 For comparison of sliding block technology in the embodiments of the present invention, Figure 13 (a) shows the working process of traditional sliding block technology. Figure 13 (b) is the working process of the partitioned look-ahead sliding block structure used in the partitioned look-ahead parallel detector of the present invention. Detailed Implementation

[0055] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings in the embodiments of the present invention.

[0056] like Figure 1 As shown, the low-latency sequence detection device for CPO photoelectric signal reception in this embodiment includes an adaptive module and multiple partitioned look-ahead parallel detectors (APLSDs). Each APLSD includes a partitioning and branching metric calculation module, a grid merging module, and a decision module connected in sequence. The partitioning and branching metric calculation module is used to calculate the input signal... The PAM4 signal partitioning and branching metrics are calculated based on the tap coefficients and ISI coefficients provided by the adaptive module; the grid merging module is used to calculate the path metrics of synchronization blocks and backtracking blocks by merging grid signals according to the obtained branch metrics; the decision module is used to evaluate the input signal based on the obtained path metrics. The decision is made to obtain a binary PAM4 signal. The adaptive module is used to determine the binary PAM4 signal output by the decision module. The generation partition and branch metric calculation module is used to perform the tap coefficients and ISI coefficients required for signal detection.

[0057] See Figure 1 In this embodiment, the partition look-ahead parallel detector uses a data sequence length of 32. The input 64 channels of 8-bit data are fed into two parallel partition look-ahead parallel detectors (APLSDs). The APLSDs employ pipelined technology, reducing the number of parallel blocks by caching the results of each add-compare-select (ACS) operation. Therefore, only two APLSDs are needed to handle half of the 64-channel data sequence. The input signals for the partition and branch metric calculation modules... For 32 channels of 8-bit data (32×8), the partitioning and branching metric calculation module calculates 4 sets of branch metrics, and the grid merging module calculates the path metrics Post for the synchronization block and backtracking block. and Pre The final decision module outputs a binary PAM4 signal with a length of 8. The signal is 64 bits. Finally, a 128-bit binary result is output through two partitioned look-ahead parallel detectors (APLSD). The path metric length for the synchronization block and the backtracking block is chosen to be 8. This is because it is exactly a power of 2 and greater than the state length of 5.

[0058] For signal constellation {0 , 1 , 2 , The PAM4 signal of 3} has a 2-state consisting of values ​​from 00 to 100. ~ There are 16 possible combinations in total (33). Traditional 2-state MLSD requires calculation based on the ISI coefficient. h 0 ,h1} Calculate the state value x Corresponding expected signals:

[0059] ;

[0060] The accuracy of the ISI estimate will affect the obtained The accuracy of the calculation affects the results of subsequent partitioning algorithms. It can be seen that the calculation of each path requires two multiplications and one addition. Alternatively, a feedforward equalizer (FFE) can be added before the MLSD, using the input signal for feedforward equalization to generate the equalized signal. ,Right now:

[0061] .

[0062] Since only one equalization signal calculation is needed for every 16 paths, the computational complexity can be greatly reduced. However, because MLSD requires the use of inter-symbol ISI information, this method results in the loss of inter-state ISI information, leading to a decrease in detection accuracy. To address the above technical problems, this embodiment proposes a hybrid method for calculating the expected signal, such as... Figure 2 As shown, the partition and branch metric calculation module includes: an expected calculation unit, two partition units, and an adder. The expected calculation unit is used to calculate according to the following formula. 16 expected levels at the time :

[0063] ;

[0064] in, and They are respectively Time and The desired level at time t, each desired level contains four combinations from 0 to 3, resulting in a total of 16 possible combinations of desired levels. The ISI coefficient;

[0065] And the input signal is based on the following formula. Calculate the expected input :

[0066] ;

[0067] in, This is the tap factor;

[0068] The partitioning unit is used to determine the expected input. 8 expected levels and the partition number of the input signal id Select two expected levels for partitioning. Used for branch metrics Calculation;

[0069] The adder is used to calculate the branch metric according to the following formula. :

[0070] ;

[0071] in, for Branching metric at time , for Expected level at time, for Expected input at time, For the desired level, For ISI parameters, The main tap coefficients are used. This hybrid parameter technique still only needs to consider the 2-state ISI signal without increasing complexity, while reducing computational complexity. Figure 2 In the middle, the expected computing unit The two digits entered are Time and Expected level at time Its output is the expected level of different branches. ,in ~ The expected levels for branches 1 through 16 are respectively. .in, ~ and ~ Enter the first partition unit, ~ and ~ Enter the second partition unit. and This indicates the numbers of grid 1 partition 0 and partition 1. and This indicates the numbers of grid 0 partition 0 and partition 1. This represents the input level at time 0. ~ This represents the expected levels for grid 0 partitions 0 to 0 and partitions 1 to 1. ~ This represents the branch metric between partitions 0 and 0 in grid 0 and between partitions 1 and 1. For a 2-state PAM4 signal with 16 branches, the traditional expected signal calculation method requires 32 multiplications and 16 additions, while the hybrid parameter calculation method only requires 17 multiplications and 16 additions, reducing the amount of multiplication calculation by approximately 47%.

[0072] like Figure 3 As shown, the partitioning unit includes:

[0073] Comparator circuit, used to compare the expected input The decision levels dLev0=-1 and dLev1=1 of the partition are compared to determine whether the signal is numbered 0 or 1 in the partition. The PAM4 signal with constellation set {-3,-1,1,3} is divided into two partitions numbered 0 and 1, where partition 0 contains {-3,1} and partition 1 contains {-1,3}.

[0074] The selection circuit is used to first select from 8 expected levels for each partition through four selectors (MUX) based on the comparison result obtained from the comparison circuit. Four expected levels are selected for subsequent calculations, for example, the eight expected levels for the first partition unit. for: ~ and ~ Then, based on the partition numbering of the input signal... id By using two selectors (MUX) to remove two expected levels based on the partitioning results at the next time step, two expected levels are ultimately selected for each partition. Used for branch metrics The calculation. Compared to traditional MLSD, which requires calculating 16 sets. bm The partitioning method proposed in this embodiment reduces the computational complexity by 75%.

[0075] like Figure 4 As shown, the mesh merging module includes a five-layer improved addition-comparison-selection operation module (MACS), and the partitioning and branching metric calculation module outputs 32 sets of branching metrics. bm and the corresponding partition number of the input signal The 16 Improved Add-Compare-Select Operation Modules (MACS) sent to the first level each use a partition lookahead algorithm to measure the two sets of branches. bm and the corresponding partition number of the input signal Merge into a set of branch metrics bm and the corresponding partition number of the input signal Then it is fed into the next level of 8 improved addition-comparison-selection operation modules (MACS), and finally a set of branch metrics is obtained through the fifth level of 1 improved addition-comparison-selection operation module (MACS). bm and the corresponding partition number of the input signal The partitioning and branching metric calculation module outputs 32 sets of branch metrics. bm and corresponding indexes The data is fed into the mesh merging module, which then merges adjacent meshes sequentially according to the partition lookahead algorithm. The data set requires 5 steps to merge, with each mesh merging step potentially running in parallel. Eight branches from adjacent meshes are sequentially processed using the ACS operation based on nodes with the same starting and ending points, resulting in four merged branches. Figure 4 Taking the 16 improved addition-compare-select operation modules (MACS) in the first level on the upper middle side as an example, its input... and This represents the branch measure of group 1 (32 groups). bm and the corresponding partition number of the input signal , and This indicates the branch measure of group 2 (group 32). bm and the corresponding partition number of the input signal Its output and This represents the branch measurement of groups 1-2 (32 groups). bm and the corresponding partition number of the input signal The merging result. Taking an improved add-compare-select operation module (MACS) at the fifth layer as an example, its input is... and This represents the 32 branch measures from group 1 to group 15. bm and the corresponding partition number of the input signal The result of the merger and This represents the branch measures of groups 16 to 32. bm and the corresponding partition number of the input signal The merged result, output and Represents 32 groups of branch metrics bm and the corresponding partition number of the input signal The result of the merger.

[0076] like Figure 5 As shown, the improved addition-comparison-selection operation module (MACS) includes:

[0077] Two 2-bit carry-lookahead adders (CLAs) are used to implement 2-bit branch addition operations, performing addition operations on the four branches of the input;

[0078] An improved two-bit carry-lookahead adder MCLA is used to implement a comparison operation with two-bit branches. The comparison operation is performed on the addition results of two two-bit carry-lookahead adders CLA to obtain the comparison result.

[0079] Two selectors MUX, one of which is used to select the smaller branch metric from the addition results of the two two-bit carry-lookahead adders CLA based on the comparison result of the improved two-bit carry-lookahead adder MCLA. bm The output, another selector MUX, is used to partition the corresponding two input signals in four branches based on the comparison result of the improved two-bit carry-lookahead adder MCLA. Choose a smaller branch metric bm Corresponding partition number Output.

[0080] Figure 5 middle, This indicates the branch metric from partition 0 to partition 0 of grid 1. This indicates the branch metric from partition 0 to partition 0 of grid 0. This indicates the branch metric from partition 0 to partition 1 in grid 0. This indicates the branch metric from partition 1 to partition 0 of grid 1. This indicates the branch number of grid 0 partition 0. This indicates the branch number of grid 0 partition 1. This represents the branch metric from partition 1 to partition 0 after merging grids 0 and 1. This indicates the branch number of partition 0 after merging grids 0 and 1.

[0081] In this embodiment, the two-bit carry-lookahead adder (CLA) is a logic gate device that can be used to calculate the addition result and carry flag.

[0082] In this embodiment, the logical expression of the improved two-bit carry-lookahead adder MCLA is as follows:

[0083] ;

[0084] in, and For the input operands, for The result obtained through an inverter, and the carry output flag of the improved two-bit carry-lookahead adder MCLA. As the selection input for the two selectors MUX, the comparison operation can be implemented by subtraction, thereby reducing logical latency by parallelizing the addition and comparison operations. In this embodiment, a low-latency improved add-compare-select operation module (MACS) is used in sequence detection, which reduces branch calculation latency compared to the traditional add-compare-select operation module ACS structure. Experiments show that the 8-bit MACS reduces the calculation latency by 31.3% compared to ACS.

[0085] Figure 6This is a schematic diagram illustrating the improved Add-Compare-Select Operation Module (MACS) and its 2-bit carry-lookahead adder CLA (CLA) and 2-bit carry-lookahead adder MCLA (MCLA) in an embodiment of the present invention. The dashed arrows indicate the circuit paths of the 2-bit CLA and 2-bit MCLA.

[0086] In the circuit section of the 2-bit CLA: pm(1) n-1 [0] and pm(1) n-1 [1] represents the 0th bit of the path metric for level 0 at time n-1 and the 1st bit of the path metric for level 1 at time n-1, bm(1,-3) n [0] and bm(1,-3) n [1] represents the 0th bit of the branch metric from level 1 to level -3 at time n and the 1st bit of the branch metric from level 1 to level -3 at time n, G0 represents the 0th bit generated signal, P0 represents the 0th bit transmitted signal, pm(1,-3) n [0] and pm(1,-3) n [1] represents the 0th bit of the path metric from level 0 to level -3 at time n and the 1st bit of the path metric from level 1 to level -3 at time n, G1 represents the 1st bit generated signal, P1 represents the 1st bit transmitted signal, C in [0] indicates the 0th bit of the input carry signal, C out This indicates the carry signal output.

[0087] In the circuit section of the improved Add-Compare-Select Operation Module (MACS), pm(1) n-1 The path metric for level 1 at time n-1 is represented by bm(1,-3). n The branch measure representing the level from 0 to -3 at time n is bm(-3,-3). n The branch measure representing the level from -3 to -3 at time n is pm(-3). n-1 The path metric pm(1,-3) represents the level of -3 at time n-1. n The path metric representing the level from 1 to -3 at time n is pm(-3,-3). n The path metric from level -3 to -3 at time n is represented by Sel, which is the selection signal (carry signal) emitted by the improved 2-bit carry-lookahead adder MCLA (2-bit MCLA), and pm(-3). n The path metric represents the level of -3 at time n.

[0088] In the 2-bit MCLA part, pm(1,-3) n [0] and pm(1,-3) n[1] represents the 0th bit of the path metric from level 1 to level -3 at time n and the 1st bit of the path metric from level 1 to level -3 at time n, pm(-3,-3) n [0] and pm(-3,-3) n [1] represents the 0th bit of the path metric from level -3 to level -3 at time n and the 1st bit of the path metric from level -3 to level -3 at time n.

[0089] like Figure 7 As shown, the decision module includes:

[0090] The Add-Compare-Select Operation Module (ACS) is used to process the path metrics output by the grid merging module. The path metrics of the synchronization block and the backtracking block are added separately. After performing an add-compare-select operation on the current path metric and the synchronization block and the backtracking block, the final four path metrics are output. The synchronization block and the backtracking block are the merged path metrics cached in the previous and next time steps, respectively.

[0091] The decision circuit is used to output the final four path metrics from the Add-Compare-Select (ACS) operation module. Make a decision to find the path with the minimum metric. Corresponding partition number Output;

[0092] The selector MUX is used to determine the corresponding partition number output by the decision circuit. Output the corresponding partition numbers for the four paths. The corresponding path is output and converted into a binary PAM4 signal. Figure 7 Post and Pre These are the synchronization block and the backtrack block, representing the merged path metrics cached at the previous and next time steps, respectively. After performing an ACS operation on the current path metric, the synchronization block, and the backtrack block, the final four results are output. Compare these four Having the smallest path The signal is output and converted to a binary PAM4 signal. (See diagram bm) 0-31 [0:3] represents the branch paths numbered 0-3 after merging grids 0-31, id 0-31

[00] ~id 0-31

[11] represents the four branch paths from partition {0,1} to partition {0,1} after merging grids 0-31, d out [0:63] represents a 64-bit binary PAM4 signal.

[0093] The partition look-ahead algorithm in this embodiment requires both tap coefficients and ISI parameters. The adaptive module in this embodiment uses a hybrid parameter adaptive algorithm to achieve simultaneous adaptive adjustment of the two parameters.

[0094] like Figure 8 As shown, the adaptive module includes:

[0095] The lookup table module is used to output the decision results in each cycle. Eight sets of 2-bit results are selected, and the input 2-bit detection results are converted into 8-bit signed numbers using a lookup table. ,For example Figure 8 3, 1, -1, and -3 are 8'h40, 8'h15, 8'heb, and 8'hc0 respectively; then follow the formula below:

[0096]

[0097] Generate an estimated signal, where For the first One expected level, For the first -1 desired level, The first ISI parameter is then set as a post-processor, and the updated tap coefficients and ISI coefficients are generated by combining them with the expected input signal according to the following formula:

[0098] ;

[0099] ;

[0100] in, The updated tap coefficients, For the input tap coefficient, These are the step parameters used to adjust the convergence speed. For input signal, The updated ISI coefficients, These are the input ISI coefficients.

[0101] In this embodiment, the adaptive module employs a hybrid parameter adaptive algorithm to simultaneously adjust two parameters. Parameter adaptation requires the use of detection results. The decision unit generates 64 detection results per cycle, theoretically producing 64 sets of updated adaptive coefficients per cycle. Since updating 64 sets of coefficients at once requires extensive computation, selecting only some parameters not only does not affect parameter convergence but also saves significant resources. This invention selects 4 sets of decision results per cycle for parameter updates. The decision results output in each cycle... In, select as Figure 11The input 2-bit detection result is converted into an 8-bit signed number using a lookup table, as shown in the figure. Then, an estimated signal is generated, and the updated tap coefficients and ISI coefficients can be generated according to the formula mentioned above. Figure 8 In the middle, d out [18:19] represents bits 18 and 19 of the binary PAM4 signal (decision result), and so on, d out [18:19] The 9th desired level can be calculated. Then according to:

[0102]

[0103] By generating the estimated signal, we can obtain .

[0104] Figure 9 The following are block diagrams of the mesh structure of MLSD and the structure of ACS used for comparison in embodiments of the present invention, wherein... Figure 9 (a) shows the mesh structure of MLSD. Figure 9 (b) is a structural block diagram of ACS. For example... Figure 9 As shown in (a), in the MLSD mesh structure, the two branches of each destination node are compared using an ACS operation to determine which branch to retain, as follows: Figure 9 (b) is shown by the solid line. The computational delay of ACS operations in the grid is defined as... For a length of A sequence whose grid contains nodes arrive A series of nodes. Path metric of time node Path metric depends on the node at the previous time step Since the detection results of MLSD depend on the path metric of the terminal node. Therefore, the calculation of path metrics exhibits causality. This makes the total latency of MLSD affected by the computational latency of each node. In addition, MLSD also needs to consider the comparison... The process involves backtracking the path to obtain the result sequence. The delay of the backtracking process can be defined as... The total latency of MLSD is defined as... Therefore, for a length of The total detection delay of the MLSD sequence can be expressed as:

[0105] ,

[0106] in, The computational delay for ACS operations, This represents the delay in the backtracking process. Based on the above formula, to improve the overall latency of the MLSD structure, two approaches can be taken: first, reduce the computational latency of ACS operations. Secondly, improve the sequential computation mode of MLSD to enhance its parallel processing capabilities, thereby eliminating the need for backtracking and reducing... By reducing the latency of individual ACS operations and increasing the parallelization of ACS operations, the total latency of path detection can be effectively reduced. According to the formula above, the total latency of MLSD is related to the computational latency of ACS operations. The relationship is linear. Therefore, reducing the computational latency of ACS operations is crucial for reducing the overall latency in the entire sequence detection process. From a circuit implementation perspective, analyzing the ACS module and optimizing its critical paths can effectively shorten its computational latency. Figure 10 This is a circuit schematic of a conventional 2-bit ACS circuit used for comparison in an embodiment of the present invention. In this circuit, 2-bit branch measurement... and path metrics The logic gates are added using a full adder, and the results are then compared in a comparator. The comparison result is used to select the optimal path. Since the delay differences between logic gates are relatively small, the number of gate levels has a more significant impact on the overall delay. Furthermore, because routing methods vary considerably under different implementations (ASIC or FPGA) and routing strategies, it is difficult to consider them using a unified model; therefore, this section focuses only on logic delay. In ACS operation, the longest logic propagation path is as follows: Figure 10 The logic devices marked in black in the middle contain a total of 10 levels of logic gates. Among them, the 2-bit full adder contains 4 levels of logic gates, the comparator contains 5 levels, and the selector contains 1 level.

[0107] In this embodiment, the following is adopted: Figure 6 The improved Add-Compare-Select (MACS) module shown, along with its 2-bit carry-lookahead adder CLA and 2-bit carry-lookahead adder MCLA, reduces logic latency by parallelizing addition and comparison operations. Specifically, the comparison operation can be implemented using subtraction, i.e. Therefore, a comparison structure consisting of adders can be constructed, where the carry output flag is set. This circuit uses a carry-look-ahead adder (CLA) to implement the adder and a modified 2-bit carry-look-ahead adder (MCLA) to implement the comparator. Unlike traditional comparators, the modified 2-bit carry-look-ahead adder (MCLA) utilizes… The ability to obtain the result without waiting for all bits to complete calculation shortens the longest logical path. In the improved Add-Compare-Select operation module (MACS), the longest logical path is as follows: Figure 3 The logic devices marked in black in the image contain a total of 8 logic gates. Specifically, the 2-bit CLA contains 2 logic gates, the comparator contains 5 levels, and the selector contains 1 level. Compared to the traditional ACS, the improved Add-Compare-Select Operation Module (MACS) reduces the number of logic gates by 2. Although the number of logic gates is only reduced by 2 in the comparison between the 2-bit ACS and MACS circuits, this approach optimizes the longest path in ACS operations. Therefore, as the operand width increases, the optimization effect on the number of logic gates becomes more significant.

[0108] The traditional Viterbi algorithm calculates branch metrics sequentially. However, research indicates that this traditional method suffers from redundant computation of shared paths, leading to additional complexity and longer latency. To address these issues, look-ahead MLSD (LA-MLSD) structures are beginning to be applied to NRZ signals.

[0109] Figure 11 As included in the embodiments of the present invention to A trellis diagram of the 3-state NRZ signal at time t. For modulation levels of t. The signal when the grid merging steps In this case, grid merging does not require consideration of the comparison-selection process. Figure 11 Step 1 in the process (i.e. Only the branches of two adjacent grids need to be merged. As you can see, originally each state node had only two incoming branches, but after merging, the number of incoming branches for each node increases to four. When The branch merging process will be described in detail using the solid lines representing steps 2 in the diagram. Before introducing branch merging, a key principle will be introduced:

[0110] Generalized optimality principle: Assume S and These represent the node states of two different paths. and These are the start node time and the arrival node time, respectively. When the starting node and the destination node of the path are the same, that is... and If the following conditions are met:

[0111] ;

[0112] in, This represents the branch metric from the S partition at time k+1 to the S partition at time k+1. The branch metric represents the distance from partition S' at time k+1 to partition S' at time k+1; then the path can be determined. It cannot be the result sequence of MLSD.

[0113] According to the principle of generalized optimality Figure 11 The merged branch metric, represented by the black line, can be expressed as:

[0114] ;

[0115] in, This represents the path metric from time n-4 to time n, from level 3 to level 3. This represents the path metric from time n-4 to time n-2, from level 3 to level 3. This represents the path metric from time n-2 to time n-0, from level 3 to level 3. This represents the path metric from time n-4 to time n, from level 3 to level 2. This represents the path metric from time n-2 to time n-0, from level 2 to level 3. It can be seen that LA-MLSD's mesh merging can start from any point in the signal sequence, regardless of previous or subsequent influences. Therefore, it can significantly reduce detection latency through parallel operation. Assuming the length of the sequence to be detected... Each ACS operation requires one clock cycle. Traditional MLSD requires calculating path metrics sequentially from the starting node. Then the judgment is completed, therefore the detection delay is... LA-MLSD reduces latency to [a lower level] through parallel operations. Here, It is not mandatory for the value to be a power of 2; it can be of any length. Both can utilize look-ahead algorithms and exhibit significant latency reduction. Although LA-MLSD effectively reduces latency, the mesh merging operation still requires substantial computation. The number of parallel paths increases with the modulation level. M The computational complexity increases exponentially. Therefore, look-ahead algorithms are currently mainly used in MLSD for NRZ signals and not in multi-level modulated signals such as PAM4.

[0116] Compared to NRZ signals, PAM4 signals result in an exponential increase in the number of branches when applied to traditional MLSD. Figure 12Taking the 2-state PAM4 signal shown in (a) as an example, it contains 16 branches. While directly using a look-ahead algorithm can effectively reduce latency, it introduces a large amount of redundant computation, thus increasing complexity. Existing ARSSD algorithms reduce complexity, but due to the need for sequential computation, their latency is comparable to that of traditional MLSD structures. To address these issues, this embodiment proposes a partitioned look-ahead algorithm that reduces both latency and complexity compared to traditional MLSD. For example... Figure 12 Taking the 2-state PAM4 signal shown in (a) as an example, according to the principle of maximizing the Euclidean distance between elements within a partition, the PAM4 signal with the signal constellation {-3,-1,1,3} can be divided into two partitions, 0 and 1, where partition 0 contains {-3,1} and partition 1 contains {-1,3}. k Input signal at time xk It will be compared with the decision levels (DLev / dLev) dLev0 and dLev1 of the two partitions respectively. Based on the comparison results, only calculation is needed. k The branch metric is calculated based on the number of branches originating from one node within each partition at any given time, resulting in a total of 8 branches. Then, based on... k The partitioning result at time +1 further eliminates the two branches for each partition, ultimately requiring only four branches to be calculated. The partitioning calculation process is as follows: Figure 12 In Figure 12 (a) Figure 12 (b) and Figure 12 As shown in (c). Compared to the traditional 2-state PAM4 MLSD structure, which requires the computation of 16 branches... bm The computational load was reduced by 75%, and the number of registers required to store path nodes was also reduced by 75%. For a length of... After the signal sequence is partitioned, a look-ahead algorithm can be used to perform grid merging on it. Figure 12 (d) Figure 12 (e) and Figure 12 (f) details the mesh merging process of the signals after partitioning. Time zone 0 from departure to entry Taking partition 0 at time 0 as an example, there are two possible paths: via arrive Japanese Classics arrive The combined result can be obtained according to the following formula. arrive The branch.

[0117]

[0118] in, This represents the branch metric from time k to k+2, from level 0 to level 0. This represents the branch metric from level 0 to level 0 at time k+1. This represents the branch metric from level 0 to level 1 at time k+1. This represents the branch metric from level 0 to level 0 at time k+2. This represents the branch metric from level 1 to level 0 at time k+2, with min indicating the minimum value. The remaining three merged branches can be obtained using this method. By analyzing... Continuously performing mesh merging on signal sequences can significantly reduce latency. The PAM4 signal achieves the same latency reduction effect as the NRZ signal by performing a partitioned lookahead algorithm, and the complexity of the ACS operation during merging is also greatly reduced due to the reduced number of branches. The partitioned lookahead algorithm can be expressed as: for the input signal... x Starting from 0 to the sequence length L- 1. Perform traversal: comparison x By finding dLev0, we can obtain i Branch vector of time partition 0 bm i ;Compare x By using dLev1, we can obtain i Branch vector of time partition 1 bm i ;according to i Time partitioning, excluding i Half of the branch vector at time 1. Then, from step 0 to N, adjacent grids are combined, retaining the minimum path metric vector. pm and the corresponding reserved partition number vector id Based on the final merged grid, the minimum value of the path metric vector is determined. pm The corresponding path number is the detection result vector. The traditional sliding block technique works as follows: Figure 13 As shown in (a), for a length of L The sequence is modified by adding a length of to both ends. N The synchronization block and backtracking block. In traditional sliding block technology, both the synchronization block and the sliding block need to perform ACS operations. This is actually to cut off the correlation between various decoding blocks through redundant calculations in order to achieve parallel operation of MLSD. In this embodiment, the partition look-ahead parallel detector uses a partition look-ahead sliding block structure whose working process is as follows: Figure 13 As shown in (b), in this embodiment, the partitioned look-ahead parallel detector combines the partitioned look-ahead algorithm with the sliding block technique to achieve a pipelined output of detection results. This method can simultaneously reduce the complexity, power consumption and latency of PAM4 signal sequence detection technology.

[0119] In summary, this embodiment of the low-latency sequence detection device for CPO photoelectric signal reception includes an adaptive module and multiple partitioned look-ahead parallel detectors (APLSDs). The APLSDs consist of sequentially connected partition and branch metric calculation modules, a grid merging module, and a decision module. They employ a partitioned look-ahead sequence detection method, combining the principles of partitioning and grid merging to reduce sequence detection latency. Then, a hybrid expected value calculation method is proposed to reduce computational load. Finally, a hybrid parameter adaptive algorithm is proposed for adaptive adjustment of the expected value parameters in the hybrid method. An improved add-compare-select operation module (MACS) is used, which reduces computational latency compared to the traditional ACS structure, thus providing an adaptive, low-latency sequence detection scheme for high-speed PAM4 signals. This embodiment completes the circuit design of the low-latency sequence detection device for CPO photoelectric signal reception and verifies it on an FPGA. Test results show that the 8-bit improved add-compare-select operation module (MACS) reduces logic latency by 31.3% compared to ACS. The 2-state partitioned look-ahead parallel detector (APLSD) reduces detection latency by 76.7% compared to MLSD, and also reduces FPGA resources by at least 52% and power consumption by approximately 71%. Adaptive parameters can converge within 2000 cycles.

[0120] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A low-delay sequence detection device for CPO photoelectric signal reception, characterized in that, It includes an adaptive module and multiple partitioned look-ahead parallel detectors. Each partitioned look-ahead parallel detector comprises a partition and branch metric calculation module, a grid merging module, and a decision module connected in sequence. The partition and branch metric calculation module is used to process the input signal... The partitioning and branching measurement of the PAM4 signal are calculated based on the tap coefficients and ISI coefficients provided by the adaptive module. The grid merging module is used to calculate the path metrics of synchronization blocks and backtracking blocks by merging grid signals based on the obtained branch metrics; The decision module is used to process the input signal based on the obtained path metric. The decision is made to obtain a binary PAM4 signal. The adaptive module is used to determine the binary PAM4 signal output by the decision module. The generation partition and branch metric calculation module is used to perform the tap coefficients and ISI coefficients required for signal detection. The partitioning and branching metric calculation module includes: a prediction calculation unit, two partitioning units, and an adder. The prediction calculation unit is used to calculate according to the following formula. 16 expected levels at the time : ; in, and They are respectively Time and The desired level at time t, each desired level contains four combinations from 0 to 3, resulting in a total of 16 possible combinations of desired levels. The ISI coefficient; And the input signal is based on the following formula. Calculate the expected input : ; in, This is the tap factor; The partitioning unit is used to determine the expected input. 8 expected levels and the partition number of the input signal id Select two expected levels for partitioning. Used for branch metrics Calculation; The adder is used to calculate the branch metric according to the following formula. : ; in, for Branching metric at time , for Expected level at time, for The expected input at any given time; The mesh merging module includes a five-layer improved addition-comparison-selection operation module, and the partitioning and branch metric calculation module outputs 32 sets of branch metrics. bm and the corresponding partition number of the input signal The 16 improved add-compare-select operation modules sent to the first level each measure the two sets of branches according to the partition look-ahead algorithm. bm and the corresponding partition number of the input signal Merge into a set of branch metrics bm and the corresponding partition number of the input signal Then it is fed into the next level of 8 improved addition-comparison-selection operation modules, and finally a set of branch metrics is obtained through the fifth level of 1 improved addition-comparison-selection operation module. bm and the corresponding partition number of the input signal .

2. The low-delay sequence detection device for CPO photoelectric signal reception according to claim 1, characterized in that, The input signal of the partition and branch metric calculation module The data consists of 32 channels of 8-bit data. The partitioning and branching metric calculation module calculates 4 sets of branch metrics. The grid merging module calculates 8 path metrics for the synchronization block and backtracking block. The final decision module outputs a binary PAM4 signal. It is a 64-bit signal.

3. The low-delay sequence detection device for CPO photoelectric signal reception according to claim 1, characterized in that, The partition unit includes: Comparator circuit, used to compare the expected input The decision levels dLev0=-1 and dLev1=1 of the partition are compared to determine whether the signal is numbered 0 or 1 in the partition. The PAM4 signal with constellation set {-3,-1,1,3} is divided into two partitions numbered 0 and 1, where partition 0 contains {-3,1} and partition 1 contains {-1,3}. The selection circuit is used to first select from 8 expected levels for each partition through four selectors (MUX) based on the comparison result obtained from the comparison circuit. Four expected levels are selected from the input signal for subsequent calculations, and then the partition numbering is used as the basis for the calculations. id By using two selectors (MUX) to remove two expected levels based on the partitioning results at the next time step, two expected levels are ultimately selected for each partition. Used for branch metrics The calculation.

4. The low-delay sequence detection device for CPO photoelectric signal reception according to claim 1, characterized in that, The improved addition-comparison-selection operation module includes: Two 2-bit carry-lookahead adders (CLAs) are used to implement 2-bit branch addition operations, performing addition operations on the four branches of the input; An improved two-bit carry-lookahead adder MCLA is used to implement a comparison operation with two-bit branches. The comparison operation is performed on the addition results of two two-bit carry-lookahead adders CLA to obtain the comparison result. Two selectors MUX, one of which is used to select the smaller branch metric from the addition results of the two two-bit carry-lookahead adders CLA based on the comparison result of the improved two-bit carry-lookahead adder MCLA. bm The output, another selector MUX, is used to partition the corresponding two input signals in four branches based on the comparison result of the improved two-bit carry-lookahead adder MCLA. Choose a smaller branch metric bm Corresponding partition number Output.

5. The low-delay sequence detection device for CPO photoelectric signal reception according to claim 4, characterized in that, The two-bit carry-lookahead adder (CLA) is a logic gate device.

6. The low-delay sequence detection device for CPO photoelectric signal reception according to claim 4, characterized in that, The logical expression for the improved two-bit carry-lookahead adder MCLA is as follows: ; in, and For the input operands, for The result obtained through an inverter, and the carry output flag of the improved two-bit carry-lookahead adder MCLA. As the selection input for the two selectors MUX.

7. The low-delay sequence detection device for CPO photoelectric signal reception according to claim 1, characterized in that, The judgment module includes: The add-compare-select operation module is used to add the path metric output by the mesh merging module. The path metrics of the synchronization block and the backtracking block are added separately. After performing an add-compare-select operation on the current path metric and the synchronization block and the backtracking block, the final four path metrics are output. The synchronization block and the backtracking block are the merged path metrics cached in the previous and next time steps, respectively. The decision circuit is used to output the final four path metrics from the add-compare-select operation module. Make a decision to find the path with the minimum metric. Corresponding partition number Output; The selector MUX is used to determine the corresponding partition number output by the decision circuit. Output the corresponding partition numbers for the four paths. The corresponding path is output and converted into a binary PAM4 signal.

8. The low-delay sequence detection device for CPO photoelectric signal reception according to claim 1, characterized in that, The adaptive module includes: The lookup table module is used to output the decision results in each cycle. Eight sets of 2-bit results are selected, and the input 2-bit detection results are converted into 8-bit signed numbers using a lookup table. Then follow the formula below: Generate an estimated signal, where For the first One expected level, For the first -1 desired level, The first ISI parameter is then set as a post-processor, and the updated tap coefficients and ISI coefficients are generated by combining them with the expected input signal according to the following formula: ; ; in, The updated tap coefficients, For the input tap coefficient, These are the step parameters used to adjust the convergence speed. For input signal, The updated ISI coefficients, These are the input ISI coefficients.