Digital filtering using combined approximate summation of partial products
By combining the summing circuit and the partial product circuit, using a carry-reserved adder tree and an approximate adder, the problem of inter-symbol interference at high symbol rates is solved, and a low-power and low-area digital filter design is realized, which improves the performance of the communication system.
Patent Information
- Application Number
- CN202111183428.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-09
- Filing Date
- 2021-10-11
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2041-10-11
AI Technical Summary
As the symbol rate increases, existing equalizers find it difficult to effectively reduce noise and ISI under limited power and area requirements when facing increasing intersymbol interference (ISI), resulting in increased communication system design complexity and cost.
The digital filter and filtering method are used, and the summing circuit and partial product circuit are combined to reduce power and area requirements through carry-reserved adder tree and approximate adder, including the use of CSA tree, CPA adder and approximate adder, and combined with cutoff and internal rounding technology, the combination of filter coefficients and signal samples is optimized.
It significantly reduces the power and area requirements of the digital filter, and effectively reduces inter-symbol interference, improves the accuracy of symbol judgment and the efficiency of the communication system.
Smart Images

Figure CN114337602B_ABST
Abstract
Description
Background Art
[0001] Digital communication occurs between a transmitting device and a receiving device through an intermediate communication medium or "channel" (e.g., a fiber optic cable or insulated copper wire). Each transmitting device typically transmits symbols at a fixed symbol rate, while each receiving device detects (possibly corrupted) symbol sequences and attempts to reconstruct the transmitted data. A "symbol" is the state or valid condition of a channel that lasts for a fixed period of time, which is called a "symbol interval." For example, a symbol can be a voltage or current level, an optical power level, a phase value, or a specific frequency or wavelength. The change from one channel state to another is called a symbol transition. Each symbol can represent (i.e., encode) one or more binary bits of data. Alternatively, data can be represented by symbol transitions, or by a sequence of two or more symbols.
[0002] Many digital communication links use only one bit per symbol; a binary "0" is represented by one symbol (e.g., a voltage or current signal within a first range), and a binary "1" is represented by another symbol (e.g., a voltage or current signal within a second range), but higher-order signal constellations are known and frequently used. In 4-level pulse amplitude modulation ("PAM4"), each symbol interval can carry any of four symbols labeled -3, -1, +1, and +3. Two binary bits can thus be represented by each symbol.
[0003] Channel imperfections create dispersion that can cause each symbol to interfere with its neighbors, a result known as inter-symbol interference (ISI). ISI can make it difficult for a receiving device to determine which symbols were sent in each interval, especially when such ISI is combined with added noise.
[0004] To combat noise and ISI, both transmitting and receiving devices can employ various equalization techniques. Linear equalizers typically must strike a balance between reducing ISI and avoiding noise amplification. Decision feedback equalizers ("DFEs") are often preferred because they can combat ISI without inherently amplifying noise. As the name implies, DFEs employ a feedback path to remove the effects of ISI from previously decided symbols. Other suitable equalizer designs are also known.
[0005] As symbol rates continue to increase, any equalizer used must contend with increasing levels of ISI and must process them within ever-decreasing symbol intervals. These twin demands for increased complexity and faster processing pose a greater challenge to communication system designers, especially when combined with the cost and power constraints required for commercial success. Summary of the Invention
[0006] Thus, disclosed herein are digital filters and filtering methods that can significantly reduce power and area requirements with minimal performance degradation. An illustrative digital filter includes: a summing circuit coupled to a plurality of partial product circuits. Each partial product circuit is configured to combine bits of a filter coefficient with bits of a corresponding signal sample to produce a set of partial products. The summing circuit produces a filter output using a carry-save adder ("CSA") tree that combines the partial products from the plurality of partial product circuits into bits for two addends. The CSA tree has multiple channels of addends, each channel being associated with a corresponding bit weight. The adders in one or more of the channels associated with the least significant bits of the filter output are approximate adders that sacrifice precision for a simpler implementation.
[0007] An illustrative receiver includes a digital filter as above coupled to a decision element that derives a sequence of symbol decisions based at least in part on the filter output.
[0008] An illustrative filtering method includes, for each of a plurality of filter coefficients, combining bits of the filter coefficient with bits of corresponding signal samples to produce a set of partial products. The method further includes using a CSA tree to produce a filter output, a carry-save adder tree to combine bits from the plurality of sets of partial products into two addends, the CSA tree having a plurality of channels of addends, each channel associated with a corresponding bit weight. The adders in one or more of the channels associated with the least significant bits of the filter output are approximate adders.
[0009] Each of the aforementioned equalizers and equalization methods may be implemented individually or in combination, and may be implemented in conjunction with any one or more of the following features in any suitable combination: 1. The summing circuit further includes a carry-propagation adder ("CPA") that sums two addends to form a filter output. 2. The CPA includes an adder for each overlapping bit weight of the two adders. 3. The CPA adders associated with the one or more channels are approximate adders. 4. When generating the filter output, the summing circuit truncates a predetermined number of least significant bits from the sum of the two addends. 5. The plurality of partial product circuits omit any partial products associated with bit weights below a predetermined rounding level, the predetermined rounding level being greater than the bit weight associated with the product of the least significant bit of the filter coefficient and the least significant bit of the corresponding signal sample. 6. The predetermined rounding level corresponds to the number R of internal rounding bits. 7. The adders in at least five channels associated with the most significant bits of the filter output are exact adders. 8. The signal samples are spatially separated. 9. The signal samples are separated in time. 10. The filter coefficients include feed-forward equalization ("FFE") filter coefficients and at least one feedback filter coefficient. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 An illustrative computer network is shown.
[0011] Figure 2 is a block diagram of an illustrative point-to-point communication link.
[0012] Figure 3 is a block diagram of an illustrative serializer-deserializer transceiver.
[0013] Figure 4 is a block diagram of an illustrative decision feedback equalizer ("DFE").
[0014] Figure 5 is a block diagram of an illustrative digital filter.
[0015] Figure 6A is a block diagram of an illustrative parallel DFE.
[0016] Figure 6B This is a timing diagram of the clock signals used to parallelize the DFE.
[0017] Figure 7A Illustrative partial product arrangement involving filter outputs.
[0018] Figure 7B An illustrative partial product arrangement with truncation and internal rounding is shown.
[0019] Figure 7C An illustrative summing circuit using approximation is shown.
[0020] Figures 8A-8H An illustrative logical equivalent of the summing circuit components is shown.
[0021] Figure 9 is a flow chart of an illustrative digital filter design method.
[0022] Figure 10 is an illustrative partial product circuit. DETAILED DESCRIPTION
[0023] Note that the specific embodiments set forth in the drawings and the following description do not limit the present disclosure. Rather, they provide a basis for one of ordinary skill in the art to identify alternatives, equivalents, and modifications that are contemplated by the inventor and encompassed within the scope of the claims.
[0024] The disclosed digital filters and filtering methods are best understood in the context of the larger environment in which they operate. Accordingly, Figure 1 An illustrative communication network is shown that includes a mobile device 102 and computer systems 103-104 coupled via a packet-switched routing network 106. Routing network 106 may be or include, for example, the Internet, a wide area network, or a local area network. Figure 1 , routing network 106 includes a network of equipment items 108, such as switches, routers, bridges, etc. Equipment items 108 are connected to each other and to computer systems 103-104 via point-to-point communication links 110 that transmit data between the various network components.
[0025] Figure 2 It can be expressed Figure 1 FIG2 is a diagram of an illustrative point-to-point communication link of link 110 in FIG2 . The illustrated embodiment includes a first node 202 (“Node A”) communicating with a second node 204 (“Node B”). Node A and Node B can each be, for example, any of mobile device 102, item of equipment 108, computer systems 103-104, or other transmitting / receiving devices suitable for high-rate digital data communication.
[0026] Node A includes a transceiver 202 coupled to node A's internal data bus 204 via a host interface 206. Similarly, node B includes a transceiver 203 coupled to its internal bus 205 via a host interface 207. Data transmission within the node can occur via, for example, a parallel 64- or 128-bit bus. When a block of data needs to be transmitted to or from a remote destination, the host interfaces 206, 207 can convert between the internal data format and the network packet format and can provide packet sorting and buffering for the transceiver. The transceivers 202, 203 serialize the data packets for transmission via a high-bandwidth communication channel 208 and process the received signals to extract the transmitted data.
[0027] A communication channel 208 extends between transceivers 202 and 203. Channel 208 may include, for example, a transmission medium such as a fiber optic cable, a twisted pair, a coaxial cable, a backplane transmission line, and a wireless communication link. (The channel may also be formed by a magnetic or optical information storage medium having a read-write transducer that acts as a transmitter and a receiver.) Bidirectional communication between node A and node B can be provided using independent unidirectional channels, or in some cases, a single channel that transmits signals in opposite directions without interference. The channel signal can be, for example, a voltage, current, optical power level, polarization, angular momentum, wavelength, frequency, phase value, or any suitable energy property that is transferred from the beginning of the channel to the end of the channel. The transceiver includes a receiver that processes the received channel signal to reconstruct the transmitted data.
[0028] Figure 3 An illustrative single-chip transceiver chip 302 is shown. Chip 302 includes a SerDes module with contacts 320 for receiving and transmitting a high-rate serial bit stream across eight lanes of communication channel 208; an additional SerDes module with contacts 322 for transmitting the high-rate serial bit stream to host interface 206; and core logic 324 for implementing the channel communication protocol while buffering the bit stream between the channel and the host interface. Also included are various support modules and contacts 326, 328, such as those for power regulation and distribution of control signals, clock generation, digital input / output lines, and a JTAG module for built-in self-test.
[0029] The "deserializer" implements the receive functionality of chip 302 and implements decision feedback equalization ("DFE") or any other suitable equalization technique, such as linear equalization or partial response equalization, which may employ a digital filter with adjustable tap coefficients. The selected equalizer operates under strict timing constraints at the envisioned symbol rates (above 50 Gbd).
[0030] Figure 4An illustrative embodiment of a DFE configured to receive a PAM4 signal is shown. An optional continuous time linear equalization ("CTLE") filter 402 provides analog filtering to band limit the signal spectrum while optionally boosting the high frequency components of the received signal RX_IN. An analog to digital converter 404 samples and digitizes the received signal. A feed forward equalization ("FFE") filter 406 operates on the digitized signal to minimize leading inter-symbol interference ("ISI") while optionally reducing the length of the channel impulse response. A summer 408 subtracts the feedback signal provided by the feedback filter 410 from the filtered signal provided by the FFE filter 406 to produce an equalized signal in which the effects of trailing ISI have been minimized. A decision element 412 (sometimes called a "slicer") operates on the equalized signal to determine which symbol it represents in each symbol interval. The resulting stream of symbol decisions is represented as A k , where k is the time index.
[0031] In the example shown, the symbols are assumed to be PAM4 (-3, -1, +1, +3), so the comparator employed by the decision element 412 uses decision thresholds of -2, 0, and +2, respectively. (For generality, the units used to express the symbols and thresholds are omitted, but volts can be assumed for the purpose of explanation.) The comparator outputs can be collectively viewed as thermometer-coded digital representations of the output symbol decisions, or a digitizer can optionally be used to convert the comparator outputs into a binary representation, for example, 00 represents -3, 01 represents -1, 10 represents +1, and 11 represents +3. Alternatively, a Gray-coded representation can be employed.
[0032] DFE utilizes a memory with recent output symbol decisions (A k-1 、A k-2 ,…,A k-N , where N is the filter coefficient F -iA feedback filter 410, comprising a series of delay elements D (e.g., latches, flip-flops, or shift registers) with a plurality of delay elements D (e.g., a plurality of delay elements D) (e.g., a plurality of latches, flip-flops, or shift registers) generates a feedback signal. A set of multipliers determines the product of each symbol using the corresponding filter coefficients, and a summing circuit 414 combines these products to obtain the feedback signal. In situations where a short symbol interval makes it impossible to implement the feedback filter 410, the decision element 412 can be modified as a pre-compensation unit to "unfold" one or more taps of the feedback filter, thereby potentially eliminating the feedback filter altogether. The pre-compensation unit provides speculative decisions to a multiplexing arrangement such as described in U.S. Patent No. 8,301,036 ("High-speed adaptive decision feedback equalizer") and U.S. Patent No. 9,071,479 ("High-speed parallel decision feedback equalizer"), each of which is incorporated herein by reference in its entirety.
[0033] Additionally, we note that actual DFE implementations typically include a timing recovery module and a filter coefficient adaptation module, but these considerations are discussed in the literature and are well known to those skilled in the art. However, we note that at least some contemplated embodiments include one or more additional comparators for comparing the combined signal to one or more of the extreme symbol values (-3, +3), thereby providing an error signal (or error polarity signal) that can be used for timing recovery and filter adaptation.
[0034] In the illustrated embodiment, the FFE filter 406 is a digital filter with adjustable tap coefficients. Figure 5 An illustrative finite impulse response ("FIR") implementation of a digital filter is shown. An input signal is provided to a sequence of delay elements. The first of the delay elements captures the input signal value once in each symbol interval while outputting the captured value from the preceding symbol interval. Each of the other delay elements captures the held value from the preceding element, and this operation is repeated to provide increasingly delayed signal samples. A set of multipliers multiply the input signal by the corresponding coefficients F i Each input value in the sequence is scaled and the scaled values are provided in the form of partial products (described further below) to summing circuit 416, which sums the scaled values to form the filter output.
[0035] The digital multiplier of FFE 406 (and the digital multiplier of feedback filter 410) can be configured to use the partial product principle according to the Baugh Wooley algorithm outlined in, for example, Keote and Karule, “Performance analysis of fixed width multiplier using Baugh Wooley Algorithm,” IOSR J. of VLSI and Signal Proc., Vol. 8, No. 3, p. 31, (2018). That is, once the filter coefficients F m Sum signal sample D k-m In binary form, their product can be expressed as the sum of the products of their bits:
[0036]
[0037] The task is then to sum the bit products while taking into account their bit weights. (For simplicity, equation (1) assumes unsigned binary numbers. Adjustments for signed numbers are discussed in the open literature. See, for example, Keote and Karule.) This approach allows the use of a carry-save adder ("CSA") design that defers the full propagation of any carry bit until the last summation. When configured in this way, the multipliers can supply the bit products to a unitary summing circuit 416, which combines the bit products to form the filter output.
[0038] This method may be applied to feedback filter 410 such that summing circuit 414 combines the bit products supplied from the multipliers in feedback filter 410. This method may optionally be further extended such that summing circuits 408, 414, and 416 are all combined into a single unitary summing circuit that implements only one complete propagation of the carry bit, delaying the one complete propagation of the carry bit until the end of the summation, at which end the equalized signal is produced.
[0039] Figure 4-Figure 5 The design of an equalizer may require performing a relatively large number of operations in each symbol interval, which becomes increasingly challenging as the symbol interval becomes smaller. Figure 6A Accordingly, a parallelized version of the equalizer is provided, which has a parallelized FFE filter, a parallelized decision element and a parallelized feedback filter to relax the timing requirements.
[0040] exist Figure 6AIn FIG, CTLE filter 402 bandwidth limits the received signal before supplying it in parallel to an array of analog-to-digital converter (ADC) elements. Each of the ADC elements has its own clock signal, each of which has a different phase, so that the elements in the array take turns sampling and digitizing the input signal. At any given time, only one of the ADC element outputs is converting. Figure 6B Illustration of how clock signals are phase-shifted relative to each other. Note that the duty cycles shown are merely illustrative; the point being conveyed is the sequential nature of the transitions in the different clock signals.
[0041] An array of FFE filters (FFE0 to FFE7) each forming a weighted sum of the ADC element outputs. The weighted sum employs filter coefficients that are cyclically shifted relative to one another. FFE0 operates on the hold signals from the three ADC elements operating before CLK0, the ADC element responsive to CLK0, and the three ADC elements operating after CLK0, such that during the assertion of CLK4, the weighted sum produced by FFE0 corresponds to the output of FFE filter 406 ( Figure 4 and Figure 5 ). FFE1 operates on the held signals from the three ADC elements operating before CLK1, the ADC elements responsive to CLK1, and the three ADC elements operating after CLK1, so that during the assertion of CLK5, the weighted sum corresponds to the output of FFE filter 406. The operation of the remaining FFE filters in the array follows the same pattern with associated phase shifts. In practice, the number of filter taps can be smaller, or the number of elements in the array can be larger, to provide a longer valid output window.
[0042] and Figure 4 As with a receiver, a summing circuit can combine the output of each FFE filter with a feedback signal to provide an equalized signal to the corresponding decision element. Figure 6A An array of decision elements (Limiter 0 to Limiter 7) is shown, each operating on an equalized signal derived from a corresponding FFE filter output. Figure 4 Like the decision element 412 shown, the decision element shown uses a comparator to determine which symbol the equalized signal most likely represents. The decision is made when the corresponding FFE filter output is valid (e.g., limiter 0 operates when CLK4 is asserted, limiter 1 operates when CLK5 is asserted, etc.). The symbol decisions are provided in parallel on the output bus to enable lower clock rates for subsequent on-chip operations.
[0043] The feedback filter array (FBF0 to FBF7) operates on the leading symbol decisions to provide feedback signals for the summer. As with the FFE filter, the input to the feedback filter is cyclically shifted and is only used when the input corresponds to the feedback filter 410 ( Figure 4 ) content, consistent with the time window of the corresponding FFE filter. In practice, the number of feedback filter taps can be less than that shown, or the number of array elements can be larger to provide a longer valid output window.
[0044] and Figure 4 The decision element 412 is the same as Figure 6B The decision elements in each of the FFE filters can each employ an additional comparator to provide timing recovery information, coefficient training information, and / or pre-compensation to expand one or more taps of the feedback filter. After accounting for cyclic shifts, the same tap coefficients can be used for each FFE filter and each feedback filter. Alternatively, the error module can calculate the residual ISI of each filter separately, enabling the adaptation module to independently adapt to the tap coefficients of different filters. The summing circuit can be configured to receive and combine bit products from its associated multipliers as described above.
[0045] Figure 7A is a graph showing the bit products P of equation (1) arranged in the lanes corresponding to their associated bit weights. m,i,j =(f m,i d k-m,j ). For illustrative purposes, this example assumes eight-bit coefficients and eight-bit signal sample values in a three-tap FFE filter. The bit resolution can be higher or lower, and the number of filter taps can vary, but will typically be greater than three. The first partial product circuit multiplies the bits of the filter coefficient F0 with the signal sample D k-0 The bits of coefficient F1 are combined to produce a first set of bit products 702. The second partial product circuit combines the bits of coefficient F1 with the signal sample D k-1 The bits of coefficient F2 are combined to produce a second set of bit products 704. The third partial product circuit combines the bits of coefficient F2 with the signal sample D k-2 The bits of are combined to produce a third set of bit products 706. Additional filter coefficient multiplications will be implemented by additional partial product circuits.
[0046] Arrow 708 represents the operation of the CSA tree structure, which reduces the bit product in each lane to the bits of two addends 710. (In the figure, symbol S j represents the least significant bit of the sum of channel j, C j represents the carry bit of the sum of lane j. ) Various reduction techniques developed for standalone binary multipliers (e.g., "carry-save array" reduction, Wallace tree reduction, and Dadda tree reduction) are also suitable for Figure 7A More details on such techniques can be found in the literature. For example, see C.S.W. Allace, "A suggestion for a fast multiplier," IEEE Transactions on Electronic Computers, Vol. 13, No. 1, pp. 14-17, February 1964. L. Dadda, "Some schemes for parallel multipliers," Alta Frequenza, Vol. 34, No. 5, pp. 349-356, May 1965.
[0047] Arrow 712 represents the operation of any suitable adder structure that sums the addends 710 to obtain a sum having bits R. j The binary representation of the result 714. Suitable options include carry propagate adder (CPA) and carry lookahead adder structures. The result may correspond to the filter output.
[0048] Although applying the CSA reduction technique to the combinatorial set of partial products already provides significant efficiency gains, further gains can be achieved through judicious use of truncation, internal rounding, and approximations, which will be discussed in detail in the next section. Figure 7B Have a discussion.
[0049] Truncation can be applied to the filter coefficients and / or the filter output. It is rarely necessary to make each of the filter coefficients scalable over the full eight bits of resolution; for most channels and systems, the magnitude of most of the filter coefficients is much less than 1. This observation can be used to reduce the number of bits used to represent those coefficients. Therefore, in Figure 7B , the three most significant bits of the first coefficient are truncated, thereby eliminating the bit products in row 716. The two most significant bits of the second coefficient are truncated, thereby eliminating row 718. The four most significant bits of the third coefficient are truncated, thereby eliminating row 720. In this way, the number of bit products to be combined can be significantly reduced.
[0050] The bit resolution of the filter output can also be truncated. For example, if the T result bits 721 are not included in the filter output, they do not need to be calculated. However, the associated carry bits may affect the result, so caution is required. Figure 7B In the case where five output bits are selected for truncation, internal rounding is only used to eliminate the calculation of the four least significant bits (channel 722).
[0051] Internal rounding ignores the bit products associated with a specified number R of bit weight channels. In theory, a fixed offset can be added to the result to minimize rounding errors, but in practice, the offset can be naturally accounted for elsewhere in the receive chain (for example, in gain and offset adaptation for analog-to-digital conversion). The number of internal rounding bits, R, is related to the number of output truncation bits, T, where T is expected to be greater than or equal to R; otherwise, the (RT) least significant bit of the output will always be zero. Conversely, if T is much greater than R, a lot of effort is invested in calculations that are increasingly unlikely to affect the result. Therefore, it can be conceived that (T-3)≤R≤T. In this way, the number of bit products to be combined can be further significantly reduced.
[0052] To further reduce power consumption and area requirements, the CSA tree structure used to combine bit products can be modified to replace exact adders with approximate adders, as described further below. The error introduced by the approximate adder masquerades as statistical noise, whose power is determined by the bit weights of the channels 724 in which the replacement is made. To limit the added noise power to be equal to or lower than the power of other expected noise sources in the system, the approximation can be limited to a specified number A of least significant bit weight channels 724. With this replacement, the gate count is significantly reduced.
[0053] Arrow 726 accordingly represents the operation of a modified CSA tree structure that employs truncation, internal rounding, and / or approximation to reduce the power and area requirements of the CSA process for summing the bitwise products to form the bits of the two addends 710. The selected adder structure may also replace the exact adders in lanes 724 with approximate adders in the addend sum operation represented by arrow 728 to achieve further efficiency.
[0054] Considering the foregoing context, Figure 7C An illustrative CSA tree structure (stages 732-738) coupled to a CPA structure (stage 739) is shown. Figure 7C The closed point on the left side represents the bit product P organized by bit weight m,i,j (Since the bit products within a given lane each have the same bit product, they can be reordered as desired).
[0055] The adder receives the bit products from multiple partial product circuits. Figure 10 , which shows the method for generating the set of partial products 706 ( Figure 7B) partial product circuit. Each coefficient bit is combined with each signal sample bit using a logical AND, except where the bit product would fall into a lane that is ignored due to internal rounding or where the coefficient bits have been truncated. The products are arranged in lanes according to bit weights and fed to an adder structure. Other partial product circuit implementations are known and would be suitable.
[0056] Return to Figure 7C , the adder structure includes here will refer to Figures 8A-8H The illustrated structure was developed using the Wallace tree reduction technique, but the principles disclosed are equally applicable to other CSA arrays and tree structures.
[0057] In the channel 724 employing the approximation, the first stage 732 provides an AND-OR element 801 that combines paired bit products using logical AND and logical OR, as shown in FIG. Figure 8A This operation changes the statistics of the bit products in a way that is beneficial to the approximation. Specifically, the bit product P m,i,j The expected probability of being "1" is about 25%, and the probability of being "0" is about 75%. Subsequently, the probability of the logical AND of the two bit products being "1" is about 6% (1 / 16), and the probability of being "0" is about 94%. The probability of the logical OR of the two bit products being "1" is about 44% (7 / 16), and the probability of being "0" is about 56%. Note that the number of "1"s is conserved between the input and output of the AND-OR element.
[0058] In the second stage 734, the A outputs of the AND-OR elements are combined by the group element 802, as shown in FIG. Figure 8B As shown in Figure 1. A group element logically ORs together up to four inputs. If each input has only a 6% probability of being asserted, the probability that the group element's output will correctly represent the sum of the inputs is close to 98%.
[0059] In discussion Figure 7C Before discussing the rest of Figure 8C-8H . Figure 8C An exact half adder 803 is shown which combines the two inputs according to the rules of binary addition to provide a sum bit S and a carry bit C. Figure 8C An equivalent logic circuit for a half adder is shown, but other implementations are also applicable.
[0060] Figure 8D An exact full adder 804 is shown that combines the three inputs according to the rules of binary addition to provide a sum bit S and a carry bit C. Figure 8D An equivalent logic circuit for an exact full adder is also shown, but other implementations are also applicable.
[0061] Figure 8E An exact compressor 805 is shown that combines the five inputs to provide a sum bit S and two carry bits C1, C2. In addition to different propagation delays, the carry bits are interchangeable and each carry bit is associated with a bit weight that is twice the bit weight of the sum bit S. Figure 8E The equivalent logic circuit for a compressor is shown, but other implementations are also applicable.
[0062] Figure 8F An approximate half adder 806 is shown, which combines two inputs to produce a sum bit S and a carry bit C. Assuming that the probability of the independent inputs being asserted is 50%, the probability of the output being accurate is 75%, and the probability of the error being 1 is 25%. The approximate half adder 806 is relatively accurate compared to the exact half adder 803 ( Figure 8C ) with a reduced gate count.
[0063] Figure 8G The approximate full adder 807 is shown, which combines the three inputs to produce the sum bit S and the carry bit C. Assuming that the probability of the independent input being asserted is 50%, the probability of the output being accurate is 75%, and the probability of the error being 1 is 25%. The approximate full adder 807 is relatively accurate compared to the exact full adder 804 ( Figure 8D ) with a reduced gate count.
[0064] Figure 8H An approximate compressor 808 is shown that combines the four inputs to produce a sum bit S and a carry bit C. Approximate compressor 808 is conceptually a replacement for exact compressor 805, but it ignores one of the inputs and does not compute the second carry bit. Assuming a 50% probability that the four independent inputs are asserted (and assuming the second carry bit remains zero), the probability that the output is accurate is 69% and the probability that the error is 1 is 31%.
[0065] In the second stage 734, the O outputs of three or four AND-OR elements are combined at once within each channel by an approximate full adder 807 or an approximate compressor 808. Each of the approximate full adders 807 or approximate compressors 808 generates a sum bit S and a carry bit C. In the non-approximate channel 725, stage 734 uses an exact full adder 804 to sum the three bit products within each channel at once, thereby generating the corresponding sum bit and carry bit. The carry bit has a higher associated bit weight and is accordingly routed to the next channel in preparation for the third stage. The second stage 734 reduces the number of bits from 90 to 43.
[0066] In the third stage 736, three or four bits of the bits in each approximate channel 724 are combined at once by an approximate full adder 807 or an approximate compressor 808. Each of the approximate full adders 807 or approximate compressors 808 produces a sum bit S and a carry bit C. In each of the non-approximate channels 725, stage 736 combines three or five bits at once using an exact full adder 804 or an exact compressor 805. Each full adder produces a sum bit and a carry bit. Each compressor produces a sum bit and two carry bits C1 and C2. The carry bit has a higher associated bit weight and is accordingly routed to the next channel in preparation for the fourth stage. The third stage 736 reduces the number of bits from 43 to 28.
[0067] In the fourth stage 738, the bits in each approximate lane 724 are combined three or four bits at a time by an approximate full adder 807 or an approximate compressor 808, each of which produces a sum bit S and a carry bit C. In each of the non-approximate lanes 725, stage 738 combines three bits at a time using an exact full adder 804, each of which produces a sum bit and a carry bit. The carry bit has a higher associated bit weight and is accordingly routed to the next lane in preparation for the CPA stage 739. The fourth stage 738 reduces the number of bits from 28 to 17. With a bit weight of 2 4 and 2 5 The associated channels each have a single summed bit S, and the fourth stage produces a summed bit with bit weight 2 13 A single carry bit. For 2 5 and 2 13 For each of the bit weights between , the fourth stage produces a single sum bit S and a single carry bit C in each channel. The output of the fourth stage can form the bits of the two addends accordingly.
[0068] Stage 739 is shown in a CPA configuration and operates to sum the addend bits. 4 and 2 5 No further action is required in the corresponding channels, so these bits are passed through without modification. 6 In the associated channel, the approximate half adder 806 sums the carry bit from the previous channel with the sum bit, thereby generating a channel with weight 2. 6 The output bit and the carry bit for the next channel. 7 In the associated channel, the approximate full adder 807 sums the carry bit from the previous channel of the CPA stage with the sum bit and carry bit from the fourth stage, thereby producing a channel with weight 2 7 The output bit and the carry bit for the next channel. 8In the associated channel, the approximate full adder 807 sums the carry bit from the previous channel of the CPA stage with the sum bit and carry bit from the fourth stage, thereby producing a channel with weight 2 8 The output bit of the channel and the carry bit for the next channel.
[0069] In with bit weight 2 9 to 2 12 In the associated lane, the precise full adder 804 sums the carry bit from the previous lane of the CPA stage with the sum bit and carry bit from the fourth stage to produce the output bit for that lane and the carry bit for the next lane. 13 , the precise half adder 803 sums the carry bit from the previous channel of the CPA stage with the carry bit from the fourth stage, thereby producing a half adder with weight 2 13 The sum of the bits has weight 2 14 The carry bit of .
[0070] although Figure 7C The description assumes unsigned binary numbers, but the principles disclosed are readily applicable to the multiplication of signed (two's complement) binary numbers using the Baugh-Wooley technique as described in R. Baugh and B.A. Wooley, "A two's complement parallel array multiplication algorithm," IEEE Trans. Comput., Vol. 22, pp. 1045-1047, December 1973, or as improved upon in subsequent literature. This method complements the partial products including the sign bit and adds a fixed offset to each filter coefficient-signal sample product. Of course, the fixed offsets can be combined to Figure 7A-7B An additional row of partial products is arranged to form a single offset and is explained during the CSA adder structure design.
[0071] Figure 9Flowchart of a method for designing an equalizer that uses truncation, internal rounding, and / or approximation to reduce filter complexity, power, and area. In block 902, a designer obtains a model of a communication channel and system that includes a transmitter and a receiver and may include expected characteristics of noise, jitter, and other performance degradation sources. In block 904, the designer develops a design for receiver-side equalization that includes at least one digital filter and may further develop a digital filter for transmit-side equalization ("pre-equalization"). The equalizer design(s) may include the equalizer type (e.g., linear equalizer, decision feedback equalizer), the number of taps per filter, a coefficient adaptation module, and a timing recovery module.
[0072] In block 906, the designer typically determines the performance of the system through simulation without truncation, internal rounding, approximation, or other constraints on the filter coefficients or filter implementation. Performance can be evaluated over a range of expected channel variations to determine a useful range of values for each filter coefficient. If system performance is insufficient, the designer will modify the system design by, for example, increasing the number of filter coefficients, increasing the bit resolution of the signal samples, adjusting the size of the signal constellation, and / or reducing the symbol rate until satisfactory performance is achieved.
[0073] In block 908, the designer uses the useful range of values identified for each coefficient to set the number of bits required for each coefficient. When combined with the bit resolution selected for the signal samples, the coefficient resolution will determine the possible range of the filter output, which is useful for determining the number of bits desired in the filter output. Using this information, the designer can select the number of bits to truncate from the filter output in block 910. For example, if the filter output nominally includes 15 bits of resolution, the designer may recognize that no more than 10 or 11 bits are needed to adequately perform the adaptation and timing recovery modules, and therefore may decide to truncate, for example, the five least significant bits from the filter output.
[0074] In block 912, the designer may set the number of filter output bits that can be safely derived using the approximate adder, and therefore the number of bits for which the exact adder should be used. The designer may initially choose to use approximation for approximately half of the filter output bits (e.g., the five least significant bits). (In Figure 7B 、 7C In the example of , an approximate adder is used for the four least significant output bits and one of the truncated bits).
[0075] In block 914, the designer can set the number of bits used for internal rounding. As previously discussed, the number of internal rounding bits, R, is related to the number of output truncation bits, T, where T is expected to be greater than or equal to R; otherwise, the (RT) least significant bit of the output will always be zero. Conversely, if T is much greater than R, a lot of effort is expended on computations that are increasingly unlikely to affect the result. Thus, it is contemplated that (T-3)≤R≤T, although there may of course be cases where it is desirable to choose R outside of this range. Figure 7B 、 Figure 7C In the example of , R is chosen to be one less than the number of truncation bits, i.e., four.
[0076] In block 916, the designer provides partial product circuits to generate the appropriate bit products through the associated bit weight arrangement (or "stack"). In block 918, the designer provides a layer of AND-OR elements to provide pairwise combinations of bit products in the channels where the approximate adders are to be used.
[0077] In block 920, the designer constructs a stack reduction tree using an array structure, a Wallace tree structure, a Dadda tree structure, or any suitable variant. (The Dadda method generally produces an adder structure with the fewest gates.) The stack reduction circuit can employ approximate adders in those channels that fall below the selected limit for using approximate adders. Typically, using such approximate adders, the number of gates in these channels can be reduced by nearly half. The final stage of the reduction structure uses CPA or other suitable structure to sum the addend bits from the previous stage, and the structure may also include approximate adders in channels where approximation is desired.
[0078] In block 922, the designer evaluates system performance over a range of expected channel variations, verifying that the use of truncation, internal rounding, and / or approximations has provided the desired efficiency gains without causing performance to fall below specification. It is expected that other noise sources and performance degradation will typically mask any degradation caused by a more efficient filter design, but if this is not the case, the designer can repeat blocks 910-922 as needed to adjust the design factors and determine the most efficient filter design that meets the performance requirements.
[0079] Once a desired design is found, layout software can be used to design and optimize process masks for fabricating an integrated circuit including the filter in block 924. In block 926, the process masks are used to fabricate the integrated circuit, which is then packaged and integrated into a cable, interface, or product for retail sale.
[0080] Although the foregoing discussion has focused on equalizers for sequences of sampled signals, the disclosed principles are also applicable to filters for spatially sampled signals used in beamformers, interferometers, and other multiple-input, multiple-output (MIMO) systems. Image and video filters, neural networks, and in fact any system, device, or method that performs matrix or vector multiplication, convolution, or weighted summation using an application-specific integrated circuit can use truncation, internal rounding, and / or approximations in the structure to more efficiently derive addends and the final sum from multiple sets of bitwise weight stacked bitwise products.
[0081] Numerous alternatives, equivalents, and modifications will become apparent to those skilled in the art once the above disclosure is fully appreciated. Where applicable, the following claims are to be interpreted as including all such alternatives, equivalents, and modifications.
Claims
1. A receiver, comprising: Digital filters, including: a plurality of partial product circuits, each partial product circuit configured to combine bits of the filter coefficients with bits of corresponding signal samples to produce a set of partial products; and a summing circuit coupled to the plurality of partial product circuits to produce a filter output using a carry-save adder tree that combines partial products from the plurality of partial product circuits into bits for two addends, the carry-save adder tree having a plurality of channels of adders, each channel being associated with a corresponding bit weight, the adders in one or more of the channels associated with least significant bits of the filter output being approximate adders; and A decision element is provided for deriving a symbol decision sequence based at least in part on the filter output.
2. The receiver according to claim 1, wherein The filter coefficients include feedforward equalization filter coefficients and at least one feedback filter coefficient.
3. The receiver according to claim 1, wherein The decision element comprises a limiter for four-level pulse amplitude modulation.
4. The receiver according to claim 1, wherein The decision element comprises a multiplexing arrangement coupled to a precompensation unit.
5. The receiver according to claim 1, wherein The summing circuit further includes a carry-propagate adder that sums the two addends to form the filter output, the carry-propagate adder having an adder for each overlapping bit weight of the two addends, wherein the carry-propagate adder associated with the one or more lanes associated with the least significant bit is an approximate adder.
6. The receiver according to claim 1, wherein The plurality of partial product circuits omit any partial products associated with bit weights below a predetermined rounding level that is greater than a bit weight associated with a partial product of a least significant bit of a filter coefficient and a least significant bit of a corresponding signal sample.
7. The receiver according to claim 6, wherein The predetermined rounding level corresponds to a number of internal rounding bits R, and wherein the summing circuit truncates at least one least significant bit from a sum of the two addends when generating the filter output.
8. The receiver according to claim 1, wherein The adders in at least five channels associated with most significant bits of the filter output are exact adders.
9. The receiver according to claim 1, wherein The signal samples are separated in space.
10. The receiver according to claim 1, wherein The signal samples are separated in time.
11. A receiving method, comprising: for each of a plurality of filter coefficients, combining bits of the filter coefficient with bits of corresponding signal samples to produce a set of partial products; generating a filter output using a carry-save adder tree that combines partial products from a plurality of sets into bits for two addends, the carry-save adder tree having a plurality of lanes of adders, each lane associated with a corresponding bit weight, the adders in one or more of the lanes associated with the least significant bits of the filter output being approximate adders; and A sequence of symbol decisions is derived based at least in part on the filter outputs.
12. The receiving method according to claim 11, wherein: The deriving includes comparing the filter output to one or more decision thresholds.
13. The receiving method according to claim 12, wherein: The one or more decision thresholds include pre-compensation.
14. The receiving method according to claim 11, wherein: The deriving includes combining the filter output with a feedback signal and comparing the combined signal to one or more decision thresholds.
15. The receiving method according to claim 11, wherein: The generating further employs a carry-propagate adder that sums the two addends to form the filter output, the carry-propagate adder having an adder for each overlapping bit weight of the two addends, wherein the carry-propagate adder associated with the one or more channels is an approximate adder, the one or more channels being associated with a least significant bit.
16. The receiving method according to claim 11, wherein: Each set of partial products omits any partial products associated with bit weights below a predetermined rounding level that is greater than the bit weight associated with the partial product of the least significant bit of the filter coefficient and the least significant bit of the corresponding signal sample.
17. The receiving method according to claim 16, wherein: The predetermined rounding level corresponds to a number R of internal rounding bits, and wherein the generating comprises truncating at least one least significant bit from the sum of the two addends.
18. The receiving method according to claim 11, wherein: The adders in at least five channels associated with most significant bits of the filter output are exact adders.
Citation Information
Patent Citations
High-speed adaptive decision feedback equalizer
US8301036B2
High-speed parallel decision feedback equalizer
US9071479B2
Low Power array multiplier
US20060212505A1