Horner form arbitrary coefficient multiplierless finite impulse response filter
Folded FIR filters with multiplexers and partial product stages address power and heat issues in quantum computing, improving qubit fidelity and operation in refrigerated environments.
Patent Information
- Application Number
- JP2025060285
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-02
- Filing Date
- 2025-04-01
- Publication Date
- 2025-11-12
AI Technical Summary
Quantum computing systems are susceptible to environmental noise and heat, and existing filter technologies consume excessive power and generate heat, complicating qubit operation and fidelity.
Implementing folded finite impulse response (FIR) filters with a series-connected arrangement of unit delays and adders, using multiplexers and partial product stages to reduce hardware complexity and power consumption, suitable for quantum computing systems.
The solution reduces hardware complexity and power consumption, enhancing qubit fidelity and reducing crosstalk interference, allowing efficient operation in refrigerated environments.
Smart Images

Figure 2025169163000001_ABST
Abstract
Description
[Technical Field]
[0001] Aspects of the present disclosure relate to implementations of folded finite impulse response (FIR) filters suitable for use in quantum computing systems and other low-power, high-speed applications. [Background technology]
[0002] Quantum computing systems are highly susceptible to environmental influences such as magnetic field fluctuations, electromagnetic interference, and thermal noise. As a result, the hardware that forms the qubits of a quantum computing system is typically operated in a shielded environment at temperatures near absolute zero (0 K). To provide error correction and reliable operation, a quantum computing system may contain hundreds or thousands of qubits operating coherently and simultaneously.
[0003]
[0003] To avoid long wiring from external test equipment to the quantum computing system, the drive electronics can be installed and operated in a refrigerated environment, such as that provided by a dilution refrigerator. The drive electronics can include waveform synthesizers that drive the qubits. In this case, each waveform synthesizer includes a clock that runs continuously at high speed to maintain the appropriate qubit control phase synchronized with the qubit phase. However, dilution refrigerators have limited ability to remove heat generated by the drive electronics while maintaining a low operating temperature. Summary of the Invention
[0004] In one aspect, the present disclosure provides a filter circuit including a series-connected arrangement of unit delays and adders in an alternating pattern, and a plurality of coefficient multipliers, each having a respective output connected to one or more of the adders, each coefficient multiplier including a multiplexing stage including one or more multiplexers addressed using one or more bits of a respective input coefficient vector, wherein a first coefficient multiplier of the plurality of coefficient multipliers includes a partial product stage configured to provide a plurality of integer partial products of an input data vector to the multiplexing stages of the plurality of coefficient multipliers.
[0005] In one aspect, in combination with any exemplary filter circuit described above or below, the one or more multiplexers include a first multiplexer and a second multiplexer. The multiplexing stage further includes a left shifter connected to the output of the first multiplexer, where one or more bits addressing the first multiplexer are more significant than one or more bits addressing the second multiplexer. The multiplexing stage further includes an adder connected to the output of the second multiplexer, the adder connected to the output of the left shifter.
[0006]
[0006] In one aspect, in combination with any of the exemplary filter circuits described above or below, the multiplexing stage further includes a unit delay connected to the output of the adder.
[0007]
[0007] In one aspect, in combination with any of the exemplary filter circuits described above or below, the partial product stage includes one or more left shifters and one or more adders configured to receive the input data vector and the output of one of the one or more left shifters.
[0008] In one aspect, in combination with any of the exemplary filter circuits described above or below, the input coefficient vectors provided to the multiple coefficient multipliers are symmetric.
[0009] In one aspect, in combination with any of the exemplary filter circuits described above or below, one or more multiplexers of the first coefficient multiplier have respective N select lines, and the partial product stages are N It is configured to provide integer partial products up to -1.
[0010] In one aspect, the present disclosure provides a coefficient multiplier circuit for a filter, the coefficient multiplier circuit including a partial product stage configured to provide a plurality of integer partial products of an input data vector, the coefficient multiplier circuit further including a multiplexing stage including one or more multiplexers configured to receive the plurality of integer partial products of the input data vector, the one or more multiplexers being addressed using one or more bits of the input coefficient vector, and outputs connected to the one or more multiplexers and connected to one or more adders of the filter.
[0011]
[0011] In one aspect, in combination with any of the exemplary coefficient multiplier circuits described above or below, the filter is one of a finite impulse response (FIR) filter, an infinite impulse response (IIR) filter, and an autoregressive moving average (ARMA) filter.
[0012] In one aspect, in combination with any exemplary coefficient multiplier circuit described above or below, the one or more multiplexers include a first multiplexer and a second multiplexer. The multiplexing stage further includes a left shifter connected to the output of the first multiplexer, where the one or more bits addressing the first multiplexer are more significant bits of the input coefficient vector than the one or more bits addressing the second multiplexer. The multiplexing stage further includes an adder connected to the output of the second multiplexer, the adder connected to the output of the left shifter.
[0013]
[0013] In one aspect, in combination with any of the exemplary coefficient multiplier circuits described above or below, the multiplexing stage further includes a unit delay connected to the output of the adder.
[0014]
[0014] In one aspect, in combination with any of the exemplary coefficient multiplier circuits described above or below, the partial product stage includes one or more left shifters and one or more adders configured to receive the input data vector and the output of one of the one or more left shifters.
[0015] In one aspect, in combination with any of the exemplary coefficient multiplier circuits described above or below, one or more multiplexers have respective N select lines, and the partial product stages are N It is configured to provide integer partial products up to -1.
[0016] In one aspect, in combination with any example coefficient multiplier circuit described above or below, the coefficient multiplier circuit further includes a second multiplexing stage including one or more second multiplexers configured to receive a plurality of integer partial products of the input data vector, the one or more second multiplexers being addressed using one or more bits of the second input coefficient vector, and second outputs connected to the one or more second multiplexers and connected to one or more other adders of the filter.
[0017]
[0017] In one aspect, the present disclosure provides a system. The system includes a refrigeration system configured to maintain a refrigerated environment, a plurality of qubits disposed within the refrigerated environment, and control circuitry connected to the plurality of qubits. The control circuitry includes, for each qubit of the plurality of qubits, a respective waveform synthesizer configured to drive the qubit. The waveform synthesizer includes an interpolation filter circuit. The interpolation filter circuit includes a series-connected arrangement of unit delays and adders in an alternating pattern, and a plurality of coefficient multipliers, each coefficient multiplier having a respective output connected to one or more of the adders, each coefficient multiplier including a multiplexing stage including one or more multiplexers addressed using one or more bits of a respective input coefficient vector. A first coefficient multiplier of the plurality of coefficient multipliers includes a partial product stage configured to provide a plurality of integer partial products of an input data vector to the multiplexing stages of the plurality of coefficient multipliers.
[0018]
[0018] In one aspect, in combination with any exemplary system described above or below, the refrigeration system includes a dilution refrigerator defining a plurality of temperature steps, the plurality of qubits being disposed at a lowest temperature step of the plurality of temperature steps, and the control circuitry being disposed at another temperature step of the plurality of temperature steps.
[0019] In one aspect, in combination with any exemplary system described above or below, the one or more multiplexers include a first multiplexer and a second multiplexer. The multiplexing stage further includes a left shifter connected to an output of the first multiplexer, wherein one or more bits addressing the first multiplexer are more significant than one or more bits addressing the second multiplexer. The multiplexing stage further includes an adder connected to an output of the second multiplexer, the adder connected to an output of the left shifter.
[0020]
[0020] In one aspect, in combination with any of the exemplary systems described above or below, the multiplexing stage further includes a unit delay connected to the output of the adder.
[0021]
[0021] In one aspect, in combination with any of the exemplary systems described above or below, the partial product stage includes one or more left shifters and one or more adders configured to receive an input data vector and an output of one of the one or more left shifters.
[0022] In one aspect, in combination with any of the exemplary systems described above or below, the input coefficient vectors provided to the multiple coefficient multipliers are symmetric.
[0023] In one aspect, in combination with any of the exemplary systems described above or below, one or more multiplexers of the first coefficient multiplier have respective N select lines, and the partial product stages are N It is configured to provide integer partial products up to -1.
[0024]
[0024] So that the above-described features of the present disclosure can be understood in detail, a more detailed description of the present disclosure than that briefly summarized above can be made by reference to several exemplary embodiments, some of which are illustrated in the accompanying drawings. [Brief explanation of the drawings]
[0025] [Figure 1]
[0025] FIG. 1 illustrates an exemplary quantum computing system, according to one or more aspects. [Figure 2]
[0026] FIG. 1 is a block diagram of an exemplary control and measurement surface, according to one or more embodiments. [Figure 3]
[0027] FIG. 1 is a diagram of an even-order convolutional finite impulse response (FIR) filter with an example coefficient multiplier circuit, according to one or more aspects. [Figure 4]
[0028] FIG. 1 is a diagram of an odd-order folding FIR filter with an exemplary coefficient multiplier circuit, according to one or more embodiments. [Figure 5]
[0029] FIG. 1 is a diagram of an example coefficient multiplier circuit for an 8-bit input coefficient vector, according to one or more embodiments. [Figure 6]
[0030] FIG. 1 is a diagram of an example coefficient multiplier circuit for a 9-bit input coefficient vector, according to one or more aspects. [Figure 7]
[0031] FIG. 1 is a diagram of an example partial product stage for a coefficient multiplier circuit, according to one or more embodiments. [Figure 8]
[0032] 1 is an exemplary method of creating a coefficient multiplier circuit, according to one or more aspects. [Figure 9]
[0033] 1 is an exemplary method of creating a coefficient multiplier circuit, according to one or more aspects. DETAILED DESCRIPTION OF THE INVENTION
[0026]
[0034] This disclosure provides multiple implementations of folded finite impulse response (FIR) filters suitable for use in quantum computing systems. The folded FIR filters use an arbitrary coefficient multiplier-less design, which, in combination with a Horner-style folded FIR filter, requires significantly less hardware (e.g., fewer half adders, full adders, and registers). The folded FIR filter can be implemented within an interpolator to operate the waveform synthesizer at a lower rate to save power. For example, the interpolator can interpolate "missing" samples to increase the synthesized waveform from 1 gigasamples per second (GSPS) to 2 GSPS. In this example, the timing margin provided by the folded FIR filter implementation allows waveform synthesis to occur without increased hardware complexity or power consumption due to memory ping-ponging.
[0027]
[0035] Furthermore, the use of FIR filters (whether folded or unfolded) can result in higher fidelity of qubits in a quantum computing system due to reduced crosstalk from other qubit signals on shared lines. Fidelity generally refers to the qubit's state angular error, which determines the reliability of the qubit's calculations. The upsampling provided by the interpolator reduces the amplitude of dominant spurs (e.g., unwanted narrowband signals) caused by the digital-to-analog converter (DAC) output. Because multiple qubits in a quantum computing system typically share the same RF signal lines (e.g., using frequency division multiplexing), spurs can interfere with other qubits and cause them to drift from their quantum states.
[0028]
[0036] Notably, the description herein focuses primarily on FIR filters (also described as "moving average" filters) to achieve linear phase. However, the various techniques described herein are not limited to FIR filters, but are also suitable for other filter topologies, such as infinite impulse response (IIR) filters (also described as "autoregressive" filters), autoregressive moving average (ARMA) filters, autoregressive integrating moving average (ARIMA) filters, etc. These filters may be one-dimensional or multidimensional.
[0029]
[0037] In the present disclosure, reference is made to various embodiments. However, it should be understood that the disclosure is not limited to the particular described embodiments. Instead, any combination of the following features and elements, whether associated with various embodiments or not, is contemplated for implementing and practicing the teachings provided herein. Furthermore, when elements of an embodiment are described in the form of "at least one of A and B," it should be understood that embodiments including element A only, element B only, and elements A and B are each contemplated. Furthermore, while some embodiments may realize other potential solutions and / or advantages over the prior art, whether or not a particular advantage is realized by a given embodiment does not limit the disclosure. Accordingly, the embodiments, features, and advantages disclosed herein are merely exemplary and should not be considered elements of or limit the scope of the appended claim(s) unless expressly recited in the claim(s). Similarly, references to "the present invention" should not be construed as a generalization of all inventive subject matter disclosed herein, and should not be considered an element of or limiting the scope of any accompanying claim(s) unless expressly recited in the claim(s).
[0030]
[0038] In some aspects, a filter circuit includes a series-connected arrangement of unit delays and adders in an alternating pattern, and a plurality of coefficient multipliers, each coefficient multiplier having a respective output connected to one or more of the adders and including a multiplexing stage with one or more multiplexers addressed using one or more bits of a respective input coefficient vector, and a first coefficient multiplier of the plurality of coefficient multipliers includes a partial product stage configured to provide a plurality of integer partial products of the input data vector to the multiplexing stage of the plurality of coefficient multipliers.
[0031]
[0039] In this manner, coefficient multiplication is performed using partial products and multiplexers instead of using Booth multipliers, Wallace tree multipliers, or other conventional multiplier architectures that are more complex and power-hungry. The partial products of the input data vector are pre-computed once for each value of the input data vector and used in all coefficient multiplications. Furthermore, because the coefficient values of the input coefficient vector do not need to be static, the filter circuit may be suitable for adaptive filtering applications, time-varying equalizer applications, etc.
[0032]
[0040] 1 illustrates an exemplary quantum computing system 100 in accordance with one or more aspects. Features illustrated in quantum computing system 100 may be used in conjunction with other aspects described herein.
[0033]
[0041] As shown, quantum computing system 100 comprises refrigeration system 105, quantum data surface 110, control and measurement surface 115, control processor surface 120, and host processor 125. In some aspects, quantum data surface 110 comprises hardware that physically forms multiple qubits 112 as well as structures used to support and / or hold multiple qubits 112. Multiple qubits 112 may have any suitable form, such as quantum dots, superconducting circuits, etc.
[0034]
[0042] In some aspects, quantum data surface 110 comprises additional circuitry that operates to measure the states of qubits 112 and manipulate the states of qubits 112 when performing operations. For example, gate operations may be performed using control signals that alter the Hamiltonians (i.e., descriptions of the total energy and state dynamics) of qubits 112. In some aspects, quantum information in qubits 112 may be stored, altered, and / or retrieved by transmitting microwave photons across qubits 112.
[0035]
[0043] Control and measurement surface 115 comprises hardware that receives digital control signals from control processor surface 120 and converts the digital control signals into analog (or wave) control signals that are read and executed at quantum data surface 110 to perform quantum operations on multiple qubits 112. In some aspects, control and measurement surface 115 comprises one or more coaxial cables or waveguides that support the transmission (and, in some cases, shielding) of signals to and from quantum data surface 110. For example, in one implementation of quantum data surface 110 having multiple qubits 112 formed in superconducting circuits, control signals may be transmitted to multiple qubits 112 through microwave waveguides that extend through refrigeration system 105.
[0036]
[0044] Control and measurement plane 115 further comprises hardware that receives analog outputs (representing measurements of qubits) from quantum data plane 110 and converts the analog outputs to digital signals that are transmitted to control processor plane 120. The hardware included within control and measurement plane 115 may include additional shielding to mitigate the effects of environmental noise on the analog outputs received from quantum data plane 110.
[0037]
[0045] Control processor surface 120 implements a quantum algorithm or sequence of quantum operations and provides corresponding instructions that are implemented in control and measurement surface 115. A host processor 125 is communicatively coupled to control processor surface 120 and provides digital signal(s) that implement and / or interact with the quantum algorithm in control processor surface 120.
[0038]
[0046] In some aspects, quantum algorithms are implemented using quantum circuits. Each quantum circuit represents a computing routine having a sequence of quantum operations on multiple qubits 112 of quantum data plane 110. In some aspects, quantum algorithms may be implemented using development tools and libraries. In some aspects, quantum algorithms include a sequence of gate operations and measurements performed on control and measurement plane 115.
[0039]
[0047] In some aspects, control processor surface 120 is implemented as one or more processors and memory (not shown). The one or more processors are any electronic circuitry, including, but not limited to, one or a combination of a microprocessor, a microcontroller, an application specific integrated circuit (ASIC), an application specific instruction set processor (ASIP), and / or a state machine, that is communicatively coupled to memory and controls the operation of control processor surface 120.
[0040]
[0048] The one or more processors may include other hardware that runs software for controlling and processing information. The one or more processors execute software stored in memory to perform any of the functions described herein. The one or more processors control the operation and management of control processor surface 120 by processing information (e.g., information received from input devices and / or communicatively coupled electronic devices).
[0041]
[0049] Memory may store, either permanently or temporarily, data, operable software, or other information for one or more processors. Memory may include any one or combination of volatile or non-volatile local or remote devices suitable for storing information. For example, memory may include random access memory (RAM), read-only memory (ROM), magnetic storage devices, optical storage devices, or any other suitable information storage device, or a combination of these devices. Software represents any suitable set of instructions, logic, or code embodied in a computer-readable storage medium, such as memory. In certain embodiments, software may include applications executable by one or more processors to perform one or more of the functions described herein.
[0042]
[0050] An exemplary implementation of control and measurement plane 115 is provided in FIG. 2. In block diagram 200, control and measurement plane 115 includes control circuitry 205 and measurement circuitry 220. Control circuitry 205 includes multiple waveform synthesizers 210. Multiple waveform synthesizers 210 include multiple interpolation filters 215. In some aspects, each waveform synthesizer in multiple waveform synthesizers 210 corresponds to a respective qubit in multiple qubits 112. In some aspects, each interpolation filter in multiple interpolation filters 215 corresponds to a respective waveform synthesizer in multiple waveform synthesizers 210. Exemplary implementations of multiple interpolation filters 215 are described below with respect to FIGS. 3 and 4.
[0043]
[0051] The multiple waveform synthesizers 210 may have any suitable implementation. For example, the multiple waveform synthesizers 210 may be implemented with a direct digital synthesizer (DDS), a phase-locked loop (PLL), etc. The multiple waveform synthesizers 210 synthesize waveforms at any suitable frequency to support the sampling rate required for the quantum computing system 100. For example, the sampling rate may be greater than 1 GSPS, such as 2 GSPS or greater.
[0044]
[0052] The multiple interpolation filters 215 may have any suitable implementation. In some aspects, the multiple interpolation filters 215 include a low-pass filter. This low-pass filter contributes “missing” samples to the synthesized waveform. This effectively increases the overall sampling rate of the multiple waveform synthesizers 210. For example, assuming a target sampling rate of 2 GSPS, the multiple waveform synthesizers 210 may operate at 1 GSPS, and the multiple interpolation filters 215 may effectively increase the sampling rate from 1 GSPS to 2 GSPS. Other sampling rates and upsampling factors are also contemplated. In this manner, the multiple waveform synthesizers 210 may be implemented with less hardware complexity and / or may have lower power consumption. This tends to reduce the cost of implementing the multiple waveform synthesizers 210, as well as the power consumed and heat generated by the multiple waveform synthesizers 210. This heat needs to be removed from the refrigerated environment 140.
[0045]
[0053] In some embodiments, control circuitry 205 transmits multiple pulse-level analog control signals over first conductive link 135-1 to quantum data surface 110. This is to configure quantum dots for the multiple qubits 112, for example, by electrostatically creating potential wells that move around and trap electrons that will form the multiple qubits 112. Although not shown, quantum data surface 110 may further include magnets (e.g., superconducting magnet coils or permanent magnets). These magnets establish a static magnetic field that defines the primary qubit state reference axis (e.g., the north-south Bloch sphere). The phase of a precisely maintained intermediate frequency defines and sets the position of the east-west axis. Quantum data surface 110 may also include smaller magnets (e.g., “micromagnets” positioned adjacent to the multiple qubits 112). These magnets are used to determine the resonant frequencies of each of the multiple qubits 112. In this manner, the control signals transmitted by control circuitry 205 to the multiple qubits 112 may be frequency-division multiplexed (FDM). Thereby, the first conductive link 135-1 may be implemented as a single coaxial cable.
[0046]
[0054] In some aspects, control circuitry 205 resets the plurality of qubits 112 after control and measurement surface 115 configures the plurality of qubits 112. Control circuitry 205 transmits RF gate pulses interspersed with pulsed analog levels (e.g., using waveform synthesizer 210 with multiple interpolation filters 215) to perform qubit gate operations that configure a subset of the plurality of qubits 112 (e.g., some or all of the plurality of qubits 112) to perform a calculation. Any type of quantum gate is contemplated, including single-qubit gates such as Poly-X, Poly-Y, and Poly-Z gates, as well as multi-qubit gates such as Hadamard gates and CNOT gates. Collectively, the quantum gates formed by the subsets may be referred to as a quantum algorithm. In some aspects, control circuitry 205 transmits the RF gate pulses using second conductive link 135-2.
[0047]
[0055] Upon completion of the computation, the result of the computation (e.g., a binary number) is encoded into the states of the subset of the plurality of qubits 112. In some aspects, control circuitry 205 transmits a combination of analog and RF pulses (e.g., using waveform synthesizer 210 having multiple interpolation filters 215) to access and read out the states of the subset of the plurality of qubits 112. In some aspects, reading out the states of the subset of the plurality of qubits 112 includes detecting a state-dependent response (e.g., a change in current corresponding to state-preferential tunneling) caused by interaction with the transmitted analog and RF pulses. In this manner, the characteristics of the analog and RF pulses are selected to minimize disturbances to the states of the qubits while still providing accurate measurements. In some aspects, control circuitry 205 transmits the pulsed analog signal using third conductive link 135-3.
[0048]
[0056] State information describing the state of the subset of the plurality of qubits 112 is transmitted over fourth conductive link 135-4 to control and measurement plane 115. In some aspects, the state information is provided to measurement circuitry 220. Measurement circuitry 220 may include hardware such as amplifiers and detectors. Measurement circuitry 220 may perform further processing to determine the state of the subset of the plurality of qubits 112. Based on the state, control and measurement plane 115 may adjust one or more control parameters for subsequent operations.
[0049]
[0057] In some embodiments, some of the components of quantum computing system 100 operate within a refrigerated environment 140 maintained by refrigeration system 105. Refrigerated environment 140 mitigates the effects of thermal noise on various components of quantum computing system 100. As shown, control and measurement surface 115 and quantum data surface 110 hardware are located within refrigerated environment 140, while host processor 125 and control processor surface 120 are located outside of refrigerated environment 140. Other embodiments may have various configurations, such as some or all of control processor surface 120 being located within refrigerated environment 140. In some embodiments, host processor 125 is considered a “classical” processor, which includes any electronic circuitry, including, but not limited to, one or a combination of a microprocessor, a microcontroller, an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), and / or a state machine.
[0050]
[0058] In some embodiments, refrigeration system 105 uses a liquefied gas (such as liquid nitrogen, liquid hydrogen, or liquid helium) in fluid communication with refrigerated environment 140 to maintain a temperature(s) below approximately 120 K (-153°C). In some embodiments, refrigerated environment 140 has a temperature(s). The temperature(s) are very close to absolute zero (0 K, -273.15°C) because superconductivity and its associated quantum properties are most pronounced at these temperatures.
[0051]
[0059] In some embodiments, the refrigeration system 105 comprises a dilution refrigerator. A dilution refrigerator typically comprises a mixing chamber, a still, and a series of heat exchangers. The dilution refrigerator defines a plurality of temperature stages 130-1, 130-2, ..., 130-5 operating at progressively lower temperatures. Five (5) stages 130-1, 130-2, ..., 130-5 are illustrated, corresponding to 35K, 4K, 900mK, 100mK, and 10mK, respectively, although different numbers of stages and different temperatures are also contemplated for dilution refrigerators.
[0052]
[0060] The dilution process begins with a liquid mixture of isotopes, e.g., helium-3 and helium-4. Exposing this mixture to low temperatures causes the helium-3 component to enter a superfluid state. As the mixture is pumped through a series of heat exchangers and stages, the helium-3 component is used to cool the helium-4 component through dilution, thereby lowering the temperature of the mixture. The mixing chamber is the coldest portion of the dilution refrigerator, e.g., located at temperature stage 130-5. In some embodiments, multiple qubits 112 are located at the lowest temperature stage of the multiple temperature stages (e.g., temperature stage 130-5), and control circuitry 205 is located at another temperature stage of the multiple temperature stages (e.g., temperature stage 130-1).
[0053]
[0061] 3 is a diagram 300 of an even-order convolutive finite impulse response (FIR) filter 302 (also referred to as "filter 302") having an example coefficient multiplier circuit, according to one or more aspects. The features illustrated in diagram 300 may be used in conjunction with other aspects. For example, filter 302 may represent an example implementation of one of the interpolation filters 215 of FIG. 2.
[0054]
[0062] The filter 302 comprises a serial arrangement 305 of unit delays 310-1, 310-2, 310-3, 310-4 and summers 315-1, 315-2, 315-3, 315-4 connected in an alternating pattern. The filter 302 is illustrated as a fourth-order folded FIR filter having four (4) unit delays 310-1, ..., 310-4 and four (4) summers 315-1, ..., 315-4, although other orders and arrangements of unit delays and summers are also contemplated. Notably, the coefficient multiplier circuitry described herein is also contemplated for use with other types of filters, such as non-folded FIR filters, IIR filters, ARMA filters, and ARIMA filters.
[0055]
[0063] The filter 302 further comprises a plurality of coefficient multipliers 320-0, 320-1, 320-2. Each coefficient multiplier has a respective output connected to one or more of the adders 315-1, ..., 315-4. Each of the plurality of coefficient multipliers 320-0, 320-1, 320-2 receives a respective input coefficient vector C0, C1, C2. As shown, the input coefficient vectors C0, C1, C2 are 8 bits long, although other lengths are contemplated.
[0056]
[0064] The filter 302 implements the Horner form of the FIR filter equation: TIFF2025169163000002.tif23170 where Y(z) represents the z-transform of the output data, N represents the order of the filter 302, and c k represents the filter coefficients (e.g., impulse response values), and z -1 represents a unit sample delay (e.g., implemented using a register), and X(z) is the delay of the input data vector x n Represents the z-transform of
[0057]
[0065] FIR filters such as filter 302 may achieve precise linear phase for distortion-free operation, which is advantageous when used in quantum computing system 100. In such cases, the input coefficient vectors C0, C1, and C2 provided to multiple coefficient multipliers 320-0, 320-1, and 320-2 are symmetric about their centers. For example, filter 302 is considered symmetric because all of the input coefficient vectors other than input coefficient vector C2 (i.e., input coefficient vectors C0 and C1) are reused. In even-order FIR filters such as filter 302, the central input coefficient vector C1 is not reused. Thus, filter 302 has a “folded” configuration, and the number of multipliers required to implement filter 302 is reduced by a factor of two, thereby reducing manufacturing costs, energy consumption, and / or heat generated by filter 302. However, in other aspects, filter 302 need not be implemented with symmetric input coefficient vectors.
[0058]
[0066] In some aspects, each of the plurality of coefficient multipliers 320-0, 320-1, 320-2 comprises a multiplexing stage (described below with reference to FIGS. 5 and 6) comprising one or more multiplexers that are addressed using one or more bits of a respective input coefficient vector C0, C1, C2. In some aspects, only the first coefficient multiplier 320-0 of the plurality of coefficient multipliers 320-0, 320-1, 320-2 provides the input data vector x to the multiplexing stage of the plurality of coefficient multipliers 320-0, 320-1, 320-2. n 5 and 6 , configured to provide a plurality of integer partial products of In this manner, the partial products are pre-multiplied by the first coefficient multiplier 320-0, and the downstream coefficient multipliers 320-1, 320-2 do not need to include hardware to perform redundant (one or more) multiplication operations.
[0059]
[0067] As shown in diagram 300, the first coefficient multiplier 320-0 multiplies the input coefficient vector C0 by the input data vector x nThe first coefficient multiplier 320-0 receives the product C0x from the unit delay 310-1 and the adder 315-4. n-3 The output provides the product C0x n-3 is the input data vector x n (e.g., corresponding to the three pipeline stages of the first multiplier 320-0 as shown in FIG. 5). The second coefficient multiplier 320-1 receives the input coefficient vector C1 and the partial products from the first coefficient multiplier 320-0. The second coefficient multiplier 320-1 outputs the product C1x to adders 315-1, 315-3. n-3 The third coefficient multiplier 320-2 receives the input coefficient vector C2 and the partial products from the first coefficient multiplier 320-0. The third coefficient multiplier 320-2 outputs the product C2x to the adder 315-2. n-3 The output provides:
[0060]
[0068] The unit delay 310-1 outputs the delay product C0x to the adder 315-1. n-4 Adder 315-1 adds the sum C0x to unit delay 310-2. n-4 +C1x n-3 The unit delay 310-2 provides the delay sum C0x to the adder 315-2. n-5 +C1x n-4 Adder 315-2 adds the sum C0x to unit delay 310-3. n-5 +C1x n-4 +C2x n-3 The unit delay 310-3 provides the delay sum C0x to the adder 315-3. n-6 +C1x n-5 +C2x n-4 Adder 315-3 adds the sum C0x to unit delay 310-4. n-6 +C1x n-5 +C2x n-4 +C1x n-3 The unit delay 310-4 provides the delay sum C0x to the adder 315-4. n-7 +C1x n-6 +C2x n-5 +C1x n-4 Adder 315-4 provides C0x as the output of filter 302. n-7 +C1xn-6 +C2x n-5 +C1x n-4 +C0x n-3 to provide.
[0061]
[0069] 4 is a diagram 400 of an odd-order folding FIR filter 402 (also referred to as "filter 402") having an example coefficient multiplier circuit, according to one or more aspects. The features illustrated in diagram 400 may be used in conjunction with other aspects. For example, filter 402 may represent an example implementation of one of the interpolation filters 215 of FIG. 2.
[0062]
[0070] The filter 402 comprises a serial arrangement 405 of unit delays 310-1, 310-2, 310-3 and summers 315-1, 315-2, 315-3 connected in an alternating pattern. The filter 402 is illustrated as a third order filter having three (3) unit delays 310-1, 310-2, 310-3 and three (3) summers 315-1, 315-2, 315-3, although other orders and arrangements of unit delays and summers are also contemplated.
[0063]
[0071] As shown, the first coefficient multiplier 320-0 multiplies the input coefficient vector C0 by the input data vector x n The first coefficient multiplier 320-0 receives the product C0x from the unit delay 310-1 and the adder 315-3. n-3 The second coefficient multiplier 320-1 receives the input coefficient vector C1 and the partial products from the first coefficient multiplier 320-0. The second coefficient multiplier 320-1 outputs the product C1x to adders 315-1 and 315-2. n-3 The output provides:
[0064]
[0072] The unit delay 310-1 outputs the delay product C0x to the adder 315-1. n-4 Adder 315-1 adds the sum C0x to unit delay 310-2. n-4 +C1x n-3 The unit delay 310-2 provides the delay sum C0x to the adder 315-2. n-5 +C1xn-4 Adder 315-2 adds the sum C0x to unit delay 310-3. n-5 +C1x n-4 +C2x n-3 The unit delay 310-3 provides the delay sum C0x to the adder 315-3. n-6 +C1x n-5 +C1x n-4 Adder 315-3 provides C0x as the output of filter 402. n-6 +C1x n-5 +C1x n-4 +C0x n-3 to provide.
[0065]
[0073] 5 is a diagram 500 of an example coefficient multiplier circuit 502 for an 8-bit input coefficient vector, in accordance with one or more aspects. The features illustrated in diagram 500 may be used in conjunction with other aspects. For example, coefficient multiplier circuit 502 may represent an example implementation of one of coefficient multipliers 320-0, 320-1, 320-2 of FIG. 3. Furthermore, although the features of coefficient multiplier circuit 502 are described with respect to a folding FIR filter, it is contemplated that coefficient multiplier circuit 502 may be used in other applications.
[0066]
[0074] The coefficient multiplier circuit 502 multiplies the input data vector x n In some aspects, the first coefficient multiplier circuit 502 includes a partial product stage 505 that provides a plurality of integer partial products of an input data vector x n to one or more subsequent coefficient multiplier circuits 502 that do not include a partial product stage 505.
[0067]
[0075] In some aspects, the partial product stage 505 includes one or more left shifters and a nand one or more adders configured to receive the output of one of the one or more left shifters. As shown, the partial product stage 505 includes one (1) left shifter 515-1 and one (1) adder 520-1, although other numbers are contemplated. In some aspects, the one or more left shifters are implemented using wiring reassignment and do not require any additional hardware components.
[0068]
[0076] The left shift operation by n bits is 2 n As shown, left shifter 515-1 performs a one-place shift (“<<1”). This is equivalent to multiplying the input data vector x n The adder 520-1 multiplies the input data vector x n and the input data vector x n and a one-place left-shifted version of . Thus, the output of adder 520-1 is 3x n is.
[0069]
[0077] In some aspects, the partial product stage 505 further comprises unit delays 522-1, 522-2. Each unit delay may be implemented as a pipeline register. The unit delay 522-1 delays the input data vector x n-1 and unit delay 522-2 provides a sum delay of 3x n-1 Thus, the partial product stage 505 provides the input data vector x n-1 The product of two integer parts (here, x n-1 and 3x n-1 The two integer partial products are provided to a multiplexing stage 510 of the coefficient multiplier circuit 502 as well as to subsequent coefficient multiplier circuits 502.
[0070]
[0078] The multiplexing stage 510 comprises one or more multiplexers configured to receive a plurality of integer partial products of an input data vector. As shown, the multiplexing stage 510 comprises four (4) multiplexers 530-1, 530-2, ..., 530-4 that are addressed using two bits. The one or more multiplexers 530-1, 530-2, ..., 530-4 receive the input coefficient vector Cn , 530-4. Although each of multiplexers 530-1, 530-2, ..., 530-4 is illustrated as a respective four-way multiplexer, in other aspects, multiplexing stage 510 may include different configurations (e.g., number and / or size) of multiplexers. Exemplary techniques for determining an efficient configuration of multiplexers are described below with respect to FIG. 8. Multiplexers 530-1, 530-2, ..., 530-4 may have any suitable implementation, such as logic gates or transmission gates.
[0071]
[0079] Each of the multiplexers 530-1, ..., 530-4 receives an input data vector x n-1 (i.e., x n-1 and 3x n-1 ) as an input. In some aspects, multiplexing stage 510 further comprises a plurality of left shifters 525-1, 525-2, ..., 525-4 corresponding to multiplexers 530-1, ..., 530-4. Each of left shifters 525-1, 525-2, ..., 525-4 receives x as another input to the corresponding multiplexer 530-1, ..., 530-4. n-1 A left-shifted version of (i.e., 2x n-1 In an alternative embodiment, the partial product stage 505 performs a one-place shift to provide x n-1 A left-shifted version of (i.e., 2x n-1 ) to the multiplexing stage 510. However, by using multiple left shifters 525-1, 525-2, ..., 525-4, x n-1 By reallocating wiring, a one-place left-shifted version of can be provided to multiplexers 530-1, ..., 530-4 without the need for additional unit delays to be included in partial product stage 505. The remaining inputs to multiplexers 530-1, ..., 530-4 are connected to ground (zero value).
[0072]
[0080] In some aspects, the one or more multiplexers of multiplexing stage 510 comprise a first multiplexer and a second multiplexer (e.g., a pair of multiplexers 530-1, 530-2, or a pair of multiplexers 530-3, 530-4). In some aspects, multiplexing stage 510 further comprises a left shifter (e.g., left shifter 535-1 or 535-2) connected to the output of the first multiplexer (e.g., multiplexer 530-1 or multiplexer 530-3). Left shifters 535-1, 535-2 provide two-place left-shifted versions of the outputs of respective multiplexers 530-1, 530-3.
[0073]
[0081] As shown, multiplexer 530-4 receives the input coefficient vector C n The first multiplexer is addressed using the two least significant bits b1, b0 of the input coefficient vector C, the second multiplexer is addressed using the next two bits b3, b2, the second multiplexer is addressed using the next two bits b5, b4, and the third multiplexer is addressed using the next two bits b7, b6. In this manner, one or more bits addressing the first multiplexer (e.g., b7, b6, or b3, b2) may be more significant than one or more bits addressing the second multiplexer (e.g., b5, b4, or b1, b0). n is a bit.
[0074]
[0082] In some aspects, multiplexing stage 510 further comprises an adder (e.g., adder 540-1 or adder 540-2) connected to the output of the second multiplexer (e.g., multiplexer 530-2 or multiplexer 530-4) and connected to the output of the left shifter (e.g., left shifter 535-1 or left shifter 535-2). Thus, the output of multiplexer 530-1 is shifted left by two positions using left shifter 535-1. Adder 540-1 receives the output of multiplexer 530-2 and the two-position left-shifted output of multiplexer 530-1. Adder 540-1 provides a sum shifted left by four bits using left shifter 545. The four-bit left-shifted output is provided to unit delay 550-1 (e.g., a pipeline register), and the delayed output is provided to adder 555. In some aspects, adders 520-1, 540-1, 540-2, 555 are implemented with full precision using carry look-ahead adders or subtractors.
[0075]
[0083] The output of multiplexer 530-3 is shifted left by two bits using left shifter 535-2. Adder 540-2 receives the output of multiplexer 530-4 and the two-bit left-shifted output of multiplexer 530-3. Adder 540-2 provides a sum that is provided to unit delay 550-2 (e.g., a pipeline register), the delayed output of which is provided to adder 555. The sum provided by adder 555 is provided to unit delay 560-1, and the output of coefficient multiplier circuit 502 is used to generate product C n x n-3 In some aspects, the output of the coefficient multiplier circuit 502 is connected to one or more multiplexers 530-1, 530-2, ..., 530-4 and connected to one or more adders of a folding FIR filter.
[0076]
[0084] In some aspects, as shown, unit delays 522-1, 522-2, 550-1, 550-2, 560-1 (e.g., pipeline registers) are included within coefficient multiplier circuit 502, thereby allowing only a single addition or subtraction operation to be performed by coefficient multiplier circuit 502 during each clock cycle. These aspects may be suitable for particular manufacturing technologies or applications, for example, where speed or transport delay is not critical. Accordingly, in other aspects, unit delays 522-1, 522-2, 550-1, 550-2, 560-1 may be omitted from coefficient multiplier circuit 502.
[0077]
[0085] Thus, in some aspects, coefficient multiplier circuit 502 pre-calculates the triple x value (also called the "triple" or "x3" partial product) once for all coefficients. In a three-coefficient system such as that shown in FIG. 3, this saves one adder, but in systems with a larger number of coefficients, the benefit is much greater.
[0078]
[0086] 6 is a diagram 600 of an example coefficient multiplier circuit 602 for a 9-bit input coefficient vector, in accordance with one or more aspects. The features illustrated in diagram 600 may be used in conjunction with other aspects. For example, coefficient multiplier circuit 602 may represent an example implementation of one of coefficient multipliers 320-0, 320-1, 320-2 of FIG. 3. Furthermore, although the features of coefficient multiplier circuit 602 are described with respect to a folding FIR filter, it is contemplated that coefficient multiplier circuit 602 may be used in other applications.
[0079]
[0087] The coefficient multiplier circuit 602 multiplies the input data vector x n In some aspects, the first coefficient multiplier circuit 602 includes a partial product stage 605 that provides a plurality of integer partial products of an input data vector x n to one or more subsequent coefficient multiplier circuits 602 that do not include the partial product stage 605 (eg, all subsequent coefficient multiplier circuits 602).
[0080]
[0088] The multiplexing stage 606 multiplexes the input data vector x n As shown, the multiplexing stage 606 comprises three (3) multiplexers 625-1, 625-2, 625-3 that are addressed using three bits. The multiplexers 625-1, 625-2, 625-3 receive the input coefficient vector C n Each of multiplexers 625-1, 625-2, 625-3 is addressed using one or more bits of 0 to 2. Although each of multiplexers 625-1, 625-2, 625-3 is illustrated as a respective 8-way multiplexer, in other aspects multiplexing stage 606 may include different configurations (e.g., number and / or size) of multiplexers. Thus, in some aspects, each of the one or more multiplexers has a respective M select line, and partial product stage 605 may address one or more of the multiplexers using a number of select lines ranging from 0 to 2. M -1 (M=3 for coefficient multiplier circuit 602) n provides the integer partial products of
[0081]
[0089] In some aspects, the partial product stage 605 includes one or more left shifters and a n and the output of one of the one or more left shifters. As shown, partial product stage 605 includes three (3) left shifters 610-1, 610-2, 615 and three (3) adders 620-1, 620-2, 620-3, although other numbers are contemplated. Left shifters 610-1, 610-2 each perform a one-place shift, corresponding to a multiplication by two. Left shifter 615 performs a two-place shift, corresponding to a multiplication by four.
[0082]
[0090] Adder 620-1 adds the input data vector x n and the input data vector x from the left shifter 610-1. n and a one-place left-shifted version of . Thus, the output of adder 620-1 is 3x n The adder 620-2 adds the input data vector xn and the input data vector x from the left shifter 615 n and a left-shifted version of . The output of adder 620-2 is therefore 5x n The left shifter 610-2 takes the output of the adder 620-1 (i.e., 3x n ) The adder 620-3 receives the input data vector x n and a one-place left-shifted version of the output from adder 620-1 (i.e., 6x n ) and the output of adder 620-3 is therefore 7x n is.
[0083]
[0091] Each of the multiplexers 625-1, 625-2, and 625-3 receives an input data vector x n (i.e., x n , 2x n , …, 7x n ) as input. The remaining inputs to multiplexers 625-1, 625-2, and 625-3 are ground (input x n As shown, multiplexer 625-3 is connected to an input coefficient vector C n The first multiplexer 625 is addressed using the three least significant bits b2, b1, b0, the second multiplexer 625-1 is addressed using the next three bits b5, b4, b3, and the third multiplexer 625-2 is addressed using the three most significant bits b8, b7, b6.
[0084]
[0092] In some aspects, multiplexing stage 606 further comprises an adder 630 connected to the output of multiplexer 625-3 and a left shifter 635 connected to the output of multiplexer 625-2. The left shifter 635 provides a three-place left-shifted version of the output to adder 630. Multiplexing stage 606 further comprises an adder 640 connected to the output of adder 630 and a left shifter 645 connected to the output of multiplexer 625-1. The left shifter 645 provides a six-place left-shifted version of the output to adder 640. In all these cases, the number of bits being shifted in multiplexing stage 606 is equal to the number of coefficient bits used in the right adder input. For example, left shifter 645 shifts by six bits because the input of right adder 640 uses six bits, b5, b4, b3, b2, b1, and b0, in calculating its right input value.
[0085]
[0093] In some alternative aspects, the partial product stage 605 and / or the multiplexing stage 606 may include a unit delay, such that only a single addition operation is performed by the coefficient multiplier circuit 602 during each clock cycle. In some aspects, a coefficient multiplier circuit, such as the coefficient multiplier circuit 502, 602, performs unsigned multiplication. The sign of the multiplication may be transferred to the product as follows: n or C n If only one of is negative, the result is negative, otherwise the result is positive. The multiplication operation is described in terms of positive integers, but floating-point filtering is also supported, since the integer values of the coefficients and the scaling of the filter's output take into account implicit decimal places.
[0086]
[0094] Beneficially, by using the partial product stage 605, the input data vector x n Integer partial products of (e.g., 0, x n , 2x n , …, 7x n) may all be pre-computed, which eliminates approximately three (3) n-bit wide carry look-ahead adders for each subsequent coefficient multiplier circuit 602, at the expense of CMOS and gate implementation in alternative forms of additional transmission gates (twice the word width transistors) or multiplexers.
[0087]
[0095] Furthermore, significant power savings in FIR filters can be realized using the partial product stage 605 due to reduced hardware requirements and the static nature of the multiplexer selection. By comparison, a conventional implementation of a fully pipelined, folded, 11th-order FIR filter incorporating a fast Wallace tree multiplier requires 297 half adders, 529 full adders, and 2418 registers. Using the techniques described herein, an 11th-order FIR filter can be implemented with 106 half adders (a 64% reduction), 350 full adders (a 34% reduction), and 1031 registers (a 57% reduction). High throughput of the FIR filter is achieved by pipelining all additions. For example, a simulation of an 11th-order FIR filter using the coefficient multiplier circuit 602 runs at 2 GSPS, significantly improving the timing margins of digital synthesis.
[0088]
[0096] 7 is a diagram 700 of an example partial product stage 702 for a coefficient multiplier circuit, in accordance with one or more embodiments. The features illustrated in diagram 700 may be used in conjunction with other embodiments. For example, partial product stage 702 may be used for a multiplexing stage that addresses its multiplexer using four bits. Furthermore, although the features of partial product stage 702 are described with respect to a coefficient multiplier circuit, it is contemplated that partial product stage 702 may be used in other applications.
[0089]
[0097] In some aspects, the partial product stage 702 includes one or more left shifters and a nand one or more adders configured to receive the output of one of the one or more left shifters (or adders). As shown, partial product stage 702 includes eight (8) left shifters 705-1, 705-2, 705-3, 705-4, 710-1, 710-2, 710-3, 715 and seven (7) adders 720-1, 720-2, ..., 720-7, although other numbers are also contemplated. Left shifters 705-1, 705-2, 705-3, 705-4 each perform a one-position shift corresponding to a multiplication by two, left shifters 710-1, 710-2, 710-3 each perform a two-position shift corresponding to a multiplication by four, and left shifter 715 performs a three-position shift corresponding to a multiplication by eight.
[0090]
[0098] Adder 720-1 adds the input data vector x n and the input data vector x from the left shifter 705-1. n and a one-place left-shifted version of . Thus, the output of adder 720-1 is 3x n The adder 720-2 adds the input data vector x n and the input data vector x from the left shifter 710-1. n and a left-shifted version of 5× n The adder 720-3 adds the input data vector x n and the input data vector x from the left shifter 715 n and a three-place left-shifted version of . Thus, the output of adder 720-3 is 9x n is.
[0091]
[0099] The left shifter 705-2 shifts the output of the adder 720-1 (i.e., 3x n ) The adder 720-4 receives the input data vector x n and a one-place left-shifted version of the output from adder 720-1 (i.e., 6x n ) and the output of adder 720-4 is therefore 7x n is.
[0092]
[0100] The left shifter 705-3 shifts the output of the adder 720-2 (i.e., 5x n ), and receives a one-place left-shifted version of the output from adder 720-2 (i.e., 10x n ) is provided by the left shifter 705-4. n ), and receives a one-place left-shifted version of the output from adder 720-4 (i.e., 14x n ) is provided.
[0093]
[0101] The left shifter 720-2 takes the output of the adder 720-1 (i.e., 3x n ), and receives a left-shifted version of the output from adder 720-1 by two (i.e., 12x n ) The left shifter 710-3 provides the output of the adder 720-1 (i.e., 3x n ), and receives a left-shifted version of the output from adder 720-1 by two (i.e., 12x n ) is provided.
[0094]
[0102] The left shifter 720-5 shifts the output from the adder 720-3 (i.e., 9x n ) and the input data vector x from the left shifter 705-1 n A left-shifted version of (i.e., 2x n ) and the output of adder 720-5 is therefore 11x n The adder 720-6 adds the input data vector x n and the output from the left shifter 710-2 (i.e., 12x n ) and the output of adder 720-6 is therefore 13x n The adder 720-7 adds the output of the adder 720-1 (i.e., 3x n ) and a two-place left-shifted version of the output from left shifter 710-3. Thus, the output of adder 720-7 is 15x n Other combinations of sums and left shifts of the partial product stage 702 values are also contemplated.
[0095]
[0103] 8 is an example method 800 for creating a coefficient multiplier circuit according to one or more aspects. Method 800 may be used in conjunction with other aspects. For example, method 800 may be used to determine an efficient configuration of multiplexers for a multiplexing stage of a coefficient multiplier circuit. This multiplexing stage may, in some cases, be performed during the design phase of the coefficient multiplier circuit. Furthermore, in some aspects, method 800 may be performed in conjunction with one or more electronic devices, whether implemented as a circuit (e.g., an integrated circuit), computer program code, or the like.
[0096]
[0104] Method 800 begins at optional block 805, where the electronic device receives user input specifying a maximum multiplexer size for the multiplexers in the coefficient multiplier circuit. Some examples of maximum multiplexer sizes include 2-way, 4-way, and 8-way (corresponding to 1-bit, 2-bit, and 3-bit addressing, respectively). The maximum multiplexer size threshold is upper bounded by the complexity of partial product calculations. At optional block 815, the electronic device receives user input specifying a threshold number of 1-bit AND gate multiplexers to be included in the coefficient multiplier circuit. A large number of 1-bit AND gate multiplexers increases the number of adders required below the multiplexers (e.g., as shown in FIGS. 5 and 6).
[0097]
[0105] In block 825, the electronic device decomposes the length of the input coefficient vector to provide integer divisions corresponding to the combinations of one or more multiplexers in the coefficient multiplier circuit. For example, for a 9-bit input coefficient vector C n Using the example of FIG. 6 with nThe length of can be decomposed into the following integer divisions: That is, {9, 8+1, 7+2, 7+1+1, 6+3, 6+2+1, 6+1+1+1, 5+4, 5+3+1, 5+2+2, 5+2+1+1, 5+1+1+1+1, 4+4+1, 4+3+2, 4+3+1+1, 4+2+2+1, 4+2+1+1+1, 4+1+1+1+1+1, 3+3+3, 3+3+2+1, 3+3+1+1+1, 3+2+2+2, 3+2+2+1+1, 3+2+1+1+1+1, 3+1+1+1+1+1, 2+2+2+2+1, 2+2+2+1+1+1, 2+2+1+1+1+1+1, 2+1+1+1+1+1+1+1, 1+1+1+1+1+1+1+1+1}. In some aspects, each "1" instance occurring in the set of values may be implemented as a respective 1-bit AND gate multiplexer in the multiplexing stage. Each non-"1" instance may be implemented as a respective multiplexer in the multiplexing stage, where the non-"1" value represents the number of address bits in the select address input. Thus, the multiplexing stage may be created as a 9-bit addressing multiplexer (512-way), an 8-bit addressing multiplexer (256-way), a 1-bit AND gate multiplexer, etc.
[0098]
[0106] In block 835, the electronic device eliminates from the integer division one or more candidate divisions corresponding to at least a threshold number of 1-bit AND gate multiplexers in the coefficient multiplier circuit. Having too many 1-bit AND gate multiplexers in the multiplexing stage may result in unacceptable hardware complexity, transport delays, and / or power consumption in the multiplexing stage. Assuming the threshold number of 1-bit AND gate multiplexers is two (indicating that only one 1-bit AND gate multiplexer is acceptable), the set of candidate divisions may be reduced to the following subset: {9, 8+1, 7+2, 6+3, 6+2+1, 5+4, 5+3+1, 5+2+2, 4+4+1, 4+3+2, 4+2+2+1, 3+3+3, 3+3+2+1, 3+2+2+2, 2+2+2+2+1}.
[0099]
[0107] In block 845, the electronic device removes from the set of candidate partitions one or more candidate partitions with the largest multiplexer size that exceeds the maximum multiplexer size (e.g., specified in block 805). If the multiplexer size is too large, the hardware complexity and power consumption of the partial product stage may become unacceptable. Assuming the maximum multiplexer size of the multiplexer is 16-way (4-bit addressing), the subsets of block 835 may be reduced to the following subsets: {4+4+1, 4+3+2, 4+2+2+1, 3+3+3, 3+3+2+1, 3+2+2+2, 2+2+2+2+1}.
[0100]
[0108] In optional block 855, the electronic device selects, for each largest multiplexer size from the set of candidate partitions, the respective candidate partition corresponding to the smallest number of addends for that largest multiplexer size. For example, the subsets of block 845 may be reduced to the following subsets: {4+4+1, 4+3+2, 3+3+3, 2+2+2+2+1}. In some aspects, the electronic device applies one or more other criteria that further reduce the number of values for a particular largest multiplexer size to one. For example, this may be in response to determining that evaluating the entire subset may be too computationally expensive. As shown, the values of the subset of block 855 {4+4+1, 4+3+2} have the largest multiplexer size with 4-bit addressing. In one embodiment, the value {4+4+1} may be selected if a 4-bit addressing multiplexer corresponds to a standard size (e.g., has less hardware complexity) or is expected to be more efficient than a combination of 3-bit and 2-bit addressing multiplexers. In another embodiment, the value {4+3+2} may be selected if a combination of 3-bit and 2-bit addressing multiplexers is expected to be more efficient than including a 1-bit AND gate multiplexer. Other criteria have also been considered.
[0101]
[0109] In block 865, the electronic device selects a first candidate partition from the set of candidate partitions that has the smallest number of addends. For example, the subset of block 855 may be reduced to the following subset: {4+4+1, 3+3+3, 2+2+2+2+1}. In another embodiment (e.g., when optional block 855 is not performed), the subset of block 845 may be reduced to the following subset: {4+4+1, 4+3+2, 3+3+3}. Beneficially, reducing the set of candidate partitions according to the various blocks of method 800 saves a substantial amount of computational resources and / or power consumption for subsequent evaluation of the candidate partitions. For example, for a 32-bit input coefficient vector Cn The FIR filter with n The FIR filter with corresponds to over 1.7 million candidate splits.
[0102]
[0110] In block 875, the electronic device evaluates each selected candidate partition based on size, transport delay, and / or power consumption. In some aspects, evaluating each selected candidate partition includes implementing the candidate partition in a high-level hardware description language (such as Verilog or VHDL), synthesizing and laying out the circuit, extracting parasitics, and running a final simulation.
[0103]
[0111] For a given clock rate, power consumption and size (e.g., layout area) are determined by the number and size of adders and pipeline registers. The presence of pipeline registers depends on the clock speed and may determine when pipelining is necessary. In some high-speed applications, there may be many pipeline registers per adder. In wide adders, there may be multiple pipeline stages for each adder. Transport delay is determined by the clock speed and is expressed in clock cycles.
[0104]
[0112] Using one example subset (e.g., determined in block 865 above), the division of the subsets can be analyzed as follows: TIFF2025169163000003.tif50170
[0105]
[0113] Based on the results of the evaluation, the electronic device selects one of the candidate partitions as the final topology for the multiplexing stage (at block 885). In some aspects, the electronic device may present each selected candidate partition to a user and receive user input selecting one of the candidate partitions. In other aspects, the electronic device may make the selection automatically according to one or more criteria of the evaluation.
[0106]
[0114] In this case, the partition {3+3+3} may be selected because a smaller-sized MUX requires fewer terms to be expanded in the partial product stage. If the {3+3+3} partition is not possible, the selection of {4+3+2} and {4+4+1} may be based on the fact that it is easier to design a single 16-way (4 address bits) MUX and reuse the design than to design 16-way, 8-way, and 2-way MUXs, respectively. An alternative design for the 16-way MUX is to combine two 8-way and 2-way MUXes, and either partition may have layout size advantages. Furthermore, because not all adder elements are the same, for example, if the number of half and full adder elements is different, design synthesis tools can be used to determine size and power and perform a final comparison. Method 800 ends after completing block 885.
[0107]
[0115] 9 is an example method 900 for creating a coefficient multiplier circuit according to one or more aspects. Method 900 may be used in conjunction with other aspects, for example, to determine a partial product stage of the coefficient multiplier circuit, which in some cases is performed during the design phase of the coefficient multiplier circuit. Furthermore, in some aspects, method 900 may be performed in conjunction with one or more electronic devices, whether implemented as a circuit (e.g., an integrated circuit), computer program code, or the like.
[0108]
[0116] Method 900 begins at optional block 905, where the electronic device receives user input specifying a maximum multiplexer address width. Some examples of maximum multiplexer address width include 1-bit addressing, 2-bit addressing, etc. This maximum multiplexer address width represents a size threshold that is upper bounded by the complexity of partial product calculations.
[0109]
[0117] In block 915, the electronic device initializes a set of computational instructions ("dsn") with a first binary value (e.g., 1). The size of the set of computational instructions is a number of bits equal to the maximum multiplexer address width (e.g., 2 width -1). In some aspects, after completion of method 900, the set of computational instructions includes an addition of an individual vector(s), a pair of vectors, and / or a left shift of an individual vector(s). At block 925, the electronic device initializes a vector of computed binary vectors ("vec") to the empty set.
[0110]
[0118] In block 935, the electronic device sequentially determines the operations necessary to unfold the partial products. In some aspects, these partial products are width Block 935 steps through all vectors (v). All vectors (v) are in binary representation with zero and base. The sequence of values for v represents the partial product coefficients of x. For each vector v, blocks 945, 955, 965, and possibly block 975 are executed to define the computational operations that generate its partial products.
[0111]
[0119] In block 945, the electronic device generates a binary vector representation of the value of v with the maximum multiplexer address width defined in block 905. For example, if width=4 and v represents the number 3, then the value of v is {0,0,1,1}. The electronic device adds the value to the collection of vectors.
[0112]
[0120] In block 955, starting with the first element in the vector of computed binary vectors (vec), the electronic device searches through the vector of computed binary vectors (vec) to find and select a binary value ("vsh") that can be left shifted by k positions to generate the current value of v, where k=1, 2, ... (width-1). The left shift operation removes the most significant (leftmost) k bits from the number in vec and adds the least significant k bits to the number backfilled with zeros.
[0113]
[0121] If the left-shift operation produces a value that matches v, the electronic device stops searching and indicates that the left-shift operation was found. If a matching left-shifted element of vec is not found, the electronic device provides an indication that the left-shift operation was not found. In block 965, when a match is found, the electronic device adds "left-shift vsh by v=k places (1 or more)" to the set of computational instructions to be executed (dsn), and the k-place left-shift operation operating on vsh is added to v=2. k Indicates that x is used to calculate the partial product of x. The method 900 returns to block 935 to proceed to the next value v.
[0114]
[0122] At optional block 975, the electronic device did not find a left-shifted value that can produce a value v derived from the computed binary vector (vec). If vec contains more than two elements, the electronic device searches all possible subsets of vec that have two elements. The electronic device creates a list of all possible pairs of values in vec and tests each pair to see if an exclusive OR (XOR) operation produces the vector v. When the first matching pair, (v1XORv2)=v, is found, it is saved in the set of computation instructions (dsn) as "v=v1+v2," and method 900 returns to block 935 to proceed to the next value v. After all of the partial products have been designed, method 900 ends.
[0115]
[0123] In some aspects, the method 900 may be implemented using the pseudocode provided below: width=4; (* Maximum MUX address width (bits) *) vectors={}; (* vec, List of binary partial product vectors to be developed *) design={}; (* Design operation list *) (* This algorithm applies for width>2. For each vector, v *) For[k=0,k<=2^width-1,k++, v=IntegerDigits[k,2,width];(* Convert to its binary representation *) vectors=AppendTo[vectors,v];(* Append v to the vector list *) (* Search each of the vectors for a left-shift solution to prioritize on left-shift vs. addition *) If[k==0, sol=v;lsFound=True;,(* The left-shift solution for the first vector equaling 0 is simply that vector *) lsFound=False; (* Search for left-shift solutions for non-zero vectors starting with the assertion that it has not been found *) For[i=2,i<=Length[vectors],i++, (* Look at each vector so far processed in increasing sequence *) sol=leftShiftFound[vectors[[i]],v]; (* Return a left shift operation if found *) If[!(sol===Null),lsFound=True;Break[]] (* If a left-shift solution has been found stop searching *) ]; ]; (* Now search for adders if shifts not found *) (* For each vector element, v, look at all previous combinations to determine if a sum exists that produces the vector *) If[lsFound==True, (* Left-shift solution found, append to the list of operations *) AppendTo[design,{ToString[TraditionalForm[Subscript["V",ToString[k]]]]<>" = "<>ToString[sol]}];, (* Left-shift not found. List all possible pairs of vectors hitherto processed *) adderFound=False; subsets=Subsets[Select[vectors,#!=Table[0,width]&],{2}];(* All subsets not containing all 0's*) (* Search through the list of pairs to find the first sum that produces v *) For[i=1,i<= Length[subsets],i++, sol=sumFound[subsets[[i,1]],subsets[[i,2]],v]; If[!(sol===""),adderFound=True;Break[];] (* When found stop looking *) ]; (*Append adder found to the operation list*) If[adderFound, AppendTo[design,{ToString[TraditionalForm[Subscript["V",ToString[k]]]]<>" = "<>ToString[sol]}];, AppendTo[design,{ToString[TraditionalForm[Subscript["V",ToString[k]]]]<>" = "<>ToString[v]}]; ]; ]; ]; design / / Flatten (* Provide the list of operations in the design *) MatrixForm[design]
[0116]
[0124] For example, using pseudocode, the set of computational instructions (dsn) can be calculated as follows: V(0) = {0, 0, 0, 0} V(1) = {0, 0, 0, 1} Shift V(2) = {0, 0, 0, 1} left by 1 place. V(3) = {0, 0, 0, 1}+{0, 0, 1, 0} Shift V(4) = {0, 0, 0, 1} left by 2 places. V(5) = {0, 0, 0, 1}+{0, 1, 0, 0} Shift V(6) = {0, 0, 1, 1} left by 1 place. V(7) = {0, 0, 0, 1}+{0, 1, 1, 0} Shift V(8) = {0, 0, 1, 0} left by 2 places. V(9) = {0, 0, 0, 1}+{0, 1, 0, 0} Shift V(10) = {0, 1, 0, 1} left by 1 place. V(11) = {0, 0, 0, 1}+{1, 0, 1, 0} Shift V(12) = {0, 0, 1, 1} left by 2 places. V(13) = {0, 0, 0, 1}+{1, 1, 0, 0} Shift V(14) = {0, 1, 1, 1} left by 1 place. V(15) = {0, 0, 0, 1}+{1, 1, 1, 0}
[0117]
[0125] Thus, in one aspect, the present disclosure provides a method for creating a coefficient multiplier circuit. The method includes decomposing the length of an input coefficient vector to provide a set of candidate partitions corresponding to one or more multiplexer combinations of the coefficient multiplier circuit. The method further includes eliminating, from the set of candidate partitions, one or more candidate partitions corresponding to at least a threshold number of 1-bit AND gate multiplexers. The method further includes eliminating, from the set of candidate partitions, one or more candidate partitions having a largest multiplexer size greater than a maximum multiplexer size. The method further includes selecting, from the set of candidate partitions, a first candidate partition having the smallest number of addends.
[0118]
[0126] In one aspect, in combination with any example method above or below, the method further includes receiving user input specifying a maximum multiplexer size.
[0119]
[0127] In one aspect, in combination with any example method above or below, the method further includes, for each largest multiplexer from the set of candidate partitions, selecting a respective candidate partition corresponding to the smallest number of addends for that largest multiplexer size, wherein the first candidate partition is one of the selected respective candidate partitions.
[0120]
[0128] In one aspect, in combination with any example method above or below, the method further includes evaluating each selected candidate partition based on one or more of size, transportation delay, and power consumption.
[0121]
[0129] In one aspect, in combination with any example method above or below, the method further includes generating a set of computational instructions corresponding to a maximum address width, wherein generating the set of computational instructions includes, for each computational instruction of the set, performing one or more of: selecting a vector, retrieving a left shift operation on a vector of a previously determined computation result of the set, and retrieving an add operation on two vectors (at least one vector is from a previously determined computation result of the set).
[0122]
[0130] As will be appreciated by one of ordinary skill in the art, aspects described herein may be embodied as a system, method, and / or computer program product. Accordingly, aspects may take the form of entirely hardware aspects, entirely software aspects (including firmware, resident software, microcode, etc.), or aspects combining software and hardware aspects, all of which may be broadly referred to herein as "circuits," "modules," or "systems." Furthermore, aspects described herein may take the form of a computer program product embodied in one or more computer-readable storage medium(s) having computer-readable program code embodied therein.
[0123]
[0131] The program code embodied in the computer readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, fiber optic cable, RF, etc., or any suitable combination thereof.
[0124]
[0132] Computer program code for carrying out operations of aspects of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code may run entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0125]
[0133] Aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to aspects of the present disclosure. It will be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose or special-purpose computer or other programmable data processing device to produce a machine. These instructions, executed via the processor of the computer or other programmable data processing device, thereby create means for performing the function(s) / acts identified in the block(s) of the flowcharts and / or block diagrams.
[0126]
[0134] These computer program instructions may also be stored on a computer-readable medium that may direct a computer, other programmable data processing apparatus, or other device to function in a particular manner. The instructions stored in the computer-readable medium thereby produce an article of manufacture. The instructions include instructions that implement the functions / acts identified in the flowchart and / or block diagram block(s).
[0127]
[0135] Computer program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be executed on the computer, other programmable data processing apparatus, or other device to generate computer-implemented processes, whereby the instructions executed on the computer, other programmable data processing apparatus, or other device provide steps for performing the functions / acts identified in the flowchart and / or block diagram block(s).
[0128]
[0136] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various aspects of the present disclosure. As such, each block in the flowcharts and block diagrams may represent a module, segment, or portion of code, including one or more executable instructions for implementing specific logical function(s). In some alternative implementations, the functions shown in the blocks need not occur in the order depicted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially simultaneously, or the blocks may be executed in reverse or out of order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a special-purpose hardware-based system that performs particular functions or functions, or by a combination of special-purpose hardware and computer instructions.
[0129]
[0137] While the foregoing is directed to aspects of the disclosure, other and further aspects of the disclosure may be devised without departing from the basic scope thereof, the scope of which is defined by the following claims.
Claims
1. A filter circuit (302, 402) comprising: a serially connected arrangement (305, 405) of unit delays (310-1, ..., 310-4) and adders (315-1, ..., 315-4) in an alternating pattern; and a plurality of coefficient multipliers (320-0, 320-1, 320-2, 502, 602), each having a respective output connected to one or more of said adders, each coefficient multiplier receiving a respective input coefficient vector (C 0 , C 1 , C 2 a plurality of coefficient multipliers, including a multiplexing stage (510, 606) including one or more multiplexers (530-1, ..., 530-4, 625-1, ..., 625-3) addressed using one or more bits of A first coefficient multiplier (320-0) of the plurality of coefficient multipliers provides an input data vector (x n ) multiple integer partial products (0, x n , ..., 15x n 1. A filter circuit comprising a partial product stage (505, 605, 702) configured to provide a
2. The one or more multiplexers include a first multiplexer (530-1, 530-3, 625-1, 625-2) and a second multiplexer (530-2, 530-4, 625-3), and the multiplexing step includes: a left shifter (535-1, 535-2, 635, 645) connected to the output of the first multiplexer, wherein the one or more bits addressing the first multiplexer are more significant than the one or more bits addressing the second multiplexer; and 2. The filter circuit of claim 1, further comprising an adder (540-1, 540-2, 630, 640) connected to the output of the second multiplexer, the adder being connected to the output of the left shifter.
3. 3. The filter circuit of claim 2, wherein the multiplexing stage further comprises a unit delay (550-1, 550-2) connected to the output of the adder.
4. The partial product step comprises: one or more left shifters (515-1, 610-1, 610-2, 615); and 2. The filter circuit of claim 1, comprising one or more adders (520-1, 620-1, 620-2, 620-3) configured to receive the input data vector and an output of one of the one or more left shifters.
5. The filter circuit of claim 1 , wherein the input coefficient vectors provided to the plurality of coefficient multipliers are symmetric.
6. The one or more multiplexers of the first coefficient multiplier each have N select lines, and the partial product stages are N 2. The filter circuit of claim 1, configured to provide integer partial products up to -1.
7. A coefficient multiplier circuit (320-0, 320-1, 320-2, 502, 602) for a filter, comprising: The input data vector (x n ) multiple integer partial products (0, x n , ..., 15x n a partial product stage (505, 605, 702) configured to provide a multiplexing step (510, 606), said multiplexing step comprising: one or more multiplexers (530-1, ..., 530-4, 625-1, ..., 625-3) configured to receive the plurality of integer partial products of the input data vector, and an input coefficient vector (C 0 , C 1 , C 2 one or more multiplexers that are addressed using one or more bits of a coefficient multiplier circuit coupled to the one or more multiplexers and including an output coupled to one or more adders (315-1, . . . , 315-4) of the filter;
8. The coefficient multiplier circuit of claim 7 , wherein the filter is one of a finite impulse response (FIR) filter, an infinite impulse response (IIR) filter, and an autoregressive moving average (ARMA) filter.
9. The one or more multiplexers include a first multiplexer (530-1, 530-3, 625-1, 625-2) and a second multiplexer (530-2, 530-4, 625-3), and the multiplexing step includes: a left shifter (535-1, 535-2, 635, 645) connected to the output of the first multiplexer, wherein the one or more bits addressing the first multiplexer are more significant bits of the input coefficient vector than the one or more bits addressing the second multiplexer; and 8. The coefficient multiplier circuit of claim 7, further comprising an adder (540-1, 540-2, 630, 640) connected to the output of the second multiplexer, the adder being connected to the output of the left shifter.
10. The multiplexing step comprises: The coefficient multiplier circuit of claim 9 further comprising a unit delay (550-1, 550-2) connected to the output of the adder.
11. The partial product step comprises: one or more left shifters (515-1, 610-1, 610-2, 615); and 8. The coefficient multiplier circuit of claim 7, comprising one or more adders (520-1, 620-1, 620-2, 620-3) configured to receive the input data vector and an output of one of the one or more left shifters.
12. The one or more multiplexers each have N select lines, and the partial product stages are N 8. The coefficient multiplier circuit of claim 7, configured to provide integer partial products up to -1.
13. Further comprising a second multiplexing step, the second multiplexing step comprising: one or more second multiplexers configured to receive the plurality of integer partial products of the input data vector, the one or more second multiplexers being addressed using one or more bits of a second input coefficient vector; and 8. The coefficient multiplier circuit of claim 7, including a second output connected to the one or more second multiplexers and connected to one or more other adders of the filter.
14. A system (100), comprising: a refrigeration system (105) configured to maintain a refrigerated environment (140); a plurality of qubits (112) disposed within the refrigerated environment; and a control circuit (205) connected to the plurality of qubits, the control circuit including, for each qubit of the plurality of qubits, a respective waveform synthesizer (210) configured to drive the qubit, the waveform synthesizer including an interpolation filter circuit (215), the interpolation filter circuit including: a serially connected arrangement (305, 405) of unit delays (310-1, ..., 310-4) and adders (315-1, ..., 315-4) in an alternating pattern; and a plurality of coefficient multipliers (320-0, 320-1, 320-2, 502, 602), each having a respective output connected to one or more of said adders, each coefficient multiplier receiving a respective input coefficient vector (C 0 , C 1 , C 2 a plurality of coefficient multipliers, including a multiplexing stage (510, 606) including one or more multiplexers (530-1, ..., 530-4, 625-1, ..., 625-3) addressed using one or more bits of A first coefficient multiplier (320-0) of the plurality of coefficient multipliers provides an input data vector (x n ) multiple integer partial products (0, x n , ..., 15x n ) a partial product stage (505, 605, 702) configured to provide
15. 15. The system of claim 14, wherein the refrigeration system includes a dilution refrigerator defining a plurality of temperature steps (130-1, 130-2, ..., 130-5), the plurality of quantum bits being disposed at a lowest temperature step (130-5) of the plurality of temperature steps, and the control circuitry being disposed at another temperature step (130-1) of the plurality of temperature steps.
16. The one or more multiplexers include a first multiplexer (530-1, 530-3, 625-1, 625-2) and a second multiplexer (530-2, 530-4, 625-3), and the multiplexing step includes: a left shifter (535-1, 535-2, 635, 645) connected to the output of the first multiplexer, wherein the one or more bits addressing the first multiplexer are more significant than the one or more bits addressing the second multiplexer; and 15. The system of claim 14, further comprising an adder (540-1, 540-2, 630, 640) connected to the output of the second multiplexer, the adder being connected to the output of the left shifter.
17. The multiplexing step comprises: The system of claim 16 further comprising a unit delay (550-1, 550-2) connected to the output of the adder.
18. The partial product step comprises: one or more left shifters (515-1, 610-1, 610-2, 615); and 15. The system of claim 14, comprising one or more adders (520-1, 620-1, 620-2, 620-3) configured to receive the input data vector and an output of one of the one or more left shifters.
19. The system of claim 14 , wherein the input coefficient vectors provided to the plurality of coefficient multipliers are symmetric.
20. The one or more multiplexers of the first coefficient multiplier each have N select lines, and the partial product stages are N 15. The system of claim 14, configured to provide integer partial products up to -1.