Machine learning for synchronizing multiple FPGA ports in quantum system
Patent Information
- Application Number
- JP2023008596
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-01-24
- Filing Date
- 2023-01-24
- Publication Date
- 2025-08-07
- Estimated Expiration
- 2043-01-24
AI Technical Summary
Conventional methods for synchronizing multiple FPGA ports in quantum systems are inefficient and require repeated calibrations, leading to potential delays and deviations that compromise the reliability of quantum computers.
A machine learning approach is employed to synchronize FPGA ports by using delay lines and test signals to determine optimal delays for each port, eliminating the need for external storage and iterative calibrations.
This method ensures precise synchronization of FPGA ports without the need for external storage or repeated calibrations, enhancing the reliability and efficiency of quantum computers.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Background technology]
[0001]
[0001] The limitations and shortcomings of the conventional uses of multiple FPGA ports will become apparent to those skilled in the art by comparing certain aspects of the present method and system presented in the remainder of this disclosure with conventional approaches, with reference to the drawings. Summary of the Invention [Means for solving the problem]
[0002]
[0002] As more fully set forth in the claims, there is provided a method and system for synchronizing multiple FPGA ports in a quantum system, substantially as shown in and / or described in conjunction with at least one of the figures. [Brief explanation of the drawings]
[0003] [Figure 1] FIG. 1 illustrates an example quantum system with multiple synchronized FPGA ports, in accordance with various example implementations of the present disclosure. [Figure 2] FIG. 1 illustrates an example system for training a quantum system to synchronize multiple FPGA ports, in accordance with various example implementations of the present disclosure. [Figure 3] 1 is a flowchart of an example method for synchronizing multiple FPGA ports, in accordance with various example implementations of the present disclosure. [Figure 4] 10 is a graph of example phase values measured over a range of delays in accordance with various example implementations of the present disclosure. [Figure 5] 10A-10C illustrate example tap / delay estimation functions as linear fits to ideal phase values for different ports with different setup and hold times, according to various example implementations of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0004]
[0008] Traditional computers operate by storing information in the form of binary digits ("bits") and processing these bits through binary logic gates. At any time, each bit can have only one of two discrete values: 0 (i.e., "off") and 1 (i.e., "on"). The logical operations performed by binary logic gates are defined by Boolean algebra, and circuit operation is governed by classical physics. In modern traditional systems, the circuits for storing bits and realizing logical operations are typically made from electrical wires capable of carrying two different voltages that represent the bits 0 and 1, and transistor-based logic gates that perform the Boolean logic operations.
[0005]
[0009] Logical operations in traditional computers are performed on fixed states. For example, at time 0, a bit is in a first state, at time 1 a logical operation is applied to the bit, and at time 2 the bit is in a second state, as determined by the state at time 0 and the logical operation. The state of the bit is typically represented by a voltage (e.g., 1V for a "1"). dc , or 0V for "0" dc ) A logical operation typically comprises one or more transistors.
[0006]
[0010] Clearly, traditional computers using single bits and single logic gates have limited effectiveness, which is why modern traditional computers, even with modest computing power, contain billions of bits and transistors. That is, traditional computers that can solve increasingly complex problems necessarily require an increasingly large number of bits and transistors and / or an increasingly large amount of time to execute their algorithms. Nevertheless, there are some problems that require an infeasibly large number of transistors and / or an infeasibly large amount of time to arrive at a solution. Such problems are called infeasible.
[0007]
[0011] Quantum computers operate by storing information in the form of quantum bits ("qubits") and processing these qubits through quantum gates. Unlike bits, which can only be in one state (0 or 1) at any time, qubits can be in a superposition of two states simultaneously. More precisely, a qubit is a system whose states live in a two-dimensional Hilbert space and are therefore written as a linear combination α|0〉+β|1〉, where |0〉 and |1〉 are two basis states and α and β are functions of |α| 2 +|β| 2 = 1, which is a complex number, usually called a probability amplitude. Using this notation, when a qubit is measured, it is measured with probability |α| 2 becomes 0, and the probability |β| 2 The basis states |0〉 and |1〉 are two-dimensional basis vectors
number
number
number
[0008]
[0012] Unlike traditional bits, qubits cannot be stored as a single voltage value on a wire. Instead, qubits are physically realized using a two-level quantum mechanical system. For example, at time 0, a qubit is
number
number
[0009]
[0013] 1 shows an example quantum system with synchronized multiple FPGA ports according to various example implementations of this disclosure. The quantum system includes a quantum programming subsystem (QPS) 101, a quantum controller (QC) 103, and a quantum processor 107.
[0010]
[0014] QPS101 is capable of generating a quantum algorithm description, including instructions that QC103 can execute to configure QC103 and execute the quantum algorithm (i.e., generate the necessary outbound quantum control pulses) with little or no human intervention during runtime. In an example implementation, QPS101 is a personal computer with a processor, memory, and other associated circuitry (e.g., an x86 or x64 chipset). QPS101 compiles the high-level quantum algorithm description into a machine-language version of the quantum algorithm description (i.e., a series of binary vectors representing instructions that QC103 can directly interpret and execute).
[0011]
[0015] The QPS101 may be coupled to the QC103 via an interconnection that may utilize, for example, a Universal Serial Bus (USB), a Peripheral Component Interconnect (PCIe) bus, wired or wireless Ethernet, or any other suitable communication protocol.
[0012]
[0016] The QC103 comprises circuitry operable to load a machine-language quantum algorithm description from the QPS101 via the interconnect. Execution of the machine language by the QC103 causes the QC103 to generate the necessary outbound quantum control pulses corresponding to the desired operation to be performed on the quantum processor 107 (e.g., to a qubit to manipulate the qubit's state, or to a readout resonator to read the qubit's state, etc.). The machine language also causes the QC103 to perform an analysis of the input signals. The analysis results may be used to determine the state of the qubit or the quantum register (quantum measurement). Depending on the quantum algorithm to be implemented, the outbound pulses for executing the algorithm can be predetermined at design time and / or may have to be determined during runtime. Runtime determination of the pulses may involve performing traditional computations and processing in the QC103 during the runtime of the algorithm (e.g., run-time analysis of inbound pulses received from the quantum processor).
[0013]
[0017] The QC 103 generates a precise series of external signals, typically pulses of electromagnetic waves and pulses of baseband voltage, to perform the desired logical operations (and thus execute the desired quantum algorithm).
[0014]
[0018] During run-time and / or upon completion of the quantum algorithm performed by QC103, QC103 can output data / results to QPS101. In example implementations, these results may be used to generate new quantum algorithm descriptions for subsequent runs of the quantum algorithm and / or to update the quantum algorithm descriptions during run-time. Additionally, QC103 can output raw or processed inbound pulses received from quantum processor 107 representing qubit state estimates, or quantum program control flow and branching information, as well as metadata representing internal variable calculations during program execution.
[0015]
[0019] QC103 comprises a plurality of pulse processors, which may be implemented with field programmable gate arrays (FPGAs), application specific integrated circuits, or the like, operable to control analog outbound pulses that drive quantum elements (e.g., one or more qubits and / or resonators) or to enable interactions between quantum elements and digital outbound pulses that can control ancillary equipment required for program execution (e.g., gating the analog outbound pulses or controlling external devices such as photon detectors).
[0016]
[0020] Quantum algorithms are implemented in quantum processor 107 when one or more qubits interact with quantum control pulses. These quantum control pulses are electromagnetic RF signals or pulses that are digitally generated at baseband in QC 103, converted to analog waveforms via multiple DACs 109-0, 109-1, 109-2, and 109-3, and upconverted by RF circuitry 105. The desired signal may be generated according to a known set of instructions with various operations, such as arithmetic or logical calculations, communications with various components, and traditional control flow operations (jumps, branches, etc.). An application layer (APP) in QC 103 controls the physical layer (PHY) to digitally generate (and further modify) samples of this analog waveform. Inbound pulses are further received by QC 103 from quantum processor 107 via RF circuitry 105 and multiple ADCs 111-0 and 111-1.
[0017]
[0021] Qubits can have lifetimes in the range of hundreds of microseconds, resulting in very low program execution runtimes. Furthermore, in a data center where a quantum computer is acting as a co-accelerator for a particular computation, there may be thousands of programs queued to use a designated quantum processor 107.
[0018]
[0022] Periodic recalibration is necessary as process, voltage, and temperature (PVT) changes occur. Therefore, a fast, robust, and independent approach to recalibration could facilitate much better use of quantum computers while minimizing dead time between programs.
[0019]
[0023] FIG. 2 shows an example system for training a quantum system to synchronize multiple FPGA ports, according to various example implementations of this disclosure.
[0020]
[0024] The quantum controller 103 of FIG. 1 may include a PCB 201-0 and an FPGA 203-0. The FPGA 203-0 may have numerous ports 207-0 and 207-1 that need to be synchronized with each other. Furthermore, the PCB 201-0 and the FPGA 203-0 may have design differences 201-1 and 203-1, respectively, such that transmissions through ports 207-0 and 207-1 of the FPGA 203-0 may not be aligned with transmissions through ports 207-0 and 207-1 of the FPGA 203-1. Synchronization is necessary so that outputs from the FPGA 203-0 or 203-1 reach all DACs 109-0 and 109-1 simultaneously and independently, without bias, despite design differences between similar system components. Any delay or misalignment in one of the signals being output from the FPGA 203-0 or 203-1 could dramatically impair the reliability of the quantum computer.
[0021]
[0025] The hardware paths from FPGA 203-0 to ports 207-0 and 207-1 may vary with FPGA 203-1 variants (e.g., FPGA 203-1 from a different batch but a revision of the same PCB 201-0). Different characteristics of FPGA 203-1 may affect the time it takes for signals to reach DACs 109-0 and 109-1, potentially destroying the calibrated synchronization for different PCBs 201. To train this machine learning model, several FPGA designs 203-0 and 203-1, each with a different layout, and several quantum control units with different PCBs 201-0 and 201-1, are used for training.
[0022]
[0026] Synchronizing ports 207-0 and 207-1 of FPGAs 203-0 and / or 203-1 incorporates delay lines 113-0 and 113-1 for each port 207-0 and 207-1, allowing the signals leaving FPGAs 203-0 and / or 203-1 to be programmatically and digitally "shifted" in regular and discrete steps so that all signals are aligned at their destinations as required by quantum control applications.
[0023]
[0027] A machine learning approach is disclosed for synchronizing all quantum FPGA ports 207-0 and 207-1 without the need to save, load, and maintain previously acquired data for each quantum control unit (i.e., without using external storage). The machine learning approach also eliminates the need for calibration using external input / output devices, as well as the need for lengthy and repeated calibrations of the quantum control platform.
[0024]
[0028] To train the machine learning model, information from different PCBs 201-0 and / or 201-1 and different FPGA logic designs 203-0 and / or 203-1 is collected for each port 207-0 and 207-1. Training is only required once.
[0025]
[0029] Test signals may be generated by generators 205-0 and 205-1. The test signals may be sinusoidal signals or any other signals with deterministic phase. For each PCB 201-0 or 201-1, a possible delay length (0 to N) is programmed into delay lines 113-0 and 113-1 for each FPGA design 203-0 or 203-1. The phase of the test signals is measured at DACs 109-0 and 109-1. This provides information on which tap / delay is needed for each port 207-0 or 207-1. The test signals are synchronized when they have the same phase at DACs 109-0 and 109-1. A formula for determining the tap / delay may be derived from the measured phase values using linear regression as a function of the port's setup and hold time. Alternatively, a nonlinear equation for determining the tap / delay may be derived to take into account the nonlinear delay line.
[0026]
[0030] FIG. 3 shows a flowchart of an example method for synchronizing multiple FPGA ports in accordance with various example implementations of this disclosure.
[0027]
[0031] Each FPGA logic design is characterized by a setup-and-hold (S / H) time, which also serves as part of the input to the training phase. At 301, the S / H time is determined for each port of the FPGA. The FPGA ports may be asynchronous.
[0028]
[0032] A test signal is generated. At 303, the test signal is sent to a destination through each of the FPGA ports, which may be a DAC. Each FPGA port is coupled to a multi-tap delay line. Each of the multiple multi-tap delay lines is initiated by setting a tap (i.e., a selectable delay).
[0029]
[0033] The phase of the test signal as received at the destination from every port is measured at 305. For example, if each of eight ports sends a sinusoidal test signal to each of eight DACs, the test signals received at the DACs are processed to determine phase values.
[0030]
[0034] It is determined whether all taps have been used at 307. If more taps are available, the next tap is selected at 309, the test signal is retransmitted at 303, and the phase value of the test signal as received at the destination from every port is measured at 305.
[0031]
[0035] Once each tap / delay has been chosen and the data collection phase has taken place, the application is operable to select an ideal tap / delay from the plurality of phase values for each of the plurality of ports at 311. The ideal tap / delay for every port will correspond to the same phase.
[0032]
[0036] It is determined whether more PCBs are available for training at 313. If more PCBs are available, new PCBs are used at 315, the taps / delays are reinitialized for each delay line, the test signal is retransmitted at 303, and the phase values of the test signal as received at the destination from every port are measured at 305.
[0033]
[0037] It is determined whether more FPGAs are available for training at 317. If more FPGAs are available, new FPGAs are used at 319, the taps / delays are reinitialized for each delay line, the test signal is retransmitted at 303, and the phase values of the test signal as received at the destination from every port are measured at 305.
[0034]
[0038] FIG. 4 shows example graphs of phase values measured over a range of delays in accordance with various example implementations of this disclosure.
[0035]
[0039] The horizontal axis is the delay line value (i.e., the amount in time units the test signal is delayed). The vertical axis is the measured phase. Such a graph can be generated for each port.
[0036]
[0040] As shown, the tap scale in each delay line may range from 0 to 63. This range is an example, as any range may be used. Ideally, each tap would correspond to a precise delay; for example, a range from 0 to 4 nsec with a resolution of 62.5 psec. However, precise delays are sometimes not possible, and a linear relationship between time and taps is not a requirement of this method.
[0037]
[0041] As shown, the test signal is a sinusoid, but any test signal with a deterministic phase may be used. The ideal tap / delay is selected as 18 because 18 corresponds to the center of the period of constant phase. Therefore, the selected ideal tap may be determined according to the midpoint between two phase changes for a particular port of the multiple asynchronous ports. In other embodiments, the test signal may include a pulse. For a test signal pulse, the selected ideal tap may be based on a phase transition (i.e., from on to off, or vice versa).
[0038]
[0042] Returning now to FIG. 3, at 321, for each of the multiple ports, a tap estimation function is generated as a function of the selected / ideal tap and the setup and hold time.
[0039]
[0043] FIG. 5 shows example tap / delay estimation functions as linear fits to ideal phase values (as discussed with respect to FIG. 4) for different ports with different setup and hold times, according to various example implementations of the present disclosure.
[0040]
[0044] Using linear regression on the training data, the coefficients a and b may be determined so that the following function fits the collected data: Delay = a(S / H) + b where S / H is the setup and hold time of the port (which can vary for each FPGA logic design), and Delay is the delay required for the port such that all ports on a particular PCB are synchronized.
[0041]
[0045] This machine learning approach can determine the optimal delay for different logic designs, different PVTs, and different batches of FPGA chips. This optimal delay may be achieved without repeated calibration using external wiring and persistent storage.
[0042]
[0046] The method and / or system may be implemented in hardware, software, or a combination of hardware and software. The method and / or system may be implemented in a centralized manner in at least one computing system, or in a distributed manner with different elements spread across several interconnected computing systems. Any kind of computing system or other apparatus adapted for performing the methods described herein is suitable. A typical implementation may include one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), and / or one or more processors (e.g., x86, x64, ARM, PIC, and / or any other suitable processor architecture) and associated support circuitry (e.g., storage, DRAM, flash, bus interface circuits, etc.). Each individual ASIC, FPGA, processor, or other circuit may be referred to as a "chip," and multiple such circuits may be referred to as a "chipset." Another implementation may comprise a non-transitory machine-readable (e.g., computer-readable) medium (e.g., flash drive, optical disk, magnetic storage disk, or the like) having stored thereon one or more lines of code that, when executed by the machine, causes the machine to perform the processes as described in this disclosure. Another implementation may comprise a non-transitory machine-readable (e.g., computer-readable) medium (e.g., flash drive, optical disk, magnetic storage disk, or the like) having stored thereon one or more lines of code that, when executed by the machine, configures the machine to operate as a system described in this disclosure (e.g., by loading software and / or firmware into its circuitry).
[0043]
[0047] As used herein, the terms "circuit" and "circuitry" refer to physical electronic components (i.e., hardware) as well as any software and / or firmware ("code") that may comprise, execute on, and / or otherwise be associated with hardware. As used herein, for example, a particular processor and memory may comprise a first "circuit" when executing a first one or more lines of code, and a second "circuit" when executing a second one or more lines of code. As used herein, "and / or" means any one or more of the items in the list joined by "and / or." As an example, "x and / or y" means any element in the three-element set {(x), (y), (x,y)}. As another example, "x, y, and / or z" means any element in the seven-element set {(x), (y), (z), (x,y), (x,z), (y,z), (x,y,z)}. As used herein, the term "exemplary" means serving as a non-limiting example, instance, or illustration. As used herein, the terms "e.g., " and "for example" emphasize a list of one or more non-limiting examples, instances, or illustrations. As used herein, a circuit device is "operable" to perform a function whenever the circuit device comprises the necessary hardware and code (if necessary) to perform the function, regardless of whether performance of the function is disabled or enabled (e.g., by a user-configurable setting, factory maintenance, etc.). As used herein, the term "based on" means "based at least in part on." For example, "x based on y" means that "x" is based at least in part on "y" (and may also be based on z, for example).
[0044]
[0048] While the present methods and / or systems have been described in terms of certain implementations, it will be understood by those skilled in the art that various modifications may be made and equivalents may be substituted without departing from the scope of the present methods and / or systems. In addition, many modifications may be made to adapt a particular situation or material to the teachings of the present disclosure without departing from the scope of the present disclosure. Therefore, it is intended that the present methods and / or systems not be limited to the particular implementations disclosed, and that the present methods and / or systems include all implementations falling within the scope of the appended claims.
Claims
1. setting one tap of a plurality of taps for each of a plurality of multi-tap delay lines; transmitting a test signal to a destination via each of a plurality of asynchronous ports, each of the plurality of asynchronous ports operatively coupled to one multi-tap delay line of the plurality of multi-tap delay lines; measuring the phase of the test signal at the destination corresponding to each of the plurality of asynchronous ports; repeating the phase measurement after setting each tap of the plurality of taps; selecting a tap for each of the plurality of asynchronous ports in response to the phase measurements for each of the plurality of asynchronous ports; generating a tap estimation function responsive to the selected taps for the plurality of asynchronous ports; A method comprising:
2. 2. The method of claim 1, wherein the tap estimation function is generated as a function of a setup and hold time for each of the plurality of asynchronous ports.
3. 2. The method of claim 1, wherein a field programmable gate array (FPGA) comprises the plurality of asynchronous ports.
4. The method of claim 3 , wherein the FPGA comprises the plurality of multi-tap delay lines.
5. The method of claim 3 , wherein the tap estimation function is generated in response to phase measurements from multiple FPGAs.
6. The method of claim 1 , wherein the destination comprises one or more digital-to-analog converters (DACs).
7. The method of claim 1 , wherein the test signal is a sine wave.
8. The method of claim 1 , wherein the test signal comprises a pulse.
9. 2. The method of claim 1, wherein a selected tap for a particular port of the plurality of asynchronous ports corresponds to a period of a constant phase.
10. 2. The method of claim 1, wherein the selected tap for a particular port of the plurality of asynchronous ports is determined as a function of one or more phase changes.
11. a signal generator operable to generate a test signal; a plurality of multi-tap delay lines operable to receive the test signals, each multi-tap delay line operable to output a delayed test signal corresponding to one of the plurality of taps; a plurality of asynchronous ports, each asynchronous port operable to transmit the delayed test signal to a destination; 1. An application for generating a tap estimation function, comprising: for each of the plurality of asynchronous ports, the application is operable to measure a plurality of phase values, the plurality of phase values corresponding to the plurality of taps; for each of the plurality of asynchronous ports, the application is operable to select a tap from the plurality of phase values; the tap estimation function is generated in response to the selected tap for each of the plurality of asynchronous ports. Applications and A system comprising:
12. 12. The system of claim 11, wherein the tap estimation function is generated as a function of a setup and hold time for each of the plurality of asynchronous ports.
13. 12. The system of claim 11, wherein a field programmable gate array (FPGA) comprises the plurality of asynchronous ports.
14. The system of claim 13 , wherein the FPGA comprises the plurality of multi-tap delay lines.
15. The system of claim 13 , wherein the tap estimation function is generated as a function of phase measurements from multiple FPGAs.
16. The system of claim 11 , wherein the destination comprises one or more digital-to-analog converters (DACs).
17. The system of claim 11 , wherein the test signal is a sine wave.
18. The system of claim 11 , wherein the test signal comprises a pulse.
19. 12. The system of claim 11, wherein a selected tap for a particular port of the plurality of asynchronous ports corresponds to a period of constant phase.
20. 12. The system of claim 11, wherein the selected tap for a particular port of the plurality of asynchronous ports is determined as a function of one or more phase changes.