Quantum error correction hardware decoder and chip

The quantum error correction hardware decoder addresses inefficiencies in neural network-based decoding by providing a scalable and flexible hardware architecture for quantum error correction, reducing delay and enhancing performance.

JP2025536342APending Publication Date: 2025-11-05TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025522634
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-04-27
Filing Date
2023-11-21
Publication Date
2025-11-05

AI Technical Summary

Technical Problem

Current decoding schemes for quantum error correction based on neural network models lack effective hardware implementations, leading to inefficiencies in decoding performance, delay, scalability, and flexibility.

Method used

A quantum error correction hardware decoder and chip architecture that includes an instruction storage device, control unit, and neural network processing units, enabling efficient decoding of error syndrome information using a neural network model, with programmable scalability and flexibility.

Benefits of technology

The hardware decoder significantly reduces decoding delay, enhances scalability, and improves flexibility by allowing configuration adjustments for various noise levels and error-correcting code formats, ensuring efficient real-time quantum error correction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536342000001_ABST
    Figure 2025536342000001_ABST
Patent Text Reader

Abstract

A quantum error correction hardware decoder and chip, relating to the fields of artificial intelligence and quantum technology, includes: an instruction storage device configured to store computer instructions; a control unit configured to read the computer instructions from the instruction storage device and control at least one neural network processing unit based on the computer instructions; the at least one neural network processing unit configured, in response to control by the control unit, to decode error syndrome information of a quantum circuit based on a neural network model to obtain an output result of the neural network model, where the error syndrome information indicates an error syndrome obtained by performing error measurements on the quantum circuit; and an error batch search unit configured to determine error information based on the output result, where the error information indicates a quantum bit in the quantum circuit where an error has occurred and an error type corresponding to the error. The programmable hardware architecture provided in this application can ensure decoding performance while sufficiently reducing decoding delay and improving scalability and flexibility.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to a Chinese patent application filed with the China Patent Office on April 27, 2023, bearing application number 202310479401.9 and entitled "Quantum error correction hardware decoder and chip," the entire contents of which are incorporated herein by reference. FIELD OF THE INVENTION The present application relates to the fields of artificial intelligence and quantum technology, and more particularly to quantum error correction hardware decoders and chips. [Background technology]

[0002] All operational processes in actual quantum computing, including quantum gates and quantum measurements, are accompanied by noise. This means that the circuits used for quantum error correction themselves are noisy. Fault-tolerant quantum error correction refers to the use of a noisy error correction circuit through clever design of the error correction circuit. Fault-tolerant quantum error correction can also achieve the goal of correcting errors and preventing them from spreading over time.

[0003] In fault-tolerant quantum error correction, a syndrome measurement is performed on a quantum circuit to obtain corresponding error syndrome information, and then the error syndrome information is decoded to determine the quantum bit in which an error occurs in the quantum circuit and the error type corresponding to the error. Related art provides a method for decoding error syndrome information based on a neural network model. The error syndrome information of the quantum circuit is input into a neural network model, and the neural network model decodes the error syndrome information to obtain an output result of the neural network model. Then, based on the output result, the quantum bit in which an error occurs in the quantum circuit and the error type corresponding to the error can be further determined.

[0004] Currently, the decoding scheme based on the neural network model requires further research into its hardware implementation. Summary of the Invention [Problem to be solved by the invention]

[0005] The embodiments of the present application provide a quantum error correction hardware decoder and chip. The technical solutions are as follows: [Means for solving the problem]

[0006] According to one aspect of an embodiment of the present application, there is provided a quantum error correction hardware decoder, the quantum error correction hardware decoder comprising: an instruction storage device; a control unit; at least one neural network processing unit; and an error batch search unit; the instruction storage device is configured to store computer instructions; the control unit is configured to read the computer instructions from the instruction storage device and control the at least one neural network processing unit based on the computer instructions; the at least one neural network processing unit is configured to, in response to control by the control unit, decode error syndrome information of a quantum circuit based on a neural network model to obtain an output result of the neural network model, the error syndrome information referring to an error syndrome obtained by performing an error measurement on the quantum circuit; The error batch search unit is configured to determine error information based on the output result, and the error information indicates a quantum bit in which an error has occurred in the quantum circuit and an error type corresponding to the error.

[0007] According to one aspect of the embodiment of the present application, there is provided a chip on which the above quantum error correction hardware decoder is arranged. [Effects of the Invention]

[0008] The technical solutions provided in the embodiments of the present application include at least the following beneficial effects:

[0009] A hardware implementation architecture for a neural network-based quantum error correction decoding algorithm is provided, where the hardware architecture is programmable and computer instructions for implementing the decoding algorithm can be pre-stored in an instruction storage device. In a real-time decoding process, a control unit reads the computer instructions and controls a neural network processing unit based on the computer instructions to decode error syndrome information of a quantum circuit based on a neural network model, ultimately obtaining error information. The hardware decoding architecture can efficiently implement the quantum error correction decoding algorithm based on a neural network model, thereby ensuring decoding performance and significantly reducing decoding delay. Furthermore, in the hardware decoding architecture, the number of neural network processing units and the number of arithmetic units included in the neural network processing unit can be designed and expanded according to actual needs, thereby providing better scalability and flexibility. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a schematic diagram of a surface of revolution code shown in one embodiment of the present application. [Figure 2] FIG. 1 is a schematic diagram of an application scenario of a solution according to one embodiment of the present application. [Figure 3] FIG. 3 is a schematic diagram of an error correction decoding process according to the application scenario of the solution shown in FIG. 2; [Figure 4] FIG. 1 is a schematic diagram of a quantum error correction hardware decoder architecture according to one embodiment of the present application. [Figure 5] 1 is a schematic diagram of a quantum error correction hardware decoder according to one embodiment of the present application; [Figure 6] FIG. 2 is a schematic diagram of a prefetch mechanism in an adder tree according to one embodiment of the present application; [Figure 7] 1 is a schematic diagram of a multi-core architecture according to one embodiment of the present application. [Figure 8] 1 is a schematic diagram of a programmable decoder and resource consumption of a specific network architecture according to one embodiment of the present application; DETAILED DESCRIPTION OF THE INVENTION

[0011] In order to more clearly describe the objectives, technical solutions and advantages of the present application, the following describes the embodiments of the present application in more detail with reference to the drawings.

[0012] Before describing the examples of this application, some terms used in this application will first be explained.

[0013] 1. Quantum Computation (QC): A method that utilizes the properties of superposition and entanglement of quantum states to perform specific computational tasks at high speed.

[0014] 2. Quantum Error Correction (QEC): This is a method of encoding a quantum state by mapping it to one subspace in the Hilbert space of a many-body quantum system. Quantum noise transfers the encoded quantum state to another subspace. By continuously observing the quantum state's location space (syndrome extraction), quantum noise can be evaluated and corrected without interfering with the encoded quantum state, thereby preventing the encoded quantum state from being interfered with by quantum noise. Specifically, a [[n,k,d]] quantum error correction code represents the encoding of k logical qubits with n physical qubits, and any error occurring in any single qubit can be corrected.

number

[0015] 3. Data quantum state: The quantum state of the data qubit for storing quantum information during quantum computing.

[0016] 4. Stabilizer generator: Also called parity check operator. The occurrence of quantum noise (error) causes some changes in the eigenvalues ​​of the stabilizer generator, and therefore quantum error correction can be performed based on this information.

[0017] 5. Stabilizer group: A stabilizer group is a group generated by stabilizer generators. For example, an abelian group generated by stabilizer generators is called a stabilizer generating group. If there are k stabilizer generators, the stabilizer group has 2 k It contains elements and is an Abelian group.

[0018] 6. Error syndrome: When there are no errors, the eigenvalue of the stabilizer generator is 0. When quantum noise occurs, the eigenvalue of the stabilizer generator (parity check operator) of some error-correcting codes that anti-commutate with errors becomes 1. A bit string consisting of these syndrome bits of 0 and 1 is called an error syndrome.

[0019] 7. Syndrome measurement: This is the measurement required to collect the error syndrome. To avoid destroying the information on the data qubits, the syndrome measurement is performed on the auxiliary qubits.

[0020] 8. Number of syndrome measurements: Because the measurements themselves contain noise, to make the measurement process fault-tolerant, it is necessary to repeat the measurements multiple times and use all the collected error syndromes for decoding. The number of syndrome measurements can be represented as T.

[0021] 9. Topological quantum error-correcting code: This is a special type of quantum error-correcting code. The quantum bits of this type of error-correcting code are distributed in a lattice-like arrangement of dimensions greater than two. The lattice constitutes a discrete structure of a higher-dimensional manifold. In this case, the stabilizer generator of the error-correcting code is defined on a limited number of geometrically adjacent quantum bits, so it is geometrically local and physically easy to measure. The quantum bits on which the logical operators of this type of error-correcting code act constitute a kind of topologically nontrivial geometric object on the lattice-like arrangement manifold.

[0022] 10. Surface Code: A surface code is a type of topological quantum error-correcting code defined on a two-dimensional manifold. Its stabilizer generators are typically supported by four qubits (two qubits at the boundary), and the logical operators are nontrivial chains that cross the array in a strip-like fashion. The specific two-dimensional structure of a surface code (7x7, containing 49 data qubits and 48 auxiliary qubits, for a total of 97 physical qubits, capable of correcting any error occurring in two qubits) is shown in Figure 1. The black circles 11 represent data qubits used in quantum computation, and the crosses 12 represent auxiliary qubits. The auxiliary qubits are initialized to the |0〉 or |+〉 state. The hatched and white blocks represent two types of stabilizer generators for detecting Z and X errors, respectively.

[0023] 11. Surface code scale L: One-quarter of the perimeter of the surface code array. The surface code array in Figure 1 has L=7, which means a total of 97 physical qubits, including 49 data qubits and 48 auxiliary qubits.

[0024] 12. X and Z errors: These are errors caused by the Pauli X and Pauli Z operators that randomly occur in the quantum state of a physical qubit. According to quantum error correction theory, if an error-correcting code can correct X and Z errors, it can also correct errors that occur in any single qubit.

[0025] 13. Fault-tolerant quantum error correction (FTQEC): All operational processes in actual quantum computing, including quantum gates and quantum measurements, are accompanied by noise. This means that the circuits used for quantum error correction themselves contain noise. Fault-tolerant quantum error correction refers to a clever design that makes it possible to correct errors even using noisy correction circuits, thereby achieving the goal of correcting errors and preventing them from spreading over time.

[0026] 14. Fault-tolerant quantum computation (FTQC): In the process of quantum computing, any physical operation, including the quantum error correction circuit itself and quantum bit measurement, is accompanied by noise. Assuming that classical operations (such as inputting instructions and decoding error-correcting codes) are noise-free, fault-tolerant quantum computing is a technical solution for effectively controlling and correcting errors in the process of quantum computing using noisy quantum bits by rationally designing a quantum error correction scheme and performing quantum gate operations in a specific manner on the encoded logical quantum state.

[0027] 15. Physical qubit: A qubit realized using an actual physical device.

[0028] 16. Logical qubit: A mathematical degree of freedom in a Hilbert subspace defined by an error-correcting code. Its quantum state typically describes a multi-body entangled state, typically a two-dimensional subspace combining multiple physical qubits and a Hilbert space. Fault-tolerant quantum computation must be performed on logical qubits protected by error-correcting codes.

[0029] 17. Quantum gates / circuits: Quantum gates / circuits that operate on physical qubits.

[0030] 18. Threshold theorem: For a quantum computing scheme that meets the requirements of fault-tolerant quantum computing, if the error rate of all operations is lower than a certain threshold, the accuracy rate of the computation can be arbitrarily approached to 1 by using better error-correcting codes, more quantum bits, and more quantum operations, and at the same time, these additional resource overheads are negligible compared to the exponential acceleration of quantum computing.

[0031] 19. Neural Network: An artificial neural network is an adaptive, nonlinear dynamic system composed of a large number of interconnected simple basic elements called neurons. While the structure and function of each neuron are relatively simple, the system behavior generated by the combination of a large number of neurons is extremely complex, and in principle, it can represent any function.

[0032] 20. Convolutional Neural Network (CNN): A convolutional neural network is a type of feedforward neural network that includes convolutional operations and has a deep structure. The convolutional layer is the core foundation of a convolutional neural network, which performs convolution operations between a discrete two-dimensional or three-dimensional filter (also called a convolution kernel, which is a two-dimensional or three-dimensional matrix, respectively) and a two-dimensional or three-dimensional data point cloud.

[0033] 21. Field Programmable Gate Array (FPGA): This is a further developed product based on programmable devices such as PAL (Programmable Array Logic) and GAL (Generic Array Logic). It has emerged as a kind of semi-custom circuit in the field of Application Specific Integrated Circuit (ASIC), solving the shortage of custom circuits and overcoming the drawback of the limited number of gate circuits in traditional programmable devices.

[0034] 22. Application Specific Integrated Circuit (ASIC): Refers to an integrated circuit designed and manufactured according to the requirements of a specific user and the needs of a specific electronic system. Designing ASIC using CPLD (Complex Programmable Logic Device) and FPGA is one of the most popular methods. What they have in common is that they are both field programmable by the user and support boundary scan technology, but they have their own characteristics in terms of integration, speed, and programming method.

[0035] 23. Single Flux Quantum (SFQ) Circuit: Also known as an RSFQ (Rapid Single Flux Quantum) circuit, this is a circuit composed of Josephson junctions (JJs) that represent "1" or "0" depending on the presence or absence of a flux quantum. In the circuit, "X" represents a Josephson junction. The top and bottom layers are made of superconductors, and the middle layer is made of a very thin insulator. It can be used for digital logic calculations.

[0036] 24. Neural Network Model Quantization: Quantization is the compression of the original network by reducing the number of bits required to represent each weight. Neural network models generally require very large storage capacity. When implementing neural network algorithms on hardware chips, this occupies a large amount of valuable on-chip memory, on-chip registers, and wiring, severely impacting computation speed. Because these parameters are floating-point numbers, conventional compression algorithms cannot effectively address this situation. If calculations can be performed using other simple numerical types (e.g., fixed-point arithmetic) within the model without affecting the accuracy of the model, the consumed hardware computing resources (including hardware calculation units and storage units) can be significantly reduced. For algorithm calculation chips in real-time feedback systems, quantization algorithms can increase the amount of calculation per unit time and reduce the delay of the decoding algorithm.

[0037] 25. Multiply Accumulate (MAC): This is a special operation in digital signal processors or some microprocessors. The hardware circuit unit that realizes this operation is called a "multiplier accumulator" or "multiplication accumulator." This operation adds the result of the multiplication to the value of the accumulator and stores it back in the accumulator.

[0038] 26. Low-Voltage Differential Signaling (LVDS): A type of electrical standard for low-voltage differential signaling, it is a method of serially transmitting data using differential signals. Its physical implementation typically uses twisted pair cables, enabling high-speed transmission at relatively low voltages. LVDS is only a physical layer specification and is not involved in protocol layer communication methods, so it has very low latency.

[0039] 27, Given an arbitrary Pauli operator P and a generator S of the stabilizer group of an error-correcting code C (we use a rotating surface code as an example in this application, but any topological error-correcting code can be defined in a similar way), the Pauli operator P acting on the physical qubits supporting the error-correcting code can be decomposed as follows: P=L(P)T(S(P)) Here, S(P) is the generator (also called the syndrome of the Pauli operator P) of the part of S that anticommutes with P, and S(P) can be considered as a bit array consisting of 0s and 1s. T(S(P)) is the Pauli operator generated by mapping based on this generator. In quantum mechanics, if the operators F and G satisfy FG=GF, then the operators F and G are said to commute, and if the operators F and G satisfy FG=-GF, then the operators F and G are said to anticommutate. T(S(P)) and S(P) have a one-to-one correspondence and are called simple representations of P. For rotation surface codes, a geometrically meaningful definition can be given to the simple representations, namely, they can be defined as the shortest Pauli operators connecting the syndromes to the boundary that anticommute with P. An X-type Pauli operator is the shortest Pauli operator connecting syndrome point a to the boundary that anticommute with P. A Z-type Pauli operator is the shortest Pauli operator connecting syndrome point b to the boundary that anticommute with P. In general, all topological error-correcting codes can have similar simple representation mappings.

[0040] L(P) is one particular operator in the logical type of error-correcting code to which P belongs (once chosen, it is fixed). A similar decomposition applies when considering another Pauli operator P'. P'=L(P')T(S(P'))

[0041] If S(P')=S(P) and L(P') and L(P) belong to the same logic type, then P' and P differ by one element of the stabilizer group, i.e., they are equivalent in terms of error correction. Then, for any Pauli operator P, we can define P C =L C (P)T(S(P)) where L C (P) is a fixed representative element of the logical type to which L(P) belongs, and P C is called the canonical representation of P, and L C (P)T(S(P)) is called the canonical decomposition of P. It converts all Pauli operators into their equivalent canonical representations. This significantly limits the unnecessary diversity of Pauli operators, which significantly reduces the difficulty of model training and improves the convergence speed of the training process, especially when Pauli operators are selected as the model output.

[0042] The technical solution of this application relates to the fields of quantum technology and artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system for using digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results. In other words, artificial intelligence is a comprehensive technology of computer science that attempts to understand the essence of intelligence and create new intelligent machines that can respond in a manner similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines so that machines can have the capabilities of perception, reasoning, and decision-making.

[0043] Artificial intelligence technology is a comprehensive academic field that encompasses a wide range of fields, including both hardware and software technologies. Fundamental technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. Artificial intelligence software technology primarily encompasses key fields such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0044] Machine learning (ML) is a multidisciplinary field that spans various fields, including probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. ML focuses on the study of how computers can simulate or realize human learning behavior to acquire new knowledge and skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and a fundamental means of endowing computers with intelligence, and is applied to various fields of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, trust networks, reinforcement learning, transfer learning, inductive learning, and teaching learning.

[0045] With the research and progress of artificial intelligence technology, it has been researched and applied in various fields, such as general smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, robots, smart medical care, and smart customer service. With the development of technology, it is believed that artificial intelligence technology will be applied in more fields and will play an increasingly important role.

[0046] The solutions provided in the embodiments of the present application relate to the application of artificial intelligence machine learning technology in the quantum technology field, specifically, the application of quantum error correcting codes in quantum error correcting code decoding algorithms, which will be described in detail with reference to the following embodiments.

[0047] Because qubits are highly susceptible to noise, achieving quantum computing directly on physical qubits is not yet practical with current technology. Developments in quantum error correction and fault-tolerant quantum computing technology offer the possibility of achieving arbitrary precision quantum computing on noisy qubits. Generally, measuring the stabilizer generator of a quantum error-correcting code (also known as a qubit parity check) requires the introduction of long-range quantum gates, while simultaneously requiring the preparation of complex quantum auxiliary states using additional qubits to achieve fault-tolerant error correction. Due to limitations in current experimental tools, people do not yet have the ability to achieve high-precision long-range quantum gates or the preparation of complex quantum auxiliary states. On the other hand, methods for fault-tolerant quantum error correction and fault-tolerant quantum computing using surface codes do not require the use of long-range quantum gates or the preparation of complex quantum auxiliary states, and are therefore considered highly likely to realize universal fault-tolerant quantum computers using current technology.

[0048] For an error-correcting code, after an error occurs, a parity check can be used to obtain an error syndrome. Based on these syndromes, the location and type of error (X error, Z error, or Y error, which includes both) must be determined based on the specific decoding algorithm of the error-correcting code. In the case of a surface code, the error and error syndrome have a specific spatial location. If an error causes a syndrome, the eigenvalue of the ancillary quantum bit at the corresponding location is 1 (a point particle can be considered to have appeared at that location). If there is no error, the eigenvalue of the ancillary quantum bit at the corresponding location is 0. In this case, decoding can be reduced to the following problem: given a spatial digital array (two or three dimensions, with values ​​of 0 or 1), based on a specific error generation model (i.e., the error probability distribution of the quantum bits), infer which quantum bit is most likely to generate an error and its specific error type, and then perform error correction based on this inference.

[0049] As mentioned above, by employing a decoding algorithm for error-correcting codes to decode the error syndrome information, corresponding error result information, including the location and type of the error, can be obtained. This application mainly focuses on the hardware implementation architecture of a neural network-based quantum error-correction decoding algorithm and proposes a neural network quantum error-correction hardware decoder based on a programmable architecture. The quantum error-correction hardware decoder described in this application refers to a hardware architecture realized by implementing a decoding algorithm related to quantum error correction on corresponding hardware resources, and is applicable to real-time quantum error correction scenarios.

[0050] The quality of a quantum error correction hardware decoder can be evaluated based on several performance indicators, such as decoding performance, decoding delay, scalability, and flexibility.

[0051] Decoding performance (accuracy): Measured using the error rate generated by the logical qubits after error-correction decoding based on a specific noise model. For the same physical qubit error rate, the lower the logical error rate, the better the decoding performance. For a hardware decoder, the upper limit of its decoding performance is determined by the decoding algorithm used. During the hardware implementation process, specific hardware resources may impose certain limitations on the implementation of the algorithm (such as quantization of neural networks), which may cause a certain loss of accuracy.

[0052] Decoding Delay: The lifetime of a superconducting qubit is approximately 150 microseconds under favorable process conditions. During the decoding process, the system is idle and errors gradually accumulate over time. Theoretically, the time consumed by the entire error correction process must be less than 1 / 1000 to 1 / 100 of the superconducting qubit lifetime, meaning the entire correction time must be less than 1.5 microseconds (us), otherwise the error rate may exceed the correction capability of the surface code. Targeted optimization of the hardware decoder plays a crucial role in achieving this delay goal.

[0053] Scalability: It is desirable that the resource consumption of a hardware decoder increases gently with increasing scale L. Otherwise, it may become impossible to realize a decoder for a larger L. This means that the decoding algorithm used must first have low algorithmic complexity, and then the resource consumption must be further reduced by optimizations on the hardware architecture.

[0054] Flexibility: In practical application scenarios, hardware decoders must adapt to various environmental configurations, such as different noise levels, surface code distances, and error-correcting code formats. Therefore, it is desirable to flexibly configure the decoder and frequently reconfigure it for experimentation to find the optimal design for different scenarios. This requirement is particularly important when implementing hardware decoders using ASICs, because the ASIC redesign and deployment process is lengthy and costly, making it impractical for engineering and implementing application scenarios.

[0055] 2 is a schematic diagram of an application scenario of the solution according to one embodiment. As shown in FIG. 2, the application scenario may be a superconducting quantum computing platform, which includes a quantum circuit 21, a dilution refrigerator 22, a control device 23, and a computer 24.

[0056] The quantum circuit 21 is a circuit that operates on physical quantum bits, and the quantum circuit 21 can be realized as a quantum chip, such as a superconducting quantum chip at near absolute zero. The dilution refrigerator 22 is used to provide an absolute zero environment for the superconducting quantum chip.

[0057] Control device 23 is used to control quantum circuit 21, and computer 24 is used to control control device 23. For example, a created quantum program is compiled into instructions by software in computer 24 and sent to control device 23 (e.g., an electronic / microwave control system), and control device 23 converts the instructions into electronic / microwave control signals and inputs them into dilution refrigerator 22 to control the 10 mK superconducting quantum bits. The readout process is the reverse of this.

[0058] The present application provides an optimized quantum algorithm suitable for real-time error correction of superconducting qubits. As shown in FIG. 3 , the real-time fault-tolerant quantum error correction algorithm works in conjunction with a control device 23 (e.g., by integrating the decoding algorithm into an electronic / microwave control system). After an overall control system 23a (e.g., a central board FPGA) of the control device 23 reads error syndrome information from the quantum circuit 21, the overall control system 23a sends an error correction command to an error correction module 23b of the control device 23, the error correction command including the error syndrome information of the quantum circuit 21. The error correction module 23b may be an FPGA or an ASIC chip, and the error correction module 23b executes a quantum error correction decoding algorithm based on a neural network to decode the error syndrome information and convert the error result information obtained by decoding in real time into an error correction control signal, which is then sent to the quantum circuit 21 for error correction.

[0059] FIG. 4 is a schematic diagram of a quantum error correction hardware decoder architecture according to one embodiment of the present application.

[0060] In some embodiments, a quantum error correcting code can be used to perform error syndrome measurements on a quantum circuit to obtain corresponding error syndrome information, which is a data array consisting of eigenvalues ​​of the stabilizer generator of the quantum error correcting code. Exemplarily, the error syndrome information is a two-dimensional or three-dimensional data array consisting of 0s and 1s. For example, if there is no error, the eigenvalue of the stabilizer generator is 0, and if an error occurs, the eigenvalue of the stabilizer generator is 1. Taking the quantum error correcting code as a surface code, for example, errors and error syndromes have specific spatial locations. When an error causes a syndrome, the eigenvalue of the ancillary quantum bit at the corresponding location is 1 (which can be considered as a point particle appearing at that location), and when there is no error, the eigenvalue of the ancillary quantum bit at the corresponding location is 0. Therefore, for a surface code, if the error of the correction process itself is not taken into account (i.e., the measurement process is perfect, in this case called a perfect syndrome), the error syndrome information can be considered as a two-dimensional data array consisting of 0s and 1s. For example, if multiple syndrome measurements are performed on a quantum circuit, error syndrome information in the form of a two-dimensional data array can be obtained with each syndrome measurement, and error syndrome information in the form of a three-dimensional data array can be obtained with multiple syndrome measurements.

[0061] In some embodiments, for a surface code with scale L, T measurements (where T is greater than L) are required to ensure fault tolerance of measurement errors. Each measurement produces a one-bit result corresponding to each ancillary qubit. The results of multiple measurements are merged into a three-dimensional data array consisting of 0s and 1s, which becomes the error syndrome information for the quantum circuit.

[0062] Optionally, the quantum error correction hardware decoder obtains the error syndrome information of the quantum circuit through an overall control system. For example, the overall control system instructs the quantum error correction hardware decoder to start quantum error correction through an error correction command. Illustratively, the error correction command includes the error syndrome information of the quantum circuit, or the quantum error correction hardware decoder reads the error syndrome information from the overall control system in response to the received error correction command. That is, the error syndrome information of the quantum circuit is sent to the quantum error correction hardware decoder, which decodes the error syndrome information to finally obtain error information of the quantum circuit, which indicates the quantum bit in which an error has occurred in the quantum circuit and the error type corresponding to the error. The error type includes at least one of an X error and a Z error.

[0063] In some embodiments, the error information of different error types is obtained by decoding the error syndrome information by different quantum error correction decoders. For example, two quantum error correction hardware decoders, denoted as a first quantum error correction hardware decoder and a second quantum error correction hardware decoder, are designed. Here, the first quantum error correction hardware decoder (i.e., the X error decoder in FIG. 4 ) is used to obtain X-type error information by decoding, where the X-type error information indicates a qubit in which an X error has occurred in the quantum circuit. The second quantum error correction hardware decoder (i.e., the Z error decoder in FIG. 4 ) is used to obtain Z-type error information by decoding, where the Z-type error information indicates a qubit in which a Z error has occurred in the quantum circuit.

[0064] In one possible embodiment, the error syndrome information is decomposed into two parts, namely, X-type error syndrome information and Z-type error syndrome information, which are input into a first quantum error correction hardware decoder and a second quantum error correction hardware decoder, respectively; the X-type error syndrome information is decoded by the first quantum error correction hardware decoder to obtain X-type error information; and the Z-type error syndrome information is decoded by the second quantum error correction hardware decoder to obtain Z-type error information, which helps to reduce the computational complexity of the decoders.

[0065] In another possible embodiment, the first quantum error correction hardware decoder and the second quantum error correction hardware decoder both receive error syndrome information, and do not distinguish between X-type and Z-type error syndrome information, but instead use all syndrome bits together to decode X and Z errors. Because X and Z errors are interrelated, the above method considers all syndrome bits together to more accurately determine the locations of X and Z errors.

[0066] In some other embodiments, error information of different error types can be obtained by decoding error syndrome information using the same quantum error correction decoder. For example, only one quantum error correction hardware decoder is designed, and error syndrome information is input to the quantum error correction hardware decoder, and the error syndrome information is decoded by the quantum error correction hardware decoder to obtain error information. The error information includes the quantum bit in which an error occurred in the quantum circuit and the error type corresponding to the error. That is, the X error and the Z error are simultaneously determined by the single quantum error correction hardware decoder.

[0067] A quantum error correction hardware decoder decodes error syndrome information based on a neural network model. A neural network model refers to an AI model built based on a neural network. The quantum error correction hardware decoder needs to store a network parameter file, which contains parameters of the neural network model (e.g., weight parameters and bias parameters of the neural network model). The calculation process of the neural network model uses a large number of parameters, all of which are obtained in advance by training with software and loaded into the network parameter file before error correction begins. These neural network model parameters are divided into different groups and loaded into the calculation unit, requiring the decoder to quickly switch between parameters in each group during the real-time decoding process.

[0068] Therefore, the present application preferably selects on-chip storage to store the network parameter file to avoid long read delays from off-chip storage. The entire network parameter file is divided into two parts according to different data structures: one part is used to store weight parameters (also called a weight matrix), and the other part is used to store bias parameters (also called a bias vector). These parameters are originally floating-point numbers, which require complex multiplication operations and consume a large amount of memory. To improve storage and computation efficiency, the present application quantizes these parameters to 8-bit signed fixed-point numbers and implements them on an FPGA. By selecting non-saturating quantization, the implementation of the computation unit and data file is simplified while causing only negligible precision loss. Of course, the above floating-point format parameters can also be quantized to signed fixed-point numbers with other bit counts. This can be determined by comprehensively considering multiple aspects, such as the decoder's computational power, precision requirements, and algorithm complexity, and the embodiments of the present application are not limited to this.

[0069] In a quantum error correction hardware decoder, the main hardware unit responsible for neural network operations is a neural network processing unit (also called a neural network processing engine (NPE)), and the number of neural network processing units may be one or more. Each neural network processing unit may include multiple arithmetic units (AUs). In some embodiments, the neural network model can be constructed based on 3D CNN and FCN (Fully Connected Neural Network). In the neural network processing unit, most arithmetic units perform a similar vector dot product operation, which can be expressed as the following equation:

number

[0070] where y n+1 represents the output data of the nth layer of the neural network model, which also serves as the input data of the n+1th layer, where n is a positive integer. j,i n is the weight coefficient corresponding to the calculation node in the jth row and ith column of the nth layer, and y i n is the i-th element of the input data vector of the n-th layer, and s j n represents the calculation result obtained by performing weighted addition on the input data vector in the nth layer based on the weight coefficient corresponding to the calculation node in the jth row of the nth layer coefficient matrix, and b j n is the bias coefficient of the jth row of the nth layer, A represents a nonlinear function, and y j n+1represents the output data of the jth row of the nth layer. The bias vector can be stored in a simple set of registers because it is used only once per iteration. The multiply-and-add operations in the neural network processing unit account for the majority of the computational resources of the entire decoder.

[0071] The quantum error correction hardware decoder further includes a batch error search unit, and the batch error search unit is configured to determine error information based on the output result of the neural network model. For example, when decoding X errors and Z errors separately, the first quantum error correction hardware decoder and the second quantum error correction hardware decoder each have a length of (L 2 The decoding result, i.e., the possible error combinations, is generated as (L −1) / 2 bits, where L is the surface code scale. This decoding result is used as an address to search the lookup table. Correspondingly, the lookup table also contains (L 2 -1) / 2 entries, each of size L 2 and corresponds to the number of data qubits. Therefore, the memory consumption of the lookup table is

number

[0072] 5 is a schematic diagram of a quantum error correction hardware decoder according to one embodiment of the present application, which includes an instruction storage device 51, a control unit 52, at least one neural network processing unit 54, and an error batch search unit 53.

[0073] The instruction storage device 51 is configured to store computer instructions.

[0074] The control unit 52 is configured to read computer instructions from the instruction storage device 51 and control, based on the computer instructions, the at least one neural network processing unit 54. In response to the control of the control unit 52, the at least one neural network processing unit 54 is configured to decode, based on the neural network model, error syndrome information of the quantum circuit to obtain an output result of the neural network model, where the error syndrome information refers to an error syndrome obtained by performing error measurement on the quantum circuit.

[0075] The error batch search unit 53 is configured to determine error information based on the output result, the error information being for indicating the quantum bit in which an error occurs in the quantum circuit and the error type corresponding to the error.

[0076] The control unit 52 is mainly responsible for instruction scheduling for the entire decoding process. Before real-time error correction decoding begins, a compiler can generate a series of assembly codes (also called computer instructions) describing the decoding algorithm and load them into the instruction storage device 51. The assembly codes are used to implement the processing flow of the entire decoding algorithm, thereby controlling at least one neural network processing unit 54 to decode the error syndrome information of the quantum circuit based on the neural network model and obtain the output result of the neural network model. After real-time error correction decoding begins, the control unit 52 reads the computer instructions from the instruction storage device 51 and performs decoding and processing. The above scheme makes the quantum error correction hardware decoder according to the embodiment of the present application programmable, allowing corresponding codes to be flexibly written according to the decoding algorithm, thereby improving the versatility of the quantum error correction hardware decoder for different error correction algorithms.

[0077] Optionally, the control unit 52 receives an error correction command sent from the overall control system and obtains error syndrome information of the quantum circuit from the error correction command. The control unit 52 transmits the error syndrome information of the quantum circuit to the neural network model, and the neural network model decodes the error syndrome information of the quantum circuit. That is, the error syndrome information of the quantum circuit is used as input data for the neural network model. The specific format of the error syndrome information of the quantum circuit can be referred to in the above embodiments and will not be described again here.

[0078] For example, the number of network layers in the neural network model and the role of each network layer are determined according to the quantum error correction algorithm actually used. For example, the neural network model is designed based on networks such as 3D CNN, FCN, and recurrent neural network (RNN), and is used to decode the type of error that occurred based on the error feature information. This application does not limit the model structure of the neural network model.

[0079] In some embodiments, the decoding process of the neural network model is divided into n subprocesses that are executed sequentially, where different subprocesses reuse the neural network processing unit 54, and n is an integer greater than 1. Optionally, the subprocesses included in the decoding process are obtained by dividing the decoding process. For example, the decoding process is divided into n subprocesses according to the role of each stage of the decoding process. For any of the n subprocesses, the subprocess first obtains input data, and during the subprocess, the subprocess processes the input data to obtain output data. After the subprocess is completed, the output data is output.

[0080] The neural network model decoding process refers to the process of decoding the error syndrome information of the quantum circuit based on the neural network model to obtain the output result of the neural network model. The decoding process can also be understood as a classification process, that is, encoding the error syndrome information to obtain feature information, and decoding based on the feature information to obtain the decoded result of the error syndrome information.

[0081] Optionally, the output result of the neural network model includes a plurality of decoding results. Illustratively, the type of the decoding result includes at least one of a probability distribution of an error type included in the error syndrome information, a representative element associated with the error type, and a canonical syndrome associated with the type. Here, the canonical syndrome refers to a canonical decomposition result of the error syndrome information. For example, the plurality of decoding results include a probability distribution of each error type included in the error syndrome information. In another example, the plurality of decoding results include a representative element associated with each error type and a canonical syndrome associated with each error type.

[0082] It should be noted that the output result of the neural network model is related to the training data selected in the neural network training process. For example, when a neural network model is trained using sample error syndrome information and a probability distribution of error types, the output result after the trained neural network model processes the error syndrome information includes a probability distribution of error types. In another example, when a neural network model is trained using sample error syndrome information and representative elements related to error types, the output result after the trained neural network model processes the error syndrome information includes a probability distribution of error types. The present application does not limit the format of the output result of the neural network model.

[0083] Optionally, based on the structure of the neural network model, the decoding process of the neural network model is divided into n subprocesses that are executed sequentially. Here, the (i+1)th subprocess is executed after the (i)th subprocess, i.e., the (i+1)th subprocess starts executing after the (i)th subprocess is executed, where i is a positive integer less than or equal to n. Exemplarily, the neural network model has a four-layer structure, including a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fully connected layer that are connected in series. The input data of the first convolutional layer includes error syndrome information of the quantum circuit. The first convolutional layer performs a convolutional operation on the error syndrome information of the quantum circuit to obtain output data of the first convolutional layer, which is used as input data of the second convolutional layer. The second convolutional layer performs a convolutional operation on its input data to obtain output data of the second convolutional layer, which is used as input data of the third convolutional layer. The third convolutional layer performs a convolution operation on its input data to obtain output data for the third convolutional layer, which is used as input data for the fully connected layer. The fully connected layer performs a vector dot product operation on its input data to obtain an output result for the neural network model. Based on the structure of the above neural network model, the decoding process of the neural network model can be divided into four subprocesses that are executed sequentially, where the first subprocess refers to the operation process of the first convolutional layer, the second subprocess refers to the operation process of the second convolutional layer, the third subprocess refers to the operation process of the third convolutional layer, and the fourth subprocess refers to the operation process of the fully connected layer.

[0084] The parameters of the neural network model used in different sub-processes may be different. Optionally, at the start of a specific sub-process, the neural network processing unit 54 acquires the parameters of the neural network model used in that sub-process. To improve the speed of acquiring the parameters of the neural network model used in that sub-process, the parameters of the neural network model are illustratively divided into pre-stored blocks, and for any parameter block, the parameters of the neural network model included in that parameter block are used in that specific sub-process. A parameter block can be understood as parameters with consecutive storage addresses or parameters with a certain regularity, and these parameters must be used in the same sub-process. For example, at the start or soon-to-start stage of a sub-process, a quantum error correction hardware decoder can acquire all parameters required for the progress of that sub-process by reading the parameters in the parameter block. For specific steps of parameter reading, see the example below.

[0085] This division allows different subprocesses to reuse the neural network processing unit 54. For example, the neural network processing unit 54 can start execution of a first subprocess, and after completing the first subprocess, the neural network processing unit 54 can start execution of a second subprocess, and so on, until the nth subprocess is completed, thereby obtaining the output result of the neural network model. By dividing the decoding process of the neural network model into multiple subprocesses that are executed sequentially and allowing different subprocesses to reuse the neural network processing unit, it is no longer necessary to assign dedicated arithmetic units to different network layers, which improves the utilization of each arithmetic unit in the decoder and reduces the complexity and cost of the decoder's hardware design.

[0086] In some embodiments, each sub-process includes a first stage, a second stage, and a third stage. The first stage is used to perform a multiplication and addition operation on input data of the sub-process to obtain at least one first calculation result. The second stage is used to perform an addition operation on the at least one first calculation result to obtain at least one second calculation result. The third stage is used to obtain output data of the sub-process based on the at least one second calculation result. Here, the input data of a first sub-process among the n sub-processes includes error syndrome information of the quantum circuit, and the output data of a last sub-process (i.e., the nth sub-process) among the n sub-processes includes an output result of the neural network model.

[0087] As can be seen from the introduction of the main calculation process of the neural network model above, the neural network model mainly involves multiplication and addition operations. Therefore, each subprocess is divided into three stages. The first stage, also called the multiplication and addition stage, is used to perform multiplication and addition operations on the input data of the subprocess to obtain at least one first calculation result. For example, the first stage mainly performs multiplication operations to obtain

number

[0088] For any of the three sequentially executed stages included in the subprocess, the operation type to be performed in the same stage is the same, the types of arithmetic units used are similar, and data transmission between chronologically adjacent stages is stored. Dividing the subprocess into multiple stages facilitates designing the layout of the arithmetic units used in each stage in the neural network processing unit and designing the wiring between the arithmetic units. This improves the regularity of the arithmetic units in the neural network processing unit, supports performance expansion of the neural network processing unit in later stages, supports larger surface code scales, and improves the adaptability of the quantum error correction hardware decoder to different error correction situations.

[0089] In some embodiments, when each sub-process is divided into the above three stages, the neural network processing unit 54 includes a first subunit 54a, a second subunit 54b, and a third subunit 54c. The first subunit 54a includes at least one arithmetic unit for performing the first stage. For example, the first subunit 54a includes one or more multiplication arithmetic units. Optionally, the first subunit 54a further includes one or more addition arithmetic units. The second subunit 54b includes at least one arithmetic unit for performing the second stage. For example, the second subunit 54b includes one or more addition arithmetic units. The third subunit 54c includes at least one arithmetic unit for performing the third stage. For example, the third subunit 54c includes an adder. Optionally, the third subunit 54c further includes a register, and the register and the adder are used in combination to perform an accumulation operation.

[0090] The type of arithmetic unit included in the subunit is related to the operation that needs to be performed at the stage of execution of the subunit, and is not limited here.

[0091] In the embodiment of the present application, the number of multiplication units and addition units included in the first subunit 54a is not particularly limited and can be designed according to actual needs. For example, if the first subunit 54a includes four multiplication units, it supports parallel calculation of four sets of multiplication operations. If the first subunit 54a includes eight multiplication units, it supports parallel calculation of eight sets of multiplication operations. The more multiplication units included in the first subunit 54a, the more parallel multiplication operations it supports, and the correspondingly more complex the structure of the first subunit 54a. In addition, the addition units included in the first subunit 54a may have a tree structure and add the multiplication results obtained by the multiplication units. In actual applications, the appropriate number of multiplication units and addition units can be selected by comprehensively considering various aspects such as the complexity of the overall decoding algorithm, decoding delay, and hardware complexity.

[0092] The operations that need to be performed at each stage are not exactly the same, and there is a chronological relationship between each stage. Therefore, by treating each stage as the smallest unit and setting a corresponding arithmetic unit for each stage, not only can the normal progress of decoding be ensured, but also the appearance of redundant arithmetic units can be avoided, thereby reducing the hardware overhead of the quantum error correction hardware decoder.

[0093] In some embodiments, the first sub-unit 54a further includes a data operation unit and a weight operation unit, where the data operation unit is configured to read input data used in the first stage from a first register, and the weight operation unit is configured to read weight parameters of a neural network model used in the first stage from a second register, and the input data and weight parameters are sent to corresponding multiplication operation units for operation.

[0094] Optionally, the process of the data manipulation unit reading the input data used in the first stage from the first register and the process of the weight manipulation unit reading the weight parameters of the neural network model used in the first stage from the second register can be performed in parallel.

[0095] For example, the control unit 52 sends a first computer instruction to the data operation unit and a second computer instruction to the weight operation unit in the same clock cycle, so that the data operation unit reads input data used in the first stage from a first register based on the first computer instruction, and the weight operation unit reads weight data used in the first stage from a second register based on the second computer instruction. This setting reduces the time difference between the weight data and the input data reaching the multiplication operation unit in the first sub-unit, thereby shortening the time required for the first stage.

[0096] Of course, the control unit 52 may send the first computer instruction and the second computer instruction in the same clock cycle, or may send the first computer instruction and the second computer instruction in different clock cycles, although the present application is not limited thereto.

[0097] Optionally, there is a wired connection between the data operation unit and the arithmetic unit included in the first sub-unit 54a, where the data operation unit transmits input data to the arithmetic unit via the wire, and there is a wired connection between the weight operation unit and the arithmetic unit included in the first sub-unit 54a, where the weight operation unit transmits weight data to the arithmetic unit via the wire.

[0098] Considering that a neural network model typically processes multidimensional input data in vector or matrix format, and different elements (understood as one numerical value in the input data) in the vector or matrix correspond to different weight data, there are hardwired connections between the data operation units and at least one arithmetic unit included in the first subunit 54a, and the data operation units transmit the above elements to the corresponding arithmetic unit via the hardwired connections.

[0099] The data operation unit and weight operation unit ensure that input data and weight parameters are correctly provided to the corresponding multiplication operation units for processing.

[0100] 5, the quantum error correction hardware decoder further includes a first register and a second register. The first register is configured to store input data used in the decoding process of the neural network model and to store output data generated in the decoding process of the neural network model. The second register is configured to store weight parameters and bias parameters of the neural network model.

[0101] Optionally, error syndrome information is stored in the first register. Illustratively, the first register obtains the error syndrome information from the error correction instruction and stores the error syndrome information. In each sub-process of the decoding process, the neural network processing unit can read data and parameters required for that sub-process from the two registers, and after performing that sub-process, store the output data obtained in that sub-process in the first register for subsequent use. By setting the two registers, the data and parameters used in the decoding process can be stored separately, thereby ensuring the efficiency of reading and writing data and parameters.

[0102] The first subunits 54a are scalable. For example, the number of multiplication and addition units in the first subunits 54a can be increased as needed to support a larger surface code scale. Optionally, the number of first subunits 54a can also be increased as needed. For example, multiple first subunits 54a can be configured to execute the operations required in the first stage in parallel, and the operation results of the multiple first subunits 54a can be finally integrated to obtain at least one first calculation result. In the first stage, which one or more first subunits 54a to select for use in the operation and which one or more multiplication and addition units to select for use in the operation can be specified in advance by computer instructions stored in the instruction storage device 51. The control unit 52 can control the corresponding first subunits 54a and operation units to perform the corresponding operations based on the computer instructions.

[0103] In the embodiment of the present application, the number of addition units included in the second subunit 54b is not particularly limited and can be designed according to actual needs. The addition units included in the second subunit 54b may have a tree structure and add at least one first calculation result. The number of addition units included in the first subunit 54a and the second subunit 54b can be determined based on the number of multiplication units included in the first subunit 54a. For example, if the number of multiplication units included in the first subunit 54a is C, the required depth of the addition tree is log2C. The number of addition units included in the first subunit 54a and the second subunit 54b can be determined based on the depth of the addition tree log2C. For example, if the number of multiplication units is 4, the depth of the addition tree is 2, and a total of 2 + 1 = 3 addition units are included. If the number of multiplication units is 8, the depth of the addition tree is 3, and a total of 4 + 2 + 1 = 7 addition units are included. If the number of multiplication operation units is 16, the depth of the addition tree is 4, and it contains a total of 8+4+2+1=15 addition operation units, and so on.

[0104] In some embodiments, the second subunit 54b further includes at least one multiplexer (a module indicated by a trapezoid in FIG. 5 ), configured to prefetch calculation results at different depths in the calculation unit tree (e.g., an addition tree) corresponding to the second stage. The calculation unit tree includes a plurality of calculation units, which have a hierarchical structure in the calculation unit tree. For a first layer other than the bottom layer of the calculation unit tree, input data of the calculation unit in the first layer is obtained from a calculation unit in a second layer. For example, the input data of the calculation unit in the first layer is the calculation result of at least one calculation unit in the second layer, and the second layer refers to a layer adjacent to and lower than the first layer in the calculation unit tree. For example, the calculation unit tree includes four layers, and the hierarchical order from highest to lowest is Layer 1, Layer 2, Layer 3, and Layer 4. If the first layer is Layer 2, the second layer is Layer 1.

[0105] In the first stage, considering that not all multiplication units need to be used every time, such a prefetch mechanism allows for flexible configuration of the number of multiplication operations. For example, assume that there are eight multiplication units, as shown in FIG. 6, corresponding to the circles numbered 1 to 8 in FIG. 6. The two operation results obtained by the multiplication units numbered 1 and 2 are added together in one addition operation unit A1 to obtain a sum. The two operation results obtained by the multiplication units numbered 3 and 4 are added together in one addition operation unit A2 to obtain a sum. The two operation results obtained by the multiplication units numbered 5 and 6 are added together in one addition operation unit A3 to obtain a sum. The two operation results obtained by the multiplication units numbered 7 and 8 are added together in one addition operation unit A4 to obtain a sum. Furthermore, the sums obtained by the addition operation units A1 and A2 are added together in one addition operation unit B1 to obtain a sum. The summation results obtained by the addition operation units A3 and A4 are added together in one addition operation unit B2 to obtain a summation result. The summation results obtained by the addition operation units B1 and B2 are added together in one addition operation unit C1 to obtain a summation result. The addition operation units A1 to A4, B1 to B2, and C1 jointly constitute an addition operation tree. In some cases, only one addition operation unit may be used to add two numbers. For example, only addition operation unit A1 may be used to add two operation results obtained by multiplication operation units numbered 1 and 2. Alternatively, in some cases, only three addition operation units may be used to add four numbers. For example, only addition operation units A1, A2, and B1 may be used to add four operation results obtained by multiplication operation units numbered 1 to 4. Therefore, the multiplexer can be configured to prefetch calculation results at different depths in the addition operation tree, which improves the adaptability of the second subunit 54b to different operation scenarios, increases flexibility, and reduces delays.

[0106] The second subunits 54b are also scalable. For example, the number of summing operation units in the second subunits 54b can be increased as needed to support a larger surface code scale. Optionally, the number of second subunits 54b can also be increased as needed. For example, multiple second subunits 54b can be configured to execute the operations required in the second stage in parallel, and the operation results of the multiple second subunits 54b can be finally integrated to obtain at least one second calculation result. In the second stage, which one or more second subunits 54b to select for use in the operation and which one or more summing operation units to select for use in the operation can be specified in advance by computer instructions stored in the instruction storage device 51. The control unit 52 can control the corresponding second subunits 54b and operation units to perform the corresponding operations based on the computer instructions.

[0107] In the embodiment of the present application, the specific type and number of arithmetic units included in the third subunit 54c are not limited, and can be designed accordingly depending on the operations required to be performed in the third stage. For example, if the third stage needs to perform an addition operation between the second calculation result and a bias coefficient, the third subunit 54c needs an addition arithmetic unit. For example, if the third stage needs to perform the calculation of several nonlinear functions, the third subunit 54c needs a multiplication arithmetic unit and an addition arithmetic unit. It should be understood that not all of the arithmetic units included in the third subunit 54c are necessarily used each time the third stage is executed. For example, the arithmetic operations required in the third stage may differ for different network layers in the neural network model. The control unit 52 only needs to control the corresponding arithmetic units in the third subunit 54c to perform the corresponding operations based on the computer instructions read from the instruction storage device 51, and some unnecessary arithmetic units can be bypassed.

[0108] In some embodiments, as shown in FIG. 5 , the control unit 52 includes an instruction decoder, a scheduler, and a manager. The instruction decoder is configured to decode the read computer instructions to obtain first-type instructions and second-type instructions, provide the first-type instructions to the scheduler, and provide the second-type instructions to the manager. The scheduler is configured to control at least one neural network processing unit based on the first-type instructions. The at least one neural network processing unit is further configured to perform an arithmetic operation in response to control of the scheduler. The manager is configured to control each register included in the quantum error correction hardware decoder based on the second-type instructions. Each register included in the quantum error correction hardware decoder is configured to perform a read or write operation in response to control of the manager.

[0109] Instructions fetched from the instruction storage device 51 are decoded by the control unit 52 and then assigned to a scheduler or manager depending on the instruction type. The instruction decoding here is similar to instruction decoding in classical computers (i.e., the fetch decode stage in classical computers), and extracts data contained in the instruction, such as the opcode, address, and operands, for subsequent processing. Some basic network information (e.g., the number of parameters of the neural network model, the addresses where the parameters are stored, and the addresses where the results are stored) is also stored in the decoder's on-chip memory in advance and accessed by the control unit 52 at runtime.

[0110] Here, the scheduler receives instructions for computational execution (i.e., the first type of instructions discussed above) and uses a finite state machine (FSM) to determine the specific operations to be performed in the neural network processing unit 54. The manager schedules communication between the registers of each stage of the neural network processing unit 54 and the input / output data collectors.

[0111] Managing the different instructions with a scheduler or manager allows the neural network processing unit 54 to perform arithmetic and read / write operations in parallel, thereby improving the efficiency of quantum error correction.

[0112] Compared to classical processors, error decoding is a static process, and the number of execution cycles for the entire neural network processing unit 54 and the computational division of each network layer can be determined in advance. Therefore, a very long instruction word (VLIW) method can be selected to reduce instruction execution delay. The control instructions of the programmable decoder of this application can be classified into two groups: computation class instructions and memory transfer class instructions, i.e., the aforementioned first and second type instructions. These two groups of instructions are used to instruct the scheduler and manager, respectively. Therefore, the design of the control instructions essentially represents a method for operating the configurable FSM in the control unit. The reason for using different group schedule instructions is that the memory read delay can be overlapped with the execution time of the neural network processing unit 54, thereby reducing the overall delay.

[0113] It is also considered that the matrix-vector calculations in a single network layer may be too large to be performed in a single subprocess. Therefore, the input data of a larger network layer is divided into multiple blocks and calculated in a scheduled order according to the instructions. For this purpose, the third stage may implement an accumulator to accumulate the execution results of different blocks. Depending on the instructions, calculations of network layers that do not require block division may bypass this accumulator. Correspondingly, in some cases, multiple small network layers may be processed in parallel within one neural network processing unit 54, and the prefetch mechanism of the second stage described above enables such parallelism.

[0114] An embodiment of the present application provides a hardware implementation architecture for a neural network-based quantum error correction decoding algorithm. The hardware architecture is a programmable architecture, and computer instructions for implementing the decoding algorithm can be pre-stored in an instruction storage device. During the real-time decoding process, a control unit reads the computer instructions and controls a neural network processing unit based on the computer instructions to decode error syndrome information of the quantum circuit based on a neural network model, ultimately obtaining error information. The hardware decoding architecture can efficiently implement the quantum error correction decoding algorithm based on a neural network model, thereby ensuring decoding performance and significantly reducing decoding delay. Furthermore, in the hardware decoding architecture, the number of neural network processing units and the number of arithmetic units included in the neural network processing unit can be designed and expanded according to actual needs, so that the hardware decoding architecture provided in the present application has better scalability and flexibility.

[0115] In addition, by dividing the decoding process of the neural network model into multiple sub-processes that are executed sequentially, and different sub-processes reuse the neural network processing units, the resource utilization of the decoder is improved, resource consumption is reduced when extending to large-scale error correction decoding, and flexibility requirements can be met when facing different experimental environments.

[0116] In the above embodiment, the decoding process of the neural network model is divided into multiple sub-processes that are executed sequentially, and different sub-processes can reuse the arithmetic units in the neural network processing unit 54 .

[0117] In some embodiments, the neural network model is a multi-task classification neural network model, and the neural network model is obtained by integrating multiple classification neural networks, and any one of the multiple classification neural networks is used to determine a decoding result corresponding to the error syndrome information. Illustratively, the multiple classification neural networks can share some network parameters.

[0118] Optionally, when divided according to function, the neural network model includes a feature encoder and an error decoder, and the working processes of the feature encoder and the error decoder belong to different sub-processes. Here, the feature encoder is used to perform feature extraction on the error syndrome information to generate feature information of the error syndrome information. The feature information represents the distribution feature of the error syndrome information in the feature space. The error decoder is used to decode the feature information to obtain a decoding result of the error syndrome information.

[0119] Illustratively, the working process of the feature encoder can be divided into at least one sub-process, and the working process of the error decoder can be divided into at least one sub-process.

[0120] Illustratively, the error decoder of the neural network model includes multiple types of decoders, and a specific type of decoder among the multiple types of decoders is used to determine whether a specific error type (e.g., a specific X-type error syndrome or a specific Z-type error syndrome) has occurred based on the feature information. The feature encoder and any one type of decoder are combined to form a complete classification neural network. That is, in this embodiment, the multiple classification neural networks share the network parameters of the feature encoder. This method can reduce the parameter size of the neural network model and the amount of calculation in the process of determining the error type and the qubit in which the error occurred.

[0121] For example, the feature encoder may be referred to as a front end or a feature extraction network. For example, the feature encoder may include multiple cascaded feature extraction subnetworks and a feature fusion subnetwork. The role of the feature extraction subnetworks is to extract local feature information in a divide-and-conquer manner, and the feature fusion subnetwork finally aggregates and compresses all the local feature information to obtain final feature information.

[0122] For example, the feature sub-network can be constructed based on a 3D CNN, and the feature fusion sub-network can be constructed based on a fully connected network. For example, the feature fusion sub-network includes one or two fully connected layers. For the feature information of the error syndrome information, the feature information is at least

number

[0123] Optionally, in the process of generating feature information, block-partitioned feature extraction can be performed on the error syndrome information. That is, when extracting feature information using a feature extraction subnetwork, block-partitioning is performed on the error syndrome information, dividing it into multiple small blocks, and then performing feature extraction on each small block. That is, block-partitioned feature extraction refers to dividing input data into blocks to obtain at least two blocks, and then using at least two feature extraction units to perform feature extraction processing on the at least two blocks in parallel. Here, the at least two blocks and the at least two feature extraction units correspond one-to-one, and each feature extraction unit is used to perform feature extraction on one block, and the number of blocks and feature extraction units is the same. Furthermore, the at least two blocks perform feature extraction in parallel, i.e., simultaneously, which helps to shorten the time required for feature extraction. When performing feature extraction on at least two blocks in parallel, the two blocks must be controlled by different processors. The feature processing unit refers to a structure within the neural network model used for control by different processors. For a specific process in which multiple processors operate in parallel, see the following examples.

[0124] Illustratively, the type decoder may also be referred to as a backend or a feature decoding network. For example, the type decoder may include an FFN layer, and the type decoder may be used to generate a decoding result corresponding to the error syndrome information.

[0125] After the decoding results generated by the multiple type decoders are obtained, error information is calculated based on these decoding results (i.e., the output results of the neural network model).

[0126] In some embodiments, the process of training the neural network model includes at least some of the following steps:

[0127] 1. Obtain sample error syndrome information and sample error result information corresponding to the sample error syndrome information.

[0128] Optionally, the sample error syndrome information is randomly generated syndrome information (which may be X syndrome or Z syndrome alone, or may include these two types of syndromes simultaneously), and the sample error result information corresponding to the sample error syndrome information is the error result corresponding to the syndrome information, which is obtained after one-hot encoding. It should be noted that different error results may correspond to the same syndrome information. By maintaining the diversity of output results at the output end, the model can finally learn an output probability distribution based on specific syndrome information during the training process.

[0129] 2. The neural network model to be trained determines the predictive decoding results for each of the error decoders based on the sample error syndrome information.

[0130] 3. Determine loss function values ​​corresponding to a plurality of feature decoding networks according to the predicted decoding results and sample error result information respectively determined by a plurality of error decoders. Optionally, the loss function values ​​can be calculated based on cross entropy.

[0131] 4. Determine a total loss function value based on the loss function values ​​corresponding to each of the plurality of error decoders.

[0132] 5. Adjusting parameters of the neural network model to be trained based on the total loss function value to obtain a trained neural network model. Optionally, storing the trained parameters of the neural network model in a quantum error correction hardware decoder. In some embodiments, the process of training the neural network model to be trained is performed by another computer.

[0133] In some embodiments, the arithmetic units in the neural network processing unit 54 may be allocated and used in the following manner: The neural network processing unit includes multiple arithmetic units, which are allocated to perform calculations on different network layers in the neural network model. The number of arithmetic units allocated to each network layer in the neural network model is determined based on the total number of arithmetic units included in the neural network processing unit and the number of operations included in each network layer. In this allocation method, dedicated arithmetic units are allocated to different network layers, and adjacent network layers are connected by hardwired connections. That is, the input and output of each adjacent network layer are directly connected by a specific data transmission line. This is an excellent solution for scenarios where the surface code scale is relatively small (e.g., L=5 or 7) and the network architecture of the neural network model is fixed. The above scenario is called a "specific network architecture."

[0134] For a "specific network architecture", this application proposes a resource allocation model to optimize decoding performance. Let C be the total number of arithmetic units included in the neural network processing unit, and N be the number of network layers included in the neural network model. l and the j-th layer of the neural network model is M j Assume that there are multiplication operations. The number of operation units allocated to the j-th layer of the neural network model is C jand define this problem as C j This is simplified to a constrained optimization that partitions the

number

[0135] In some embodiments, the quantum error correction hardware decoder may employ a single-core architecture or a multi-core architecture, where single-core architecture and multi-core architecture refer to the number of processors (or processing cores, also called cores) included.

[0136] When adopting a single-core architecture, the quantum error correction hardware decoder includes a single processing core, and the entire decoding algorithm is performed by the single processing core. When the surface code scale L is small, the single-core architecture can withstand this computational complexity, but when the surface code scale L is large, the single-core architecture has its limitations. Therefore, this application proposes a multi-core architecture solution.

[0137] In some embodiments, the quantum error correction hardware decoder includes multiple processing cores, each of which is used to process a portion of the operations of the neural network model, and two processing cores having a data dependency relationship are connected by one or more data lines. In the embodiments of the present application, in a multi-core architecture, the number of processing cores included in the quantum error correction hardware decoder is not limited, and can be specifically designed taking into account the size of the surface code scale L and the computational complexity, and the entire decoding algorithm can be performed on the premise that the computing power of each processing core is fully utilized.

[0138] When a multi-core architecture is adopted, any two processing cores that do not have a data dependency relationship have parallelism, thereby maximizing the computational power of each processor and shortening the decoding time. Furthermore, any two processing cores that do not have a data dependency relationship can execute sequentially. For example, FIG. 7 shows a schematic diagram of a multi-core architecture. Processors 1 to p may not be connected to each other, and these p processors can execute in parallel, for example, using a division and processing method to process different blocks of error syndrome information in parallel. The calculation results obtained by processors 1 to p are sent to processor p+1, which performs further calculations on the calculation results provided by processors 1 to p. The calculation results of processor p+1 are then input to processors p+2 to N, and processors p+2 to N may not have a data dependency relationship with each other, and these multiple processors can execute in parallel. Finally, a certain processor collects the output results of the entire neural network model, and error information can be determined based on the output results.

[0139] Here, multiple processing cores form a tree-like structure, with each processing core responsible for a portion of the computations in the 3D CNN, or fully connected network. The hierarchical 3D CNN employed here allows the inputs of different processing cores to be largely independent, requiring almost no complex communication between processing cores except for input-output data transmission. The same is true for the fully connected network used for classification. Therefore, when each processing core is fully utilized, this architecture maintains a constant decoding latency by adding more cores to expand the parallel computation scale.

[0140] In such a multi-core architecture, only input and output data transmission between the networks of each layer is required, and this data exists in the form of vectors. Because the amount of data in these vectors is not very large, the latency of a single data transmission is more important than the data transmission bandwidth. This architecture uses a method based on the LVDS standard to establish data transmission between different processing cores. First, complex serial-to-parallel conversion protocols are not added to the input and output terminals. Because high-speed communication interfaces generally transmit serial data, the output terminal must convert logical parallel data to serial data and transmit it, while the input terminal must convert the received serial data to parallel data. This requires serial-to-parallel conversion protocols commonly used in the industry, such as a general SerDes IP. However, these protocols are relatively complex to operate and incur large latency overhead, making them unsuitable for this scenario. In the real-time error correction decoding scenario described in this application, a training code method is used to pre-train the alignment between each data line. At the same time, the data rate is maintained at an appropriate level, e.g., 1G to 1.5G, to avoid data rates that are too high and make synchronization difficult. Furthermore, the number of channels between the data paths of each group can be increased. The data path here is the communication interface mentioned above, and in Figure 7, it is the black arrow between each processor. The number of channels is the number of data lines arranged in parallel on one data path. The method we use is to send multiple sets of serial data in parallel using multiple data lines. By controlling the data rate, the data on each data line does not need to go through the complex communication protocol mentioned above. Instead, we simply use a training code to pre-train these data lines before official communication to ensure data alignment. This method avoids the delay overhead of the communication protocol and guarantees overall bandwidth by using multiple lines in parallel.

[0141] Measurements have shown that this data transmission method can achieve very low transmission latency, enabling the scalability of multi-core architectures. Even as the surface code scale L increases and the amount of data that needs to be transmitted between cores increases, the Hamming weights of these vectors can be expected to remain relatively low by utilizing the sparsity of the syndrome bits. In this case, using a compression method to reduce the amount of data between each transmission group can further reduce latency. Specifically, the Hamming weights within the data are first calculated, and if the Hamming weight is determined to be below a certain threshold, the data is sent to the compression module. The compression module can incorporate multiple data compression methods. For example, dynamic zero compression divides data into k blocks and uses k flag bits to indicate whether the corresponding block is all zero. If it is zero, the corresponding position is set to 1; if it is non-zero, it is set to 0. The data of that block is then added to the compressed result, where k is an integer greater than 1. Since different compression methods have different effects on different types of transmission data, the final stage of the compression module compares the compression rates of different compression methods and selects the method with the highest compression rate to be used for the final data communication.

[0142] The following describes experimental data related to the technical solutions of the present application.

[0143] First, we implemented decoders with L=5 and L=7 using the specific network architecture described above on an FPGA platform and verified the delay results. The configuration of the neural network model used and the delay results are shown in Table 1 below. [Table 1] Here, L represents the surface code scale, T represents the number of syndrome measurements, and N represents the number of parameters of the neural network model.

[0144] The delay results of the three decoders implemented are all smaller than the expected delay value (1.5 us), which satisfies the real-time error correction requirements of modern superconducting quantum computers. Here, the L=7 decoder is the largest and has the smallest delay among currently known hardware implementations of a real-time decoder.

[0145] The programmable decoder described in the above embodiment was also implemented using the same hardware platform. Figure 8 is a schematic diagram of the resource consumption of a programmable decoder and a specific network architecture. The top three figures show the overall FPGA situation, and the bottom table shows the consumption of various resources. It can be seen that for two major hardware resources, the DSP (Digital Signal Processor) and the ALM (Adaptive Logic Module), the programmable microarchitecture in the embodiment of this application reduces resource consumption by 2.4 times and 3.0 times compared to the specific network architecture. At the same time, unlike the specific network architecture, which requires separate hardware implementations for each different neural network configuration, the programmable decoder can implement various different network configurations with the same hardware implementation, thereby meeting the flexibility requirement and becoming the optimal choice for large-scale neural network decoders.

[0146] An exemplary embodiment of the present application further provides a chip, in which the quantum error correction hardware decoder provided in the above embodiment is disposed, and optionally, the chip is an FPGA chip or an ASIC chip.

[0147] In this specification, "plurality" refers to two or more than two. The term "and / or" describes only a relational relationship and indicates that three relationships may exist. For example, A and / or B can indicate three cases: A exists independently, both A and B exist, and B exists independently. The symbol " / " typically indicates that the relation between related objects is an "or" relationship. Furthermore, the numbering of the steps described in this specification merely exemplifies one possible execution order between the steps. In some other embodiments, the steps may not be executed in numerical order. For example, two differently numbered steps may be executed simultaneously, or two differently numbered steps may be executed in the reverse order of the illustrated steps. The embodiments of the present application are not limited thereto.

Claims

1. A quantum error correction hardware decoder, comprising: an instruction storage device; a control unit; at least one neural network processing unit; and an error batch search unit; the instruction storage device is configured to store computer instructions; the control unit is configured to read the computer instructions from the instruction storage device and control the at least one neural network processing unit based on the computer instructions; the at least one neural network processing unit is configured to, in response to control by the control unit, decode error syndrome information of a quantum circuit based on a neural network model to obtain an output result of the neural network model, the error syndrome information referring to an error syndrome obtained by performing an error measurement on the quantum circuit; The error batch search unit is configured to determine error information based on the output result, and the error information is for indicating a quantum bit in which an error has occurred in the quantum circuit and an error type corresponding to the error.

2. the decoding process of the neural network model is divided into n sub-processes that are executed sequentially, different sub-processes reuse the neural network processing units, and n is an integer greater than 1; 2. The quantum error correction hardware decoder of claim 1.

3. Each sub-process includes a first stage, a second stage, and a third stage; the first stage is used to perform a multiplication and addition operation on input data of the sub-process to obtain at least one first calculation result; the second step is used to perform an addition operation on the at least one first calculation result to obtain at least one second calculation result; the third step is used to obtain output data of the sub-process based on the at least one second calculation result; 3. The quantum error correction hardware decoder of claim 2.

4. the neural network processing unit comprises a first subunit, a second subunit, and a third subunit; the first subunit includes at least one arithmetic unit for performing the first stage; the second subunit includes at least one arithmetic unit for performing the second stage; the third subunit includes at least one arithmetic unit for performing the third step; 4. The quantum error correction hardware decoder of claim 3.

5. The first subunit further comprises a data operation unit and a weight operation unit, wherein the data operation unit is configured to read input data used in the first stage from a first register, and the weight operation unit is configured to read weight parameters of the neural network model used in the first stage from a second register.

5. The quantum error correction hardware decoder of claim 4.

6. The second subunit further includes at least one multiplexer configured to prefetch calculation results at different depths in a calculation unit tree corresponding to the second stage.

5. The quantum error correction hardware decoder of claim 4.

7. input data of a first sub-process among the n sub-processes includes the error syndrome information, and output data of a last sub-process among the n sub-processes includes the output result; 3. The quantum error correction hardware decoder of claim 2.

8. the control unit includes an instruction decoder, a scheduler, and a manager; the instruction decoder is configured to decode the read computer instructions to obtain first type instructions and second type instructions, provide the first type instructions to the scheduler, and provide the second type instructions to the manager; the scheduler is configured to control the at least one neural network processing unit based on the first type instructions; the at least one neural network processing unit is further configured to perform arithmetic operations in response to control of the scheduler; the manager is configured to control each register included in the quantum error correction hardware decoder based on the second type instruction; each register included in the quantum error correction hardware decoder is configured to perform read and write operations in response to control of the manager; 2. The quantum error correction hardware decoder of claim 1.

9. the quantum error correction hardware decoder further includes a first register and a second register; the first register is configured to store input data used in a decoding process of the neural network model and to store output data generated in the decoding process of the neural network model; the second register is configured to store weight parameters and bias parameters of the neural network model; 2. The quantum error correction hardware decoder of claim 1.

10. the neural network processing unit includes a plurality of arithmetic units, the plurality of arithmetic units being assigned to perform calculations of different network layers in the neural network model, and the number of arithmetic units assigned to each of the network layers in the neural network model being determined based on the total number of arithmetic units included in the neural network processing unit and the number of calculations included in each of the network layers; 2. The quantum error correction hardware decoder of claim 1.

11. the quantum error correction hardware decoder includes a plurality of processing cores, each processing core being used to process a part of the operations of the neural network model, and two processing cores having a data dependency relationship being connected by one or more data lines; 2. The quantum error correction hardware decoder of claim 1.

12. A chip on which the quantum error correction hardware decoder according to any one of claims 1 to 11 is arranged.

13. The chip is a field programmable gate array (FPGA) chip or an application specific integrated circuit (ASIC) chip. The chip of claim 12.

Citation Information

Patent Citations

  • Arithmetic processor and arithmetic processing method

    JP2009080693A

  • Audio signal processing in high frequency reconstruction.

    JP2022141919A

  • Neural network-based quantum error correction decoding method, device, chip, computer device, and computer program

    JP2022532466A

  • Syndrome Data Compression for Quantum Computing Devices

    JP2022543085A