Quantum decoder
The method simplifies indexing for quantum error correction decoders by reducing the number of bits required for representing and transmitting execution schemes, addressing storage and bandwidth challenges, and enabling flexible reconfiguration for improved scalability and efficiency in quantum computing systems.
Patent Information
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-03
- Publication Date
- 2026-03-11
AI Technical Summary
Current quantum computing systems face challenges in efficiently decoding error correction codes due to high memory and bandwidth requirements for indexing and storing execution schemes of processing elements, which hinder parallelization and increase storage overhead, limiting the scalability and flexibility of quantum error correction methods.
A computer-implemented method and apparatus that uses simplified indexing variables to identify processing elements in a quantum error correction decoder, reducing the number of bits required for representing and transmitting execution schemes, thereby lowering bandwidth and storage overhead, and enabling flexible reconfiguration for various decoding tasks.
The method reduces memory and bandwidth needs, allowing for efficient parallel processing and reconfiguration of quantum error correction algorithms, enhancing the scalability and flexibility of quantum computing systems.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
This disclosure relates to methods, systems, and apparatuses for use in decoding errors in a quantum computer system. Background Quantum computers hold the potential to revolutionize various fields of science and technology. However, today's quantum computers cannot realise these transformational possibilities because their fundamental components, qubits, are highly error-prone. To unlock the transformative possibilities of quantum computing, the error rates for operating on quantum data must be dramatically reduced. This issue may be addressed in part by hardware improvements as physicists and engineers get better at building more stable qubits, but these advancements alone won't be enough to enable algorithms to run with millions or billions of operations reliably. There is a need for quantum error correction methods, however these methods present significant challenges. Fault-tolerant quantum computation will likely involve millions of physical data qubits, generating a huge amount of data which must be processed extremely quickly and in real-time. Both classical and quantum error correction methods seek to identify and address errors in information storage and processing. However, these general approaches operate in fundamentally different ways due to the unique principles of quantum mechanics. For example, while classical information is encoded using classical bits which exist in a well-defined state (either 0 or 1) at any given time, quantum information is encoded using a register of quantum devices such as a register of qubits, which can exist in superpositions of states (both 0 and 1 simultaneously) and exhibit entanglement. In addition, since directly measuring the state of a qubit causes the state to collapse, quantum error correction techniques must measure certain properties of the quantum system to infer the existence of errors, while mitigating the effects of measurements on the quantum state. Some quantum error correction codes involve decoding a "syndrome", which can be considered to be a signature associated with an error state of the physical qubits. A decoder is used to identify an error (or errors) which could have caused the syndrome. A "decoding graph" can be used to facilitate decoding of the syndrome by grouping "defects" in the syndrome. These defects generally provide an indication of end points of chains of errors on physical data qubits in the error correction code. Errors generally create a pair of defects, and an objective of quantum error correction techniques is to identify and match up these defects. After the defects in the syndrome have been grouped, a next step can be determined. For example, a correction for the error state can be determined and applied. Decoding algorithms designed to operate on this type of data can be run by decoding hardware comprising a plurality of processing elements (PEs). It is beneficial to parallelise the execution of multiple PEs at some or all steps of the decoding algorithm, for example using a distributed decoding algorithm. Distributed decoding algorithms comprise at least some parallelisation in the execution of PEs. Often these PEs are connected using interface links in hardware. Each node in the decoding graph may be associated with a different PE according to a 1:1 mapping, in order to enable parallelisation. This grid of PEs can be referred to as a processing grid. Connections are created between neighbouring PEs, i.e. PEs associated with nodes which neighbour each other in the decoding graph, via interface links in the processing grid. These interface links enable a PE to communicate, i.e. query and share data, with other PEs in the processing grid. In this way, data can be propagated through the network of PEs in the processing grid. The execution of the PEs can be implemented using a processing schedule or execution scheme, detailing which PEs will be processed (or executed) at a given time. In current implementations, a PE in the processing grid may be recognised in the execution scheme by a unique index number or index variable. For example, in a 6x6 grid of PEs, the PE's may each be assigned an index between 0-36. This approach requires multiple bits in each index. Storing an execution scheme in which PEs are indexed in this way requires a large amount of memory and, in turn, a large amount of bandwidth when the execution scheme is uploaded or transferred between devices, e.g. at run-time. Another drawback to the current implementations is the associated large storage overhead of the decoding hardware. Decoding algorithms require information to be stored at each Pe regarding certain other PEs, logic, and scheduling information, and each PE has a dedicated memory to enable storage of this information. The number of bits required to store PE indices has an impact on the memory requirements for each PE, which in turn has a significant impact on the size and cost of the decoding apparatus. The present application seeks to address these and other disadvantages encountered in the prior art by providing an improved computer-implemented quantum error correction method for decoding a quantum error correction code from a quantum computer, and an improved system suitable for implementing such a method. Summary According to an aspect of the invention, a computer-implemented method is provided, for controlling the execution of a plurality of distributed processing elements, PEs, in a decoder apparatus of a quantum computer system. The quantum computer system further comprises a register of quantum devices. The method comprises controlling the plurality of PEs to perform a distributed decoding algorithm according to an execution scheme. The execution scheme identifies which PEs of the plurality of PEs should be executed at different stages of the distributed decoding algorithm. Each PE is identifiable within the execution scheme using one or more indexing variables. Each indexing variable of the one or more indexing variables is comprised of one or more bits, with each bit being associated with the execution of one or more PEs of the plurality of distributed PEs. With each bit of the one or more indexing variables being associated with the execution of one or more PEs in the execution scheme, the execution scheme of PEs is simplified compared to prior methods. Fewer bits are required to represent an execution scheme of PEs, meaning that fewer bits are needed to stored and / or transmitted at runtime. Having fewer bits at run time allows for the use of lower bandwidth communication in the decoder apparatus. Furthermore, fewer bits needing to be stored in the decoder apparatus leads to lower storage overhead. The present method also allows for flexibility in reconfiguration. The ability to easily reconfigure the execution of PEs is useful for new-age distributed decoders used in quantum computers. The flexibility allows for highly configurable architectures for quantum compilers and opens avenues to performing lattice surgery logical operations with distributed decoders. Furthermore, flexibility in reconfiguration allows for a configurable decoder that can address various surface-code patch sizes via resizing. Optionally, each bit in the one or more indexing variables is associated with the execution of one or more PEs, independent of the other bits in the indexing variable. Optionally, each PE is uniquely identifiable within the execution scheme using the one or more indexing variables. Optionally, each bit of the one or more indexing variables is associated with a different respective subset of the plurality of distributed PEs. Optionally, the one or more indexing variables comprise a first indexing variable and a second indexing variable. Optionally, a particular PE is identifiable within the execution scheme via the combination of a first bit in the first indexing variable and a second bit in the second indexing variable. The first bit may be associated with a first subset of PEs and the second bit may be associated with a second subset of PEs. The particular PE may be present in both the first and the second subset of PEs. Optionally, the one or more indexing variables comprise a first dimension indexing variable and a second dimension indexing variable. The number of bits in the first dimension indexing variable may correspond with a number of PEs in a first dimension, and the number of bits in the second dimension indexing variable may correspond with a number of PEs in a second dimension. Optionally, the decoder apparatus further comprises a controller coupled to each PE of the plurality of PEs. The method may further comprise executing, by the controller, one or more PEs of the plurality of distributed PEs according to the execution scheme. Optionally, each bit of the one or more indexing variables is associated with a different respective subset of the plurality of distributed PEs. Each bit may further identify a control line which couples the controller to its associated subset of PEs. Optionally, the controller controls which PEs are activated according to the execution scheme thereby enabling the plurality of PEs to perform the distributed decoding algorithm. Controlling the PEs may comprise determining which control lines are identified by the one or more indexing variables in the execution scheme at a particular stage of the decoding algorithm and activating the identified control lines. Control lines connect one or more PEs to one or more controllers, allowing communication between the one or more PEs and the one or more controllers. Optionally, each PE of the plurality of PEs is coupled to the controller via a different control line, and the one or more indexing variables is a single indexing variable comprised of a number of bits equal to the number of PEs in the plurality of PEs. Each bit in the single indexing variable may uniquely identify a respective PE. Optionally, the computer-implemented method comprises receiving, at the decoder apparatus, syndrome data representative of an error state of the quantum devices in the register of quantum devices. The syndrome data may comprise a plurality of defects. The syndrome data may be representable as a decoding hypergraph comprising a plurality of nodes connected by hyperedges representing error mechanisms associated with the plurality of quantum devices. Each PE of the plurality of distributed PEs may be associated with one or more nodes of the decoding hypergraph. Performing the distributed decoding algorithm may comprise determining a correction for the error state. Optionally, the computer-implemented method comprises measuring a logical state encoded in the quantum devices of the quantum computer to obtain a logical state measurement and applying the correction for the error state to the logical state measurement. Optionally, the decoding algorithm is a clustering algorithm, and the clustering algorithm grows and merges clusters of nodes based on the number of defects in each cluster until a final cluster state is reached. Determining a correction for the error state may be based on the final cluster state. According to another aspect of the present invention, a decoder apparatus is provided, comprising a controller, a plurality of distributed processing elements, PEs, and computer memory. The computer memory may store an execution scheme which, when implemented by the controller, controls which PEs of the plurality of PEs are executed at different stages of a distributed decoding algorithm. Each PE may be identifiable within the execution scheme using one or more indexing variables, wherein each indexing variable of the one or more indexing variables is comprised of one or more bits, with each bit being associated with the execution of one or more PEs of the plurality of distributed PE. The computer memory may store instructions which, when implemented by the controller, cause the controller to perform the method of any preceding claim. According to another aspect of the present invention, a quantum computer system is provided, comprising a register of quantum devices and a decoder apparatus according to any of the methods set out above or disclosed herein. According to another aspect of the present invention, a computer-readable medium is provided, comprising instructions which, when executed by a controller of a decoder apparatus, cause the controller to perform the methods set out above or disclosed herein. Figures Specific embodiments are now described, by way of example only, with reference to the drawings, in which: Figures la-e show an example of a process for decoding a patch of surface code using a clustering decoder; Figure 2 is a flowchart depicting a local clustering decoder algorithm according to the present disclosure; Figure 3a-h show an example of a process for decoding a patch of surface code using a local clustering decoder algorithm; Figure 4 show a graph depicting a processing grid of processing elements; Figure 5 depicts a decoder apparatus according to the prior art; Figure 6 is a graph depicting a method according to the present disclosure; Figure 7 is a graph depicting a method according to the present disclosure; Figure 8 depicts a graph depicting a processing grid of processing elements; Figure 9 depicts a decoder apparatus; Figure 10 depicts a graph depicting a processing grid of processing elements; Figure 11 depicts a decoder apparatus; Figure 12 is an example embodiment of a quantum computer system; Figure 13 is an example embodiment of a computer program product. Detailed Description In overview, and without limitation, the application discloses a computer-implemented method for controlling the execution of a plurality of distributed processing elements, PEs, in a decoder apparatus of a quantum computer system. In approaches of the present disclosure, the execution of PEs is encoded in an execution scheme. The execution scheme may state which PEs are to be executed, or "activated", at each stage of a distributed decoding algorithm. The execution scheme comprises bits that are used to identify which PEs are to be executed. Each PE is identifiable in the execution scheme by one or more indexing variables and each indexing variable is represented by one or more bits. For example, if the PEs are arranged in a square grid, each PE can be assigned separate indexing variables for rows and columns, where each bit in the indexing variable is associated with a row in the grid and / or a column in the grid. Each bit being "on" or "off" in the row / column indexing variable therefore represents PEs in a row / column being switched "on" or "off". Combining these indexing variables then makes it possible to index PEs individually. Methods of the present disclosure require fewer bits to represent the execution scheme of PEs in a decoding algorithm compared to prior methods. This means that fewer bits need to be stored and / or fewer bits need to be transferred at run-time. The use of fewer bits leads to improvement in the bandwidth requirements of the decoder, as the execution scheme can be efficiently transferred at run-time. In addition, the methods of the present disclosure allow an improved storage overhead. A low storage overhead is possible due to the small amount of memory needed to store the execution scheme. Due to the low bandwidth and low storage overhead associated with the present method, reconfiguring the decoder for different decoding problems becomes simpler. This fast reconfiguration makes the present disclosure suitable for future quantum logical operations, for example performing logical gates on logical qubits. Several terms and concepts will be discussed in the present application, and though the skilled person will be familiar with this terminology and these concepts, the following section serves to provide additional context to the reader. A quantum computing system (also referred to herein as a quantum computer or quantum computer system) is a computing system that exploits quantum mechanical phenomena (i.e. using quantum devices). The quantum devices may be any quantum devices capable of storing quantum information (i.e. any devices suitable for encoding information using quantum computational states). The quantum devices may be qubits. Alternatively, the quantum devices may be other devices capable of storing quantum information, such as qudits or qutrits. While the description herein will primarily refer to qubits, any reference herein to qubits should be understood to also encompass other types of quantum devices unless explicitly stated otherwise. Building a useful fault-tolerant quantum computer will require quantum error correction hardware that can receive and process enormous amounts of error information (i.e. syndrome data) in real-time almost instantaneously. A delay in decoding can lead to the creation of a backlog that grows exponentially with the size of the computation, which will ultimately lead to failure of the quantum computation. The speed of the decoder acts as a bottleneck to the number of qubits in a quantum error correction code (and therefore also as a bottleneck to reducing logical error rates). Improvements to decoding hardware and algorithms help to prevent this backlog, thereby enabling quantum computers with higher numbers of qubits and lower error rates because faster decoders can handle quantum error correction codes involving more data qubits, and using more data qubits leads to a reduction in logical error rates when performing fault-tolerant quantum computation. Quantum error correction codes generally involve decoding a "syndrome", which can be considered to be a signature associated with an error state of the physical qubits in the code (two different errors can potentially have the same syndrome). A "decoder" is then used to identify an error which could have caused the syndrome (or possibly just a correct operation that can be used to correct the logical qubit states encoded in the code e.g. a single bit representing whether the eventual logical measurement outcomes need to be flipped). Decoding hypergraphs (especially decoding graphs) are used in many error correction codes (such as topological error correction codes including the surface code) to facilitate decoding of the syndrome by grouping "defects" in the syndrome (these defects generally provide an indication of end points of chains of errors on physical data qubits in the error correction code). A decoding hypergraph is a hypergraph (in the mathematical sense) comprising hyperedges representing error mechanisms, and nodes (or vertices) representing differences in successive syndrome measurements (or more generally, a decoding hypergraph comprises nodes representing detectors, which are measurement results that sum to zero (e.g. modulo 2 sum) during perfect (i.e. error-free) operation of the quantum computing system). The decoding hypergraph may have one or more boundaries, which involve hyperedges extending beyond the hypergraph (e.g. to one or more virtual boundary nodes), i.e. the decoding hypergraph may be a sub-hypergraph of a larger hypergraph including the virtual boundary node(s)). Conventionally, some error correction literature has referred to "rough" boundaries and "smooth boundaries". However, the "smooth" boundaries in such literature are not actually boundaries in the above sense, and they will not be referred to as boundaries in the present disclosure. Accordingly, the boundaries referred to herein are synonymous with the "rough" boundaries in such literature. Clustering algorithms decode syndromes by clustering defects into groups. A cluster of decoding graph edges is formed around each defect (defects are located at a subset of nodes of the decoding graph), and these clusters are grown by including additional edges until all defects are included in a cluster containing an even number of defects (or the cluster touches a boundary of the decoding graph), with overlapping clusters being merged after each growth stage. Correction(s) to the encoded logical state can then be determined based on the clustering of defects. Since an objective of this category of quantum error correction approaches is to correct errors by grouping defects, some algorithms seek to decode syndromes by clustering defects. An example of this approach is described in (Delfosse, N. and Nickerson, N.H. (2021) 'Almost-linear time decoding algorithm for topological codes', Quantum, 5, p. 595. doi:10.22331 / q-2021-12-02-595). Every 'odd' cluster, i.e. every grouping of nodes (or 'vertices') that is associated with an odd number of defects, grows until it becomes an "even" cluster, for example a cluster associated with an even number of defects. The number of defects in the cluster is a primary factor which affects the "parity" of a cluster. Each node which is part of an odd cluster grows outward until it encounters another cluster, at which point the clusters merge. If the parity of the new, merged cluster is even, then it stops growing. Otherwise, it continues to grow. This continues until all clusters have an even parity. There may be other criteria which cause a cluster to stop growing, such as if a cluster touches a boundary of the decoding graph. Next steps can then be determined based on the decoded syndrome. For example, correction(s) to the encoded logical state can be determined based on the clustering of defects. This overview has been simplified for brevity and to aid quick understanding, and the limitations of this simplified overview will be understood by the skilled person. A syndrome (also referred to as syndrome data) is a collection of values (e.g. measurement values, generally based on qubit measurements, in particular syndrome qubit measurement) representative of an error state of physical data qubits in the quantum computer. The syndrome may also include decoding hypergraph location information for each syndrome value - e.g. a coordinate or index value. Syndrome data may be obtained by measuring a plurality of syndrome qubits (e.g. surface code stabiliser measurements). The decoder may receive the syndrome data as raw (e.g. analogue) measurement data, or the syndrome data may be pre-processed (e.g. processed into digital form by a control system). A defect (also referred to as an excitation or measurement event) generally represents the end of a chain (or hyperchain) of errors in the decoding hypergraph (the chain of errors may span both space-like and temporal dimensions of the decoding hypergraph). Defects are non-trivial syndrome values, and they may correspond to a change in value of a syndrome qubit measurement outcome between successive rounds of syndrome measurement. A decoding system (also referred to herein as a decoder or quantum error decoding system) is a classical computing system that decodes syndromes and provides one or both of (i) possible error locations (i.e. which data qubits may have experienced an error), and (ii) a correction for the qubit error state. It is possible to determine a correction during decoding without determining error locations, and the correction may be a single bit representing whether a logical error has occurred. The correction can generally be tracked by a classical computer (e.g. by the decoder or a control system) and does not generally need to be applied to the quantum devices. The decoder may be a dedicated hardware device (e.g. implemented using an FPGA or ASIC or similar) or it may be a software component implemented using a CPU. The quantum error correction code may be a surface code (e.g. planar code) error correction procedure. Alternatively, the quantum error correction code may be any other error correction code that utilises a decoding hypergraph (e.g. a decoding graph), such as other topological quantum error correction codes. The decoding hypergraph may be a decoding graph, and the hyperedges may be edges. A hypergraph is a generalisation of a graph in which edges ("hyperedges") can be connected to more than two nodes (graphs are a specific type of hypergraph in which each edge connects to two nodes). The methods of the present invention apply equally to decoding hypergraphs. Accordingly, any reference herein to decoding graphs and edges should be understood to also encompass decoding hypergraphs and hyperedges respectively. The quantum computing system (e.g. a control system of the quantum computing system) may be further configured to measure a logical state encoded in the quantum devices and apply the correction to the measured logical state (the correction may be applied at the control system, at the decoding system or at some other subsystem of the quantum computing device, such as a device operating at the algorithmic / application layer of the quantum stack). The clusters may also be referred to as sets / groups / collections, or any similar term that refers to a grouping of defects. The quantum error correction method may be a surface code (e.g. planar code) error correction procedure. Alternatively, the quantum error correction method may be any other error correction procedure that can be decoded by grouping defects on a decoding graph, such as other topological quantum error correction codes. One skilled in the art will appreciate that the details of how the correction is identified, and how it is based on the clusters of defects, will depend upon the configuration of the error correction method in question. For example, identifying the correction for the error state may comprise: identifying a logical operator for the quantum error correction method involving physical qubits of the quantum computer system that are associated with the boundary of the decoding graph; determining a parity of a total number of clusters that the logical operator intersects that contain an odd total number of defects, wherein an even parity indicates that the logical operator is in a correct logical state and an odd parity indicates that the logical operator is in an incorrect logical state. The parity of a cluster can be described in terms of different bases, for example a parity can be discussed as being even / odd, or 0 / 1. These definitions may be used interchangeably herein. Parity can be discussed in terms of nodes and clusters, and may be used somewhat interchangeably. Initially, the parity of a node, and therefore the parity of its cluster, is dependent on whether the node is a defect or not. For example, a node that is a defect will have an odd parity and a node that is not a defect will have an even parity. As discussed above, the distributed decoding algorithm of the present disclosure may be a clustering algorithm. A clustering algorithm grows and merges clusters of nodes based on the number of defects in each cluster until a final cluster state is reached. The correction for the error state can then be determined based on the final cluster state in a known way. Figures la-e depict an example of a process for decoding a patch of surface code using a clustering decoder. The aim of clustering decoders is to cluster defects together into a set of decodable clusters. Several different variants of clustering algorithms exist, and one skilled in the art will appreciate that other clustering decoding algorithms may be used instead of or as well as that shown in Figures la-e. Figures la-e depict decoding graphs 110,120,130,140 and 150. Figures la-e depict stages of the clustering process, where Figure la is the first stage and Figure le is the last stage of this example clustering process. Conventional decoding algorithms primarily focus on error mechanisms that leave qubits in a computational basis state (e.g. some superposition of the |0) and |1) states), such as bit-flip and phase-flip errors. The circles on the decoding graphs 110,120, 130, 140, 150 depict nodes of the decoding graph. The nodes of the decoding graphs 110,120,130,140,150 correspond to syndrome qubits. With no errors, syndrome qubits are in the |0) state, as illustrated by the empty circles in Figures la-e. The hashed circles in Figures la-e correspond to defects, i.e. syndrome qubits in the |1) state. In Figures la-e, the defects are labelled as 112a-f in the decoding graphs. The numbering of the rows (from 0-6) and columns (from 0-7) of the nodes on the decoding graph is included for ease of reference. The edges of the decoding graphs 110,120, 130,140,150 correspond to data qubits, with weights determined by the error probability. The decoding graphs have some edges that only connect to a single node; these edges represent boundaries of the decoding graph. The boundaries can conceptually be considered to all connect to one or more virtual nodes / boundary nodes. One skilled in the art will appreciate that the nuances of the decoding graph will depend upon the error correction code being implemented, and that some decoding graphs (e.g. those for toric codes) do not have boundaries. The concept of boundaries can be extended to hyperedges, in which hyperedges at a boundary may connect to a virtual node in addition to one or more nodes of the decoding hypergraph. While the graphs in Figure la-e can be used as basic decoding graphs, it is also possible to use more complex decoding graphs, for example with an extra dimension representing time, in which there is not necessarily a one-to-one correspondence between data qubits and decoding graph edges. One skilled in the art will appreciate that the physical qubits do not necessarily need to be physically arranged as shown in Figures la-e. Figure la shows a first graph 110, depicting an example decoding graph. In this example, there are 6 defects 112a-f in the decoding graph. Decoding algorithms such as clustering algorithms may be used to decode the defects 112a-f. The illustrated clustering algorithm begins by placing each defect 112a-f in its own cluster. The defects 112a-f may be referred to as the first defect 112a, second defect 112b, third defect 112c, fourth defect 112d, fifth defect 112e and sixth defect 112f. This naming convention is used for ease of reference, and not indicative of any order associated with the defects. Figure lb shows a second graph 120, depicting the result of a first stage of growth using the clustering algorithm. In graph 120, each cluster has grown out by a half-edge in each of the four directions of the decoding graph. The growth of each cluster is depicted by bold black lines along the edges of the graph. Each of these clusters has grown due to the presence of an odd number of defects in each cluster, specifically one in each cluster, in this case. In Figure lb, by growing each cluster out by half-edges of the graph, two pairs of clusters have connected; the cluster containing the first defect 112a is in contact with the cluster containing the second defect 112b, and the cluster containing the third defect 112c is in contact with the cluster containing the fourth defect 112d. Any clusters that are touching will merge into a larger cluster, such that there is a cluster containing the first defect 112a and second defect 112b, and another cluster containing the third defect 112c and fourth defect 112d, in this example. In each iterative round of the clustering algorithm, each cluster containing an odd number of defects may extend outwards by a half-edge of the graph, dependent on whether or not the cluster is connected to a boundary. Each cluster containing an even number of defects stops growing. As both of these clusters now contain an even number of defects, they will stop growing. As the clusters containing the fifth defect 112e and sixth defect 112f remain isolated (containing an odd number of defects), they will keep growing. Clusters with an even number of defects can be said to have an even parity, and clusters with an odd number of defects can be said to have an odd parity. Figure lc shows a third graph 130, depicting the result of a second stage of growth using the clustering algorithm. As the clusters containing the fifth defect 112e and sixth defect 112f contained an odd number of defects in graph 120, in the second growth stage both of these clusters grew by half an edge in each of the four connected edges of the decoding graph. After the second stage of growth, the cluster containing the fifth defect 112e still only contains one defect. However, the cluster containing the sixth defect 112f has reached the boundary at point (0, 2) on the decoding graph. A cluster that has reached a boundary will not grow further, and therefore the cluster containing the sixth defect 112f will stop growing. Figure Id shows a fourth graph 140, depicting the result of a third stage of growth using the clustering algorithm. After the third stage of growth, the cluster containing the fifth defect 112e touches the cluster containing the third defect 142c and fourth defect 142d. Therefore a larger cluster is formed containing the third defect 112c, fourth defect 112d and fifth defect 112e. As this cluster has an odd number of defects and is not touching the boundary, this cluster will keep growing. Figure le shows a fifth graph 150, depicting the result of a fourth stage of growth using the clustering algorithm. After the fourth stage of growth, there are two clusters in the decoding graph 150. Both of the clusters in the decoding graph 150 are "neutral" clusters (which may also be referred to as stable clusters): a first neutral cluster 152 and a second neutral cluster 154. Each of the two remaining clusters is "neutral" due to having an even number of defects and / or having met the boundary. The first neutral cluster 152 contains the third defect 112c, fourth defect 112d, fifth defect 112e and sixth defect 112f, This first cluster 152 is formed due to the cluster containing the third defect 112c, fourth defect 112d and fifth defect 112e meeting the cluster containing the sixth defect 112f. The second neutral cluster 154 contains the first defects 112a and the second defect 112b. The second neutral cluster 154 has not grown since the first stage of growth, depicted in Figure lb. The first neutral cluster 152 contains 4 defects and the second neutral cluster 154 contains 2 defects. As the first neutral cluster 152 and the second neutral cluster 154 both contain an even number of defects, both the first neutral cluster 152 and the second neutral cluster 154 stop growing. It can also be noted that the first neutral cluster 152 is also connected to the boundary and this alone could be enough to make the cluster "neutral", regardless of whether the cluster contains an odd or even number of defects. As all the defects are in "neutral" clusters, the clusters can now be decoded by any conventional means. For example, the error can be decoded by defining a logical operator involving edges at a boundary of the decoding graph (e.g. the boundary edges on either the left or the right side of the decoding graph in Figure la-e) and counting how many clusters this logical operator intersects that contain an odd total number of defects. If the parity of this count is even, then the defined logical operator is considered to be free from error. However, if the parity is odd then the defined logical operator is considered to be in an error state, and its logical value should be flipped when it is measured. In this way, a single bit can be used to track the error state of the defined logical operator (i.e. it is not necessary to determine physical qubit error locations and physical qubit corrections). Upon using prior art approaches to cluster decoding, data conflicts can occur when processing elements associated with nearby or neighbouring nodes attempt to access the same information. For example, with reference to the example described above, when the cluster containing the first defect 112a met the cluster containing the second defect 112b, the edges of the two clusters may be growing out at the same rate, and therefore each cluster may require to 'know' about the other cluster's presence at the same time. With reference to figures 1-e, and assuming a 1:1 mapping between processing elements (on the decoding apparatus) and nodes (on the hypergraph), it is possible that, as the clusters grow using the prior algorithm, a first processing element associated with a first node may have requested information from a second, neighbouring processing element associated with a second node, at the same time the second processing element made the same or a similar request for information to the first processing element. Because the hardware implementing the prior art method must account for this possibility, separate input and output connections between each neighbouring processing element are required so that a node can both make queries and receive them at the same time. In addition, while the first processing element requests information from the second processing node, it is also possible that a third processing element associated with a third node requested information from the second processing node at the same time. Because the hardware implementing the prior art method must account for this possibility too, each processing element must have sufficient memory requirements to enable a request from a node to be stalled. For example, in this scenario, the second processing element must have sufficient memory to stall the request from the third processing element while the request from the first processing element is dealt with. While reference has been made primarily to requests for information, processing elements associated with neighbouring nodes may also write (not just read) information in each other's memories while performing the decoding algorithm. Therefore, the possibility arises that a first processing element may be stalled in its attempt to update information in a neighbouring second processing element's memory, while a request from a third processing element reads that same information. This scenario stalls the ability of nodes to update their neighbouring nodes with correct information, which in turn enables incorrect information to be passed through the network of nodes, further slowing down the decoding process. Therefore distributed decoding algorithms which make use of parallelisation of PEs can be used to implement conflict-free scheduling. An example of a parallelised algorithm that can be used in this conflict-free scheduling method is a local clustering decoder (LCD) algorithm. LCD is a distributed parallelised clustering decoder. Figure 2 depicts, at a high-level, the stages of the LCD algorithm 200. The flowchart of figure 2 could also be described as a state transition diagram. The new algorithm may be described as a parallelised LCD algorithm. The algorithm 200 comprises several stages, including an initialisation or initialising stage 205, a growing or growth stage 210, a merging stage 215, a picking stage 220, a syncing stage 225 and a stopping or exiting stage 230. The algorithm 200 comprises several steps, where each step is associated with one of the stages depicted in Figure 2. Each processing element is configured to perform the steps of algorithm 200 in respect of its associated node or nodes, such that all of the processing elements perform the LCD algorithm collectively, and together. Examples of code which could be used to implement each stage of the algorithm 200 are provided below, on the final pages of the description. The algorithm comprises several conceptual similarities with the algorithm described above with respect to figures la-e, and reference to the accompanying description of figures la-e may aid understanding of the algorithm 200. Each step of the algorithm 200 may be performed by one or more nodes at any given time. The description of the algorithm 200 may refer to a node in question, or a 'particular' node, where the node in question may be any single node in the decoding graph at a point in time. Also, the description of the algorithm 200 may sometimes refer to a node taking action, such as a node checking information with a neighbouring node. The skilled person will understand that this is shorthand for the processing element(s) associated with the node(s) taking the action. In addition, the present disclosure may refer to a node "growing", or "merging with its neighbour", and the like. Again, the skilled person will understand that this is shorthand for the associated processing elements performing actions such as updating a growth parameter associated with a particular node, or updating a cluster index associated with a particular node, and the like. In order to obtain parallelisation of a distributed decoder, PEs can be grouped. In particular, the PEs are grouped into a plurality of groups of PEs. These groups of PEs perform the steps of each stage of the clustering algorithm sequentially and successively, such that each group moves through the stages of the clustering algorithm. Groups of PEs may perform the steps of each stage of the clustering algorithm sequentially; for example, in a first stage of the clustering algorithm, the PEs in a first group of PEs may perform one or more steps associated with the first stage, then PEs in the second group perform the one or more steps associated with the first stage, then PEs in the third group perform the one or more steps associated with the first stage, and so on until each group of PEs has performed the one or more steps. In a subsequent second stage, the PEs in the first group may perform one of more steps associated with the second stage, then PEs in the second group may perform the one or more steps associated with the second stage, and so on. In this way, conflicts between neighbouring nodes can be avoided. The PEs are grouped such that conflict is avoided between neighbouring nodes (i.e. nodes connected to each other by one hyperedge on the hypergraph). For the effects of the grouping to be maximised, the groups should be formed such that no node has a neighbouring node which is associated with a different PE, where those PEs are in the same group. In other words, for any particular node on the hypergraph, the neighbouring nodes of the particular node must be either: i) associated with the same PE as that associated with the particular node (in which case the PE can schedule tasks for the neighbouring nodes to avoid conflict, for example by assigning each node in its 'batch' of nodes the steps of the algorithm successively as will be explained); or ii) associated with a different PE to the PE associated with the particular node, where these PEs are in different groups (in which case, since the groups of PEs perform tasks successively rather than at the same time, the neighbouring nodes will never send conflicting requests to one another). Further to point i), these neighbouring nodes form part (or all) of the "batch" of nodes associated with that PE, as will be explained. In some implementations of the present disclosure, each of the plurality of PEs is associated with a respective batch of nodes of the decoding hypergraph, where each batch contains multiple nodes. In this implementation, each PE performs one or more first steps of the clustering algorithm by performing them in respect of each of the nodes in its batch of nodes successively. In other words, each PE with a batch of nodes performs tasks in respect of each of its associated nodes in sequence. In this way, conflict between nodes in the same batch is avoided. Figure 2 depicts an algorithm 200 for clustering nodes in a decoding hypergraph or graph. The growth of clusters on a decoding graph is often described in terms of tree growth, where a cluster may be referred to as a "tree". Continuing this analogy, each cluster has a "root" node, whereby all nodes in a cluster are descended from the root node. As a cluster grows, the root node may become the "parent" to other nodes in the cluster, as more nodes are added to the edge of the cluster. The edge of the cluster may also be known to the skilled person as the "boundary" of the cluster. Similarly, a new node added to a cluster may be described as a "child" node to a parent node. Therefore a large cluster will exhibit a series of child-parent relationships along the branches of the cluster tree, all the way back to the root node. If a child node has an odd parity, it may be described herein as an odd child. At the beginning of the algorithm 200, each node is the root of its own cluster tree, as there is one node per cluster. As a cluster grows, the root of the cluster may change, for example upon merging with one or more other clusters. Clusters may be associated with an activity status. The activity status of a cluster is either active or inactive. An active cluster is a cluster that comprises an odd number of defects and is not touching a boundary of the decoding graph. A node is active if its cluster is active. Therefore, an active cluster and an active node may be used somewhat interchangeably. A cluster is inactive if there are an odd number of defect nodes in the cluster or the cluster meets the boundary of the decoding graph. When a cluster is inactive, all the nodes within that cluster will be inactive, such that all the nodes in the cluster will stop growing. Therefore an inactive cluster and an inactive node may also be used somewhat interchangeably. A node that is a defect will initially be active and a node that is not a defect will initially be inactive. Therefore, the activity of a node, and hence a cluster, may be related to the node's parity. Processing elements and / or nodes may be deemed busy or not busy. Multiple processing elements may be connected to a controller, in which the controller controls the processing of each processing element it is connected to. For example, the controller instructs processing elements to be in certain stages of the method 200. A processing element becoming busy is used to flag to the controller that something has changed. In the situation where each processing element is associated with a plurality of nodes, a processing element is busy when at least one of its nodes is busy. For example, a node is busy if its data changes during the merging or syncing stages, as described below. If a node is busy, the controller must re-run the current stage to allow the other processing elements to process the change that has occurred to the node. In implementations of the present disclosure, one node may be assigned to each processing element, e.g. in a 1:1 mapping. Alternatively, multiple nodes may be assigned to each processing element. As set out above, the multiple nodes assigned to a processing element may be referred to as a "batch" of nodes. Equivalently, similar words such as "group" may be used to refer to the multiple nodes assigned to a processing element. When multiple nodes are associated with each processing element, the stages of the method 200 may be performed with an extra outer loop, such that each stage of the method 200 is performed for each node in the batch. This is best appreciated by inspection of the code listings provided toward the end of this description. At step 205, the LCD algorithm is in the "initialising" stage. The initialising stage may include setting up initial parameters associated with the decoding hypergraph used in the algorithm 200. Each PE comprises its own dedicated memory, and this memory is used to store attributes relating to the one or more nodes associated with the PE. For PEs which have a batch of nodes, the dedicated memory is used to store information associated with each node in the batch. The attributes associated with each node may comprise one or more of an index of the node, a cluster index, a growth parameter, a defect flag indicating whether the node is associated with a defect; an activity flag indicating whetherthe node and / or a cluster to which the node belongs is active; a parity flag indicating a parity of the node, and a busyness flag indicating whether the node is busy, and the like. Stage 205 comprises initialising these attributes to starting values. For example, step 205 may comprise assigning each node a node index. At step 205, each node in the decoding graph is assigned to its own cluster, so that each node is initialised with its own cluster index. Therefore at step 205, the node index and cluster index may be the same. The parent of each node may be initialised to be the node itself. In other words, each node is initialised to be the root of its own singleton cluster. In an example, non-limiting implementation, attributes for each node, and their initialised values, may be as follows: nindex The index of the node. Must be in [2, N), where N is the number of nodes in the decoding graph. The indices 0 and 1 are reserved for the boundary nodes on the left and right-hand side of the patch cindex The index of the node's cluster. Must be in [0, N), where N is the number of nodes in the decoding graph. Initially, node.cindex = node.nindex parent The parent of the node in the tree of its cluster. Must be in [0, N), where N is the number of nodes in the decoding graph. Initially, node.parent = node Growth parameter The degree to which the node has grown. Must be in [0, 2], where 0 represents un-grown, 1 half-grown and 2 fully-grown. Initially, node.growth = 0 defect True if the node is defective. Must be in [0, 1], since this is a Boolean flag. Initially, node.defect = syndrome[node.nindex] active True if the node is active. Must be in [0, 1], since this is a Boolean flag. Initially, node.active = node.defect parity True if the node is odd. Must be in [0,1], since this is a Boolean flag. Initially, node.parity = node.defect busy True if the node is busy. Must be in [0,1], since this is a Boolean flag. Initially, node.busy = 0 As will be appreciated from the following description, algorithm 200 involves neighbouring PEs reading data from and / or writing data to the dedicated memory of their neighbouring PEs, and this comprises reading an attribute and / or updating an attribute in the dedicated memory of the neighbouring PE. The dedicated memory of each PE stores attributes relating to the one or more nodes associated with the PE. At stage 210, the algorithm 200 is in the "growing" stage. In the growing stage, active nodes (i.e. those nodes in an active cluster) which haven't already reached a maximum growth parameter value have their growth parameter increased. Therefore, at stage 210, the size of one or more clusters in the decoding graph grows. In the growth stage, a PE may update a growth parameter associated with at least one of its associated nodes by one unit, and store the updated growth parameter in its dedicated memory. This may only happen if the node is active. Active nodes form part of an active cluster, i.e. a cluster with an odd number of defects and which has not merged with one or more boundary nodes of the hypergraph. A "growth parameter" is used to define the degree of growth of each node in a cluster. In practice, the growth parameter may be considered to be quantity such as a radius around a node. The growth parameter value may take values which correspond to "ungrown", "half grown", and "fullygrown", at any point in the growth cycle of a cluster. In this example, "ungrown" is a minimum value and "fully grown" is a maximum value of the growth parameter value. The increase of the growth parameter value from ungrown to half-grown may be referred to as one unit, and similarly the increase from half-grown to fully-grown. In an example, the growth parameter can be increased from 0 (minimum value), to 1 (half-grown), to 2 (a maximum value). Once at the maximum value, the nodes cannot be "grown" any further. A half-grown growth parameter of a node may span half an edge along each of its incident edges of the decoding graph. A fully-grown growth parameter of a node may span a full edge along each of its incident edges of a hypergraph. Therefore if two neighbouring nodes are half-grown, they are connected by one fully grown edge on the decoding graph. If instead two nodes are separated by a distance of two edges on the decoding graph, they may be connected at the stage of both of those nodes having fully-grown edges. At stage 210, any active nodes that have an ungrown or half-grown growth parameters associated with them in the decoding graph will increase their growth parameter by a unit. Any active nodes with a fully-grown growth parameter will remain with a fully-grown growth parameter, and the edge will not continue to grow. At stage 215, the algorithm 200 is in the "merging" stage. In the merge stage, one or more PEs query the dedicated memory of a neighbouring PE to determine whether a sum of the growth parameters of neighbouring nodes meets a growth parameter threshold. Two inactive clusters, and therefore two inactive nodes, cannot be merged in the merge stage. However, the merge stage may be performed on inactive nodes in order to accurately propagate information. Only one cluster needs to be active for a merge to occur. In other words, an active cluster can merge with another active cluster, or with an inactive cluster. A growth parameter threshold is met when a fully-grown edge is formed between two neighbouring nodes. The growth parameter threshold may be met when the sum of the growth parameter values of two neighbouring nodes is equal to or greater than 2 units. For example, two neighbouring nodes would meet this criteria if they both had a half-grown edge, and were therefore connected. In the merge stage, different clusters are merged when the growth parameter threshold is met between neighbouring nodes in different clusters. When it has been determined, that two clusters should be merged, it then needs to be determined which cluster index to assign to the new cluster. Therefore, if the growth parameter threshold is met, the one or more PEs additionally determine whether one or more merging criteria are met. A purpose of the merging criteria is to determine which cluster index value to use for the new cluster. The one or more merging criteria are based on a cluster index value stored in the dedicated memory of the neighbouring PE. One or more of the merging criteria may be based on the cluster index of the neighbouring node being less than the cluster index of the node in question. If the one or more merging criteria are met for a particular cluster, the one or more PEs merge the clusters by updating a cluster index in their dedicated memories for those nodes associated with the particular cluster. If the one or more merging criteria are not met, the cluster indices are not updated for those nodes. The cluster index of multiple nodes may be updated within the dedicated memory of the same PE within the merging stage. In summary, a node will (typically) adopt the cluster index of a neighbour connected to it by a fully-grown edge if the neighbour has a cluster index less than the cluster index of the node. At the beginning of stage 215, the status of the node in question may be set to not busy. At stage 215, clusters that met during stage 210 may merge into a single, larger cluster. Typically, two clusters will meet by one fully-grown edge and they will become one single, resultant cluster during the merge. Alternatively, three or more clusters may meet at the same time by a fully-grown edge between each cluster. Alternatively, two clusters could meet by two or more fully-grown edges at one time, and so on. As the clusters merge, the status of the node is set to busy. During the merging stage 215, the cluster index, parent or parity of a node can change. When two or more clusters merge, the cluster index of one or more of the clusters may change, such that the resultant cluster has one cluster index. In this example of the algorithm 200, the cluster index of the resultant cluster is taken to be the cluster index of the merging cluster that has the lowest cluster index. However, in alternative implementations, the resultant cluster may take on another cluster index, such as the largest cluster index of the merging clusters, for example. When the node in question is connected to a neighbour by a fully-grown edge, the cluster index of the node in question is updated to be the cluster index of the neighbour if the cluster index of the neighbour is lower than that of the node in question. The parent node of the node in question may then change as the clusters merge. The node in question may become a child node to the neighbouring node. In other words, the neighbouring node may become the parent node to the node in question as the clusters merge. Once the resultant cluster is formed, a root of the resultant cluster may be allocated. The root of the resultant cluster will be decided based on the parenthood relationships in the merging clusters. For example, the root node of the resultant cluster may be the root of the merging cluster with the lowest cluster index. The root of the new larger cluster may be another node, such as the root node of the merging cluster with the largest cluster index. The parity of the resultant cluster may be decided after the merge. Each child node in the cluster relays its parity onto its parent, if the child node in question is odd. This relaying of the parity may in practice be an addition modulo two. For example, for a child node with an odd parity (i.e. a parity of 1) and its parent node with an even parity (i.e. parity of 0), the child node may relay its odd parity such that the child node has a parity of 0 and the parent node has a parity of 1. This operation will not be performed for the node which is the root of the cluster, as the root node may be defined as its own parent node. As data is being changed in this process, the node in question is set to busy when relaying the parity information to its parent. This relaying of parity information is repeated for each node within the cluster until the parity of each child node in the cluster is even. This process may be repeated multiple times for each node. For example, if at least one child node in the resultant cluster starts with an odd parity and therefore relays this parity information, this process will need to happen at least twice for each node. This is because the odd parity child node will be in the busy state. Each node in the cluster will then need to be checked again until all of the child nodes are not busy, in order to continue to the next stage. The parity of the root node will therefore represent the parity of the cluster. At the end of the merging state, the parity of each cluster will be equal to the parity of the root node, i.e. the node with the lowest index in the cluster. At stage 220, the algorithm 200 is in the "picking" stage. In the picking stage, one or more PEs update the activity of the root node associated with one or more clusters formed by one or more of their associated nodes in their dedicated memories. The parity of the root node of a cluster determines the activity of the cluster. If the root of the resultant cluster has an odd parity, the one or more PEs put the root node associated with this resultant cluster into an active state. Therefore this resultant cluster is said to be active. If the root of the resultant cluster has an even parity, the one or more PEs put the root node associated with this resultant cluster into an inactive state. Therefore this resultant cluster is said to be inactive. At the end of the picking stage, the nodes in a cluster that are not the root node are set to be inactive. Therefore, at the end of the picking stage, only the root node of a cluster will be active. At step 225, the algorithm 200 is in the "syncing" stage. In stage 225, one or more PEs update the activity flag and / or busyness flag associated with one or more nodes in their dedicated memory. The one or more PEs propagate the activity status of the root node of one or more clusters to other nodes of the one or more clusters. If the node in question is connected by a fullygrown edge to a neighbour that is active, the node in question becomes busy. By determining that these nodes are connected by a fully grown edge, we are considering nodes that are in the same cluster. The node in question will then become active if it is busy or if it was already active before this step. The picking stage ensures that all nodes in an odd parity cluster must be active. As represented by the arrow looping from the end of the syncing stage back to the start, the syncing stage is repeated until none of the nodes are busy. The explanations of stages 205 225 have generally been given in terms of one resultant cluster. However, the algorithm 200 is scalable to multiple clusters on the decoding graph. For example, there may be multiple clusters growing simultaneously or at different times during the stages of the algorithm 200. If one or more clusters in the decoding graph are still active at the end of stage 225, the algorithm 200 goes back to the growing stage at step 210. At stage 230, the algorithm 200 is in the "exiting" stage. The exiting stage will begin once a stopping criterion is reached. The groups of PEs perform the steps of each stage of the clustering algorithm until the stopping criterion is reached. This stopping criterion defines the final cluster state. The stopping criterion is reached when each cluster of the decoding hypergraph has even parity, i.e. either comprises an even number of defects, or has reached the boundary of the decoding hypergraph. In other words, if there are no active clusters remaining in the decoding graph in the syncing stage, the algorithm 200 will continue to the exiting stage 230, in which the algorithm stops. This could include clusters that have an odd parity but have met the boundary, and therefore will not continue growing. At step 230, the decoding graph will contain one or more neutral clusters in which the errors can be decoded. Correction(s) to the encoded logical state can then be determined based on the final cluster state, i.e. based on the clustering of defects. Figures 3a-h show decoding graphs 310, 320, 330, 340, 350, 360, 370 and 380. Figures 3a-h depict a simplified example of a process for decoding a patch of surface code using the algorithm 200 depicted in figure 2. Figure 3a depicts a first stage, and Figure 3h depicts a last stage of decoding a patch of surface code. Each stage of the decoding process shown in Figures 3a h may show the result of a stage of the algorithm 200, shown in Figure 2, as detailed below. However, stages of the decoding process may also not correspond to a stage of the algorithm 200 or may show the result of multiple stages of the algorithm 200. The circles and squares on the decoding graphs both represent nodes. The circles represent inactive nodes and the squares represent active nodes. Empty nodes represent nodes with an even parity and filled-in nodes represent nodes with an odd parity. The nodes of the decoding graphs depicted in Figures 3a-h may correspond to syndrome qubits in a quantum computing system. Each node is labelled with its cluster index above it and its node index below it. The cluster index of a node depicts the number of the cluster that the node belongs to. In Figure 3a the nodes are labelled with cluster indices 0 to 12, meaning each node belongs to a cluster numbered 0 to 12. The cluster index of a node may change through the clustering algorithm, as will be seen by inspection of Figures 3a-h. The node index of a node depicts a number associated with the node, in order to easily refer to different nodes in a decoding graph. Each node in the decoding graphs of Figures 3a g will have a different node index to every other node in the decoding graph. The exception to this is the boundary nodes which may have the same node index as at least one other node in the decoding graph. The node index of a node stays constant through the clustering algorithm, and therefore through Figures 3a-h. In order for the decoding apparatus to perform the decoding algorithm 200, certain PEs must be activated at each stage of the decoding algorithm 200. The determination of which PEs are activated is dependent on the execution scheme. The execution scheme is executed by a controller of the quantum computing system. Suitable hardware arrangements are disclosed below. In Figures 3a-h, possible connections between nodes are represented using dotted lines. While the graphs in Figure 3a-h can be used as basic decoding graphs, it is also possible to use more complex decoding graphs, for example with an extra dimension representing time, in which there is not necessarily a one to one correspondence between qubits and decoding graph edges. One skilled in the art will appreciate that the physical qubits do not necessarily need to be physically arranged as shown in Figures 3a-h. In Figures 3a-h, the block lines between nodes show the extent of the growth parameter of a node on the decoding graph at a particular time, i.e. ungrown, half-grown or fully-grown. The arrows on fully-grown edges display parenthood relationships between nodes, i.e. an arrow pointing from node 1 to node 2 signifies that node 2 is a parent of node 1 and similarly that node 1 is the child of node 2. Figures 3a-h contain boundary nodes 312a-c, 314a-c. The boundary nodes 312a-c, 314a-c show the nodes on the boundary of the decoding graph. The boundary nodes 312a-c are the boundary nodes on the left boundary of the decoding graph and the boundary nodes 314a-c are the boundary nodes on the right boundary of the decoding graph. The boundary nodes 312a-c each have a node index of 0 and the boundary bodes 314a-c each have a node index of 1. One skilled in the art will appreciate that the nuances of the decoding graph will depend upon the error correction code being implemented, and that some decoding graphs (e.g. those for toric codes) do not have boundaries. The graphs 310, 320, 330, 340, 350, 360, 370 and 380 may each represent information derived from the hypergraph at one of the stages of the algorithm 200 described above in relation to Figure 2. Figures 3a-h schematically depict the general process of the method 200 for this example decoding graph, however not every step of the process is shown in Figures 3a-h for simplicity and brevity. In the example decoding graph given in Figures 3a-h, all of the 12 nodes are assigned to one single processing element for simplicity. In other words, the 12 nodes are part of one PE's batch of nodes. Figure 3a shows a first graph 310. The graph 310 depicts an example decoding graph. In this example decoding graph, there are 3 defects present. The defects are at node / cluster indices 6, 10 and 11. The defects are demonstrated by filled-in square nodes. The square shape of the node means that the node is active and the fact that the node is filled-in means that the node has an odd parity. All of the nodes in the graph 310 have ungrown edges. The graph 310 is included here to demonstrate the decoding problem, and therefore does not correspond directly to the result of a stage of the algorithm 200. Figure 3b shows a second graph 320. The graph 320 depicts the result of a first growing stage, correlating with stage 210 of method 200. In this first growing stage, the PE updates the growth parameter associated with the nodes at node indices 6,10 and 11, i.e. the defects. The PE updates the growth parameter of these nodes by one unit and stores the updated value of the growth parameter in its dedicated memory. In the graph 320, this updated growth parameter for defects at node indices 6, 10 and 11 is shown by each of these nodes growing out by half an edge. This half-edge growth is along each of the four edges incident to each of these nodes on the decoding graph. Therefore the extent of clusters 6, 10 and 11 have grown in this step. As the current stage is not syncing, the controller puts the PE into the next stage, i.e. the merging stage at stage 215 in method 200. Due to the lack of fully-grown edges, no cluster indices change. Also since there are no nodes with odd children, no parities change. Similarly, the PE moves through the subsequent picking and syncing stages (stage 220 and 225, respectively) without changing the activity of the nodes. After the syncing stage, the PE has not reached the stopping criterion as multiple clusters have an odd parity. The PE remains active and so the controller puts the PE back into the growing stage (stage 210). Figure 3c shows a third graph 330. The graph 330 depicts the result of a second growing stage, once again correlating with stage 210 of method 200. In this second growing stage, each defect / cluster (at node indices 6, 10 and 11) has grown out its growth parameter by a half-edge along all possible edges surrounding it on the decoding graph. In other words, the PE associated with each of the nodes at node indices 6,10 and 11 increases a growth parameter associated with these nodes by one unit in its memory. Therefore the extent of clusters 6, 10 and 11 have grown in this step. The three clusters in the graph containing defects are now connected to each other by at least one fully-grown edge. At the same time, the cluster with cluster index 11 has met the boundary at node 0. The PE associated with the nodes of the decoding graph will then update the cluster index of the nodes associated with the clusters 6, 10 and 11. Figure 3d shows a fourth graph 340. The graph 340 depicts the result of a merging stage, correlating with step 215 of method 200. In this merging stage attributes for multiple nodes are updated in the memory of the PE. This merging stage starts a flood, where many cluster indices and parent nodes change as the union often clusters take shape. As explained in step 215, the parent-child relationship between nodes will change based on the cluster index of neighbouring nodes. For example, the defect node at node index 10 is now connected by a fully-grown edge to the neighbours with node indices 8, 9, 12 and 13. As neighbouring node 8 has the lowest cluster index, defect node 10 will have an updated cluster index in the memory of the PE to that of cluster 8, and node 8 will become its parent. During this change, the node will be put into the busy state. As explained with regards to stage 215, the defects at node indices 6,10 and 11 relay their odd parity to their parents. During this change, the node will also be put into a busy state, if it was not already in a busy state. At the end of this merge, by cycling through all the nodes in the cluster, the nodes with node index 4 and 8 have an odd parity. Some of the edges between nodes in Figure 3d do not show parent-child relations due to the order in which the nodes have been processed in this example. As multiple changes of parent and parity have occurred during this merge, at least one node is still in the busy state, and therefore the merge will continue further. Figure 3e,f and g show graphs 350, 360, and 370, respectively. These graphs depict further merge steps, once again correlating with stage 215 of method 200. The merge continues due to at least one node still being in the busy state, as child-parent relationships and node parities continue to be updated in the memory of the PE whilst cycling through the nodes in the resultant cluster. Equivalently, the merging stage re-runs until all children in the resultant cluster have an even parity. Graph 370 depicts the final merge step, in which the child-parent relationships and node parities of the cluster reach their final merged state. In the final merged state, each node in the resultant cluster has the cluster index 0 stored in the memory of the PE. The root of the cluster tree is the node with node index 0, which has an odd parity. This can be seen by all the child-parent relationships in the cluster pointing back to cluster 0. In this example, the resultant cluster has an overall odd parity, however it ceases to continue growing as the resultant cluster has met the boundary. Figure 3h shows an eighth graph 380. The graph 380 depicts picking and syncing stages after the final merge stage, correlating to step 220 and step 225 of method 200 respectively. In the picking stage, the PE updates the activity of the root node associated with the resultant cluster. In the syncing stage, the PE updates the activity status of the child nodes based on the activity of the root node of the resultant cluster. The PE updates the busyness flag of multiple nodes associated with the resultant cluster in its dedicated memory . If one or more nodes are still busy at the end of the syncing stage, this stage may be repeated one or more times, as described above. These final picking and syncing stages deactivate the defect nodes 6,10 and 11. The LCD algorithm proceeds into the exiting stage, corresponding with step 230 of method 200. The logical correction equals the parity of the topmost boundary node on the left i.e. a parity of 1. This is equivalent to the sum modulo 2 of the number of defects in cluster 0. Distributed decoding algorithms of the type described above can be performed using a plurality of distributed PEs, for example implemented using FPGAs and ASICs. The resulting decoder can be configured to perform, for example, a clustering algorithm such as the LCD algorithm. The PEs may be arranged in an ordered arrangement. The ordered arrangement may take the form of an array or grid, for example. The plurality of PEs in the decoder is often depicted in a square grid manner, and in some implementations this may reflect an actual hardware implementation. However, it should be understood that many different arrangements are possible and while reference is frequently made herein to a "grid" and many diagrams depict a grid, it should be understood that this is not an essential feature of the hardware. Even when a "grid" is used in hardware, implementations other than square or rectangle are also possible, such as in other shapes of two dimensional arrangements and three dimensional structures. Each PE of the plurality of distributed PEs is associated with one or more nodes of the decoding hypergraph. Each PE in the processing grid may be associated with a different node in the decoding graph according to a 1:1 mapping, for example, or each PE in the processing grid may be associated with more than one node in the decoding graph. Herein, a grid of PEs may also be referred to as an array of elements. Figure 4 shows a graph depicting an example processing grid 400 of PEs 402. In Figure 4, the processing grid 400 comprises a 6 by 6 grid of PEs 402, such that there are 36 PEs 402. Each of the PEs 402 are represented by a square with an index inside. Each index is in the form 'PEX' where 'PE' denotes the phrase processing element and X denotes the index of the processing element in the processing grid 400. The indices have been included in the representation of the PEs for ease of reference. Figure 4 shows an example where the bottomright processing element is labelled PEO. The index increments by 1 in the right-to-left direction until hitting the boundary. The indexing then continues from the right-most element in the next line in similar manner. This process is repeated until all elements of the grid are exhausted and have appropriate index. Figure 4 shows an example processing grid 400 which could be used to decode a rotated planar surface code patch of size 5x5x6. Each processing element 402 in the processing grid 400 is essentially a small decoding engine that can process several nodes in a serial manner and / or in a parallel manner at stages of the decoding algorithm. Figure 4 also depicts data links 404 between PEs. Data links 404 are used to directly couple PEs in hardware. The data links 404 in Figure 4 can be classified into four different types of data link: spatial links, temporal links, short hook links and long hook links. Data links 404 are associated with measurement errors in the decoding graph, of which the skilled person is familiar. Decoding graphs may be made up of one or more layers, wherein each layer is associated with a decoding round. Therefore, each layer in the decoding graph may be associated with different moments in time and may be referred to as a "time slice". A decoding algorithm may be performed on each layer of the decoding graph, and then repeated again in each decoding round. As would be understood by the skilled person, the purpose of having multiple decoding rounds is to allow the decoder to understand and / or correct measurement errors that may have occurred between decoding rounds, and hence between time slices. Figure 5 depicts a decoder apparatus 500. The decoder apparatus 500 comprises a controller 504 and a plurality of distributed PEs 402. The decoder apparatus 500 further comprises a plurality of control lines 506. Each control line 506 in the decoder apparatus 500 couples the controller 504 to a respective PE 402. A control line 506 may be used to activate PEs according to an execution scheme. The decoder apparatus 500 depicted in Figure 5 is in line with prior art arrangements. An execution scheme can be implemented, by the decoder apparatus 500, in which each PE 402 that needs to be executed at a particular stage of the distributed decoding algorithm can be identified using the indices assigned to the PEs. At any stage of the distributed decoding algorithm, one or more PEs may be activated in order to perform the stage of the algorithm. For example, a plurality of PEs may be activated when performing the merge stage 215 of the LCD algorithm 200 due to an execution scheme. In this merge stage, for example, PEs 0, 3, 4 may need to be activated. The execution scheme may comprise the information related to the indices of these PEs, for example the indices 0, 3 and 4, respectively. The number of indices is equal to the number of PEs in the processing grid. The PE indices may be labelled from 0 to N-l, where N is the number of PEs in the processing grid. Alternatively, the PE indices may be labelled from 1 to N, or other similar representations. For example in Figure 5, the PE indices range from 0 to 8 for the 9 PEs. The indices of the PEs to be executed according to this prior art method may be stored as binary indices. The number of binary digits needed to store an index of a PE depends on the number of PEs. In the execution scheme according to the processing grid 500 of Figure 5, 4 bits may be needed to refer to a particular PE 402 in the processing grid 500. For example, PEO may be represented in 3 bits as 000 (or 4 bits as 0000), PE5 may be represented in 3 bits as 101 (or 4 bits as 0101) and PE8 may be represented in 4 bits as 1000. The processing grid 400 of Figure 4 has 35 PEs. In the execution scheme according to the processing grid 400 of Figure 4, 6 bits may be needed to refer to a particular PE 402 in the processing grid 400. For example, PE35 may be represented in 6 bits as 100011. Figure 5 shows control lines 506 connecting PEs 402 to a controller 504. Each control line 506 may transmit information between the controller 504 and the PEs 402 connected to the particular control line 506. The indices of the PEs 402 to be activated are broadcast to all PEs 402. This information may comprise a list of the binary indices of the PEs to be activated. For example, if only PE5 in the processing grid 500 is to be activated according to an execution scheme, all PEs in the processing grid 500 will receive the binary index 101. PE5 will then determine that it is a PE to be activated. Each control line 402 therefore needs to have a high bandwidth in order to broadcast the one or more indices of PEs 402. The number of PEs that are to be activated according to the execution scheme impacts the amount of storage needed to store the execution scheme. Therefore, the bandwidth required to communicate the execution scheme to the controller 504 may be affected. A large number of PEs to be activated in an execution scheme is associated with a large amount of storage of the controller 504 needed to store the execution scheme and / or a large amount of bandwidth to communicate the execution scheme to the controller 504. The execution scheme may be hardcoded in the hardware or transferred at run-time of the hardware. It follows that the number of PEs to be activated in an execution scheme has a direct impact on the number of bits required to store the execution scheme. For example, 6 bits may be required to address one index of a PE in hardware, as in the processing grid 400 of Figure 4. In the example where the PEs 0, 3 and 4 of Figure 5 are to be activated using an execution scheme, the execution scheme may require (3*6=) 18 bits to store all 3 indices in the execution scheme. The amount of storage required to store the execution scheme increases with the number of PEs that need to be activated, which further adds complexity. It is evident from the above example that storing the execution scheme in this manner is highly unoptimized due to the amount of storage required to store the execution scheme based on indices of individual PEs. This is particularly unoptimized in the situation where only one control line links the controller to the PEs, as in Figure 5. This single control line will need to have a high bandwidth in order to broadcast the PE indices to be activated. The following methods of the present disclosure show improvements to the prior methods. Control lines 506 are different to the data links between PEs 402. Data links connect one PE 402 to another, whereas control lines 506 connect at least one PE 402 to a controller 504. There may be one or more controllers 504 in the decoding apparatus 500. For example, each PE 402 could be connected via a control line 506 to a separate controller 504. Alternatively multiple PEs 402 may be connected a first controller 504, and multiple other PEs 402 may be connected to a second controller 504 via control lines 506. Figure 6 depicts a method 600 according to the present disclosure. Optional steps of the method 600 are shown with a dashed line in Figure 6. The method 600 may comprise steps 602, 604, 606 and 608. The method 600 is a computer-implemented method for controlling the execution of a plurality of distributed PEs, in a decoder apparatus of a quantum computer system, for example a quantum computer system 1200 as depicted in Figure 12. The quantum computer apparatus further comprises a register of quantum devices. The decoder apparatus is configured to perform the method 600. The distributed decoding algorithm is performed by the plurality of PEs. The distributed decoding algorithm requires communication between linked PEs of the plurality of distributed PEs. The decoder apparatus may comprise a controller and computer memory. There may instead be a separate controller coupled to each PE, as will be described later with respect to Figs 9-11, which acts to move each PE through the various stages of the clustering algorithm. At step 602, the decoder apparatus receives syndrome data. Step 602 is an optional step in the method 600. The syndrome data is representative of an error state of the quantum devices in the register of quantum devices, the syndrome data comprising a plurality of defects. The syndrome data is representable as a decoding hypergraph comprising a plurality of nodes connected by hyperedges representing error mechanisms associated with the plurality of quantum devices. Each PE of the plurality of distributed PEs is associated with one or more nodes of the decoding hypergraph. At step 604, the plurality of PEs are controlled to perform a distributed decoding algorithm according to an execution scheme. The PEs of the plurality of PEs are distributed PEs and each PE is controlled by one or more controllers 504. The execution scheme may be stored by the quantum computing system, and optionally at the one or more controllers. Alternatively, the execution scheme may be generated by the one or more controllers at run-time. There may be one or more aspects that determine which PEs should be activated in the execution scheme for different decoding scenarios, for example for different distributed decoding algorithms and / or different decoding graph arrangements. The execution scheme identifies which PEs of the plurality of PEs should be executed at different stages of the distributed decoding algorithm. The execution scheme of the merging stage 215 may be similar to or the same as the execution scheme of the syncing stage 225. The execution scheme of the picking stage 220 may be similar to or the same as the execution of the growing stage 210. Each different stage of the distributed decoding algorithm may be associated with a different execution scheme. With reference to the clustering algorithm 200 as depicted in Figure 2, the growing stage 210 and picking stage 220 may be associated with one execution scheme and the merging stage 215 and syncing stage 225 may be associated with another execution scheme. The indices included in the execution scheme will change depending on the number and identification of the PEs that need to be execution at each stage. Each PE is identifiable within the execution scheme using one or more indexing variables, as will be described with respect to Figures 9-11. Each indexing variable of the one or more indexing variables is comprised of one or more bits, with each bit being associated with the execution of one or more PEs of the plurality of distributed PEs. Each bit in the one or more indexing variables is associated with the execution of one or more PEs, independent of the other bits in the indexing variable. Each PE may be uniquely identifiable within the execution scheme using the one or more indexing variables, as will be described later with respect to Figures 9-11. Performing the distributed decoding algorithm may comprise determining a correction for the error state. The correction corrects for the error state of the quantum devices, and may be based on a final cluster state. The final cluster state is reached when the clustering algorithm stops or exists. For example is using a clustering algorithm, once the final clustering state has been reached, a correction can be determined in a known way. At step 606, the logical state encoded in the quantum devices of the quantum computer may be measured to obtain a logical state measurement. Step 606 is an optional step in the method 600. At step 608, the correction for the error state may be applied to the logical state measurement. Step 608 is an optional step in the method 600. Figure 7 depicts a method 700 according to the present disclosure. The optional steps of the method 700 are shown with a dashed outline in Figure 7. The method 700 comprises steps 702, 704 and 706. Steps 702, 704 and 706 are steps which may be performed as part of step 604 of method 600 in Figure 6. The 604 of method 600 in Figure 6 may comprise one or more of the steps 702, 704 or 706. The method 700 therefore depicts optional steps in controlling the plurality of PEs to perform a distributed decoding algorithm according to an execution scheme. Method 700 is performed at step 604 when the method 600 is performed in connection with a particular hardware implementation. As set out above, the controller controls which PEs are activated according to the execution scheme, thereby enabling the plurality of PEs to perform the distributed decoding algorithm. The controller is coupled to PEs via control lines. In the indexing scheme described herein, each bit is associated with the execution of one or more PEs. In some hardware implementations, this is accomplished by virtue of each bit of the indexing variable being associated with a different respective subset of the plurality of distributed PEs. In this scenario, each bit can further identify a control line which couples the controller to the associated subset of PEs. At step 702, the controller determines which control lines are identified by the execution scheme, e.g. at a particular stage of the algorithm. In this way, the execution scheme indicates to the controller which control lines, and hence which PEs, to activate. Each control line may connect a single PE to a controller. This will be discussed in more detail with respect to Figure 11. Alternatively, a control line may connect more than one PE to a controller. This will be discussed in more detail with respect to Figure 9. At step 704, the identified control lines are activated by the controller. The controller determines the one or more PEs that need to be activated according to the execution scheme. The controller determines which one or more control lines are identified by the one or more indexing variables in the execution scheme. The controller sends a signal to the PEs that need to be activated via the identified control lines. This signal may comprise a single bit of information. The single bit of information is received at the particular one or more PEs to be activated via the one or more control lines. The one or more PEs to be activated then identify that they are to be activated according to this signal. At step 706, the one or more PEs of the plurality of distributed PEs are executed, by the controller, according to the execution scheme. In the example where the only PE to be executed according to the execution scheme is PE5, PE5 will then run the necessary step / s of the algorithm. Figure 8 depicts a graph depicting a processing grid 800 of processing elements 402. Specifically, the processing grid 800 of Figure 8 depicts the processing grid 400 of Figure 4, but with a different indexing system schematically displayed in the figure. One or more indexing variables may comprise a first dimension indexing variable and a second dimension indexing variable. For example, the first dimension may refer to rows and a second dimension may refer to columns. Furthermore, a first dimension indexing variable may refer to rows of a processing grid and a second dimension indexing variable may refer to columns of a processing grid. The number of bits in the first dimension indexing variable may correspond with a number of PEs in a first dimension, and the number of bits in the second dimension indexing variable may correspond with a number of PEs in a second dimension. For example, the number of bits in the first dimension indexing variable may equal the number of rows in a processing grid and / or the number in the second dimension indexing variable may equal the number of columns in a processing grid. The first dimension indexing variable in the processing grid 800 of Figure 8 is a row indexing variable and the second dimension indexing variable in the processing grid is a column indexing variable. The rows of the processing grid 800 are labelled with a row label 806 varying from 0-5. The columns of the processing grid 800 are labelled with a column label 808 varying from 0-5. As such, PE35 in the processing grid 800 of Figure 8 has a row label of 5 and a column label of 5. As previously described, in the execution scheme, each bit is associated with the execution of one or more PEs of the plurality of distributed PEs. For example, with respect to the processing grid 800 of Figure 8, each bit of the execution scheme refers to a row or a column of the processing grid 800, as will be further explained below. As previously described, each PE is identifiable within the execution scheme using the one or more indexing variables. The one or more indexing variables may comprise a first indexing variable and a second indexing variable. In the example given in Figure 8, the first indexing variable may be a row indexing variable and the second indexing variable may be a column indexing variable. The first indexing variable is associated with the row label 806 in Figure 8. The second indexing variable is associated with the column label 808 in Figure 8. Alternatively, the first indexing variable could be associated with columns and the second indexing variable could be associated with rows. The first indexing variable may also comprise a different number of bits to the second indexing variable, for example in a rectangular grid rather than a square grid. As described previously with respect to Figure 4, the processing grid is displayed in Figure 8 as a square ordered grid, however there are many other alternatives. For example, the grid may have more rows than columns, or more columns than rows. Another alternative is that the grid may not be in an array structure at all. For example, the PEs could be physically arranged in the shape of a circle. Further, the PEs could be arranged in a completely random shape with no inherent pattern or array. In such arrangements, row labels and column labels can still be assigned to PEs in a similar way. The arrangement of PEs and the connection of PEs via control lines may be based upon capabilities of the hardware used, for example the capabilities of an FPGA or ASIC. A first indexing variable can still be associated with row labels 806 and a second indexing variable can still be associated with column labels 808. For example, the row labels 806 and the column labels 808 of the PEs in a non square grid can equivalently be referred to as the row labels and column labels of control lines, as will be described below with respect to Figure 9. Therefore, a first indexing variable and second indexing variable can equally be assigned to PEs in a processing grid that is not in a square-shape, or even an ordered array. The processing grid 800 of Figure 8 has 6 rows, where the rows have a row label 806 varying from 0-5 and 6 columns, where the columns have a column label 808 varying from 0-5. Each row can be encoded as a 1 bit entry in a first indexing variable. Each column can be encoded as a 1 bit entry in a second indexing variable. The number of bits required to store the first indexing variable may be equal to the number of rows and the number of bits required to store the second indexing variable may be equal to the number of columns in the processing grid 800. This may not be the case, for example, in a disordered processing grid where the PEs are not in a regular array. The one or more indexing variables may be in the form of a bit string. A bit being in a 0 state in an indexing variable may depict that the one or more PEs associated with the bit are not activated in the execution scheme. A bit being in a 1 state may depict that the one or more PEs associated with the bit are activated in the execution scheme. Alternatively, the 1 state could represent an off state of a bit and the 0 state could represent an on state of a bit. Furthermore, '0' and '1' could be replaced with 'on' and 'off' states, '+1' and '-1' states, or any suitable representation to differentiate a 'go' from a 'no-go' state. For example, the one or more indexing variables may be in a form such as '000000' or '6'b000000' for a 6 bit indexing variable. Similarly, if an indexing variable comprised 4 bits, it could be in the form '0000' or '4'b0000'. This example depicts each bit being in an 0 state, or off state. One or more of the indexing variables may be in different forms to one or more of the other indexing variables. The identification of PEs in an indexing variable could be from left to right, or right to left, or by other means. As can be appreciated, each bit of the one or more indexing variables is associated with a different respective subset of the plurality of distributed PEs. The first subset of PEs may be associated with rows of PEs and the second subset of PEs may be associated with columns of PEs, as in Figure 8. For example, the last bit in the first indexing variable '000000' may be associated with a first row of PEs. Similarly, the last bit in the second indexing variable '000000' may be associated with a first column of PEs. The states of the bits in the one or more indexing variables depict which subsets of PEs are to be activated in the execution scheme. For example, the first indexing variable '000100' could indicate that only PEs in row 2 are to be executed, with the last '0' depicting row 0 and the penultimate '0' depicting row 1 of the PEs. Similarly, the second indexing variable '100000' could indicate that only PEs in column 5 are to be executed. To schedule execution of PEs in rows 0, 4 and 5 together, the first indexing variable could be '110001'. A particular PE may be identifiable within the execution scheme via the combination of a first bit in the first indexing variable and a second bit in the second indexing variable. The first bit is associated with a first subset of PEs and the second bit is associated with a second subset of PEs. For example, a first bit in the first indexing variable may determine which row the particular PE is in, and a second bit in the second indexing variable may determine which column the particular PE is in. The particular PE is present in both the first and the second subset of PEs. For example, with respect to Figure 8, PE34 is in row 5 (a first subset) and column 4 (a second subset) and therefore the particular PE34 is present in both subsets. This means that this particular PE, PE34, is identifiable by using the first indexing variable and the second indexing variable. PE34 may be activated in an execution scheme where the first indexing variable is 000010 and the second indexing variable is 010000. Furthermore, multiple PEs may be activated at once using an execution scheme. For example, an execution scheme comprising a first indexing variable of 001010 and a second indexing variable of 000100 may activate PE8 and PE20 in Figure 8. In a distributed decoding algorithm such as the algorithm 200, as detailed above with respect to Figure 2, multiple stages of the algorithm may comprise PEs being executed in parallel with multiple or all other PEs. To enable parallel execution of all PEs in the processing grid 800 of Figure 8, the first indexing variable could be 111111 and the second indexing variable could be 111111. Therefore, parallel execution of all PEs only requires two 6 bit indexing variables, and therefore the storage of 12 bits altogether. Prior methods may require the storage of 6 bit variables for each of the 36 PEs, such that 6*36=216 bits must be stored to execute the same parallelisation of PEs. Therefore this solution could achieve an 18 times reduction in the amount of required storage. This number may vary based on the arrangement of the PEs and based on the one or more indexing variables used. For example, there may be further improvement when using more than 36 PEs. A grid capable of processing a 23x23x24 size of Rotated Planar Surface Code Patch has 576 PEs. In prior approaches, each PE may require 10 bits to store the index of each PE, and therefore 10 * 576 = 5760 bits would be needed to store the parallelised schedule for all the PEs. In the proposed method, the same functionality can be achieved with 48 bits altogether, which is reduction of a factor of 120 in the required storage compared to prior approaches. The one or more indexing variables have so far been discussed in terms of a first indexing variable and a second indexing variable, however there may be further indexing variables. There may be a third indexing variable, a fourth indexing variable, and so on. For example, the execution scheme may comprise a third indexing variable for a three dimensional arrangement of PEs. Figure 9 depicts a decoder apparatus 900. Figure 9 depicts a decoder apparatus 900 that may be suitable for an indexing scheme as described with respect to Figure 8, for example. Figure 9 depicts a decoder apparatus 900 that may be suitable for what will be referred to herein as "dense encoding". Each PE 402 of the plurality of PEs may be coupled to the controller 504 via a control line. Each control line may be coupled to a subset of PEs. In other words, multiple PEs may be coupled to the controller 504 via the same control line. For example, in Figure 9, PEO, PE3 and PE6 are all connected a single control line. PEO, PE3 and PE6 are therefore a subset of PEs. In the execution scheme, each bit of the one or more indexing variables may be associated with a different respective subset of the plurality of distributed PEs. This has been previously described with respect to Figure 8, where each bit in a first indexing variable was associated with a row of PEs and each bit in a second indexing variable was associated with a column of PEs. In Figure 8, the subset of PEs comprising PEO, PE3 and PE6 is a subset associated with a first column of PEs. Therefore one bit in the second indexing variable is associated with a subset of PEs comprising PEO, PE3 and PE6. As previously mentioned, the first indexing variable has been described as associated with rows and the second indexing variable has been described as associated with columns, however this could be the other way round, or different indexing variables could be used. For example, the subset of PEs comprising PEO, PE3 and PE6 could instead be associated with one bit in a first indexing variable. Each bit further identifies a control line which couples the controller to its associated subset of PEs. For example, in Figure 8, there is a single control line that connects the PEs in the subset comprising PEO, PE3 and PE6 to the controller. Therefore the one bit in the second indexing variable that is associated with the subset of PEs comprising PEO, PE3 and PE6, also identifies the single control line that connects this subset of PEs to the control. The PEs 402 in Figure 9 are arranged in an ordered squared grid. However it is simple to use alternative arrangements with the encoding schemes described herein. For example, as each bit may be associated with a subset of PEs and therefore a control line in this way, the PEs do not need to be in an ordered arrangement. The PEs could be in a random sparse arrangement. The PEs in a subset do not need to be next to each other in a processing grid. The PEs in a subset may be separated from each other and / or in random positions. As the control line connects each of the PEs in the subset, there is no need for the PEs to be in an ordered array. A PE may be in one or more subsets of PEs of the plurality of PEs. This means that each PE may be coupled to the controller via multiple control lines. Each control line may be associated with a different subset of PEs in the plurality of PEs. For example, in Figure 9, PEO, PE3 and PE6 are in a first subset and PE6, PE7 and PE8 are in a second subset of PEs. Both the first subset and second subset of PEs comprise PE6. PE6 has two separate control lines that couple it to the controller 504. A particular PE may only be executed if all the necessary control lines that connect the particular PE to the controller are activated. The control lines are activated according to step 704 of the method 700. For example, in Figure 9, PE6 will only be executed if both of the control lines connected to PE6 are activated by the controller. In this way, PEs are uniquely identifiable by the activation of control lines according to an execution scheme. The execution schedule of PEs may be associated with a single stage of a distributed decoding algorithm, such as the clustering algorithm 200 in Figure 2. The initialising stage, the growing stage and the picking stage may involve all PEs executing in parallel with one another. In other words, an execution scheme may not be needed. The merging and syncing stages may require an execution schedule. When activated by the execution scheme, a control line may pass information from the controller to the one or more PEs connected to the control line in order to activate the PEs according to the execution scheme. Each of the control lines may only need to transmit one bit of information to a PE. This is because the control line may only need to transmit whether a PE needs to be activated or not according to the execution scheme. The bandwidth of each control line is therefore independent of the number of PEs using this type of indexing variable. In prior methods, dedicated logic needs to be implemented to convert from the index stored in the control schedule of the PEs in order to enable signals that connect to the relevant PE. In the proposed method, the execution scheme can read the stored schedule value and the controller can send signals directly to the relevant PEs. In this way, each PE may only receive two one-bit signals. By simply toggling the status of the control lines, any PE can be controlled. This leads to a reduced bus-width of each control line and logic simplification. These benefits further help to improve implementation methodologies of possible hardware, such as "Place &Route" step of FPGA / ASIC methodology, for example. As previously mentioned, there may be multiple controllers in the decoder apparatus. Each PE in a subset may be coupled to one controller via a single control line, and each PE in a second subset of PEs may be coupled to another controller via a different single control line. A PE that is in more then one subset of PEs may be connected to multiple different controllers via multiple different control lines. Multiple PEs in the processing grid may be associated with the same one or more indexing variables. For example, PEs associated with nodes in different time slices of the decoding hypergraph may be associated with the same one or more indexing variables. For example, two PEs with the same one or more indexing variables may be associated with nodes in the decoding hypergraph that are a few layers apart, such that these PEs are not neighbouring PEs. In other words, these two PEs are not associated with nodes in the decoding hypergraph that are neighbouring nodes. PEs with the same one or more indexing variables may be in the same subset of PEs. Furthermore, PEs with the same one or more indexing variables may be associated with the same control lines coupled to the controller. Figure 10 depicts a graph depicting a processing grid 1000 of processing elements 402. Figure 10 depicts how the methods disclosed herein can be used with "patches" of PEs in a processing grid. Specifically, Figure 10 depicts a graph of a processing grid 1000 wherein the PEs are categorised into either a PE associated with a "grid patch" 1004 or a PE associated with a "dummy patch" / "inactive PE patch" 1002. The grid patch 1004 and dummy patch 1002 are distinguished using grey shading in Figure 10, as detailed by the key in the bottom right corner of Figure 10. The PEs associated with a dummy patch 1002 may by referred to as "dummy PEs" that are not instantiated physically. The dummy PEs may not physically exist in hardware. Alternatively, each of the dummy PEs, if they physically exist in hardware, may effectively be disabled at any point in the decoding algorithm, such that they cannot be executed by an execution scheme. The PEs associated with a grid patch 1004 may be executed or not according to an execution scheme. Therefore, "patches" of PEs in a processing grid can be executed. In this way, the number of PEs in a group of executable PEs and the shape configuration of this group in a processing grid is flexible and reconfigurable, which makes the design of the processing grid more adaptable to different decoding problems. An example of dummy PEs in a processing grid 1000 is shown in Figure 10. The rows and columns of PE in the processing grid 1000 are labelled with row labels and column labels, respectively, as in Figure 8. Column 5 of the processing grid 1000 comprises six PEs which are "dummy PEs": PE5, PE11, PE17, PE23, PE29 and PE35. These six PEs are not activated according to an execution scheme. The PEs in column 5 of the processing grid 1000 may be associated with one subset of PEs. This subset of PEs is represented in an execution scheme by a single bit. For example, an execution scheme may comprise a second indexing variable wherein the first bit is associated with column 5 of the PEs in the processing grid 1000 (a subset of PEs). The bit associated with this subset of PEs in the second indexing variable may be 0 in the execution scheme, as the PEs are not activated according to the execution scheme. This is effectively disabling all the PEs in column 5 of the processing grid 1000. Similarly, the bit associated with the PEs in the subset associated with row 5 may also be 0 in a first indexing variable, as all the PEs in row 5 are also dummy PEs. As such, the methods described herein are able to support rectangular patches of PEs in a processing grid and in particular, arbitrary configurations of PEs in a processing grid can be supported. Furthermore, dummy PEs may be initialised at any time, such that they are no longer dummy PEs and can be activated according to an execution scheme. Therefore, the methods disclosed herein allow for improved flexibility and reconfiguration of PEs. Figure 11 depicts a decoder apparatus 1100. Figure 11 depicts an alternative decoding apparatus 1100 to the decoding apparatus 900 of Figure 9. Figure 11 depicts a decoder apparatus 1100 that may be suitable for what will be referred to herein as the "sparse encoding", for example. Each PE 402 of the plurality of PEs may be coupled to the controller 504 via a different control line 506. Alternatively, each PE 402 could be coupled to one or more controllers via a different control line 506. The execution scheme for the PEs 402 in Figure 11 may comprise a single indexing variable. The single indexing variable is comprised of a number of bits equal to the number of PEs in the plurality of PEs. In the decoder apparatus 1100 of Figure 11, 9 PEs are used and each are labelled with an index from PE0 8. Each bit in the single indexing variable uniquely identifies the respective PE. A single indexing variable may be a bit string comprising 9 bits. For example, similarly to the description with respect to Figure 8, an execution scheme comprising an indexing variable of 000000001 may refer to only PEO being executed in the execution scheme from Figure 11. Alternatively, the indexing variable may be in a matrix form. For example, with respect to Figure 11, an execution scheme may comprise an indexing variable of the following form: 0 10 10 0 0 0 0 An indexing variable in the above form may correlate with the arrangement of PEs in the decoder apparatus 1100, such that PES and PE7 are activated in this execution scheme. Each bit of the single indexing variable is associated with a different respective subset of the plurality of distributed PEs. This is because each PE is identified as being in a different subset in this arrangement. An indexing variable of this form may allow the activation of many arbitrary patterns of PEs. In this example, the number of control lines is the same as the number of PEs. Therefore, each of the control lines need only transmit one bit of information to a PE, similarly to the description of Figure 9. A schematic of an exemplary quantum computing system 1200 suitable for performing the method of the present disclosure is shown in Figure 12. The quantum computer system 1200 may comprise a register of quantum devices and a decoder apparatus. The quantum computing system 1200 comprises a plurality of physical qubits 1206 (unless specified otherwise, reference herein to qubits should be understood to refer to physical qubits rather than logical qubits). The qubits 1206 include data qubits used to encode logical qubit states, and syndrome qubits (or auxiliary qubits) used to perform syndrome measurements for quantum error correction. While the exemplary quantum computing system 1200 uses qubits 1206, one skilled in the art will appreciate that the invention described herein is also applicable to quantum computing systems that use other quantum devices, such as qutrits and qudits. Accordingly, it should be understood that any reference herein to qubits is applicable to any type of quantum devices that can be used to encode quantum information. The qubits 1206 are controlled by a control system 1204 having one or more classical processors. The control system 1204 transmits control signals (e.g. RF pulses) to the qubits 1206 for performing operations on the qubits 1206 (including measurement operations) and receives measurement information from the qubits 1206. The measurement information will generally be analogue data signals, although the analogue signals may alternatively be converted to digital signals before being transmitted to the control system 1204 in some implementations (e.g. the qubits 1206 may be provided with one or more analogue to digital converters). The control system 1204 may receive high-level instructions from an algorithmic system or similar (not shown) and convert these high-level instructions (such as logic gates) into low-level qubit instructions (e.g. microwave pulses etc.), which may be in analogue format. The quantum computing system 1200 also comprises a decoding system 1202 (also referred to herein as a decoder or decoder apparatus). The decoding system 1202, which is generally a classical computing system, receives an error syndrome (also referred to as syndrome data) obtained from measurements of syndrome qubits. The error syndrome may comprise raw analogue measurement data, or it may alternatively be pre-processed (e.g. into digital format) by the control system 1204. The decoding system 1202 may be connected to the control system 1204 and receive the error syndrome via the control system 1204 as illustrated in Figure 12 (potentially via one or more additional intermediary systems), or in alternative examples the decoding system 1202 may be connected directly to the qubits 1206 and receive the error syndrome from the qubits 1206 (e.g. as raw analogue signals or digital measurement values). The decoding system 1202 uses a decoding process / algorithm to decode the error syndrome to determine a correction for an error state of the qubits 1206 associated with the error syndrome (i.e. an error state that causes the measured error syndrome). The decoding system 1202 comprises a plurality of distributed PEs, in the manner described extensively above. At compile-time, nodes may be assigned to processing elements of the hardware; for example so that each PE has a batch of nodes and / or such that each PE is grouped according to the grouping rules described herein. The decoder apparatus may comprise a computer memory. The computer memory may store instructions which, when implemented by the controller, cause the controller to perform any of the methods disclosed herein. One skilled in the art will appreciate that the quantum computing system 1200 may also comprise additional intermediary components positioned between the illustrated components, and that the illustrated components may be connected in a different configuration (e.g. the decoding system 1202 may be connected directly to the qubits 1206 as previously described). Figure 13 depicts a computer-readable medium according to the present disclosure. The various methods described above may be implemented by a computer program. The computer-readable medium may include computer code (e.g. instructions) 1310 which, when executed by a controller of a decoder apparatus, cause the controller to perform any of the methods disclosed herein. The steps of the methods described above may be performed in any suitable order. For example, the steps of the clustering algorithm 200 may be performed in any suitable order. The computer program and / or the code 1310 for performing such methods may be provided to an apparatus, such as a computer, on one or more computer readable media or, more generally, a computer program product 1300)), depicted in Figure 13. The computer readable media may be transitory or non-transitory. The one or more computer readable media 1300 could be, for example, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, ora propagation medium for data transmission, for example for downloading the code over the Internet. Alternatively, the one or more computer readable media could take the form of one or more physical computer readable media such as semiconductor or solid state memory, magnetic tape, a removable computer diskette, a random access memory (RAM), a read-only memory (ROM), a rigid magnetic disc, and an optical disk, such as a CD-ROM, CD-R / W or DVD. The instructions 1310 may also reside, completely or at least partially, within the memory and / or within the controller circuitry 1204 during execution thereof by the computing system 1210, the memory and the controller circuitry 1204 also constituting computer-readable storage media. It is to be understood that the above description is intended to be illustrative, and not restrictive. Many other implementations will be apparent to those of skill in the art upon reading and understanding the above description. Although the present disclosure has been described with reference to specific example implementations, it will be recognized that the disclosure is not limited to the implementations described, but can be practiced with modification and alteration within the spirit and scope of the appended claims. Accordingly, the specification and drawings are to be regarded in an illustrative sense rather than a restrictive sense. The scope of the disclosure should, therefore, be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.
Claims
1. A computer-implemented method for controlling the execution of a plurality of distributed processing elements, PEs, in a decoder apparatus of a quantum computer system, the quantum computer system further comprising a register of quantum devices; the method comprising:controlling the plurality of PEs to perform a distributed decoding algorithm according to an execution scheme;wherein the execution scheme identifies which PEs of the plurality of PEs should be executed at different stages of the distributed decoding algorithm, and each PE is identifiable within the execution scheme using one or more indexing variables, wherein each indexing variable of the one or more indexing variables is comprised of one or more bits, with each bit being associated with the execution of one or more PEs of the plurality of distributed PEs.
2. The method of claim 1, wherein each bit in the one or more indexing variables is associated with the execution of one or more PEs, independent of the other bits in the indexing variable.
3. The method of claim 1 or claim 2, wherein each PE is uniquely identifiable within the execution scheme using the one or more indexing variables.
4. The method of any preceding claim, wherein each bit of the one or more indexing variables is associated with a different respective subset of the plurality of distributed PEs.
5. The method of any preceding claim, wherein the one or more indexing variables comprise a first indexing variable and a second indexing variable.
6. The method of claim 5, wherein a particular PE is identifiable within the execution scheme via the combination of a first bit in the first indexing variable and a second bit in the second indexing variable; wherein the first bit is associated with a first subset of PEs and the second bit is associated with a second subset of PEs, and wherein the particular PE is present in both the first and the second subset of PEs.
7. The method of claim 5 or claim 6, wherein the one or more indexing variables comprise a first dimension indexing variable and a second dimension indexing variable; wherein the number of bits in the first dimension indexing variable corresponds with a number of PEs in a first dimension, and the number of bits in the second dimension indexing variable corresponds with a number of PEs in a second dimension.
8. The method of any preceding claim, the decoder apparatus further comprising a controller coupled to each PE of the plurality of PEs, and the method further comprising executing, by the controller, one or more PEs of the plurality of distributed PEs according to the execution scheme.
9. The method of claim 8, wherein each bit of the one or more indexing variables is associated with a different respective subset of the plurality of distributed PEs, and each bit further identifies a control line which couples the controller to its associated subset of PEs.
10. The method of claim 9, wherein the controller controls which PEs are activated according to the execution scheme thereby enabling the plurality of PEs to perform the distributed decoding algorithm; wherein controlling the PEs comprises determining which control lines are identified by the one or more indexing variables in the execution scheme at a particular stage of the decoding algorithm; and activating the identified control lines.
11. The method of claim 10, wherein each PE of the plurality of PEs is coupled to the controller via a different control line, and the one or more indexing variables is a single indexing variable comprised of a number of bits equal to the number of PEs in the plurality of PEs; wherein each bit in the single indexing variable uniquely identifies a respective PE.
12. The method of any preceding claim, further comprising:receiving, at the decoder apparatus, syndrome data representative of an error state of the quantum devices in the register of quantum devices, the syndrome data comprising a plurality of defects, wherein the syndrome data is representable as a decoding hypergraph comprising a plurality of nodes connected by hyperedges representing error mechanisms associated with the plurality of quantum devices, and wherein each PE of the plurality of distributed PEs is associated with one or more nodes of the decoding hypergraph; andwherein performing the distributed decoding algorithm comprises determining a correction for the error state.
13. The method of claim 12, further comprising:measuring a logical state encoded in the quantum devices of the quantum computer to obtain a logical state measurement; andapplying the correction for the error state to the logical state measurement.
14. The method of claim 12 or claim 13, wherein the decoding algorithm is a clustering algorithm, and the clustering algorithm grows and merges clusters of nodes based on the number of defects in eachcluster until a final cluster state is reached; optionally wherein determining a correction for the error state is based on the final cluster state.
15. A decoder apparatus comprising a controller, a plurality of distributed processing elements, PEs, and5 computer memory storing:an execution scheme which, when implemented by the controller, controls which PEsof the plurality of PEs are executed at different stages of a distributed decoding algorithm, and each PE is identifiable within the execution scheme using one or more indexing variables, wherein each indexing variable of the one or more indexing variables is comprised of one or more bits, with each bit being10 associated with the execution of one or more PEs of the plurality of distributed PE; andinstructions which, when implemented by the controller, cause the controller to perform the method of any preceding claim.
16. A quantum computer system comprising a register of quantum devices and a decoder apparatus15 according to claim 15.
17. A computer-readable medium comprising instructions which, when executed by a controller of a decoder apparatus, cause the controller to perform the method of any of claims 1 to 14.Application No: GB2412888.6 Examiner: Alessandro PotenzaClaims searched: 1-17Date of search: 31 January 2025Patents Act 1977: Search Report under Section 17Documents considered to be relevant: Category Relevant to claims Identity of document and passage or figure of particular relevance X 1-17 J. Sign. Proc. Syst. (2014) 77: 5-29, Boppu S et al., "Compact Code Generation for Tightly-Coupled Processor Arrays", DOI 10.1007 / sl 1265-014-0891-2 (BOPPU) see figure 1, 5, 9-10, and section 4.1.2 X 1-17 Springer, Cham, Advances in Soft Computing, MIC AI 2019, Lecture Notes in Computer Science, vol 11835, Velarde Martinez A, et al, "Parallel Graph Task Scheduling Based on the Internal Structure", DOI: https: / / doi.org / 10.l007 / 978-3-030-33749-0 22 (VELARDE MATINEZ) see figures 1-2 and sections 7 and 7.1 X 1-17 9th IEEE International Symposium on Applied Machine Intelligence and Informatics, SAMI 2011, Mados B et al, "Data Flow Graph Mapping Techniques of Computer Architecture with Data Driven Computation Model", DOI: 10.1109 / SAMI.2011.5738905 (MADOS) see figure 5 and section V A - Arxiv.org, Barber B et al, "A real-time, scalable, fast and highly resource efficient decoder for a quantum computer", 2023, available from https: / / arxiv.org / pdf / 2309.05558vl (BARBER) A - Arxiv.org, LIYANAGE N et al, "FPGA-based Distributed Union-Find Decoder for Surface Codes", 20 March 2024, available from https: / / arxiv.org / pdi72406.08491vl (LIYANAGE) A IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), Heer MJ et al, "Novel Union-Find-based Decoders for Scalable Quantum Error Correction on Systolic Arrays", 2023, DOI 10.1109 / IPDPSW59300.2023.00092 (HEER) A - US 2006 / 0156291 Al (DELL) Categories: X Document indicating lack of novelty or inventive A Document indicating technological background and / or state step of the art. Y Document indicating lack of inventive step if P Document published on or after the declared priority date but combined with one or more other documents of before the filing date of this invention.same category.& Member of the same patent family E Patent document published on or after, but with priority dateearlier than, the filing date of this application.Field of Search:International Classification:Subclass Subgroup Valid From G06N 0010 / 70 01 / 01 / 2022 G06F 0008 / 41 01 / 01 / 2018 G06F 0009 / 50 01 / 01 / 2006
Citation Information
Patent Citations
System and method for managing processor execution in a multiprocessor system
US20060156291A1