System and method for CMOS-based decoding
The QuBRIM architecture extends Ising machines to handle multi-body interactions, enabling efficient and exact solutions for LDPC decoding and other NP-complete problems by directly addressing higher-order terms in optimization functions.
Patent Information
- Application Number
- PCT/US2024/045962
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-22
- Filing Date
- 2024-09-10
- Publication Date
- 2025-09-04
AI Technical Summary
Existing Ising machines are limited to quadratic spin variables and struggle to efficiently solve NP-complete problems like LDPC decoding due to the need for order reduction, which increases hardware and algorithmic overhead.
A QuBRIM architecture that implements multi-body interactions using a network of computation nodes, logic gates, coupling units, and digital-to-analog converters to directly handle cubic and higher-order terms in optimization functions, allowing for efficient decoding of LDPC codes and other NP-complete problems.
The QuBRIM architecture provides exact solutions to LDPC decoding and other NP-complete problems without the overhead of order reduction, enhancing scalability and efficiency.
Smart Images

Figure US2024045962_04092025_PF_FP_ABST
Abstract
Description
Attorney docket # 204606-0168-00WO SYSTEM AND METHOD FOR CMOS-BASED DECODING CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to US Provisional Patent Application No.63 / 581,787, filed on September 11, 2023, and US Provisional Patent Application No.63 / 623,561, filed on January 22, 2024, both of which are incorporated herein by reference in their entirety. STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0002] This invention was made with government support under FA8650-23-C-7312 awarded by the Air Force Research Laboratory. The government has certain rights in the invention. BACKGROUND OF THE INVENTION
[0003] The Ising model was originally proposed to study the interactions between magnetic dipoles, also known as spins, in ferromagnetism (see E. Ising, Zeitschrift für Physik, 1925). The interactions between the coupled spins in a ferromagnetic material affect the dynamics of the system, which naturally minimizes the system energy represented by the Ising Hamiltonian (see T. Wang et al., arXiv:1903.07163, 2019). Although finding the minimum of the Ising Hamiltonian is difficult using conventional algorithms for large problems, physical systems naturally minimize their own Hamiltonian with relative ease. However, of particular interest is that many non-deterministic polynomial time (NP) hard combinatorial optimization problems (COPs) can be modeled using the Ising Hamiltonian (see A. Lucas, Frontiers in Physics, 2014). As a result, physical implementations of Ising machines are attracting increased interest for their potential to solve COPs with high efficiency. Several implementations of such Ising machines exist, such as D-wave's quantum-based annealers (see Z. Bian, et al, 2010; and "The D-Wave 2000Q Quantum Computer.") and coherent Ising machines based on optical cavities (see H.Attorney docket # 204606-0168-00WO Takesue, et al., Journal of the Physical Society of Japan, 2019; T. Inagaki, et al., Science, 2016; Y. Yamamoto, et al., npj Quantum Information, 2017; P. L. McMahon, et al., Science, 2016; K. Takata, et al., Scientific Reports, 2016; and F. Böhm, et al., Nature Communications, 2019)
[0004] Despite the improved efficiency of Ising machines, scaling such systems is challenging. Recently, more focus has been shifted towards CMOS-based Ising machines to improve scalability and cost-efficiency. For example, coupled oscillator networks have been used to implement Ising machines (see T. Wang et al., arXiv:1903.07163, 2019; S. Dutta, et al., Nature Electronics, 2021; and W. Moy, et al., Nature Electronics, 2022). One such implementation uses a system of coupled oscillators, that have been shown to minimize the Ising Hamiltonian when the oscillator phases are binarized using second order sub-harmonic injection locking (SHIL) (see T. Wang et al., arXiv:1903.07163, 2019). However, this design uses bulky LC oscillators, and the proposed implementation requires externally supplied injection locking signals.
[0005] Bistable Resistively-coupled Ising Machine (BRIM) (see R. Afoakwa, et al., in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), Feb. 2021) is an all-to-all CMOS-based Ising machine design. In BRIM, the spins are generated using bistable nodes, each designed as a capacitor controlled by a feedback circuit. The spins are represented as voltage values rather than phases of an oscillator. A modified design has been disclosed in QuBRIM (see Y. Zhang, et al., in Proceedings of the 41st IEEE / ACM International Conference on Computer-Aided Design, Dec.2022) which uses quantized nodal interactions, further reducing the design complexity and allowing for more straightforward node interactions.
[0006] Another family of Ising machines is based on simulated annealing where spins are represented as bits in memory (see M. Yamaoka, et al., in 2015 IEEE International Solid-State Circuits Conference - (ISSCC) Digest of Technical Papers, Feb 2015; T. Takemoto, et al., in IEEE International Solid- State Circuits Conference, Feb.2019; and K. Yamamoto, et al., IEEE Journal of Solid-State Circuits, Jan.2021). In these designs, the Ising model is implemented using algorithms rather than physical interactions. The randomness in these systems is exploited to implement the annealing process. For more detailed information on Ising machines and a compilation of various implementations see R. Afoakwa, et al., in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), Feb.2021.
[0007] Despite the significant potential of Ising machines, preliminary research has mostly focused on Max-Cut problems, which are easily mapped to an Ising machine (see T. Wang et al.,Attorney docket # 204606-0168-00WO arXiv:1903.07163, 2019; R. Afoakwa, et al., in 2021 IEEE International Symposium on High- Performance Computer Architecture (HPCA), Feb.2021; and Y. Zhang, et al., in Proceedings of the 41st IEEE / ACM International Conference on Computer-Aided Design, Dec.2022). Similarly, LDPC (see R. Gallager, IRE Transactions on Information Theory, 1962) decoding can be treated as a COP (see A. Hareedy, et al., IEEE Transactions on Information Theory, 2019) and mapped to Ising machines. Since LDPC decoding is an NP-complete problem, conventional computational methods do not guarantee exact solutions to LDPC decoding for long codewords (or large number of variables). However, by mapping the LDPC decoding problem to the energy landscape of an Ising machine, exact solutions can be obtained when the Ising Hamiltonian is minimized (see M. Tawada, et al., in 2020 IEEE International Conference on Consumer Electronics (ICCE), Jan.2020). Therefore, using Ising machines to solve LDPC decoding can provide higher quality solutions than existing methods, and with potentially better scalability.
[0008] The primary challenge in the Ising machine formulation is that it is limited to functions quadratic to the spin variables. In the quadratic unconstrained binary optimization (QUBO) formulation of an LDPC COP, higher order terms exist that cannot be directly mapped to an Ising machine. A methodology has previously been proposed in M. Tawada, et al., in 2020 IEEE International Conference on Consumer Electronics (ICCE), Jan.2020 to convert the LDPC decoding COP into a QUBO problem for Ising machines by implementing order reduction for products that are higher than quadratic. The order reduction adds significant overhead both to the hardware implementation and speed of the algorithm.
[0009] Therefore, there is a need in the art for a framework for implementing multi-body interactions using the QuBRIM architecture. Such a framework extends the applicability of Ising machines from well-known QUBO problems to those that include cubic, quartic, and higher cost terms in the optimization function. The extension of Ising machines to multi-body makes it suitable (and practical) for not only LDPC decoding applications but other NP-complete problems such as Boolean Satisfiability problems (e.g.3-SAT), Traveling Salesman Problem (TSP), etc. Such a methodology would allow for solving binary optimization problems with higher-than-quadratic cost terms directly without order reduction.Attorney docket # 204606-0168-00WO SUMMARY OF THE INVENTION
[0010] In one aspect, a network comprises a plurality of computation nodes, a plurality of logic gates, a plurality of coupling units, a plurality of digital-to-analog converters, each of the plurality of computation nodes comprising an input configured to received currents from the coupling units in the network and currents from at least one of the digital-to-analog converters, an output configured to generate at least two discrete output voltages, a capacitor configured to store internal state as a voltage, a comparator configured to compare the voltage on the capacitor against a threshold to produce a first discrete voltage value at its output when the voltage is above the threshold, and a second discrete voltage value at its output when the voltage is below the threshold, with the output connected to the output of the computation node, and a current conveyor circuit having an input connected to the input of the computation node and an output connected to the capacitor, the current conveyor configured to hold its input at a constant voltage and mirror a scaled version of the current received at the input into the capacitor.
[0011] In one embodiment, each logic gate comprises an input connected to the output of at least one computation node, an output configured to generate one of at least two discrete output voltages, the output connected to an input of at least one coupling unit of the plurality of coupling units. In one embodiment, at least one logic gate is selected from an XOR gate or an XNOR gate. In one embodiment, each coupling unit comprises an input connected to the voltage output of at least one logic gate of the plurality of logic gates, an output configured to generate one of at least two discrete output current values, the output connected to an input of at least one computation node of the plurality of computation nodes, and means to convert voltage values at the input to currents proportional to the input voltage at the output of the coupling unit.
[0012] In one embodiment, each digital-to analog converter comprises an input configured to receive a digital sample representing a noisy message or parity bit received by a receiver circuit in a communication system, means to produce current proportional to the digital input and supply it to the output of the digital-to-analog converter, and an output configured to supply the current to the input of at least one computation node of the plurality of computation nodes. In one embodiment, at least one coupling unit of the plurality of coupling units comprises aAttorney docket # 204606-0168-00WO resistive element configured to convert an input voltage to the coupling unit into a current supplied at its output.
[0013] In one embodiment, at least one coupling unit of the plurality of coupling units comprises at least one transistor configured as a current source, with a drain terminal connected to the output of the coupling unit through switches controlled by the voltage at the input of the coupling unit. In one embodiment, the network further comprises a spin perturbation control circuit configured to change a polarity of capacitor voltages in at least one computation node during the network’s operation. In one embodiment, the spin perturbation control circuit is configured to generate perturbation signals from a random source or a pseudo-random source. In one embodiment, the spin perturbation control circuit is configured to generate perturbation signals such that the rate of spin perturbation events diminishes over time. In one embodiment, the network is configured such that the rate of spin perturbation events has an exponential decay over time. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The foregoing purposes and features, as well as other purposes and features, will become apparent with reference to the description and accompanying figures below, which are included to provide an understanding of the invention and constitute a part of the specification, in which like numerals represent like elements, and in which: Fig.1 is an illustration of the working principal of a QuBRIM node. The currents from the other nodes are summed through a current conveyor. The summed current is mirrored and charges the node capacitor. Fig.2 is a high level block diagram of a QuBRIM system containing 3 nodes for the sake of brevity. Fig.3 is a Tanner graph for the LDPC code represented by a parity matrix. Fig.4 is an illustration of a node of a QuBRIM system for LDPC decoding. Fig.5 is a Simplified QuBRIM diagram for an LDPC decoder, with coupling resistors Rcat the output of each XOR gate omitted for brevity.Attorney docket # 204606-0168-00WO Fig.6 is a circuit design of a single-ended QuBRIM node, which includes a class-AB current conveyor, node capacitor C, and comparator circuit. Fig.7 is a diagram of a pseudo-resistor used in the coupling units of the QuBRIM system using PMOS gates. Fig.8 is a QuBRIM LDPC decoder system diagram. Fig.9 is an exemplary computing device. Fig.10 is a graph of QuBRIM internal node voltages for (128,32) LDPC decoding with SNR = 4 dB. Fig.11 is a set of histograms of the time required for different QuBRIM LDPC implementations to reach a stable state. The results were obtained for 100 different codewords. Vertical dashed lines represent the mean. Fig.12 is a graph showing the change in BER across 100 Monte Carlo samples for (32,8) LDPC QuBRIM. Mean = 0.6225, standard deviation = 0.1325. Fig.13 is a set of graphs depicting a demonstration of the annealing schedule using bit flips and thermal noise to decode one message. Fig.14 is a graph of BER vs. SNR for 1 / 2 coding rate (672,336) LDPC. Fig.15 is a set of graphs showing an improved annealing schedule with higher frequency bit flips. The top plot shows the steps taken before the early termination signal is triggered for each received message, and the bottom plot shows the annealing schedule. Fig.16 is a graph showing a BER comparison for 1 / 2 coding rate (672,336) LDPC between the disclosed design QuBRIM (QB) and other LDPC algorithms using (1) iteration and (4) iterations for each.Attorney docket # 204606-0168-00WO DETAILED DESCRIPTION
[0015] It is to be understood that the figures and descriptions of the present invention have been simplified to illustrate elements that are relevant for a clear understanding of the present invention, while eliminating, for the purpose of clarity, many other elements found in related systems and methods. Those of ordinary skill in the art may recognize that other elements and / or steps are desirable and / or required in implementing the present invention. However, because such elements and steps are well known in the art, and because they do not facilitate a better understanding of the present invention, a discussion of such elements and steps is not provided herein. The disclosure herein is directed to all such variations and modifications to such elements and methods known to those skilled in the art.
[0016] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, exemplary methods and materials are described.
[0017] As used herein, each of the following terms has the meaning associated with it in this section.
[0018] The articles “a” and “an” are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element.
[0019] “About” as used herein when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of ±20%, ±10%, ±5%, ±1%, and ±0.1% from the specified value, as such variations are appropriate.
[0020] Throughout this disclosure, various aspects of the invention can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example,Attorney docket # 204606-0168-00WO description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 2.7, 3, 4, 5, 5.3, 6 and any whole and partial increments therebetween. This applies regardless of the breadth of the range.
[0021] In some aspects of the present invention, software executing the instructions provided herein may be stored on a non-transitory computer-readable medium, wherein the software performs some or all of the steps of the present invention when executed on a processor.
[0022] Aspects of the invention relate to algorithms executed in computer software. Though certain embodiments may be described as written in particular programming languages, or executed on particular operating systems or computing platforms, it is understood that the system and method of the present invention is not limited to any particular computing language, platform, or combination thereof. Software executing the algorithms described herein may be written in any programming language known in the art, compiled or interpreted, including but not limited to C, C++, C#, Objective-C, Java, JavaScript, MATLAB, Python, PHP, Perl, Ruby, or Visual Basic. It is further understood that elements of the present invention may be executed on any acceptable computing platform, including but not limited to a server, a cloud instance, a workstation, a thin client, a mobile device, an embedded microcontroller, a television, or any other suitable computing device known in the art. The Ising Model
[0023] The Hamiltonian ^^σ^ for an Ising model can be written as (see T. Wang et al., arXiv:1903.07163, 2019): ^^σ^ = −^^^^ ⋅ σ^σ^ − ^ℎ^ ⋅ σ^where ^^^ represents the coefficient of coupling interaction between the spins σ^ and σ^ ,σ ∈{−1,1}, and ℎ^ is the coefficient of the external magnetic field. In the absence of an externalmagnetic field, the Hamiltonian becomesAttorney docket # 204606-0168-00WO ^^σ^ = −^^^^ ⋅ σ^σ^^,^
[0024] In the case of Max-Cut problems, the maximum cut can be defined as ar ^ ^^^ ^^where ^ ∈ ^ା,^ି represents the vertices of a graph with edges E, and ^^^ represents the weightof each edge. The maximum cut equation in Equation 3 can be easily mapped to the Ising equation of Equation 2. Then, by minimizing the resulting Ising Hamiltonian to the ground state, the solution to the NP-hard problem of Max-Cut can be obtained. Such solutions are not guaranteed to be the most optimal, however, high quality solutions are obtained significantly faster than by using conventional computation. QuBRIM
[0025] For the case of QuBRIM, the dynamics of the system can be explained by first referring to the QuBRIM node, as shown in Fig.1, which represents a single node ^^. In this figure, the currents from the neighboring nodes are summed, then mirrored into the node capacitor 101, charging it into the spin voltage ^^. The spin voltage ^^is quantized as ^(^^) through a comparator 102 where^(^ ) = −1, ^^ < 0^ ^
[0026] The quantized spin voltage ^(^^)is then coupled back to the neighbouring nodes through the resistors 103 which controls the coupling weight and can be written asAttorney docket # 204606-0168-00WO ^^^^ = ,ห^^^ห5where R is a constant scaling term. In this system, the polarity of ^(^^) represents the spin valueσ ∈ {−1,1}, and the sign of ^^^ controls whether the coupling is negative or positive. Due toquantization of the node voltages, the polarity of the coupling ^^^൫^^^൯ can be implemented with a simple inverter gate, or the lack thereof.
[0027] The current i going into the node capacitor C with voltage ^^ can be expressed as^^^ ^^^൫^ ൯ ∗ ^൫^ ൯= 1^ ^^ ^^^ ^ ^^^
[0028] Due to current mirroring, the current from neighboring nodes ^^ஷ^does not depend on ^^,but only on the quantized voltages ^൫^^ஷ^൯. By substituting Equation 5, Equation 6 becomes:1^^ = ∗ ^൫^^൯where τ = ^^. From Equation 7, a Lyapunov function exists of the form^(^) = −^^^^^(^^)^൫^^൯
[0029] This Lyapunov function in Equation 8 has the same form as the Ising Hamiltonian in Equation 2. Furthermore, carrying out Lyapunov stability analysis as in Y. Zhang, et al., in Proceedings of the 41st IEEE / ACM International Conference on Computer-Aided Design, Dec. 2022 demonstrates that the time derivative of Equation 8 isAttorney docket # 204606-0168-00WO ^^(^) ( )= −τ^ ^^^ ^^ ^^^^ ^ ∗ ^ ≤ 0^ ^^ ^^which is a non-positive value, therefore the Ising machine represented by Equation 8 evolves toward equilibrium by lowering the system Hamiltonian ^(^).
[0030] A block diagram of an exemplary QuBRIM system is shown in Fig.2. The QuBRIM nodes 201 are labelled as^^, and the coupling units (CUs) 202 controlling the interactionsbetween neighbouring nodes (i.e., the coupling resistors) are labelled as ^ ^^^. Auxiliary andcontrol circuit blocks are also shown in the figure. Multi-body Interaction through QuBRIM
[0031] In QuBRIM, multi-body interactions are carried out by multiplication of the node variables. Because the node variables are quantized into ^(^^), the multiplication is carried out by using either XOR or XNOR gates. The choice of logic gate depends on the number of spin variables in each term and the polarity of the term. In some embodiments, for terms with an even number of spin variables, their multiplication is carried out with an XNOR gate, while for terms with an odd number of spin variables, an XOR gate is used. This is readily verifiable with a truth table with the assumption that logic ‘0’ represents negative spin (-1), and logic ‘1’ represents positive spin (1). The negative and positive sign in front of each term represents the polarity of the coupling coefficient. The sign is implemented by the presence (or lack of) an inverter following the logic gate to implement negative (positive) coupling. A negative coupling will convert an XOR to XNOR and vice versa. For example, consider the following objectivefunction for a system of 4 nodes^ = −^^ଶଷσ^σଶσଷ − ^^ଶସσ^σଶσସ − ^^ଷସσ^σଷσସ − ^ଶଷସσଶσଷσସEquation 10 where each term is cubic and represents multi-body interaction between 3 spins. In order to minimize the objective function H, a QuBRIM system can be constructed fromAttorney docket # 204606-0168-00WO ^^ ^^^^ ^^൫^^൯,^(^^)^ − ^^ெ^ ^^^ = ^^^^^^where ^^^^ = ^ / ^^^^,^^^^(^, ^) is a 2-input XNOR logic gate with logic ‘0’ and logic ‘1’outputs of 0V andௗ^ௗ, respectively, and^^ெis a common-mode voltage typically equal to the half of the power supply voltageௗ^ௗ.
[0032] Note that the numerator term under the double summation in Equation 11 takes only twodiscrete values േ ௗ^ௗ / 2 (i.e., േ1 times a constant) and its similarity with the numerator term inthe set of differential equations in Equation 6.
[0033] The QuBRIM construction process (i.e., construction of Equation 11 from Equation 10 iscarried out by using the following construction rule: ௗ௩^ ^ புௗ௧ = −த ப^^ , 1 ≤ ^ ≤ ^ followed bysubstituting σ^with ^(^^), as further explained in
[0034] Objective functions including terms higher than cubic, or even a mix of different order- terms, can be minimized using the same approach. LDPC
[0035] In the communication system considered, signals are represented as base-band signalsunder binary phase-shift keying, where each bit ^^ is represented as ^^ ∈ {^1,−1}. The receivedsignal r is received from an Additive White Gaussian Noise (AWGN) channel where bits aredistorted according to^^ = ^^ ^ ^^Equation 12 where ^^is the noise component following a normal distribution.
[0036] An (n,k) LDPC is an error-correcting code in which a codeword w of length n is generated from a message v of length k. The codeword is generated with a number of parity bitsAttorney docket # 204606-0168-00WO^ = ^ − ^. For the error-correcting mechanism of the LDPC code, an m-by-n parity matrix P isdefined such that the summation of certain bits (mod 2) of the codeword is always zero, as stated in matrix format Pw = 0Equation 13 In order to generate codewords that satisfy this condition, a k-by-n generator matrix G is defined such that PGT = 0m×kEquation 14
[0037] The codeword w is obtained from the message v according to w= GT^Equation 15
[0038] The goal of LDPC decoding is to obtain the codeword w that best satisfies Equation 13 given the received noisy message r. The process of obtaining the codeword with the highest probability of being sent is known as maximum likelihood decoding. The decoding of LDPC using maximum likelihood is a problem known to be NP-hard (see E. Berlekamp, et al., IEEE Transactions on Information Theory, 1978; J. Bruck, et al., IEEE Transactions on Information Theory, 1990). As a result, alternative methods such as those based on belief-propagation or its derivatives (see J. Chen, et al., IEEE Transactions on Communications, 2005; D. MacKay, IEEE Transactions on Information Theory, 1999) are employed. While these approximate algorithms have historically been easier to implement in practical circuits, these algorithms do not provide an exact solution and lack theoretical guarantees (see A. Riazanov, et al., in 2018 Annual American Control Conference (ACC), 2018).
[0039] The present invention relates to a network of coupled computation nodes implemented in Complementary Metal-oxide-semiconductor (CMOS) technology for decoding Low DensityAttorney docket # 204606-0168-00WO Parity Codes (LDPC). In one embodiment of the present invention, the network of coupled computation nodes comprises computation nodes, where each computation node comprises an input receiving currents from coupling units of the network, a current summing circuit connected to the input of the node, a capacitor to store a voltage state of the node connected to the output of the current summing circuit, a comparator that compares the voltage state on the capacitor against a threshold producing two distinct output voltage values depending on the comparison result, and the output of the comparator connected to at least one output of the node. The network of coupled computation nodes may further comprise an array of logic gates performing mod 2 addition (e.g., exclusive NOR or exclusive OR gates) each receiving at its input terminals the output from at least one computation node. The network of coupled computation nodes further comprises an array of coupling units where at least one of the coupling units comprising at least one input to receive at least one output from one of the logic gates, means to generate current proportional to the voltage at the input of the coupling unit and produce the generated current at the output of the coupling unit that is then connected to the input of at least one node in the network. The network further comprises an array of digital-to-analog converters (DAC) that receive noisy signal from the communication channel and set the initial voltage on the capacitors of the computation nodes to values proportional to the received signal. The outputs from the DACs are also used to produce currents proportional to the noisy received signal that are sent to the inputs of the corresponding computation nodes. The LDPC decoding operation of the network comprising N computation nodes is as follows: (1) Convert an N-sample-long word containing message bits and parity bits received from a noisy communication channel to analog voltages and charge the capacitors in the computation nodes to the corresponding voltage values. (2) Enable inputs and outputs of the computation nodes permitting interactions between the nodes through the array of mod 2-addition logic gates and the array of coupling units as well as permitting nodes to receive currents from the DAC array proportional to the codeword received from the noisy communication channel. (3) After a time period of^^௧^^(where^^௧^^may be determined from the number of nodes N, capacitance value of the nodes’ capacitor as well as peak current generated by the coupling units), the output of the computation nodes is sampled to produce LDPC decoder output.Attorney docket # 204606-0168-00WO
[0040] In another embodiment of the present invention, a network of computation nodes described above also comprises means to generate N perturbation signals that are sent to the corresponding computation nodes during the network's operation (i.e., the nodal interaction phase described in step (2) above), where each perturbation signal contains pulses whose rising (or falling edge) indicates a beginning of perturbation event. During each perturbation event, a capacitor in a computation node is charged to a voltage value complementary to the voltage currently present at the output of the computation node. The perturbation signals may be generated by a source of N random signals or a pseudo-random source and the perturbation signals may be generated such that the average frequency of perturbation events diminishes over time. In one example, the average frequency of perturbation events per node may start at 0.1 perturbations / ns and exponentially decay to 0 perturbation / ns at time instance^^௧^^). In yet another embodiment of the present invention, the perturbation event comprises of charging the capacitor to a voltage value chosen at random.
[0041] In another embodiment of the present invention, a network of computation nodes for decoding LDPC codes, as described in the first embodiment, further comprises a network of logic gates that receives outputs from the computation nodes during the network’s operation (i.e., during step (2) above), performs parity checks on the outputs of the computation nodes as defined by the LDPC code parity matrix and generates an early termination signal when all parity checks are satisfied. The early termination signal is then used to terminate the nodal interaction phase in step (2) (i.e., defining^^௧^^as described in step (3) above).
[0042] For the purpose of implementing the LDPC decoder on a QuBRIM Ising machine, the constraint on the mod 2 operation can be expressed as ^ି^ ^ି^^ ^^^^ 0
[0043] Next, a spin ^^is associated to each bit in a codeword w such thatAttorney docket # 204606-0168-00WO ^−1, ^^ = 0^ = ^ 1, ^^ = 1
[0044] A vector ^^,^of length ^(^)containing the elements of the multiplication of ^^by the non- zero terms of^^,^is defined. The constraint in Equation 16 can be written as the following Lyapunov function ^ି^ ^(^)ି^^ = ^ ^ ^ ^ ∗ (−1)^(^)ା^
[0045] Equation 18 always returns -1 for even terms in the summation, ensuring that the lowest energy level satisfies the parity condition set by the parity matrix. To complete the Lyapunov function, an objective term is added that depends on the received noisy message ^= α^(^ − )ଶ^^^. ^ ^^where α is a scale factor and ^^is the received message bit signal. The Lyapunov function for the LDPC code then becomes ^ି^ ^(^)ି^^ ^ ^ ^ ^ ^ ^ ^ ଶ
[0046] It is clear from Equation 20 that if a system that can be described with this Lyapunov function is found, the system would exhibit stable local minima whose corresponding spin state would satisfy the condition in Equation 16. This system is determined by calculating partial derivatives of H with respect to each spin and using the following relationshipAttorney docket # 204606-0168-00WO ^^^ ∂^(^)^^ = −^∂^^where k is some constant. A set of differential equations can be written, describing the system that satisfies Equation 20 and Equation 21, where ^^is a capacitor voltage corresponding to thespin ^^^^^ ∂^(^)^^ = −^∂^^Equation 22
[0047] The previous methodology is applied on an example LDPC code for clarification. Consider a (6,3) LDPC with a parity matrix P, where 11 1 1 0 0^ = ^0 0 1 1 0 1൩1 0 0 1 1 0Equation 23 The Tanner graph for this code is shown in Fig. 3 with 6 bit nodes 301 and 3 check nodes 302.The codeword for this LDPC is of the form^ = ^^^ ^ଶ ^ଷ ^ସ ^ହ ^^^்Equation 24 The Lyapunov function that has local minima for the permitted codewords is ^(^) = + + + ^ ^(^^ − ^^)ଶEquation 25
[0048] By equating ௗ௩^ = −^ பு(௩), in the case of ^^:Attorney docket # 204606-0168-00WO ^^^^^ = ^(^ଶ^ଷ^ସ − ^ସ^ହ) + ^^^^where ^^ = 2^ and ^ = 1 / ^^. The same operation is repeated for ^ଶ: ^^ to obtain a system ofdifferential equations. Although this system of differential equations minimizes the energy ^(^) in Equation 25, it does not guarantee that the stable solution will have nodal voltages ^^that are at the vertexes, e.g. +1 / -1 V. Additionally, the multiplication operation of continuous values in Equation 26 in analog electronics is a challenging and expensive task. Both of these challenges are resolved by quantizing all nodal states to the nearest value of +1 and -1 before they areapplied in all of the multiplicative terms, as in^ଶ^ଷ^ସ → ^(^ଶ)^(^ଷ)^(^ସ)Equation 27 resulting in a set of differential equations describing a QuBRIM system where ^(^^) is as in Equation 4. Once the states are quantized to +1 / -1, the multiplication operation representing multi-body interactions may be implemented using XOR logic gates. The system of differential equations then becomesAttorney docket # 204606-0168-00WO ^^^ ( ) ( ) ( ) (^^ = ^൫^ ^ଶ ^ ^ଷ ^ ^ସ − ^ ^ସ)^(^ହ)൯ + ^^^^
[0049] The choice of XOR gates follows the discussion above. In the LDPC system of equations in Equation 28, all the positive terms contain an odd number of variables, which are implemented with XOR gates. All the negative terms contain an even number of variables. Although terms with an even number of variables can be implemented with XNOR gates, an inverter may be needed to implement the negative coupling. This inverter effectively converts the XNOR gates to XOR gates, and so therefore, all terms in Equation 28 may be implemented with XOR gates to account for the number of variables and the polarity of the coupling. A summary of the LDPC decoding algorithm on QuBRIM is shown in Algorithm 1 below.Attorney docket # 204606-0168-00WO Algorithm 1: QuBRIM LDPC decoding summary Input: Received signal r Output: Decoded codeword w 1 Initialize weights {k, kα} 2 for Iteration i = 1 to 10 do 3 Initialize nodes’ v with received signal r4 Update node values: vi = ^∑^^௪^^ ൫∏^^^^^(^) ^^൫^^ஷ^൯൯ + ^ఈ^^^^5 Perform annealing:6 Hard decision ^^ = ^1, ^^ > 07 Check8 Circuit Level Implementation
[0050] The circuit level implementation for the QuBRIM LDPC decoder design is presented in this section. The diagram of a QuBRIM node for node ^^in the system described by Equation 28is shown in Fig. 4. As a reference, the differential equation for node ^^ is^^^^^ = ^൫^(^ଶ)^(^ଷ)^(^ସ) − ^(^ସ)^(^ହ)൯ + ^^^^
[0051] The first two products in Equation 29 are implemented using XOR gates to represent the mod 2 constraint of the parity matrix. The XOR gates are connected to the current summing node through coupling resistance ^^. The last term ^^^^, hereon referred to as a “bias” term, is a function of the noisy received signal ^^. The bias term is implemented as a binary-weighted current-steering DAC, whose design is adopted from W.-T. Lin, et al., in 2013 IEEE International Solid-State Circuits Conference Digest of Technical Papers, 2013; with modifications described below. Unlike the design in Lin, that drives a resistive load producing a voltage output, the disclosed system is a current-output DAC steering output currents into a virtual ground resulting in much reduced power and increased linearity (e.g., for the peak outputAttorney docket # 204606-0168-00WOcurrent of 14 μ^ and ^^ெ = ௗ^ௗ / 2 = 0.55^, the static power of each DAC in this design is only15.4 μ^). Another modification from Lin is the use of programmable unit / reference current toenable an annealing schedule with time-varying ^^as described herein, and shown in Fig.15.
[0052] The remaining nodes in the system described by Equation 28 are implemented similar to node ^^. The system diagram containing the XOR network for all nodes is shown in Fig.5. Note the similarity between the XOR network in Fig.5 and the parity matrix P in Equation 23. The general methodology for converting a parity matrix to an XOR network in QuBRIM is as follows: For each bit ‘1’ in column i and row j of the parity matrix P, place an XOR gate with its output connected to the input of node Ni. The inputs of this particular XOR gate are the outputsof the nodes located at the positions of bit ‘1’ in row j where the column index is ≠ ^. Based onthe previous discussion, the column degree ^^^∘(number of 1’s in a column) column idetermines the number of XOR gates connected to each node Ni. The row degree ^^^∘of eachrow j determines the size of XOR gates in that row, where the size of each XOR gate is ^^^∘^ −1. If there is a row with degree 2, then the gates become simple inverters or buffers depending onthe polarity of the coupling.
[0053] The current summing circuit for the QuBRIM node is implemented as a class AB current conveyor formed by the transistors ^^−^଼, as shown in Fig.6. Contrary to the original QuBRIM implementation (see Y.et al., in Proceedings of the 41st IEEE / ACM International Conference on Computer-Aided Design, Dec.2022), a single-ended node design is used in this work. This saves roughly half the area required for the QuBRIM implementation. Inthe current conveyor, the biasing voltages for ^^,^ସ are set by the current sources with current^^. Using current mirrors, the same amount ofpasses in the branches ^^,^ସ and^ହ,^ଶ,^ଷ,^^. The input current is then mirrored to the output branchand stored in thecapacitor at node Z. Since Node X is fixed atௗ^ௗ / 2 and node Y is a virtual ground, the current from the connected XOR gates depends only on the neighboring nodes. A class AB current conveyor was originally chosen for this circuit due to its high slew-rate capability. Additionally, it allows the input and output currents to be higher than the bias current, which is suitable in case of a large fan-in to the QuBRIM input.Attorney docket # 204606-0168-00WO
[0054] Since the output current has the opposite polarity of the input current, thecomparator / quantizer is implemented as an inverter to obtain ^(^^)from(−^^). A 3-stagetapered inverter chain is used to ensure full voltage swing when driving the XOR gates. The riseand fall times are balanced through proper gate sizing. One can easily obtain ൫−^( ^^)൯ from thesecond stage of the 3-stage inverter if needed.
[0055] The coupling resistance between the XOR gates and QuBRIM nodes is designed as shown in Fig.7. Two PMOS transistors 701, 702 are connected together forming a programmable pseudo-resistor (see E. Guglielmi, et al., IEEE Journal of Solid-State Circuits, 2020) which is more area efficient than physical resistors. The resistance value of the coupling resistors are set by^^௧^^which is programmed with a Digital-to-Analog Converter (DAC). Because the voltage magnitude across the terminals of the pseudo-resistor is fixed atௗ^ௗ / 2, the resistor exhibits good linearity. Thanks to the current coupling mode of the QuBRIM design, large voltage swings across the terminals of the pseudo-resistor are avoided which would otherwise cause non-linearity errors.
[0056] The system diagram for the complete QuBRIM decoder is shown in Fig 8. The decoder contains QuBRIM nodes 801, as in Fig.4, with their number equal to the code length n. The system includes an XOR gate and a CU block 802 for each bit ‘1’ in the parity matrix. Because all CUs have the same resistance value, only one DAC is needed to program all CUs.
[0057] As mentioned at the beginning of this section, the bias terms may be implemented as current-steering DACs, one for each node. The same array of DACs may also be used to set the initial state of the spin capacitors, by integrating the output current from corresponding DACs onto the nodal capacitors for a fixed time interval. For the purpose of this disclosure, 8-bit resolution DACs with settling time of 1ns were used, adding an acceptable 4% overhead to the overall decoding time.
[0058] An early termination signal can be implemented by checking the parity condition of the QuBRIM decoder. The global minimum is satisfied when all the parity checks result in an output of 0. The parity condition for each row in the LDPC parity matrix can be calculated by adding a 2-input XOR gate for each row in the XOR network in Fig.5. The added gate in each row will XOR the output of one XOR gate in that row with the spin variable in the column associatedAttorney docket # 204606-0168-00WO with it. This process needs to be done only once per row and choosing any XOR gate in a given row will give the same result. The early termination signal can be easily produced by NORing the outputs of all the added 2-input XOR gates. Additionally, an annealing control block is needed to carry out the annealing schedule discussed herein. The annealing block contains random number generators (RNGs) that can be implemented as linear feedback shift registers.
[0059] Parts of this invention are described as software running on a computing device. Though software described herein may be disclosed as operating on one particular computing device (e.g. a dedicated server or a workstation), it is understood in the art that software is intrinsically portable and that most software running on a dedicated server may also be run, for the purposes of the present invention, on any of a wide range of devices including desktop or mobile devices, laptops, tablets, smartphones, watches, wearable electronics or other wireless digital / cellular phones, televisions, cloud instances, embedded microcontrollers, thin client devices, or any other suitable computing device known in the art.
[0060] Similarly, parts of this invention are described as communicating over a variety of wireless or wired computer networks. For the purposes of this invention, the words “network”, “networked”, and “networking” are understood to encompass wired Ethernet, fiber optic connections, wireless connections including any of the various 802.11 standards, cellular WAN infrastructures such as 3G, 4G / LTE, or 5G networks, Bluetooth®, Bluetooth® Low Energy (BLE) or Zigbee® communication links, or any other method by which one electronic device is capable of communicating with another. In some embodiments, elements of the networked portion of the invention may be implemented over a Virtual Private Network (VPN).
[0061] Fig.9 and the following discussion are intended to provide a brief, general description of a suitable computing environment in which the invention may be implemented. While the invention is described above in the general context of program modules that execute in conjunction with an application program that runs on an operating system on a computer, those skilled in the art will recognize that the invention may also be implemented in combination with other program modules.
[0062] Generally, program modules include routines, programs, components, data structures, and other types of structures that perform particular tasks or implement particular abstract data types.Attorney docket # 204606-0168-00WO Moreover, those skilled in the art will appreciate that the invention may be practiced with other computer system configurations, including hand-held devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, and the like. The invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote memory storage devices.
[0063] Fig.9 depicts an illustrative computer architecture for a computer 900 for practicing the various embodiments of the invention. The computer architecture shown in Fig.9 illustrates a conventional personal computer, including a central processing unit 950 (“CPU”), a system memory905, including a random access memory 910 (“RAM”) and a read-only memory (“ROM”) 915, and a system bus 935 that couples the system memory 905 to the CPU 950. A basic input / output system containing the basic routines that help to transfer information between elements within the computer, such as during startup, is stored in the ROM 915. The computer 900 further includes a storage device 920 for storing an operating system 925, application / program 930, and data.
[0064] The storage device 920 is connected to the CPU 950 through a storage controller (not shown) connected to the bus 935. The storage device 920 and its associated computer-readable media provide non-volatile storage for the computer 900. Although the description of computer- readable media contained herein refers to a storage device, such as a hard disk or CD-ROM drive, it should be appreciated by those skilled in the art that computer-readable media can be any available media that can be accessed by the computer 900.
[0065] By way of example, and not to be limiting, computer-readable media may comprise computer storage media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, RAM, ROM, EPROM, EEPROM, flash memory or other solid state memory technology, CD-ROM, DVD, or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, orAttorney docket # 204606-0168-00WO any other medium which can be used to store the desired information and which can be accessed by the computer.
[0066] According to various embodiments of the invention, the computer 900 may operate in a networked environment using logical connections to remote computers through a network 940, such as TCP / IP network such as the Internet or an intranet. The computer 900 may connect to the network 940 through a network interface unit 945 connected to the bus 935. It should be appreciated that the network interface unit 945 may also be utilized to connect to other types of networks and remote computer systems.
[0067] The computer 900 may also include an input / output controller 955 for receiving and processing input from a number of input / output devices 960, including a keyboard, a mouse, a touchscreen, a camera, a microphone, a controller, a joystick, or other type of input device. Similarly, the input / output controller 955 may provide output to a display screen, a printer, a speaker, or other type of output device. The computer 900 can connect to the input / output device 960 via a wired connection including, but not limited to, fiber optic, Ethernet, or copper wire or wireless means including, but not limited to, Wi-Fi, Bluetooth, Near-Field Communication (NFC), infrared, or other suitable wired or wireless connections.
[0068] As mentioned briefly above, a number of program modules and data files may be stored in the storage device 920 and / or RAM 910 of the computer 900, including an operating system 925 suitable for controlling the operation of a networked computer. The storage device 920 and RAM 910 may also store one or more applications / programs 930. In particular, the storage device 920 and RAM 910 may store an application / program 930 for providing a variety of functionalities to a user. For instance, the application / program 930 may comprise many types of programs such as a word processing application, a spreadsheet application, a desktop publishing application, a database application, a gaming application, internet browsing application, electronic mail application, messaging application, and the like. According to an embodiment of the present invention, the application / program 930 comprises a multiple functionality software application for providing word processing functionality, slide presentation functionality, spreadsheet functionality, database functionality and the like.Attorney docket # 204606-0168-00WO
[0069] The computer 900 in some embodiments can include a variety of sensors 965 for monitoring the environment surrounding and the environment internal to the computer 900. These sensors 965 can include a Global Positioning System (GPS) sensor, a photosensitive sensor, a gyroscope, a magnetometer, thermometer, a proximity sensor, an accelerometer, a microphone, biometric sensor, barometer, humidity sensor, radiation sensor, or any other suitable sensor. EXPERIMENTAL EXAMPLES
[0070] The invention is further described in detail by reference to the following experimental examples. These examples are provided for purposes of illustration only, and are not intended to be limiting unless otherwise specified. Thus, the invention should in no way be construed as being limited to the following examples, but rather, should be construed to encompass any and all variations which become evident as a result of the teaching provided herein.
[0071] Without further description, it is believed that one of ordinary skill in the art can, using the preceding description and the following illustrative examples, make and utilize the system and method of the present invention. The following working examples therefore, specifically point out the exemplary embodiments of the present invention, and are not to be construed as limiting in any way the remainder of the disclosure.
[0072] The details of the circuit-level simulations are described in Example #1 below. The circuit-level implementation is intended as a proof of concept and to evaluate effects of PVT variations and mismatch on the decoder’s performance, therefore smaller designs are considered with no annealing circuit. The results of the high-level simulations are discussed in Example #2, where the system is implemented on a more practical and larger LDPC problem. The annealing process is implemented in the high-level model to showcase the full potential of the proposed algorithm. A comparison with several standard algorithms and state-of-the-art implementations is provided in Example #3. Example #1 – Circuit-Level Simulation
[0073] To validate the proposed LDPC QuBRIM methodology and design, the corresponding circuit was implemented on Cadence Virtuoso Analog Design Environment (ADE) in 45 nmAttorney docket # 204606-0168-00WO generic PDK provided by Cadence. LDPC parity matrices were generated using the software provided in R. Neal, “Software for Low Density Parity Check Codes.” The generated parity matrices were converted to circuit schematics using SKILL script to automate component placement. The implemented LDPC parity matrices were uniform with fixed column and row degrees of 3 and 4 respectively, and a coding rate of 1 / 4.
[0074] To estimate performance metrics of the QuBRIM-based LDPC decoder under both PVT variations and mismatch meaningfully (e.g., Bit Error Rate, or BER, for a large SNR value of 4dB under different process corners), a large set of decoding problems needed to be tested. Circuit-level simulations are time-consuming in this regard, therefore, for the purpose of verification and to improve simulation time, some simplification steps were taken. First, only small LDPC matrices were implemented using circuit-level simulation. Although small LDPC matrices may not serve a practical purpose in communication, they are used for the verification of the QuBRIM decoding algorithm. A larger LDPC decoding problem is confronted below in the high-level simulations in Example #2. Because programmable PMOS resistors are used mainly for area saving and are not essential to the operation of the QuBRIM LDPC, poly resistors were used instead to further improve the simulation efficiency. Additionally, eachcurrent-steering DAC was replaced with a bias voltage source with ^^^ = ^^ and connected to thecurrent summing input of node^^through a bias resistor ^^, such that the resulting current was equal to that of the corresponding DAC. The resistors ^^were implemented as poly resistors to capture the effects of PVT variations and mismatch in DACs. The bias sources^^^were also used to charge the node capacitors ^^to an initial condition proportional to the magnitude of the received noisy signal ^^. Finally, the circuit-level results were reported with no annealing mechanism in-place to escape the local minima.
[0075] The lack of annealing results in a deteriorated BER, nonetheless the BER was expected to be higher than the uncoded case. MATLAB was used to generate sets of 100 randomly generated noisy codewords, each with a Signal-to-Noise Ratio (SNR) of 4 dB. Each noisy signal was then loaded into the QuBRIM schematic as the initial condition for the node capacitors, and on the voltage source connected to the current summing node through the resistance ^^. The QuBRIM nodes then interacted with each other and their states evolved toward their stable valuesAttorney docket # 204606-0168-00WO representing a minimum. The final settling values of the QuBRIM nodes were then compared with the correct non-noisy message sequence.
[0076] The LDPC QuBRIM was simulated using 32, 64 and 128-bit LDPC parity matrices with a coding rate of 1 / 4 and SNR of 4 dB. The result from one such run on the (128,32) LDPC matrix is shown in Fig.10, where the voltage of the node capacitor ^^(before quantization) was plotted as a function of time. The initial condition for each node capacitor was set as the received noisy signal ^^. The QuBRIM node voltages then evolved towards the minima located around the initial condition. Given the analog nature of this design, the time it takes for the circuit to become stable after starting from each different initial condition varied. The histogram of the settling time for 100 different codewords for the 32, 64 and 128-bit QuBRIM LDPC implementations is shown in Fig.11. Despite the difference in the sizes of the LDPC problems implemented, they all reached stable states within about 3ns. This shows that the disclosed QuBRIM LDPC decoder scales well for larger code lengths.
[0077] The histograms in Fig. 11 were obtained after sweeping the values of the resistors ^^ and^^ for each LDPC problem size. The final values were selected in order to optimize the BER aswell as the maximum and mean delay for the 100 codewords simulated. The optimized resistance values are tabulated in Table 1. A comparison between the implementation of the three different LDPC sizes is shown in Table 2. The throughput was estimated based on the worst-case delay from Fig.11, which was almost the same for each of them. However, since larger LDPC decoders contain more message bits k, the throughput scales with k and is higher for larger decoders. The Energy Efficiency (EE) improves slightly for 64-bit LDPC compared to the 32-bit version. For the 128-bit implementation, the EE shows more improvement, however, it is worth noting that the average delay is higher for this larger sized decoder. The results, again, show promising scaling potential for the disclosed QuBRIM LDPC. LDPC size (32,8) (64,16) (128,32) Rc(Ω) 60k 40k 60k Rb(Ω) 20k 40k 40k Table 1Attorney docket # 204606-0168-00WO LDPC (32,8) (64,16) (128,32) Min. throughput (Gbps) 10.4 23 46.6 Avg. power (mW) 1.91 4.04 7.35 Peak Power 1.96 5.66 11.95 EE (pJ / bit) 0.18 0.18 0.16 BER (%) 0.59 0.33 0.10 Areaα(mm2) 0.025 0.05 0.1 Table 2
[0078] The QuBRIM LDPC was tested under mismatch and process variations using the (32,8) LDPC matrix. The 100 messages used in the results for Table 2 were simulated across 100 Monte Carlo samples, and the histogram for the change in BER is reported in Fig.12. The mean BER was 0.62% which is close to the nominal BER of 0.59%. The standard deviation was 0.1325%. These results show no significant change under mismatch and process variations.
[0079] The disclosures of each and every patent, patent application, and publication cited herein are hereby incorporated herein by reference in their entirety. While this invention has been disclosed with reference to specific embodiments, it is apparent that other embodiments and variations of this invention may be devised by others skilled in the art without departing from the true spirit and scope of the invention. The appended claims are intended to be construed to include all such embodiments and equivalent variations.
[0080] An LDPC code with a variable row degree results in an Ising Hamiltonian with mixed- order terms. These mixed-order terms require XOR gates of different sizes. The unequal delays from these XOR gates may prolong the settling time of the QuBRIM nodes, and can potentially push the state of the QuBRIM further away from the correct global minima leading to a high error rate. This issue can be solved by using equally sized XOR gates for all terms and then grounding the unused input terminals to obtain lower order XOR gates. Additionally, it is possible to increase the size of the node capacitor, which slows down the charging and discharging rate of all nodes. This increases the stability of the circuit since the gate delay is no longer the dominating factor, leading to faster QuBRIM settling time and higher result accuracy. Example #2 – High-Level Simulation
[0081] In order to show the annealing process, a MATLAB-based high-level model of QuBRIM was implemented. The results from the high level model are compared with the circuit-level results in Table 2 and their consistency was verified. A 1 / 2 coding rate (672,336) LDPC matrixAttorney docket # 204606-0168-00WO based on IEEE802.11ad standard was used with the high-level model. The high-level model used did not account for signal propagation delays. In the high-level model, two mechanisms ofperturbation were applied in the annealing process. First, an error term with normal distributionσ௧^^^^^^ ^^^^^ was introduced in the comparator stage. This term reflects the presence of a noisycomparator. The second mechanism was carried out by flipping random spin bits as the QuBRIM reached equilibrium, as shown in Fig.13. In the top plot 1301, the voltage of all QuBRIM nodes is shown as the system seeks equilibrium. In the middle plot 1302, the Hamiltonian of the same system is shown. In the Hamiltonian plot, zero represents the level where all polarity checks are satisfied coinciding with the global minimum. The QuBRIM reached equilibrium within 20 time steps signified by no change in the Hamiltonian, however, since the Hamiltonian is higher than zero, the system was stuck at a local minimum. A small number of randomly selected spin bits were then flipped to push the QuBRIM outside the local minima, and the process was repeated several times, as shown in the bottom plot 1303 of Fig.13. In this example, the QuBRIM nodes reached the ground level at the end of the graph, where all polarity checks are satisfied and the decoded message matches the original message.
[0082] For the annealing process, a notable difference between the Max-Cut problem and the LDPC decoding problem was present. In a Max-Cut problem, more than one global minimum exists and all of them are considered as equally correct solutions. In the LDPC decoding problem, all correct codewords are global minima in the energy landscape, however, only one of them is the correct solution. Fortunately, for the LDPC problem using the noisy received signal as the initial condition places the system closer to the correct global minimum, while for Max- Cut the initial condition is usually random. A possible annealing schedule for Max-Cut is to use high perturbations in the beginning that are later tuned down until the system settles to its final state. In the case of LDPC decoding, too much perturbation would push the system toward the vicinity of an incorrect global minimum, resulting in a poor solution quality.
[0083] The number of flipped bits may be adjusted to obtain high solution quality determined by BER. It is worth noting that the adjustments made to the annealing schedule in the present disclosure were based on trial and error. There were many parameters that could be considered during the improvement process and several parameters are discussed below, but it should be noted that this is not an exhaustive list. One such parameter is the number of bits flipped. If theAttorney docket # 204606-0168-00WO number of flipped bits is too small, the QuBRIM may not be able to escape the local minimum. However, if the number of flipped bits is too high, the QuBRIM state may jump further away from the vicinity of the correct global minimum. In this case, the final solution will have high BER even if all polarity checks are satisfied.
[0084] Another parameter is the frequency of bit flips. In Fig.13, the system was allowed to reach a stable state before each subsequent bit flipping event. This can be described as low- frequency bit-flipping. Low-frequency bit-flipping can be combined with a higher number of bits flipped without pushing the QuBRIM too far away from the targeted global minimum, as in the case of Fig.12, where ∼ 4% of bits were flipped each time. The frequency of perturbation can be increased by applying bit flips before the system reaches a stable state. In this case, the ratio of bits that could be flipped while keeping the QuBRIM within the range of the targeted global minimum is lower. After several simulations, it was determined that in some embodiments, with higher frequency bit flips, the system can reach the global minimum faster. Once the system reaches the global minimum, further bit flips have a lower probability of pushing the QuBRIM outside the minimum.
[0085] In some embodiments, the annealing schedule could be further improved by flipping a variable number of bits during each step of the annealing process. For simplicity, only a fixed number of bits were flipped in the current example while their frequency and quantity were adjusted between iterations.
[0086] In another embodiment of the present invention, annealing is achieved by using a problem-aware perturbation heuristic to escape local minima in search of a global minimum in the vicinity of the received word. With reference to Fig.5, the disclosed annealing method is demonstrated whereby each row of XOR gates defines one parity condition. Initially, the machine is initialized with a smaller sub-set of rows (e.g., one row) of XOR gates selected at random to be active (i.e. powered) while the remaining XOR gates are off (i.e. not powered) with floating outputs, which enforces only those parity conditions corresponding to the active row(s). Then, after a certain time (e.g., a few hundreds of picoseconds), the initially selected sub-set is expanded by supplying power to more rows of XOR gates (i.e., introducing more parity conditions to be enforced), also chosen at random. This last step is repeated until all parityAttorney docket # 204606-0168-00WO conditions are enforced at which time the annealing process is terminated and the final state is read out as the decoded word. In this way, the machine starts with a lower-dimensional search space and then the dimensionality is increased by adding more parity conditions. In this annealing schedule, the random walks are always in the right direction. This is potentially more efficient than random spin perturbations in terms of decoding time and energy, where sometimes the machine might walk away from the desired global minimum or minima.
[0087] The LDPC decoding results for BER vs. SNR using the QuBRIM with different annealing conditions are shown in Fig.14. As shown in Fig.14, using random bit flips significantly improves BER. In standard LDPC algorithms, accuracy is further improved through iterations. Each iteration incrementally improves the accuracy of the decoded message. In QuBRIM, the results could be further improved through repeated sampling. Each sample is a rerun of the decoding process where the spin capacitors’ initial condition is set to the received noisy message. The difference between samples is the random thermal noise generated in the system that influences the equilibrium seeking process. Additionally, the bit flip signal generated from the random number generator is different for repeated samples. The QuBRIM LDPC system may then detect if a sample reaches global minimum by performing a parity check. The repeated samples for the same message in QuBRIM are henceforth referred to as QuBRIM iterations.
[0088] The results in Fig.14 were obtained using a maximum of 10 iterations per message based on the annealing schedule in Fig.13. A well designed annealing schedule results in an improved BER while requiring a lower number of maximum iterations. An improved annealing schedule is shown in the bottom plot 1502 of Fig.15. In the first 25 steps, a reduced value of k and ^^(refer to Equation 29) was used to slow down the system, while no bit flips were taking place. After 25 steps, k and ^^were increased and bit flip perturbations were enabled at a higher frequency and lower bit number per flip than in Fig. L. Using this annealing schedule, 100K random noisy messages were decoded with an SNR of 4dB. In the top plot 1501 of Fig.15, a histogram of the number of steps needed to reach the global minimum within one iteration is shown. An average of 190.1 steps were consumed by the QuBRIM system to determine the majority of the codewords. In some embodiments, if the accuracy is not sufficient, the remaining codewords can be re-attempted in the next iteration.Attorney docket # 204606-0168-00WO
[0089] The BER results for this annealing schedule is further discussed in Example #3. Example #3 - Comparison
[0090] Using high-level modeling, a 1 / 2 coding rate (672,336) LDPC decoder was implemented with several decoding algorithms from literature. The results using QuBRIM (QB) with the annealing schedule in Fig.15 were compared with Normalized Min-Sum (NMS), Offset Min- Sum (OMS) (see J. Chen, et al., in Proceedings. International Symposium on Information Theory, 2005), Propagation Belief (BP) (see R. G. Gallager, M.I.T. Press, 2005) and Layered Propagation Belief (LBP) (see D. Hocevar, in IEEE Workshop on Signal Processing Systems, 2004), using (1) and (4) iterations for each, as shown in Fig.16. The accuracy for QuBRIM using one iteration was significantly higher as compared to the remaining algorithms. Using 4 iterations, the QuBRIM accuracy improved, however, the accuracy of LBP, NMS, and OMS rapidly caught up, and slightly surpassed QuBRIM. It is worth noting that for QuBRIM, each iteration is an independent sample and can be run concurrently. Therefore, the QuBRIM iterations could be implemented in parallel on different QuBRIM modules, creating a trade-off between throughput, area, and power.
[0091] The implementation of the QuBRIM LDPC decoder is compared with various state-of- the-art decoders from literature in Table 3 below. The parameters used for QuBRIM were based on circuit-level simulations and estimates. The plots in Fig.11 are reproduced for the (672,336) LDPC decoder using circuit-level and high-level simulations. The circuit- and high-level implementations were both optimized for BER, and the average time in the circuit-level histogram was equalized with the average simulation step in the high-level simulation. Based on this assumption, each simulation step was estimated to be 106ps. The average throughput of QuBRIM with early termination was estimated to be 8.32Gbps based on Fig.15, as\begin{equation*} Throughput = ^(avg. # of steps) ∗ (time per step) ∗ ^^^௫Equation 30Attorney docket # 204606-0168-00WO where n is the code length 672, and ^^^௫is the maximum number of iterations. The average power was estimated from the circuit-level simulation. The normalized energy efficiency (NEE) was normalized to 45nm CMOS technology according to R. Ghanaatian, et al., IEEE Transactions on Very Large Scale Integration (VLSI) Systems, Feb.2018, as shown in Table 3. In the comparison, 1-QuBRIM refers to a single module implementation where each iteration is carried out in series. The parameters for 1-QuBRIM with a maximum iteration of one are included in the comparison due to its high accuracy making it relevant.4-QuBRIM refers to an implementation with four QuBRIM modules working in parallel, each decoding the same message. Each module carries out one iteration with a different seed for the random number generator responsible for bit flips. If one of the 4 modules triggers the early termination signal for satisfying the parity checks, the system chooses the decoded word from that module and terminates the decoding process. This implementation would have the same throughput as the single iteration QuBRIM, but with more accuracy, power, and area. Present Disclosure Lopez Milicevic Ajaz Li Technology 45 28 28 65 90 (nm) Supply Voltage 1.1 0.9 0.9 1.1 1.05 Standard IEEE 802.11ad IEEE 802.15.3c Coding Rate ½ 1 / 2, 3 / 4 1 / 2 1 / 2 1 / 2, 5 / 8, 3 / 4, 7 / 8 Architecture 1-QuBRIM 4-QuBRIM SAMS Flooding Layered Layered MS NMS Max Iterations 4 1 1 5 10 7 5 Throughput 7.9* 31.6* 31.6* 10.5 6.78 9.25 5.28 (Gbps) Power (mW) 38.3* 38.3* 153.2* 81 279 272.9 182 NEEnorm** 1.21 1.21 4.84 4.9 9.2 2.9 3.8 (pJ / bit / iter) Area (mm2) 0.55a0.55a2.2a0.14 1.99 0.575 2.25 Table 3
[0092] With reference to Table 3 above, the * indicates values that are based on estimates from circuit and behavioral simulations. The ** indicates a value normalized to 45-nm CMOS technology, where ^^^norm= NEE x (1 / SU2), S=௧^^^^^^^^௬is the relative dimension to 45nm,Attorney docket # 204606-0168-00WO and U=^௨^^^௬^.^is the relative voltage to 1.1V(see R. Ghanaatian, et al., IEEE Transactions on Very Large Scale Integration (VLSI) Systems, Feb.2018).aindicates that the values are a crude estimate based on conservative assumptions of 30% overhead on top of the standard cell sizes. Lopez refers to H. Lopez, et al., IEEE Transactions on Very Large Scale Integration (VLSI) Systems, Apr.2020; Milicevic refers to M. Milicevic et al., IEEE Transactions on Very Large Scale Integration (VLSI) Systems, 2018; Ajaz refers to S. Ajaz et al., in 2014 IEEE Asia Pacific Conference on Circuits and Systems (APCCAS), Nov.2014; and Li refers to M.-R. Li, et al., IEEE Journal of Solid-State Circuits, Feb.2017. References
[0093] The following publications are incorporated herein by reference in their entirety:
[0094] E. Ising, “Beitrag zur theorie des ferromagnetismus,” Zeitschrift für Physik, vol.31, no. 1, pp.253–258, 1925.
[0095] T. Wang and J. Roychowdhury, “OIM: Oscillator-based Ising Machines for Solving Combinatorial Optimisation Problems,” arXiv:1903.07163, 2019.
[0096] A. Lucas, “Ising formulations of many np problems,” Frontiers in Physics, vol.2, p.5, 02 2014.
[0097] Z. Bian, F. Chudak, W. G. Macready, and G. Rose, “The Ising model: teaching an old problem new tricks,” p.32, 2010.
[0098] “The D-Wave 2000Q Quantum Computer.” [Online]. Available: {https: / / www.dwavesys.com / sites / default / files / D- Wave\%202000Q\%20Tech\%20Collateral\0117F.pdf}
[0099] H. Takesue, T. Inagaki, K. Inaba, T. Ikuta, and T. Honjo, “Large-scale Coherent Ising Machine,” Journal of the Physical Society of Japan, vol.88, no.6, p.061014, 2019.Attorney docket # 204606-0168-00WO
[0100] T. Inagaki, Y. Haribara, K. Igarashi, T. Sonobe, S. Tamate, T. Honjo, A. Marandi, P. L. McMahon, T. Umeki, K. Enbutsu, O. Tadanaga, H. Takenouchi, K. Aihara, K.-i. Kawarabayashi, K. Inoue, S. Utsunomiya, and H. Takesue, “A coherent ising machine for 2000-node optimization problems,” Science, vol.354, no.6312, pp.603–606, 2016.
[0101] Y. Yamamoto, K. Aihara, T. Leleu, K.-i. Kawarabayashi, S. Kako, M. Fejer, K. Inoue, and H. Takesue, “Coherent ising machines—optical neural networks operating at the quantum limit,” npj Quantum Information, vol.3, no.1, p.49, 2017.
[0102] P. L. McMahon, A. Marandi, Y. Haribara, R. Hamerly, C. Langrock, S. Tamate, T. Inagaki, H. Takesue, S. Utsunomiya, K. Aihara, R. L. Byer, M. M. Fejer, H. Mabuchi, and Y. Yamamoto, “A fully programmable 100-spin coherent ising machine with all-to-all connections,” Science, vol.354, no.6312, pp.614–617, 2016.
[0103] K. Takata, A. Marandi, R. Hamerly, Y. Haribara, D. Maruo, S. Tamate, H. Sakaguchi, S. Utsunomiya, and Y. Yamamoto, “A 16-bit coherent ising machine for one-dimensional ring and cubic graph problems,” Scientific Reports, vol.6, no.1, p.34089, 2016.
[0104] F. Böhm, G. Verschaffelt, and G. Van der Sande, “A poor man’s coherent ising machine based on opto-electronic feedback systems for solving optimization problems,” Nature Communications, vol.10, no.1, p.3538, 2019. [Online]. Available: https: / / doi.org / 10.1038 / s41467-019-11484-3
[0105] S. Dutta, A. Khanna, A. S. Assoa, H. Paik, D. G. Schlom, Z. Toroczkai, A. Raychowdhury, and S. Datta, “An ising hamiltonian solver based on coupled stochastic phase- transition nano-oscillators,” Nature Electronics, vol.4, no.7, pp.502–512, 2021.
[0106] W. Moy, I. Ahmed, P.-w. Chiu, J. Moy, S. S. Sapatnekar, and C. H. Kim, “A 1,968-node coupled ring oscillator circuit for combinatorial optimization problem solving,” Nature Electronics, May 2022.
[0107] R. Afoakwa, Y. Zhang, U. K. R. Vengalam, Z. Ignjatovic, and M. Huang, “BRIM: Bistable Resistively-Coupled Ising Machine,” in 2021 IEEE International Symposium on High- Performance Computer Architecture (HPCA), Feb.2021, pp.749–760.Attorney docket # 204606-0168-00WO
[0108] Y. Zhang, U. K. R. Vengalam, A. Sharma, M. Huang, and Z. Ignjatovic, “QuBRIM: A CMOS Compatible Resistively-Coupled Ising Machine with Quantized Nodal Interactions,” in Proceedings of the 41st IEEE / ACM International Conference on Computer-Aided Design, ser. ICCAD ’22. New York, NY, USA: Association for Computing Machinery, Dec.2022, pp.1–8.
[0109] M. Yamaoka, C. Yoshimura, M. Hayashi, T. Okuyama, H. Aoki, and H. Mizuno, “24.3 20k-spin ising chip for combinational optimization problem with cmos annealing,” in 2015 IEEE International Solid-State Circuits Conference - (ISSCC) Digest of Technical Papers, Feb 2015, pp.1–3.
[0110] T. Takemoto, M. Hayashi, C. Yoshimura, and M. Yamaoka, “2.6 a 2 by 30k-spin multichip scalable annealing processor based on a processing-in-memory approach for solving large-scale combinatorial optimization problems,” in IEEE International Solid- State Circuits Conference, February 2019.
[0111] K. Yamamoto, K. Kawamura, K. Ando, N. Mertig, T. Takemoto, M. Yamaoka, H. Teramoto, A. Sakai, S. Takamaeda-Yamazaki, and M. Motomura, “STATICA: A 512-Spin 0.25M-Weight Annealing Processor With an All-Spin-Updates-at-Once Architecture for Combinatorial Optimization With Complete Spin–Spin Interactions,” IEEE Journal of Solid- State Circuits, vol.56, no.1, pp.165–178, Jan.2021.
[0112] R. Gallager, “Low-density parity-check codes,” IRE Transactions on Information Theory, vol.8, no.1, pp.21–28, 1962.18
[0113] A. Hareedy, C. Lanka, N. Guo, and L. Dolecek, “A combinatorial methodology for optimizing non-binary graph-based codes: Theoretical analysis and applications in data storage,” IEEE Transactions on Information Theory, vol.65, no.4, pp.2128–2154, 2019.
[0114] M. Tawada, S. Tanaka, and N. Togawa, “A New LDPC Code Decoding Method: Expanding the Scope of Ising Machines,” in 2020 IEEE International Conference on Consumer Electronics (ICCE), Jan.2020, pp.1–6.Attorney docket # 204606-0168-00WO
[0115] E. Berlekamp, R. McEliece, and H. van Tilborg, “On the inherent intractability of certain coding problems (corresp.),” IEEE Transactions on Information Theory, vol.24, no.3, pp.384– 386, 1978.
[0116] J. Bruck and M. Naor, “The hardness of decoding linear codes with preprocessing,” IEEE Transactions on Information Theory, vol.36, no.2, pp.381–385, 1990.
[0117] J. Chen, A. Dholakia, E. Eleftheriou, M. Fossorier, and X.-Y. Hu, “Reduced-complexity decoding of ldpc codes,” IEEE Transactions on Communications, vol.53, no.8, pp.1288–1299, 2005.
[0118] D. MacKay, “Good error-correcting codes based on very sparse matrices,” IEEE Transactions on Information Theory, vol.45, no.2, pp.399–431, 1999.
[0119] A. Riazanov, Y. Maximov, and M. Chertkov, “Belief propagation min-sum algorithm for generalized min-cost network flow,” in 2018 Annual American Control Conference (ACC), 2018, pp.6108–6113.
[0120] W.-T. Lin and T.-H. Kuo, “A 12b 1.6GS / s 40mW DAC in 40nm CMOS with >70dB SFDR over entire Nyquist bandwidth,” in 2013 IEEE International Solid-State Circuits Conference Digest of Technical Papers, 2013, pp.474–475.
[0121] E. Guglielmi, F. Toso, F. Zanetto, G. Sciortino, A. Mesri, M. Sampietro, and G. Ferrari, “High-value tunable pseudo-resistors design,” IEEE Journal of Solid-State Circuits, vol.55, no. 8, pp.2094–2105, 2020.
[0122] R. Neal, “Software for Low Density Parity Check Codes.” [Online]. Available: http: / / radfordneal.github.io / LDPC-codes /
[0123] J. Chen, R. Tanner, C. Jones, and Y. Li, “Improved min-sum decoding algorithms for irregular ldpc codes,” in Proceedings. International Symposium on Information Theory, 2005. ISIT 2005., 2005, pp.449–453.
[0124] R. G. Gallager, Low-density parity-check codes. M.I.T. Press, 2005.Attorney docket # 204606-0168-00WO
[0125] D. Hocevar, “A reduced complexity decoder architecture via layered decoding of ldpc codes,” in IEEE Workshop onSignal Processing Systems, 2004. SIPS 2004., 2004, pp.107–112.
[0126] H. Lopez, H.-W. Chan, K.-L. Chiu, P.-Y. Tsai, and S.-J. J. Jou, “A 75-Gb / s / mm2 and Energy-Efficient LDPC Decoder Based on a Reduced Complexity Second Minimum Approximation Min-Sum Algorithm,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol.28, no.4, pp.926–939, Apr.2020.
[0127] M. Milicevic and P. G. Gulak, “A multi-gb / s frame-interleaved ldpc decoder with path- unrolled message passing in 28-nm cmos,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol.26, no.10, pp.1908–1921, 2018.
[0128] S. Ajaz and H. Lee, “Multi-Gb / s multi-mode LDPC decoder architecture for IEEE 802.11ad standard,” in 2014 IEEE Asia Pacific Conference on Circuits and Systems (APCCAS), Nov.2014, pp.153–156.
[0129] M.-R. Li, C.-H. Yang, and Y.-L. Ueng, “A 5.28-Gb / s LDPC Decoder With Time-Domain Signal Processing for IEEE 802.15.3c Applications,” IEEE Journal of Solid-State Circuits, vol. 52, no.2, pp.592–604, Feb.2017.
[0130] R. Ghanaatian, A. Balatsoukas-Stimming, T. C. Müller, M. Meidlinger, G. Matz, A. Teman, and A. Burg, “A 588-Gb / s LDPC Decoder Based on Finite-Alphabet Message Passing,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, no.2, pp.329–340, Feb. 2018.
Claims
Attorney docket # 204606-0168-00WO CLAIMS What is claimed is:
1. A network, comprising: a plurality of computation nodes; a plurality of logic gates; a plurality of coupling units; a plurality of digital-to-analog converters; each of the plurality of computation nodes comprising: an input configured to received currents from the coupling units in the network and currents from at least one of the digital-to-analog converters; an output configured to generate at least two discrete output voltages; a capacitor configured to store internal state as a voltage; a comparator configured to compare the voltage on the capacitor against a threshold to produce a first discrete voltage value at its output when the voltage is above the threshold, and a second discrete voltage value at its output when the voltage is below the threshold, with the output connected to the output of the computation node; a current conveyor circuit having an input connected to the input of the computation node and an output connected to the capacitor, the current conveyor configured to hold its input at a constant voltage and mirror a scaled version of the current received at the input into the capacitor.
2. The network of claim 1, each logic gate comprising an input connected to the output of at least one computation node; an output configured to generate one of at least two discrete output voltages, the output connected to an input of at least one coupling unit of the plurality of coupling units.
3. The network of claims 1 or 2, wherein at least one logic gate is selected from an XOR gate or an XNOR gate.
4. The network of any of claims 1-3, each coupling unit comprising:Attorney docket # 204606-0168-00WO an input connected to the voltage output of at least one logic gate of the plurality of logic gates; an output configured to generate one of at least two discrete output current values, the output connected to an input of at least one computation node of the plurality of computation nodes; and means to convert voltage values at the input to currents proportional to the input voltage at the output of the coupling unit.
5. The network of any of claims 1-4, each digital-to analog converter comprising: an input configured to receive a digital sample representing a noisy message or parity bit received by a receiver circuit in a communication system; means to produce current proportional to the digital input and supply it to the output of the digital-to-analog converter; and an output configured to supply the current to the input of at least one computation node of the plurality of computation nodes.
6. The network of any of claims 1-5, wherein at least one coupling unit of the plurality of coupling units comprises a resistive element configured to convert an input voltage to the coupling unit into a current supplied at its output.
7. The network of any of claims 1-5, wherein at least one coupling unit of the plurality of coupling units comprises at least one transistor configured as a current source, with a drain terminal connected to the output of the coupling unit through switches controlled by the voltage at the input of the coupling unit.
8. The network of any of claims 1-7, further comprising a spin perturbation control circuit configured to change a polarity of capacitor voltages in at least one computation node during the network’s operation.
9. The network of claim 8, wherein the spin perturbation control circuit is configured to generate perturbation signals from a random source or a pseudo-random source.Attorney docket # 204606-0168-00WO 10. The network of claim 8 or 9, wherein the spin perturbation control circuit is configured to generate perturbation signals such that the rate of spin perturbation events diminishes over time.
11. The network of computation nodes of Claim 10, configured such that the rate of spin perturbation events has an exponential decay over time.
12. The network of any of claims 1-11, further comprising: a plurality of logic gates arranged in a set of rows, each row corresponding to a parity check; and a parity check perturbation control circuit configured to supply power to a first subset of the rows of logic gates at a first time.
13. Thee network of claim 12, wherein the first time corresponds to the initialization of the network.
14. The network of claim 12 or 13, wherein the parity check perturbation control circuit is further configured to supply power to a second subset of rows different from the first subset of rows at a second time after the first time.
15. The network of claim 14, wherein the first and second subsets overlap.
16. The network of claim 14, wherein the first and second subsets do not intersect.
17. The network of any of claims 12-16, wherein the first subset of rows of logic gates is chosen randomly or pseudo-randomly.