Bipolar single FLUX quantum weight circuit and weighting a unipolar single FLUX quantum input pulse
The bipolar single flux quantum weight circuit addresses the inefficiencies of unipolar SFQ circuits by employing a superconducting weight storage loop and oppositely coupled SQUIDs to achieve efficient bipolar synaptic weights, improving circuit efficiency and reducing complexity in SFQ-based neural networks.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE GOVERNMENT OF THE UNITED STATES OF AMERICA AS REPRESENTED BY THE SECRETARY DEPARTMENT OF HEALTH & HUMAN SERVICES
- Filing Date
- 2025-04-07
- Publication Date
- 2026-05-28
AI Technical Summary
Conventional single flux quantum (SFQ) circuits face challenges in implementing bipolar synaptic weights due to their unipolar nature, leading to increased circuit complexity, energy consumption, and inefficiencies in storing and programming analog weight values, which are crucial for advanced artificial neural network functions.
A bipolar single flux quantum weight circuit architecture utilizing a superconducting weight storage loop, mutually coupled SQUIDs with opposite polarities, and inductive dividers to achieve differential signal modulation, enabling both positive and negative synaptic weights efficiently.
The circuit allows for compact, low-power, and high-speed implementation of bipolar synaptic weights, reducing circuit complexity and energy consumption by using a single analog weight storage loop to control two opposing signal paths, thereby enhancing the functionality of SFQ-based neuromorphic systems.
Smart Images

Figure IB2025000686_28052026_PF_FP_ABST
Abstract
Description
BIPOLAR SINGLE FLUX QUANTUM WEIGHT CIRCUIT AND WEIGHTING A UNIPOLAR SINGLE FLUX QUANTUM INPUT PULSESTATEMENT REGARDING FEDERALLY SPONSORED RESEARCH
[0001] This invention was made with United States Government support from the National Institute of Standards and Technology (NIST), an agency of the United States Department of Commerce. The Government has certain rights in this invention.CROSS-REFERENCE TO RELATED APPLICATION
[0002] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 575,253 (filed April 5, 2024), which is herein incorporated by reference in its entirety.BACKGROUND
[0003] The present invention generally relates to the field of superconducting electronics, and more particularly to techniques for implementing bipolar synaptic weight functions in single flux quantum circuit architectures for artificial neural networks.
[0004] Artificial neural networks (ANNs), inspired by biological nervous systems, have emerged as powerful computational paradigms, demonstrating remarkable success in diverse applications ranging from pattern recognition to natural language processing. The performance of these networks, however, often relies on computationally intensive training procedures executed on conventional software platforms, typically utilizing graphical processing units (GPUs). While effective, this software-centric approach faces significant limitations in terms of processing speed and energy consumption, particularly as network complexity and dataset sizesescalate. Consequently, there is a burgeoning interest in developing dedicated hardware implementations of ANNs, often termed neuromorphic computing, which promise orders-of-magnitude improvements in speed and energy efficiency by directly mapping neural computations onto physical device characteristics.
[0005] Among the candidate technologies for hardware ANNs, superconducting electronics, particularly circuits based on Josephson junctions (JJs) operating with single flux quantum (SFQ) pulses, offer compelling advantages. SFQ circuits leverage the quantum mechanical properties of superconductors to achieve extremely high operating frequencies (hundreds of gigahertz) and exceptionally low energy dissipation per operation (sub-attojoule levels), potentially surpassing biological efficiency even when accounting for cryogenic cooling overhead. The natural spiking behavior of JJs mirrors neuronal activity, and SFQ pulses propagate nearly losslessly on superconducting transmission lines, analogous to axons. However, realizing complex ANN functionalities, especially the crucial synaptic weighting required for both inference and learning, within the constraints of SFQ technology presents substantial challenges. A primary difficulty arises from the inherently unipolar nature of standard SFQ pulses; a Josephson junction either produces a voltage pulse of a specific polarity or produces no pulse.
[0006] This unipolar characteristic complicates the implementation of bipolar synaptic weights, the ability for a connection between neurons to be either excitatory (positive weight) or inhibitory (negative weight), which is fundamental to the functionality and learning capability of many sophisticated ANN models. Prior approaches within SFQ architectures to achieve bipolar weighting often involve cumbersome solutions, such as duplicating synaptic circuitry for separate positive and negative pathways or pre-assigning synaptic polarity, significantly increasing circuit footprint, complexity, and energy consumption for a given network size. Furthermore, storing and precisely programming the analog weight values, representing synaptic strength, is notoriously difficult in any analog hardware system due to inherent fabrication variations. While digital SFQ approaches exist, realizing the high density and potential efficiency of analog weight storage remains desirable. Efficiently implementing learning algorithms directly in SFQ hardware is also problematic, asmany conventional algorithms like Error Backpropagation require non-local information flow ill-suited to direct, scalable hardware mapping.
[0007] It is therefore an objective of the present invention to provide a bipolar single flux quantum weight circuit architecture that enables the representation and programmable adjustment of both positive and negative synaptic weights using inherently unipolar SFQ signals, thereby overcoming the above-mentioned disadvantages of the prior art at least in part by reducing circuit complexity and enhancing the efficiency of implementing neural network functions in superconducting hardware. Accordingly, circuits and methods for realizing efficient and programmable bipolar synaptic weighting within SFQ-based neuromorphic systems would be advantageous and would be favorably received in the art.BRIEF DESCRIPTION
[0008] One aspect of the present invention relates to a bipolar single flux quantum weight circuit. A bipolar circuit may be understood as one capable of processing or representing information with both positive and negative values or polarities. Single flux quantum (SFQ) refers to a digital logic family based on superconducting electronics, utilizing quantized pulses of magnetic flux carried by Josephson junctions for ultra-fast, low-power computation. A weight circuit, in the context of neural networks, functions as an artificial synapse, modulating the influence of an input signal based on a stored weight value.
[0009] It may be provided that the circuit includes a superconducting weight storage loop comprising a weight storage inductor. A superconducting loop is a closed path made of material exhibiting zero electrical resistance below a critical temperature, capable of sustaining persistent currents indefinitely. A weight storage inductor within such a loop allows the stored persistent current, representing the synaptic weight, to generate a corresponding magnetic flux. This arrangement permits the non-volatile storage of an analog weight value directly as a physical current magnitude within the superconducting loop, offering potential advantages in terms of energy efficiency for maintaining the weight state compared to conventional memory technologies.
[0010] It may be provided that the circuit includes a first superconducting quantum interference device (SQUID) mutually coupled to the superconducting weight storage loop with a first polarity. A SQUID is a device utilizing superconducting loops and Josephson junctions, known for its extreme sensitivity to magnetic flux; its effective inductance can be modulated by external magnetic fields or circulating currents. Mutual coupling implies that the magnetic flux generated by the current in the weight storage loop influences the first SQUID. The first polarity defines the directional sense of this interaction. This feature leverages the tunable inductance property of the SQUID, allowing the stored weight current to directly control a circuit parameter (inductance) of the first SQUID, forming a basis for modulating signal flow.
[0011] It may be provided that the circuit includes a second SQUID mutually coupled to the superconducting weight storage loop with a second polarity opposite the first polarity, wherein an inductance of the first SQUID and an inductance of the second SQUID vary based upon a current circulating within the superconducting weight storage loop. Providing a second SQUID, also sensitive to the weight storage loop current but coupled with opposite polarity, establishes a differential mechanism. As the weight current changes, the inductance of the first SQUID might increase while the inductance of the second SQUID decreases (or vice versa), depending on the SQUID design and operating point. This differential variation in inductance based on a single stored weight current is fundamental to enabling the bipolar response from a unipolar weight representation, improving circuit efficiency by using one storage element to control two opposing paths.
[0012] It may be provided that the circuit includes a first inductive divider electrically connected to an input terminal, the first inductive divider comprising a first leg including the first SQUID connected to ground and a second leg including a first coupling inductor. An inductive divider splits an incoming current based on the inductance ratio of its parallel paths or legs. Connecting the input terminal allows an incoming SFQ pulse to enter this divider. The first leg routes a portion of the input current through the first SQUID to ground, a reference potential. The second leg routes the remaining portion through the first coupling inductor. The variable inductance of the first SQUID, controlled by the weight current, dynamically alters the division ratioof this first divider, modulating how much signal passes through the first coupling inductor.
[0013] It may be provided that the circuit includes a second inductive divider electrically connected to the input terminal, the second inductive divider comprising a first leg including the second SQUID connected to ground and a second leg including a second coupling inductor. This second divider mirrors the first but incorporates the second SQUID, whose inductance varies oppositely to the first SQUID in response to the weight current. It receives the same input signal and divides it between the second SQUID path to ground and the second coupling inductor path. This structure ensures that as the signal flow through the first coupling inductor is modulated in one direction (e.g., increased), the signal flow through the second coupling inductor is simultaneously modulated in the opposite direction (e.g., decreased), based on the single weight current value.
[0014] It may be provided that the circuit includes a soma element comprising a soma inductor line, the first coupling inductor mutually coupled to the soma inductor line with a third polarity, and the second coupling inductor mutually coupled to the soma inductor line with a fourth polarity opposite the third polarity, the soma element further comprising a soma output Josephson junction connected to the soma inductor line. The soma element acts as a summation point, analogous to a neuron's cell body. A soma inductor line is a structure, likely a series of inductors, where signals can be combined via mutual coupling. The first coupling inductor transfers its portion of the signal to the soma line with one polarity (e.g., positive contribution), while the second coupling inductor transfers its signal portion with the opposite polarity (e.g., negative contribution). This differential coupling to the soma, combined with the differential signal division controlled by the weight current, achieves the bipolar weighting: the net signal induced in the soma can be positive, negative, or near zero depending on the stored weight. The soma output Josephson junction acts as a threshold device, firing an output SFQ pulse only if the summed signal in the soma inductor line exceeds its critical current, completing the basic neuron computation of weighted sum and threshold activation with inherent high speed and low power characteristic of SFQ technology. This integrated structure provides acompact and efficient hardware realization of a bipolar synapse feeding into a neuron soma.
[0015] One aspect of the present invention relates to a method for weighting a unipolar single flux quantum (SFQ) input pulse utilizing a bipolar single flux quantum weight circuit. A method may be understood as a particular procedure for accomplishing or approaching something. Weighting involves modifying the influence of a signal based on a specific value or factor. A unipolar pulse exhibits only one polarity relative to a baseline, typical of standard SFQ signals which are either present (one polarity) or absent. Single flux quantum (SFQ) refers to a superconducting digital logic paradigm using quantized pulses of magnetic flux. An input pulse is the signal entering the system. The bipolar single flux quantum weight circuit is a specific hardware arrangement, detailed previously, designed to apply both positive and negative effective weights to such unipolar inputs, comprising elements like a superconducting weight storage loop (200), oppositely coupled SQUIDs (201 , 202), inductive dividers (203, 204), coupling inductors (205, 206), and a soma element (207) with an output Josephson junction (209).
[0016] It may be provided that the method includes applying the unipolar SFQ input pulse (x) to an input terminal (210a) connected to the first inductive divider (203) and the second inductive divider (204). Applying the pulse initiates the weighting operation. The input terminal (210a) serves as the entry point for the signal to be weighted. Connection to both inductive dividers (203, 204), which form parallel processing paths, ensures the input signal is simultaneously directed into the differential structure necessary for bipolar weighting. This arrangement allows the circuit to begin processing the input signal immediately along both complementary pathways, contributing to the high-speed operation potential of the SFQ architecture.
[0017] It may be provided that the method includes dividing the SFQ input pulse (x) via the first inductive divider (203), wherein a first portion is directed through the first SQUID (201 ) and a second portion is directed through the first coupling inductor (205), the division ratio being dependent on an inductance of the first SQUID (201 ) influenced by the weight current. Dividing the pulse is the core mechanism for modulating its strength. The first inductive divider (203) acts as a current splitter. The path through the first SQUID (201) leads to ground, while the path through the firstcoupling inductor (205) leads toward the summation stage. The division ratio, determining how much current flows into the coupling inductor (205), is dynamically controlled by the inductance of the first SQUID (201 ). This inductance is, in turn, influenced by the analog weight current stored in the superconducting weight storage loop (200). This arrangement translates the stored analog weight into a variable impedance, enabling analog modulation of the signal component destined for positive contribution to the output, thereby performing a part of the synaptic weighting function efficiently.
[0018] It may be provided that the method includes dividing the SFQ input pulse (x) via the second inductive divider (204), wherein a third portion is directed through the second SQUID (202) and a fourth portion is directed through the second coupling inductor (206), the division ratio being dependent on an inductance of the second SQUID (202) influenced by the weight current. This step executes the complementary division process in the parallel path. The second inductive divider(204) operates similarly to the first, but involves the second SQUID (202) whose inductance is also controlled by the same weight current, but varies oppositely to the first SQUID'S inductance. Consequently, as the signal portion directed through the first coupling inductor (205) increases or decreases, the signal portion directed through the second coupling inductor (206) does the opposite. This coordinated, differential division based on a single weight current value forms the physical basis for generating opposing signal components necessary for bipolar representation, enhancing the circuit's functional capability without requiring separate positive and negative weight storage elements.
[0019] It may be provided that the method includes inducing a first signal in the soma inductor line (207a) from the second portion via the first coupling inductor(205) with a first signal polarity. Inducing the signal transfers the modulated current from the first divider's second leg into the summation circuit. The soma inductor line (207a) acts as the input structure for the summation. Mutual coupling via the first coupling inductor (205) allows this transfer efficiently at SFQ speeds. The defined first signal polarity (e.g., positive) establishes this path's contribution as potentially excitatory. This arrangement facilitates the collection of the positively weighted signal component.
[0020] It may be provided that the method includes inducing a second signal in the soma inductor line (207a) from the fourth portion via the second coupling inductor (206) with a second signal polarity opposite the first signal polarity. This step transfers the modulated signal from the second divider's path into the same summation structure but with opposite polarity. The second coupling inductor (206) is arranged to induce this opposing signal (e.g., negative) in the soma inductor line (207a). This provides the potentially inhibitory contribution. Having both oppositely polarized signals induced into the same soma line allows for their direct summation or cancellation.
[0021] It may be provided that the method includes generating a net signal in the soma element (207) based on a summation of the first induced signal and the second induced signal, wherein the soma output Josephson junction (209) emits an output SFQ pulse (y) if the net signal exceeds a predetermined threshold. Generating the net signal represents the culmination of the weighting process. The soma element (207), particularly the soma inductor line (207a), passively sums the oppositely polarized induced signals. The resulting net signal's amplitude and polarity thus directly reflect the bipolar weight determined by the current in the weight storage loop (200). The soma output Josephson junction (209) then performs a thresholding operation; it acts as a non-linear activation function inherent to SFQ circuits. If the summed analog current (net signal) surpasses the junction's critical current (the predetermined threshold), it fires, emitting an output SFQ pulse (y). Otherwise, no pulse is emitted. This completes the synaptic computation (weighted sum and thresholding) rapidly and efficiently, translating the bipolar weighted analog sum into a digital SFQ output suitable for propagation in a larger superconducting neural network. This process allows a unipolar input to effectively be weighted by a bipolar value implemented through the differential hardware structure.BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The following description cannot be considered limiting in any way. Various objectives, features, and advantages of the disclosed subject matter can be more fully appreciated with reference to the following detailed description of thedisclosed subject matter when considered in connection with the following drawings, in which like reference numerals identify like elements.
[0023] FIG. 1 shows, according to some embodiments, a block diagram representation of a small-scale 2-input, 2-hidden-unit, 1 -output-unit neural network architecture, alongside a graph illustrating the network's probability of correct output versus time during a continuous SPICE simulation where the network learns three distinct logical functions sequentially based on changing target outputs.
[0024] FIG. 2 shows, according to some embodiments, graphs illustrating the effect of stochastic weight exploration magnitude on the learning process for a specific function, comparing the probability of correct output and the evolution of synaptic weight storage currents over time for simulations with low versus higher stochastic excitation levels applied to hidden layer units.
[0025] FIG. 3 shows, according to some embodiments, results from training the network to perform the XOR function, including a graph of probability correct versus time, a graph of the corresponding weight storage currents evolving over time, and visualizations of the input and output spiking patterns both at the beginning of training and after the network has successfully learned the function.
[0026] FIG. 4 shows, according to some embodiments, a graph illustrating the projected time scaling performance of the proposed superconducting neuromorphic architecture, plotting the total time delay for a combined inference and learning / weight update cycle against the number of layers in the network for several different network widths or layer depths.
[0027] FIG. 5 shows, according to some embodiments, a block diagram illustrating the logic circuitry designed to implement the weight update rule for a synaptic weight within the output layer of the neural network, detailing the flow of signals representing network output, target output, and synapse input.
[0028] FIG. 6 shows, according to some embodiments, a conceptual circuit diagram of the bipolar synaptic circuit including distinct positive and negative branches coupled to a soma element, accompanied by SPICE simulation results demonstratinghow varying the current in the weight storage loop affects the polarity and magnitude of the signal coupled into the soma element in response to input spikes.
[0029] FIG. 7 shows, according to some embodiments, an alternative block diagram representation of the weight update logic specifically designed for an output layer synapse.
[0030] FIG. 8 shows, according to some embodiments, a block diagram illustrating a variation of the output layer weight update logic incorporating D-Flip-Flop (DFF) elements for signal storage or timing.
[0031] FIG. 9 shows, according to some embodiments, comparative block diagrams illustrating the weight update logic for synapses in the output layer alongside the distinct logic required for synapses in hidden layers within the reinforcement learning framework.
[0032] FIG. 10 shows, according to some embodiments, a circuit diagram illustrating the hardware implementation of a stochastic element, designed similarly to the bipolar synapse but utilizing a smaller storage loop and potentially thermally excitable Josephson junctions to generate random weight fluctuations.
[0033] FIG. 11 shows, according to some embodiments, a detailed circuit schematic representing the complete SPICE simulation model for a 5-node neural network, encompassing all input, hidden, and output units along with their associated bipolar synapses, soma elements, and reinforcement learning update logic circuits.DETAILED DESCRIPTION
[0034] A detailed description of one or more embodiments is presented herein by way of exemplification and not limitation.
[0035] Conventional approaches to implementing artificial neural networks using single flux quantum (SFQ) superconducting electronics encounter significant hurdles, particularly in representing and manipulating synaptic weights. The inherent unipolar nature of SFQ pulses makes the direct realization of bipolar weights (synapticconnections that can be either excitatory or inhibitory) inefficient. Existing solutions often necessitate duplicating circuit pathways for positive and negative weights, substantially increasing the component count, physical footprint, and overall energy dissipation of the network. Furthermore, reliably storing and precisely programming analog weight values in any hardware implementation, including superconducting circuits, faces challenges due to fabrication variability, which can compromise network performance and complicate training. Training algorithms themselves, especially those requiring global information or complex backpropagation calculations, are difficult to map efficiently onto the local connectivity constraints and operational characteristics of SFQ hardware.
[0036] The bipolar single flux quantum weight circuit overcomes these limitations related to implementing bipolar synaptic functions within unipolar SFQ architectures and managing analog weight storage variability. It has been discovered that a bipolar single flux quantum weight circuit, structured according to the principles outlined herein, provides an efficient mechanism for achieving programmable positive and negative weighting effects using standard SFQ signals and components. One technical advantage stems from the use of a superconducting weight storage loop comprising a weight storage inductor; this permits the analog synaptic weight to be stored persistently as a circulating current, minimizing static power consumption associated with weight maintenance. The differential modulation scheme, employing a first SQUID mutually coupled to the weight storage loop with a first polarity and a second SQUID mutually coupled with an opposite second polarity, allows a single stored current value to simultaneously control the inductances of both SQUIDs in opposing directions. This differential inductance variation is then translated into differential signal division through a first inductive divider connected to an input terminal and incorporating the first SQUID in one leg and a first coupling inductor in another, and a corresponding second inductive divider connected to the same input terminal and incorporating the second SQUID and a second coupling inductor. This structure avoids the need for separate circuits or storage elements for positive and negative weights, offering a more compact and potentially less complex implementation compared to duplicated pathway approaches. The final bipolar summation occurs within a soma element comprising a soma inductor line, where the first coupling inductor is mutually coupled with a third polarity and the second couplinginductor is mutually coupled with an opposite fourth polarity. This arrangement ensures that the differentially modulated signal components from the two dividers contribute oppositely to the net signal in the soma. The inclusion of a soma output Josephson junction connected to the soma inductor line provides the necessary thresholding function, emitting an output SFQ pulse only when the net weighted sum exceeds its critical current, thereby integrating the bipolar synaptic weighting directly with the neuronal activation mechanism characteristic of SFQ neuromorphic circuits. This architecture facilitates a direct hardware implementation of synaptic weighting capable of bipolar operation, enhancing the functional capability of SFQ-based neural networks while leveraging their inherent high-speed and low-power potential.
[0037] In an embodiment, a bipolar single flux quantum weight circuit comprises: a superconducting weight storage loop (200) comprising a weight storage inductor (200a); a first superconducting quantum interference device (SQUID) (201 ) mutually coupled to the superconducting weight storage loop (200) with a first polarity; a second SQUID (202) mutually coupled to the superconducting weight storage loop (200) with a second polarity opposite the first polarity, wherein an inductance of the first SQUID (201 ) and an inductance of the second SQUID (202) vary based upon a current circulating within the superconducting weight storage loop (200); a first inductive divider (203) electrically connected to an input terminal (210a), the first inductive divider (203) comprising a first leg including the first SQUID (201 ) connected to ground (GND) and a second leg including a first coupling inductor (205); a second inductive divider (204) electrically connected to the input terminal (210a), the second inductive divider (204) comprising a first leg including the second SQUID (202) connected to ground (GND) and a second leg including a second coupling inductor (206); and a soma element (207) comprising a soma inductor line (207a), the first coupling inductor (205) mutually coupled to the soma inductor line (207a) with a third polarity, and the second coupling inductor (206) mutually coupled to the soma inductor line (207a) with a fourth polarity opposite the third polarity, the soma element (207) further comprising a soma output Josephson junction (209) connected to the soma inductor line (207a). In an embodiment, the superconducting weight storage loop (200) further comprises a first Josephson junction (200b) and a second Josephson junction (200c) arranged to permit injection of single flux quantum pulses for increasing or decreasing the current circulating within the superconducting weight storage loop(200). In an embodiment, the second leg of the first inductive divider (203) further comprises a first resistor (205a) in series with the first coupling inductor (205) connected to ground (GND), and the second leg of the second inductive divider (204) further comprises a second resistor (206a) in series with the second coupling inductor (206) connected to ground (GND). In an embodiment, the soma element (207) further comprises a soma resistor (207b) connecting the soma inductor line (207a) to ground (GND), and wherein the soma output Josephson junction (209) connects the soma inductor line (207a) to an output Josephson transmission line (208). In an embodiment, the first SQUID (201) is physically positioned relative to the superconducting weight storage loop (200) differently than the second SQUID (202) to achieve the first polarity and the second polarity of mutual coupling. In an embodiment, the mutual coupling of the first coupling inductor (205) to the soma inductor line (207a) and the mutual coupling of the second coupling inductor (206) to the soma inductor line (207a) result in signals induced in the soma element (207) that are opposite in polarity relative to each other upon application of an input pulse to the input terminal (210a). In an embodiment, the circuit further comprises an input splitter (210) connecting the input terminal (210a) to both the first inductive divider (203) and the second inductive divider (204). In an embodiment, the first SQUID (201 ) and the second SQUID (202) are biased to remain in a non-spiking state across a range of the current circulating within the superconducting weight storage loop (200). In an embodiment, the circuit further comprises a bias weight circuit (211 ) comprising a bias weight storage loop (212) and a bias weight inductive divider leg (213), the bias weight inductive divider leg (213) mutually coupled to the soma inductor line (207a). In an embodiment, the superconducting weight storage loop (200), the first SQUID (201 ), the second SQUID (202), the first coupling inductor (205), the second coupling inductor (206), the soma inductor line (207a), and the soma output Josephson junction (209) comprise superconducting materials and Josephson junctions.
[0038] The bipolar single flux quantum weight circuit comprises a superconducting weight storage loop (200) including a weight storage inductor (200a). This loop (200), formed from superconducting materials operating at cryogenic temperatures, functions to maintain a persistent current whose magnitude represents the analog synaptic weight. The weight storage inductor (200a) within the loop (200) determines the range and resolution (effective bit depth) of the storable weight valueand generates a magnetic flux proportional to this stored current. Implementation involves patterning superconducting thin films (like Niobium) into a closed loop geometry containing the inductor (200a). A primary benefit of this element is the ability to store the analog weight value non-volatilely with essentially zero static power dissipation, contributing significantly to the potential energy efficiency of the neuromorphic system. Variations might include different inductor geometries or incorporating materials to enhance flux focusing, although the basic structure remains a superconducting loop with inductance. In operation, learning circuitry adjusts the persistent current in this loop (200) to update the synaptic weight.
[0039] The circuit incorporates a first superconducting quantum interference device (SQUID) (201 ) which is mutually coupled to the superconducting weight storage loop (200) with a first polarity. The first SQUID (201 ) functions as a magnetically tunable inductor. Its effective inductance is highly sensitive to the magnetic flux generated by the current circulating in the weight storage loop (200). Mutual coupling, achieved through proximity between the SQUID loop and the weight storage loop (200), enables this control. The 'first polarity' refers to the specific directional relationship of this coupling. SQUIDs (201) are typically implemented using one or two Josephson junctions within a small superconducting loop, fabricated using standard superconducting circuit processes. This arrangement advantageously translates the stored analog weight current into a variable impedance (inductance) within one path of the circuit, providing a means to modulate signal flow based on the weight value, leveraging the high sensitivity and speed of SQUIDs. Alternative implementations could involve different SQUID geometries or junction types. As an example, a clockwise current in the weight storage loop (200) might generate a flux that increases the inductance presented by the first SQUID (201 ).
[0040] Complementing the first SQUID, the circuit includes a second SQUID (202) mutually coupled to the superconducting weight storage loop (200) with a second polarity opposite the first polarity. The inductance of both the first SQUID (201 ) and the second SQUID (202) vary based upon the current circulating within the superconducting weight storage loop (200). This second SQUID (202) also functions as a variable inductor controlled by the weight loop current. Crucially, its coupling polarity is opposite to that of the first SQUID (201 ), often achieved by placing theSQUIDs physically on opposite sides of the weight loop conductor or using opposite winding directions in the coupling design. This ensures that as the weight loop current changes, the inductance of the second SQUID (202) varies inversely compared to the first SQUID (201 ). This differential variation is key to the circuit's ability to produce a bipolar output from a single unipolar weight current; it provides improved circuit efficiency and potentially enhanced noise immunity compared to non-differential schemes. Variations could explore different physical arrangements for achieving opposite coupling. For instance, if the clockwise weight loop current increases the inductance of the first SQUID (201 ), it would simultaneously decrease the inductance of the second SQUID (202).
[0041] The circuit features a first inductive divider (203) electrically connected to an input terminal (210a). This divider (203) comprises a first leg including the first SQUID (201) connected to ground (GND) and a second leg including a first coupling inductor (205). The function of this divider (203) is to split an incoming SFQ pulse current arriving at the input terminal (210a). The current divides between the path through the first SQUID (201 ) to ground and the path through the first coupling inductor (205). The ratio of this division is determined by the relative inductances of these two paths, specifically the dynamically varying inductance of the first SQUID (201 ) controlled by the weight current. It is implemented as a branching point in the superconducting wiring. This structure provides the mechanism for modulating the portion of the input signal that will contribute positively (due to the coupling inductor's polarity) to the final sum, directly implementing part of the weight multiplication. Alternative divider designs might exist, but this structure is typical for SFQ current division. If the first SQUID (201 ) inductance is high, a larger fraction of the input current pulse flows through the first coupling inductor (205).
[0042] Similarly, a second inductive divider (204) is electrically connected to the input terminal (210a), comprising a first leg including the second SQUID (202) connected to ground (GND) and a second leg including a second coupling inductor (206). This divider (204) performs the complementary current splitting function for the parallel path. Connected to the same input terminal (210a), often via a splitter, it divides the input current based on the inductance of the second SQUID (202) relative to the path containing the second coupling inductor (206). Since the second SQUID'Sinductance varies oppositely to the first, the current division here is complementary to that in the first divider (203). Implemented similarly with superconducting wiring, this structure modulates the portion of the input signal destined for negative contribution to the sum. The coordinated action of both dividers (203, 204), driven differentially by the single weight current, enables the bipolar weighting functionality efficiently. If the second SQUID (202) inductance is low, a smaller fraction of the input current flows through the second coupling inductor (206).
[0043] The circuit incorporates a soma element (207) comprising a soma inductor line (207a). This element (207) functions as the integration or summation point for the weighted signals, analogous to a biological neuron's soma. The soma inductor line (207a) is typically a patterned superconducting trace or a series of inductors designed to allow signals from multiple synapses (represented by coupling inductors) to be combined via mutual inductance. Its implementation involves standard superconducting layer patterning. This provides a common structure where the opposing signal contributions can effectively sum or subtract, forming the net weighted input current. Variations could include different inductor line geometries or the inclusion of damping elements. This line (207a) collects the outputs from the coupling inductors (205, 206).
[0044] The first coupling inductor (205) is mutually coupled to the soma inductor line (207a) with a third polarity. This inductor (205), carrying the modulated current from the first divider (203), transfers this signal into the soma inductor line (207a). Mutual coupling, achieved by placing inductor (205) near the soma line (207a), allows efficient, contactless signal transfer. The 'third polarity' defines the sign (e.g., positive or excitatory) of the contribution induced in the soma line (207a) by current in the first coupling inductor (205). This facilitates the summation of the positively weighted component. Different physical layouts can vary the mutual inductance M, affecting the gain.
[0045] Correspondingly, the second coupling inductor (206) is mutually coupled to the soma inductor line (207a) with a fourth polarity opposite the third polarity. Carrying the modulated current from the second divider (204), this inductor (206) also transfers its signal to the soma line (207a), but the coupling is arranged (via relative position or winding) to induce a signal of opposite polarity (e.g., negative orinhibitory) compared to the first coupling inductor (205). This arrangement is fundamental to achieving the bipolar summation, allowing the signal from the second path to effectively subtract from the signal from the first path within the soma line (207a). The net effect is determined by the balance set by the weight current acting on the dividers.
[0046] The soma element (207) further comprises a soma output Josephson junction (209) connected to the soma inductor line (207a). This junction (209) serves as the thresholding element of the artificial neuron. It monitors the net current integrated in the soma inductor line (207a). If this net current exceeds the junction's intrinsic critical current (plus any applied bias), the junction switches from the superconducting state to the resistive state, generating an output SFQ voltage pulse. If the threshold is not reached, it remains superconducting, and no output pulse is produced. Implemented as a standard Josephson junction integrated at the output of the soma line (207a), it performs the crucial non-linear activation function. This converts the summed analog weighted current into a digital SFQ pulse suitable for transmission to subsequent stages in the neural network, completing the neuron's computation with high speed and low energy consumption. Different junction parameters allow tuning of the activation threshold.
[0047] The interaction between these components facilitates the circuit's core function: applying a programmable bipolar weight to an incoming unipolar SFQ pulse. The superconducting weight storage loop (200) establishes a controllable magnetic flux that differentially modulates the inductances of the first SQUID (201 ) and the second SQUID (202) due to their opposite coupling polarities. This differential inductance change directly influences the current division ratios within the first inductive divider (203) and the second inductive divider (204). Consequently, the portions of the input signal directed through the first coupling inductor (205) and the second coupling inductor (206) are adjusted in opposite directions based on the single stored weight current. The opposing polarities (third and fourth) with which these coupling inductors (205, 206) interact with the soma inductor line (207a) ensure that these differentially modulated signal portions contribute additively or subtractively to the net current in the soma element (207). The soma output Josephson junction (209)then applies a threshold to this net current, producing an output pulse only when the integrated signal is sufficiently large.
[0048] Operationally, an SFQ pulse arriving at the input terminal (210a) is simultaneously fed into both the first inductive divider (203) and the second inductive divider (204). Within the first divider (203), the pulse current splits between the path through the first SQUID (201 ) to ground and the path through the first coupling inductor(205). The proportion flowing through the first coupling inductor (205) is determined by the inductance of the first SQUID (201), which is set by the weight current in the loop (200). Concurrently, within the second divider (204), the pulse current splits between the path through the second SQUID (202) to ground and the path through the second coupling inductor (206), with this division ratio controlled by the inversely varying inductance of the second SQUID (202). The current pulse portion traversing the first coupling inductor (205) induces a signal of a specific polarity (e.g., positive) in the soma inductor line (207a), while the portion traversing the second coupling inductor(206) induces a signal of the opposite polarity (e.g., negative). These induced signals summate within the soma inductor line (207a). If the magnitude of this summed current exceeds the threshold of the soma output Josephson junction (209), an output SFQ pulse is generated.
[0049] The bipolar nature of the weighting is achieved through this differential structure's response to the stored weight current. When the current in the superconducting weight storage loop (200) is near zero, the inductances of the first SQUID (201 ) and the second SQUID (202) may be designed to be approximately equal, leading to roughly equal current division in both inductive dividers (203, 204). The oppositely coupled signals induced in the soma inductor line (207a) via the coupling inductors (205, 206) then largely cancel, resulting in a near-zero net signal and likely no output pulse from the soma output Josephson junction (209), effectively representing a zero weight. If a significant current circulates in the weight storage loop(200) in one direction, it might, for example, increase the inductance of the first SQUID(201 ) and decrease the inductance of the second SQUID (202). This would cause a larger portion of the input pulse to flow through the first coupling inductor (205) (inducing a positive signal) and a smaller portion through the second coupling inductor (206) (inducing a negative signal), resulting in a net positive current in the somainductor line (207a) and potentially triggering the output junction (209), representing a positive weight. Conversely, circulating current in the opposite direction in the loop (200) would reverse the inductance changes in the SQUIDs (201 , 202), leading to a smaller positive contribution and a larger negative contribution, resulting in a net negative (inhibitory) current in the soma inductor line (207a), representing a negative weight.
[0050] This bipolar single flux quantum weight circuit architecture provides notable advantages for constructing SFQ-based neuromorphic systems. It offers improved functionality by enabling both excitatory and inhibitory synaptic connections using standard unipolar SFQ signals, a capability often required for complex neural network algorithms. The design achieves this bipolarity efficiently, utilizing a single analog weight storage loop (200) to differentially control two signal paths, thereby potentially reducing the circuit area and complexity compared to approaches requiring duplicated components for positive and negative weights. The integration of tunable SQUIDs (201 , 202) as variable impedance elements and the use of inductive dividers (203, 204) and mutual coupling allows for analog modulation and summation operations compatible with the high-speed, low-power characteristics of SFQ technology. The direct coupling into a soma element (207) with an output Josephson junction (209) seamlessly integrates the synaptic weighting function with the neuron's thresholding activation, providing a fundamental building block for scalable superconducting neural networks.
[0051] The superconducting weight storage loop (200) may further incorporate a first Josephson junction (200b) and a second Josephson junction (200c), specifically arranged to facilitate the injection of single flux quantum (SFQ) pulses. This arrangement functions as the mechanism for programming or updating the synaptic weight stored as current in the loop (200). By selectively triggering SFQ pulses into the loop (200) via either the first junction (200b) or the second junction (200c), the circulating persistent current can be incrementally increased or decreased. These junctions (200b, 200c) are implemented as standard Josephson junctions integrated within the storage loop (200), often biased appropriately. This feature provides a digitally controlled method for adjusting the analog weight value using the native SFQ pulses generated by learning circuitry, allowing for precise, quantizedupdates to the synapse's strength and enabling on-chip learning capabilities within the SFQ architecture. Variations could involve different junction configurations or biasing schemes. For instance, applying a pulse via junction (200b) might increment the clockwise current, while a pulse via junction (200c) might decrement it (or induce counter-clockwise current).
[0052] It may be provided that the second leg of the first inductive divider(203) further comprises a first resistor (205a) in series with the first coupling inductor (205) connected to ground (GND), and the second leg of the second inductive divider(204) further comprises a second resistor (206a) in series with the second coupling inductor (206) connected to ground (GND). These resistors (205a, 206a) serve primarily to set the L / R time constant, or decay rate, for signals passing through the coupling inductor paths (205, 206) of the dividers (203, 204). They also help prevent unintended persistent current buildup or ringing within these divider legs. Implementation involves integrating resistive material layers in series with the coupling inductors (205, 206) during fabrication. A benefit is improved signal integrity and potentially relaxed timing constraints for input pulses by controlling the effective duration of the signal coupled into the soma element (207). Alternative implementations might use different resistive materials or configurations. These resistors help ensure that the induced signals in the soma element (207) decay appropriately after the input pulse passes.
[0053] The soma element (207) may further comprise a soma resistor (207b) connecting the soma inductor line (207a) to ground (GND), and the soma output Josephson junction (209) may connect the soma inductor line (207a) to an output Josephson transmission line (208). The soma resistor (207b) establishes an L / R time constant for the entire soma element (207), defining how quickly the summed current in the soma inductor line (207a) decays. This 'leakiness' allows the soma (207) to integrate signals arriving slightly asynchronously due to wiring delays or jitter, improving robustness. The resistor (207b) is typically implemented as a shunt resistor connecting the end of the soma line (207a) opposite the output junction (209) to the ground plane. Connecting the soma output Josephson junction (209) to an output Josephson transmission line (208) provides the pathway for propagating the generated output SFQ pulse (if the threshold is exceeded) to subsequent stages ofthe neural network. The transmission line (208) is a standard SFQ interconnect. These features refine the temporal integration properties of the soma (207) and ensure proper signal output and propagation, contributing to reliable network operation.
[0054] It may be provided that the first SQUID (201 ) is physically positioned relative to the superconducting weight storage loop (200) differently than the second SQUID (202) specifically to achieve the first polarity and the second, opposite polarity of mutual coupling. This geometric arrangement is a practical method for realizing the differential coupling. For example, placing the conductive trace of the weight storage loop (200) between the loops of the first SQUID (201 ) and the second SQUID (202), perhaps with one SQUID above the trace and the other below in a multi-layer process, inherently results in the magnetic flux from the loop current threading the two SQUIDs in opposite directions. This physical layout directly implements the opposite coupling polarities using standard fabrication techniques, providing a reliable way to construct the differential inductance modulation mechanism central to the bipolar weight function. Alternative physical layouts that achieve opposite flux threading could also be employed.
[0055] It may be provided that the mutual coupling of the first coupling inductor (205) to the soma inductor line (207a) and the mutual coupling of the second coupling inductor (206) to the soma inductor line (207a) result in signals induced in the soma element (207) that are opposite in polarity relative to each other upon application of an input pulse to the input terminal (210a). This describes the functional outcome of the specific coupling arrangement between the divider outputs and the soma input. By designing the physical placement or winding direction of the coupling inductors (205, 206) relative to the soma line (207a) appropriately, the currents flowing through them induce opposing voltage or current signals in the soma line (207a). This engineered opposite polarity induction ensures that the contributions from the two divider paths subtract from each other, directly implementing the summation of positive and negative weighted components required for bipolar synaptic function. This is achieved through layout design during the circuit fabrication process.
[0056] The circuit may further comprise an input splitter (210) connecting the input terminal (210a) to both the first inductive divider (203) and the second inductive divider (204). The input splitter (210) serves to duplicate the incoming SFQpulse from the input terminal (210a) and deliver it simultaneously to the inputs of both the first (203) and second (204) inductive dividers. This is typically implemented using a simple T-junction or a more sophisticated SFQ splitter gate designed to maintain signal integrity and proper impedance matching. This ensures that both differential paths receive the same input signal concurrently, which is necessary for the differential division process based on the SQUID inductances (201 , 202) to accurately reflect the stored weight value. It enables the parallel processing essential for the differential weighting scheme.
[0057] It may be provided that the first SQUID (201 ) and the second SQUID (202) are biased to remain in a non-spiking state across a range of the current circulating within the superconducting weight storage loop (200). This operational condition ensures that the SQUIDs (201 , 202) function purely as variable inductors, smoothly changing their inductance in response to the weight current, rather than acting as thresholding devices themselves by firing SFQ pulses. This is achieved by setting the bias currents applied to the Josephson junctions within the SQUIDs (201 , 202) sufficiently below their critical currents, considering the additional flux contributed by the weight storage loop (200). Maintaining the non-spiking state provides a smooth, analog variation in the divider ratios as the weight current changes, allowing for finer control over the synaptic weight and contributing to the circuit's function as an analog multiplier followed by summation. This bias condition is set during circuit operation.
[0058] The circuit may further comprise a bias weight circuit (211 ), including a bias weight storage loop (212) and a bias weight inductive divider leg (213), where the bias weight inductive divider leg (213) is mutually coupled to the soma inductor line (207a). This bias weight circuit (211 ) functions as an adjustable offset or bias input to the soma element (207). It operates similarly to one branch (e.g., the positive branch) of the main bipolar synapse but has its own independently programmable weight stored in its bias weight storage loop (212). Its output, via the bias weight inductive divider leg (213), is coupled (typically with a fixed positive polarity) into the same soma inductor line (207a) as the main synaptic inputs. This allows the effective firing threshold of the soma output Josephson junction (209) to be adjusted dynamically by the learning process, independently of the main synaptic inputs. Thisadjustable bias is often crucial for enabling neural networks to learn certain functions effectively. Its implementation mirrors a single branch of the bipolar synapse.
[0059] Finally, it may be provided that the superconducting weight storage loop (200), the first SQUID (201 ), the second SQUID (202), the first coupling inductor (205), the second coupling inductor (206), the soma inductor line (207a), and the soma output Josephson junction (209) comprise superconducting materials and Josephson junctions. Utilizing superconducting materials (like niobium, aluminum, etc.) enables near-lossless current flow and the operation of Josephson junctions, which are essential for SFQ pulse generation / detection and the function of SQUIDs. Josephson junctions, typically superconductor-insulator-superconductor tunnel junctions, provide the non-linear switching behavior required. This material basis is what enables the ultra-high speed and ultra-low power operation potential of the entire circuit architecture. The specific choice of superconducting materials and junction parameters affects operating temperature and performance characteristics.
[0060] In an embodiment, a method for weighting a unipolar single flux quantum (SFQ) input pulse (x) utilizing a bipolar single flux quantum weight circuit comprising a superconducting weight storage loop (200) containing a variable weight current, a first SQUID (201 ) and a second SQUID (202) oppositely mutually coupled to the weight storage loop (200), a first inductive divider (203) including the first SQUID (201 ) and a first coupling inductor (205), a second inductive divider (204) including the second SQUID (202) and a second coupling inductor (206), and a soma element (207) having a soma inductor line (207a) mutually coupled with opposite polarities to the first coupling inductor (205) and the second coupling inductor (206) and connected to a soma output Josephson junction (209), comprises: applying the unipolar SFQ input pulse (x) to an input terminal (210a) connected to the first inductive divider (203) and the second inductive divider (204); dividing the SFQ input pulse (x) via the first inductive divider (203), wherein a first portion is directed through the first SQUID (201 ) and a second portion is directed through the first coupling inductor (205), the division ratio being dependent on an inductance of the first SQUID (201 ) influenced by the weight current; dividing the SFQ input pulse (x) via the second inductive divider (204), wherein a third portion is directed through the second SQUI D (202) and a fourth portion is directed through the second coupling inductor (206), the division ratio beingdependent on an inductance of the second SQUID (202) influenced by the weight current; inducing a first signal in the soma inductor line (207a) from the second portion via the first coupling inductor (205) with a first signal polarity; inducing a second signal in the soma inductor line (207a) from the fourth portion via the second coupling inductor (206) with a second signal polarity opposite the first signal polarity; and generating a net signal in the soma element (207) based on a summation of the first induced signal and the second induced signal, wherein the soma output Josephson junction (209) emits an output SFQ pulse (y) if the net signal exceeds a predetermined threshold. In an embodiment, the method further comprises programming the variable weight current in the superconducting weight storage loop (200) by injecting one or more SFQ current pulses into the superconducting weight storage loop (200) via a first Josephson junction (200b) or a second Josephson junction (200c) associated with the loop (200). In an embodiment, injecting the one or more SFQ current pulses adjusts the variable weight current to increase or decrease a magnitude of the net signal generated in the soma element (207) in response to the SFQ input pulse (x). In an embodiment, when the variable weight current in the superconducting weight storage loop (200) is substantially zero, the first induced signal and the second induced signal substantially cancel each other, resulting in the net signal being substantially zero. In an embodiment, adjusting the variable weight current in a first direction increases the magnitude of the first induced signal and decreases the magnitude of the second induced signal, resulting in a net signal with the first signal polarity. In an embodiment, adjusting the variable weight current in a second direction, opposite the first direction, decreases the magnitude of the first induced signal and increases the magnitude of the second induced signal, resulting in a net signal with the second signal polarity. In an embodiment, the method further comprises operating the circuit within a reinforcement learning framework, including: performing an inference cycle by applying the SFQ input pulse (x) and generating the output SFQ pulse (y) or no pulse; storing a value representing the SFQ input pulse (x) and a value representing the output SFQ pulse (y) in memory elements (215); receiving a learning clock signal (217) to initiate a learning cycle; calculating a weight update value based on the stored value representing the SFQ input pulse (x), the stored value representing the output SFQ pulse (y), a target output value (y*), and potentially a reinforcement signal (r) using learning rule logic (214); and adjusting the variable weight current in the superconducting weight storage loop (200) based on the calculated weight updatevalue by injecting SFQ pulses. In an embodiment, calculating the weight update value for hidden layers within a neural network utilizes a change in the reinforcement signal (Ar) between a previous learning cycle and a current learning cycle and a change in the output SFQ pulse (Ay) between a previous inference cycle and a current inference cycle. In an embodiment, the method further comprises applying a stochastic SFQ pulse generated by a stochastic hardware element (216) to the soma element (207) during the inference cycle to enhance weight exploration, wherein the stochastic hardware element (216) utilizes thermal fluctuations at an operating temperature to generate random SFQ pulses. In an embodiment, the method forms part of a neural network computation, the net signal representing a weighted synaptic input to a neuron, and the output SFQ pulse (y) representing a neuron firing.
[0061] The method commences with applying the unipolar SFQ input pulse (x) to an input terminal (210a) connected to the first inductive divider (203) and the second inductive divider (204). This step serves to introduce the signal that requires weighting into the parallel processing paths of the circuit. The input pulse (x), typically representing an activation from a preceding neural stage, propagates to the input terminal (210a), which acts as the circuit's signal entry point. From this terminal, the signal is directed, often via an input splitter, concurrently into both the first (203) and second (204) inductive dividers. This simultaneous application ensures that the differential weighting mechanism begins processing the input without temporal skew between the paths, supporting the high-speed operation inherent in SFQ circuits. Variations could involve different means of delivering the SFQ pulse (x) to the terminal (210a). In practice, this step represents the arrival of information at the synapse.
[0062] Following input application, the method involves dividing the SFQ input pulse (x) via the first inductive divider (203). During this division, a first portion of the pulse's current is directed through the first SQUID (201 ) towards ground, while a second portion is directed through the first coupling inductor (205). The critical aspect here is that the division ratio — the proportion of current flowing through the coupling inductor (205) versus the SQUID (201) — is not fixed but depends dynamically on the inductance presented by the first SQUID (201 ). This inductance is directly influenced and controlled by the magnitude of the persistent weight current circulating within the superconducting weight storage loop (200). The physical implementation relies on theinherent current-splitting behavior of parallel inductive paths. This step provides a mechanism for analog modulation of the signal component that will represent the positive contribution to the weighted sum, effectively performing one half of the analog multiplication defined by the stored weight. The use of the tunable SQUID inductance allows for this modulation to occur efficiently at SFQ speeds. If the weight current causes the first SQUID (201 ) inductance to be high, a larger second portion flows via the first coupling inductor (205).
[0063] Concurrently, the method includes dividing the SFQ input pulse (x) via the second inductive divider (204). Here, a third portion of the current flows through the second SQUID (202) to ground, and a fourth portion flows through the second coupling inductor (206). Similar to the first division, the ratio is dependent on the inductance of the second SQUID (202), which is also influenced by the same weight current from the loop (200). However, due to the opposite mutual coupling polarity, the inductance of the second SQUID (202) varies inversely compared to the first SQUID (201 ). This complementary division ensures that as the second portion (through inductor 205) increases or decreases, the fourth portion (through inductor 206) does the opposite. Implemented through the parallel inductive paths of the second divider (204), this step provides the complementary analog modulation for the signal component representing the negative contribution to the weighted sum. This differential division driven by a single weight current is a key benefit, enabling bipolar weighting without requiring separate weight storage. If the weight current causes the second SQUID (202) inductance to be low, a smaller fourth portion flows via the second coupling inductor (206).
[0064] The method proceeds with inducing a first signal in the soma inductor line (207a) from the second portion via the first coupling inductor (205) with a first signal polarity. This step transfers the modulated signal component from the first divider's output path into the central summation structure. The current constituting the second portion, flowing through the first coupling inductor (205), generates a timevarying magnetic flux that is mutually coupled to the soma inductor line (207a). This coupling induces a corresponding signal (voltage or current pulse) within the soma line (207a). The physical arrangement dictates the 'first signal polarity', conventionally representing the positive or excitatory influence. This use of mutual inductanceprovides an efficient, high-speed mechanism for signal transfer and contribution to the summation. Different coupling efficiencies can be achieved through design variations. This step effectively delivers the positively weighted signal component to the neuron's input integration stage.
[0065] Simultaneously or nearly simultaneously, the method involves inducing a second signal in the soma inductor line (207a) from the fourth portion via the second coupling inductor (206) with a second signal polarity opposite the first signal polarity. This mirrors the previous step but for the complementary path. The current constituting the fourth portion, flowing through the second coupling inductor (206), induces a signal in the same soma inductor line (207a) via mutual coupling. Critically, the design ensures this induced signal has the opposite polarity (e.g., negative or inhibitory) compared to the first induced signal. This opposing polarity is achieved by the relative physical arrangement or winding sense of the second coupling inductor (206) with respect to the soma line (207a). This step delivers the negatively weighted signal component to the same integration stage, setting up the direct summation of opposing influences necessary for bipolar operation.
[0066] Finally, the method culminates in generating a net signal in the soma element (207) based on a summation of the first induced signal and the second induced signal. The soma inductor line (207a) acts as the physical location where these two oppositely polarized induced signals combine linearly. The resulting net signal's amplitude and polarity represent the final bipolar weighted value of the input pulse (x), determined by the stored weight current. The method concludes with the action of the soma output Josephson junction (209); this junction evaluates the net signal against its intrinsic critical current (acting as a predetermined threshold). If this threshold is exceeded by the net signal, the junction (209) switches and emits an output SFQ pulse (y). If the net signal is below the threshold (including being zero or negative), the junction (209) remains in its superconducting state, and no output pulse is generated. This step performs the essential neuronal functions of summing weighted inputs and applying a non-linear activation threshold, producing a digital SFQ output suitable for further processing in a network while efficiently leveraging the physics of the superconducting components for speed and low power. This processsuccessfully transforms the unipolar input pulse (x) into an output (y or null) that reflects bipolar weighting based on the stored analog value.
[0067] The method may further involve programming the variable weight current residing within the superconducting weight storage loop (200). This programming is accomplished by the directed injection of one or more single flux quantum (SFQ) current pulses into the loop (200), utilizing either a first Josephson junction (200b) or a second Josephson junction (200c) specifically incorporated within or associated with the loop (200) for this purpose. This step provides the essential mechanism for modifying the synapse's strength. The junctions (200b, 200c), acting as controlled entry points, allow external learning circuitry to incrementally add or remove flux quanta, thereby precisely adjusting the persistent current magnitude that represents the analog weight. This SFQ-based programming method integrates seamlessly with the overall circuit operation, enabling efficient, digitally controlled updates to the analog weight value, which is fundamental for implementing on-chip learning algorithms within the superconducting hardware. For example, a pulse sent to junction (200b) could increase the stored current, while a pulse to junction (200c) could decrease it.
[0068] Injecting these one or more SFQ current pulses serves to adjust the variable weight current such that it increases or decreases the magnitude of the net signal ultimately generated within the soma element (207) in response to a subsequent SFQ input pulse (x). This step clarifies the direct functional consequence of the weight programming action. By altering the persistent current in the loop (200), the magnetic flux influencing the first SQUID (201 ) and the second SQUID (202) is changed. This, in turn, modifies their inductances and consequently alters the current division ratios in the inductive dividers (203, 204). The result is a change in the relative magnitudes of the first and second induced signals within the soma inductor line (207a), leading to an adjustment in the overall net signal amplitude and potentially its polarity. This establishes the direct link between the digitally controlled weight update mechanism and the resulting analog modulation performed by the synapse circuit, confirming its role in adaptive synaptic weighting. An increase in positive weight current, achieved via pulse injection, will lead to a larger positive net signal for a given input (x).
[0069] It may be provided that when the variable weight current within the superconducting weight storage loop (200) is substantially zero, the first induced signal and the second induced signal resulting from an input pulse (x) substantially cancel each other. This cancellation leads to the net signal in the soma element (207) being substantially zero. This describes the baseline behavior of the circuit when representing a neutral or zero synaptic weight. This functionality arises from the symmetric design of the differential structure, where, in the absence of a controlling current in the loop (200), the inductances of the SQUIDs (201 , 202) may be designed to be similar, leading to comparable signal portions being directed through the coupling inductors (205, 206). Since these coupling inductors induce signals of opposite polarity in the soma line (207a), their effects largely negate each other. This capability to accurately represent a zero weight is significant for unbiased operation in many neural network applications. Minor imbalances due to fabrication variations might prevent perfect cancellation, but the net signal would be close to zero.
[0070] Adjusting the variable weight current in a first direction (e.g., clockwise) causes an increase in the magnitude of the first induced signal (originating from the first coupling inductor 205) and a simultaneous decrease in the magnitude of the second induced signal (from the second coupling inductor 206). This imbalance results in a net signal within the soma element (207) that exhibits the first signal polarity (e.g., positive). This operational mode demonstrates the circuit's capacity to implement positive synaptic weights. The current flow in the first direction alters the SQUID inductances (201 , 202) differentially, causing the first inductive divider (203) to pass a larger signal portion to the first coupling inductor (205) while the second divider (204) passes a smaller portion to the second coupling inductor (206). Given the opposite coupling polarities to the soma line (207a), the positive contribution outweighs the negative one. This provides the mechanism for excitatory synaptic action within the SFQ framework.
[0071] Conversely, adjusting the variable weight current in a second direction, opposite the first direction (e.g., counter-clockwise), leads to a decrease in the magnitude of the first induced signal and an increase in the magnitude of the second induced signal. This reversed imbalance results in a net signal in the soma element (207) possessing the second signal polarity (e.g., negative). Thisdemonstrates the implementation of negative synaptic weights. The opposite current flow reverses the differential inductance changes in the SQUIDs (201 , 202), now favoring signal flow through the second coupling inductor (206) over the first (205). Consequently, the negative contribution induced in the soma line (207a) dominates the positive one, yielding a net negative signal. This capability for inhibitory action is crucial for the computational power of many neural networks and is achieved here efficiently through the differential structure.
[0072] The method may further encompass operating the bipolar weight circuit as part of a broader reinforcement learning framework integrated within the hardware. This operation involves distinct cycles. First, an inference cycle occurs where the SFQ input pulse (x) is applied, processed by the weight circuit, and potentially results in an output SFQ pulse (y) from the soma output Josephson junction (209) (or no pulse). Following inference, essential state information — specifically, a value representing the input pulse (x) occurrence (e.g., 1 or 0) and a value representing the output pulse (y) occurrence — is stored in dedicated memory elements (215). Subsequently, if learning is enabled, a learning clock signal (217) arrives, initiating a learning cycle. During this cycle, learning rule logic circuitry (214) calculates a weight update value. This calculation utilizes the stored input (x) and output (y) values, compares the actual output (y) with a desired or target output value (y*), and potentially incorporates a global or local reinforcement signal (r). Finally, based on the calculated update value, the learning rule logic (214) triggers the injection of the appropriate number and direction of SFQ pulses into the superconducting weight storage loop (200), thereby adjusting the variable weight current. This sequence enables the circuit to adapt its synaptic weight based on performance feedback directly within the SFQ hardware, facilitating self-training and potentially continuous learning, offering significant advantages in speed and efficiency over software-based training paradigms.
[0073] Within the reinforcement learning framework, the step of calculating the weight update value for synapses associated with hidden layers (neurons not directly connected to the final output) may specifically utilize a change in the reinforcement signal (Ar) and a change in the neuron's own output (Ay). The change Ar compares the reinforcement signal received in the current learning cycle with thatfrom the previous cycle, indicating whether the network's overall performance improved orworsened. The change Ay compares the hidden neuron's output (y) (spike or no spike) in the current inference cycle with its output in the previous cycle. This calculation is typically performed by the learning rule logic (214) using stored values. Employing these change signals (Ar, Ay) along with the stored input (x) value allows the implementation of correlative learning rules, such as those derived from the REINFORCE algorithm family. These rules adjust weights based on the correlation between local activity changes (Ay) and global performance changes (Ar), providing a mechanism for credit assignment in multi-layer networks that is more readily implementable in hardware than full backpropagation, thus enabling efficient training of deeper superconducting networks.
[0074] The method may also include the step of applying a stochastic SFQ pulse to the soma element (207) during the inference cycle. This stochastic pulse is generated by a separate stochastic hardware element (216), which is designed to leverage thermal fluctuations present at the cryogenic operating temperature (e.g., 4 K) to produce random SFQ pulse events. This element (216), possibly resembling the synapse circuit but using components like low-critical-current Josephson junctions highly sensitive to thermal energy, introduces controlled randomness into the neuron's summation process. The addition of this stochastic pulse to the deterministic weighted sum in the soma inductor line (207a) serves to enhance the exploration of the possible weight configurations during the learning process. This intentional noise injection helps the learning algorithm avoid getting trapped in suboptimal local minima of the error landscape, promoting convergence towards better overall network solutions, particularly for complex, non-linear problems. This utilizes a naturally available source of randomness (thermal noise) to improve the robustness and effectiveness of the on- chip learning.
[0075] Finally, the described method inherently forms part of a larger neural network computation. The net signal generated within the soma element (207) directly corresponds to the weighted synaptic input arriving at a neuron from a specific connection. The subsequent action of the soma output Josephson junction (209) — either emitting an output SFQ pulse (y) or remaining quiescent — represents the neuron firing (activation) or not firing, based on the comparison of the integrated weightedinput against its threshold. This contextualization highlights the direct mapping between the physical processes occurring in the bipolar SFQ weight circuit and the fundamental operations of synaptic weighting and neuronal activation within an artificial neural network model. It underscores the utility of the method as a building block for constructing functional, high-performance neuromorphic computing systems based on superconducting electronics.
[0076] In an embodiment, a method for fabricating a bipolar single flux quantum weight circuit comprises: forming a superconducting weight storage loop (200) including a weight storage inductor (200a) from superconducting material layers; forming a first superconducting quantum interference device (SQUID) (201 ) and a second SQUID (202) adjacent to the superconducting weight storage loop (200); arranging the first SQUID (201) relative to the superconducting weight storage loop (200) to establish a mutual inductive coupling exhibiting a first polarity; arranging the second SQUID (202) relative to the superconducting weight storage loop (200) to establish a mutual inductive coupling exhibiting a second polarity substantially opposite to the first polarity; forming a first inductive divider (203) comprising a first leg including the first SQUID (201 ) and a second leg including a first coupling inductor (205); forming a second inductive divider (204) comprising a first leg including the second SQUID (202) and a second leg including a second coupling inductor (206); forming a soma element (207) comprising a soma inductor line (207a) and a soma output Josephson junction (209) connected thereto; arranging the first coupling inductor (205) relative to the soma inductor line (207a) to establish a mutual inductive coupling exhibiting a third polarity; arranging the second coupling inductor (206) relative to the soma inductor line (207a) to establish a mutual inductive coupling exhibiting a fourth polarity substantially opposite to the third polarity; and electrically connecting the first inductive divider (203) and the second inductive divider (204) to an input terminal (210a), connecting the first leg of the first inductive divider (203) and the first leg of the second inductive divider (204) to ground (GND), and connecting the second leg of the first inductive divider (203) and the second leg of the second inductive divider (204) to ground (GND). In an embodiment, the method further comprises forming a first Josephson junction (200b) and a second Josephson junction (200c) within the superconducting weight storage loop (200), positioned to enable bidirectional current injection. In an embodiment, the method further comprisesforming a first resistor (205a) in series with the first coupling inductor (205) in the second leg of the first inductive divider (203), and forming a second resistor (206a) in series with the second coupling inductor (206) in the second leg of the second inductive divider (204). In an embodiment, the method further comprises forming a soma resistor (207b) connecting the soma inductor line (207a) to ground (GND), and electrically connecting the soma output Josephson junction (209) to an output Josephson transmission line (208). In an embodiment, arranging the first SQUID (201 ) and the second SQUID (202) relative to the superconducting weight storage loop (200) involves positioning conductive layers of the first SQUID (201 ) and the second SQUID (202) on opposing sides of a conductive layer of the superconducting weight storage loop (200). In an embodiment, arranging the first coupling inductor (205) and the second coupling inductor (206) relative to the soma inductor line (207a) involves configuring winding directions or relative positions to achieve the third and fourth opposite polarities. In an embodiment, the method further comprises forming an input splitter (210) configured to electrically connect the input terminal (210a) simultaneously to the first inductive divider (203) and the second inductive divider (204). In an embodiment, the method further comprises forming electrical connections for applying bias currents to the first SQUID (201) and the second SQUID (202) sufficient to maintain them in a non-spiking state during operation. In an embodiment, the method further comprises: forming a bias weight circuit (211 ) including a bias weight storage loop (212) and a bias weight inductive divider leg (213); and arranging the bias weight inductive divider leg (213) relative to the soma inductor line (207a) to establish a mutual inductive coupling. In an embodiment, forming the superconducting weight storage loop (200), the first SQUID (201 ), the second SQUID (202), the first coupling inductor (205), the second coupling inductor (206), the soma inductor line (207a), and the soma output Josephson junction (209) utilizes deposition and patterning of superconducting thin films, insulating layers, and resistive materials.
[0077] The process for fabricating a bipolar single flux quantum weight circuit commences with forming a superconducting weight storage loop (200), which includes a weight storage inductor (200a), utilizing layers of superconducting material. This initial step establishes the core component responsible for storing the synaptic weight as a persistent current. The implementation typically involves standard microfabrication techniques such as sputter deposition of a superconducting thin film(e.g., Niobium) onto a substrate, followed by photolithography and etching (e.g., reactive ion etching) to define the closed-loop geometry incorporating the inductor structure (200a). The formation of this superconducting loop (200) provides the benefit of enabling non-volatile analog weight storage with minimal static power dissipation, a significant advantage for energy-efficient neuromorphic computing. Variations could involve using different superconducting materials or employing multi-layer processes to create more complex inductor geometries for optimizing inductance or minimizing footprint. This step lays the foundation for the analog memory element of the synapse.
[0078] Subsequently, the method involves forming a first superconducting quantum interference device (SQUID) (201 ) and a second SQUID (202) adjacent to the previously formed superconducting weight storage loop (200). This step creates the variable impedance elements controlled by the weight current. Forming the SQUIDs (201 , 202) requires additional deposition and patterning steps to create the small superconducting loops containing one or two Josephson junctions (typically superconductor-insulator-superconductor tunnel junctions). These SQUIDs (201 , 202) are positioned in close proximity to the weight storage loop (200) to facilitate mutual inductive coupling. This provides the tunable inductors necessary for modulating the signal paths based on the stored weight, enabling the core function of synaptic weighting. Alternative SQUID designs or junction types could be employed. This fabrication step introduces the magnetically sensitive components that will read out the stored weight value.
[0079] The process then includes arranging the first SQUID (201 ) relative to the superconducting weight storage loop (200) to establish a mutual inductive coupling exhibiting a first polarity. This involves precise control over the physical layout and relative positioning of the first SQUID (201 ) and the weight loop (200) during the fabrication sequence. Achieving a specific coupling polarity might involve placing the SQUID loop (201 ) on one side (e.g., above) of the weight loop conductor (200a) or configuring the overlap geometry in a specific way. This deliberate arrangement ensures that the magnetic flux from the weight loop (200) influences the first SQUID (201 ) in a defined directional sense, which is necessary for predictable control over its inductance and forms one half of the differential control mechanism. The precise layout provides improved control over the circuit's operational characteristics.
[0080] Complementary to the previous step, the method involves arranging the second SQUID (202) relative to the superconducting weight storage loop (200) to establish a mutual inductive coupling exhibiting a second polarity substantially opposite to the first polarity. This step completes the differential coupling structure. It requires positioning the second SQUID (202) relative to the weight loop (200) such that the induced flux linkage is opposite to that of the first SQUID (201). A common implementation places the second SQUID (202) on the opposite side (e.g., below) of the weight loop conductor (200a) compared to the first SQUID (201), often utilizing insulating layers between them in a multi-layer fabrication process. Establishing this opposite coupling polarity is fundamental to the bipolar operation, ensuring that the inductances of the two SQUIDs (201 , 202) vary inversely with the weight current, enabling differential signal modulation from a single storage element, thereby enhancing circuit compactness and functionality.
[0081] Next, the method includes forming a first inductive divider (203) comprising a first leg that incorporates the first SQUID (201) and a second leg that includes a first coupling inductor (205). This step constructs one of the two parallel signal paths where the input signal will be modulated. It involves patterning superconducting traces to create a branching structure where current can flow either through the path containing the first SQUID (201) or the path containing the first coupling inductor (205). The coupling inductor (205) itself is formed as a patterned superconducting trace designed for mutual inductance with the soma element. This divider structure provides the physical means for splitting the input current based on the variable inductance of the first SQUID (201), directly enabling the weightdependent modulation of the signal portion carried by the first coupling inductor (205).
[0082] Similarly, the process involves forming a second inductive divider (204) comprising a first leg including the second SQUID (202) and a second leg including a second coupling inductor (206). This step creates the complementary signal path. Patterning superconducting traces forms the branching structure incorporating the second SQUID (202) in one path and the second coupling inductor (206) in the other. The second coupling inductor (206) is also patterned for mutual inductance with the soma element, but potentially with a different geometry or orientation compared to the first coupling inductor (205). Forming this second divider(204) provides the pathway for modulating the signal component that will contribute oppositely to the first path, completing the differential structure required for bipolar weighting based on the single weight stored in loop (200).
[0083] The method proceeds with forming a soma element (207) comprising a soma inductor line (207a) and a soma output Josephson junction (209) connected thereto. This constructs the signal summation and thresholding stage of the artificial neuron. Forming the soma inductor line (207a) involves patterning a superconducting trace, potentially with multiple turns or segments, designed to effectively receive induced signals via mutual coupling. The soma output Josephson junction (209) is fabricated and integrated at one end of this line (207a), typically using standard junction fabrication steps involving deposition of superconductor-insulator- superconductor layers and patterning. This provides the physical structure for integrating the weighted inputs and applying the non-linear activation function necessary for neuronal computation, directly realized in SFQ hardware.
[0084] A subsequent step involves arranging the first coupling inductor (205) relative to the soma inductor line (207a) to establish a mutual inductive coupling exhibiting a third polarity. This requires careful layout design to position the superconducting trace forming the first coupling inductor (205) in proximity to the soma inductor line (207a) such that current flowing in inductor (205) induces a signal of the desired polarity (e.g., positive) in the soma line (207a). This arrangement ensures that the signal component modulated by the first divider path (203) contributes constructively (or in a defined excitatory manner) to the summed signal in the soma element (207).
[0085] Correspondingly, the method includes arranging the second coupling inductor (206) relative to the soma inductor line (207a) to establish a mutual inductive coupling exhibiting a fourth polarity substantially opposite to the third polarity. This arrangement positions the second coupling inductor (206) near the soma line (207a), potentially with a different relative orientation or winding sense compared to the first coupling inductor (205), such that current flowing through it induces a signal of the opposite polarity (e.g., negative) in the soma line (207a). Establishing this opposite coupling polarity is essential for the bipolar summation, allowing the signal from thesecond divider path (204) to effectively subtract from the signal from the first path, thereby realizing the bipolar weighting function within the soma element (207).
[0086] Finally, the fabrication method includes electrically connecting the various components. The first inductive divider (203) and the second inductive divider (204) are connected to an input terminal (210a), typically via patterned superconducting interconnects and potentially an input splitter. The first leg of the first inductive divider (203) (containing the first SQUID 201) and the first leg of the second inductive divider (204) (containing the second SQUID 202) are connected to the circuit's ground reference (GND), often a superconducting ground plane. Similarly, the second leg of the first inductive divider (203) (containing the first coupling inductor 205) and the second leg of the second inductive divider (204) (containing the second coupling inductor 206) are also connected to ground (GND), potentially through series resistors. These connections are typically formed using patterned superconducting wiring layers and vias (connections between layers). Completing these electrical connections ensures the proper flow of SFQ pulses and bias currents throughout the circuit, enabling the intended operation of the inductive dividers, SQUIDs, and soma element according to the design principles for bipolar weighting. This step integrates the separately formed components into a functional circuit.
[0087] Further enhancing the programmability of the circuit, the method may include forming a first Josephson junction (200b) and a second Josephson junction (200c) within the superconducting weight storage loop (200). These junctions are strategically positioned to enable bidirectional current injection into the loop (200). Implementation involves integrating these Josephson junction structures (superconductor-insulator-superconductor layers) at appropriate points along the patterned superconducting loop (200) during the fabrication sequence, along with necessary control lines. This addition provides the crucial benefit of allowing the persistent weight current stored in the loop (200) to be incrementally increased or decreased by applying external single flux quantum (SFQ) pulses via the respective junctions (200b, 200c). This enables direct, on-chip modification of the synaptic weight based on signals from learning circuitry, facilitating adaptive behavior and hardwarebased training protocols. Different junction types or placement configurations might bepossible alternatives. This step provides the physical mechanism for digitally updating the analog weight value.
[0088] The fabrication process may further involve forming a first resistor (205a) in series with the first coupling inductor (205) within the second leg of the first inductive divider (203), and similarly forming a second resistor (206a) in series with the second coupling inductor (206) in the second leg of the second inductive divider (204). This involves depositing and patterning a layer of resistive material (such as molybdenum, palladium, or a suitable alloy) to create discrete resistive elements (205a, 206a) electrically connected in series with their respective coupling inductors (205, 206) and to the ground connection. Functionally, these resistors establish specific L / R time constants for these signal paths, which helps control the decay rate of signals coupled into the soma and can prevent unwanted oscillations or persistent current build-up in the divider legs. This contributes to improved signal integrity and potentially makes the circuit operation less sensitive to minor variations in input pulse timing or shape. While series resistors are specified, alternative damping mechanisms could exist, but series resistors offer a straightforward implementation. These resistors refine the temporal response characteristics of the signal pathways leading to the soma.
[0089] Additionally, the method may include forming a soma resistor (207b) that connects the soma inductor line (207a) to ground (GND), and also electrically connecting the soma output Josephson junction (209) to an output Josephson transmission line (208). The soma resistor (207b) is typically formed as a patterned resistive element shunting the soma line (207a) to the ground plane. Its function is to set the integration time constant (leak rate) for the entire soma element (207), allowing it to effectively sum inputs arriving within a certain time window while preventing charge build-up over longer periods. Connecting the output junction (209) to a Josephson transmission line (208) (a standard patterned superconducting waveguide for SFQ pulses) is achieved through superconducting interconnects. This connection provides the necessary path for propagating the neuron's output signal (an SFQ pulse if generated) to subsequent network stages. Forming the resistor (207b) enhances the circuit's robustness to input timing jitter, while forming the output connection (208) enables integration into larger computational systems.
[0090] A specific technique for arranging the first SQUID (201) and the second SQUID (202) relative to the superconducting weight storage loop (200) involves positioning conductive layers associated with the first SQUID (201 ) and the second SQUID (202) on opposing sides of a conductive layer forming part of the superconducting weight storage loop (200). This is typically implemented in a multilayer fabrication process where, for example, the weight loop trace (200a) is formed in one superconducting layer, the first SQUID (201 ) is formed predominantly in a layer above it (separated by an insulator), and the second SQUID (202) is formed predominantly in a layer below it (also separated by an insulator). This vertical separation inherently causes the magnetic flux from the weight loop current to pass through the two SQUID loops in opposite directions, directly establishing the required opposite coupling polarities. This provides a manufacturable and spatially efficient method for realizing the differential inductance control mechanism.
[0091] Similarly, arranging the first coupling inductor (205) and the second coupling inductor (206) relative to the soma inductor line (207a) may involve configuring their winding directions or relative physical positions to achieve the third and fourth opposite polarities of mutual coupling. This is accomplished through careful layout design during the patterning of the superconducting layers forming these inductors (205, 206) and the soma line (207a). For instance, placing the inductor traces parallel to the soma line but offset in specific ways, or orienting them on opposite sides of the soma line, can result in the currents inducing opposing fluxes and thus opposite polarity signals in the soma line (207a). Achieving these opposite coupling polarities is essential for enabling the subtractive summation in the soma element (207) that underlies the bipolar weighting function. Precise control over this layout ensures the intended differential signal combination.
[0092] The fabrication method may also include forming an input splitter (210) that is configured to electrically connect the input terminal (210a) simultaneously to both the first inductive divider (203) and the second inductive divider (204). This splitter (210) function is typically realized by patterning superconducting traces in a T or 'Y' junction geometry, or by implementing a more complex active SFQ splitter circuit if fan-out amplification is needed. Its purpose is to duplicate the incoming SFQ pulse from the terminal (210a) and distribute it with minimal timing difference to the inputs ofthe parallel divider paths (203, 204). Forming this splitter ensures that the differential weighting mechanism operates on the identical input signal concurrently, which is necessary for the resulting net signal in the soma (207) to accurately reflect the influence of the stored weight.
[0093] Furthermore, the method can comprise forming electrical connections specifically designed for applying bias currents to the first SQUID (201 ) and the second SQUID (202). These connections involve patterning additional superconducting lines that route external DC bias currents to the appropriate points on the SQUID circuits, typically near their Josephson junctions. The purpose of these bias currents is to set the operating point of the SQUIDs such that they remain in the non-spiking (superconducting) state across the expected range of flux modulation from the weight storage loop (200). Forming these bias connections allows for external control to ensure the SQUIDs function solely as smooth, variable inductors, which is beneficial for achieving predictable analog modulation of the signal dividers based on the stored weight.
[0094] The method may also extend to forming a complete bias weight circuit (211 ), which itself includes forming a bias weight storage loop (212) and a bias weight inductive divider leg (213). Subsequently, this bias weight inductive divider leg (213) is arranged relative to the main soma inductor line (207a) to establish a mutual inductive coupling. This involves fabricating structures analogous to one of the main synaptic branches (e.g., loop 212 similar to loop 200, leg 213 similar to inductor 205 with its associated SQUID and divider structure, though potentially simplified). Positioning the coupling leg (213) appropriately near the soma line (207a) allows this separate circuit (211 ) to inject an independently programmable DC offset or bias current into the soma element (207). Forming this dedicated bias circuit (211) provides enhanced functionality by allowing the neuron's effective activation threshold to be adjusted through its own stored weight in loop (212), improving the network's learning capacity for various tasks.
[0095] Fundamentally, the process of forming the key components — the superconducting weight storage loop (200), the first SQUID (201 ), the second SQUID (202), the first coupling inductor (205), the second coupling inductor (206), the soma inductor line (207a), and the soma output Josephson junction (209) — utilizesestablished microfabrication techniques involving the deposition and patterning of superconducting thin films, insulating layers, and potentially resistive materials. This encompasses a sequence of steps such as substrate preparation, sputtering or evaporation of materials (e.g., niobium for superconductors, aluminum oxide or silicon dioxide for insulators, molybdenum or similar for resistors), application of photoresist, exposure through masks defining the circuit layout, development of the resist, etching (e.g., reactive ion etching, wet etching) to remove unwanted material, and resist stripping. Repeating these steps for multiple layers allows the construction of the complex three-dimensional structure of the circuit. Relying on these standard deposition and patterning processes makes the fabrication of the bipolar SFQ weight circuit compatible with existing superconducting integrated circuit manufacturing capabilities, providing a pathway to realizing these devices.
[0096] FIG. 1 illustrates aspects of a small-scale neural network architecture and its simulated learning performance. Panel (a) depicts a block diagram of a specific network configuration, representing a 2-2-1 feedforward network structure. This structure comprises an input layer with two input nodes receiving external signals denoted as Xu and xi,2. These inputs propagate to a hidden layer containing two units (neurons). The connections between the input layer and the hidden layer are represented by synaptic weights, specifically W211 and W212 connecting inputs xi,i and xi,2 respectively to the first hidden unit, and W221 and W222 connecting inputs xi,i and xi ,2 respectively to the second hidden unit. Each hidden unit also receives a bias input weighted by W210 and W220 respectively, and incorporates a stochastic input represented by elements d2i and d22. The weighted inputs, bias, and stochastic term are summed within each hidden unit's soma element, denoted S21 and S22. The outputs of the hidden layer units, X2,i and X2,2, then serve as inputs to the output layer, which consists of a single unit. Connections to this output unit are weighted by W311 (from hidden unit 1 ) and W312 (from hidden unit 2), and there is a bias weight W310. The summation in the output unit's soma is denoted S31 , producing the final network output ys,i . The diagram also indicates the presence of Learning Rule Logic, which receives the network output ys,i and a target output signal y*. This logic interfaces with the weight elements (w) to facilitate network training and adaptation based on the comparison between the actual and target outputs, likely implementing a reinforcement learning algorithm. The interconnectivity shown suggests a fullyconnected architecture between adjacent layers in this specific example. Implementation of this architecture is simulated using SPICE models representing the underlying bipolar single flux quantum weight circuits and associated logic.
[0097] Panel (b) presents results from a continuous SPICE simulation of the network depicted in panel (a) undergoing training for three different logical functions sequentially over a period of 1500 nanoseconds (ns). The vertical axis represents the probability of the network producing the correct output, calculated as a running average, indicating learning success. The horizontal axis represents time in nanoseconds, with each nanosecond potentially corresponding to one inference and learning cycle in the simulation. During the first phase (0 ns to 500 ns), the target output y* required the network to output a pulse whenever input xi was active (representing the function xi). The graph shows the probability correct rapidly converging to and maintaining a value of 1.0, indicating successful learning of this function. In the second phase (501 ns to 1000 ns), the target output y* was changed to require an output pulse when either input xi or input X2 was active (representing the logical OR function). The simulation continued without resetting weights, and the graph shows the network adapting, with some transient fluctuations in accuracy, before again converging to a probability correct of 1 .0, demonstrating its ability to learn the new OR function. In the third phase (1001 ns to 1500 ns), the target output y* was further changed to require an output only when both input xi and input X2 were active (representing the logical AND function). Again, the network adapted from its previous state and successfully learned the AND function, reaching and maintaining a probability correct of 1 .0. This simulation demonstrates the operability and functionality of the network architecture combined with the reinforcement learning approach, showcasing its ability to learn different functions dynamically and rapidly (within hundreds of nanoseconds per function) based solely on changes in the supervisory target signal y*. One benefit highlighted is the network's adaptability and potential for continual learning without requiring external intervention or complete retraining from random initial states. The extremely fast learning time scale (order 1 ns per cycle) is a direct consequence of the high-speed nature of the underlying simulated SFQ hardware components and the efficiency of the hardware-implemented learning rules, representing a significant performance advantage over conventional software-based neural network training. Variations could involve different network sizes, learning rules,or input patterns. This figure exemplifies the use of the bipolar SFQ weight circuits within a functional, self-training neuromorphic system.
[0098] FIG. 2 presents comparative simulation results elucidating the impact of stochastic weight excitation on the learning performance of the previously described neural network architecture, specifically when tasked with learning to identify input 2 (output y should equal input X2). The simulations depicted utilize the same network configuration as in FIG. 1 but employ a particularly challenging input sequence where input 2 alternates between 0 and 1 on consecutive cycles (010101...). This rapid alternation can obscure the reinforcement signal needed for effective weight updates in the hidden layer, potentially hindering learning convergence. Panels (a) and (b) illustrate the network's behavior under conditions of low stochastic excitation, set at approximately 1 % of the maximum possible value of a single weighted synaptic input. Panel (a) plots the probability correct (averaged over recent cycles) against simulation time in nanoseconds. The graph reveals that the network fails to consistently learn the target function; the probability correct oscillates significantly and remains far below unity, indicating poor performance and an inability to converge to the correct solution. Panel (b) complements this by showing the corresponding time evolution of the currents within the weight storage loops for each synapse in the network (weights W312, W311 , W310, W222, W221 , W220, W212, W211 , W210, offset vertically for clarity). These weight currents are observed to oscillate periodically around certain values without settling, suggesting that the learning process is trapped in a local minimum of the error landscape or is being driven cyclically by the alternating input pattern, preventing the discovery of the weights required to correctly implement the X2 function. The physical implementation represented is a SPICE simulation incorporating the dynamics of the bipolar synapses, soma elements, and learning logic with this minimal stochastic influence.
[0099] In contrast, panels (c) and (d) depict the simulation results for the identical network and input sequence, but with the magnitude of the stochastic weight excitation increased substantially to approximately 20% of the maximum weighted input value. Panel (c), plotting probability correct versus time, now shows a markedly different behavior. The network performance steadily improves over time, eventually converging to and maintaining a probability correct near 1.0, signifying successfullearning of the target function y = X2 despite the challenging input sequence. Panel (d) illustrates the corresponding evolution of the weight storage loop currents for this scenario. Unlike the oscillations observed in panel (b), the weight currents in panel (d) are seen to explore more broadly initially and then converge towards stable values. These final weight values represent a configuration of the network that successfully implements the desired function. The comparison between the two scenarios clearly demonstrates the operability and functionality of stochastic excitation as a mechanism to enhance learning. The added randomness allows the network to escape the local minima or cyclic behavior observed with low stochasticity, enabling effective exploration of the weight space and convergence to a correct solution. The implementation involves simulating the stochastic hardware element (potentially as shown in FIG. 10) providing random ±1 equivalent inputs to the hidden layer somas, scaled to the higher, 20% magnitude level.
[0100] The benefit illustrated by FIG. 2 is the significantly improved robustness and effectiveness of the on-chip reinforcement learning process achieved by incorporating an appropriate level of stochastic excitation. This technique allows the hardware network to successfully train on tasks or input sequences that might otherwise prevent convergence. It provides a practical method, realizable with superconducting hardware leveraging thermal noise, to overcome common challenges in optimizing complex systems like neural networks. Variations could involve tuning the magnitude of the stochastic excitation; the presented results suggest a range is effective, potentially from 10% to 30% or even higher, depending on the specifics of the network and task, while levels as low as 1 % may be insufficient in some cases. This figure provides a concrete example of how carefully controlled noise injection, implemented via dedicated stochastic elements within the SFQ architecture, enhances the network's learning capabilities and overall utility.
[0101] FIG. 3 provides simulation results illustrating the successful training of the 2-2-1 neural network architecture, configured with approximately 20% stochastic weight excitation in the hidden layer, to perform the exclusive OR (XOR) function. The XOR function is a classic benchmark problem in machine learning because it is not linearly separable, meaning its solution requires the computational capability provided by intermediate or hidden layers within a network, thus demonstrating the network'scapacity for non-trivial computations. The training utilized an input pattern consisting of a repeating sequence of 16 quasi-random input vectors presented over 1000 nanoseconds, corresponding to 1000 training cycles. Panel (a) plots the probability that the network's output correctly matches the target XOR output (outputting a T or spike if and only if exactly one of the two inputs is '1') against simulation time. The probability correct is shown as a running average, likely over four cycles as mentioned in related descriptions. The plot shows an initial period of low accuracy, followed by a relatively rapid increase, reaching and sustaining a probability correct near 1.0 in less than 400 nanoseconds. This trajectory clearly indicates that the simulated network, employing the bipolar SFQ weight circuits and associated reinforcement learning logic, successfully learns the XOR function. The ability to learn this non-linear function demonstrates the computational completeness of the architecture at this scale.
[0102] Panel (b) offers insight into the internal adaptation process by displaying the currents within the superconducting weight storage loops for the various synapses in the network (weights labeled wijk corresponding to the diagram in FIG. 1 a) as a function of time during the XOR training simulation. The currents, which represent the analog synaptic weights, are shown with vertical offsets for visual distinction. Initially, the weights fluctuate as the network explores different configurations driven by the learning algorithm and stochastic perturbations. However, correlating with the increasing accuracy shown in panel (a), the weight currents are observed to converge towards specific, stable values by approximately 400 ns. These final, converged current levels represent the learned synaptic strengths that collectively enable the network to implement the XOR logical operation. This panel visually confirms the operation of the weight programming mechanism, showing the dynamic adjustment of the analog weights stored within the superconducting loops in response to the reinforcement learning signals. The relatively fast convergence time underscores the efficiency of the hardware-based learning process.
[0103] Panel (c) provides a snapshot of the network's spiking behavior during the initial phase of training, specifically covering the first 50 nanoseconds when the probability correct shown in panel (a) is still low (around 0.5 or less). This plot displays the timing of spikes for Input 1 (xi, orange), Input 2 (X2, cyan), the target XOR output (y*, blue), and the actual network output (y, black), with traces offset verticallyfor clarity. By comparing the actual output trace (y) with the target output trace (y*) for the corresponding input patterns (xi, X2), numerous discrepancies can be observed. For example, the network might fail to spike when it should (false negative) or spike when it should not (false positive), reflecting its initial inability to correctly compute the XOR function. This visual representation confirms the low accuracy state before significant learning has occurred.
[0104] Panel (d) presents the analogous spiking patterns but for a later time window, from 950 ns to 1000 ns, after the network has successfully learned the XOR function, corresponding to the high probability correct plateau in panel (a). In this plot, the actual network output trace (y, now shown in green) exhibits a near-perfect match with the target output trace (y*, blue) across the various input combinations presented. Specifically, an output spike (y=1) occurs when either Input 1 or Input 2 is active ('1 ') but not both, and no output spike (y=0) occurs when both inputs are active or neither is active. This visual evidence confirms that the network, through the adaptation of its internal weights shown in panel (b), has converged to a state where it accurately implements the XOR function. Comparing panel (d) with panel (c) highlights the functional transformation achieved through the simulated learning process. Overall, FIG. 3 provides compelling evidence for the functionality and learning capability of the simulated bipolar SFQ circuit architecture, demonstrating its ability to solve non- linearly separable problems rapidly and accurately through hardware-based reinforcement learning, showcasing its utility for implementing advanced computational tasks.
[0105] FIG. 4 presents a graphical analysis of the projected time scaling characteristics of the proposed superconducting neuromorphic architecture, illustrating the relationship between network size and operational speed. The vertical axis represents the total time required for one complete cycle, encompassing both an inference pass and a subsequent learning plus weight update phase, measured in nanoseconds (ns). The horizontal axis represents the depth of the neural network, quantified as the number of layers. Four distinct data series are plotted on the graph, each corresponding to a different network width, defined as the number of processing units (neurons) per layer: 1 unit wide (gray squares), 200 units wide (green downward triangles), 10,000 units wide (red circles), and 100,000 units wide (blue upwardtriangles). This figure, therefore, provides insight into how the architecture's processing time is expected to scale as networks become both deeper and wider, based on performance parameters derived from detailed SPICE simulations of the constituent single flux quantum (SFQ) circuits, including the bipolar synapses and learning logic. The plots demonstrate the operability and projected performance of the architecture across various scales. A primary observation is the dependence of the cycle time on the number of layers. For all network widths shown, the time per full cycle increases as the layer depth increases. This relationship appears to be approximately linear, particularly for deeper networks (e.g., beyond 10-20 layers). This linear scaling with depth is largely attributable to the inference latency, as signals must propagate sequentially through each layer during the forward pass. The slope of this linear increase represents the effective delay added per layer, which includes soma processing time, fanout delays for signal distribution within and between layers, and potentially wiring delays.
[0106] FIG. 4 also reveals the impact of network width on the cycle time. While wider networks inherently involve more computations and signal distribution points, leading to longer absolute cycle times compared to narrower networks of the same depth, the increase in cycle time with width appears to be significantly sub-linear. For instance, increasing the width by orders of magnitude (from 200 to 100,000) results in only a modest increase in the total cycle time for a given depth. This favorable scaling with width is attributed in the accompanying text to the efficient fanout characteristics of SFQ circuits, where signal distribution time may scale logarithmically with the number of receiving units, and to the parallel and asynchronous nature of the learning updates, which do not necessarily incur delays proportional to the number of units.
[0107] A significant benefit highlighted by FIG. 4 is the remarkably fast overall cycle times projected even for very large and deep networks. For a network with 100 layers and 100,000 units per layer, the total time for a full inference and learning cycle is projected to be less than 20 nanoseconds according to the plotted data points. This suggests operational frequencies in the range of 50 million cycles per second or higher, encompassing both forward computation and weight adaptation. This level of performance represents a potential speedup of many orders of magnitudecompared to conventional software-based training on GPUs or CPUs. This enhanced performance stems directly from the ultra-high speed of SFQ logic elements, the rapid propagation of signals on superconducting lines, and the efficient hardware implementation of the parallel learning rules. Variations in actual performance would depend on specific circuit optimizations, fabrication technology parameters, and detailed physical layout including wiring lengths, but the trends shown indicate excellent scalability potential. This figure exemplifies the projected utility of the superconducting neuromorphic architecture for tackling large-scale machine learning problems demanding high throughput and rapid training or adaptation.
[0108] FIG. 5 presents a schematic block diagram outlining the structure and operation of the logic circuitry designed to compute and apply weight updates for a synaptic connection located in the output layer of the neural network. This logic implements a specific learning rule, transforming signals representing the network's performance into commands that adjust the physical weight stored, likely as a current, in the associated bipolar synapse's weight storage loop. The diagram receives several input signals: a 'Learn Clock' signal which initiates the update process, the actual output 'y' produced by the neuron during the preceding inference phase, the target or desired output 'y*' for that phase, and the input 'x' that was delivered to the specific synapse being updated during inference. The logic circuit processes these inputs through a sequence of standard single flux quantum (SFQ) logic gates, including XOR (exclusive OR) gates and DRO (Destructive Readout) gates, along with delay elements, ultimately producing output signals directed to the increment ('inc') and decrement ('dec') inputs of the synapse's weight storage loop.
[0109] The operational flow begins upon arrival of the Learn Clock signal, potentially passing through an initial 'Long delay' block which may serve to synchronize the learning cycle initiation relative to the completion of the inference phase and stabilization of the y and y* signals. The core logic is conceptually divided into determining if an update is necessary ('Update not zero') and, if so, determining the direction of the update ('Update direction'). The 'Update not zero' stage employs an XOR gate receiving the actual output 'y' and the target output 'y*' as inputs. These inputs 'y' and 'y*' likely represent stored binary values (spike or no spike) from the inference cycle. The XOR gate outputs a high signal (representing a '1 ' or an SFQpulse) if and only if y and y* are different, signifying an error occurred. This output signal serves as a conditional clock ('elk') for the subsequent logic stages. If y equals y*, the XOR output remains low, no clock pulse is generated, and the weight update process is effectively gated off for this cycle, providing a hardware-efficient implementation of the principle that no weight change is needed when the output is correct. A 'reset' signal path is also indicated, suggesting a mechanism to clear stored states after the update calculation is complete or aborted.
[0110] If an update is warranted (y != y*), the generated clock signal ('elk') enables the 'Update direction' logic. This section determines whether the weight should be increased or decreased to reduce the observed error. It utilizes DRO gates to access the stored value of the synapse input 'x' and likely implicitly accesses stored y or y* information via the paths shown leading into the second XOR gate. Destructive Readout implies these values are retrieved from temporary storage elements (e.g., SFQ flip-flops) and potentially cleared in the process. A 'delay' element is shown processing the output from the initial DRO handling input x, likely ensuring proper timing alignment of signals arriving at the final XOR gate. The final XOR gate computes the direction based on the relationship between the synapse input 'x' and the error condition (represented implicitly by the enabled clock and potentially the y / y* state derived through the DROs). The specific logic implemented by this final XOR, combined with the inputs it receives, determines the sign of the required weight adjustment according to the chosen learning rule. For instance, the configuration might be designed such that if the XOR output 'Q' is high, a decrement signal ('Q dec') is sent to the weight storage loop, and if the complementary output 'Q_bar' is high, an increment signal ('Q_bar inc') is sent.
[0111] The implementation of this logic relies on standard SFQ gate designs, optimized for high-speed, low-power operation. A notable benefit of this specific logical arrangement is its efficiency; by using the error signal (y != y*) from the first XOR gate as a clock for the rest of the logic, the circuit avoids unnecessary computation and signal propagation when no weight update is required. This clock gating approach reduces the average switching activity and power consumption. The use of DRO gates interfaces the logic with the necessary memory elements holding the state from the inference cycle. The entire circuit is designed to execute the weightupdate calculation extremely rapidly, consistent with the multi-gigahertz clock speeds achievable with SFQ technology, enabling very fast on-chip training cycles. Variations could involve alternative SFQ gate implementations for the XOR or DRO functions, different delay values for timing optimization, or modifications to the exact logic to implement slightly different learning rule variants. This logic block exemplifies how specific learning algorithms can be translated into efficient superconducting hardware for controlling the bipolar SFQ synaptic weights.
[0112] FIG. 6 illustrates both the conceptual circuit design and the simulated operational behavior of the bipolar single flux quantum (SFQ) synaptic weight circuit. Panel (a) provides a schematic diagram outlining the core components and their interconnections. An input signal, represented as an SFQ pulse 'x', arrives at an input splitter (210), which duplicates the signal and directs it simultaneously to two parallel processing pathways: a 'positive branch' highlighted in a blue box and a 'negative branch' highlighted in an orange box. Each branch functions as an inductive divider structure. The positive branch incorporates circuit elements corresponding to the first inductive divider (203), including a variable inductance element likely representing the first SQUID (201) connected in one leg, and a coupling inductor corresponding to the first coupling inductor (205) in the other leg, potentially with a series resistor (like 205a). Similarly, the negative branch incorporates elements corresponding to the second inductive divider (204), including a variable inductance element representing the second SQUID (202) and a coupling inductor representing the second coupling inductor (206), again potentially with a series resistor (like 206a). Centrally located between these branches is the implied superconducting weight storage loop (200), depicted with light blue wiring in related source material though not explicitly boxed here, which contains the stored weight current. This weight loop (200) is shown receiving increment ('Inc') and decrement ('Dec') SFQ pulse inputs from Reinforcement Learning (RL) logic circuitry (214, implied input). The current in this loop (200) mutually couples to and controls the inductance of the variable inductance elements (SQUIDs 201 , 202) within the positive and negative branches, respectively, with opposite polarities. The coupling inductors (205, 206) from both the positive and negative branches are, in turn, mutually coupled to a common 'Soma' element (207), highlighted in a green box. The coupling polarities are arranged such that the positive branch induces a signal of one polarity (e.g., positive) in the soma (207), while thenegative branch induces a signal of the opposite polarity (e.g., negative). The soma element (207) integrates these induced signals along its soma inductor line (207a, implied structure within the box) and produces an output 'y' via its soma output Josephson junction (209, implied output stage). The implementation relies on standard SFQ circuit elements like Josephson junctions (represented by 'x' symbols), inductors (coils), resistors (zig-zag symbols), and current biases (circles with arrows), all fabricated using superconducting materials. The benefit of this structure is its ability to realize a bipolar synaptic weight function using a single weight storage loop (200) by differentially modulating two parallel paths and combining their outputs subtractively.
[0113] Panel (b) presents the results of a SPICE (Simulation Program with Integrated Circuit Emphasis) simulation, graphically demonstrating the circuit's dynamic behavior over time in response to input pulses and weight adjustments. Four traces are shown: the current in the weight storage loop (200) (top trace, gray), the current induced in the soma (207) by the negative branch coupling inductor (206) (second trace, orange), the current induced in the soma (207) by the positive branch coupling inductor (205) (third trace, blue), and the net current within the soma element (207) (bottom trace, green). Initially, at time t=0, the storage loop current is zero. An input pulse 'x' arrives, causing transient currents of equal magnitude but opposite polarity to be induced by the positive (blue) and negative (orange) branches, resulting in a net soma current (green) of approximately zero, correctly representing a zero weight. Subsequently, between approximately 1 ns and 2 ns, a series of 32 SFQ decrement pulses are applied to the weight storage loop (200), causing its current to decrease to a minimum negative saturation value (around -75 microamperes for the simulated 250 pH inductor). When a second input pulse 'x' arrives around 3 ns, the negative branch (orange) now induces a significantly larger magnitude negative signal in the soma (207), while the positive branch (blue) induces a smaller positive signal, due to the altered SQUID inductances (201 , 202). This results in a distinct net negative spike in the soma current (green), demonstrating negative weighting. Following this, between approximately 4 ns and 6 ns, 64 SFQ increment pulses are applied, driving the storage loop current through zero to its maximum positive saturation value (around +75 microamperes). An input pulse 'x' applied around 7 ns now results in a large positive signal induced by the positive branch (blue) and a smaller negative signal induced by the negative branch (orange), yielding a net positive spike in the somacurrent (green), demonstrating positive weighting. Finally, 30 decrement pulses applied between 8 ns and 10 ns return the storage loop current to near zero. A final input pulse 'x' around 11 ns again produces cancelling contributions from the positive and negative branches, resulting in near-zero net soma current. This simulation powerfully illustrates the circuit's operability, confirming its ability to represent and dynamically program both positive and negative synaptic weights using SFQ pulses to adjust the analog current in the storage loop (200), and validating the design concept for achieving bipolar functionality from unipolar SFQ components. The results highlight the relationship between the weight current (ranging effectively from -30 to +30 SFQ quanta for this implementation) and the resultant weighted output polarity and magnitude, showcasing the circuit's utility as a programmable synapse. Variations in the storage inductor size would alter the effective bit depth or resolution of the weight. The near-perfect cancellation at zero weight shown in simulation provides a desirable baseline, though functional operation may tolerate some offset.
[0114] FIG. 7 provides a block diagram illustrating the logical structure and signal flow for implementing the weight update mechanism specifically tailored for synapses residing in the output layer of the neural network. This diagram, labeled "Output Layer update," details how the circuit processes key signals derived from the network's performance to generate increment or decrement commands for adjusting the corresponding synaptic weight, which is physically stored, for instance, as a persistent current within the bipolar synapse's superconducting weight storage loop. The inputs to this logic module are clearly indicated: a 'Learn Clock' pulse that triggers the learning cycle computation, the actual output 'y' (spike or no spike, representing 1 or 0) generated by the output neuron during the preceding inference step, the corresponding target output 'y*' provided externally, and the specific input 'x' (spike or no spike) that was presented to this particular synapse during that inference step. The diagram shows these inputs feeding into a network of interconnected functional blocks, primarily standard logic gates such as XOR and DRO (Destructive Readout) gates, interspersed with 'delay' elements, culminating in outputs labeled 'Q (dec)1and 'Q_bar (inc)' directed towards the weight storage loop adjustment mechanism. An associated equation w += (y* - y)(x - 0.5) summarizes the intended learning rule implemented by this logic, although the circuit implements a ternary version (+1 , -1 , 0) of this concept suitable for SFQ hardware. The functionality depicted begins with the arrival of theLearn Clock signal, which may pass through an initial delay element for timing purposes. The core operation proceeds in stages. First, the actual output 'y' and the target output 'y*' are fed into an XOR gate. This gate acts as an error detector; its output signal (labeled '(y* - y != 0)' which functionally represents the condition y != y*) becomes active (e.g., generates an SFQ pulse) only if the network's output differs from the desired target. This error signal is crucial as it serves as a conditional gate or clock for the subsequent stages. If y equals y*, indicating correct performance for this cycle, the error signal remains inactive, and no further processing occurs, thus preventing unnecessary weight adjustments. This clock gating based on the error signal provides a significant benefit in terms of computational efficiency and reduced power consumption, as updates are only computed when necessary. A 'reset' line connected after a delay suggests that stored input values might be cleared after the update calculation or if no update occurs.
[0115] If an error is detected (y != y*), the error signal enables the remainder of the logic circuit. Destructive Readout (DRO) gates are employed to retrieve the stored binary values of the synapse input 'x' and potentially the output 'y' and target 'y*' states from memory elements where they were latched during the inference phase. The use of DRO implies that reading the value might also reset the memory element. The retrieved input signal 'x' (after passing through a DRO) and a signal derived from 'y' and 'y*' (after passing through their DROs and potentially the first XOR gate) are then processed by a second XOR gate, potentially after passing through appropriately timed delay elements to ensure signal arrival synchronization. This final XOR gate, along with the specific connections feeding it, implements the logic to determine the direction of the weight update required to correct the error, based on the relationship between the input 'x' and the nature of the error (e.g., needed a spike but didn't get one, or vice-versa). The output 'Q' of this XOR gate directly drives the decrement ('dec') input of the weight storage loop, while its complementary output 'Q_bar' drives the increment ('inc') input. Therefore, depending on the state of 'x' and the error condition, either an increment pulse or a decrement pulse (but not both) is sent to adjust the stored weight current.
[0116] The implementation of this logic utilizes standard, well-characterized SFQ logic gate families and relies on associated SFQ memory cells (like flip-flops,implied by the DRO usage) to store the necessary state variables (x, y, y*) between the inference and learning cycles. The entire logic structure is designed for compatibility with the high-speed operation of SFQ circuits, allowing the weight update calculation to be performed extremely rapidly, potentially within a few hundred picoseconds, following the inference phase. This rapid update capability is essential for achieving the fast overall training times projected for the architecture. Variations on this specific logic diagram could involve different arrangements of SFQ gates to achieve the same logical function, alternative memory element interfaces, or adjustments to delay elements based on detailed timing analysis of a specific physical implementation. This diagram provides a concrete example of how a correlation-based learning rule suitable for reinforcement learning in the output layer can be efficiently translated into a functional superconducting hardware circuit module controlling the bipolar SFQ synapse.
[0117] FIG. 8 presents an alternative block diagram representation detailing the logic structure for implementing the weight update mechanism for synapses within the output layer of the neural network. Similar to FIG. 7, this diagram illustrates how signals related to network performance are processed to generate adjustment commands for the synaptic weight stored in the corresponding weight storage loop. The primary inputs remain the 'Learn Clock' signal, the neuron's actual output 'y' from the inference phase, the target output 'y*', and the synapse's specific input 'x' during inference. The outputs are again the signals driving the increment ('inc') and decrement ('dec') adjustments of the weight storage loop. However, this specific embodiment explicitly incorporates D-Flip-Flop (DFF) logic elements within the processing pathway, suggesting a different approach to signal timing, synchronization, or state storage compared to the DRO-based logic depicted in FIG. 7.
[0118] The operational sequence commences with the Learn Clock initiating the update cycle, potentially passing through an initial delay element. The actual output 'y' and the target output 'y*' are compared using a first XOR gate. The output of this XOR gate, representing the error condition (active if y != y*), is then fed into the data (D) input of a first DFF. This DFF likely captures the error state or computes a related value upon the arrival of a clock edge (implicitly triggered by the Learn Clock or a derived signal). The complementary output (Q_bar) of this first DFF is notablylabeled 'r = 1', suggesting it might represent or directly generate the reinforcement signal T', indicating a positive reward or need for update when an error occurs (since Q_bar would be high if the XOR output, representing error, was high and latched). Meanwhile, the direct error condition signal (y* - y != 0) appears to propagate through a delay element, potentially serving as an enable or gating signal for the subsequent logic, ensuring that the update direction is only computed if an error actually occurred.
[0119] The logic then proceeds to determine the update direction. The synapse input signal 'x' and the target output 'y*' (or potentially the reinforcement signal 'r' derived from the first DFF) are fed into a second XOR gate. The output of this second XOR gate represents the calculated update direction based on the relationship between the input and the desired outcome or error signal. This direction signal is then fed into the data (D) input of a second DFF. This second DFF serves to latch or synchronize the computed direction signal before it is applied to the weight storage loop. The Q output of this second DFF drives the decrement ('Q dec') signal, while the complementary Q_bar output drives the increment ('Q_bar inc') signal. Therefore, the latched state of this second DFF determines whether an SFQ pulse is sent to increase or decrease the stored weight current.
[0120] The implementation of this logic circuit utilizes standard SFQ components, including XOR gates, delay elements, and specifically DFFs. SFQ DFFs are fundamental building blocks for sequential logic and memory in superconducting circuits, capable of storing a binary state between clock cycles. Incorporating DFFs explicitly in the diagram suggests a registered logic design, which can offer benefits in managing timing and ensuring reliable state transitions in high-speed circuits. One advantage of using DFFs might be improved robustness against timing variations or hazards compared to purely combinatorial logic or asynchronous approaches relying solely on delays and DROs. It also explicitly shows how intermediate results, like the reinforcement signal 'r', can be generated and held stable for use in subsequent calculations within the same learning cycle. Variations could involve using different types of SFQ flip-flops, altering the clocking scheme, or modifying the combinational logic feeding the DFFs. This diagram provides another concrete example of how output layer learning rules can be mapped onto functional SFQ hardware, emphasizing the use of sequential logic elements for state management andsynchronization during the weight update process. The logic ensures that weight adjustments occur based on a comparison of actual versus target performance correlated with the synapse's input activity.
[0121] FIG. 9 presents a comparative view of the block diagrams representing the distinct weight update logic circuits designed for synapses residing in the output layer versus those in the hidden layers of the neural network architecture. This side-by-side presentation highlights the necessary differences in computational structure arising from the distinct roles and available information at different stages within the network during the reinforcement learning process. Both diagrams depict logic modules receiving inputs such as a 'Learn Clock', the neuron's own output 'y', and the specific synapse's input 'x', ultimately producing increment ('inc') and decrement ('dec') signals to adjust the corresponding weight storage loop. However, the internal processing logic and additional required inputs differ significantly between the two layers.
[0122] The left side of FIG. 9 reiterates the logic for an "Output Layer" synapse, consistent with the embodiment shown in FIG. 8. It takes the actual output 'y', the target output 'y*', and the synapse input 'x' as primary data inputs, initiated by the Learn Clock. The core function involves first detecting an error by comparing 'y' and 'y*' using an XOR gate. The result of this comparison, potentially latched by a D- Flip-Flop (DFF), determines if an update is needed and may contribute to generating a reinforcement signal (T = T is indicated as the complement of the latched error signal). The error signal also gates the subsequent calculation. The direction of the update is then determined by combining the synapse input 'x' with the target 'y*' (or the reinforcement signal 'r') using another XOR gate, whose output is latched by a second DFF before driving the 'inc' / 'dec' lines of the weight storage loop. This output layer logic directly utilizes the known target 'y*' to calculate the error and determine the appropriate weight adjustment based on correlating the input 'x' with this error. The implementation uses standard single flux quantum (SFQ) gates like XORs and DFFs, along with delay elements.
[0123] The right side of FIG. 9 illustrates the more complex logic required for updating weights of synapses connected to a "Hidden Layer" neuron. Crucially, hidden layers do not have an explicit target output corresponding to 'y*'. Instead, theirweight updates rely on correlating local activity with a global reinforcement signal 'r' (which is typically derived from the performance error calculated at the output layer) and changes in that signal over time. The inputs shown are the Learn Clock, the hidden neuron's output 'y', the synapse input 'x', and the global reinforcement signal T. The logic explicitly calculates the change in the neuron's output ('Ay != O') by comparing the current output 'y' with its value from the previous cycle (requiring memory, not explicitly shown but implied), likely using an XOR gate. Similarly, it calculates the change in the reinforcement signal ('Ar != O') by comparing the current 'r' with the previous 'r' (again implying memory), using another XOR. The decision of whether to perform an update depends on a combination of factors. An AND gate combines a signal representing the global error condition (labeled '(y* - y != 0)', likely referring to the output layer error that generated T) with the change in reinforcement signal 'Ar != O'. The output of this AND gate seems to act as a primary enabling condition for the update. The direction of the update is determined by correlating the synapse input 'x' with the change in the hidden neuron's local output 'Ay != O' using a final XOR gate. The output of this XOR is latched by a DFF, which then generates the ' i nc' / 'dec' signals for the hidden layer synapse's weight storage loop.
[0124] Comparing the two diagrams highlights the increased complexity needed for hidden layer updates. The hidden layer logic must infer the correct weight adjustment direction indirectly by correlating local changes (Ay) with changes in the global performance indicator (Ar), a process central to correlative reinforcement learning algorithms like REINFORCE. This requires additional logic for calculating these change signals (Ay, Ar) and combining them appropriately with the global error context and local input 'x'. The implementation necessitates more logic gates (including XORs and an AND gate) and memory elements (implied for storing previous 'y' and T values) compared to the output layer logic. The benefit of this more intricate hidden layer logic is that it enables the training of multi-layer networks by providing a mechanism for credit assignment (adjusting weights in earlier layers based on their contribution to the final output error, as reflected in the reinforcement signal) without requiring the direct backpropagation of error gradients. This figure clearly illustrates the distinct hardware structures needed to implement layer-specific learning rules within a unified reinforcement learning framework realized using SFQ circuits, showcasing the adaptability of the architecture to different network depths.
[0125] FIG. 10 presents a circuit schematic, derived from a SPICE simulation environment, illustrating the design of the stochastic hardware element (216) referenced in the network architecture (e.g., as d2i , d22 in FIG. 1 a). The physical structure depicted shares significant similarities with the bipolar synaptic weight circuit shown in FIG. 6a, employing superconducting circuit components such as Josephson junctions (represented by 'X' symbols within circles, often with associated bias currents indicated by yellow circles with arrows), inductors (coiled symbols, labeled with inductance values like 'LX'), and resistors (zig-zag symbols, labeled with resistance values like ' RX'). However, key design modifications differentiate this circuit to enable its function as a generator of stochastic signals rather than a programmable analog weight. Specifically, as indicated in the source text description associated with this figure, the central superconducting storage loop incorporates a deliberately small inductance ('LX' value likely corresponds to a small physical size). This small inductance limits the amount of magnetic flux the loop can stably store to approximately one single flux quantum (<$>o) in either the clockwise or counterclockwise direction. Furthermore, the Josephson junctions positioned within or flanking this storage loop (analogous to 200b, 200c) are potentially designed with low critical currents (Ic). This low critical current makes these junctions highly susceptible to being triggered by the ambient thermal energy present at the cryogenic operating temperature, typically around 4 Kelvin for niobium-based circuits. The circuit includes inputs labeled 'Inc' and 'Dec', analogous to the weight update inputs of the synapse, though their function here might relate to resetting or interacting with the stochastic state. Outputs labeled 'Sto+' and 'Sto-' suggest connections designed to deliver the positive and negative components of the stochastic signal to the soma element (207) of the neuron it is associated with, likely via coupling inductors similar to those in the bipolar synapse (205, 206).
[0126] Functionally, this circuit (216) is designed to generate random fluctuations that effectively act as a synaptic input with a weight randomly taking on values approximating -1 , 0, or +1. The operability relies on the thermally induced, random switching of the low-critical-current Josephson junctions associated with the small-inductance storage loop. Thermal energy fluctuations at 4K can provide enough energy to momentarily exceed the low critical current of these sensitive junctions, causing them to fire an SFQ pulse spontaneously and unpredictably. Such a firingevent injects a single flux quantum into the storage loop, setting its state to +1<$>o (e.g. , clockwise current) or -1 <t>o (e.g., counter-clockwise current), depending on which junction fires. The small loop inductance prevents the storage of multiple flux quanta, enforcing the ternary state (-1 , 0, +1 ). These asynchronous, thermally driven events continuously alter the flux state stored in the loop. The term "storing the asynchronous thermally driven events in a 'stochastic synapse'" implies that the loop holds the most recent randomly set state (+1 , 0, or -1 flux quantum). To utilize this random state synchronously within the neural network's operation cycle, an input pulse (likely triggered concurrently with the main synaptic inputs 'x' during the inference phase) is applied to the input stage of this stochastic element (216). This input pulse then interacts with the stored state (-1 , 0, or +1 ) via the divider and coupling structures (analogous to the bipolar synapse's paths), generating a corresponding output signal (+, 0, or -) coupled into the neuron's soma (207) via the 'Sto+' and 'Sto-' outputs. This allows the inherently asynchronous thermal noise to be sampled and injected as a synchronous stochastic signal perturbation during each inference cycle.
[0127] The implementation leverages standard superconducting fabrication techniques but tailors component parameters (small loop inductance, low junction Ic) to enhance sensitivity to thermal noise at the operating temperature. This provides a significant benefit by harnessing a naturally occurring physical phenomenon — thermal fluctuations — as a source of true randomness, avoiding the need for complex and potentially power-hungry pseudo-random number generator circuits. The resulting stochastic signal, taking values of -1 , 0, or +1 , provides the controlled noise injection needed for the hidden layer reinforcement learning algorithm (as discussed in relation to FIG. 2 and FIG. 9) to effectively explore the weight space and escape local minima during training. This enhances the overall functionality and reliability of the learning process. Variations could involve adjusting the junction critical currents or loop inductance to modify the probability distribution or frequency of the random fluctuations. Alternative methods for generating stochasticity in SFQ circuits might exist, but this approach offers an elegant and potentially very low-power solution integrated directly within a synapse-like structure. This circuit (216) exemplifies a specialized hardware component designed to improve the performance and robustness of on-chip learning in the superconducting neuromorphic architecture.
[0128] FIG. 11 presents a comprehensive circuit schematic, rendered from a SPICE simulation environment, depicting the detailed implementation of a complete 5-node neural network, consistent with the 2-input, 2-hidden-unit, 1 -output-unit (2-2- 1 ) architecture discussed in relation to FIG. 1 a. This schematic provides a low-level view, illustrating the intricate network of individual circuit components (Josephson junctions, inductors, resistors, and current sources) that constitute the functional blocks of the neuromorphic system, including the bipolar synapses, soma elements, stochastic inputs, and the associated reinforcement learning logic circuitry. The visual complexity of this diagram underscores the component count and interconnectivity involved in realizing even a small-scale, fully functional, self-training superconducting neural network. Distinct groupings of components can be discerned, corresponding roughly to the functional units shown in block diagrams. Input signals arrive at the left, propagating towards structures representing the synapses connecting to the hidden layer. These feed into larger blocks likely representing the hidden layer soma elements. Associated with each hidden soma can be a stochastic input circuit (Sto), consistent with the design shown in FIG. 10. The outputs from the hidden layer somas then connect via another set of synaptic circuits to the final output soma element, which generates the network's ultimate output signal. Interspersed throughout are additional circuit blocks representing the learning rule logic (similar to FIG. 9) that calculate and deliver weight update pulses (increment / decrement signals) back to the respective synaptic weight storage loops embedded within the synapse circuits.
[0129] The operability demonstrated by this detailed schematic lies in its use as the basis for the accurate SPICE simulations whose results are presented in FIGS. 1 , 2, 3, and 6b. This level of detailed modeling, capturing the non-linear dynamics of Josephson junctions and the interactions within the complex network of superconducting loops and resistors, allows for realistic prediction of the circuit's electrical behavior, including the propagation of single flux quantum (SFQ) pulses during inference, the analog summation within the soma elements, the thresholding action of output junctions, and the precise timing and effect of weight update pulses generated by the learning logic. The schematic implicitly defines the physical structure and interconnectivity required for fabrication, representing a blueprint for translating the conceptual design into a physical superconducting integrated circuit.
[0130] A significant benefit provided by visualizing this full network schematic is the appreciation of the scale and integration level achieved. It illustrates how the novel bipolar synaptic circuit (detailed in FIG. 6a) integrates seamlessly with the soma elements and the distinct learning logic modules (output layer vs. hidden layer, FIG. 9) and stochastic elements (FIG. 10) to form a cohesive, self-contained computational system capable of both inference and on-chip learning. The density of components highlights the importance of efficient circuit designs, such as the bipolar synapse which avoids pathway duplication, for managing complexity and area requirements in larger networks. While this represents a small 5-node example, containing perhaps close to 1000 Josephson, it serves as a concrete instance demonstrating the feasibility of constructing such integrated neuromorphic circuits. Variations would naturally involve scaling this architecture to larger numbers of inputs, hidden units, layers, and outputs, requiring careful layout and potentially hierarchical design strategies. The successful simulation of this detailed model provides strong support for the projected high-speed operation and learning capabilities of the architecture.
[0131] Existing superconducting circuit approaches often struggle to efficiently represent synaptic connections capable of both positive (excitatory) and negative (inhibitory) weighting, a crucial feature for many complex neural network computations, given the inherently unipolar nature of standard single flux quantum (SFQ) pulses. Some techniques might involve duplicating entire synaptic circuits to handle positive and negative weights separately, leading to considerable overhead in terms of device count, circuit area, and power consumption. Other methods might predefine connections as either excitatory or inhibitory, limiting the network's flexibility and learning capacity. While SQUIDs have been explored as variable elements in SFQ circuits, including synaptic contexts, the specific configuration enabling programmable bipolar output from a single weight storage element often remains cumbersome or inefficient in those schemes. Furthermore, integrating efficient on-chip learning mechanisms that can directly adjust these weights based on network performance presents additional challenges within the SFQ paradigm.
[0132] The described bipolar single flux quantum weight circuit departs from these techniques through a distinct structural and operational arrangement. Central tothis approach is the utilization of a single superconducting weight storage loop (200), containing a persistent current representing the synaptic weight. This single analog value simultaneously controls two separate signal pathways via a differential mechanism. Specifically, a first SQUID (201) and a second SQUID (202) are mutually coupled to this weight storage loop (200), but with precisely opposite coupling polarities. This deliberate opposite coupling ensures that as the stored weight current varies, the effective inductances of the first SQUID (201 ) and the second SQUID (202) change inversely.
[0133] These dynamically varying SQUID inductances form key components within two parallel inductive dividers (203, 204). The first divider (203) splits an incoming unipolar SFQ pulse between a path through the first SQUID (201 ) and a path through a first coupling inductor (205). The second divider (204) simultaneously splits the same input pulse between a path through the second SQUID (202) and a path through a second coupling inductor (206). Because the SQUID inductances vary differentially, the proportion of the input signal directed through the first coupling inductor (205) changes oppositely to the proportion directed through the second coupling inductor (206) in response to adjustments in the weight loop current.
[0134] The final step in achieving the bipolar output involves how these modulated signals are combined. Both the first coupling inductor (205) and the second coupling inductor (206) are mutually coupled to a common soma inductor line (207a) within a soma element (207). Critically, this coupling is also arranged with opposite polarities (the signal induced by the first inductor (205) has one polarity (e.g., positive), while the signal induced by the second inductor (206) has the opposite polarity (e.g., negative)). Consequently, the soma element (207) performs a subtractive summation of the signals from the two differentially modulated pathways. The resulting net signal in the soma inductor line (207a), which can be positive, negative, or near zero depending on the stored weight current, is then thresholded by a soma output Josephson junction (209).
[0135] This specific combination (a single weight storage loop (200), differential control of two SQUIDs (201 , 202) via opposite polarity coupling, dual inductive dividers (203, 204) modulated by these SQUIDs, and differential summation into the soma element (207) via opposite polarity coupling inductors (205, 206))constitutes a distinguishing technical characteristic. It provides a compact and efficient means to translate a unipolar SFQ input signal and a single stored analog weight value into an output that reflects programmable bipolar weighting, overcoming the limitations associated with circuit duplication or inflexibility in other approaches. The seamless integration with SFQ-based weight update mechanisms (using junctions 200b, 200c) and neuronal thresholding (junction 209) further distinguishes this system-level solution.
[0136] It is contemplated that the bipolar circuitry can include the properties, functionality, hardware, and process steps described herein and embodied in any of the following non-exhaustive list: a process (e.g., a computer-implemented method including various steps; or a method carried out by a computer including various steps); an apparatus, device, or system (e.g., a data processing apparatus, device, or system including means for carrying out such various steps of the process; a data processing apparatus, device, or system including means for carrying out various steps; a data processing apparatus, device, or system including a processor adapted to or configured to perform such various steps of the process); a computer program product (e.g., a computer program product including instructions which, when the program is executed by a computer, cause the computer to carry out such various steps of the process; a computer program product including instructions which, when the program is executed by a computer, cause the computer to carry out various steps); computer-readable storage medium or data carrier (e.g., a computer-readable storage medium including instructions which, when executed by a computer, cause the computer to carry out such various steps of the process; a computer- readable storage medium including instructions which, when executed by a computer, cause the computer to carry out various steps; a computer-readable data carrier having stored thereon the computer program product; a data carrier signal carrying the computer program product);a computer program product including comprising instructions which, when the program is executed by a first computer, cause the first computer to encode data by performing certain steps and to transmit the encoded data to a second computer; or a computer program product including instructions which, when the program is executed by a second computer, cause the second computer to receive encoded data from a first computerand decode the received data by performing certain steps.
[0137] The articles and processes herein are illustrated further by the following Example, which is non-limiting.EXAMPLE
[0138] A self-training spiking superconducting neuromorphic architecture
[0139] Neuromorphic computing takes biological inspiration to the device level aiming to improve computational efficiency and capabilities. One of the major issues that arises is the training of neuromorphic hardware systems. Typically training algorithms require global information and are thus inefficient to implement directly in hardware. In this paper we describe a set of reinforcement learning based, local weight update rules and their implementation in superconducting hardware. Using SPICE circuit simulations, we implement a small-scale neural network with a learning time of order one nanosecond per update. This network can be trained to learn new functions simply by changing the target output for a given set of inputs, without the need for any external adjustments to the network. Further, this architecture does not require programing explicit weight values in the network, alleviating a critical challenge with analog hardware implementations of neural networks.
[0140] Neuromorphic computing hardware draws inspiration from the way the human brain processes information to improve computational efficiency and broaden the potential methods of computation beyond typical digital logic. There are many reasons to look to the brain for computational inspiration. Using roughly 20 Watts of power, the human brain is an extremely energy efficient computational system. In addition, with its highly parallel architecture, the brain is both fault tolerant and can perform certain tasks with surprising speed given the typical neuron firing rate of a few hundred hertz. Perhaps the most intriguing properties of the brain are its adaptability and capability for learning new tasks with minimal supervision and incomplete training data. With these capabilities in mind, we have developed an architecture that is selftraining, generalizable, fast, and energy efficient.
[0141] Superconducting hardware lends itself to a neuromorphic approach in part because of the natural spiking behavior of the Josephson junction (J J), similar to a neuron, and the near lossless propagation of these spikes on superconducting transmission lines, similar to axons. Because of potential improvements in energy efficiency and speed, a superconducting digital logic family has been developed using these JJ based spikes, called single flux quantum (SFQ) logic. These superconductingcircuits can operate at speeds in excess of 100 GHz and are extremely energy efficient. For example, the typical spiking energies of the JJs modeled here are less than one attojoule. Once a cryogenic system is scaled beyond the initial research and development phase, it takes less than 1000 W of wall power to cool 1 W of power dissipated at 4 K . This results in SFQ spiking energies that are still more than an order of magnitude lower than the human brain's 10 femtojoules per spike, even when accounting for the energy overhead required to operate at 4 K .
[0142] In part because of the characteristics noted above, there has been a recent interest in superconducting neuromorphic computing. Several different implementations of superconducting neuron or soma cells have been simulated and implemented in hardware. JJ neurons that take advantage of the high-speed dynamical interactions for greater biological realism have also been implemented. Various superconducting synaptic cells have been proposed including those based on hybrid superconducting spintronics, biased superconducting quantum interference devices (SQUIDs), and stochastic JJ synapses where the weight is transformed into a probability. Using these building blocks small JJ based circuits have been shown to exhibit spike timing dependent plasticity. There have also been small-scale demonstrations of superconducting networks and classifiers. The work presented here is distinct from these previous efforts because it is self-training and scalable to deeper neural networks. In addition, because of the nature of the self-training approach this architecture has the potential to be used for continual learning.
[0143] We take advantage of the SFQ logic family to create simple circuits that can implement reinforcement learning rules and have modified the typical learning rules to follow the constraints of what can be efficiently made with the hardware. In order to complete a superconducting neuromorphic architecture, we have developed a new synaptic function circuit, which controls the strength of the connection between two neurons. We demonstrate that modulating the amplitude of neuronal SFQ pulses can be used to implement the synaptic weighting. Further the weight value of our novel synaptic circuit can be programmed via SFQ spikes that are the output of reinforcement learning logic circuits. With these building blocks, we create a biologically inspired self-training architecture, demonstrate its basic functionality, and explore its ability to be scaled to larger tasks.
[0144] While algorithmic neural networks have proven to be useful for a multitude of tasks they are typically limited by the energy and time that it takes to train them30. Reinforcement learning has been shown to be a viable method for training neural networks. In its most basic form, reinforcement learning establishes a reward function based on the comparison between the desired and actual outputs of a system. We have modified the typical reinforcement learning algorithm to be compatible with our hardware constraints to create this self-training superconducting architecture. We test this architectural approach using accurate SPICE models that have been previously validated with fabricated neuromorphic circuit elements. We demonstrate the basic functionality of the neural network by having it learn several small problems using these SPICE models, demonstrating robust learning capability. We extend these results by implementing in Python the hardware-constrained learning rules in a neural network that learns the MNIST handwritten digit benchmark. The most striking result from this architecture is the training speed of order 1 ns per learning cycle, which we show has favorable scaling to large networks.
[0145] Results
[0146] We first implemented a SPICE model of a small but complete spiking, reinforcement learning-based self-training superconducting network. This network requires bipolar synaptic circuits, threshold soma circuits, and learning rule logic circuits, the details of which are described in the Methods section. FIG. 1 a shows a basic block diagram of this network. There are two inputs that each take an external binary signal and convert it to a single flux quantum (SFQ) spike for a 1 and no spike for a 0 . In addition, there is an input clock (not shown in the diagram) that spikes on every input cycle and is used for both a bias weight w0to adjust the threshold and the stochastic excitations (d) that are summed along with the weighted inputs in the soma.
[0147] The test network modeled in SPICE has two input units fully connected to a hidden layer of two units, which is subsequently fully connected to an output layer with one unit. The block diagram of the network can be seen in Fig. 1 a. The details of the weight circuits ( w ), the stochastic excitation circuits (d), and summation circuits ( s ) are described in the methods section. Briefly, the weights are bipolar synaptic circuits labeled as w( u ;where I is the layer, u is the unit, and i is the input, with i = 0 being the bias of a given unit. The summation circuit adds all incomingweighted signals with the bias and stochastic excitation. If the fixed threshold is exceeded the unit spikes and the signal is propagated to the next layer in the network.
[0148] The common Error Backpropagation learning rules that are used to adjust the weights in neural networks are based on stochastic gradient descent to minimize the mean squared error between the network's output and its desired outputs. In this work, we instead implement a correlative update algorithm involving a global reinforcement signal simultaneously sent to all hidden layer units. The algorithm, described in detail in the methods section, is based on Williams35definition of his non-episodic REINFORCE algorithm with consideration given to the efficient implementation in superconducting circuits. While this type of algorithm generally requires more updates than Error Backpropagation, it considerably reduces the complexity of the circuits required for training. It thus provides advantages in both size and speed in a direct hardware implementation.
[0149] FIG. 1 b shows the results from a single SPICE simulation of the network. The network is trained to classify an input vector by providing the desired output y* for a series of input vectors. In this simulation input 1 had an input series of 0011 repeated and input 2 had an input series of 0101 repeated. The time for inference plus learning in this example was 1 ns . The target output was divided into three phases. In phase 1 , which covers the first 500 inputs ( 500 ns in time), the network learned to spike whenever input 1 spikes (Xt) regardless of input 2 . In the second phase (from input 501 ns to 1000 ns ) the network learned to spike when either input 1 or input 2 spiked (OR). In the third phase, from 1001 ns to 1500 ns , the network learned to spike only when input 1 and input 2 spiked (AND). It should be noted that this was a continuous simulation where the only change between the phases was the target output y*. This helps illustrate the robustness of the training algorithm to any effect of the initial weight state. The effective training of these three phases can be seen in the graph of the probability correct of Fig. 1 b . The probability correct is a running average of four cycles of the probability that the output was correct for a given input vector, meaning that a value of 1 implies that the output of the network matched the target output for at least four cycles. FIG. 1 I shows that given a new desired output, the network can adjust its own weights without intervention to learn this function. This ability for the network to retrain is quite promising for continual learning applications.
[0150] The previous functions could be robustly learned without stochasticity. However, most non-trivial classifications benefit from the enhanced weight exploration that the additional stochastic term in the hidden layer units provides. FIG. 2 shows the same network training to spike whenever input 2 spikes ( X2). In this example, input 2 was ted an alternating signal of 010101... Without providing the same input for two samples in a row the change in the reward signal Ar and the change in the output Ay no longer provide a well-defined reinforcement signal for the weight updates to follow. FIG. 2a, b have a stochastic excitation of the hidden layer units that was set to < 1% of the maximum value of a synaptic weight. FIG. 2a shows the probability that the network outputs the correct value. FIG. 2 b shows the stored weights as current in the storage loop in the synapse. The subscripts of the weights correspond to the numbering scheme described in Fig. 1. It can be seen in this case that the periodic oscillation of the value of input 2 leads to an oscillation of the weight values around a local minimum in the error, which prevents the network from converging on the correct output behavior.
[0151] A stochastic excitation can enable training around this flawed input sample, even in this case where there is not a clear reinforcement signal for the hidden units to follow. In our SPICE simulations we have a stochastic input to each soma that is a bipolar synapse as described in the methods section. This stochastic input is tuned to have only three states available -1, 0, 1. The values are randomly selected in the simulation and could be physically implemented using a thermally excitable JJ. For the simulations the synapse gets an input at every cycle, similar to the bias input. The results of adding a stochastic excitation to the soma circuit can be seen in Fig. 2c, d which have the same input pattern and network as Fig. 2a, b. In the simulations for Fig. 2c, d the stochastic excitation synapse has an increased coupling to the soma with a magnitude of approximately 20% of the maximum weighted input value of the standard inputs. As can be seen in Fig. 2c the network now correctly learns to identify input 2 regardless of the input 1 value. In general, we find that a stochastic excitation of between 20% and 50% of the maximum weighted input is an effective magnitude for weight space exploration. It should be noted that if the inputs were sampled twice, e.g. 00110011 ... ,X2 could be correctly learned without the additional stochastic excitation, as expected. Also, we note for completeness that all cases in Fig. 1 can be learned with a stochastic excitation of 20% to 50% of the maximum weighted input. Ingeneral, as the problems become more complex the additional weight exploration provided by the stochastic excitation to the hidden units may become more beneficial for faster learning convergence.
[0152] FIG. 3 shows the same network as in Figs. 1 and 2 with a stochastic excitation of about 20% of the maximum weighted input learning the XOR function. This is a nonlinearly separable problem, which therefore requires the hidden layer and demonstrates the generality of both the network and its training rules. In this example we use an input pattern that is a set of 16 quasi-random input vectors that is repeated for 1000 total cycles. FIG. 3a shows the probability that the output correctly spikes when only input 1 or input 2 spikes, averaged over 4 cycles. FIG. 3b shows the weights of the synapses as the network correctly learns the XOR function in less than 400 ns. Weight subscripts are the same as defined in Fig. 1 . FIG. 3c shows the spiking pattern of the inputs, the desired output y*, and the actual network output for the first 50 ns of training when the probability correct is approximately 25%. FIG. 3d shows the spiking patterns from 950 ns to 1000 ns of the SPICE simulation where one can see that the network has correctly learned the XOR function.
[0153] Extension to MNIST
[0154] To investigate the suitability of the reinforcement learning approach for larger problems, simulations were run with a python model implementing the learning rules described in the methods section and applied to the problem of classifying images of handwritten digits in the modified national institute of standards and technology (MNIST) dataset. We tried to put limits in the simulations to reflect the constraints of the SPICE models above, these include a weight bit depth of 6 bits and a fan-in limitation of 50 . However, details such as cross talk and wiring layout were not considered. The results in Table 1 were obtained with neural networks containing one, two, or three hidden layers, each with 200 units. The networks were trained using 50,000 images for training, 10,000 images for validation and 10,000 images for test as is typical. The training, validation, and testing procedure was performed five times with different random initial weights each time and also a different set of random connections, that arise from the fan-in limitation of 50 detailed below, each time. Table 1 shows the minimum and maximum values of accuracy for the training, validation and test datasets.
[0155] The current hardware design of the soma has a limit of 50 incoming connections. This was modeled in the python implementation in the following way. Each unit in the first hidden layer was connected to a randomly chosen 7 x 7 patch of the 28 x 28 original image. While this is like the kernels in convolutional neural networks, here convolution is not performed. Each unit only gets the 49 intensities from the single 7 x 7 patch randomly assigned to it when the neural network is defined. The units in the remaining hidden layers receive inputs from 49 randomly selected outputs from the previous hidden layer. Each of the 10 units in the output layer receives 49 randomly selected outputs from the last hidden layer. Thus, each unit in the neural network has 49 weights corresponding to its 49 inputs, plus one bias weight, for a total of 50 weights in each unit.Table 1 | The percentage of train and test set images for MN I ST that are correctly classified by networks having no, 1 , 2, or 3 hidden layers of 200 units each
[0156] Classification is performed by picking the unit with the largest value among the 10 units in the output layer. Several possibilities exist for implementing this "argmax" operation in the circuit, including a winner-take-all operation based on recurrent connections in the output layer. SPICE modeling of this part of the circuit is beyond the scope of this paper and will be investigated in future work.
[0157] Scaling
[0158] In the above examples we show SPICE simulations of a small proof- of principle self-training network that fully implements reinforcement learning using a common global reward signal. The network is a mixed digital and analog set of circuits that is fully self-contained. The bipolar synapse and soma constitute the analog portion of the network. The analog weights are adjusted via injection of SFQ pulses into a superconducting weight storage loop. SFQ pulses can be injected on either side of the loop, resulting in the ability to increase or decrease the strength of the weight in either direction. The bipolar synapse architecture further enables any synapse to be either inhibitory or excitatory in nature. This bipolar synaptic network style enables a more efficient hardware implementation since synapses do not need to be duplicated or assigned a sign in advance of the training. Because the learning rules determine the direction of a given weight update, there is no need to precisely set the weight valueor threshold in the analog portions of the network. This alleviates one of the main challenges in the use of analog hardware, since the typical process variations inherent in all hardware can be accommodated by the learning rules.
[0159] Table 2 shows the timing delays based on our SPICE simulations. These timings could likely be reduced with further circuit optimization. In addition, it is worth noting that the 150 ps delay in inference per layer is adjustable based on an L / R time constant in the soma circuit. We choose a 50 ps time constant, which is the approximate time that all incoming spikes to a given soma should arrive. We choose this number to easily account of any jitter or wiring layout non-idealities. Further, this architecture is globally asynchronous. This alleviates many of the issues of clocking SFQ circuits when scaled up.Table 2 | Timing delays based on SPICE simulations.
[0160] The total 750 ps learning delay listed in Table 2 comes from three logic updates that each take 250 ps in their current form. These logic cells are the output reward update, the local soma update, and the local synapse update. Since the update rules do not need to be applied sequentially between layers or neurons, the reward signal can be broadcast back to all somas at the same time and the soma rule can be broadcast to all incoming synapses at the same time.
[0161] The real advantage of the local learning rules and architecture can be seen when examining how the time delays scale up. We present the 3-layer MNIST network timing as a detailed example. Inference latency is 10 ps input +48 ps fanout to hidden layert+ lSOpssomaSi + 48ps fanout to hidden layer2+ 150ps somas2+ 48ps fanout to hidden layer3+ 150ps somas3+ 23ps fanout to output layer 100 ps wiring. This gives us 727 ps for the inference latency. The learning update delay is still 750 ps because all updates are independent +60 ps fanout to all neurons delay +36 ps fanout to the incoming synapses delay +100 ps wiring. This gives us 946 ps learning update delay. This gives us a total learning time of 1673 ps per image. Given the 60,000 training images shown twice each for 500 epochs we would expect a total hardware training time of 100.4 ms .
[0162] From these two examples we can see three main reasons for the favorable time scaling of the proposed architecture. First, superconducting wiring transmits SFQ signals extremely fast and is not subject to RC time constants. Second, the fanout scales as log2meaning that even a fanout to 1,000,000 units would only add a time delay of 120 ps. Finally, this architecture takes advantage of true local learning rules meaning that the weight updates are independent and globally asynchronous.
[0163] FIG. 4 shows the total time delay for a full inference plus learning cycle versus the number of layers in the network for four different network widths. There is a linear scaling of the inference latency on the number of layers, which dominates the total time as the network grows in size. However, the total time for a learning cycle is still relatively short because of the ability to broadcast the reward signal to the entire network at the same time. While the wiring for this type of networkwill be a challenge, it does not incur the usual RC time constant or energy issues of typical resistive wiring. It is worth noting that there will be additional timing delays as the network grows due to wiring, but these have a minimal impact on the time scaling. The time to propagate SFQ pulses in Nb is roughly c / 3 (where c is the speed of light in vacuum), or 10ps / mm40. While repeater JJs will be needed for JTLs, passive superconducting transmission lines can be used for longer wires. Thus, the additional timing overhead of routing the broadcast signal should be much less than 1 ns . This gives us a 1 million unit, 100 layer deep learning time of about 31 ns or about 30 million cycles per second and favorable time scaling as the network gets larger.
[0164] Finally, it is worth noting that million unit networks are not within reach of current fabrication densities. However, our unoptimized example network contained a little less than 1000 JJs for 5 units in total. Using 200 JJs per unit as a very rough estimation, networks with order 1000 units should be achievable with current fabrication technology. This includes the 3 hidden layer MNIST network in our python simulations. If implemented in our JJ reinforcement learning architecture, the 3 hidden layer MNIST network would take about 100 ms as described above. However, after training, image classification could be run at a rate of 1 image every 150 ps ( 6.7 GHz).
[0165] Discussion
[0166] Superconducting neuromorphic computing is relatively undeveloped compared to other emerging devices. There are many potential reasons for this. However, the most obvious limitation is the requirement of cryogenic operation, and this should be considered. In this paper we assume a Nb technology that is operational at 4 K. Closed cycle cryogenics has come a long way in recent years expanding potential applications. However, edge computing, which many other emerging devices are targeting, is not a good fit for cryogenic operation.
[0167] In this paper we have demonstrated the basic components needed to make a complete self-training superconducting neuromorphic network. This includes the analog and digital circuits that are required for basic neural network inference. The digital circuits that are required for learning, and the learning rules that can be efficiently implemented with this hardware. Together these components comprise a superconducting neuromorphic architecture, which we have shown hasthe potential to scale to more interesting applications. Scaling this architecture past the MNIST level will likely require some further innovations. Of particular interest is finding the best way to work with the fan-in limitation of 50 . There are likely architectural choices that mimic the dendritic processing of the brain that can be used to work with this limitation. Future work to increase the capability of these systems has compelling potential.
[0168] Even with these limitations, we find that there are many potential advantages that make superconducting neuromorphic computing worth pursuing. High-speed, energy efficiency (if scaled), and natural spiking behavior make superconducting neuromorphic computing a compelling technology to investigate. In this article we have shown a mixed analog and digital neuromorphic SFQ architecture that is self-training using reinforcement learning. The learning rules implemented were constrained for efficient implementation in superconducting hardware. We use SPICE simulations to demonstrate a small-scale network capable of classifying a non-linearly separable problem. We have also demonstrated robust learning with this architecture in part due to the use of stochastic weight exploration in the hidden layer units. We have extended the same reinforcement learning logic demonstrated in the SPICE simulations to python where the network can be scaled to perform MNIST level classifications.
[0169] The highlight of this architecture is the ability for the network to selftrain given the target output value. This approach mitigates the non-idealities in device variations that is problematic in analog neuromorphic hardware. In addition, the fastlearning time leads to a cycle time of about 1 ns for small networks and excellent timescaling for larger networks. For example, projections for the total time to train a 3 hidden-layer network to learn the MNIST benchmark are on the order of 100 ms . In addition, after the learning phase, the network can be run in inference only mode, where the same network can operate at one image every 150 ps. These demonstrations prove out the basic building blocks and architecture for a scalable, high-speed network that has inherent fault tolerance. In addition to the speed, the spiking energies of this network are sub-attojoule providing energy efficiency even when accounting for the cooling overhead. Together this architecture provides a clear path and a compelling argument for superconducting neuromorphic computing.
[0170] Methods
[0171] Learning rules
[0172] All experiments here are supervised-learning problems for which the desired outputs are known. The common approach for training neural networks to solve supervised learning problems is to use Error Backpropagation to back-propagate the gradient of the mean squared error with respect to every weight. T o avoid the extra time and circuits required to perform this backwards flow of information through the multiple hidden layers of a neural network, we instead implement a correlative update algorithm involving a global reinforcement signal simultaneously sent to all hidden layer units. While this requires more updates than Error Backpropagation, it considerably reduces the complexity of the circuits required for training.
[0173] The basic definitions and notation used for the network are: a network with L layers, each layer I having Ntunits; weights, wbUiifor layer I, unit u, and input t; bias weights, wlu 0for layer I, unit u; inputs, xl tfor layer I, and input i; weighted sums, sl>ufor layer I, unit u; output yl ufor layer I, unit u, which becomesthe inputs to the following layer; stochastic perturbation dLiUfor layer I, unit u, where di,u= {-5, 0, 5}; global reward r; weight update wbu i+, with increment size p; and an input sample consisting of N±inputs to the first layer, layer 1 , with the corresponding desired output y „, for unit u in the last layer, layer L.
[0174] The equations of the network are slightly different between the hidden layers and the output layer. The output layer is where the global binary reinforcement signal is determined and the target of the output of the units is known directly. The output layer uses its current inputs xL i, outputs yL uand desired outputs y£„ t° update its weights. The hidden layers need the current inputs and outputs, a global reinforcement signal, and a memory of the previous output state in order to determine the weight update direction. They also require a stochastic excitation or inhibition input to ensure that the network effectively samples the available states. We find that best practice for learning with this network is for a sample (inputs and desired output) to remain fixed for two time increments before changing conditions. We define the equations of the network as follows:
[0175] Equations for the output layer units
[0176] Equations for hidden layer units*i,i = yi-i,i for I = 1,2,3, ... (7)
[0177] The above equations for updating the hidden layer units are based on Williams335definition of his non-episodic REINFORCE algorithm, given by:
[0178] Some describe these equations as pertaining to a bias weight having a constant input of 1 . With the addition of variable inputs, xJtand using y = 1, these equations reduce to wj = a(r - rt-1)(y - yt-1)(x / - l / 2), (16) which correspond to the weight update equations given above for the hidden layer units. Letting W represent all of the weights in the network, Williams proves for his REINFORCE class of algorithms, that the inner product of E{ W I W) and VE{r I VIZ] is nonnegative. While this is not a convergence proof, it does prove that the change in weight values, Aw, tends to increase the expected value of r.
[0179] A known limitation of Williams's REINFORCE algorithm is that it learns more slowly than a version that includes a learned value function. In future work, a learned value function could be added using a second neural network implemented and trained similarly to our current neural network implementation.
[0180] Single flux quantum circuit implementation of the learning rules
[0181] The weight update rules for the output layer can be implemented with standard rapid single flux quantum (RSFQ) digital logic gates and a learning clock signal whose only requirement is to arrive any time after the inference output. While there are many potential ways to implement the logic, we used the following basic approach. Weight update values that are potentially zero, for example (y* - y), are used as the clock signal for the logic gates that are closer to the synapse. The absence of this clock ensures that there is no unwanted weight update. The sign of any prefactor is determined as close as possible to the output layer. For example, ArAy is determined at the soma output where ynowis stored during inference. The final signof the update is determined at the synapse since it depends on the input x to the synapse. Finally, all values that are stored from the inference step are reset at the end of the learning cycle to prepare for the next inference cycle. Approaching the update logic in this way concentrates the number of logic gates at the output and soma level, minimizing the number of gates at the synapse and the amount of information that must be broadcast to each synapse.
[0182] FIG. 5 shows an example of the nominal weight update logic for an output layer synapse. The desired output y' and network output y are fed into an exclusive or (XOR) gate to determine if there is a weight update. If the values are the same, (y* - y) = 0 there should be no update. Otherwise, the signal from this XOR is used as a clock for the rest of the update logic. To determine the sign of the weight update, the input to the synapse x and the desired output y* are fed into an XOR gate via destructive readout (DRO) gates to determine the direction of the update. The output and complementary output of this XOR feed a single flux quantum (SFQ) pulse into the decrement or increment side respectively of the weight storage loop described below in the synapse section. A similar logical flow is followed for the other weight updates.
[0183] Bipolar synapse and soma circuit While the neuron and axon building blocks are easily implemented with simple superconducting elements, the required weight and threshold functions are implemented using an analog circuit that we break up into the synapse (weighting operation) and soma (summation operation)5 / L1. The synaptic subcircuit in Fig. 6a is composed of a bipolar "synaptic" weight, meaning that the weight can be either positive or negative, connected to a summation "soma." The soma contains a line of mutually coupled inductors with a resistor to ground at one end and the input to a Josephson transmission line at the other. The inputs are set to arrive at the same time where they are weighted, and a portion of the incoming spikes are coupled into the soma. The resistor to ground at the bottom of the soma can be used to adjust the L / R time constant of this summation operation to account for wiring and jitter on the incoming signals, relaxing the timing condition on spikes arriving to the soma. The top of the soma feeds a biased J J, which then applies the threshold activation function defined above as yL u. One technical difference is that in the circuit implementation, the threshold value is shifted from zero current in the soma to apositive current value. The threshold is set by a combination of the fixed bias current of the JJ at the top of the soma that is the input to the Josephson transmission line, and the w0bias weight, which is adjusted by the reinforcement learning logic. The effect however is the same; if the threshold is exceeded a single flux quantum (SFQ) pulse is emitted by the soma (1), otherwise nothing is emitted (0).
[0184] A more detailed look at the synapse shown in Fig. 6a follows: The weighting is accomplished using an inductive divider, where one path passes through an adjustable superconducting quantum interference device to ground, and the other path passes through an inductor in series with a resistor to ground. This second path's inductor is mutually coupled to the soma along with the other synapses that are connected to that soma. The resistor on the second leg of the divider is used to prevent unintended current buildup and to set the L / R time constant (leak rate) of the input signal to a synapse. The weight storage loop (shown in grey wiring), controlled by the learning logic, affects the inductive divider by changing the amount of current circulating in adjustable SQUID, which in turn changes its inductance. Each synapse has a positive and negative branch. Both branches have the same inductive divider topology described above but with the coupling directions changed.
[0185] FIG. 6 b shows the behavior of the synapse from SPICE simulations. The storage loop starts at zero current and an input pulse x is applied to the synapse. In this case the positive and negative branches have equal and opposite current coupled into the soma, which is seen as no spike appearing in the soma at the bottom of Fig. 6b. Note that while perfect cancelation is accomplished in the simulation, it is not required to train the network. After the first input spike, a series of 32 SFQ pulses is applied to the decrement side of the weight storage loop in about 1 ns . Note that the number of SFQ pulses stored in the weight storage loop has a maximum of 30 for the simulated circuit shown here with a storage inductor of 250 pH . When the maximum number of SFQ pulses for the weight storage loop is exceeded, the additional SFQ pulses are expelled via the JJ on the other side of the loop and the maximum current value in the weight storage loop is maintained. After this minimum value of the weight storage loop is reached, a subsequent input pulse x is applied to the synapse at 3 ns and now a larger portion of the negative branch and a smaller portion of the positive branch are coupled into the soma. This is seen in the soma atthe bottom of Fig. 6 b as the net negative spike. After this, 64 SFQ pulses are applied to the increment side of the weight storage loop. Subsequently an input pulse x is applied to the synapse, which results in a larger positive and smaller negative spike. This is seen in the soma as a net positive spike from the input x. Finally, 30 pulses are applied to the decrement side of the weight storage loop to return the weight storage loop to near zero current. The final input pulse x is applied to the synapse and once again the positive and negative branches cancel out and no spike is observed in the soma.
[0186] One other element that is used is a w0, which acts as an adjustable bias for the soma. The w0circuit is the same as the positive branch of the bipolar synapse circuit. Each soma circuit has one such w0to adjust the threshold level. There is a separate weight loop and its own learning rule logic to set and store this iv0for each soma. In addition, there is a iv0input that spikes at every inference cycle. This results in the bias value being independent of whether an input signal was sent to a particular neuron, and no external adjustment of the soma bias is needed as described in the network equations above.
[0187] The weight storage is a superconducting loop with a biased JJ on either side, shown in light blue in the middle of Fig. 6a. Current can be injected in either direction in this loop following the result of the learning rules. The size of the inductor sets the effective bit depth of the weight storage. In addition, the adjustable SQUIDs in the synapses are set so that they always remain in the non-spiking state regardless of the weight value, which results in a smooth variation in weight value. In the simulations described in this article we use a 250 pH storage inductor, which results in just under a 6 -bit weight. The bit depth is approximately linear with inductor size, meaning that one can use a 500 pH storage inductor for a weight depth of about 6 bits.
[0188] The processes described herein may be embodied in, and fully automated via, software code modules executed by a computing system that includes one or more general purpose computers or processors. The code modules may be stored in any type of non-transitory computer-readable medium or other computer storage device. Some or all the methods may alternatively be embodied in specializedcomputer hardware. In addition, the components referred to herein may be implemented in hardware, software, firmware, or a combination thereof.
[0189] Many other variations than those described herein will be apparent from this disclosure. For example, depending on the embodiment, certain acts, events, or functions of any of the algorithms described herein can be performed in a different sequence, can be added, merged, or left out altogether (e.g. , not all described acts or events are necessary for the practice of the algorithms). Moreover, in certain embodiments, acts or events can be performed concurrently, e.g., through multithreaded processing, interrupt processing, or multiple processors or processor cores or on other parallel architectures, rather than sequentially. In addition, different tasks or processes can be performed by different machines and / or computing systems that can function together.
[0190] Any logical blocks, modules, and algorithm elements described or used in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, and elements have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. The described functionality can be implemented in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the disclosure.
[0191] The various illustrative logical blocks and modules described or used in connection with the embodiments disclosed herein can be implemented or performed by a machine, such as a processing unit or processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A processor can be a microprocessor, but in the alternative, the processor can be a controller, microcontroller, or state machine, combinations of the same, or the like. A processor can include electrical circuitry configured to process computer-executable instructions.In another embodiment, a processor includes an FPGA or other programmable device that performs logic operations without processing computer-executable instructions. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Although described herein primarily with respect to digital technology, a processor may also include primarily analog components. For example, some or all of the signal processing algorithms described herein may be implemented in analog circuitry or mixed analog and digital circuitry. A computing environment can include any type of computer system, including, but not limited to, a computer system based on a microprocessor, a mainframe computer, a digital signal processor, a portable computing device, a device controller, or a computational engine within an appliance, to name a few.
[0192] The elements of a method, process, or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module stored in one or more memory devices and executed by one or more processors, or in a combination of the two. A software module can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of non- transitory computer-readable storage medium, media, or physical computer storage known in the art. An example storage medium can be coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor. The storage medium can be volatile or nonvolatile.
[0193] While one or more embodiments have been shown and described, modifications and substitutions may be made thereto without departing from the spirit and scope of the invention. Accordingly, it is to be understood that the present invention has been described by way of illustrations and not limitation. Embodiments herein can be used independently or can be combined.
[0194] All ranges disclosed herein are inclusive of the endpoints, and the endpoints are independently combinable with each other. The ranges are continuous and thus contain every value and subset thereof in the range. Unless otherwise statedor contextually inapplicable, all percentages, when expressing a quantity, are weight percentages. The suffix (s) as used herein is intended to include both the singular and the plural of the term that it modifies, thereby including at least one of that term (e.g., the colorant(s) includes at least one colorants). Option, optional, or optionally means that the subsequently described event or circumstance can or cannot occur, and that the description includes instances where the event occurs and instances where it does not. As used herein, combination is inclusive of blends, mixtures, alloys, reaction products, collection of elements, and the like.
[0195] As used herein, a combination thereof refers to a combination comprising at least one of the named constituents, components, compounds, or elements, optionally together with one or more of the same class of constituents, components, compounds, or elements.
[0196] All references are incorporated herein by reference.
[0197] The use of the terms “a,” “an,” and “the” and similar referents in the context of describing the invention (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. It can further be noted that the terms first, second, primary, secondary, and the like herein do not denote any order, quantity, or importance, but rather are used to distinguish one element from another. It will also be understood that, although the terms first, second, etc. are, in some instances, used herein to describe various elements, these elements should not be limited by these terms. For example, a first current could be termed a second current, and, similarly, a second current could be termed a first current, without departing from the scope of the various described embodiments. The first current and the second current are both currents, but they are not the same condition unless explicitly stated as such.
[0198] The modifier about used in connection with a quantity is inclusive of the stated value and has the meaning dictated by the context (e.g., it includes the degree of error associated with measurement of the particular quantity). The conjunction or is used to link objects of a list or alternatives and is not disjunctive;rather the elements can be used separately or can be combined together under appropriate circumstances.
Claims
What is claimed is:1 . A bipolar single flux quantum weight circuit comprising: a superconducting weight storage loop (200) comprising a weight storage inductor (200a); a first superconducting quantum interference device (SQUID) (201) mutually coupled to the superconducting weight storage loop (200) with a first polarity; a second SQUID (202) mutually coupled to the superconducting weight storage loop (200) with a second polarity opposite the first polarity, wherein an inductance of the first SQUID (201 ) and an inductance of the second SQUID (202) vary based upon a current circulating within the superconducting weight storage loop (200); a first inductive divider (203) electrically connected to an input terminal (210a), the first inductive divider (203) comprising a first leg including the first SQUID (201 ) connected to ground (GND) and a second leg including a first coupling inductor (205); a second inductive divider (204) electrically connected to the input terminal (210a), the second inductive divider (204) comprising a first leg including the second SQUID (202) connected to ground (GND) and a second leg including a second coupling inductor (206); and a soma element (207) comprising a soma inductor line (207a), the first coupling inductor (205) mutually coupled to the soma inductor line (207a) with a third polarity, and the second coupling inductor (206) mutually coupled to the soma inductor line (207a) with a fourth polarity opposite the third polarity, the soma element (207) further comprising a soma output Josephson junction (209) connected to the soma inductor line (207a).
2. The circuit of claim 1 , wherein the superconducting weight storage loop (200) further comprises a first Josephson junction (200b) and a second Josephson junction (200c) arranged to permit injection of single flux quantum pulses for increasing ordecreasing the current circulating within the superconducting weight storage loop (200).
3. The circuit of claim 1 , wherein the second leg of the first inductive divider(203) further comprises a first resistor (205a) in series with the first coupling inductor (205) connected to ground (GND), and the second leg of the second inductive divider(204) further comprises a second resistor (206a) in series with the second coupling inductor (206) connected to ground (GND).
4. The circuit of claim 1 , wherein the soma element (207) further comprises a soma resistor (207b) connecting the soma inductor line (207a) to ground (GND), and wherein the soma output Josephson junction (209) connects the soma inductor line (207a) to an output Josephson transmission line (208).
5. The circuit of claim 1 , wherein the first SQUID (201 ) is physically positioned relative to the superconducting weight storage loop (200) differently than the second SQUID (202) to achieve the first polarity and the second polarity of mutual coupling.
6. The circuit of claim 1 , wherein the mutual coupling of the first coupling inductor (205) to the soma inductor line (207a) and the mutual coupling of the second coupling inductor (206) to the soma inductor line (207a) result in signals induced in the soma element (207) that are opposite in polarity relative to each other upon application of an input pulse to the input terminal (210a).
7. The circuit of claim 1 , further comprising an input splitter (210) connecting the input terminal (210a) to both the first inductive divider (203) and the second inductive divider (204).
8. The circuit of claim 1 , wherein the first SQUID (201 ) and the second SQUID (202) are biased to remain in a non-spiking state across a range of the current circulating within the superconducting weight storage loop (200).
9. The circuit of claim 1 , further comprising a bias weight circuit (21 1 ) comprising a bias weight storage loop (212) and a bias weight inductive divider leg (213), the bias weight inductive divider leg (213) mutually coupled to the soma inductor line (207a).
10. The circuit of claim 1 , wherein the superconducting weight storage loop (200), the first SQUID (201 ), the second SQUID (202), the first coupling inductor (205), the second coupling inductor (206), the soma inductor line (207a), and the soma output Josephson junction (209) comprise superconducting materials and Josephson junctions.
11. A method for weighting a unipolar single flux quantum (SFQ) input pulse (x) utilizing a bipolar single flux quantum weight circuit comprising a superconducting weight storage loop (200) containing a variable weight current, a first SQUID (201 ) and a second SQUID (202) oppositely mutually coupled to the weight storage loop (200), a first inductive divider (203) including the first SQUI D (201 ) and a first coupling inductor (205), a second inductive divider (204) including the second SQUID (202) and a second coupling inductor (206), and a soma element (207) having a soma inductor line (207a) mutually coupled with opposite polarities to the first coupling inductor (205) and the second coupling inductor (206) and connected to a soma output Josephson junction (209), the method comprising: applying the unipolar SFQ input pulse (x) to an input terminal (210a) connected to the first inductive divider (203) and the second inductive divider (204);dividing the SFQ input pulse (x) via the first inductive divider (203), wherein a first portion is directed through the first SQUID (201 ) and a second portion is directed through the first coupling inductor (205), the division ratio being dependent on an inductance of the first SQUID (201 ) influenced by the weight current; dividing the SFQ input pulse (x) via the second inductive divider (204), wherein a third portion is directed through the second SQUID (202) and a fourth portion is directed through the second coupling inductor (206), the division ratio being dependent on an inductance of the second SQUID (202) influenced by the weight current; inducing a first signal in the soma inductor line (207a) from the second portion via the first coupling inductor (205) with a first signal polarity; inducing a second signal in the soma inductor line (207a) from the fourth portion via the second coupling inductor (206) with a second signal polarity opposite the first signal polarity; and generating a net signal in the soma element (207) based on a summation of the first induced signal and the second induced signal, wherein the soma output Josephson junction (209) emits an output SFQ pulse (y) if the net signal exceeds a predetermined threshold.
12. The method of claim 1 1 , further comprising programming the variable weight current in the superconducting weight storage loop (200) by injecting one or more SFQ current pulses into the superconducting weight storage loop (200) via a first Josephson junction (200b) or a second Josephson junction (200c) associated with the loop (200).
13. The method of claim 12, wherein injecting the one or more SFQ current pulses adjusts the variable weight current to increase or decrease a magnitude of the net signal generated in the soma element (207) in response to the SFQ input pulse (x).
14. The method of claim 11 , wherein when the variable weight current in the superconducting weight storage loop (200) is substantially zero, the first induced signal and the second induced signal substantially cancel each other, resulting in the net signal being substantially zero.
15. The method of claim 11 , wherein adjusting the variable weight current in a first direction increases the magnitude of the first induced signal and decreases the magnitude of the second induced signal, resulting in a net signal with the first signal polarity.
16. The method of claim 11 , wherein adjusting the variable weight current in a second direction, opposite the first direction, decreases the magnitude of the first induced signal and increases the magnitude of the second induced signal, resulting in a net signal with the second signal polarity.
17. The method of claim 11 , further comprising operating the circuit within a reinforcement learning framework, the method including: performing an inference cycle by applying the SFQ input pulse (x) and generating the output SFQ pulse (y) or no pulse; storing a value representing the SFQ input pulse (x) and a value representing the output SFQ pulse (y) in memory elements (215); receiving a learning clock signal (217) to initiate a learning cycle; calculating a weight update value based on the stored value representing the SFQ input pulse (x), the stored value representing the output SFQ pulse (y), a target output value (y*), and potentially a reinforcement signal (r) using learning rule logic (214); andadjusting the variable weight current in the superconducting weight storage loop (200) based on the calculated weight update value by injecting SFQ pulses.
18. The method of claim 17, wherein calculating the weight update value for hidden layers within a neural network utilizes a change in the reinforcement signal (Ar) between a previous learning cycle and a current learning cycle and a change in the output SFQ pulse (Ay) between a previous inference cycle and a current inference cycle.
19. The method of claim 17, further comprising applying a stochastic SFQ pulse generated by a stochastic hardware element (216) to the soma element (207) during the inference cycle to enhance weight exploration, wherein the stochastic hardware element (216) utilizes thermal fluctuations at an operating temperature to generate random SFQ pulses.
20. The method of claim 11 , wherein the method forms part of a neural network computation, the net signal representing a weighted synaptic input to a neuron, and the output SFQ pulse (y) representing a neuron firing.
21. A method for fabricating a bipolar single flux quantum weight circuit, the method comprising: forming a superconducting weight storage loop (200) including a weight storage inductor (200a) from superconducting material layers; forming a first superconducting quantum interference device (SQUID) (201 ) and a second SQUID (202) adjacent to the superconducting weight storage loop (200); arranging the first SQUID (201 ) relative to the superconducting weight storage loop (200) to establish a mutual inductive coupling exhibiting a first polarity;arranging the second SQUID (202) relative to the superconducting weight storage loop (200) to establish a mutual inductive coupling exhibiting a second polarity substantially opposite to the first polarity; forming a first inductive divider (203) comprising a first leg including the first SQUID (201 ) and a second leg including a first coupling inductor (205); forming a second inductive divider (204) comprising a first leg including the second SQUID (202) and a second leg including a second coupling inductor (206); forming a soma element (207) comprising a soma inductor line (207a) and a soma output Josephson junction (209) connected thereto; arranging the first coupling inductor (205) relative to the soma inductor line (207a) to establish a mutual inductive coupling exhibiting a third polarity; arranging the second coupling inductor (206) relative to the soma inductor line (207a) to establish a mutual inductive coupling exhibiting a fourth polarity substantially opposite to the third polarity; and electrically connecting the first inductive divider (203) and the second inductive divider (204) to an input terminal (210a), connecting the first leg of the first inductive divider (203) and the first leg of the second inductive divider (204) to ground (GND), and connecting the second leg of the first inductive divider (203) and the second leg of the second inductive divider (204) to ground (GND).
22. The method of claim 21 , further comprising forming a first Josephson junction (200b) and a second Josephson junction (200c) within the superconducting weight storage loop (200), positioned to enable bidirectional current injection.
23. The method of claim 21 , further comprising forming a first resistor (205a) in series with the first coupling inductor (205) in the second leg of the first inductive divider (203), and forming a second resistor (206a) in series with the second coupling inductor (206) in the second leg of the second inductive divider (204).
24. The method of claim 21 , further comprising forming a soma resistor (207b) connecting the soma inductor line (207a) to ground (GND), and electrically connecting the soma output Josephson junction (209) to an output Josephson transmission line (208).
25. The method of claim 21 , wherein arranging the first SQUID (201 ) and the second SQUID (202) relative to the superconducting weight storage loop (200) involves positioning conductive layers of the first SQUID (201 ) and the second SQUID (202) on opposing sides of a conductive layer of the superconducting weight storage loop (200).
26. The method of claim 21 , wherein arranging the first coupling inductor (205) and the second coupling inductor (206) relative to the soma inductor line (207a) involves configuring winding directions or relative positions to achieve the third and fourth opposite polarities.
27. The method of claim 21 , further comprising forming an input splitter (210) configured to electrically connect the input terminal (210a) simultaneously to the first inductive divider (203) and the second inductive divider (204).
28. The method of claim 21 , further comprising forming electrical connections for applying bias currents to the first SQUID (201) and the second SQUID (202) sufficient to maintain them in a non-spiking state during operation.
29. The method of claim 21 , further comprising:forming a bias weight circuit (211) including a bias weight storage loop (212) and a bias weight inductive divider leg (213); and arranging the bias weight inductive divider leg (213) relative to the soma inductor line (207a) to establish a mutual inductive coupling.
30. The method of claim 21 , wherein forming the superconducting weight storage loop (200), the first SQUID (201), the second SQUID (202), the first coupling inductor (205), the second coupling inductor (206), the soma inductor line (207a), and the soma output Josephson junction (209) utilizes deposition and patterning of superconducting thin films, insulating layers, and resistive materials.