Analog memory-based complex multiply-accumulate (MACC) calculation engine
The analog memory-based complex MACC computation engine addresses the limitation of traditional crossbar arrays by enabling efficient complex-valued MACC operations, improving adaptability and generalization in deep learning with minimal modifications, thus extending energy efficiency and throughput advantages.
Patent Information
- Application Number
- JP2025523006
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-31
- Filing Date
- 2023-09-06
- Publication Date
- 2025-10-30
AI Technical Summary
Existing crossbar arrays primarily focus on real-valued operations and lack the capability to efficiently perform complex-valued multiplication-accumulation (MACC) operations necessary for complex-valued neural networks, limiting their adaptability and generalization in deep learning applications.
An analog memory-based complex MACC computation engine is introduced, utilizing pulse width modulators and differential circuits with configurable resistors to facilitate simultaneous computation of real and imaginary terms, enabling complex-valued MACC operations with minimal modifications to traditional crossbar arrays.
This solution allows for efficient execution of complex-valued MACC operations, enhancing adaptability and generalization in deep learning, while maintaining energy efficiency and throughput advantages of traditional crossbar arrays with low overhead pulse width modulation rerouting.
Smart Images

Figure 2025535927000001_ABST
Abstract
Description
[Background technology]
[0001] The present invention relates generally to electrical, electronic and computer technology, and more particularly to computational engines.
[0002] The multiplication-accumulation (MACC) operation in machine learning typically uses only real values. However, complex-valued neural networks are attracting more and more research interest. For example, one conventional application provides automatic differentiation for implementing the backpropagation algorithm, which supports a large number of complex-valued layers and activation functions.
[0003] Crossbar arrays are also traditionally used to implement neural networks and numerical computation units, providing efficient in-memory computation. Supervised and unsupervised learning can be implemented using the artificial behavior of synapses and neurons provided by neural networks implemented on crossbar arrays. Crossbar arrays also support the performance of MACC operations, which are based on the physics of Ohm's law and Kirchhoff's laws. Crossbar arrays can accelerate the performance of MACC operations, but typically focus only on real-valued operations. Summary of the Invention
[0004] The principles of the present invention provide an analog memory-based complex multiply-accumulate (MACC) computation engine. In one aspect, an exemplary circuit comprises: a first pulse width modulator configured to generate a first pulse based on a first input; a second pulse width modulator configured to generate a second pulse based on a second input; a first differential circuit including a first transistor, a second transistor, a first resistor, and a second resistor; a second differential circuit including a first transistor, a second transistor, a first resistor, and a second resistor, wherein a gate of the first transistor of the first differential circuit and a gate of the second transistor of the first differential circuit; and a gate of the first transistor of the second differential circuit and a gate of the second transistor of the second differential circuit are configured to be controlled by the first pulse width modulator and the second pulse width modulator based on the first input and the second input; the first resistor of the first differential circuit has a configurable resistance, a first terminal of the first resistor is coupled to a voltage supply node and a second terminal of the first resistor is coupled to a voltage supply node and a second terminal of the first resistor is coupled to a voltage supply node and a second terminal of the first transistor of the first differential circuit. the first terminal of the second resistor of the first differential circuit has a configurable resistance, the first terminal of the second resistor of the first differential circuit is coupled to the voltage via the voltage supply node, and the second terminal of the second resistor of the first differential circuit is coupled to a first terminal of the second transistor of the first differential circuit; the first resistor of the second differential circuit has a configurable resistance, the first terminal of the first resistor of the second differential circuit is coupled to the voltage via the voltage supply node, and the second differential circuit a second terminal of the first resistor coupled to a first terminal of the first transistor of the second differential circuit; and the second resistor of the second differential circuit having a configurable resistance, the first terminal of the second resistor of the second differential circuit coupled to the voltage via the voltage supply node and the second terminal of the second resistor of the second differential circuit coupled to a first terminal of the second transistor of the second differential circuit; the circuit further comprises an analog-to-digital converter having an input coupled to the voltage supply node.
[0005] In one aspect, the circuit includes a differential circuit having a first transistor, a second transistor, a first resistor, and a second resistor; a first pulse width modulator configured to generate pulses based on a first input, the output of the first pulse width modulator being configured to be selectively coupled to one of a gate of the first transistor of the differential circuit and a gate of the second transistor of the differential circuit; the first resistor of the differential circuit having a configurable resistance, a first terminal of the first resistor coupled to a voltage via a voltage supply node, and a second terminal of the first resistor coupled to a first terminal of the first transistor; a second pulse width modulator configured to generate pulses based on the input of the differential circuit, the output of the second pulse width modulator configured to be selectively coupled to one of the gate of the first transistor of the differential circuit and the gate of the second transistor of the differential circuit; a second resistor of the differential circuit having a configurable resistance, a second terminal of the second resistor coupled to the voltage via the voltage supply node and coupled to the second terminal of the first resistor and coupled to the first terminal of the second transistor; and an analog-to-digital converter having an input coupled to the voltage supply node.
[0006] As used herein, "facilitating" an action includes performing the action, making the action easier, assisting in the performance of the action, or having the action performed. Thus, by way of example and not limitation, instructions executing on a processor may facilitate an action performed by semiconductor manufacturing equipment by sending appropriate data or commands to cause or assist in the performance of the action. An action is performed by an entity or combination of entities even if the actor facilitates the action by something other than performing the action.
[0007] The techniques disclosed herein can provide substantial beneficial technical effects. Some embodiments may not have these potential advantages, and these potential advantages are not necessarily required for all embodiments. By way of example only, and not by way of limitation, one or more embodiments may provide one or more of the following:
[0008] Implementation of complex-valued MACC operations in crossbar arrays;
[0009] A complex MACC calculation engine that enables the execution of complex-valued MACC operations in addition to conventional real-valued MACC operations.
[0010] More adaptability, more deep learning degrees of freedom, and wider generalization of the MACC computation engine;
[0011] Extending the energy efficiency and throughput per area advantages of traditional crossbar arrays to the complex domain;
[0012] Very low overhead pulse width modulation (PWM) rerouting;
[0013] the ability to handle complex-valued synaptic weights and complex-valued activations;
[0014] An implementation requiring only minor unit cell modifications and additional PWM routing; and
[0015] Usefulness for training deep neural networks (DNNs).
[0016] Some embodiments may not have these potential advantages, and these potential advantages are not necessarily required in all embodiments. These and other features and advantages will become apparent from the following detailed description of illustrative embodiments of the invention when read in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0017] The following drawings are presented by way of example only, and not by way of limitation, and like reference numerals (when used) refer to corresponding elements throughout the several views.
[0018] [Figure 1A] FIG. 2 is a circuit diagram for a first exemplary embodiment of a circuit for a MACC calculation engine, in accordance with an exemplary embodiment.
[0019] [Figure 1B] 1B is a circuit diagram for a first exemplary embodiment of a circuit for the MACC calculation engine of FIG. 1A configured for phase 2, which is the calculation of imaginary terms, in accordance with an exemplary embodiment.
[0020] [Figure 1C] FIG. 1 illustrates an exemplary representation of a conventional crossbar array.
[0021] [Figure 2A] FIG. 10 is a circuit diagram for a second exemplary embodiment of a circuit for a MACC calculation engine, in accordance with an exemplary embodiment.
[0022] [Figure 2B] 2B is a circuit diagram for a second exemplary embodiment of a circuit for the MACC calculation engine of FIG. 2A configured for phase 2, the calculation of imaginary terms, in accordance with an exemplary embodiment.
[0023] [Figure 3A] FIG. 10 is a circuit diagram for a third exemplary embodiment of a circuit for a MACC calculation engine, in accordance with an exemplary embodiment.
[0024] [Figure 3B] FIG. 3B is a circuit diagram for a third exemplary embodiment of a circuit for the MACC calculation engine of FIG. 3A configured for phase 2, which is the calculation of imaginary terms, in accordance with an exemplary embodiment.
[0025] [Figure 4A] FIG. 10 is a circuit diagram for a fourth exemplary embodiment of a circuit for a MACC calculation engine, in accordance with an exemplary embodiment.
[0026] [Figure 4B] FIG. 3C is a circuit diagram for a fourth exemplary embodiment of a circuit for the MACC calculation engine of FIG. 3B configured for phase 2, which is the calculation of imaginary terms, in accordance with an exemplary embodiment.
[0027] [Figure 5] FIG. 2 illustrates a high-level diagram of an implementation of a MACC calculation engine, according to an example embodiment.
[0028] [Figure 6] 1 illustrates a computing environment according to one embodiment of the present invention.
[0029] It should be understood that elements in the figures are shown for simplicity and clarity, and that common but well-understood elements that may be useful or necessary in commercially feasible embodiments may not be shown in order to lessen obstruction of the view of the illustrated embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0030] The principles of the invention are described herein in the context of exemplary embodiments. Moreover, it will be apparent to one skilled in the art, given the teachings herein, that numerous modifications to the illustrated embodiments can be made that are within the scope of the claims. Thus, no limitations with respect to the embodiments shown and described herein are intended or should be inferred.
[0031] In one or more embodiments, the crossbar array is advantageously adapted to perform complex-valued MACC operations. The resulting complex MACC computation engine can also be used to perform traditional real MACC operations. Thus, the disclosed embodiments represent a more generalized MACC computation engine that can handle both real-valued and complex-valued MACC operations, providing more degrees of freedom, greater adaptability, and broader generalization for deep learning. In one exemplary embodiment, the crossbar array is adapted to store complex-valued synaptic weights and activations in rectangular coordinates. Aspects of the disclosed technique extend the energy efficiency and throughput-per-area advantages of traditional crossbar arrays to the complex domain. In one exemplary embodiment, the MACC operations can be handled by introducing additional rerouting circuitry.
[0032] Multiplication operation
number
number
number
number
number
[0033] In one exemplary embodiment, the complex-valued MACC operation is implemented in a crossbar array. Circuitry is disclosed that allows for simultaneous computation of real and imaginary terms, thus requiring only two computation phases.
[0034] 1A is a circuit diagram of a first exemplary embodiment of a circuit 100 for a MACC computation engine, according to an exemplary embodiment. FIG. 1A shows the configuration for phase 1, the computation of real terms (where complex MACC is composed of four terms). The terms a, b, c, and d can be positive or negative, with c = c + - c - and d = d + - d - representing weights w and implemented by resistor pairs 136-1, 136-2, 136-3, 136-4 in conjunction with transistors 140-1, 140-2, 140-3, 140-4, such as field-effect transistors (FETs). As explained in more detail below, the gates of each transistor 140-1, 140-2 are driven by an a: magnitude signal, and the gates of each transistor 140-3, 140-4 are driven by a b: magnitude signal. In one embodiment, the drain of each transistor 140-1, 140-2, 140-3, and 140-4 is driven by a selected voltage determined by the sign of the corresponding term of activation, and the source of each transistor 140-1, 140-2, 140-3, and 140-4 is coupled to a resistor 136-1, 136-2, 136-3, and 136-4, respectively. Note that representative read voltages of 0.2V, 0.4V, and 0.6V are used as examples (other representative read voltages are contemplated).
[0035] During phase 1, 'a' input register 104-1 stores the activation a term, and 'b' input register 104-2 stores the activation b term. Pulse width modulators 108-1, 108-2 generate a-magnitude pulses and b-magnitude pulses, respectively, which are applied to the gates of circuit pairs 112-1, 112-2, respectively. Because the widths of the pulses are proportional to the activation a term and the activation b term, respectively, current flows through resistor pairs 136-1, 136-2, 136-3, 136-4 (c+, c- and d+, d-) for an amount of time proportional to the pulse width (and therefore proportional to the activation a term and the activation b term, respectively). Note that voltage supply node 132 (also referred to herein as node 132) is maintained at 0.4 volts (V) as shown, and the current indicated by the arrow flows in one direction for positive values of the activation term and in the opposite direction for negative values of the activation term. The amount of current flowing toward analog-to-digital converter (ADC) 116 is representative of the calculated real term, converted to a binary value by current-based ADC 116, and stored in output register 120-2. By concatenating nodes 132 of other circuit 100 (representing other activations and weights from other multiplication operations), the sum of the MACC operation can be input to ADC 116 and stored in output register 120-2. Note that output register 120-1 is described below in connection with FIG. 1B. Note that in one or more embodiments, the absolute value of each resistor pair 136-1, 136-2, 136-3, 136-4 is not as important as the shape of the distribution of resistance values. In one or more embodiments, the software-trained unitless weight distributions are rescaled by a certain scale factor to suitable hardware units (microsiemens or μS conductance) so that the shape of the weight distributions is not distorted. Those skilled in the art will be familiar with the use of pulse units and suitable controllers to set the weights (conductance) of memory stick elements to desired set, reset, and / or intermediate states. In one or more embodiments, activation is encoded as pulse width only, with no amplitude differences.Furthermore, in one or more embodiments, only the relative resistance values are important and it is only required to preserve the weight / resistance distribution. Furthermore, in one or more embodiments, the weight distribution is obtained from software (unitless) and is to be programmed in hardware (microsiemens). Therefore, a certain rescaling factor is applied such that the units change but the shape of the distribution remains the same.
[0036] FIG. 1B is a circuit diagram of a first exemplary embodiment of the circuit 100 for the MACC computation engine of FIG. 1A configured for phase 2, the computation of the imaginary terms, according to an exemplary embodiment. As shown in FIG. 1B, the application of the outputs of the 'a' input register 104-1 and the 'b' input register 104-2 to the pulse-width modulators 108-1, 108-2, the application of the a magnitude pulse and the b magnitude pulse to the gates of the circuit pairs 112-1, 112-2, and the application of the 'a' sign and the 'b' sign to the circuit pairs 112-1, 112-2 are inverted. As will be appreciated by those skilled in the art, in the example of FIG. 1B, a:sign and b:sign are analog voltages applied to the bottom terminals of the gating transistors and may be governed by a corresponding digital sign bit (not shown). (Note that the first voltage pair corresponds to a positive sign and the second voltage pair corresponds to a negative sign.) For the avoidance of doubt, in one or more embodiments, a sign bit is present. However, for simplicity of explanation, the figures show analog voltages. In one or more embodiments, there is some circuitry, omitted for simplicity of explanation, that converts this sign bit to corresponding analog voltages, shown throughout the figures as 0.2V and 0.6V. Those skilled in the art, given the teachings herein, will be able to implement such circuitry using known techniques.
[0037] As in Phase 1, 1) the width of the pulse is proportional to the activation a and b terms, respectively; 2) current flows through resistor pairs 136-1, 136-2, 136-3, and 136-4 (c+, c- and d+, d-) for an amount of time proportional to the pulse width (and thus proportional to the activation a and b terms, respectively); and the current indicated by the arrows flows in one direction for positive values of the activation terms and in the opposite direction for negative values of the activation terms. The amount of current flowing to ADC 116 is representative of the calculated imaginary term, which is converted to a binary value by ADC 116 and stored in output register 120-1. (Output register 120-2 holds the calculated real term during Phase 2, as indicated by the shaded area.) After Phase 2, the real and imaginary terms can be transported to the inputs of another array, passed through an activation function, and so on. It should be noted that instead of swapping the application of the outputs of the 'a' input register 104-1 and the 'b' input register 104-2 to the pulse width modulators 108-1, 108-2, the application of the outputs of the pulse width modulators 108-1, 108-2 to the gates may be swapped (applicable to all embodiments). Those skilled in the art will understand that the value of the first voltage pair of the sign bit corresponds to a positive activation term.
[0038] FIG. 1C shows an example representation of a conventional crossbar array. Input values V1 through VN generate output currents I1 through IM. Note that circuit 100 of FIGS. 1A and 1B provides support for complex conductance (G) values for crossbar array 300. Thus, in one or more embodiments, the G values are complex, and the inputs (activity) are encoded using voltages and pulse durations, and may be designated, for example, as complex input value 1, complex input value 2, and complex input value N, rather than V1, V2, ..., VN as in the prior art. Thus, FIG. 1C, now representing a standard MACC block or tile, relates to FIGS. 1A and 1B in that each input and conductance element is complex-valued in rectangular coordinates, and input activity and weight are now meant to each include two values (real and imaginary). As such, the high-level FIG. 1C relates to the more specific circuit implementation shown in 1A and 1B.
[0039] FIG. 2A is a circuit diagram for a second exemplary embodiment of a circuit 200 for a MACC calculation engine, according to an exemplary embodiment. FIG. 2A illustrates the configuration for Phase 1, the calculation of real terms, where the complex MAC is composed of four terms. Because terms a, b, c, and d are restricted to positive values, there is no sign bit or corresponding circuit. Because all terms are positive, each weight can be represented by a single resistor 136-1, 136-2, and current always flows through resistors 136-1, 136-2 in the same direction. Note that representative read voltages of 0.2V, 0.4V, and 0.6V are used as examples (other representative read voltages are also contemplated). Note that the 'b' sign is negative to implement a subtraction operation in the sum of the bidi terms.
[0040] During phase 1, the 'a' input register 104-1 stores the activation a term, and the 'b' input register 104-2 stores the activation b term. Multiplying two imaginary numbers results in -1. Thus, if j is an imaginary number, then j*j=-1 (by definition). For complex-valued MACC operations, see the equations above. Pulse width modulators 108-1, 108-2 generate a magnitude pulse and a b magnitude pulse, respectively, which are applied to the gates of circuit pair 212-1. Because the widths of the pulses are proportional to the activation a term and the activation b term, respectively, current flows through resistor pair 136-1, 136-2 (c and d) for an amount of time proportional to the pulse width (and therefore proportional to the activation a term and the activation b term, respectively). Note that node 132 is maintained at 0.4 volts (V) as shown, and the current indicated by the arrow flows in one direction for positive values of the activation term and in the opposite direction for negative values of the activation term. The amount of current flowing toward analog-to-digital converter (ADC) 116 is representative of the calculated real term, converted to a binary value by current-based ADC 116, and stored in output register 120-2. By concatenating nodes 132 of other circuits 100 (representing other activations and weights from other multiplication operations), the sum of the MACC operation can be input to ADC 116 and stored in output register 120-2. Note that output register 120-1 is described below in connection with FIG. 2B.
[0041] 2B is a circuit diagram of a second exemplary embodiment of the circuit 200 for the MACC calculation engine of FIG. 2A configured for Phase 2, the calculation of the imaginary term, according to an exemplary embodiment. As shown in FIG. 2B, the application of the outputs of 'a' input register 104-1 and 'b' input register 104-2 to pulse-width modulators 108-1, 108-2, and the application of the a-magnitude pulse and the b-magnitude pulse to the gate of circuit pair 212-1 are reversed. As in Phase 1, 1) the width of the pulse is proportional to the activation a term and the activation b term, respectively, and 2) current flows through resistor pair 136-1, 136-2 (c and d) for an amount of time proportional to the pulse width (and thus proportional to the activation a term and the activation b term, respectively). The amount of current flowing to ADC 116 is representative of the calculated imaginary term and is converted to a binary value by ADC 116 and stored in output register 120-1. (Output register 120-2 holds the calculated real terms during phase 2, as indicated by the shaded area.) After phase 2, the real and imaginary terms can be transported to the inputs of another array, passed through an activation function, and the like.
[0042] 3A is a circuit diagram for a third exemplary embodiment of a circuit 300 for a MACC calculation engine, according to an exemplary embodiment. FIG. 3A illustrates the configuration for phase 1, the calculation of real terms, where the complex MAC is composed of four terms. Terms a and b can be either positive or negative, while c and d are positive. Because weights are always positive, each weight can be represented by a single resistor 136-1, 136-2; however, because activation can be positive or negative, the direction of current through resistors 136-1, 136-2 is controlled by a sign bit. Note that representative read voltages of 0.2V, 0.4V, and 0.6V are used as examples (other representative read voltages are also contemplated).
[0043] During phase 1, 'a' input register 104-1 stores the activation a term, and 'b' input register 104-2 stores the activation b term. Pulse width modulators 108-1 and 108-2 generate a magnitude pulse and a magnitude pulse, respectively, which are applied to the gate of circuit pair 212-1. Because the width of the pulse is proportional to the activation a term and the activation b term, respectively, current flows through resistor pair 136-1 and 136-2 (c and d) for an amount of time proportional to the pulse width (and therefore proportional to the activation a term and the activation b term, respectively). Note that node 132 is maintained at 0.4 volts (V), as shown, and that current, indicated by the arrow, flows in one direction for positive values of the activation term and in the opposite direction for negative values of the activation term. The amount of current flowing towards analog-to-digital converter (ADC) 116 is representative of the calculated real term and is converted to a binary value by current-based ADC 116 and stored in output register 120-2. By concatenating nodes 132 (representing other activations and weights from other multiplication operations) of other circuits 100, the sum of the MACC operation can be input to ADC 116 and stored in output register 120-2. Note that output register 120-1 is described below in connection with FIG. 3B.
[0044] 3B is a circuit diagram of a third exemplary embodiment of circuit 300 for the MACC computation engine of FIG. 3A configured for phase 2, the computation of the imaginary terms, according to an exemplary embodiment. As with the embodiment of FIG. 1B, the application of the outputs of 'a' input register 104-1 and 'b' input register 104-2 to pulse width modulators 108-1, 108-2, the application of the a and b magnitude pulses to the gates of circuit pair 212-1, and the application of the 'a' and 'b' sign bits to circuit pair 212-1 are reversed. As in Phase 1, 1) the width of the pulse is proportional to the activation a and b terms, respectively; 2) current flows through resistor pairs 136-1, 136-2 (c+, c- and d+, d-) for an amount of time proportional to the pulse width (and thus proportional to the activation a and b terms, respectively); and the current flow, indicated by the arrows, flows in one direction for positive values of the activation terms and in the opposite direction for negative values of the activation terms. The amount of current flowing toward ADC 116 is representative of the calculated imaginary term, which is converted to a binary value by ADC 116 and stored in output register 120-1. (Output register 120-2 holds the calculated real term during Phase 2, as indicated by the shaded area.) After Phase 2, the real and imaginary terms can be transported to the inputs of another array, passed through an activation function, and so on.
[0045] 4A is a circuit diagram for a fourth exemplary embodiment of a circuit 400 for a MACC calculation engine, according to an exemplary embodiment. FIG. 4A illustrates the configuration for phase 1, the calculation of real terms where the complex MAC is composed of four terms. The terms a, b, c, and d can be positive or negative, where c' = c - csh and d' = d - dsh. In the embodiment of FIG. 4A, resistors 136-3, 136-4 of circuit pair 212-2 are shared by all circuit pairs 212-1 assigned to the same MACC operation and generate their own real and imaginary components via current-based ADC 128 for storage in output registers 124-2, 124-1, respectively. Similarly, resistors 136-1, 136-2 of circuit pair 212-1 together generate respective real and imaginary components for MACC calculation via current-based ADC 116 for storage in output registers 120-2, 120-1, respectively, during phases 1 and 2. Note that representative read voltages of 0.2V, 0.4V, and 0.6V are used as an example (other representative read voltages are contemplated).
[0046] 4B is a circuit diagram of a fourth example embodiment of circuit 400 for the MACC computation engine of FIG. 3B configured for phase 2, the computation of the imaginary terms, according to an example embodiment. Subsequently, in phase 2, the values stored in real output registers 120-2, 124-2 are summed to generate the final real component, and the values stored in real output registers 120-1, 124-1 are summed to generate the final imaginary component. At this point, the real and imaginary terms may be transported to the inputs of another array, passed through an activation function, and so on. In one or more embodiments, the sum is used to determine the final result.
[0047] FIG. 5 is a high-level diagram 500 of an implementation of a MACC computation engine according to an exemplary embodiment. In one exemplary embodiment, engine 500 includes an array 504 of circuits 508-1, 508-2, such as circuits 100, 200, 300, and 400. As described above, each of circuits 508-1, 508-2 receives an activity a1, b1 and is programmed with a weight c1, d1. As will be understood by those skilled in the art, in machine learning, weights represent learned values that govern the behavior of a neural network. To perform the computation, the imaginary component is calculated by inverting the sign and rerouting the pulses. In an exemplary embodiment, as described above, each complex x term is represented using two pulse trains, and each complex w term is represented using two non-volatile memory (NVM) devices. It should be noted that proof-of-concept experiments have been performed and satisfactory results obtained (the proof-of-concept work, which required training a complex-valued neural net, was carried out using a known machine learning framework).
[0048] One or more embodiments include suitable peripheral circuitry 512, 516 and a suitable controller 520 for input / output, weight programming, etc. The peripheral circuitry 512, 516 controls pulse-by-pulse programming, activity (including components a1, b1), and the like.
number
[0049] In one exemplary embodiment, an analog memory-based complex weight unit cell includes four source follower circuits, each having an analog memory device with one terminal connected to a source and another terminal connected to a common node NA between all analog memory elements held at a voltage VA (e.g., 0.4 V), where the conductances in the first two source follower circuits define the real part of the complex weight in different ways, and the conductances in the next two source follower circuits define the imaginary part of the complex weight in different ways. The use of source followers is preferred in one or more embodiments because it is desirable to maintain stable source voltages across a variety of different conductance values for operational calculations.
[0050] In one exemplary embodiment, a pulse duration proportional to the real part of the complex-valued input activation is applied to the gates of a first differential pair of source follower circuits representing the real part of the complex weight.
[0051] In one exemplary embodiment, a pulse duration proportional to the imaginary part of the complex-valued input activation is applied to the gates of the second differential pair of source follower circuits representing the imaginary part of the complex weight.
[0052] In one exemplary embodiment, the drain voltage of the source follower circuit is configurable to a range of voltages above and / or below VA.
[0053] In one exemplary embodiment, two or more unit cells are tied together at a common node NA.
[0054] In one exemplary embodiment, the current at the common node NA is measured by accumulating charge on a capacitor and measuring the charge using an analog-to-digital converter. In one exemplary embodiment, a pulse duration proportional to the real part of the activity is applied to a source follower gate corresponding to the imaginary part of the complex weight, and a pulse duration proportional to the imaginary part of the activity is applied to a source follower gate corresponding to the real part of the complex weight.
[0055] In light of the foregoing discussion, in general terms, an exemplary circuit according to an embodiment of the present invention may include a first pulse width modulator 108-1 configured to generate a first pulse based on a first input; a second pulse width modulator 108-2 configured to generate a second pulse based on a second input; a first differential circuit 112-1 including a first transistor 140-1, a second transistor 140-2, a first resistor 136-1, and a second resistor 136-2; a first differential circuit 112-2 including a first transistor 140-3, a second transistor 140-4, a first resistor 136-3, and a second resistor 136-4; 36-4, wherein a gate of the first transistor 140-1 of the first differential circuit 112-1 and a gate of the second transistor 140-2 of the first differential circuit 112-1; and a gate of the first transistor 140-3 of the second differential circuit 112-2 and a gate of the second transistor 140-4 of the second differential circuit 112-2 are configured to be controlled by the first pulse width modulator and the second pulse width modulator based on the first input and the second input; The first resistor 136-1 of the differential circuit 112-1 has a configurable resistance, a first terminal of the first resistor 136-1 is coupled to a voltage supply node 132, and a second terminal of the first resistor 136-1 is coupled to a first terminal of the first transistor 140-1 of the first differential circuit 112-1; the second resistor 136-2 of the first differential circuit 112-1 has a configurable resistance, and a first terminal of the second resistor 136-2 of the first differential circuit 112-1 is coupled to the voltage supply node 132, and a second terminal of the first resistor 136-1 is coupled to a first terminal of the first transistor 140-1 of the first differential circuit 112-1; a second terminal of the second resistor 136-2 of the first differential circuit 112-1 coupled to a first terminal of the second transistor 140-2 of the first differential circuit 112-1; a first resistor 136-3 of the second differential circuit 112-2 having a configurable resistance, a first terminal of the first resistor 136-3 of the second differential circuit 112-2 coupled to the voltage via the voltage supply node 132, and a second terminal of the first resistor 136-3 of the second differential circuit 112-2 coupled to a first terminal of the first transistor 140-3 of the second differential circuit 112-2;and the second resistor 136-4 of the second differential circuit 112-2 has a configurable resistance, a first terminal of the second resistor 136-4 of the second differential circuit 112-2 is coupled to the voltage via the voltage supply node 132, and a second terminal of the second resistor 136-4 of the second differential circuit 112-2 is coupled to a first terminal of the second transistor 104-4 of the second differential circuit 112-2; it will be understood that the circuit further includes an analog-to-digital converter 116 having an input coupled to the voltage supply node 132. As will be understood by those skilled in the art, the voltage supply node 132 is configured to be connected to a suitable voltage supply. During operation of the circuit, a voltage from the voltage supply is applied to the voltage supply node 132. Typically, a constant voltage is applied to the voltage supply node 132. The constant voltage may be supplied from a constant voltage supply or from a variable voltage supply configured to supply a constant voltage.
[0056] In one exemplary embodiment, the first resistor 136-1 and the second resistor 136-2 of the first differential circuit 112-1 are configured based on a first corresponding weight of the multiply-accumulate operation, and the first resistor 136-3 and the second resistor 136-4 of the second differential circuit 112-2 are configured based on a second corresponding weight of the multiply-accumulate operation. In one exemplary embodiment, the circuit further includes a first input register 104-1 coupled to an input of the first pulse-width modulator 108-1 and configured to store a first input; a second input register 104-2 coupled to an input of the second pulse-width modulator 108-2 and configured to store a second input; and a controller circuit configured to control the loading of the first input register 104-1 and the second input register 104-2.
[0057] In one exemplary embodiment, the circuit further comprises a real component output register 120-2, the input of which is coupled to the output of the analog-to-digital converter 116; an imaginary component output register 120-1, the input of which is coupled to the output of the analog-to-digital converter 116; and a controller circuit configured to control the loading of the real component output register 120-2 and the imaginary component output register 120-1.
[0058] In one exemplary embodiment, the second terminal of the first transistor 140-1 of the first differential circuit 112-1, the second terminal of the second transistor 140-2 of the first differential circuit 112-1, the second terminal of the first transistor 140-3 of the second differential circuit 112-2, and the second terminal of the second transistor 140-4 of the second differential circuit 112-2 are configured to be selectively coupled between a voltage representing the sign of the first input and a voltage representing the sign of the second input, and the circuit further comprises a controller circuit configured to control the selective coupling between the voltage representing the sign of the first input and the voltage representing the sign of the second input.
[0059] In one exemplary embodiment, analog-to-digital converter 116 comprises one or more current mirrors and one or more capacitive elements configured to accumulate charge from voltage supply node 132 .
[0060] In one exemplary embodiment, a first terminal of a first transistor 140-1 of the first differential circuit 112-1, a first terminal of a second transistor 140-2 of the first differential circuit 112-1, a first terminal of a first transistor 140-3 of the second differential circuit 112-2, a first terminal of a second transistor 140-4 of the second differential circuit 112-2, a second terminal of the first transistor 140-1 of the first differential circuit 112-1, The second terminal of the second transistor 140-2 of the second differential circuit 112-1, the second terminal of the first transistor 140-3 of the second differential circuit 112-2, and the second terminal of the second transistor 140-4 of the second differential circuit 112-2 each correspond to one of the source terminal and the drain terminal; i.e., each of the (field effect) transistors 140-1, 140-2, 140-3, 140-4 has a source terminal and a drain terminal.
[0061] In one embodiment, the circuit 200 includes a differential circuit 212-1 including a first transistor 140-1, a second transistor 140-2, a first resistor 136-1, and a second resistor 136-2; a first pulse width modulator 108-1 configured to generate pulses based on a first input, an output of the first pulse width modulator 108-1 configured to be selectively coupled to one of a gate of the first transistor 140-1 of the differential circuit 212-1 and a gate of the second transistor 140-2 of the differential circuit 212-1; a first resistor 108-1 having a configurable resistance, a first terminal of the first resistor 136-1 coupled to a voltage via a voltage supply node 132, and a second terminal of the first resistor 136-1 coupled to the first terminal of the first transistor 140-1. a second pulse width modulator 108-2 configured to generate pulses based on the output of the second pulse width modulator 108-2 configured to be selectively coupled to one of a second input, a gate of the first transistor 140-1 of the differential circuit 212-1 and a gate of the second transistor 140-2 of the differential circuit 212-1; a configurable resistance, a second resistor 136-2 of the differential circuit 212-1 coupled to a voltage via the voltage supply node 132, and having a second terminal of the second resistor 136-2 coupled to the second terminal of the first resistor 136-1 and the second terminal of the second resistor 136-2 coupled to the first terminal of the second transistor 140-2; and an analog-to-digital converter 116 having an input coupled to the voltage supply node 132.
[0062] In one exemplary embodiment, the first resistor 136-1 of the first differential circuit 212-1 is configured based on a first corresponding weight of the multiply-accumulate operation, and the second resistor 136-2 of the first differential circuit 212-1 is configured based on a second corresponding weight of the multiply-accumulate operation. In one exemplary embodiment, the circuit further includes a first input register 104-1 coupled to an input of the first pulse-width modulator 108-1 and configured to store the first input, a second input register 104-2 coupled to an input of the second pulse-width modulator 108-2 and configured to store the second input, and a controller circuit configured to control the loading of the first input register 104-1 and the second input register 104-2.
[0063] In one exemplary embodiment, the circuit further includes a real component output register 120-2 having an input coupled to the output of the analog-to-digital converter 116, an imaginary component output register 120-1 having an input coupled to the output of the analog-to-digital converter 116, and a controller circuit configured to control the loading of the real component output register 120-2 and the imaginary component output register 120-1. In one exemplary embodiment, the second terminal of the first transistor 140-1 and the second terminal of the second transistor 140-2 are configured to be selectively coupled between a voltage input representing the sign bit of the first input and a voltage input representing the sign bit of the second input, and the circuit further includes a controller circuit configured to control the selective coupling between the voltage input representing the sign of the first input and the voltage input representing the sign of the second input. (See FIGS. 3A-3B.)
[0064] In one exemplary embodiment, the first terminal of the first transistor 140-1, the first terminal of the second transistor 140-2 of the first differential circuit 112-1, the second terminal of the first transistor 140-1, and the second terminal of the second transistor 140-2 each correspond to one of a source terminal and a drain terminal. In one exemplary embodiment, the circuit further comprises a first shared transistor 140-3 and a second shared transistor 140-4, where the first shared resistor 136-3 has a configurable resistance, a first terminal of the first shared resistor 136-3 coupled to a voltage via a voltage supply node 132, and a second terminal of the first shared resistor 136-3 coupled to the first terminal of the first shared transistor 140-3; a second terminal of the second shared resistor 136-4 coupled to a first terminal of the shared transistor 140-4; and a second analog-to-digital converter 128 having an input coupled to the voltage supply node 132; wherein the output of the first pulse-width modulator 108-1 is configured to be selectively coupled to one of the gates of the first shared transistor 140-3 and the second shared transistor 140-4, and the output of the second pulse-width modulator 108-2 is configured to be selectively coupled to one of the gates of the first shared transistor 140-3 and the second shared transistor 140-4 (see FIGS. 4A-4B).
[0065] In one exemplary embodiment, the circuit further includes a real component shared output register 124-2 having an input coupled to the output of the analog-to-digital converter 128, and an imaginary component shared output register 124-1 having an input coupled to the output of the analog-to-digital converter 128. In one exemplary embodiment, the second terminal of the first shared transistor 140-3 and the second terminal of the second shared transistor 140-4 are configured to be selectively coupled between a voltage input representing the sign bit of the first input and a voltage input representing the sign bit of the second input. In one exemplary embodiment, the first terminal of the first shared transistor 140-3, the first terminal of the second shared transistor 140-4, the second terminal of the first shared transistor 140-3, and the second terminal of the second shared transistor 140-4 each correspond to one of the source terminal and the drain terminal.
[0066] Various aspects of the present disclosure are described by text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in embodiments of a computer program product (CPP). For any flowchart, depending on the technology involved, operations may be performed in an order different from that shown in a given flowchart. For example, again depending on the technology involved, two operations shown in successive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or in an at least partially overlapping manner.
[0067] A computer program product embodiment ("CPP embodiment" or "CPP") is a term used in this disclosure to describe any set of one or more storage media (also referred to as "media") collectively contained in a set of one or more storage devices that collectively contain machine-readable code corresponding to instructions and / or data for performing the computer operations specified in a given CPP claim. A "storage device" is any tangible device that can hold and store instructions for use by a computer processor. The computer-readable storage medium may be, but is not limited to, an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded device (such as pits / lands formed on the major surface of a punch card or disk), or any suitable combination of the foregoing. Computer-readable storage media, as the term is used in this disclosure, is not to be construed as storage in the form of a transient signal per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, light pulses passing through fiber optic cables, electrical signals communicated over wires, and / or other transmission media. As will be appreciated by those skilled in the art, data is typically moved at some infrequent time during the normal operation of a storage device, such as during access, defragmentation, or garbage collection, but the foregoing does not qualify a storage device as transient because the data is not transient while it is stored.
[0068] Reference is now made to FIG. 6. Computing environment 100 includes an example of an environment for execution of at least a portion of computer code 200 involved in performing the method of the present invention, such as training unitless weight distributions in software using forward and back propagation for neural networks and rescaling the same so that suitable conductances in microsiemens or μS can be programmed into suitable hardware units in accordance with embodiments of the present invention, as is well known to those skilled in the art; the code may also be used for control aspects or synthesize digital circuits to implement control aspects in a known manner. Also, the analog MAC engine circuit disclosed herein can be used within a general-purpose machine, such as the one shown, as a hardware accelerator. In addition to block 200, computing environment 100 includes, for example, computer 101, wide area network (WAN) 102, end user device (EUD) 103, remote server 104, public cloud 105, and private cloud 106. In this embodiment, computer 101 includes a set of processors 110 (including processing circuitry 120 and cache 121), a communications fabric 111, volatile memory 112, persistent storage 113 (including operating system 122 and block 200 shown above), a set of peripheral devices 114 (including a set of user interface (UI) devices 123, storage 124, and a set of Internet of Things (IoT) sensors 125), and a network module 115. Remote server 104 includes a remote database 130. Public cloud 105 includes a gateway 140, a cloud orchestration module 141, a set of host physical machines 142, a set of virtual machines 143, and a set of containers 144.
[0069] Computer 101 may take the form of a desktop computer, a laptop computer, a tablet computer, a smartphone, a smartwatch or other wearable computer, a mainframe computer, a quantum computer, or any other form of computer or mobile device now known or later developed that is capable of executing programs, accessing a network, or querying a database, such as remote database 130. As is well understood in the field of computer technology, and depending on the technology, execution of a computer-implemented method may be distributed among multiple computers and / or among multiple locations. However, in this description of computing environment 100, for purposes of brevity, the detailed discussion focuses on a single computer, specifically computer 101. Computer 101 may be located in a cloud, although it is not depicted in FIG. 6 within the cloud. However, computer 101 is not required to reside in a cloud except to any extent that may be expressly indicated.
[0070] Processor set 110 includes one or more computer processors of any type now known or later developed. Processing circuitry 120 may be distributed across multiple packages, e.g., multiple linked integrated circuit chips. Processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. Cache 121 is memory located within the processor chip package and is typically used for data or code that should be available for fast access by threads or cores executing on processor set 110. Cache memory is typically organized into multiple levels depending on relative proximity to the processing circuitry. Alternatively, some or all caches for a processor set may be located “off-chip.” In some computing environments, processor set 110 may be designed to operate with qubits and perform quantum computations.
[0071] Computer-readable program instructions are typically loaded onto computer 101 and cause processor set 110 of computer 101 to perform a series of operational steps, thereby implementing a computer-implemented method. As a result, the instructions so executed instantiate the method set forth in the flowcharts and / or descriptions of the computer-implemented method (collectively, the "methods of the present invention") contained herein. These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by processor set 110 to control and direct the execution of the methods of the present invention. In computing environment 100, at least some of the instructions for executing the methods of the present invention may be stored in block 200 within persistent storage 113.
[0072] Communications fabric 111 is the signal-conducting pathway that allows various components of computer 101 to communicate with one another. Typically, this fabric is made of switches and conductive pathways, such as switches and conductive pathways that make up buses, bridges, physical input / output ports, etc. Other types of signal communication pathways may be used, such as fiber optic and / or wireless communication pathways.
[0073] Volatile memory 112 may be any type of volatile memory now known or later developed. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory 112 is characterized by random access, although this is not required unless expressly indicated. In computer 101, volatile memory 112 is located in a single package and is internal to computer 101; however, alternatively or additionally, volatile memory may be distributed across multiple packages and / or located external to computer 101.
[0074] Persistent storage 113 is any form of non-volatile storage for a computer, now known or later developed. The non-volatility of this storage means that stored data is maintained regardless of whether power is supplied to computer 101 and / or directly to persistent storage 113. While persistent storage 113 may be read-only memory (ROM), typically at least a portion of persistent storage allows data to be written, data to be deleted, and data to be rewritten. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. Operating system 122 may take several forms, such as various known proprietary operating systems employing a kernel or open-source Portable Operating System Interface-type operating systems. The code contained in block 200 typically includes at least some of the computer code involved in performing the method of the present invention.
[0075] Peripheral device set 114 includes a set of peripheral devices of computer 101. Data communication connections between peripheral devices and other components of computer 101 may be implemented in various forms, such as Bluetooth connections, near field communication (NFC) connections, connections made by cables (such as universal serial bus (USB)-type cables), insertion-type connections (e.g., Secure Digital (SD) cards), connections made through local area communication networks, and even connections made through wide area networks such as the Internet. In various embodiments, UI device set 123 may include components such as display screens, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. Storage 124 is external storage, such as an external hard drive, or insertable storage, such as an SD card. Storage 124 may be persistent and / or volatile. In some embodiments, storage 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where computer 101 is required to have a large amount of storage (e.g., where computer 101 stores and manages large databases locally), this storage may be provided by a peripheral storage device designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple, geographically distributed computers. IoT sensor set 125 is made up of sensors that may be used in Internet of Things applications. For example, one sensor may be a thermometer and another sensor may be a motion detector.
[0076] Network module 115 is a collection of computer software, hardware, and firmware that enables computer 101 to interact with other computers over WAN 102. Network module 115 may include hardware such as a modem or Wi-Fi signal transceiver, software for packetizing and / or depacketizing data for communication network transmission, and / or web browser software for communicating data over the Internet. In some embodiments, the network control and network forwarding functions of network module 115 are performed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing Software-Defined Networking (SDN)), the control and forwarding functions of network module 115 are performed on physically separate devices, such that the control function manages several different network hardware devices. Computer-readable program instructions for performing the methods of the present invention may be downloaded to computer 101 from an external computer or external storage device, typically via a network adapter card or network interface included in network module 115.
[0077] WAN 102 is any now known or later developed wide area network (e.g., the Internet) capable of communicating computer data between remote locations by any technology for communicating computer data. In some embodiments, WAN 102 may be replaced and / or supplemented by a local area network (LAN) designed to communicate data between devices located in a local area, such as a Wi-Fi network. WANs and / or LANs typically include copper transmission cables, optical fiber transmissions, wireless transmissions, and computer hardware such as routers, firewalls, switches, gateway computers, and edge servers.
[0078] End-user device (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of the enterprise operating computer 101) and may take any of the forms described above with respect to computer 101. EUD 103 typically receives useful and useful data from the operation of computer 101. For example, in the hypothetical case where computer 101 is designed to provide recommendations to end users, the recommendations would typically be communicated from computer 101's network module 115 over WAN 102 to EUD 103. In this manner, EUD 103 can display or otherwise present the recommendations to the end user. In some embodiments, EUD 103 may be a client device such as a thin client, a heavy client, a mainframe computer, a desktop computer, and the like.
[0079] Remote server 104 is any computer system that provides at least some data and / or functionality to computer 101. Remote server 104 may be controlled and used by the same entity that operates computer 101. Remote server 104 represents a machine that collects and stores useful and useful data for use by other computers, such as computer 101. For example, in the hypothetical case where computer 101 is designed and programmed to provide recommendations based on historical data, then this historical data may be provided to computer 101 from remote database 130 of remote server 104.
[0080] A public cloud 105 is any computer system available for use by multiple entities that provides on-demand availability of computer system resources and / or other computer capacity, particularly data storage (cloud storage) and computing capacity, without requiring direct, active management by users. Cloud computing typically leverages resource sharing to achieve coherence and economies of scale. Direct, active management of the public cloud 105's computing resources is performed by computer hardware and / or software in a cloud orchestration module 141. The computing resources provided by the public cloud 105 are typically implemented by virtual computing environments running on various computers comprising a host physical machine set 142, which is the universe of physical computers within and / or available to the public cloud 105. The virtual computing environments (VCEs) typically take the form of virtual machines from a virtual machine set 143 and / or containers from a container set 144. It is understood that these VCEs may be stored as images and transferred among and between various hosts of physical machines either as images or after instantiation of the VCE. Cloud orchestration module 141 manages the transfer and storage of images, deploys new instantiations of VCEs, and manages active instantiations of VCE deployments. Gateway 140 is a collection of computer software, hardware, and firmware that enables public cloud 105 to communicate over WAN 102.
[0081] We now provide some further explanation of virtual computing environments (VCEs). A VCE can be stored as an "image." A new, active instance of a VCE can be instantiated from the image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to a feature of an operating system in which the kernel allows the existence of multiple isolated user space instances called containers. These isolated user space instances typically behave as actual computers from the perspective of the programs running within them. A computer program running on a typical operating system can utilize all of the computer's resources, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and the devices assigned to the container; this feature is known as containerization.
[0082] Private cloud 106 is similar to public cloud 105, except that the computing resources are available only for use by a single enterprise. While private cloud 106 is shown in communication with WAN 102, in other embodiments, the private cloud may be completely disconnected from the Internet and accessible only through a local / private network. A hybrid cloud is a composite of multiple clouds of different types (e.g., private, community, or public cloud types), often implemented by different vendors. While each of the multiple clouds remains a separate, discrete entity, the larger hybrid cloud architecture is linked by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the constituent clouds. In this embodiment, both public cloud 105 and private cloud 106 are part of a larger hybrid cloud.
[0083] The description of various embodiments of the present invention has been presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terms used herein have been selected to best explain the principles, practical applications, or technical improvements of the embodiments over technologies found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A circuit comprising: a first pulse width modulator configured to generate a first pulse based on a first input; a second pulse width modulator configured to generate a second pulse based on a second input; a first differential circuit having a first transistor, a second transistor, a first resistor, and a second resistor; a second differential circuit having a first transistor, a second transistor, a first resistor, and a second resistor; Equipped with a gate of the first transistor of the first differential circuit and a gate of the second transistor of the first differential circuit; and a gate of the first transistor of the second differential circuit and a gate of the second transistor of the second differential circuit are configured to be controlled by the first pulse width modulator and the second pulse width modulator based on the first input and the second input; the first resistor of the first differential circuit has a configurable resistance, a first terminal of the first resistor is coupled to a voltage supply node, and a second terminal of the first resistor is coupled to a first terminal of the first transistor of the first differential circuit; the second resistor of the first differential circuit has a configurable resistance, a first terminal of the second resistor of the first differential circuit is coupled to the voltage via the voltage supply node, and a second terminal of the second resistor of the first differential circuit is coupled to a first terminal of the second transistor of the first differential circuit; the first resistor of the second differential circuit has a configurable resistance, a first terminal of the first resistor of the second differential circuit is coupled to the voltage via the voltage supply node, and a second terminal of the first resistor of the second differential circuit is coupled to a first terminal of the first transistor of the second differential circuit; and the second resistor of the second differential circuit has a configurable resistance, a first terminal of the second resistor of the second differential circuit is coupled to the voltage via the voltage supply node, and a second terminal of the second resistor of the second differential circuit is coupled to a first terminal of the second transistor of the second differential circuit; the circuit further comprises an analog-to-digital converter having an input coupled to the voltage supply node; circuit.
2. 2. The circuit of claim 1, wherein the first resistor and the second resistor of the first differential circuit are configured based on a first corresponding weight of a multiply-accumulate operation, and the first resistor and the second resistor of the second differential circuit are configured based on a second corresponding weight of the multiply-accumulate operation.
3. a first input register coupled to an input of the first pulse width modulator and configured to store the first input; a second input register coupled to an input of the second pulse width modulator and configured to store the second input; and a controller circuit configured to control the loading of the first input register and the second input register; The circuit of claim 1 further comprising:
4. a real component output register, an input of the real component output register being coupled to an output of the analog-to-digital converter; an imaginary component output register, the input of which is coupled to the output of the analog-to-digital converter; and a controller circuit configured to control the loading of the real component output register and the imaginary component output register; The circuit of claim 1 further comprising:
5. 2. The circuit of claim 1, wherein the second terminal of the first transistor of the first differential circuit, the second terminal of the second transistor of the first differential circuit, the second terminal of the first transistor of the second differential circuit, and the second terminal of the second transistor of the second differential circuit are configured to be selectively coupled between a voltage representing a sign of the first input and a voltage representing a sign of the second input, the circuit further comprising a controller circuit configured to control the selective coupling between the voltage representing the sign of the first input and the voltage representing the sign of the second input.
6. 2. The circuit of claim 1, wherein the analog-to-digital converter comprises one or more current mirrors and one or more capacitive elements configured to accumulate charge from the voltage supply node.
7. 2. The circuit of claim 1, wherein the first terminal of the first transistor of the first differential circuit, the first terminal of the second transistor of the first differential circuit, the first terminal of the first transistor of the second differential circuit, the first terminal of the second transistor of the second differential circuit, the second terminal of the first transistor of the first differential circuit, the second terminal of the second transistor of the first differential circuit, the second terminal of the first transistor of the second differential circuit, and the second terminal of the second transistor of the second differential circuit are each a corresponding one of a source terminal and a drain terminal.
8. a differential circuit having a first transistor, a second transistor, a first resistor, and a second resistor; a first pulse width modulator configured to generate pulses based on a first input, an output of the first pulse width modulator configured to be selectively coupled to one of a gate of the first transistor of the differential circuit and a gate of the second transistor of the differential circuit; the first resistor of the differential circuit having a configurable resistance, a first terminal of the first resistor coupled to a voltage supply node, and a second terminal of the first resistor coupled to a first terminal of the first transistor; a second pulse width modulator configured to generate pulses based on a second input, the output of the second pulse width modulator configured to be selectively coupled to one of the gate of the first transistor of the differential circuit and the gate of the second transistor of the differential circuit; the second resistor of the differential circuit having a configurable resistance, the second terminal of the second resistor coupled to the voltage via the voltage supply node and to the second terminal of the first resistor, and the second terminal of the second resistor coupled to the first terminal of the second transistor; and an analog-to-digital converter having an input coupled to the voltage supply node; A circuit comprising:
9. 9. The circuit of claim 8, wherein the first resistor of the first differential circuit is configured based on a first corresponding weight of a multiply-accumulate operation, and the second resistor of the first differential circuit is configured based on a second corresponding weight of the multiply-accumulate operation.
10. a first input register coupled to an input of the first pulse width modulator and configured to store the first input; a second input register coupled to an input of the second pulse width modulator and configured to store the second input; and a controller circuit configured to control the loading of the first input register and the second input register; The circuit of claim 8 further comprising:
11. a real component output register having an input coupled to the output of the analog-to-digital converter; an imaginary component output register having an input coupled to the output of the analog-to-digital converter; and a controller circuit configured to control the loading of the real component output register and the imaginary component output register; The circuit of claim 8 further comprising:
12. 9. The circuit of claim 8, wherein the second terminal of the first transistor and the second terminal of the second transistor are configured to be selectively coupled between a voltage input representing a sign bit of the first input and a voltage input representing a sign bit of the second input, the circuit further comprising: a controller circuit configured to control the selective coupling between the voltage input representing the sign of the first input and the voltage input representing the sign of the second input.
13. 9. The circuit of claim 8, wherein the first terminal of the first transistor, the first terminal of the second transistor of the first differential circuit, the second terminal of the first transistor, and the second terminal of the second transistor each correspond to one of a source terminal and a drain terminal.
14. a first shared transistor; a second shared transistor; the first shared resistor having a configurable resistance, a first terminal of the first shared resistor coupled to the voltage via the voltage supply node, and a second terminal of the first shared resistor coupled to a first terminal of the first shared transistor; the second shared resistor having a configurable resistance, a first terminal of the second shared resistor coupled to the voltage supply node and a second terminal of the second shared resistor coupled to a first terminal of the second shared transistor; and a second analog-to-digital converter having an input coupled to the voltage supply node; Furthermore, the output of the first pulse width modulator is configured to be selectively coupled to one of a gate of the first shared transistor and a gate of the second shared transistor; and The output of the second pulse width modulator is configured to be selectively coupled to one of the gate of the first shared transistor and the gate of the second shared transistor.
9. The circuit of claim 8.
15. a real component shared output register having an input coupled to the output of the analog-to-digital converter; and an imaginary component shared output register having an input coupled to the output of the analog-to-digital converter; The circuit of claim 14 further comprising:
16. 15. The circuit of claim 14, wherein the second terminal of the first shared transistor and the second terminal of the second shared transistor are configured to be selectively coupled between a voltage input representing a sign bit of the first input and a voltage input representing a sign bit of the second input.
17. 15. The circuit of claim 14, wherein the first terminal of the first shared transistor, the first terminal of the second shared transistor, the second terminal of the first shared transistor, and the second terminal of the second shared transistor each correspond to one of a source terminal and a drain terminal.