Complex product accumulation (MACC) computation engine based on analog memory
By designing a complex MACC calculation engine based on analog memory, the problem of difficulty in performing complex MACC operations in the prior art is solved, the support for complex values and flexibility in deep learning is achieved, and the energy efficiency advantages of crossbar arrays are expanded.
Patent Information
- Application Number
- CN202380074760.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-31
- Filing Date
- 2023-09-06
- Publication Date
- 2025-06-06
AI Technical Summary
The existing crossbar array is mainly used for MACC operations of real-valued numbers, and it is difficult to effectively perform complex MACC operations.
A complex multiplication accumulation (MACC) calculation engine based on analog memory is designed to support complex values by configuring a pulse width modulator and a differential circuit.
The ability to perform complex MACC operations in crossbar arrays is realized, providing more freedom and flexibility for deep learning, and extending the energy efficiency and throughput per unit area of conventional crossbar arrays to the complex domain.
Smart Images

Figure CN120113153A_ABST
Abstract
Description
Background Art
[0001] The present invention relates generally to electrical, electronic and computer technology, and more particularly to computing engines.
[0002] In machine learning, multiply-accumulate (MACC) operations usually use only real-valued numbers. However, complex-valued neural networks are gaining more research attention. For example, a common application provides automatic differentiation to implement the backpropagation algorithm, and now supports multiple complex layers and activation functions.
[0003] In addition, crossbar arrays are conventionally used to implement neural networks and numerical computing units because they provide efficient in-memory computing. Using the synapses and artificial behaviors of neurons provided by the neural network implemented by the crossbar array, supervised learning and unsupervised learning can be achieved. The crossbar array also supports the execution of MACC operations based on the physical principles of Ohm's law and Kirchhoff's law. Although the crossbar array can accelerate the execution of MACC operations, they usually only focus on real-valued numbers. Summary of the invention
[0004] The principles of the present invention provide an analog memory-based complex multiply-accumulate (MACC) computation engine. In one aspect, an exemplary circuit includes: a first pulse width modulator configured to generate a first pulse based on a first input; a second pulse width modulator configured to generate a second pulse based on a second input; a first differential circuit including a first transistor, a second transistor, a first resistor, and a second resistor; a second differential circuit including a first transistor, a second transistor, a first resistor, and a second resistor, wherein a gate of the first transistor of the first differential circuit and a gate of the second transistor of the first differential circuit, and a gate of the first transistor of the second differential circuit and a gate of the second transistor of the second differential circuit, are configured to be controlled by the first pulse width modulator and the second pulse width modulator based on the first input and the second input, wherein: the first resistor of the first differential circuit has a configurable resistance, a first end of the first resistor is coupled to a voltage via a voltage supply node, and a second end of the first resistor is coupled to a first resistor of the first differential circuit. a first terminal of a transistor of the first differential circuit; a second resistor of the first differential circuit having a configurable resistance, a first terminal of the second resistor of the first differential circuit being coupled to a voltage via a voltage supply node, and a second terminal of the second resistor of the first differential circuit being coupled to a first terminal of a second transistor of the first differential circuit; a first resistor of the second differential circuit having a configurable resistance, a first terminal of the first resistor of the second differential circuit being coupled to a voltage via a voltage supply node, and a second terminal of the first resistor of the second differential circuit being coupled to a first terminal of a first transistor of the second differential circuit; and a second resistor of the second differential circuit having a configurable resistance, a first terminal of the second resistor of the second differential circuit being coupled to a voltage via a voltage supply node, and a second terminal of the second resistor of the second differential circuit being coupled to a first terminal of a second transistor of the second differential circuit; and the circuit further comprises an analog-to-digital converter having an input coupled to the voltage supply node.
[0005] In one aspect, a circuit includes: a differential circuit including a first transistor, a second transistor, a first resistor, and a second resistor; a first pulse width modulator configured to generate a pulse based on a first input, an output of the first pulse width modulator configured to be selectively coupled to one of a gate of the first transistor of the differential circuit and a gate of the second transistor of the differential circuit, the differential circuit first resistor having a configurable resistance, a first end of the first resistor coupled to a voltage via a voltage supply node, and a second end of the first resistor coupled to a first end of the first transistor; a second pulse width modulator configured to generate a pulse based on a second input, an output of the second pulse width modulator configured to be selectively coupled to one of a gate of the first transistor of the differential circuit and a gate of the second transistor of the differential circuit; the second resistor of the differential circuit having a configurable resistance, a second end of the second resistor coupled to a voltage via a voltage supply node and coupled to a second end of the first resistor, and a second end of the second resistor coupled to the first end of the second transistor; and an analog-to-digital converter having an input coupled to the voltage supply node.
[0006] As used herein, "facilitating" an action includes performing the action, making the action easier, assisting in performing the action, or causing the action to be performed. Thus, by way of example and not limitation, instructions executed on a processor may facilitate the performance of an action by semiconductor manufacturing equipment by sending appropriate data or commands to cause or assist the action to be performed. When a party performing an action facilitates an action by means other than performing the action, the action is still performed by some entity or combination of entities.
[0007] The technology disclosed herein can provide significant beneficial technical effects. Some embodiments may not have these potential advantages, and these potential advantages are not necessarily required for all embodiments. By way of example only and not limitation, one or more embodiments may provide one or more of the following:
[0008] Implementing complex-valued MACC operations in a crossbar array;
[0009] A complex MACC calculation engine capable of performing complex-valued MACC operations as well as conventional real-valued MACC operations;
[0010] Increased flexibility, more freedom for deep learning, and wider generalization of the MACC computation engine;
[0011] Extend the energy efficiency and throughput per unit area advantages of conventional crossbar arrays to the complex domain;
[0012] Ultra-low overhead pulse width modulation (PWM) rerouting;
[0013] The ability to handle complex-valued synaptic weights and complex-valued activations;
[0014] An implementation that requires only minor modifications to the unit cell and additional PWM routing; and
[0015] Practicality for training deep neural networks (DNNs).
[0016] Some embodiments may not have these potential advantages, and these potential advantages are not necessarily required by all embodiments. These and other features and advantages will become apparent from the following detailed description of the exemplary embodiments which is to be read in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The following drawings are presented by way of example only and not limitation, wherein like reference numerals (when used) designate corresponding elements throughout the several views, and wherein:
[0018] Figure 1A is a circuit diagram of a first example embodiment of a circuit for a MACC calculation engine according to an example embodiment;
[0019] Figure 1B is configured according to an example embodiment for use in stage 2, calculating the imaginary term Figure 1A A circuit diagram of a first exemplary embodiment of a circuit for a MACC calculation engine;
[0020] Figure 1C An example representation of a conventional crossbar array is shown;
[0021] Figure 2A is a circuit diagram of a second example embodiment of a circuit for a MACC calculation engine according to an example embodiment;
[0022] Figure 2B is configured according to an example embodiment for use in stage 2, calculating the imaginary term Figure 2A A circuit diagram of a second exemplary embodiment of a circuit for a MACC calculation engine;
[0023] Figure 3A is a circuit diagram of a third example embodiment of a circuit for a MACC calculation engine according to an example embodiment;
[0024] Figure 3B is configured according to an example embodiment for use in stage 2, calculating the imaginary term Figure 3A A circuit diagram of a third exemplary embodiment of a circuit for a MACC calculation engine in FIG.
[0025] Figure 4A is a circuit diagram of a fourth example embodiment of a circuit for a MACC calculation engine according to an example embodiment;
[0026] Figure 4Bis configured according to an example embodiment for use in stage 2, calculating the imaginary term Figure 3B A circuit diagram of a fourth example embodiment of a circuit for a MACC calculation engine in FIG.
[0027] Figure 5 is a high-level diagram of an implementation of a MACC calculation engine according to an example embodiment; and
[0028] Figure 6 A computing environment according to an embodiment of the invention is described.
[0029] It should be appreciated that elements in the figures are illustrated for simplicity and clarity. Common but well-understood elements that may be useful or necessary in a commercially feasible embodiment may not be shown to facilitate a less obstructed view of the illustrated embodiments. DETAILED DESCRIPTION
[0030] The inventive principles described herein will be in the context of exemplary embodiments. In addition, it will be apparent to those skilled in the art, based on the teachings herein, that various modifications may be made to the illustrated embodiments within the scope of the claims. That is, no limitation to the embodiments illustrated and described herein is intended or should be inferred.
[0031] In one or more embodiments, the crossbar array is advantageously adjusted to perform MACC operations of complex-valued numbers. The resulting complex MACC computing engine can also be used to perform conventional real-valued MACC operations. Therefore, the disclosed embodiments represent a more generalized MACC computing engine that can handle both real-valued and complex-valued MACC operations and provide more degrees of freedom, increased flexibility, and wider generalization for deep learning. In an example embodiment, the crossbar array is adjusted to store complex-valued synaptic weights and activations in a rectangular coordinate system. Aspects of the disclosed technology extend the energy efficiency and unit area throughput advantages of conventional crossbar arrays to the complex domain. In an example embodiment, the MACC operations can be processed by introducing additional rerouting circuits.
[0032] Perform complex-valued MACC operations, such as multiplication The summation of the results usually requires four different MACC terms (where, for example, By the real part a n and the imaginary part jb n Respectively, and By the real part c n and imaginary part jd n Respectively), to combine into the real part and imaginary part of the output MACC:
[0033]
[0034] in:
[0035]
[0036] Each of these consists of a real term R e and the imaginary term I m In the above equation, x represents activation and w represents weight.
[0037] In one exemplary embodiment, the complex-valued MACC operation is implemented in a crossbar array. A circuit is disclosed that is capable of computing real and imaginary terms simultaneously, thereby requiring only two computational stages.
[0038] Figure 1A is a circuit diagram of a first example embodiment of a circuit 100 for a MACC calculation engine according to example embodiments. Figure 1A The configuration for stage 1 of computing real terms is shown, where the complex MACC consists of 4 terms. The terms a, b, c, and d can be positive or negative, where c = c + -c - and d = d + -d - Represents weight w, and is implemented by resistor pairs 136-1, 136-2, 136-3, 136-4 together with transistors 140-1, 140-2, 140-3, 140-4 (such as field effect transistors (FETs)). As described in more detail below, the gate of each transistor 140-1, 140-2 is driven by an a: amplitude signal, and the gate of each transistor 140-3, 140-4 is driven by a b: amplitude signal. In an example embodiment, the drain of each transistor 140-1, 140-2, 140-3, 140-4 is driven by a selected voltage determined according to the sign of the corresponding item activated, and the source of each transistor 140-1, 140-2, 140-3, 140-4 is coupled to resistors 136-1, 136-2, 136-3, 136-4, respectively. Note that representative read voltages of 0.2V, 0.4V, and 0.6V are used as examples (other representative read voltages are also contemplated).
[0039] During phase 1, the "a" input register 104-1 stores the activated a term and the "b" input register 104-2 stores the activated b term. Pulse width modulators 108-1, 108-2 generate pulses of amplitude a and b, respectively, which are applied to the gates of circuit pairs 112-1, 112-2, respectively. Since the pulse widths are proportional to the activated a term and the activated b term, respectively, the pulse widths are proportional to the activated a term and the activated b term, respectively, through resistor pairs 136-1, 136-2, 136-3, 136-4 (c + ,c - and d+ ,d - ) will flow for an amount of time that is proportional to the pulse width (and therefore proportional to the activated a and b terms, respectively). Note that, as shown, voltage supply node 132 (also referred to herein as node 132) is maintained at 0.4 volts (V) and that current (represented by the arrows) will flow in one direction for positive values of the activated term and in the opposite direction for negative values of the activated term. The amount of current flowing to analog-to-digital converter (ADC) 116 will represent the calculated real term and will be converted to a binary value by the current-based ADC 116 and stored in output register 120-2. By coupling nodes 132 of other circuits 100 (representing other activations and weights from other multiplication operations), the sum of the MACC operation can be input to ADC 116 and stored in output register 120-2. Note that output register 120-1 is described below in conjunction with Figure 1B Description. It should be noted that in one or more embodiments, the absolute value of each resistor pair 136-1, 136-2, 136-3, 136-4 is not as important as the shape of the distribution of resistance values. In one or more embodiments, the unitless weight distribution trained in the software is rescaled to the appropriate hardware unit (conductance expressed in microSiemens or μS) by a constant scaling factor. In this way, the shape of the weight distribution will not be distorted. Those skilled in the art will be familiar with the use of pulse units and suitable controllers to set the weights (conductance) of the memristive elements to the desired set, reset and / or intermediate states. In one or more embodiments, activation is encoded only as pulse width, and the amplitude does not differ. In addition, in one or more embodiments, only the relative resistance value is important; only the distribution of weights / resistances needs to be maintained. In addition, in one or more embodiments, the weight distribution is obtained from the software (unitless) and will be programmed in the hardware (microSiemens). Accordingly, a constant scaling factor is applied so that the unit changes, but the shape of the distribution remains unchanged.
[0040] Figure 1B is configured according to an example embodiment for use in stage 2, calculating the imaginary term Figure 1A FIG. 1 is a circuit diagram of a first exemplary embodiment of a circuit 100 for a MACC calculation engine. Figure 1B As shown, the application of the outputs of the "a" input register 104-1 and the "b" input register 104-2 to the pulse width modulators 108-1 and 108-2, the application of the a amplitude pulse and the b amplitude pulse to the gates of the circuit pair 112-1 and 112-2, and the application of the "a" sign and the "b" sign to the circuit pair 112-1 and 112-2 are reversed. As will be appreciated by those skilled in the art, Figure 1BIn the example of , a:sign and b:sign are analog voltages applied to the bottom terminal of the gate transistor, and the analog voltages may be governed by corresponding digital sign bits (not shown). (Note that the first pair of voltages corresponds to a positive sign, and the second pair of voltages corresponds to a negative sign.) For the avoidance of doubt, in one or more embodiments, there is a sign bit; however, for simplicity of presentation, the analog voltages are shown in the diagram. In one or more embodiments, there is some circuitry (omitted for simplicity) that converts the sign bit to corresponding analog voltages, which are shown in the diagram as 0.2V and 0.6V. Those skilled in the art may implement the circuitry using known techniques based on the teachings herein.
[0041] As in stage 1, 1) the width of the pulse is proportional to the activated item a and the activated item b respectively; 2) through the resistor pairs 136-1, 136-2, 136-3, 136-4 (c + ,c - and d + ,d - ) will flow for an amount of time that is proportional to the pulse width (and therefore proportional to the activated a and b terms, respectively); and the current (represented by the arrows) will flow in one direction for positive values of the activated term, and in the opposite direction for negative values of the activated term. The amount of current flowing to ADC 116 will represent the calculated imaginary term, and is converted to a binary value by ADC 116 and stored in output register 120-1. (As shown by the shaded area, during stage 2, output register 120-2 maintains the calculated real term.) After stage 2, the real and imaginary terms may be transferred to the input of another array, transferred through an activation function, etc. It is noted that rather than swapping the application of the outputs of the "a" input register 104-1 and the "b" input register 104-2 to the pulse width modulators 108-1, 108-2, the application of the outputs of the pulse width modulators 108-1, 108-2 to the gates may be swapped (applicable to all embodiments). One skilled in the art will recognize that the first voltage value of the sign bit corresponds to the positive activated term.
[0042] Figure 1C An example representation of a conventional crossbar array is shown. Input value V 1 To V N Generates output current I 1 to I M . Please note that Figure 1A and Figure 1B The circuit 100 in FIG. 1 provides support for complex conductance (G) values for the crossbar array 300. Thus, in one or more embodiments, the G value is complex, and the input (activation) is encoded using voltage and pulse duration, and can be specified as a complex input value, for example 1 , complex input value2 , complex input value N , instead of V in the prior art 1 、V 2 ,……,V N . Figure 1C Therefore, it is expressed Figure 1A and Figure 1B The relevant point is that now each input and conductance element is a complex value in rectangular coordinates, which means that the input activation and weight now each contain two values (real and imaginary). In this way, the high-level Figure 1C and Figure 1A and Figure 1B The more specific circuit implementation shown is relevant.
[0043] Figure 2A is a circuit diagram of a second example embodiment of a circuit 200 for a MACC calculation engine according to an example embodiment. Figure 2A The configuration of stage 1 for calculating real terms is shown, where the complex MACC includes 4 terms. The terms a, b, c, and d are constrained to be positive; therefore, there is no sign bit or corresponding circuit. Since all terms are positive, each weight can be represented by a single resistor 136-1, 136-2, and the current always flows in the same direction through resistors 136-1, 136-2. It is noted that representative read voltages 0.2V, 0.4V, 0.6V are used as examples (other representative read voltages are considered). Note that the sign "b" is negative to achieve b i d i Subtraction operation in the sum of terms.
[0044] During Phase 1, the "a" input register 104-1 stores the activated a term, and the "b" input register 104-2 stores the activated b term. If two imaginary numbers are multiplied, -1 is obtained. Therefore, if the imaginary number is represented by j, then j*j=-1 (by definition). Refer to the equation for the complex-valued MACC operation above. The pulse width modulators 108-1, 108-2 generate an a-amplitude pulse and a b-amplitude pulse, respectively, which are applied to the gates of the circuit pair 212-1. Since the pulse width is proportional to the activated a term and the activated b term, respectively, the current through the resistor pair 136-1, 136-2 (c and d) will flow for an amount of time proportional to the pulse width (and therefore proportional to the activated a term and b term, respectively). It should be noted that, as shown, node 132 is maintained at 0.4 volts (V), and the current (represented by the arrow) will flow in one direction for positive values of the activated term and in the opposite direction for negative values of the activated term. The amount of current flowing to the analog-to-digital converter (ADC) 116 will represent the calculated real term and will be converted to a binary value by the current-based ADC 116 and stored in the output register 120-2. By coupling nodes 132 of other circuits 100 (representing other activations and weights in other multiplication operations), the sum of the MACC operation can be input to the ADC 116 and stored in the output register 120-2. Note that the output register 120-1 is hereinafter combined with Figure 2B describe.
[0045] Figure 2B is configured according to an example embodiment for use in stage 2, calculating the imaginary term Figure 2A A circuit diagram of a second exemplary embodiment of a circuit 200 for a MACC calculation engine is shown in FIG. Figure 2B As shown, the application of the outputs of the "a" input register 104-1, the "b" input register 104-2 to the pulse width modulators 108-1, 108-2, and the application of the a-amplitude pulse and the b-amplitude pulse to the gate of the circuit pair 212-1 are reversed. As in stage 1, 1) the width of the pulse is proportional to the activated a-term and the activated b-term, respectively; 2) the current through the resistor pairs 136-1, 136-2 (c and d) will flow for an amount of time proportional to the pulse width (and therefore proportional to the activated a-term and b-term, respectively). The amount of current flowing to the ADC 116 will represent the calculated imaginary term, and will be converted to a binary value by the ADC 116 and stored in the output register 120-1. (As shown in the shaded area, during stage 2, the output register 120-2 maintains the calculated real term.) After stage 2, the real and imaginary terms can be transmitted to the input of another array, transmitted through an activation function, etc.
[0046] Figure 3Ais a circuit diagram of a third example embodiment of a circuit 300 for a MACC calculation engine according to an example embodiment. Figure 3A The configuration of stage 1 for calculating real terms is shown, where the complex MACC includes 4 terms. Terms a, b can be positive or negative, while c and d are positive. Since the weight is always positive, each weight can be represented by a single resistor 136-1, 136-2; however, since activation can be positive or negative, the direction of the current flowing through resistors 136-1, 136-2 is controlled by the sign bit. It should be noted that representative read voltages 0.2V, 0.4V, 0.6V are used as examples (other representative read voltages are considered).
[0047] During Phase 1, the "a" input register 104-1 stores an activated a term, and the "b" input register 104-2 stores an activated b term. The pulse width modulators 108-1, 108-2 generate an a-amplitude pulse and a b-amplitude pulse, respectively, which are applied to the gates of the circuit pair 212-1, respectively. Since the pulse width is proportional to the activated a term and the activated b term, respectively, the current through the resistor pair 136-1, 136-2 (c and d) will flow for an amount of time proportional to the pulse width (and therefore proportional to the activated a term and b term, respectively). It is noted that, as shown, node 132 is maintained at 0.4 volts (V) and the current (represented by the arrows) will flow in one direction for positive values of the activated term and in the opposite direction for negative values of the activated term. The amount of current flowing to the analog-to-digital converter (ADC) 116 will represent the calculated real term and will be converted to a binary value by the current-based ADC 116 and stored in the output register 120-2. By coupling nodes 132 of other circuits 100 (representing activations and weights in other multiplication operations), the sum of the MACC operation can be input to ADC 116 and stored in output register 120-2. Note that output register 120-1 is referred to below in conjunction with Figure 3B describe.
[0048] Figure 3B is configured according to an example embodiment for use in stage 2, calculating the imaginary term Figure 3A FIG. 3 is a circuit diagram of a third exemplary embodiment of a circuit 300 for a MACC calculation engine. Figure 1B, the application of the outputs of the "a" input register 104-1, the "b" input register 104-2 to the pulse width modulators 108-1, 108-2, the application of the a amplitude pulse and the b amplitude pulse to the gate of the circuit pair 212-1, and the application of the "a" sign bit and the "b" sign bit to the circuit pair 212-1 are reversed. As in stage 1, 1) the width of the pulse is proportional to the activated a term and the activated b term, respectively; 2) the current through the resistor pair 136-1, 136-2 (c and d) will flow for an amount of time proportional to the pulse width (and therefore proportional to the activated a term and the activated b term, respectively); and the current (indicated by the arrows) will flow in one direction for positive values of the activated term and in the opposite direction for negative values of the activated term. The amount of current flowing to the ADC 116 will represent the calculated imaginary term and will be converted to a binary value by the ADC 116 and stored in the output register 120-1. (As shown by the shaded area, during stage 2, output register 120-2 holds the computed real terms.) After stage 2, the real and imaginary terms may be transmitted to the input of another array, passed through an activation function, and so on.
[0049] Figure 4A is a circuit diagram of a fourth example embodiment of a circuit 400 for a MACC calculation engine according to an example embodiment. Figure 4A The configuration for stage 1 of computing real terms is shown, where the complex MACC consists of 4 terms. The terms a, b, c, and d can be positive or negative, where c′=cc s h and d′ = dd sh .exist Figure 4A In the embodiment of the present invention, the resistors 136-3 and 136-4 in the circuit pair 212-2 are shared by all circuit pairs 212-1 assigned to the same MACC operation and generate their own real and imaginary parts via the current-based ADC 128 to be stored in the output registers 124-2 and 124-1, respectively. Similarly, during phase 1 and phase 2, the resistors 136-1 and 136-2 in the circuit pair 212-1 jointly generate their own real and imaginary parts for the MACC operation via the current-based ADC 116 to be stored in the output registers 120-2 and 120-1, respectively. It should be noted that the representative read voltages 0.2V, 0.4V, and 0.6V are used as examples (other representative read voltages are considered).
[0050] Figure 4B is configured according to an example embodiment for use in stage 2, calculating the imaginary term Figure 3B4. A circuit diagram of a fourth example embodiment of a circuit 400 for a MACC computing engine is shown in FIG. Next, in stage 2, the values stored in real output registers 120-2, 124-2 are added together to generate the final real part, and the values stored in real output registers 120-1, 124-1 are added together to generate the final imaginary part. At this point, the real and imaginary terms may be transferred to the input of another array, transferred through an activation function, etc. In one or more embodiments, addition is used to determine the final result.
[0051] Figure 5 5 is a high-level diagram 500 of an implementation of a MACC computation engine according to an example embodiment. In an example embodiment, the engine 500 includes an array 504 of circuits 508-1, 508-2, such as circuits 100, 200, 300, and 400. As described above, each input of the circuits 508-1, 508-2 activates a 1 、b 1 , and use weight c 1 ,d 1 Programming. As will be appreciated by those skilled in the art, in machine learning, weights represent learned values that govern the behavior of a neural network. To perform the calculation, the imaginary part is calculated by flipping the sign and rerouting the pulses. In the example embodiment, as described above, each complex x term is represented using two pulse trains, and each complex w term is represented using two non-volatile memory (NVM) devices. It is noted that proof-of-concept experiments have been performed and satisfactory results have been obtained (proof-of-concept work requiring training of complex-valued neural networks was performed using known machine learning frameworks).
[0052] One or more embodiments include appropriate peripheral circuits 512, 516 and a suitable controller 520 for input / output, programming weights, etc. The peripheral circuits 512, 516 may include a pulse unit for programming, a pulse unit for input activation ... (including component a 1 ,b 1) and an output buffer for storing the result of the multiply-add operation, etc., as well as other components. According to the teachings herein, by adopting known techniques, such as integrators based on operational amplifiers and capacitors, digital logic circuits, field programmable gate arrays (FPGAs) or application specific integrated circuits (ASICs) or specific macros on memory chips, etc., those skilled in the art will be able to provide any additional desired / required peripheral circuits, voltage / power supplies (which can be controlled by a controller to supply appropriate voltages as described, which can be part of the controller 520, or a separate unit), elements for interfacing with peripheral circuits, and controllers. One or more embodiments therefore include an input vector peripheral circuit (e.g., 512) coupled to multiple activation inputs; and a control circuit 520, which is configured to control the input vector peripheral circuit 512 to perform a multiply-add operation with multiple memristor units having neural network weights stored therein. In order to implement any digital circuit described herein, computer-aided semiconductor integrated circuit (IC) logic design, simulation, testing, layout and / or manufacturing can be used. The computerized design process can represent functional and / or structural design features in a design structure generated using electronic computer-aided design (ECAD). A suitable hardware description language (HDL) may be used. Those skilled in the art may use, for example, Figure 6 Known computer-aided design techniques are implemented on the machine described in the drawings to synthesize digital logic circuits to perform the desired control and other functions.
[0053] In one example embodiment, the analog memory based complex weight cell includes four source follower circuits, wherein each source follower circuit has an analog memory device with one end connected to the source and the other end connected to a common node N between all analog memory elements. A , the node is kept at voltage V A (e.g., 0.4V), wherein the conductances in a first set of two source follower circuits differentially define a real portion of a complex weight, and wherein the conductances in a second set of two source follower circuits differentially define an imaginary portion of a complex weight. The use of source follower circuits is appropriate in one or more embodiments because it is desirable for the source voltage to remain stable at a variety of different conductance values in order for computations to be performed.
[0054] In one example embodiment, a pulse duration proportional to the real portion of the complex-valued input activation is applied to the gates of a first differential pair of source follower circuits representing the real portion of the complex weights.
[0055] In one example embodiment, a pulse duration proportional to the imaginary portion of the complex-valued input activation is applied to the gates of the second differential pair of source follower circuits representing the imaginary portion of the complex weight.
[0056] In one example embodiment, the drain voltage of the source follower circuit can be configured to be a range of voltages above and / or below V A voltage.
[0057] In one example embodiment, two or more cells are tied together at a common node N. A superior.
[0058] In one example embodiment, the common node N is measured by accumulating charge on a capacitor and measuring the charge using an analog-to-digital converter. A In one example embodiment, a pulse duration proportional to the real part of the activation is applied to the source follower gate corresponding to the imaginary part of the complex weight, and a pulse duration proportional to the imaginary part of the activation is applied to the source follower gate corresponding to the real part of the complex weight.
[0059] In view of the discussion thus far, it can be appreciated that, in general, according to aspects of the present invention, an example circuit includes a first pulse width modulator 108-1 configured to generate a first pulse based on a first input; a second pulse width modulator 108-2 configured to generate a second pulse based on a second input; a first differential circuit 112-1 including a first transistor 140-1, a second transistor 140-2, a first resistor 136-1, and a second resistor 136-2; and a second differential circuit 112-2 including a first transistor 140-3, a second transistor 140-4, a first resistor 136-3, and a second resistor 136-4, wherein: the first differential circuit 112-1 includes a first transistor 140-1, a second transistor 140-2, a first resistor 136-1, and a second resistor 136-2. The gate of the first transistor 140-1 of the first differential circuit 112-1 and the gate of the second transistor 140-2 of the first differential circuit 112-1, as well as the gate of the first transistor 140-3 of the second differential circuit 112-2 and the gate of the second transistor 140-4 of the second differential circuit 112-2 are configured to be controlled by the first pulse width modulator and the second pulse width modulator based on the first input and the second input; wherein: the first resistor 136-1 of the first differential circuit 112-1 has a configurable resistance, a first end of the first resistor 136-1 is coupled to a voltage via the voltage supply node 132, and a second end of the first resistor 136-1 is coupled to the first differential circuit 112-1 The first terminal of the second transistor 140-1 of the first differential circuit 112-1 is coupled to the first terminal of the second transistor 140-1 of the first differential circuit 112-1; the second resistor 136-2 of the first differential circuit 112-1 has a configurable resistance, the first terminal of the second resistor 136-2 of the first differential circuit 112-1 is coupled to the voltage via the voltage supply node 132, and the second terminal of the second resistor 136-2 of the first differential circuit 112-1 is coupled to the first terminal of the second transistor 140-2 of the first differential circuit 112-1; the first resistor 136-3 of the second differential circuit 112-2 has a configurable resistance, the first terminal of the first resistor 136-3 of the second differential circuit 112-2 is coupled to the voltage via the voltage supply node 132, The circuit also includes an analog-to-digital converter 116 having an input coupled to the voltage supply node 132. As will be appreciated by those skilled in the art, the voltage supply node 132 is configured to be connected to a suitable voltage source. During circuit operation, a voltage from a voltage source is applied to the voltage supply node 132. A constant voltage is typically applied to the voltage supply node 132.The constant voltage may be supplied by a constant voltage source, or provided by a variable voltage source configured to supply a constant voltage.
[0060] In an example embodiment, the first resistor 136-1 and the second resistor 136-2 of the first differential circuit 112-1 are configured based on a first corresponding weight of the multiply-accumulate operation, and the first resistor 136-3 and the second resistor 136-4 of the second differential circuit 112-2 are configured based on a second corresponding weight of the multiply-accumulate operation. In an example embodiment, the circuit also includes a first input register 104-1 coupled to an input of the first pulse width modulator 108-1 and configured to store a first input; a second input register 104-2 coupled to an input of the second pulse width modulator 108-2 and configured to store a second input; and a controller circuit configured to control loading of the first input register 104-1 and the second input register 104-2.
[0061] In an example embodiment, the circuit also includes: a real output register 120-2, the input of which is coupled to the output of the analog-to-digital converter 116; an imaginary output register 120-1, the input of which is coupled to the output of the analog-to-digital converter 116; and a controller circuit configured to control the loading of the real output register 120-2 and the imaginary output register 120-1.
[0062] In an example embodiment, a second end of the first transistor 140-1 of the first differential circuit 112-1, a second end of the second transistor 140-2 of the first differential circuit 112-1, a second end of the first transistor 140-3 of the second differential circuit 112-2, and a second end of the second transistor 140-4 of the second differential circuit 112-2 are configured to selectively couple between a voltage representing a sign of the first input and a voltage representing a sign of the second input, and the circuit further includes a controller circuit configured to control the selective coupling between the voltage representing the sign of the first input and the voltage representing the sign of the second input.
[0063] In an example embodiment, the analog-to-digital converter 116 includes one or more current mirrors and one or more capacitive elements configured to accumulate charge from the voltage supply node 132 .
[0064] In an example embodiment, each of a first end of the first transistor 140-1 of the first differential circuit 112-1, a first end of the second transistor 140-2 of the first differential circuit 112-1, a first end of the first transistor 140-3 of the second differential circuit 112-2, a first end of the second transistor 140-4 of the second differential circuit 112-2, a second end of the first transistor 140-1 of the first differential circuit 112-1, a second end of the second transistor 140-2 of the first differential circuit 112-1, a second end of the first transistor 140-3 of the second differential circuit 112-2, and a second end of the second transistor 140-4 of the second differential circuit 112-2 is a corresponding one of a source terminal and a drain terminal; that is, each (field effect) transistor 140-1, 140-2, 140-3, 140-4 has a source terminal and a drain terminal.
[0065] In one aspect, the circuit 200 includes a differential circuit 212-1 including a first transistor 140-1, a second transistor 140-2, a first resistor 136-1, and a second resistor 136-2; a first pulse width modulator 108-1 configured to generate a pulse based on a first input, an output of the first pulse width modulator 108-1 configured to be selectively coupled to one of a gate of the first transistor 140-1 of the differential circuit 212-1 and a gate of the second transistor 140-2 of the differential circuit 212-1; a first resistor 136-1 of the differential circuit 212-1 having a configurable resistance, a first end of the first resistor 136-1 coupled to a voltage via a voltage supply node 132, and a second end of the first resistor 136-1 coupled to the first transistor 136-1. 40-1 is coupled; a second pulse width modulator 108-2 configured to generate a pulse based on a second input, the output of the second pulse width modulator 108-2 is configured to be selectively coupled to one of the gate of the first transistor 140-1 of the differential circuit 212-1 and the gate of the second transistor 140-2 of the differential circuit 212-1; a second resistor 136-2 of the differential circuit 212-1 has a configurable resistance, a second end of the second resistor 136-2 is coupled to a voltage via a voltage supply node 132 and is coupled to a second end of the first resistor 136-1, and a second end of the second resistor 136-2 is coupled to a first end of the second transistor 140-2; and an analog-to-digital converter 116 having an input coupled to the voltage supply node 132.
[0066] In one exemplary embodiment, the first resistor 136-1 of the first differential circuit 212-1 is configured based on a first corresponding weight of the multiply-accumulate operation, and the second resistor 136-2 of the first differential circuit 212-1 is configured based on a second corresponding weight of the multiply-accumulate operation. In one exemplary embodiment, the circuit further includes a first input register 104-1 coupled to an input of the first pulse width modulator 108-1 and configured to store the first input, a second input register 104-2 coupled to an input of the second pulse width modulator 108-2 and configured to store the second input, and a controller circuit configured to control loading of the first input register 104-1 and the second input register 104-2.
[0067] In an exemplary embodiment, the circuit further includes a real output register 120-2 having an input coupled to the output of the analog-to-digital converter 116, an imaginary output register 120-1 having an input coupled to the output of the analog-to-digital converter 116, and a controller circuit configured to control loading of the real output register 120-2 and the imaginary output register 120-1. In an exemplary embodiment, the second end of the first transistor 140-1 and the second end of the second transistor 140-2 are configured to selectively couple between a voltage input representing a sign bit of the first input and a voltage input representing a sign bit of the second input, and the circuit further includes a controller circuit configured to control selective coupling between the voltage input representing the sign of the first input and the voltage input representing the sign of the second input (see Figure 3A-3B ).
[0068] In an exemplary embodiment, each of the first end of the first transistor 140-1, the first end of the second transistor 140-2, the second end of the first transistor 140-1, and the second end of the second transistor 140-2 of the first differential circuit 112-1 is a corresponding one of a source terminal and a drain terminal. In an exemplary embodiment, the circuit further includes a first shared transistor 140-3 and a second shared transistor 140-4, a first shared resistor 136-3 having a configurable resistance, a first end of the first shared resistor 136-3 being coupled to a voltage via a voltage supply node 132, and a second end of the first shared resistor 136-3 being coupled to the first end of the first shared transistor 140-3; a second shared resistor 136-4 having a configurable resistance, a first end of the second shared resistor 136-4 being coupled to a voltage via a voltage supply node 132, and a second end of the second shared resistor 136-4 being coupled to the first end of the first shared transistor 140-3. A second end of the shared resistor 136-4 is coupled to a first end of the second shared transistor 140-4; and a second analog-to-digital converter 128 having an input coupled to the voltage supply node 132; wherein: the output of the first pulse width modulator 108-1 is configured to be selectively coupled to the gate of the first shared transistor 140-3 and the gate of the second shared transistor 140-4, and the output of the second pulse width modulator 108-2 is configured to be selectively coupled to the gate of the first shared transistor 140-3 and the gate of the second shared transistor 140-4 (see Figure 4A-4B ).
[0069] In an exemplary embodiment, the circuit further includes a real part shared output register 124-2 having an input coupled to the output of the analog-to-digital converter 128, and an imaginary part shared output register 124-1 having an input coupled to the output of the analog-to-digital converter 128. In an exemplary embodiment, the second end of the first shared transistor 140-3 and the second end of the second shared transistor 140-4 are configured to selectively couple between a voltage input representing a sign bit of the first input and a voltage input representing a sign bit of the second input. In an exemplary embodiment, each of the first end of the first shared transistor 140-3, the first end of the second shared transistor 140-4, the second end of the first shared transistor 140-3, and the second end of the second shared transistor 140-4 is a corresponding one of a source terminal and a drain terminal.
[0070] Various aspects of the present disclosure are described by narrative text, flow charts, block diagrams of computer systems, and / or machine logic block diagrams included in computer program product (CPP) embodiments. With respect to any flow chart, depending on the technology involved, the operations may be performed in an order different from that shown in a given flow chart. For example, again depending on the technology involved, two operations shown in successive flow chart blocks may be performed in reverse order, as a single integrated step, in parallel, or in a manner that overlaps at least partially in time.
[0071] Computer program product embodiments ("CPP embodiments" or "CPP") are terms used in this disclosure to describe any collection of one or more storage media (also referred to as "media") that are collectively contained in a collection of one or more storage devices that collectively contain machine-readable code corresponding to instructions and / or data for performing computer operations defined in a given CPP claim. A "storage device" is any tangible device that can hold and store instructions for use by a computer processor. Without limitation, a computer-readable storage medium may be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include: a magnetic disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device (such as a punch card or pits / land formed on a major surface of an optical disk), or any suitable combination of the foregoing. As the term is used in this disclosure, computer-readable storage media should not be construed as storage in the form of transient signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides, light pulses transmitted through optical fibers, electrical signals communicated through wires, and / or other transmission media. As will be appreciated by those skilled in the art, data is typically moved at some occasional point in time during the normal operation of the storage device (e.g., during access, defragmentation, or garbage collection), but this does not make the storage device transient because the data is not transient when it is stored.
[0072] Reference now Figure 6. The computing environment 100 includes an example of an environment for executing at least some of the computer code 200 involved in performing the method of the present invention, such as training a unitless weight distribution for a neural network using forward propagation and back propagation in software as is well known to those skilled in the art, and also rescaling the unitless weight distribution so that according to various aspects of the present invention, appropriate conductance (in micro-Siemens or μS) can be programmed into appropriate hardware units; the code can also be used for control aspects, or for synthesizing digital circuits to perform control aspects in a known manner. In addition, the analog MAC engine circuit disclosed herein can be used in a general-purpose machine as a hardware accelerator. In addition to box 200, the computing environment 100 also includes, for example, a computer 101, a wide area network (WAN) 102, an end-user device (EUD) 103, a remote server 104, a public cloud 105, and a private cloud 106. In this embodiment, the computer 101 includes a processor set 110 (including a processing circuit 120 and a cache 121), a communication structure 111, a volatile memory 112, a persistent storage device 113 (including an operating system 122 and the above-mentioned frame 200), a peripheral device set 114 (including a user interface (UI) device set 123, a storage device 124, and an Internet of Things (IoT) sensor set 125), and a network module 115. The remote server 104 includes a remote database 130. The public cloud 105 includes a gateway 140, a cloud orchestration module 141, a host physical machine set 142, a virtual machine set 143, and a container set 144.
[0073] Computer 101 may be a desktop computer, a laptop computer, a tablet computer, a smartphone, a smartwatch or other wearable computer, a mainframe computer, a quantum computer, or any other form of computer or mobile device currently known or to be developed in the future that is capable of running programs, accessing a network, or querying a database (e.g., remote database 130). As is well known in the art of computer technology, and depending on the technology, the execution of the computer-implemented method may be distributed among multiple computers and / or among multiple locations. On the other hand, in the presentation of the present computing environment 100, in order to simplify the presentation as much as possible, the detailed discussion focuses on a single computer, particularly computer 101. Computer 101 may be located in the cloud, although Figure 6 On the other hand, the computer 101 need not be located in the cloud unless otherwise indicated.
[0074] Processor set 110 includes one or more computer processors of any type currently known or to be developed in the future. Processing circuit 120 may be distributed in multiple packages, such as in multiple collaborative integrated circuit chips. Processing circuit 120 may implement multiple processor threads and / or multiple processor cores. Cache 121 is a memory located in the processor chip package, typically used for data or code that threads or cores running on processor set 710 should be able to access quickly. Cache memory is typically organized into multiple levels depending on the relative proximity to the processing circuit. Alternatively, some or all of the caches of the processor set may be located "off chip". In some computing environments, processor set 110 may be designed to process quantum bits and perform quantum computing.
[0075] Computer-readable program instructions are typically loaded into the computer 101 to cause the processor set 110 of the computer 101 to perform a series of operating steps to achieve a computer-implemented method, such that the instructions so executed will instantiate the method specified in the flowchart and / or narrative description of the computer-implemented method included in this document (collectively referred to as the "inventive method"). These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by the processor set 110 to control and direct the execution of the inventive method. In the computing environment 100, at least some of the instructions for executing the inventive method can be stored in the box 200 of the persistent storage device 113.
[0076] The communication fabric 111 is a signal conduction path that allows the various components of the computer 101 to communicate with each other. Typically, such a fabric is composed of switches and conductive paths, such as those that form a bus, a bridge, physical input / output ports, etc. Other types of signal communication paths may also be used, such as fiber optic communication paths and / or wireless communication paths.
[0077] The volatile memory 112 is any type of volatile memory currently known or to be developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, the volatile memory 112 has random access characteristics, but unless explicitly specified, it is not required to be random access. In the computer 101, the volatile memory 112 is located in a single package and is internal to the computer 101, but alternatively or additionally, the volatile memory can be distributed over multiple packages and / or located externally relative to the computer 101.
[0078] Persistent storage 113 is any form of non-volatile storage for computers currently known or to be developed in the future. The non-volatility of the storage means that the stored data is maintained regardless of whether power is supplied to the computer 101 and / or directly supplied to the persistent storage 113. Persistent storage 113 can be a read-only memory (ROM), but typically at least a portion of the persistent storage allows writing data, deleting data, and rewriting data. Some common forms of persistent storage include disks and solid-state storage devices. Operating system 122 can take several forms, such as various known proprietary operating systems or open source portable operating system interface type operating systems using a kernel. The code contained in box 200 typically includes at least some of the computer code involved in executing the inventive method.
[0079] The peripheral device set 114 includes a collection of peripheral devices of the computer 101. The data communication connection between the peripheral devices and other components of the computer 101 can be achieved in various ways, such as a Bluetooth connection, a near field communication (NFC) connection, a connection through a cable (such as a universal serial bus (USB) type cable), a plug-in connection (such as a secure digital (SD) card), a connection through a local area communication network, or even a connection through a wide area network such as the Internet. In various embodiments, the UI device set 123 may include components such as display screens, speakers, microphones, wearable devices (such as goggles and smart watches), keyboards, mice, printers, trackpads, game controllers, and tactile devices. The storage device 124 is an external storage device, such as an external hard drive, or a pluggable storage device, such as an SD card. The storage 124 can be persistent and / or volatile. In some embodiments, the storage device 124 can take the form of a quantum computing storage device for storing data in the form of quantum bits. In embodiments where computer 101 requires a large amount of storage (e.g., where computer 101 locally stores and manages a large database), storage may be provided by an external storage device designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple geographically dispersed computers. IoT sensor set 125 consists of sensors that can be used in IoT applications. For example, one sensor may be a thermometer and another sensor may be a motion detector.
[0080] The network module 115 is a collection of computer software, hardware, and firmware that allows the computer 101 to communicate with other computers via the WAN 102. The network module 115 may include hardware such as a modem or a Wi-Fi signal transceiver, software for packaging and / or unpacking data for communication network transmission, and / or a network browser software for communicating data on the Internet. In some embodiments, the network control function and the network forwarding function of the network module 115 may be executed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing software defined networks (SDN)), the control function and the forwarding function of the network module 115 are executed on physically separated devices, so that the control function manages several different network hardware devices. Computer-readable program instructions for executing the inventive method can typically be downloaded to the computer 101 from an external computer or an external storage device via a network adapter card or a network interface included in the network module 115.
[0081] WAN 102 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances by any technology currently known or to be developed in the future for transmitting computer data. In some embodiments, WAN 102 may be replaced and / or supplemented by a local area network (LAN) designed to transmit data between devices located within a local area (e.g., a Wi-Fi network). WANs and / or LANs typically include computer hardware such as copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers.
[0082] End-user device (EUD) 103 is any computer system used and controlled by an end-user (e.g., a customer of an enterprise operating computer 101), and may take any form discussed above in connection with computer 101. EUD 103 typically receives helpful and useful data from the operation of computer 101. For example, in the hypothetical case where computer 101 is designed to provide recommendations to an end-user, the recommendations would typically be transmitted from network module 115 of computer 101, over WAN 102, to EUD 103. In this way, EUD 103 may display or otherwise present the recommendations to the end-user. In some embodiments, EUD 103 may be a client device, such as a thin client, a fat client, a mainframe computer, a desktop computer, etc.
[0083] Remote server 104 is any computer system that provides at least some data and / or functionality to computer 101. Remote server 104 may be controlled and used by the same entity that operates computer 101. Remote server 104 represents a machine that collects and stores helpful and useful data for use by other computers, such as computer 701. For example, in the hypothetical situation where computer 101 is designed and programmed to provide recommendations based on historical data, then this historical data may be provided to computer 101 from remote database 130 of remote server 104.
[0084] The public cloud 105 is any computer system that is available to multiple entities and provides on-demand availability of computer system resources and / or other computing capabilities, particularly data storage (cloud storage) and computing capabilities, without the need for direct active management by users. Cloud computing typically uses resource sharing to achieve consistency and economies of scale. The direct and active management of the computing resources of the public cloud 105 is performed by computer hardware and / or software of the cloud orchestration module 141. The computing resources provided by the public cloud 105 are typically implemented through virtual computing environments that run on various computers that constitute the computers of the host physical machine set 142, which is the totality of physical computers in and / or available to the public cloud 105. The virtual computing environment (VCE) typically takes the form of a virtual machine from a virtual machine set 143 and / or a container from a container set 144. It is understood that these VCEs can be stored as images and can be transmitted between various physical machine hosts as images or after instantiation of the VCE. The cloud orchestration module 141 manages the transmission and storage of images, deploys new VCE instances, and manages active instances of VCE deployments. Gateway 140 is a collection of computer software, hardware, and firmware that allows public cloud 105 to communicate over WAN 102 .
[0085] A further explanation of a virtual computing environment (VCE) will now be provided. A VCE can be stored as an "image". A new active instance of a VCE can be instantiated from an image. Two common types of VCEs are virtual machines and containers. A container is a type of VCE that uses operating system-level virtualization. This refers to a feature of an operating system in which the kernel allows the existence of multiple isolated user space instances, called containers. These isolated user space instances generally behave like real computers from the perspective of the programs running in them. A computer program running on a normal operating system can utilize all of the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and the devices assigned to the container, a feature known as containerization.
[0086] Private cloud 106 is similar to public cloud 105, except that the computing resources are only available to a single enterprise. Although private cloud 106 is depicted as communicating with WAN 102, in other embodiments, the private cloud can be completely disconnected from the Internet and accessed only through a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private cloud, community cloud, or public cloud types), typically implemented separately by different vendors. Each of the multiple clouds remains as a separate and discrete entity, but the larger hybrid cloud architecture is bound together by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the multiple constituent clouds. In this embodiment, both public cloud 105 and private cloud 106 are part of a larger hybrid cloud.
[0087] The description of various embodiments of the present invention has been presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, practical applications or technical improvements over commercially available technologies, or to allow others of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A circuit, include: a first pulse width modulator configured to generate a first pulse based on a first input; a second pulse width modulator configured to generate a second pulse based on a second input; a first differential circuit including a first transistor, a second transistor, a first resistor, and a second resistor; A second differential circuit includes a first transistor, a second transistor, a first resistor, and a second resistor, wherein: a gate of the first transistor of the first differential circuit and a gate of the second transistor of the first differential circuit; and a gate of the first transistor of the second differential circuit and a gate of the second transistor of the second differential circuit configured to be controlled by the first pulse width modulator and the second pulse width modulator based on the first input and the second input; in: the first resistor of the first differential circuit having a configurable resistance, a first end of the first resistor coupled to a voltage supply node, and a second end of the first resistor coupled to a first end of the first transistor of the first differential circuit; the second resistor of the first differential circuit having a configurable resistance, a first end of the second resistor of the first differential circuit coupled to the voltage via the voltage supply node, and a second end of the second resistor of the first differential circuit coupled to a first end of the second transistor of the first differential circuit; the first resistor of the second differential circuit having a configurable resistance, a first end of the first resistor of the second differential circuit coupled to the voltage via the voltage supply node, and a second end of the first resistor of the second differential circuit coupled to a first end of the first transistor of the second differential circuit; and the second resistor of the second differential circuit having a configurable resistance, a first end of the second resistor of the second differential circuit coupled to the voltage via the voltage supply node, and a second end of the second resistor of the second differential circuit coupled to a first end of the second transistor of the second differential circuit; Also included is an analog-to-digital converter having an input coupled to the voltage supply node.
2. The circuit of claim 1 , wherein the first resistor and the second resistor of the first differential circuit are configured based on first corresponding weights of a multiply-add operation, and the first resistor and the second resistor of the second differential circuit are configured based on second corresponding weights of the multiply-add operation.
3. The circuit according to claim 1, further comprising: include: a first input register coupled to an input of the first pulse width modulator and configured to store the first input; a second input register coupled to an input of the second pulse width modulator and configured to store the second input; as well as A controller circuit is configured to control loading of the first input register and the second input register.
4. The circuit according to claim 1, further comprising: include: a real output register having an input coupled to an output of the analog-to-digital converter; an imaginary output register, an input of the imaginary output register being coupled to the output of the analog-to-digital converter; as well as A controller circuit is configured to control loading of the real output register and the imaginary output register.
5. The circuit of claim 1 , wherein the second end of the first transistor of the first differential circuit, the second end of the second transistor of the first differential circuit, the second end of the first transistor of the second differential circuit, and the second end of the second transistor of the second differential circuit are configured to selectively couple between a voltage representing a sign of the first input and a voltage representing a sign of the second input, and further comprising a controller circuit configured to control the selective coupling between the voltage representing the sign of the first input and the voltage representing the sign of the second input. 6 . The circuit of claim 1 , wherein the analog-to-digital converter comprises one or more current mirrors and one or more capacitive elements configured to accumulate charge from the voltage supply node.
7. The circuit of claim 1 , wherein each of the first terminal of the first transistor of the first differential circuit, the first terminal of the second transistor of the first differential circuit, the first terminal of the first transistor of the second differential circuit, the first terminal of the second transistor of the second differential circuit, the second terminal of the first transistor of the first differential circuit, the second terminal of the second transistor of the first differential circuit, the second terminal of the first transistor of the second differential circuit, and the second terminal of the second transistor of the second differential circuit is a corresponding one of a source terminal and a drain terminal.
8. A circuit, include: a differential circuit comprising a first transistor, a second transistor, a first resistor, and a second resistor; a first pulse width modulator configured to generate a pulse based on a first input, an output of the first pulse width modulator configured to be selectively coupled to one of a gate of the first transistor of the differential circuit and a gate of the second transistor of the differential circuit; The first resistor of the differential circuit has a configurable resistance, a first end of the first resistor is coupled to a voltage supply node, and a second end of the first resistor is coupled to a first end of the first transistor; a second pulse width modulator configured to generate pulses based on a second input, an output of the second pulse width modulator configured to be selectively coupled to one of a gate of the first transistor of the differential circuit and a gate of the second transistor of the differential circuit; the second resistor of the differential circuit having a configurable resistance, a second end of the second resistor coupled to the voltage via the voltage supply node and to the second end of the first resistor, and a second end of the second resistor coupled to the first end of the second transistor; as well as An analog-to-digital converter has an input coupled to the voltage supply node.
9. The circuit of claim 8, wherein the first resistor of the first differential circuit is configured based on a first corresponding weight of a multiply-add operation, and the second resistor of the first differential circuit is configured based on a second corresponding weight of the multiply-add operation.
10. The circuit according to claim 8, further comprising: include: a first input register coupled to an input of the first pulse width modulator and configured to store the first input; a second input register coupled to an input of the second pulse width modulator and configured to store the second input; as well as A controller circuit is configured to control loading of the first input register and the second input register.
11. The circuit according to claim 8, further comprising: include: a real output register having an input coupled to an output of the analog-to-digital converter; an imaginary output register having an input coupled to the output of the analog-to-digital converter; as well as A controller circuit is configured to control loading of the real output register and the imaginary output register.
12. The circuit of claim 8 , wherein the second end of the first transistor and the second end of the second transistor are configured to selectively couple between a voltage input representing a sign bit of the first input and a voltage input representing a sign bit of the second input, and further comprising a controller circuit configured to control the selective coupling between the voltage input representing the sign of the first input and the voltage input representing the sign of the second input.
13. The circuit of claim 8, wherein each of the first terminal of the first transistor, the first terminal of the second transistor of the first differential circuit, the second terminal of the first transistor, and the second terminal of the second transistor is a corresponding one of a source terminal and a drain terminal.
14. The circuit according to claim 8, further comprising: include: a first shared transistor; a second shared transistor; a first shared resistor having a configurable resistance, a first end of the first shared resistor being coupled to the voltage via the voltage supply node and a second end of the first shared resistor being coupled to a first end of the first shared transistor; a second shared resistor having a configurable resistance, a first end of the second shared resistor coupled to the voltage supply node and a second end of the second shared resistor coupled to the first end of the second shared transistor; as well as a second analog-to-digital converter having an input coupled to the voltage supply node; in: the output of the first pulse width modulator being configured to be selectively coupled to one of a gate of the first shared transistor and a gate of the second shared transistor; as well as The output of the second pulse width modulator is configured to be selectively coupled to one of the gate of the first shared transistor and the gate of the second shared transistor.
15. The circuit according to claim 14, further comprising: include: a real part shared output register having an input coupled to an output of the analog-to-digital converter; as well as An imaginary shared output register has an input coupled to the output of the analog-to-digital converter.
16. The circuit of claim 14, wherein the second terminal of the first shared transistor and the second terminal of the second shared transistor are configured to selectively couple between a voltage input representing a sign bit of the first input and a voltage input representing a sign bit of the second input. 17 . The circuit of claim 14 , wherein each of the first terminal of the first shared transistor, the first terminal of the second shared transistor, the second terminal of the first shared transistor, and the second terminal of the second shared transistor is a corresponding one of a source terminal and a drain terminal.