Scalable Switched-Capacitor Compute Core for Accurate and Efficient Deep Learning Inference

An automatic search algorithm determines optimal scaling factors for ADC truncation in switched-capacitor computation cores, ensuring accurate neural network inference by adjusting amplifiers or multipliers, thus maintaining precision despite using low-precision ADCs.

JP2025541591APending Publication Date: 2025-12-22INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025525261
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-29
Filing Date
2023-11-23
Publication Date
2025-12-22

AI Technical Summary

Technical Problem

The use of low-precision ADCs in switched-capacitor computation cores leads to reduced accuracy in neural network inference due to truncation of analog outputs, which affects the precision of multiply-and-accumulate operations.

Method used

An automatic search algorithm is employed to determine an optimal integer scalar for ADC truncation, which is incorporated into the hardware through amplifiers, input multipliers, or charge sharing to maintain accuracy during neural network operations.

Benefits of technology

This approach enables the use of low-precision ADCs without compromising accuracy, achieving performance comparable to high-precision ADCs by optimizing the scaling factors for each layer of the neural network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025541591000001_ABST
    Figure 2025541591000001_ABST
Patent Text Reader

Abstract

The apparatus includes: a first plurality of inputs representing activation input vectors; a second plurality of inputs representing weight input vectors; an analog multiply-accumulate multiplier and accumulator that generates a first analog voltage representing a first multiply-accumulate multiplication and accumulation result for the first input and the second input; a voltage multiplier that receives the first analog voltage and generates a second analog voltage representing a second multiply-accumulate multiplication and accumulation result by multiplying the first analog voltage by at least one scaling factor; an analog-to-digital converter configured to convert the second analog voltage multiply-accumulate multiplication and accumulation result to a digital signal using limited precision arithmetic during a neural network inference operation; and a hardware controller configured to determine at least one scaling factor based on the first multiply-accumulate multiplication and accumulation result or a software controller configured to determine at least one scaling factor based on the first multiply-accumulate multiplication and accumulation result.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE INVENTION Exemplary embodiments described herein relate generally to machine learning hardware device design and integrated circuit design, and more particularly to scalable switched-capacitor computational cores for accurate and efficient deep learning inference. Summary of the Invention

[0002] In one aspect, an apparatus includes a first plurality of inputs representing activation input vectors; a second plurality of inputs representing weight input vectors; an analog multiplier-and-accumulator configured to generate a first analog voltage representing a first multiply-and-accumulate result for the first input and the second input; a voltage multiplier configured to receive the first analog voltage and generate a second analog voltage representing a second multiply-and-accumulate result by multiplying the first analog voltage by at least one scaling factor; an analog-to-digital converter configured to convert the second analog voltage multiply-and-accumulate result to a digital signal using limited precision arithmetic during a neural network inference operation; and a hardware controller configured to determine at least one scaling factor based on the first multiply-and-accumulate result or a software controller configured to determine at least one scaling factor based on the first multiply-and-accumulate result.

[0003] In another aspect, an apparatus includes a first plurality of inputs representing an original activation input vector; a plurality of voltage multipliers receiving the first plurality of inputs and generating a second plurality of inputs by multiplying voltages of the original activation input vector by at least one scaling factor; a third plurality of inputs representing a weight input vector; an analog multiplier and accumulator generating analog voltages representing multiplication and accumulation results for the second and third inputs; an analog-to-digital converter configured to convert the analog voltage multiplication and accumulation results to digital signals using limited-precision operations during neural network inference operations; and a hardware controller configured to determine at least one scaling factor based on the multiplication and accumulation results, or a software controller configured to determine at least one scaling factor based on the multiplication and accumulation results.

[0004] In another aspect, a method includes receiving a first plurality of inputs representing an activation input vector; receiving a second plurality of inputs representing a weight input vector; generating, using analog multipliers and accumulators, analog voltages representing multiplication and accumulation results of the first plurality of inputs and the second plurality of inputs; converting, using an analog-to-digital converter, the analog voltage multiplication and accumulation results to digital signals using limited precision arithmetic during inference operations of the neural network; and determining at least one scaling factor used to amplify the first plurality of inputs or the analog voltage multiplication and accumulation results during training or calibration of the neural network. [Brief explanation of the drawings]

[0005] The foregoing and other aspects of the exemplary embodiments will become more apparent from the following detailed description when read in conjunction with the accompanying drawings.

[0006] [Figure 1] FIG. 1 shows a high-level diagram of a mixed-signal switched-capacitor multiplier and accumulator.

[0007] [Figure 2] A 16-bit accumulator and an 8-bit accumulator with exact truncation are shown.

[0008] [Figure 3] A 16-bit accumulator and an 8-bit accumulator with a scaled distribution of values ​​as input to the ADC are shown.

[0009] [Figure 4] Shows the processing of scalars in the DNN layer.

[0010] [Figure 5] FIG. 1 is a flow diagram of an automated search algorithm for determining the optimal scalar.

[0011] [Figure 6] 1 is an exemplary software implementation of an auto-scaling algorithm according to embodiments described herein.

[0012] [Figure 7] 1 illustrates an exemplary implementation of the truncation portion of the auto-scaling algorithm described herein.

[0013] [Figure 8] 1 illustrates an exemplary implementation of a portion of the auto-scaling algorithm described herein.

[0014] [Figure 9A] 1 illustrates the use of an amplifier to scale an analog signal in switched-capacitor hardware.

[0015] [Figure 9B] 1 illustrates the use of an input multiplier to scale an analog signal in switched-capacitor hardware.

[0016] [Figure 9C]1 illustrates charge sharing for scaling analog signals in switched-capacitor hardware.

[0017] [Figure 10] Schematic for machine learning hardware.

[0018] [Figure 11] FIG. 2 is a circuit diagram illustrating a first embodiment of the implementation described herein involving multiplication and amplification of the accumulation result.

[0019] [Figure 12] FIG. 1 is a circuit diagram of an embodiment of a sum multiplier using switched capacitors.

[0020] [Figure 13] FIG. 10 is a circuit diagram illustrating a second embodiment of the implementation described herein with a voltage multiplier at the input.

[0021] [Figure 14] FIG. 1 is a circuit diagram of one embodiment of an input multiplier using a level shifter.

[0022] [Figure 15] FIG. 10 is a circuit diagram illustrating a first operational phase of a third embodiment of an example implementation described herein that uses a sum multiplier with parallel-connected capacitors to implement voltage sampling.

[0023] [Figure 16] FIG. 10 is a circuit diagram illustrating a second operational phase of a third embodiment of an example implementation described herein that implements voltage multiplication using a sum multiplier with capacitors reconfigured to be connected in series.

[0024] [Figure 17] 10 is a graph showing NN accuracy performance results comparing results with and without implementation of an embodiment described herein.

[0025] [Figure 18] 10 is another graph showing the NN accuracy performance results.

[0026] [Figure 19] 10 is a graph showing the convergence of quantization-aware training without implementing the embodiments described herein.

[0027] [Figure 20] 10 is a graph showing a comparison of performance results with and without automatic search scaling.

[0028] [Figure 21] 1 is a logic flow diagram for implementing a method in accordance with embodiments described herein.

[0029] [Figure 22] 1 is a logic flow diagram for implementing a method in accordance with embodiments described herein. DETAILED DESCRIPTION OF THE INVENTION

[0030] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. All embodiments described in this detailed description are exemplary embodiments provided to enable any person skilled in the art to make or use the invention, and do not limit the scope of the invention, which is defined by the claims.

[0031] To limit the energy consumption of the ADC and realize an energy-efficient switched-capacitor computation core, a low-precision ADC (<16 bits) is required. Such a low-precision ADC truncates the analog output of the switched-capacitor MACC when the analog output falls outside a predetermined voltage range, providing a digital output represented in fewer than 16 bits. This truncation operation reduces the precision of the analog MACC output, potentially resulting in reduced accuracy during neural network inference. Therefore, hardware and software are needed that enable ADC truncation without reducing the accuracy of neural network inference.

[0032] Therefore, this specification describes a method for determining an optimal integer scalar for ADC truncation via an automatic search algorithm ("autoscale") and associated hardware implementation. This specification discloses process steps for determining at least one optimal integer scalar, a method for processing MACC per layer, and a method for handling overflow. This specification describes three options for incorporating the determined scalar in hardware by modifying the analog signal during switched-capacitor analog calculation: (1) an amplifier (input or output), (2) an input multiplier, and (3) charge sharing.

[0033] The problem addressed in the embodiments described herein is that SC-PT core ADC truncation impacts accuracy. The embodiments described herein fully utilize the accuracy of the SC-PT core ADC. The MACC input or output is scaled by an integer factor, with various implementation options.

[0034] Figure 1 shows a high-level diagram of a mixed-signal switched-capacitor multiplier and accumulator (10). The mixed-signal switched-capacitor multiplier and accumulator (10) can store 512 four-bit values, i.e., 512 x 4b[

number

number

number

number

number

number

number

number

number

[0035] The output (13) can be based on several factors, such as applying different DNN layers, such as linear and BMM layers. Linear layers can be based on scaling outputs or adding values ​​in activations. BMM layers can be based on scaling outputs.

[0036] The mixed-signal switched-capacitor multiplier and accumulator (10) can be coupled to an amplifier circuit (9) with amplifiers supporting different gains. Specifically, amplifiers can be added to scale analog signals with software-defined gains.

[0037] Figure 2 shows a 16-bit accumulator 20 and an 8-bit accumulator 21. As shown in Figure 2, the 16-bit accumulator 20 includes bits 24-1, 24-2, 24-3, 24-4, 24-5, 24-6, 24-7, 24-8, 24-9, 24-10, 24-11, 24-12, 24-13, 24-14, 24-15, and 24-16. The 8-bit accumulator 21 includes bits 24-2, 24-3, 24-4, 24-5, 24-6, 24-7, 24-8, and 24-9. In FIG. 2, an example of the distribution of values ​​at the input to the ADC with strict truncation is shown by bits 24-6, 24-7, 24-8, 24-9, 24-10, 24-11, 24-12, 24-13, 24-14, 24-15, and 24-16.

[0038] FIG. 2 illustrates a 16-bit accumulator (20) that does not implement the embodiments described herein. The bits labeled 24-6, 24-7, 24-8, 24-9, 24-10, 24-11, 24-12, 24-13, 24-14, 24-15, and 24-16 represent the virtual limits (range) of the MACC output (the MACC corresponds to a multiplication and accumulation operation). Two truncation thresholds (MSB truncation 22 and LSB truncation 23) determine the conversion of the MACC output from an analog voltage to a digital representation following processing by a low-precision ADC (in this example, an 8-bit ADC). In this scenario, many bits are truncated (i.e., the MACC output value is approximated with a highly truncated representation after digital conversion). This results in a larger MACC error compared to a non-approximated MACC, reducing the accuracy of the neural network. As an example, referring to the experimental results in FIG. 17, 7-bit LSB truncation (plot 806) results in an F1 performance (F1 is a measure of precision) of −6.9% compared to the no-truncation results obtained with a 16-bit ADC (plot 804).

[0039] Figure 3 shows a 16-bit accumulator (25) and an 8-bit accumulator (26). Shown are MSB truncation threshold 27 and LSB truncation threshold 28, which determine the conversion of the MACC output from an analog voltage to a digital representation following processing by the low-precision ADC. The 16-bit accumulator (25) includes bits 29-1, 29-2, 29-3, 29-4, 29-5, 29-6, 29-7, 29-8, 29-9, 29-10, 29-11, 29-12, 29-13, 29-14, 29-15, and 29-16. The 8-bit accumulator 26 includes bits 29-2, 29-3, 29-4, 29-5, 29-6, 29-7, 29-8, and 29-9. The scaled distribution of values ​​at the input to the ADC is indicated by bits 29-2, 29-3, 29-4, 29-5, 29-6, 29-7, 29-8, 29-9, 29-10, and 29-11.

[0040] Figure 3 shows a 16-bit accumulator (25) implementing an embodiment described herein. In Figure 3, the MACC value is scaled by an integer factor, then truncated by analog-to-digital conversion of a low-precision ADC, and then the result is shifted downward. This results in significant performance improvements, as shown in Figures 17 and 18. For example, 8-bit LSB truncation (plot 802) results in only -0.1% F1 performance (slightly worse) compared to the no-truncation result (plot 804) obtained with a 16-bit ADC.

[0041] The output distribution is different for each DNN layer. Fixed ADC truncation causes significant degradation. ADC power savings are primarily achieved through LSB truncation. For example, to save ADC power, it is preferable to truncate the LSB rather than the MSB.

[0042] Shaded bits 24-6, 24-7, 24-8, 24-9, 24-10, 24-11, 24-12, 24-13, 24-14, 24-15, and 24-16 in Figure 2 and shaded bits 29-2, 29-3, 29-4, 29-5, 29-6, 29-7, 29-8, 29-9, 29-10, and 29-11 in Figure 3 represent bits required to cover a hypothetical distribution of the input to the ADC, or equivalently, the output of the MACC (labeled 13 in Figure 1). For example, the input to the ADC may have a low value and occupy the least significant bits (24-6 through 24-16 in Figure 2). The other bits are unused (24-1 through 24-5 in Figure 2). As a result, when the LSB bit is truncated, performance is degraded because multiple used bits (shaded bits 24-6, 24-7, 24-8, 24-9, 24-10, 24-11, 24-12, 24-13, 24-14, 24-15, and 24-16 in FIG. 2) are truncated.

[0043] "Scaled" in Figure 3 corresponds to amplification: each value of the input to the ADC is multiplied, or scaled up, by an amplification factor. There can be a distribution of inputs to the ADC (or MACC outputs, which are the same thing). Performing many MACC operations with different inputs results in a different MACC output for each operation performed. Each distribution represents a virtual set of MACC outputs, either amplified or unamplified.

[0044] Figure 4 shows the general processing of scalars in a DNN layer. A DNN layer 35 with an accumulation size of N (e.g., N is an integer) is divided (36) into swcap operation 37 with L accumulations (e.g., L is an integer), swcap operation 38 with L accumulations, and swcap operation 39 with L accumulations. Swcap operation 37 is associated with scalar A (40), swcap operation 38 is associated with scalar B (41), and swcap operation 39 is associated with scalar C (42).

[0045] The swcap operations (37, 38, 39) perform atomic MACC within the swcap core. The accumulation length L is fixed by the hardware (e.g., L=512).

[0046] The GEMM performed by the DNN layer may require N>L accumulations, in which case the MACC layer is split into multiple swcap atomic MACCs.

[0047] Each swcap operation (37, 38, 39) may have its own independent integer (INT) scalar (40, 41, 42) that is associated with the corresponding swcap operation (37, 38, 39) during compilation.

[0048] Alternatively, all individual swcap MACC scalars (40, 41, 42) can be merged into a single layer-wise scalar (e.g., selecting the minimum value across all scalars), which is shared by all swcap MACC operations (37, 38, 39) within a given layer.

[0049] 5 is a flow diagram of an auto-search algorithm 45 for determining the optimal scalar. The "auto-search" algorithm may also be called an "auto-scale" algorithm. The algorithm 45 automatically searches for the optimal scaling / amplification factor.

[0050] Algorithm 45 includes a training / calibration (SW) portion 46 and an inference (HW) portion 47. A scalar 53 is provided to software scaling 55, which also receives input values ​​54. Scalar 53 may be a user-provided initialization value or the result of a previous loop of the automatic search algorithm for training 46. Software scaling 55 generates scaled values ​​57, which are provided to swcap analog MACC 56. Swcap analog MACC 56 generates MACC output 58, which is provided to ADC truncation 59.

[0051] ADC truncation 59 generates truncated output 60. At 61, it is determined whether MSB truncation is necessary based on truncated output 60. If MSB truncation occurs at 61 (e.g., "yes"), the method proceeds to 49. If MSB truncation does not occur at 61 (e.g., "no"), the method proceeds to 52. At 49, the INT scalar is reduced, and the method proceeds to 48. At 52, it is determined how many times MSB truncation at 61 did not occur. For example, at 52, it is determined whether "no" occurred at 61 "X" times or more, where "X" is a user-defined threshold. If "no" determinations at 61 occurred "X" times or more (e.g., "yes"), the method proceeds to 50. If "no" determinations at 61 did not occur "X" times or more (e.g., "no"), the method proceeds to 51. At 50, the INT scalar is incremented and the method transitions from 50 to 51. At 51, the method moves to the next batch and from 51, the method transitions to 48. At 48, the INT scalar moving average is updated and this INT scalar moving average will be used during inference.

[0052] Thus, if an overflow occurs during training / calibration time (output * scalar > threshold), the method involves reducing the scalar (49) and redoing the iteration, with or without updating the NN parameters. If MSB truncation occurs (61), the maximum threshold is exceeded, so the batch is repeated with a lower amplification in the next loop iteration. See item 85 in Figure 7, i.e., "if(P_abs>max_val)".

[0053] During inference 47 by inference hardware 44, input value 62 and optimal INT scalar 63 are provided to controller and programmable gain amplifier 64. The optimal scalar 63 used in inference 47 is the result of determining the INT scalar moving average determined at 48 during training 46. The moving average determined at 48 is truncated to determine the scalar 63 used in inference 47. Controller and programmable gain amplifier 64 generates a scaled value 65, which is provided as an input to swcap analog MACC 66. swcap analog MACC 66 generates a MACC output 67, which is provided as an input to ADC truncation 68. ADC truncation 68 generates a truncated output 69.

[0054] FIG. 6 is an exemplary software implementation 70 of an auto-scaling algorithm according to embodiments described herein.

[0055] Software 70 determines the INT scalar value for each layer (or atomic swcap operation) during training and / or calibration (either QAT or PTQ). If the GEMM output does not exceed a preselected ADC limit (e.g., max=32), the scalar is incremented by 1. If the GEMM output exceeds a threshold (preselected ADC limit), the scalar is decremented by 1. The optimal scalar to use during inference is the static value of the training scalar, determined as a running average truncated to an integer. A suitable static scalar for DNN inference with ADC truncation has been found.

[0056] Figure 7 shows an exemplary implementation of the truncation part of the automatic scaling algorithm described in this specification.

[0057] Figure 8 shows an exemplary Python implementation of part of the automatic scaling algorithm described in this specification. If no overflow occurs during batch processing, one or more parameters of the neural network and the learning rate (71) are updated within the part shown in Figure 8. Conversely, if an overflow occurs during batch processing, the batch is processed again using a lower amplification. The update of one or more parameters of the neural network refers to item 83, i.e., line optimizer.step(freeze_w_update=(global_step<args.freeze_g_steps+args.freeze_w_steps)). The process flow of the neural network includes sending a batch of examples through the network, obtaining the output and gradients, updating the parameters using the gradients, updating the learning rate according to the schedule, and processing the next batch. After batch processing, there are two options: 1) if there is no overflow, move to the next batch, or 2) if there is an overflow, repeat the same batch but use a lower amplification (see Figure 5, steps 49 and 51). [[ID=⑥]] [[ID=⑦]]

[0058] [[ID=⑧]] [[ID=⑨]]Figures 9A, 9B, and 9C each show options for scaling an analog signal with swcap hardware [[ID=⑩]] [[ID=⑪]]

[0059] [[ID=⑫]] [[ID=⑬]]Figure 9A shows the use of an amplifier for scaling an analog signal with switch capacitor hardware. Specifically, there are amplifiers (7⑤, 7⑥) at the input, so amplifier 7⑤ is applied to input 11 and / or amplifier 7⑥ is applied to input 12 and / or there is an amplifier (7⑦) at output 13. [[ID=⑭]] [[ID=⑮]]

[0060] [[ID=⑯]] 9B illustrates the use of an input multiplier (78) to scale an analog signal in switched-capacitor hardware. Input multiplier 78 may include Vdd / scale to represent the signal, where the scale is determined by QAT. Input multiplier 78 may be applied to either input 11 or input 12.

[0061] 9C shows charge sharing for scaling analog signals with switched-capacitor hardware. There is an accumulator or stored charge (79) for N iterations, where QAT determines N.

[0062] 10 is a circuit diagram of a circuit 100 for machine learning hardware. Circuit 100 includes N activation inputs 101 (K bits), N weights 104 (either from local storage or external inputs) (K bits), multipliers 110 (digital input, analog output), multiplier outputs 120—current or charge, summing bit lines 130, current or charge to voltage converters 140 (e.g., resistors, transimpedance amplifiers, or capacitors), sum voltages 141, analog-to-digital converters 150, and digital outputs 160 (M bits).

[0063] 11 is a circuit diagram of a circuit 200 illustrating a first embodiment of the implementation described herein, involving multiplication and amplification of the accumulation results. In circuit 200, there is a multiplier in the sum. Circuit 200 includes N activation inputs 201 (K bits), N weights 204 (either from local storage or external inputs) (K bits), multiplier 210 (digital input, analog output), multiplier output 220—current or charge, summing bit line 230, current-or-charge-to-voltage converter 240 (e.g., resistor, transimpedance amplifier, or capacitor), sum voltage 241, programmable gain amplifier 242, amplified voltage 244, amplifier gain controller 246, computer program / method 248 for determining optimal gain settings, A / D converter 250, and digital output 260 (M bits). Computer program 248 includes an auto-scaling algorithm 261.

[0064] Figure 12 is a circuit diagram of one embodiment of a switched-capacitor sum multiplier 280. Sum multiplier 280 is an exemplary implementation of programmable gain amplifier 242 shown in Figure 11. Sum multiplier 280 is composed of N capacitors to implement a multiplication factor "N." Capacitors 292-1, 292-2, 292-3, 292-4, and 292-N are shown.

[0065] The capacitors can be connected in two different ways: parallel and series. First, all capacitors are connected in parallel (285). The voltage input (282) is sampled by all capacitors simultaneously. The voltage across each capacitor is the same as the voltage input (282). If these capacitors are then configured in series (291) by one or more switches (295-1, 295-2, 295-3, 295-N-1, 295-N-2), the voltages across the capacitors are stacked, and the final output voltage (283) is N times the voltage input (282), thus realizing a sum multiplier 280. If the output is connected to an intermediate node, for example, the output of the Kth capacitor, the output voltage is K times the input voltage (see 2Vin, 3Vin, 4Vin, and S*Vin). Since K can take any value between 1 and N, the circuit has a programmable multiplication factor between 1 and N.

[0066] In Figure 12, when the switch is up (connecting the upper node of the nth capacitor to the upper node of the n+1th capacitor), the capacitors are in parallel. When the switch is down (connecting the upper node of the nth capacitor to the lower node of the n+1th capacitor), the capacitors are in series.

[0067] 13 is a circuit diagram of a circuit 300 illustrating a second embodiment of the implementation described herein with a voltage multiplier at the input. Circuit 300 has a multiplier at the input. Circuit 300 includes N activation inputs 301 (K bits) [range: 0 to V], voltage multiplier 302, multiplication activation input 303 [range: 0 to S*V] (S: scaling factor), N weights 304 (from either local storage or external input) (K bits), voltage multiplier controller 305, multiplier 310 (digital input, analog output), multiplier output 320—current or charge, summing bit line 330, current or charge to voltage converter 340 (e.g., resistor, transimpedance amplifier, or capacitor), sum voltage 341, A / D converter 350, and digital output 360 (M bits).

[0068] FIG. 14 is a circuit diagram of one embodiment of an input multiplier 400 using a level shifter. The input multiplier 400 is an exemplary implementation of the current or charge-to-voltage converter 340 shown in FIG. 13. If the input voltage 410 is zero, the output voltage 412 is also zero. If the input voltage 410 is V, the output voltage 412 depends on the circuit's supply voltage, which is the output voltage of the multiplexer 402. For example, if the output of the multiplexer 402 is K*V, the output voltage 412 is also K*V. Thus, the output voltage 412 is the input voltage 410 multiplied by K. Since K ranges from 1 to Smax, the input multiplier 400 has a multiplication factor ranging from 1 to Smax. The input multiplier 400 includes a multiplexer 402 that selects an input from among V 404, 2*V 406, and a maximum of Smax*V 408.

[0069] Figure 15 is a circuit diagram of a circuit 500-1 illustrating a first phase of operation of a third embodiment of the example described herein, which implements voltage sampling using a sum multiplier with parallel-connected capacitors. In Figure 15, the sum multiplier is used to amplify the multiplication and accumulation result 541. Figure 15 corresponds to circuit state 285 of Figure 12 (see item 585).

[0070] Figure 16 is a circuit diagram of circuit 500-2 illustrating a second phase of operation of a third embodiment of the example described herein that implements voltage multiplication using a sum multiplier with capacitors reconfigured to be connected in series. In Figure 16, the sum multiplier is used to amplify the multiplication and accumulation result 541. Figure 16 corresponds to circuit state 291 in Figure 12 (see item 591).

[0071] 15 and 16, given S input vectors (t=0...S-1), V(t) samples the sum of the t-th input vector (X(t)). Circuits (500-1, 500-2) include N activation inputs 501 (K bits), N weights 504 (from either local storage or external inputs) (K bits), summing bit lines 530, a current or charge-to-voltage converter 540 (e.g., a resistor, transimpedance amplifier, or capacitor), a sum voltage 541, capacitors 592-1, 592-2, 592-S-1, and 592-S, and an M-bit ADC converter 550. Circuit 500-1 includes an amplified voltage 544 generated from summing multiplier configuration 585, resulting in a digital output 560 after analog-to-digital conversion using ADC 550. Circuit 500-2 includes an amplified voltage 545 generated from a summing multiplier configuration 591, resulting in a digital output 561 after analog to digital conversion using ADC 550.

[0072] FIG. 17 is a graph showing NN accuracy on the evaluation set during training of a BERT-based INT4 model. FIG. 17 compares results from a baseline run (plot 804) with results without implementing embodiments described herein (plot 806) and results with implementing embodiments described herein (plot 802). The y-axis is F1. The x-axis is training iterations. Plot 804 corresponds to accuracy results without ADC truncation. Plot 802 corresponds to accuracy results with 7-bit ADC LSB truncation when autoscaling is implemented. Plot 806 corresponds to accuracy results with 7-bit ADC LSB truncation when autoscaling is not implemented. Plot 802 has a peak F1 value of 87.5%, plot 804 has a peak F1 value of 87.5%, and plot 806 has a peak F1 value of 81.8%. Therefore, in the embodiments described herein, a low precision ADC (with truncation) can be used to match the precision performance of a high precision ADC (without truncation).

[0073] Figure 18 is another graph showing NN accuracy on the evaluation set during training of a BERT-based INT4 model. The y-axis is F1, and the x-axis is training iterations. Plot 902 shows a fixed F1 value of 87.7%, plot 904 corresponds to LSB=8, MSB=0, plot 906 corresponds to LSB=6, MSB=1, and plot 908 corresponds to LSB=10, MSB=0. When LSB=6, MSB=1 (plot 906) is used, the results are closer to the int4 baseline of 87.7% (plot 902).

[0074] Figure 19 is a graph showing the convergence of quantization-aware training for the MobileNet-v1 (MB1) model. Plot 1002 corresponds to LSB=6 truncation without automatic search. Without automatic search, MB1 QAT fails to converge using LSB=6 truncation. The y-axis is training error, and the x-axis is training epochs. Figure 19 shows the convergence region 1004.

[0075] Figure 20 is a graph comparing performance results with and without automatic search scaling for post-training quantization of a BERT-based INT8 model. Plot 1202 corresponds to the implementation without automatic search scaling, and plot 1204 corresponds to the implementation with automatic search scaling. Plot 1201 corresponds to the baseline F1 score. Up to 32x amplification allows for equal accuracy for more aggressive LSB truncation.

[0076] FIG. 21 is a logic flow diagram for implementing method 1300 according to embodiments described herein. At 1310, the method includes determining a respective integer scalar value for a layer of a neural network among a plurality of layers of the neural network, where a plurality of respective integer scalar values ​​are determined for the plurality of layers of the neural network. At 1320, the method includes determining a matrix multiplication output of the neural network. At 1330, the method includes incrementing the respective integer scalar value by one if the matrix multiplication output does not exceed a threshold value of an analog-to-digital converter. At 1340, the method includes decrementing the respective integer scalar value by one if the matrix multiplication output exceeds a threshold value of an analog-to-digital converter. At 1350, the method includes determining a running average of the respective integer scalar values ​​determined for the plurality of layers. At 1360, the method includes determining a final integer scalar by truncating the running average to an integer, where the final integer scalar is used for amplification prior to analog-to-digital truncation during inference using the neural network.

[0077] The method 1300 may further include determining integer scalar values ​​for layers of the neural network during training of the neural network, where the training includes training that takes quantization into account.

[0078] The method 1300 may further include determining integer scalar values ​​for layers of the neural network during calibration of the neural network, where calibration includes post-training quantization.

[0079] The method 1300 may further include a step in which a layer of the neural network includes switched-capacitor operation.

[0080] The method 1300 may further include, in response to an overflow occurring during training or calibration, reducing the scalar and redoing the iteration of training the neural network or calibrating the neural network, with or without updating at least one parameter of the neural network.

[0081] The method 1300 may further include determining a first value associated with most-significant bit truncation; determining a threshold based on the second value associated with the bit accumulator and the first value associated with most-significant bit truncation; and determining an overflow when a third value associated with the matrix multiplication output exceeds the threshold. For example, after amplification, if the MACC output is 14 bits, the accumulator is 16, and the MSB truncation is 3 bits, the threshold is 16-3=13, and the 14-bit MACC output exceeds the 13-bit threshold. Thus, the method 1300 may further include the threshold being determined by subtracting the first value associated with most-significant bit truncation from the second value associated with the bit accumulator, where the first value is a first number of bits, the second value is a second number of bits, and the third value is a third number of bits.

[0082] The method 1300 may further include determining whether to apply most-significant-bit truncation to the matrix multiplication output; decreasing the respective integer scalar value in response to determining to apply most-significant-bit truncation to the matrix multiplication output; and increasing the respective integer scalar value in response to determining not to apply most-significant-bit truncation to the matrix multiplication output more than a threshold number of times.

[0083] 22 is a logic flow diagram for implementing a method 1400 in accordance with embodiments described herein. At 1410, the method includes receiving a first plurality of inputs (11, 201, 301, 501) representing activation input vectors. At 1420, the method includes receiving a second plurality of inputs (12, 204, 304, 504) representing weight input vectors. At 1430, the method includes generating, using an analog multiplier and accumulator (10, 240, 340, 540), an analog voltage representing a multiplication and accumulation result (13, 241, 244, 341, 541, 544) of the first plurality of inputs (11, 201, 301, 501) and the second plurality of inputs (12, 204, 304, 504). At 1440, the method includes using an analog-to-digital converter (14, 250, 350, 550) to convert analog voltage multiplication and accumulation results (13, 241, 244, 341, 541, 544) into digital signals (15, 260, 360, 560) using limited precision arithmetic during an inference operation (47) of the neural network. At 1450, the method includes determining (48, 49, 50, 70, 248, 261) scaling factors (53, 63, 80) used to amplify (64, 75, 76, 302) a first plurality of inputs (11, 201, 301, 501) or to amplify (9, 77, 242, 280, 400, 585, 591) an analog voltage multiplication and accumulation result (13, 241, 244, 341, 541, 544) during training or calibration (46) of the neural network.

[0084] Now, with reference to all figures, the following examples are disclosed herein.

[0085] Example 1 The apparatus includes: a first plurality of inputs representing activation input vectors; a second plurality of inputs representing weight input vectors; an analog multiplier and accumulator configured to generate first analog voltages representing first multiplication and accumulation results for the first inputs and the second inputs; a voltage multiplier configured to receive the first analog voltages and generate second analog voltages representing second multiplication and accumulation results by multiplying the first analog voltages by at least one scaling factor; an analog-to-digital converter configured to convert the second analog voltage multiplication and accumulation results to digital signals using limited precision arithmetic during a neural network inference operation; and a hardware controller configured to determine at least one scaling factor based on the first multiplication and accumulation results or a software controller configured to determine at least one scaling factor based on the first multiplication and accumulation results.

[0086] Example 2 2. The apparatus of example 1, wherein the at least one scaling factor comprises a plurality of independent scaling factors determined during training of the neural network, one independent scaling factor for each switched-capacitor operation of a layer of the neural network including multiple layers.

[0087] Example 3 3. The apparatus of any of Examples 1-2, wherein the apparatus determines at least one scaling factor during training of the neural network.

[0088] Example 4 4. The apparatus of example 3, wherein at least one scaling factor determined during training and used during inference is an integer value.

[0089] Example 5 5. The apparatus of any one of Examples 1 to 4, further comprising an accumulated storage charge configured to accumulate a charge corresponding to the second analog voltage multiplication and accumulation result over multiple iterations.

[0090] Example 6 6. The apparatus of any of embodiments 1-5, further comprising a programmable controller configured to control the voltage multiplier based on at least one scaling factor.

[0091] Example 7 7. The apparatus of any one of embodiments 1 to 6, wherein the voltage multiplier comprises a plurality of switched capacitors configured in series or parallel.

[0092] Example 8 The apparatus includes a first plurality of inputs representing original activation input vectors; a plurality of voltage multipliers that receive the first plurality of inputs and generate a second plurality of inputs by multiplying voltages of the original activation input vectors by at least one scaling factor; a third plurality of inputs that represent weight input vectors; an analog multiplier and accumulator that generate analog voltages that represent multiplication and accumulation results for the second and third inputs; an analog-to-digital converter configured to convert the analog voltage multiplication and accumulation results into digital signals using limited precision arithmetic during neural network inference operations; and a hardware controller configured to determine at least one scaling factor based on the multiplication and accumulation results, or a software controller configured to determine at least one scaling factor based on the multiplication and accumulation results.

[0093] Example 9 9. The apparatus of example 8, wherein the at least one scaling factor comprises a plurality of independent scaling factors, one independent scaling factor for each switched-capacitor operation of a layer of a neural network including multiple layers.

[0094] Example 10 10. The apparatus of example 9, wherein the multiple independent scaling coefficients are determined during training of the neural network.

[0095] Example 11 11. The apparatus of any of Examples 8-10, wherein the apparatus determines at least one scaling factor during training of the neural network.

[0096] Example 12 12. The apparatus of example 11, wherein at least one scaling factor determined during training and used during inference is an integer value.

[0097] Example 13 13. The apparatus of any one of Examples 8 to 12, further comprising an accumulated storage charge configured to accumulate a charge corresponding to the analog voltage multiplication and accumulation result over multiple iterations.

[0098] Example 14 An apparatus described in any of Examples 8 to 13, further comprising at least one programmable controller configured to control the plurality of voltage multipliers based on at least one scaling factor.

[0099] Example 15 The method includes receiving a first plurality of inputs representing an activation input vector; receiving a second plurality of inputs representing a weight input vector; generating, using analog multipliers and accumulators, analog voltages representing multiplication and accumulation results of the first plurality of inputs and the second plurality of inputs; converting, using an analog-to-digital converter, the analog voltage multiplication and accumulation results to digital signals using limited precision arithmetic during inference operations of the neural network; and determining at least one scaling factor used to amplify the first plurality of inputs or the analog voltage multiplication and accumulation results during training or calibration of the neural network.

[0100] Example 16 16. The method of example 15, further comprising determining a plurality of independent scaling coefficients, one independent scaling coefficient for each switched-capacitor operation of a layer of the neural network including a plurality of layers, wherein at least one scaling coefficient comprises the plurality of independent scaling coefficients.

[0101] Example 17 17. A method according to any one of Examples 15 to 16, wherein the step of amplifying the first plurality of inputs comprises using a plurality of voltage multipliers to generate the amplified first plurality of inputs by multiplying the voltages of the activation input vector by at least one scaling factor, and the method further comprises using an analog multiplier and accumulator to generate an analog voltage multiplication and accumulation result for the amplified first plurality of inputs.

[0102] Example 18 18. The method of any of Examples 15-17, wherein amplifying the analog voltage comprises using a voltage multiplier to generate an amplified analog voltage multiplication and accumulation result by applying at least one scaling factor to the analog voltage multiplication and accumulation result, and the method further comprises using an analog-to-digital converter to convert the amplified analog voltage multiplication and accumulation result to a digital signal using limited precision arithmetic during an inference operation of the neural network.

[0103] Example 19 20. The method of embodiment 18, further comprising: configuring the multiple switched capacitors of the voltage multiplier in series; or configuring the multiple switched capacitors of the voltage multiplier in parallel.

[0104] Example 20 20. The method of any of Examples 15-19, further comprising accumulating a charge corresponding to an analog voltage multiplication and accumulation result over multiple iterations.

[0105] References to "computer," "processor," etc. should be understood to encompass computers having different architectures, such as single / multi-processor architectures and sequential or parallel architectures, as well as specialized circuitry such as field-programmable gate arrays (FPGAs), application specific circuits (ASICs), signal processing devices, and other processing circuitry. References to computer programs, instructions, code, etc. should be understood to encompass software for a programmable processor, or firmware, e.g., either instructions for a processor or the programmable content of a hardware device, such as the configuration settings of a fixed function device, gate array, or programmable logic device.

[0106] A memory as described herein may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, non-transitory memory, transient memory, fixed memory, and removable memory. A memory may include a database for storing data.

[0107] As used herein, a circuit may refer to (a) a hardware circuit implementation, such as an analog and / or digital circuit implementation, and (b) a combination of circuitry and software (and / or firmware), such as (if applicable) (i) a combination of a processor, or (ii) a portion of a processor / software including a digital signal processor, software, and memory that work together to cause a device to perform various functions, and (c) a circuit that requires software or firmware to operate, even if the software or firmware is not physically present, such as a microprocessor or portion of a microprocessor. As a further example, as used herein, a circuit also covers implementations of only one processor (or multiple processors), or a portion of a processor and its (or their) accompanying software and / or firmware. For example, a circuit also covers, where applicable to the particular element, a baseband integrated circuit or an application processor integrated circuit for a mobile phone, or a similar integrated circuit in a server, cellular network device, or another network device.

[0108] For example, a list of abbreviations, where abbreviations may be appended to each other using dashes or hyphens ("-"), or other characters may be appended. AD Analog to Digital ADC Analog-to-Digital Converter ASIC Application Specific Integrated Circuit b bits (e.g., 8b) Bidirectional Encoder Representation from the BERT Transformer BMM Batch Matrix Multiplication Capacitor DNN Deep Neural Network ep Epoch F1: Harmonic mean of precision and recall FPGA Field Programmable Gate Array GEMM Generalized Matrix Multiplication HW Hardware INT integer LSB least significant bit MACC Multiply and Accumulate MB1 MobileNet-v1 (neural network model) MSB Most Significant Bit NN neural network prec. precision PTQ Post-Training Quantization QAT Quantization-aware training SC-PT Switched Capacitor Processing Tile (core hardware component) swcap Switch Capacitor SW Software trunc truncation V Voltage Vdd power supply voltage W is the weight vector - input to the multiplication and accumulation (MACC) operation X input vector - input to the multiply and accumulate (MACC) operation

[0109] In the foregoing description, numerous specific details have been set forth, such as particular structures, components, materials, dimensions, process steps, etc., to provide a thorough understanding of the exemplary embodiments disclosed herein. However, those skilled in the art will understand that the exemplary embodiments disclosed herein may be practiced without these specific details. Furthermore, details of well-known structures or process steps may be omitted or not described in order to avoid obscuring the presented embodiments.

[0110] The description of the present invention has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the precise form disclosed. Many modifications and variations will become apparent to those skilled in the art without departing from the scope of the invention. The embodiments have been chosen and described to best explain the principles and practical applications of the invention and to enable those skilled in the art to understand the invention in various embodiments with various modifications suited to the particular uses contemplated.

Claims

1. a first plurality of inputs representing activation input vectors; a second plurality of inputs representing a weight input vector; an analog multiplier and accumulator that generates a first analog voltage representing a first multiplication and accumulation result for the first input and the second input; a voltage multiplier that receives the first analog voltage and multiplies the first analog voltage by at least one scaling factor to generate a second analog voltage that represents a second multiplication and accumulation result; an analog-to-digital converter configured to convert the second analog voltage multiplication and accumulation result to a digital signal using limited precision arithmetic during a neural network inference operation; and a hardware controller configured to determine the at least one scaling factor based on the first multiplication and accumulation result, or a software controller configured to determine the at least one scaling factor based on the first multiplication and accumulation result. An apparatus comprising:

2. 2. The apparatus of claim 1, wherein the at least one scaling factor comprises a plurality of independent scaling factors determined during training of the neural network, one independent scaling factor for each switched-capacitor operation of a layer of the neural network including multiple layers.

3. The apparatus of claim 1 , wherein the apparatus determines the at least one scaling factor during training of a neural network.

4. The apparatus of claim 3 , wherein the at least one scaling factor determined during training and used during inference is an integer value.

5. The apparatus of claim 1 , further comprising an accumulated charge store configured to accumulate a charge corresponding to the second analog voltage multiplication and accumulation result over multiple iterations.

6. The apparatus of claim 1 , further comprising a programmable controller configured to control the voltage multiplier based on the at least one scaling factor.

7. 10. The apparatus of claim 1, wherein the voltage multiplier comprises a plurality of switched capacitors configured in series or parallel.

8. a first plurality of inputs representing the original activation input vector; a plurality of voltage multipliers that receive the first plurality of inputs and generate a second plurality of inputs by multiplying the voltages of the original activation input vector by at least one scaling factor; a third plurality of inputs representing a weight input vector; an analog multiplier and accumulator for generating an analog voltage representing a multiplication and accumulation result for said second input and said third input; an analog-to-digital converter configured to convert the analog voltage multiplication and accumulation results to digital signals using limited precision arithmetic during neural network inference operations; and a hardware controller configured to determine the at least one scaling factor based on the multiplication and accumulation result, or a software controller configured to determine the at least one scaling factor based on the multiplication and accumulation result. An apparatus comprising:

9. 9. The apparatus of claim 8, wherein the at least one scaling factor comprises a plurality of independent scaling factors, one independent scaling factor for each switched-capacitor operation of a layer of a neural network including multiple layers.

10. The apparatus of claim 9 , wherein the plurality of independent scaling factors are determined during training of a neural network.

11. The apparatus of claim 8 , wherein the apparatus determines the at least one scaling factor during training of a neural network.

12. The apparatus of claim 11 , wherein the at least one scaling factor determined during training and used during inference is an integer value.

13. The apparatus of claim 8 , further comprising an accumulated charge store configured to accumulate a charge corresponding to the analog voltage multiplication and accumulation result over multiple iterations.

14. The apparatus of claim 8 , further comprising at least one programmable controller configured to control the plurality of voltage multipliers based on the at least one scaling factor.

15. receiving a first plurality of inputs representing an activation input vector; receiving a second plurality of inputs representing a weight input vector; using an analog multiplier and accumulator to generate an analog voltage representing the result of multiplying and accumulating the first plurality of inputs and the second plurality of inputs; using an analog-to-digital converter to convert the results of the analog voltage multiplication and accumulation into digital signals using limited precision arithmetic during the inference operation of the neural network; and determining at least one scaling factor used to amplify the first plurality of inputs or amplify the analog voltage multiplication and accumulation result during training or calibration of the neural network; A method for providing the above.

16. 16. The method of claim 15, further comprising determining a plurality of independent scaling factors, one independent scaling factor for each switched-capacitor operation of a layer of a neural network including a plurality of layers, wherein the at least one scaling factor comprises the plurality of independent scaling factors.

17. 16. The method of claim 15, wherein amplifying the first plurality of inputs comprises using a plurality of voltage multipliers to generate amplified first plurality of inputs by multiplying voltages of the activation input vector by the at least one scaling factor, the method further comprising using the analog multiplier and accumulator to generate the analog voltage multiplication and accumulation results for the amplified first plurality of inputs.

18. 16. The method of claim 15, wherein amplifying the analog voltage comprises using a voltage multiplier to generate an amplified analog voltage multiplication and accumulation result by applying the at least one scaling factor to the analog voltage multiplication and accumulation result, the method further comprising using the analog-to-digital converter to convert the amplified analog voltage multiplication and accumulation result to the digital signal using the limited precision arithmetic during the inference operation of the neural network.

19. configuring a plurality of switched capacitors of the voltage multiplier in series; or 20. The method of claim 18, further comprising configuring the plurality of switched capacitors of the voltage multiplier in parallel.

20. 16. The method of claim 15, further comprising accumulating a charge corresponding to the analog voltage multiplication and accumulation result over multiple iterations.

21. A computer program product comprising instructions, said instructions executable by a processor to cause said processor to perform the method of any of claims 15 to 20.