Scalable switched capacitor calculation core for accurate and efficient deep learning inference

Through the automatic search algorithm to determine the optimal integer scalar and implement its merging in the hardware, the problem of low-precision ADC truncation affects the accuracy of neural networks is solved, and the same neural network performance as high-precision ADC is achieved.

CN120226019APending Publication Date: 2025-06-27INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380081907.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-29
Filing Date
2023-11-23
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, when using low-precision ADCs for neural network inference, the truncation operation of the ADC will reduce the accuracy of simulated MACC output, resulting in a decrease in the accuracy of neural network inference.

Method used

Automatic search algorithm (auto-scaling) is used to determine the optimal integer scalar for ADC truncation, and by modifying the analog signal during the switching capacitor simulation calculation, the determined scalars are combined into the hardware, specifically through an amplifier, input multiplier or charge sharing.

Benefits of technology

Effective execution of ADC truncation without decreasing the accuracy of neural network inference is significantly improved, for example, during training of BERT-based INT4 models, automatic scaling achieves the same F1 performance as high-precision ADCs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226019A_ABST
    Figure CN120226019A_ABST
Patent Text Reader

Abstract

An apparatus comprising: a first plurality of inputs representing an activation input vector; a second plurality of inputs representing weighted input vectors; an analog multiplier and accumulator for generating a first analog voltage representing a first multiplication and accumulation result of the first input and the second input; a voltage multiplier acquiring the first analog voltage and generating a second analog voltage representing a second multiplication and accumulation result by multiplying at least one scaling factor by the first analog voltage; an analog-to-digital converter configured to convert the second analog voltage multiplication and accumulation result into a digital signal using a finite precision operation during a neural network inference operation; and a hardware controller configured to determine at least one scaling factor based on the first multiplication and accumulation result, or a software controller configured to determine at least one scaling factor based on the first multiplication and accumulation result.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION

[0001] The exemplary embodiments described herein generally relate to machine learning hardware device design and integrated circuit design, and more particularly, to a scalable switched-capacitor computing core for accurate and efficient deep learning inference. SUMMARY OF THE INVENTION

[0002] In one aspect, a device includes: a first plurality of inputs representing an activation input vector; a second plurality of inputs representing a weight input vector; an analog multiplier and accumulator for generating a first analog voltage representing a first multiplication and accumulation result of the first input and the second input; a voltage multiplier for obtaining the first analog voltage and generating a second analog voltage representing a second multiplication and accumulation result by multiplying at least one scaling factor with the first analog voltage; an analog-to-digital converter configured to convert the second analog voltage multiplication and accumulation result into a digital signal using finite-precision arithmetic during a neural network inference operation; and a hardware controller or a software controller, the hardware controller being configured to determine at least one scaling factor based on the first multiplication and accumulation result, and the software controller being configured to determine at least one scaling factor based on the first multiplication and accumulation result.

[0003] In another aspect, a device includes: a first plurality of inputs representing an original activation input vector; a plurality of voltage multipliers that obtain the first plurality of inputs and generate a second plurality of inputs by multiplying at least one scaling factor with the voltage of the original activation input vector; a third plurality of inputs representing a weight input vector; an analog multiplier and accumulator for generating an analog voltage representing a multiplication and accumulation result of the second input and the third input; an analog-to-digital converter configured to convert the analog voltage multiplication and accumulation result into a digital signal using finite-precision arithmetic during a neural network inference operation; and a hardware controller or a software controller, the hardware controller being configured to determine the at least one scaling factor based on the multiplication and accumulation result, and the software controller being configured to determine the at least one scaling factor based on the multiplication and accumulation result.

[0004] In another aspect, a method includes: receiving a first plurality of inputs representing an activation input vector; receiving a second plurality of inputs representing a weight input vector; generating, using an analog multiplier and accumulator, an analog voltage representing a multiplication and accumulation result of the first plurality of inputs and the second plurality of inputs; converting, using an analog-to-digital converter, the analog voltage multiplication and accumulation result into a digital signal using finite-precision arithmetic during an inference operation of a neural network; and determining, during training or calibration of the neural network, at least one scaling factor for amplifying the first plurality of inputs or the analog voltage multiplication and accumulation result. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] The foregoing and other aspects of the exemplary embodiments become more apparent in the following detailed description when read in conjunction with the accompanying drawings, in which:

[0006] Figure 1 A high-level diagram of a mixed-signal switched-capacitor multiplier and accumulator is depicted;

[0007] Figure 2 A 16-bit accumulator and an 8-bit accumulator with severe truncation are depicted;

[0008] Figure 3 A 16-bit accumulator with a scaled distribution of values as the input to an ADC, and an 8-bit accumulator, are depicted;

[0009] Figure 4 Processing of scalars in a DNN layer is depicted;

[0010] Figure 5 A flowchart of an automatic search algorithm for determining an optimal scalar;

[0011] Figure 6 An example software implementation of an automatic scaling algorithm based on the examples described herein;

[0012] Figure 7 An example implementation of a truncated portion of the automatic scaling algorithm described herein is depicted;

[0013] Figure 8 An example implementation of a portion of the automatic scaling algorithm described herein is depicted;

[0014] Figure 9A Using an amplifier to scale an analog signal in switched-capacitor hardware is depicted;

[0015] Figure 9B Using an input multiplier to scale an analog signal in switched-capacitor hardware is depicted;

[0016] Figure 9C Charge sharing to scale an analog signal in switched-capacitor hardware is depicted;

[0017] Figure 10 A circuit diagram for machine learning hardware;

[0018] Figure 11 A circuit diagram showing a first embodiment of the example described herein, in which the multiplication and accumulation results are amplified;

[0019] Figure 12 A circuit diagram of an embodiment using a switched-capacitor sum and multiplier;

[0020] Figure 13is a circuit diagram showing a second embodiment of the examples described herein, which has a voltage multiplier at the input;

[0021] Figure 14 is a circuit diagram of one embodiment of an input multiplier using a level shifter;

[0022] Figure 15 is a circuit diagram showing the first operation phase of the third embodiment of the examples described herein, which uses a sum multiplier with capacitors connected in parallel to implement voltage sampling;

[0023] Figure 16 is a circuit diagram showing the second operation phase of the third embodiment of the examples described herein, which uses a sum multiplier to implement voltage multiplication, where the capacitors are reconfigured to be connected in series;

[0024] Figure 17 is a graph showing the NN accuracy performance results, comparing the results with and without the implementation of the examples described herein;

[0025] Figure 18 is another graph showing the NN accuracy performance results;

[0026] Figure 19 is a graph showing the quantization-aware training convergence without implementing the examples described herein;

[0027] Figure 20 is a graph showing the comparison of the performance results with and without automatic search scaling;

[0028] Figure 21 is a logic flow diagram of an implementation method based on the examples described herein; and

[0029] Figure 22 is a logic flow diagram of an implementation method based on the examples described herein. Detailed Description

[0030] The term "exemplary" is used herein to mean "serving as an example, instance, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. All embodiments described in this detailed description are provided to enable those skilled in the art to make or use the exemplary embodiments of the present invention, rather than to limit the scope of the present invention defined by the claims.

[0031] A low-precision ADC (<16 bits) is required to limit ADC power consumption and implement a high-energy-efficiency switched-capacitor computing core. When the analog output of the switched-capacitor MACC falls outside a predefined voltage range, this low-precision ADC truncates the analog output of the switched-capacitor MACC and provides a digital output represented by fewer than 16 bits. This truncation operation reduces the precision of the analog MACC output and may lead to a reduction in precision during neural network inference. Therefore, there is a need for hardware and software that can perform ADC truncation without reducing the accuracy of neural network inference.

[0032] Accordingly, this document describes a method for determining an optimal integer scalar for ADC truncation via an automatic search algorithm ("auto-scaling") and related hardware implementations. This document discloses the process steps for determining at least one optimal integer scalar, how to process layer-by-layer MACC, and how to handle overflows. This document describes three options for incorporating the determined scalar into the hardware by modifying the analog signal during switched-capacitor analog computing: (1) an amplifier (at the input or at the output), (2) an input multiplier, and (3) charge sharing.

[0033] The challenge addressed by the examples described in this document is that SC-PT core ADC truncation affects accuracy. The examples described in this document make full use of the SC-PT core ADC precision. The MACC input or output is scaled up by an integer factor, and there are various implementation options for this.

[0034] Figure 1 A high-level diagram of a mixed-signal switched-capacitor multiplier and accumulator (10) is depicted. The mixed-signal switched-capacitor multiplier and accumulator (10) takes as input an input vector including 512 values each of 4 bits or 512 × 4b (11) and a weight vector including 512 values each of 4 bits or 512 × 4b (12). The input 11 may include added shifts. The output (13) from the mixed-signal switched-capacitor multiplier and accumulator (10) is provided to a low-precision analog-to-digital converter 14. The low-precision analog-to-digital converter 14 may be an 8-bit ADC. For example, the ADC being 8-bit is an illustrative example since the size of the ADC can be a size corresponding to other limited or low precision. The output 13 of the mixed-signal switched-capacitor multiplier and accumulator (10) is an analog voltage representing the MACC result. The result of the low-precision analog-to-digital converter (15) (e.g., 8-bit) has the following form , where R corresponds to the result, is an input vector that includes N values, each of 4 bits (in this example, N = 512), is a weight vector that includes N values, each of 4 bits (in this example, N = 512), and the subscript identifies the accumulation over the N products such that and each element are first multiplied and then all N products are summed together. This operation as a whole is the so-called MACC (multiply-and-accumulate).

[0035] The output (13) can be based on several factors, such as the application to different DNN layers (such as linear layers and BMM layers). The linear layer can be based on a scale output or add numbers upon activation. The BMM layer can be based on a scale output.

[0036] The mixed-signal switched-capacitor multiplier and accumulator (10) can be coupled to an amplifier circuit (9) having amplifiers that support different magnifications. In particular, an amplifier can be added to scale the analog signal with a software-defined magnification.

[0037] Figure 2 depicts a 16-bit accumulator (20) and an 8-bit accumulator (21). As Figure 2 shown, the 16-bit accumulator 20 includes bits 24-1, 24-2, 24-3, 24-4, 24-5, 24-6, 24-7, 24-8, 24-9, 24-10, 24-11, 24-12, 24-13, 24-14, 24-15, and 24-16. The 8-bit accumulator 21 includes bits 24-2, 24-3, 24-4, 24-5, 24-6, 24-7, 24-8, and 24-9. In Figure 2 is shown an example of the distribution of values in the input of an ADC with severe truncation, via bits 24-6, 24-7, 24-8, 24-9, 24-10, 24-11, 24-12, 24-13, 24-14, 24-15, and 24-16.

[0038] Figure 2Shows a 16-bit accumulator (20) that does not implement the examples described herein. Bits indicated as 24-6, 24-7, 24-8, 24-9, 24-10, 24-11, 24-12, 24-13, 24-14, 24-15, and 24-16 show the hypothetical degree (range) of the value of the MACC output (where MACC corresponds to a multiply-accumulate operation). Two truncation thresholds (MSB truncation 22 and LSB truncation 23) determine the conversion of the MACC output from an analog voltage to a digital representation, which is then processed by a low-precision ADC (in this example, an 8-bit ADC). In this case, many bits are truncated (i.e., after the digital conversion, the MACC output value is approximated by a highly truncated representation). This results in a high MACC error and poor neural network accuracy compared to a non-approximated MACC. As an example, refer to Figure 17 the experimental results in, where compared to the non-truncated results (curve 804) obtained using a 16-bit ADC, 7-bit LSB truncation (curve 806) gives a ~6.9% F1 performance (F1 is a measure of accuracy).

[0039] Figure 3 Depicts a 16-bit accumulator (25) and an 8-bit accumulator (26). Depicted are the MSB truncation threshold 27 and the LSB truncation threshold 28, which determine the conversion of the MACC output from an analog voltage to a digital representation, which is subsequently processed by a low-precision ADC. The 16-bit accumulator (25) includes bits 29-1, 29-2, 29-3, 29-4, 29-5, 29-6, 29-7, 29-8, 29-9, 29-10, 29-11, 29-12, 29-13, 29-14, 29-15, and 29-16. The 8-bit accumulator 26 includes bits 29-2, 29-3, 29-4, 29-5, 29-6, 29-7, 29-8, and 29-9. Bits 29-2, 29-3, 29-4, 29-5, 29-6, 29-7, 29-8, 29-9, 29-10, and 29-11 show the scaled distribution of the values having an input to the ADC.

[0040] Figure 3 Shows a 16-bit accumulator (25) with the implementation of the examples described herein. In Figure 3 , the MACC value is scaled up by an integer factor, then truncated by the analog-to-digital conversion of a low-precision ADC, and then the result is shifted down. This significantly improves performance, as shown in Figure 17 and Figure 18 shown. For example, compared to the non-truncated results (curve 804) obtained using a 16-bit ADC, 8-bit LSB truncation (curve 802) gives only a ~0.1% F1 performance (a slight degradation).

[0041] The output distribution varies at each DNN layer. Fixed ADC truncation results in severe degradation. ADC power savings mainly come from LSB truncation. It is beneficial to truncate the LSB rather than the MSB, for example, to save ADC power.

[0042] Figure 2 The shaded bits 24 - 6, 24 - 7, 24 - 8, 24 - 9, 24 - 10, 24 - 11, 24 - 12, 24 - 13, 24 - 14, 24 - 15, and 24 - 16 in Figure 3 and the shaded bits 29 - 2, 29 - 3, 29 - 4, 29 - 5, 29 - 6, 29 - 7, 29 - 8, 29 - 9, 29 - 10, and 29 - 11 in Figure 1 represent the bits required to cover the assumed distribution of the input to the ADC, or equivalently, the output of the MACC ( Figure 2 the marker 13 in Figure 2 ). For example, the input to the ADC can have low values and occupy the lowest bits ( Figure 2 24 - 6 to 24 - 16 in

[0043] Figure 3 ). The other bits (

[0044] Figure 4 24 - 1 to 24 - 5 in

[0045] are not used. When the LSB bits are truncated, this results in poor performance because several of the used bits ( Figure 2 the shaded bits 24 - 6, 24 - 7, 24 - 8, 24 - 9, 24 - 10, 24 - 11, 24 - 12, 24 - 13, 24 - 14, 24 - 15, and 24 - 16) are truncated.

[0043] Figure 3 "Scaling" in

[0044] Figure 4 corresponds to magnification, where each value of the input to the ADC is multiplied by a magnification factor or scaling magnification. There may be a distribution to the input of the ADC (or the MACC output, which is the same). If many MACC operations are performed with different inputs, the MACC output is different for each individual operation performed. Each distribution represents a set of assumed MACC outputs, magnified or not magnified.

[0045] The scalar general processing of the DNN layer is depicted. The DNN layer 35 with an accumulation size of N (e.g., N is an integer) is divided 36 into a swcap operation 37 with L accumulations (e.g., L is an integer), a swcap operation 38 with L accumulations, and a swcap operation 39 with L accumulations. The swcap operation 37 is associated with the scalar A (40), the swcap operation 38 is associated with the scalar B (41), and the swcap operation 39 is associated with the scalar C (42).

[0046] The GEMM performed by the DNN layer may require a number of accumulations where N > L. If so, the MACC layer is split into several swcap atomic MACCs.

[0047] Each swcap operation (37, 38, 39) can have its own independent integer (INT) scalar (40, 41, 42), which is associated with the corresponding swcap operation (37, 38, 39) during compilation.

[0048] Alternatively, all individual SWCAP_MACC scalars (40, 41, 42) can be combined into a single per-layer scalar (e.g., selecting the minimum across all scalars), which is shared by all SWCAP_MACC operations (37, 38, 39) in a given layer.

[0049] Figure 5 is a flowchart of an automatic search algorithm 45 for determining the optimal scalar. The "automatic search" algorithm can also be referred to as the "automatic scaling" algorithm. Algorithm 45 automatically searches for the optimal scaling / magnification factor.

[0050] Algorithm 45 includes a training / calibration (SW) section 46 and an inference (HW) section 47. A scalar 53 is provided to a software scaling 55, which also receives an input value 54. The scalar 53 is either a user-provided initialization value or the result of a previous cycle of the automatic search algorithm for training 46. The software scaling 55 generates a scaled value 57, which is provided to a swcap simulated MACC 56. The swcap simulated MACC 56 generates a MACC output 58 that is provided to an ADC truncation 59.

[0051] The ADC truncation 59 generates a truncated output 60. At 61, it is determined whether there is an MSB truncation based on the truncated output 60. If there is an MSB truncation at 61 (e.g., "yes"), the method transitions to 49. If there is no MSB truncation at 61 (e.g., "no"), the method transitions to 52. At 49, the INT scalar is decreased, and the method transitions to 48. At 52, the number of times it is determined at 61 that the MSB truncation has not occurred is determined. For example, at 52, a determination is made as to whether "no" has been determined to have occurred more than 'X' times at 61, where 'X' is a user-defined threshold. If at 52 it is determined that the "no" determination at 61 has occurred more than "X" times (e.g., "yes"), the method transitions to 50. If at 52 it is determined that the "no" determination at 61 has not occurred more than "X" times (e.g., "no"), the method transitions to 51. At 50, the INT scalar is increased, and the method transitions from 50 to 51. At 51, the method moves to the next batch, and the method transitions from 51 to 48. At 48, the INT scalar moving average is updated, which will be used during inference.

[0052] Thus, in the case of an overflow during training / calibration time (output * scalar > threshold), the method includes reducing the scalar (49) with or without updating the NN parameters and repeating the iteration. If an MSB truncation has occurred (61), it has exceeded the maximum threshold, and the batch is repeated with a lower magnification in the next loop iteration. Refer to Figure 7 item 85 of, or "if (P_abs > max_val)".

[0053] During inference 47 using the inference hardware 44, the input value 62 and the optimal INT scalar 63 are provided to the controller and programmable gain amplifier 64. The optimal scalar 63 used at inference 47 is the result of the determination of the INT scalar moving average determined at 48 during training 46. The moving average determined at 48 is truncated in order to determine the scalar 63 to be used at inference 47. The controller and programmable gain amplifier 64 generate a scaled value 65, which is provided as an input to the swcap analog MACC 66. The swcap analog MACC 66 generates a MACC output 67, which is provided as an input to the ADC truncation 68. The ADC truncation 68 generates a truncated output 69.

[0054] Figure 6 is an example software implementation 70 of an automatic scaling algorithm based on the examples described herein.

[0055] Software 70 decides the INT scalar value for each layer (or atomic swcap operation) during training and / or calibration (QAT or PTQ). If the GEMM output does not exceed the preselected ADC limit (e.g. max=32), the scalar is increased by 1. If the GEMM output exceeds the threshold (preselected ADC limit), the scalar is decreased by 1. The best scalar to use at inference time is a static value, which is determined as the moving average of the training scalars, truncated to an integer. Find a suitable static scalar for DNN inference with ADC truncation.

[0056] Figure 7 Depicted is an example implementation of a truncation portion of the autoscaling algorithm described herein.

[0057] Figure 8 Depicts an example Python implementation of a portion of the autoscaling algorithm described in this paper. If no overflow occurs during batch processing, then Figure 8 The one or more parameters of the neural network and the learning rate are updated in the portion shown (71). If, on the other hand, overflow occurs during the processing of a batch, the batch is processed again using a lower augmentation. The update of one or more parameters of the neural network is described in item 83, or the line optimizer.step(freeze_w_update=(global_step <args.freeze_g_steps+args.freeze_w_steps))。神经网络处理流程包括通过网络发送一批示例,获得输出和梯度,使用梯度更新参数,根据时间表更新学习速率,以及处理下一批。在处理批次之后,存在2个选项:1)如果没有溢出:移动到下一批次,或2)如果溢出:重复相同批次但使用较低扩增(参见 Figure 5 , steps 49 and 51).

[0058] Figure 9A , Figure 9B and Figure 9C Each depicts options for scaling analog signals in swcap hardware.

[0059] Figure 9A Depicted is the use of amplifiers to scale analog signals in switched capacitor hardware. In particular, there are amplifiers (75, 76) on the inputs, such that amplifier 75 is applied to input 11 and / or amplifier 76 is applied to input 12, and / or there is an amplifier (77) at output 13.

[0060] Figure 9BDepicts the use of an input multiplier (78) to scale an analog signal in switched capacitor hardware. The input multiplier 78 may include Vdd / scale to represent the signal, where the scale is determined by the QAT. The input multiplier 78 may be applied to input 11 or input 12.

[0061] Figure 9C Depicts charge sharing to scale an analog signal in switched capacitor hardware. There is an accumulator or charge storage (79) for N iterations, where N is determined by the QAT.

[0062] Figure 10 Is a circuit diagram of circuit 100 for machine learning hardware. Circuit 100 includes N activation inputs 101 (K bits), N weights 104 (from local storage or external input) (K bits), a multiplier 110 (digital input, analog output), multiplier output 120 - current or charge, a summing bit line 130, a current or charge to voltage converter 140 (such as a resistor, transimpedance amplifier or capacitor), a summing voltage 141, an AD converter 150 and a digital output 160 (M bits).

[0063] Figure 11 Is a circuit diagram of circuit 200 showing a first embodiment of the example described herein, where the multiplication and accumulation results are amplified. In circuit 200, there is a multiplier at the sum. Circuit 200 includes N activation inputs 201 (K bits), N weights 204 (from local storage or external input) (K bits), a multiplier 210 (digital input, analog output), multiplier output 220 - current or charge, a summing bit line 230, a current or charge to voltage converter 240 (such as a resistor, transimpedance amplifier or capacitor), a summing voltage 241, a programmable gain amplifier 242, an amplified voltage 244, an amplifier gain controller 246, a computer program / method 248 for determining the optimal gain setting, an AD converter 250 and a digital output 260 (M bits). The computer program 248 includes an automatic scaling algorithm 261.

[0064] Figure 12 Is a circuit diagram of one embodiment using a switched capacitor sum multiplier 280. The sum multiplier 280 is Figure 11 An example implementation of the programmable gain amplifier 242 shown. The sum multiplier 280 includes N capacitors to achieve a multiplication factor of "N". Capacitors 292 - 1, 292 - 2, 292 - 3, 292 - 4 and 292 - N are shown.

[0065] Each capacitor can be connected in two different ways - in parallel and in series. First, all the capacitors are connected in parallel (285). The voltage input (282) is sampled simultaneously by all the capacitors. The voltage across each capacitor is the same as the voltage input (282). Then, if one or more of these capacitors are configured in series by switches (295-1, 295-2, 295-3, 295-N-1, 295-N-2) (291), each voltage across the capacitors is stacked such that the final output voltage (283) becomes N times the voltage input (282), thus implementing a summing multiplier 280. When the output is tapped at an intermediate node (e.g., the output of the Kth capacitor), the output voltage becomes K times the input voltage (refer to 2Vin, 3Vin, 4Vin, and S*Vin). Since K can be any value between 1 and N, the circuit has a programmable multiplication factor between 1 and N.

[0066] In Figure 12 , when the switch is on (connecting the top node of the nth capacitor to the top node of the n+1th capacitor), the capacitors are in parallel. When the switch is down (connecting the top node of the nth capacitor to the bottom node of the n+1th capacitor), the capacitors are in series.

[0067] Figure 13 is a circuit diagram of circuit 300, which shows a second embodiment of the example described herein, where there is a voltage multiplier at the input. In circuit 300, there is a multiplier at the input. Circuit 300 includes N activation inputs 301 (K bits) [range: 0-V], a voltage multiplier 302, the multiplied activation inputs 303 [range: 0-s*V] (S: scaling factor), N weights 304 (from local storage or external input) (K bits), a voltage multiplier controller 305, a multiplier 310 (digital input, analog output), a multiplier output 320 - current or charge, a summing bit line 330, a current or charge-to-voltage converter 340 (e.g., a resistor, a transimpedance amplifier, or a capacitor), a summing voltage 341, an AD converter 350, and a digital output 360 (M bits).

[0068] Figure 14 is a circuit diagram of an embodiment of an input multiplier 400 using a level shifter. The input multiplier 400 is Figure 13An example implementation of the current or charge to voltage converter 340 as shown. When the input voltage 410 is zero, the output voltage 412 is zero. When the input voltage 410 is V, the output voltage 412 depends on the supply voltage of the circuit that is the output voltage of the multiplexer 402. For example, when the output of the multiplexer 402 is K*V, the output voltage 412 is also K*V. Thus, the output voltage 412 is the input voltage 410 multiplied by K. When K is between 1 and Smax, the input multiplier 400 can have a multiplication factor from 1 to Smax. The input multiplier 400 includes a multiplexer 402 that selects an input from among V 404, 2*V 406 up to Smax*V 408.

[0069] Figure 15 FIG. is a circuit diagram of the circuit 500-1 showing the first operating phase of the third embodiment of the example described herein, which uses a sum multiplier with capacitors connected in parallel to implement voltage sampling. In Figure 15 there is amplification of the multiply and accumulate result 541 using a sum multiplier. Figure 15 corresponds to Figure 12 the circuit state 285 (reference item 585).

[0070] Figure 16 FIG. is a circuit diagram of the circuit 500-2 showing the second operating phase of the third embodiment of the example described herein, which uses a sum multiplier with capacitors reconfigured to be connected in series to implement voltage multiplication. In Figure 16 there is amplification of the multiply and accumulate result 541 using a sum multiplier. Figure 16 corresponds to Figure 12 the circuit state 291 (reference item 591).

[0071] Refer to Figure 15 and Figure 16, given S input vectors, where t = 0…S-1, V(t) samples the sum of the t-th input vector (X(t)). The circuits (500-1, 500-2) include N activation inputs 501 (K bits), N weights 504 (from local storage or external input (K bits)), a sum bit line 530, a current or charge to voltage converter 540 (such as a resistor, transimpedance amplifier, or capacitor), a sum voltage 541, capacitors 592-1, 592-2, 592-S-1, and 592-S, and an M-bit ADC converter 550. Circuit 500-1 includes an amplified voltage 544 generated from the configuration 585 of the sum multiplier and generates a digital output 560 after analog-to-digital conversion using ADC 550. Circuit 500-2 includes an amplified voltage 545 generated from the configuration 591 of the sum multiplier and generates a digital output 561 after analog-to-digital conversion using ADC 550.

[0072] Figure 17 is a graph showing the NN accuracy on the evaluation set during the training of the BERT-base INT4 model. Figure 17 The results of the reference run (curve 804) are compared with the results without implementing the examples described herein (curve 806) and the results with the examples described herein (curve 802). The y-axis is F1. The x-axis is the training iteration. Curve 804 corresponds to the accuracy result without ADC truncation. Curve 802 corresponds to the accuracy result when auto-scaling is implemented and 7-bit ADC LSB truncation is used. Curve 806 corresponds to the accuracy result when auto-scaling is not implemented and 7-bit ADC LSB truncation is used. Curve 802 has a peak F1 value of 87.5%, curve 804 has a peak F1 value of 87.5%, and curve 806 has a peak F1 value of 81.8%. Thus, with the examples described herein, low-precision ADCs (with truncation) can be used to match the accuracy performance of high-precision (without truncation) ADCs.

[0073] Figure 18 is another graph showing the NN accuracy on the evaluation set during the training of the BERT-base INT4 model. The y-axis is F1, and the x-axis is the training iteration. Curve 902 shows a fixed F1 value of 87.7%, curve 904 corresponds to LSB = 8, MSB = 0, curve 906 corresponds to LSB = 6, MSB = 1, and curve 908 corresponds to LSB = 10, MSB = 0. When using LSB = 6, MSB = 1 (curve 906), the results are close to the 87.7% int4 baseline (curve 902).

[0074] Figure 19It is a graph showing the quantization-aware training convergence of the MobileNet-v1 (MB1) model. Curve 1002 corresponds to LSB=6 truncation without auto-search. Without auto-search, MB1 QAT cannot converge with LSB=6 truncation. The y-axis is the training error, and the x-axis is the training epoch. Figure 19 Shows the convergence region 1004.

[0075] Figure 20 It is a graph showing a comparison of the performance results of post-training quantization for the BERT base INT8 model with and without auto-search scaling. Curve 1202 corresponds to the implementation without auto-search scaling, and curve 1204 corresponds to the implementation with auto-search scaling. Curve 1201 corresponds to the baseline F1 value. Up to 32x magnification enables equal precision with more aggressive LSB truncation.

[0076] Figure 21 It is a logic flow chart of an example implementation method 1300 described herein. At 1310, the method includes determining corresponding integer scalar values of a layer of a neural network among multiple layers of the neural network, where multiple corresponding integer scalar values are determined for the multiple layers of the neural network. At 1320, the method includes determining the matrix multiplication output of the neural network. At 1330, the method includes increasing the corresponding integer scalar value by 1 when the matrix multiplication output does not exceed the analog-to-digital converter threshold. At 1340, the method includes decreasing the corresponding integer scalar value by 1 when the matrix multiplication output exceeds the analog-to-digital converter threshold. At 1350, the method includes determining a moving average of the corresponding integer scalar values determined for the multiple layers. At 1360, the method includes determining the final integer scalar as the moving average truncated to an integer, and the final integer scalar is used for magnification before modulus truncation during inference using the neural network.

[0077] Method 1300 may further include determining the integer scalar value of the layer of the neural network during training of the neural network, where the training includes quantization-aware training.

[0078] Method 1300 may further include determining the integer scalar value of the layer of the neural network during calibration of the neural network, where the calibration includes post-training quantization.

[0079] Method 1300 may further include where the layer of the neural network includes switched-capacitor operation.

[0080] Method 1300 may further include: in response to an overflow during training or calibration, reducing the scalar and re-doing the iteration of training the neural network or calibrating the neural network with or without updating at least one parameter of the neural network.

[0081] Method 1300 may further include: determining a first value associated with most significant bit truncation; determining a threshold based on a second value associated with a bit accumulator and the first value associated with most significant bit truncation; and determining an overflow when a third value associated with a matrix multiplication output exceeds the threshold. For example, after amplification, when the MACC output is 14 bits, the accumulator is 16, and the MSB truncation is 3 bits, the threshold is 16 - 3 = 13, and the 14-bit MACC output exceeds the 13-bit threshold. Thus, method 1300 may further include wherein the threshold is determined to be the second value associated with the bit accumulator minus the first value associated with most significant bit truncation, and wherein the first value is a first number of bits, the second value is a second number of bits, and the third value is a third number of bits.

[0082] Method 1300 may further include: determining whether to apply most significant bit truncation to a matrix multiplication output; in response to determining to apply the most significant bit truncation to the matrix multiplication output, reducing the corresponding integer scalar value; and in response to determining that the application of most significant bit truncation to the matrix multiplication output does not exceed a threshold number of times, increasing the corresponding integer scalar value.

[0083] Figure 22 is a logic flow diagram of implementing method 1400 based on the examples described herein. At 1410, the method includes receiving a first plurality of inputs (11, 201, 301, 501) representing an activation input vector. At 1420, the method includes receiving a second plurality of inputs (12, 204, 304, 504) representing a weight input vector. At 1430, the method includes generating an analog voltage representing a multiplication and accumulation result (13, 241, 244, 341, 541, 544) of the first plurality of inputs (11, 201, 301, 501) and the second plurality of inputs (12, 204, 304, 504) using an analog multiplier and accumulator (10, 240, 340, 540). At 1440, the method includes using a finite precision operation during an inference operation (47) of a neural network and converting the analog voltage multiplication and accumulation result (13, 241, 244, 341, 541, 544) to a digital signal (15, 260, 360, 560) using an analog-to-digital converter (14, 250, 350, 550). At 1450, the method includes determining (48, 49, 50, 70, 248, 261) a scaling factor (53, 63, 80) for amplifying (64, 75, 76, 302) the first plurality of inputs (11, 201, 301, 501) or amplifying (9, 77, 242, 280, 400, 585, 591) the analog voltage multiplication and accumulation result (13, 241, 244, 341, 541, 544) during training or calibration (46) of the neural network.

[0084] Referring now to all the drawings, the following embodiments are disclosed herein.

[0085] Example 1. An apparatus comprising: a first plurality of inputs representing an activation input vector; a second plurality of inputs representing a weight input vector; an analog multiplier and accumulator for generating a first analog voltage representing a first multiplication and accumulation result of the first input and the second input; a voltage multiplier for obtaining the first analog voltage and generating a second analog voltage representing a second multiplication and accumulation result by multiplying the first analog voltage by at least one scaling factor; an analog-to-digital converter configured to convert the second analog voltage multiplication and accumulation result into a digital signal using finite precision arithmetic during neural network inference operations; and a hardware controller or a software controller, the hardware controller being configured to determine at least one scaling factor based on the first multiplication and accumulation result, and the software controller being configured to determine at least one scaling factor based on the first multiplication and accumulation result.

[0086] Example 2. The apparatus according to Example 1, wherein the at least one scaling factor comprises a plurality of independent scaling factors determined during the training of the neural network, and each switched capacitor of a layer of a neural network comprising a plurality of layers operates with an independent scaling factor.

[0087] Example 3. The apparatus according to any one of Examples 1 to 2, wherein the apparatus determines the at least one scaling factor during the training of the neural network.

[0088] Example 4. The apparatus according to Example 3, wherein at least one scaling factor determined during training and used during inference is an integer value.

[0089] Example 5. The apparatus according to any one of Examples 1 to 4, further comprising: an accumulative storage charge configured to accumulate charges corresponding to the second analog voltage multiplication and accumulation results for a number of iterations.

[0090] Example 6. The apparatus according to any one of Examples 1 to 5, further comprising: a programmable controller configured to control the voltage multiplier based on the at least one scaling factor.

[0091] Example 7. The apparatus according to any one of Examples 1 to 6, wherein the voltage multiplier comprises a plurality of switched capacitors in series or parallel configuration.

[0092] Example 8. An apparatus includes: a first plurality of inputs representing an original activation input vector; a plurality of voltage multipliers that receive the first plurality of inputs and generate a second plurality of inputs by multiplying a voltage of the original activation input vector by at least one scaling factor; a third plurality of inputs representing a weight input vector; an analog multiplier and accumulator for generating an analog voltage representing a multiplication and accumulation result of the second input and the third input; an analog-to-digital converter configured to convert the analog voltage multiplication and accumulation result into a digital signal using finite precision arithmetic during a neural network inference operation; and a hardware controller or a software controller, the hardware controller being configured to determine the at least one scaling factor based on the multiplication and accumulation result, and the software controller being configured to determine the at least one scaling factor based on the multiplication and accumulation result.

[0093] Example 9. The apparatus according to Example 8, wherein the at least one scaling factor includes a plurality of independent scaling factors, and each switched capacitor of a layer of a neural network including a plurality of layers operates on an independent scaling factor.

[0094] Example 10. The apparatus according to Example 9, wherein the plurality of independent scaling factors are determined during training of the neural network.

[0095] Example 11. The apparatus according to any one of Examples 8 to 10, wherein the apparatus determines the at least one scaling factor during training of the neural network.

[0096] Example 12. The apparatus according to Example 11, wherein the at least one scaling factor determined during training and used during inference is an integer value.

[0097] Example 13. The apparatus according to any one of Examples 8 to 12 further includes: an accumulative storage charge configured to accumulate charges corresponding to the analog voltage multiplication and accumulation results for a number of iterations.

[0098] Example 14. The apparatus according to any one of Examples 8 to 13 further includes: at least one programmable controller configured to control the plurality of voltage multipliers based on at least one scaling factor.

[0099] Example 15. A method includes: receiving a first plurality of inputs representing an activation input vector; receiving a second plurality of inputs representing a weight input vector; generating, using an analog multiplier and accumulator, an analog voltage representing a multiplication and accumulation result of the first plurality of inputs and the second plurality of inputs; using an analog-to-digital converter to convert the analog voltage multiplication and accumulation result to a digital signal using finite precision arithmetic during an inference operation of a neural network; and determining, during training or calibration of the neural network, at least one scaling factor for amplifying the first plurality of inputs or the analog voltage multiplication and accumulation result.

[0100] Example 16. The method of Example 15 further includes: determining a plurality of independent scaling factors, including determining one independent scaling factor for each switched-capacitor operation of a layer of a neural network including a plurality of layers, wherein at least one scaling factor includes the plurality of independent scaling factors.

[0101] Example 17. The method of any one of Examples 15 to 16, wherein amplifying the first plurality of inputs includes: using a plurality of voltage multipliers to generate amplified first plurality of inputs by multiplying the at least one scaling factor by a voltage of the activation input vector, the method further including: using the analog multiplier and accumulator to generate the analog voltage multiplication and accumulation result for the amplified first plurality of inputs.

[0102] Example 18. The method of any one of Examples 15 to 17, wherein amplifying the analog voltage includes: using a voltage multiplier to generate an amplified analog voltage multiplication and accumulation result by applying the at least one scaling factor to the analog voltage multiplication and accumulation result, the method further including: using the analog-to-digital converter to convert the amplified analog voltage multiplication and accumulation result to the digital signal using the finite precision arithmetic during the inference operation of the neural network.

[0103] Example 19. The method of Example 18 further includes: configuring a plurality of switched capacitors of the voltage multiplier in series; or configuring a plurality of switched capacitors of the voltage multiplier in parallel.

[0104] Example 20. The method of any one of Examples 15 to 19 further includes: accumulating charge corresponding to the analog voltage multiplication and accumulation results for multiple iterations.

[0105] References to "computers", "processors", etc. should be understood to cover not only computers with different architectures such as single / multi-processor architectures and sequential or parallel architectures, but also dedicated circuits such as field-programmable gate array (FPGA) dedicated circuits (ASIC), signal processing devices, and other processing circuits. References to computer programs, instructions, code, etc. should be understood to cover software or firmware for programmable processors, such as, for example, the programmable content of a hardware device, whether instructions for a processor or configuration settings for a fixed-function device, gate array, or programmable logic device, etc.

[0106] The memory as described herein can be implemented using any suitable data storage technology such as semiconductor-based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, non-transitory memory, transitory memory, fixed memory, and removable memory. The (one or more) memories can include a database for storing data.

[0107] As used herein, a circuit can refer to the following: (a) a hardware circuit implementation such as an implementation in analog and / or digital circuits, and (b) a combination of a circuit and software (and / or firmware), such as, if applicable: (i) a combination of processors or (ii) a portion of a processor / software, including a digital signal processor, software, and memory that work together to cause a device to perform various functions, and (c) a circuit such as a microprocessor or a portion of a microprocessor that requires software or firmware to operate even if the software or firmware is not physically present. As another example, as used herein, a circuit will also cover an implementation of only a processor (or processors) or a portion of a processor and its (or their) accompanying software and / or firmware. For example and if applicable to a particular element, a circuit will also cover a baseband integrated circuit or an application processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, or another network device.

[0108] A list of abbreviations that can be appended to each other or to other characters using, for example, a dash or hyphen ("-"):

[0109] AD Analog-to-Digital

[0110] ADC Analog-to-Digital Converter

[0111] ASIC Application-Specific Integrated Circuit

[0112] b bit (e.g., 8b)

[0113] BERT Bidirectional Encoder Representations from Transformers

[0114] BMM Batch Matrix Multiplication

[0115] Capacitor

[0116] DNN (Deep Neural Network)

[0117] ep (Epoch)

[0118] F1 (Harmonic mean of precision and recall)

[0119] FPGA (Field Programmable Gate Array)

[0120] GEMM (Generalized Matrix Multiplication)

[0121] HW (Hardware)

[0122] INT (Integer)

[0123] LSB (Least Significant Bit)

[0124] MACC (Multiply-Accumulate)

[0125] MB1 (MobileNet-v1 Neural Network Model)

[0126] MSB (Most Significant Bit)

[0127] NN (Neural Network)

[0128] Prec. (Precision)

[0129] PTQ (Post-Training Quantization)

[0130] QAT (Quantization-Aware Training)

[0131] SC-PT (Switching Capacitor Processing Tile, a core hardware component)

[0132] swcap (Switching Capacitor)

[0133] SW (Software)

[0134] trunc (Truncate)

[0135] V (Voltage)

[0136] Vdd (Supply Voltage)

[0137] W (Weight Vector - Input to the Multiply-Accumulate (MACC) operation)

[0138] X (Input Vector - Input to the Multiply-Accumulate (MACC) operation)

[0139] In the foregoing description, numerous specific details such as particular structures, components, materials, dimensions, processing steps, and techniques have been set forth in order to provide a thorough understanding of the exemplary embodiments disclosed herein. However, one of ordinary skill in the art will understand that the exemplary embodiments disclosed herein may be practiced without these specific details. Additionally, details of well-known structures or processing steps may have been omitted or may not have been described in order to avoid obscuring the presented embodiments.

[0140] The description of the invention has been presented for purposes of illustration and description, but is not intended to be exhaustive or to limit the disclosed forms. Many modifications and variations will be apparent to one of ordinary skill in the art without departing from the scope of the invention. The embodiments were chosen and described in order to best explain the principles of the invention and the practical application, and to enable one of ordinary skill in the art to understand the invention with various modifications suited to the particular use contemplated.

Claims

1. A device, comprising: a first plurality of inputs representing activation input vectors; a second plurality of inputs representing weight input vectors; an analog multiplier and accumulator for generating a first analog voltage representing a first multiplication and accumulation result of the first input and the second input; a voltage multiplier that obtains the first analog voltage and generates a second analog voltage representing a second multiplication and accumulation result by multiplying at least one scaling factor with the first analog voltage; an analog-to-digital converter configured to convert the second analog voltage multiplication and accumulation result into a digital signal using finite-precision arithmetic during neural network inference operations; and a hardware controller or a software controller, the hardware controller being configured to determine the at least one scaling factor based on the first multiplication and accumulation result, the software controller being configured to determine the at least one scaling factor based on the first multiplication and accumulation result.

2. The device according to claim 1, wherein The at least one scaling factor includes a plurality of independent scaling factors determined during the training of the neural network, and each switched capacitor of a layer of a neural network including multiple layers operates on an independent scaling factor.

3. The device according to claim 1, wherein The device determines the at least one scaling factor during the training of the neural network.

4. The apparatus according to claim 3, wherein, The at least one scaling factor determined during training and used during inference is an integer value.

5. The device according to claim 1, further comprising: a cumulative storage charge configured to accumulate charges corresponding to the second analog voltage multiplication and accumulation results for a number of iterations.

6. The device according to claim 1, further comprising: a programmable controller configured to control the voltage multiplier based on the at least one scaling factor.

7. The apparatus according to claim 1, wherein The voltage multiplier includes a plurality of switched capacitors in series or parallel configuration.

8. A device, comprising: a first plurality of inputs representing original activation input vectors; a plurality of voltage multipliers that obtain the first plurality of inputs and generate a second plurality of inputs by multiplying at least one scaling factor with the voltage of the original activation input vector; a third plurality of inputs representing weight input vectors; an analog multiplier and accumulator for generating an analog voltage representing a multiplication and accumulation result of the second input and the third input; an analog-to-digital converter configured to convert the analog voltage multiplication and accumulation result into a digital signal using finite-precision arithmetic during neural network inference operations; and a hardware controller or a software controller, the hardware controller being configured to determine the at least one scaling factor based on the multiplication and accumulation result, the software controller being configured to determine the at least one scaling factor based on the multiplication and accumulation result.

9. The device according to claim 8, wherein, The at least one scaling factor includes a plurality of independent scaling factors, and each switched capacitor of a layer of a neural network including multiple layers operates on an independent scaling factor.

10. The device according to claim 9, wherein, A plurality of independent scaling factors are determined during the training of the neural network.

11. The device according to claim 8, wherein, The device determines the at least one scaling factor during the training of the neural network.

12. The apparatus according to claim 11, wherein, The at least one scaling factor determined during training and used during inference is an integer value.

13. The device according to claim 8, further comprising: A cumulative storage charge, configured to cumulatively store charges corresponding to the analog voltage multiplication and accumulation results for a number of iterations.

14. The apparatus according to claim 8, further comprising: At least one programmable controller, configured to control the plurality of voltage multipliers based on the at least one scaling factor.

15. A method, comprising: Receiving a first plurality of inputs representing an activation input vector; Receiving a second plurality of inputs representing a weight input vector; Generating, using an analog multiplier and accumulator, an analog voltage representing the multiplication and accumulation results of the first plurality of inputs and the second plurality of inputs; Using an analog-to-digital converter to convert the analog voltage multiplication and accumulation results into a digital signal using finite precision arithmetic during an inference operation of a neural network; And During training or calibration of the neural network, determining at least one scaling factor for amplifying the first plurality of inputs or the analog voltage multiplication and accumulation results.

16. The method according to claim 15, further comprising: Determining a plurality of independent scaling factors, including determining an independent scaling factor for each switched-capacitor operation of a layer of a neural network including a plurality of layers, wherein the at least one scaling factor includes the plurality of independent scaling factors.

17. The method according to claim 15, wherein, Amplifying the first plurality of inputs includes: using a plurality of voltage multipliers to generate amplified first plurality of inputs by multiplying the at least one scaling factor with the voltages of the activation input vector, the method further comprising: using the analog multiplier and accumulator to generate the analog voltage multiplication and accumulation results for the amplified first plurality of inputs.

18. The method according to claim 15, wherein, Amplifying the analog voltage includes: using a voltage multiplier to generate amplified analog voltage multiplication and accumulation results by applying the at least one scaling factor to the analog voltage multiplication and accumulation results, the method further comprising: using the analog-to-digital converter to convert the amplified analog voltage multiplication and accumulation results into the digital signal using the finite precision arithmetic during the inference operation of the neural network.

19. The method according to claim 18, further comprising: Configuring a plurality of switched capacitors of the voltage multiplier in series; Or Configuring a plurality of switched capacitors of the voltage multiplier in parallel.

20. The method according to claim 15, further comprising: Cumulatively storing charges corresponding to the analog voltage multiplication and accumulation results for multiple iterations.

21. A computer program product, comprising instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 15-20.