Method and apparatus for calibrating analog circuits for performing neural network calculations

By calibrating the analog circuit and configuring the normalization layer with the statistical data of the calibration output, the accuracy reduction problem caused by the circuit non-ideality of the analog circuit in neural network calculation is solved, achieving higher calculation accuracy and reducing the need for retraining.

CN114819051BActive Publication Date: 2025-05-06MEDIATEK SINGAPORE PTE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210062183.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-01-06
Filing Date
2022-01-19
Publication Date
2025-05-06
Estimated Expiration
2042-01-19

AI Technical Summary

Technical Problem

The analog circuits are less accurate in computing due to the non-ideality of the circuit when performing neural network calculations, and retraining the neural network suitable for each chip is expensive.

Method used

By providing calibration input to the pretrained neural network, the calibration output statistics of the analog circuit are calculated, and the normalization operation is determined and configured in the normalization layer, combined with the calibration output statistics to be written to the memory while keeping the pretrained weight unchanged.

Benefits of technology

Improves the accuracy of analog circuits when performing neural network calculations, avoids the reduction in calculation accuracy caused by the non-ideality of the circuit, and does not require retraining for each manufacturing chip.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114819051B_ABST
    Figure CN114819051B_ABST
Patent Text Reader

Abstract

The present invention discloses a calibration method for an analog circuit for performing neural network calculations, comprising: providing a calibration input to a pre-trained neural network, the neural network comprising at least one given layer, the given layer having pre-trained weights stored in the analog circuit; calculating statistics of the calibration output from the analog circuit, the analog circuit performing tensor operations of the given layer using the pre-trained weights; a normalization layer after the given layer determining a normalization operation to be performed during neural network reasoning, wherein the normalization operation is combined with statistics of the calibration output; and writing the configuration of the normalization operation to a memory while keeping the pre-trained weights unchanged. The present invention implements calibration of the analog circuit, thereby making the results of the analog circuit in performing neural network calculations more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of circuit technology, and in particular to a calibration method and device for an analog circuit for performing neural network calculations. Background Art

[0002] A deep neural network (DNN) is a neural network with an input layer, an output layer, and one or more hidden layers between the input and output layers. Each layer performs operations on one or more tensors. A tensor is a mathematical object that can be zero-dimensional (aka scaler), one-dimensional (aka vector), two-dimensional (aka matrix), or multi-dimensional. The operations performed by these layers are numerical computations, including but not limited to: convolution, deconvolution, fully-connected operation, normalization, activation, pooling, resizing, element-wise arithmetic, concatenation, slicing, etc. Some layers apply filter weights to tensors, such as in convolution operations.

[0003] Neural network computations are computationally intensive and often result in high power consumption. Therefore, neural network inference on edge devices needs to be fast and low power. Well-designed analog circuits can speed up inference and improve energy efficiency compared to digital circuits. However, analog computing is more susceptible to circuit non-ideality than digital computing, such as process variations. Circuit non-ideality reduces the accuracy of neural network computations. However, retraining a neural network suitable for each manufactured chip is both expensive and infeasible. Therefore, improving the accuracy of analog neural network computations is a challenge. Summary of the invention

[0004] In view of this, the present invention provides a calibration method and apparatus for an analog circuit for performing neural network calculations to solve the above-mentioned problem.

[0005] According to a first aspect of the present invention, a method for calibrating an analog circuit for performing neural network calculations is disclosed, comprising:

[0006] providing a calibration input to a pre-trained neural network, the neural network comprising at least a given layer, the given layer having pre-trained weights stored in an analog circuit;

[0007] computing statistics from a calibrated output of the analog circuit that performs tensor operations of the given layer using the pre-trained weights;

[0008] A normalization layer following the given layer determines a normalization operation to be performed during neural network inference, wherein the normalization operation incorporates statistics of the calibration output; and

[0009] The configuration of the normalization operation is written to the memory while keeping the pre-trained weights unchanged.

[0010] According to a second aspect of the present invention, a method for calibrating an analog circuit for performing neural network calculations is disclosed, comprising:

[0011] performing, by the analog circuit, tensor operations on the calibration inputs using pre-trained weights stored in the analog circuit to generate a calibrated output for a given layer of the neural network;

[0012] receiving a configuration of a normalization layer following the given layer, wherein the normalization layer is defined by a normalization operation that incorporates statistics of the calibration output; and

[0013] Perform neural network inference, including tensor operations for a given layer using the pretrained weights and normalization operations for normalized layers.

[0014] According to a third aspect of the present invention, there is disclosed an apparatus for performing neural network calculations, comprising:

[0015] Analog circuitry for storing pre-trained weights for at least a given layer of a neural network, wherein the analog circuitry is configured to: generate calibrated outputs from the given layer by performing tensor operations on calibration inputs using the pre-trained weights during calibration; and perform neural network inference using the pre-trained weights, including the tensor operations for the given layer; and

[0016] A digital circuit is configured to receive a configuration of a normalization layer following the given layer, wherein the normalization layer is defined by a normalization operation including statistics of the calibration output, and to perform the normalization operation of the normalization layer during inference of the neural network.

[0017] The method for calibrating an analog circuit to perform neural network calculations of the present invention includes: providing a calibration input to a pre-trained neural network, the neural network including at least one given layer, the given layer having pre-trained weights stored in the analog circuit; calculating statistics of the calibration output from the analog circuit, the analog circuit using the pre-trained weights to perform tensor operations of the given layer; a normalization layer after the given layer determines a normalization operation to be performed during neural network inference, wherein the normalization operation is combined with the statistics of the calibration output; and writing the configuration of the normalization operation to a memory while keeping the pre-trained weights unchanged. The present invention implements calibration of the analog circuit, thereby making the results of the analog circuit in performing neural network calculations more accurate. Therefore, the solution proposed by the present invention avoids reducing the accuracy of neural network calculations due to circuit non-ideality, and does not require retraining for each manufactured chip to suit neural network calculations. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a block diagram illustrating a system operable (or operable) to perform neural network computations according to one embodiment.

[0019] Figure 2 is a diagram illustrating the mapping between DNN layers and hardware circuits according to one embodiment.

[0020] Figure 3 is a block diagram illustrating an analog circuit according to one embodiment.

[0021] Figure 4 is a flow chart illustrating a calibration process according to one embodiment.

[0022] Figure 5 The operation performed by a normalization layer according to the first embodiment is illustrated.

[0023] Figure 6 The operation performed by the normalization layer according to the second embodiment is shown.

[0024] Figure 7 is a flow chart illustrating a method for calibrating analog circuits for neural network computations according to one embodiment.

[0025] Figure 8 is a flow chart illustrating a method of calibrating an analog circuit for neural network calculations according to another embodiment. DETAILED DESCRIPTION

[0026] In the following detailed description of the embodiments of the invention, reference is made to the accompanying drawings, which form a part of the present invention and in which are shown by way of illustration certain preferred embodiments in which the invention may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice them, and it is understood that other embodiments may be utilized and mechanical, structural and procedural changes may be made without departing from the spirit and scope of the invention. Therefore, the following detailed description should not be construed as limiting, and the scope of the embodiments of the invention is limited only by the appended claims.

[0027] It will be understood that although the terms "first," "second," "third," "primary," "secondary," etc. may be used herein to describe various elements, components, regions, layers, and / or portions, these elements, components, regions, layers, and / or portions should not be limited by these terms. These terms are only used to distinguish one element, component, region, layer, or portion from another region, layer, or portion. Therefore, without departing from the teachings of the present inventive concept, the first or primary element, component, region, layer, or portion discussed below may be referred to as a second or secondary element, component, region, layer, or portion.

[0028] In addition, for ease of description, spatially relative terms such as "below," "under," "under," "above," and "over" may be used herein to describe the relationship of an element or feature to another element or feature as shown. In addition to the orientations described in the figures, spatially relative terms are intended to cover different orientations of the device in use or operation. The device can be oriented in other ways (rotated 90 degrees or in other orientations), and the spatially relative descriptors used herein can be interpreted accordingly. In addition, it will be understood that when a "layer" is referred to as being "between" two layers, it can be the only layer between the two layers, or one or more intermediate layers can also be present.

[0029] The terms "approximately", "roughly" and "about" generally mean within the range of ±20% of a specified value, or ±10% of the specified value, or ±5% of the specified value, or ±3% of the specified value, or ±2% of the specified value, or ±1% of the specified value, or ±0.5% of the specified value. The specified values ​​of the present invention are approximate values. When not specifically described, the specified values ​​include the meanings of "approximately", "roughly" and "about". The terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention. As used herein, the singular terms "one", "an" and "the" are also intended to include plural forms, unless the context clearly indicates otherwise. The terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the inventive concept. As used herein, the singular forms "one", "a kind of" and "the" are also intended to include plural forms, unless the context clearly indicates otherwise.

[0030] It will be understood that when an “element” or “layer” is referred to as being “on,” “connected to,” “coupled to,” or “adjacent to” another element or layer, it can be directly on, connected to, coupled to, or adjacent to the other element or layer, or intervening elements or layers may be present. In contrast, when an element is referred to as being “directly on,” “directly connected to,” “directly coupled to,” or “immediately adjacent to” another element or layer, there are no intervening elements or layers.

[0031] Note that: (i) like features will be indicated by like reference numerals throughout the figures and will not necessarily be described in detail in each figure in which they appear, and (ii) a series of figures may show different aspects of a single item, each aspect being associated with various reference labels which may appear throughout the sequence or may only appear in selected figures of the sequence.

[0032] Embodiments of the present invention provide an apparatus and method for calibrating analog circuits to improve the accuracy of analog neural network calculations. The apparatus may include analog circuits and digital circuits for performing neural network calculations according to a deep neural network (DNN) model. The DNN model includes a first set of layers ("A layers") mapped to analog circuits and a second set of layers ("D layers") mapped to digital circuits. Each layer is defined by a corresponding operation (or operation). For example, a convolution layer is defined by corresponding filter weights and parameters for performing convolution. The DNN model is pre-trained before being loaded into the apparatus. However, analog circuits made on different chips may have different non-ideal characteristics. Therefore, the same set of pre-trained filter weights and parameters may cause different analog circuits to produce different outputs. The calibration described herein eliminates or reduces the differences between different chips.

[0033] Calibration is performed offline after the DNN is trained on the output of each A layer. During calibration, the calibration input is fed to the DNN and statistics are collected for the calibration output of each A layer. The calibration input can be a subset of the training data used for DNN training. Calibration is different from retraining because the parameters and weights learned in training remain unchanged during and after calibration.

[0034] In some embodiments, statistics of the calibrated output of each A layer are used to modify or replace some operations defined in the DNN model. The statistics can be used to modify a batch normalization (BN) layer that is located after the A layer in the DNN model. Alternatively, the statistics can be used to define a set of multiply-and-add operations applicable to the output of the A layer. In the description below, the term "normalization layer" refers to a layer that is immediately after the A layer and applies a normalization operation to the output of the A layer. The normalization operation is determined based on the statistics of the calibrated output of the A layer. After calibrating and configuring the normalization layer, the device performs reasoning based on the calibrated DNN model containing the normalization layer.

[0035] In one embodiment, the tensor operation performed by the A layer and the D layer can be a convolution operation. The convolutions performed by the A layer and the D layer can be the same type of convolutions or different types of convolutions. For example, the A layer can perform ordinary convolutions and the D layer can perform depth convolutions, or vice versa. The channel size is the same as the depth size. When performing depth convolution, the channel size of the input is the same as the channel size of the output. Assume that a convolution layer receives an input tensor of M channels and produces an output tensor of N channels, where M and N can be the same number or different numbers. In "ordinary convolution" using N filters, each filter is convolved with the M channels of the input tensor to produce M outputs. The M outputs are added to generate one of the N channels of the output tensor. In "depth convolution", M=N, and there is a one-to-one correspondence between the M filters used in the convolution and the M channels of the input tensor, where each filter is convolved with one channel of the input tensor to obtain one channel of the output tensor. Ordinary convolution can also be called normal convolution. In previous solutions, there is no calibration method. The inventor of the present invention creatively proposes a calibration solution to solve the problem of the prior art, rather than ignoring the problem. The present invention implements calibration of the analog circuit, so that the analog circuit performs neural network calculations more accurately. Therefore, the solution proposed by the present invention avoids reducing the accuracy of neural network calculations due to circuit non-ideality, and does not require retraining for each manufactured chip to adapt to neural network calculations.

[0036] Figure 11 is a block diagram illustrating a device 100 operable (operable) to perform neural network computations according to one embodiment. The device 100 includes one or more general and / or special purpose digital circuits 110, such as a central processing unit (CPU), a graphics processing unit (GPU), a digital processing unit (DSP), a field-programmable gate array (FPGA), a neural processing unit (NPU), an arithmetic and logic unit (ALU), an application-specific integrated circuit (ASIC), and other digital circuits. The device 100 also includes one or more analog circuits 120 that perform mathematical operations; the mathematical operations are, for example, tensor operations (or operations). In one embodiment, the analog circuit 120 can be an analog compute-in-memory (ACIM) device, which includes a cell array with storage and embedded computing capabilities. For example, the cell array of the ACIM device can store the filter weights of a convolutional layer. When input data arrives at the cell array, the cell array performs convolution by generating an output voltage level corresponding to the convolution of the filter weights and the input data.

[0037] In one embodiment, digital circuit 110 is coupled to memory 130, which may include memory devices such as dynamic random-access memory (DRAM), static random access memory (SRAM), flash memory, and other non-transitory machine-readable storage media; for example, volatile or non-volatile storage devices. For simplicity of illustration, memory 130 is represented as a block; however, it should be understood that memory 130 may represent a hierarchy of memory components, such as cache memory, system memory, solid-state or magnetic storage devices, etc. Digital circuit 110 executes instructions stored in memory 130 to perform operations (or operations) such as tensor operations (or operations) and normalization on one or more neural network layers.

[0038] In one embodiment, the apparatus (or device) 100 further includes a controller 140 for scheduling and assigning operations defined in the DNN model to the digital circuit 110 and the analog circuit 120. In one embodiment, the controller 140 may be part of the digital circuit 110. In one embodiment, the apparatus 100 further includes a calibration circuit 150 for performing calibration of the analog circuit 120. The calibration circuit 150 is shown in dashed outline to show that it may be located in an alternative location. The calibration circuit 150 may be on the same chip as the analog circuit 120; alternatively, the calibration circuit 150 may be on a different chip than the analog circuit 120, but in the same apparatus 100. In yet another embodiment, the calibration circuit 150 may be in another system or device, such as a computer or server.

[0039] The device 100 may also include a network interface (network interface) 160 for communicating with another system or device via a wired and / or wireless network. It is understood that for simplicity of description, the device 100 may include Figure 1 In one embodiment, digital circuit 110 may execute instructions stored in memory 130 to perform operations of controller 140 and / or calibration circuit 150 .

[0040] Figure 2 2 is a diagram illustrating the mapping between a DNN model 200 and a hardware circuit according to one embodiment. The term "mapping" refers to the assignment of tensor operations (or operations) defined in the DNN model to hardware circuits that perform the operations. In this example, the DNN model includes a plurality of convolutional layers (e.g., CONV1-CONV5). Also refer to Figure 1 , the operations of CONV1, CONV2, and CONV3 ("A layers") may be assigned to analog circuitry 120, while the operations of CONV4 and CONV5 ("D layers") may be assigned to digital circuitry 110. The assignment of convolutional layers to analog circuitry 120 or digital circuitry 110 may be guided by criteria such as computational complexity, power consumption, accuracy requirements, etc. The filter weights of CONV1, CONV2, and CONV3 are stored in analog circuitry 120, and the filter weights of CONV3 and CONV4 are stored in a memory device accessible by digital circuitry 110 (e.g., Figure 1 The DNN model 200 may include additional layers (e.g., pooling, ReLU (Rectified Linear Unit, linear rectifier function), etc.). To simplify the description, Figure 2 These contents are omitted.

[0041] Figure 2The DNN model 200 in is a calibrated DNN; that is, it includes normalization layers (N1, N2, and N3) resulting from calibration. Each normalization layer is placed at the output of a corresponding A layer. In a first embodiment, the normalization layer can be a modified BN layer modified by statistics of the calibration output of the previous A layer. In a second embodiment, the normalization layer can apply a deep convolution to the output of the previous A layer, where the filter weights are at least partially obtained from the statistics of the calibration output of the previous A layer. The filter weights associated with CONV1-CONV5 learned from training are stored in the device 100 (e.g., the analog circuit 120 and the memory 130), and they do not change during and after calibration.

[0042] Figure 3 1 is a block diagram illustrating analog circuit 120 according to one embodiment. Analog circuit 120 may be an ACIM device including a cell array for data storage and memory calculation. There are various designs and implementations of ACIM devices; it is understood that analog circuit 120 is not limited to a particular type of ACIM device. In this example, the cell array of analog circuit 120 includes a plurality of cell array portions (e.g., cell array portions 310, 320, and 330) that store filter weights of convolutional layers CONV1, CONV2, and CONV3, respectively. For example, cell array portion 310 stores filter weights of convolutional layer CONV1, cell array portion 320 stores filter weights of convolutional layer CONV2, and cell array portion 320 stores filter weights of convolutional layer CONV2. Analog circuit 120 is coupled to input circuit 350 and output circuit 360, which buffer input data and output data of convolution operations, respectively. Input circuit 350 and output circuit 360 may also include conversion circuits for converting between analog and digital data formats.

[0043] Figure 4 4 is a flow chart illustrating a calibration process 400 according to one embodiment. The calibration process 400 begins with a training step 410, when a DNN (e.g., Figure 2 The DNN model 200 in the example is trained using a set of training data by a digital circuit; the digital circuit is, for example, a CPU in a computer. The training generates the filter weights (or filter weights) of the convolution and the parameters of batch normalization (or standardization) (e.g., β and γ). The value ε is used to avoid zero division. The training methods of convolution and batch normalization are known in the field of neural network computing. In step 420, the filter weights and parameters are loaded into a device (e.g., Figure 1The device 100 in FIG. 1 ). The first set of filter weights is stored in a memory accessible to (or accessible by) the digital circuit, while the second set of filter weights is stored in the analog circuit. Steps 430-450 are calibration steps. In step 430, a calibration input is provided to the DNN, at which point the DNN is trained and uncalibrated. In one embodiment, the calibration input may be a subset of the training data used in step 410. In step 440, the calibration output of each A layer is collected, and statistics of the calibration output are collected and calculated. In one embodiment, the statistics may include the mean and / or standard deviation of the calibration output. Statistics (e.g., mean and / or standard deviation) may be calculated for each calibration output startup, including all dimensions (i.e., height, width, and depth). Alternatively, for each calibration output startup across the height and width dimensions, statistics may be calculated by depth (i.e., each channel). The actions performed in step 440 may also be described as collecting calibration outputs for each layer ("A layer") mapped to the analog circuit, and calculating statistics for the calibration outputs.

[0044] The calculation of the statistics may be performed by an on-chip processor or circuit; alternatively, the calculation may be performed by off-chip hardware or other devices such as computers or servers. In step 450 for each A layer, the statistics are incorporated into a normalization operation or operation that defines a normalization layer after the A layer in the DNN. In step 450, for each A layer, a normalization operation for a corresponding normalization layer after the A layer is determined, wherein the normalization operation contains statistical information for calibrating the output. Reference Figure 5 and 6 A non-limiting example of a normalization operation is provided. The DNN including the normalization layer determined in step 450 is referred to as a calibrated DNN (calibrated DNN). In step 460, the calibrated DNN is stored in the device, wherein the calibrated DNN includes a corresponding normalization layer for each A layer (the calibrated DNN is stored, and the calibrated DNN includes a corresponding normalization layer for each A layer in the device (or apparatus)). In the reasoning step 470 (step 470 is a reasoning step), the device performs neural network reasoning according to the calibrated DNN (performing DNN reasoning according to the calibrated DNN). The filter weights obtained from the training of step 410 remain unchanged and are used for neural network reasoning.

[0045] Figure 5 FIG. 5 illustrates a normalization layer 500 according to a first embodiment. Figure 2In the example of , the normalization layer 500 may be any one of N1, N2, and N3. The normalization layer 500 may be a modified BN layer. In the trained DNN, the unmodified BN layer immediately follows the A layer 510 (e.g., any one of CONV1, CONV2, and CONV3). During training, the parameters of the unmodified BN layer (e.g., β, γ, and ε) are learned. When the trained DNN is loaded into the apparatus 100 ( Figure 1 ) after which the calibration process 400 is performed ( Figure 4 ) to calibrate the layers mapped to the analog circuit 120, the layers of the analog circuit 120 including the A layer 510.

[0046] The normalization layer 500 is defined by a normalization operation applied to the output tensor (represented by the solid cube 550) output from the A layer 510. During calibration, this tensor is referred to as the calibration output or calibration output start. The tensor has a height dimension (H), a width dimension (W), and a depth dimension (C), which is also called the channel dimension. The normalization operation converts each x i (represented by the elongated cube in dashed outline) into The tensor output by layer A 510 (represented by a solid cube 550) is normalized or computed before being sent to the next layer 520. i and All extend across the entire depth dimension or depth dimension C. Figure 5 In the example of , the normalization layer 500 incorporates the mean μ and the standard deviation (standard deviation) σ into the normalization operation (or operation) (that is, the normalization operation uses parameters including the mean μ and the standard deviation σ). In another embodiment, the normalization layer 500 can incorporate one of μ and σ into the normalization operation (or operation). The mean μ and the standard deviation σ are calculated from the calibration output of the A layer 510, which includes data points on all dimensions (H, W, and C). In addition, the normalization layer 500 also combines the parameters of the unmodified BN layer learned during training (e.g., β and γ). Therefore, the normalization layer 500 is also called a modified BN layer, which is modified to at least include (or combine, merge) the mean values ​​calculated on all dimensions of the calibration output. That is, the statistical data is calculated to include the mean values ​​of all dimensions of the calibration output.

[0047] Figure 6 The operation performed by the normalization layer 600 according to the second embodiment is illustrated. Figure 2In the example in FIG. 1 , the normalization layer 600 can be any one of N1, N2, and N3. The normalization layer 600 can replace the BN layer located after the A layer 610 (e.g., any one of CONV1, CONV2, and CONV3) in the uncalibrated DNN. During training, the depth parameters (e.g., β) of each channel across the depth dimension (depth-wise) are learned. k , γ k and ε), where the run index k identifies a specific channel. The trained DNN is loaded into the device 100 ( Figure 1 ) after which the calibration process 400 is performed ( Figure 4 ) to calibrate the layers mapped to the analog circuit 120, the layers of the analog circuit 120 including the A layer 610.

[0048] The normalization layer 600 is defined by a normalization operation applied to the tensor output from the A layer 610 (represented by each cube 650 in the solid line). During calibration, this tensor is referred to as the calibration output or calibration output start. The tensor has a height dimension (H), a width dimension (W), and a depth dimension (C), which is also called the channel dimension. The normalization operation (or operation) converts each F k,i,j (represented by a slice of the elongated cube with dashed outline) into where the run index k identifies a specific channel. k,i,j and are per-channel tensors. The tensors from layer A 610 (represented by each cube 650 in the solid line) are normalized or computed before being sent to the next layer 620. Figure 6 In the example of , the normalization layer 600 incorporates the per-channel mean value and the per-channel standard deviation into the normalization operation (or operation). In another embodiment, the normalization layer 600 can incorporate (or include, combine) one of the per-channel mean and the per-channel standard deviation into the normalization operation. The per-channel mean and the per-channel standard deviation are calculated based on the calibration output of the A layer 610 in the H and W dimensions of each channel in the C dimension. In addition, the normalization layer 500 also incorporates (or includes) the depth parameters (e.g., β) learned during training k , γ k and ε). Figure 6As shown, the normalization operation includes a depth-wise multiply-and-add operation (or operation), which at least includes (or combines) the depth (i.e., per-channel) average value calculated from each channel of the calibration output. That is, statistics are calculated to include the depth average value of the calibration output for each of the multiple channels in the depth dimension. Since the multiplication matrix shown in the normalization layer 600 is a diagonal matrix, the depth-wise (or depth-dimensional) multiply-and-add operation in this example is also called a 1x1 depth-wise (or depth-dimensional) convolution operation.

[0049] Figure 7 is a flow chart illustrating a method 700 for calibrating an analog circuit to perform neural network calculations according to one embodiment. The method 700 may be performed by a calibration circuit (e.g., Figure 1 The calibration circuit 150 is performed by the calibration circuit, which can be on the same chip as the analog circuit, on a different chip, or in a different device where the analog circuit is located.

[0050] Method 700 begins at step 710, when the calibration circuit sends a calibration input to a pre-trained neural network that includes at least a given layer having pre-trained weights stored in an analog circuit (providing a calibration input to a pre-trained neural network that includes at least a given layer having pre-trained weight values ​​stored in an analog circuit (pre-trained weights)). At step 720, the calibration circuit calculates statistics of the calibration output from the analog circuit, the analog circuit performing tensor operations (or operations) of the given layer on the calibration input using the pre-trained weights (pre-trained weights) (calculating statistics of the calibration output from the analog circuit, the analog circuit performing tensor operations of the given layer on the calibration input using the pre-trained weights). At step 730, the calibration circuit determines a normalization operation (or operation) to be performed at a normalization layer after the given layer during neural network inference. The normalization operation (or operation) includes (or is combined with) statistics of the calibration output (determining a normalization operation to be performed during neural network inference at a normalization layer after the given layer, wherein the normalization operation includes statistics of the calibration output). At step 740, the calibration circuit writes the configuration of the normalization operation into the memory. The pre-trained weights (pre-trained weights) remain unchanged after calibration (writing the configuration of the normalization operation into the memory while keeping the pre-trained weights unchanged).

[0051] Figure 8 8 is a flow chart illustrating a method 800 for calibrating an analog circuit for neural network computing according to one embodiment. The method 800 may be performed by a device including an analog circuit for neural network computing; for example, Figure 1 device 100.

[0052] Method 800 begins at step 810, at which time the analog circuit performs a tensor operation (operation) on the calibration input using pre-trained weights stored in the analog circuit and the calibration input. By performing the tensor operation, the analog circuit generates a calibration output for a given layer of the neural network (the analog circuit performs a tensor operation on the calibration input using the pre-trained weights stored in the analog circuit to generate a calibration output for the given layer of the neural network). At step 820, the device receives a configuration of a normalization layer after the given layer. The normalization layer is defined by a normalization operation (operation) that includes (or combines) statistical information of the calibration output (receiving a configuration of a normalization layer after the given layer, wherein the normalization layer is defined by a normalization operation that includes statistical information of the calibration output). At step 830, the device performs neural network inference using the pre-trained weights and the normalization operation of the normalization layer, and the neural network inference includes tensor operations (or operations) for the given layer (performing neural network inference using pre-trained weights and the normalization operation of the normalization layer, and the neural network inference includes tensor operations for the given layer).

[0053] In one embodiment, during neural network inference, analog circuits are assigned to perform tensor operations (or operations) of a given layer using pre-trained weights, and digital circuits in the device are assigned to perform normalization operations of a normalization layer. In other words, tensor operations of a given layer are assigned to analog circuits for execution; and normalization operations of a normalization layer are assigned to digital circuits for execution during neural network inference.

[0054] Various functional components or blocks have been described herein. As will be appreciated by those skilled in the art, the functional blocks will preferably be implemented by circuits (special purpose or general purpose circuits, which operate under the control of one or more processors and coded instructions), which typically include transistors that are configured to control the operation of the circuits according to the functions and operations described herein.

[0055] Figure 4 , 7 The operation of the flowchart of 8 has been referred to Figure 1 However, it should be understood that Figure 4 , 7 and 8. The operation of the flowchart can be performed by Figure 1 to perform embodiments of the present invention other than the embodiments of Figure 1 Embodiments of the invention may perform operations different from those discussed with reference to the flowcharts. Figure 4 , 7The flowcharts of and 8 show a specific order of operations performed by certain embodiments of the present invention, but it should be understood that such order is exemplary (for example, alternative embodiments may perform operations in a different order, combine certain operations, overlap certain operations, etc.).

[0056] Those skilled in the art will readily observe that many modifications and variations of the apparatus and method can be made while maintaining the teachings of the present invention.Accordingly, the above disclosure should be interpreted as being limited only by the metes and bounds of the appended claims.

Claims

1. A method for calibrating an analog circuit for performing neural network calculations, characterized in that include: providing a calibration input to a pre-trained neural network, the neural network comprising at least a given layer, the given layer having pre-trained weights stored in an analog circuit; wherein the neural network model includes a first set of layers mapped to analog circuits and a second set of layers mapped to digital circuits; During calibration, the calibration input is fed to the neural network and statistics of the calibration output of each of the first set of layers are collected; The calibration input is a subset of the training data used for training the neural network; computing statistics from a calibrated output of the analog circuit that performs tensor operations of the given layer using the pre-trained weights; A normalization layer after the given layer determines a normalization operation to be performed during neural network inference, wherein the normalization operation incorporates statistics of the calibration output; as well as The configuration of the normalization operation is written to the memory while keeping the pre-trained weights unchanged.

2. The calibration method according to claim 1, characterized in that: The analog circuit is an analog memory computing device.

3. The calibration method according to claim 1, characterized in that: Calculating statistics of the calibrated output from the analog circuit further includes: The statistical data is calculated to include at least one of a standard deviation and a mean of the calibration output.

4. The calibration method according to claim 1, characterized in that: The calibrated output has a height dimension, a width dimension, and a depth dimension, and calculating statistics of the calibrated output from the analog circuit further comprises: The statistics are calculated to include the means of all dimensions of the calibration output.

5. The calibration method according to claim 4, characterized in that: The normalization layer at least includes batch normalization of the mean value.

6. The calibration method according to claim 1, characterized in that: The calibrated output has a height dimension, a width dimension, and a depth dimension, and calculating statistics of the calibrated output from the analog circuit further includes: The statistics are calculated to include a depth average of the calibrated outputs for each of a plurality of channels in the depth dimension.

7. The calibration method according to claim 6, characterized in that: The normalization operation includes a multiplication-addition operation in the depth direction, which combines at least the average value in the depth direction of each channel.

8. The calibration method according to claim 1, characterized in that: The calibration of the analog circuit is performed on the same chip as the analog circuit; or, the calibration of the analog circuit is performed on a different chip or a different device than the analog circuit.

9. A method for calibrating an analog circuit for performing neural network calculations, characterized in that: include: performing, by the analog circuit, tensor operations on the calibration inputs using pre-trained weights stored in the analog circuit to generate a calibrated output for a given layer of the neural network; wherein the neural network model includes a first set of layers mapped to analog circuits and a second set of layers mapped to digital circuits; During calibration, the calibration input is fed to the neural network and statistics of the calibration output of each of the first set of layers are collected; The calibration input is a subset of the training data used for training the neural network; receiving a configuration of a normalization layer subsequent to the given layer, wherein the normalization layer is defined by a normalization operation that incorporates statistics of the calibration output; as well as Perform neural network inference, including tensor operations for a given layer using the pretrained weights and normalization operations for normalized layers.

10. The calibration method according to claim 9, characterized in that: Also includes: Assigning the tensor operation of the given layer to the simulation circuit for execution; as well as The normalization operation of the normalization layer is assigned to digital circuits to perform during neural network inference.

11. A device for performing neural network calculations, characterized in that include: Analog circuitry for storing pre-trained weights for at least a given layer of a neural network, wherein the analog circuitry is configured to: generate a calibrated output from the given layer by performing tensor operations on calibration inputs using the pre-trained weights during calibration; and perform neural network inference using the pre-trained weights, including the tensor operations for the given layer; as well as digital circuitry for receiving a configuration of a normalization layer following the given layer, wherein the normalization layer is defined by a normalization operation including statistics of the calibration output, and performing the normalization operation of the normalization layer during inference of the neural network; wherein the neural network model includes a first set of layers mapped to the analog circuit and a second set of layers mapped to the digital circuit; During calibration, the calibration input is fed to the neural network and statistics of the calibration output of each of the first set of layers are collected; The calibration input is a subset of the training data used for training the neural network.

Citation Information

Patent Citations

  • Large-Scale Artificial Neural-Network Accelerators Based on Coherent Detection and Optical Data Fan-Out

    US20210357737A1

  • Calibration of analog circuits for neural network computing

    US20220230064A1