Neural network device performing floating-point arithmetic and its operating method

The neural network device addresses the challenge of performing floating-point operations by converting inputs and weights into block floating-point format, allowing for efficient floating-point operations with reduced power consumption and maintained accuracy.

JP7683915B2Active Publication Date: 2025-05-27SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021094974
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-06-12
Filing Date
2021-06-07
Publication Date
2025-05-27
Estimated Expiration
2041-06-07

AI Technical Summary

Technical Problem

Existing neural network devices struggle to perform floating-point operations efficiently, particularly when using in-memory computing circuits that are limited to fixed-point operations, leading to accuracy loss and increased power consumption during quantization and relearning processes.

Method used

A neural network device and method that convert weights and input activations into block floating-point format, allowing for floating-point operations by separating the exponent part for digital arithmetic and using an analog crossbar array for fractional part operations, thereby minimizing power consumption and accuracy loss.

Benefits of technology

Enables efficient floating-point operations in neural network devices while reducing power consumption and maintaining accuracy, by leveraging block floating-point formats and in-memory computing circuits effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007683915000008
    Figure 0007683915000008
  • Figure 0007683915000009
    Figure 0007683915000009
  • Figure 0007683915000010
    Figure 0007683915000010
Patent Text Reader

Abstract

To provide a neural network apparatus performing floating-point operations and an operating method of the same.SOLUTION: A neural network apparatus performs MAC operations with respect to fractions of weights and input activations in a block floating-point format by using an analog crossbar array, performs addition operations with respect to shared exponents of weights and input activations in a block floating-point format by using a digital computing circuit, and outputs a partial sum of floating-point output activations by combining the result of the MAC operations and the result of the addition operations.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a neural network device that performs floating-point arithmetic and a method of operating the same. [Background technology]

[0002] There is increasing interest in neuromorphic processors that perform neural network operations. For example, research is being conducted to realize neuromorphic processors that include neuron circuits and synapse circuits. Such neuromorphic processors are also used in neural network devices for driving various neural networks such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and feedforward neural networks (FNNs), and are also used in fields including data classification and image recognition. Summary of the Invention [Problem to be solved by the invention]

[0003] The problem to be solved by the present invention is to provide a neural network device that performs floating-point arithmetic and an operating method thereof. The technical problem to be solved by the present invention is not limited to the above-mentioned technical problem, and other technical problems can be inferred from the following embodiments. [Means for solving the problem]

[0004] According to one aspect, a method for operating a neural network device performing floating-point arithmetic includes the steps of: determining, for each weight kernel, a first shared exponent representing a weight included in the weight kernel, and obtaining a first block floating point format weight including a first fraction adjusted based on the first shared exponent; determining, for each of a plurality of input stripes included in an input feature map, a second shared exponent representing input activations included in the input stripe, and obtaining a second block floating point format input activation including a second fraction adjusted based on the second shared exponent; performing a multiply-accumulate (MAC) operation on the first fraction and the second fraction using an analog crossbar array, and converting the first fraction and the second fraction into an analog to digital (ADC) output. converting a result of the MAC operation into a digital signal using a digital arithmetic circuit; and performing an addition operation involving the first shared exponent and the second shared exponent using a digital arithmetic circuit, combining the result of the MAC operation and the result of the addition operation, and outputting a partial sum of floating-point output activations included in a channel of an output feature map.

[0005] According to another aspect, a neural network apparatus for performing floating-point arithmetic may include at least one control circuit for determining, for each weight kernel, a first shared exponent representative of a weight included in the weight kernel, obtaining weights in a first block floating-point format including a first fractional part scaled based on the first shared exponent, and for each of a plurality of input stripes included in an input feature map, determining a second shared exponent representative of input activations included in the input stripe, obtaining input activations in a second block floating-point format including a second fractional part scaled based on the second shared exponent; an analog crossbar array for performing a MAC operation on the first fractional part and the second fractional part; an in-memory computing circuit including an ADC for converting a result of the MAC operation into a digital signal; and a digital arithmetic circuit for performing an addition operation on the first shared exponent and the second shared exponent, combining a result of the MAC operation and a result of the addition operation, and outputting a partial sum of floating-point output activations included in a channel of an output feature map. [Brief description of the drawings]

[0006] [Figure 1] 1 is a diagram illustrating the architecture of a neural network according to some embodiments. [Diagram 2] 1 is a diagram illustrating operations performed in a neural network according to some embodiments. [Diagram 3] 1 illustrates an in-memory computing circuit according to some embodiments. [Figure 4] FIG. 2 is a schematic diagram illustrating the overall process of a neural network device performing floating-point arithmetic according to some embodiments. [Diagram 5] 11 is a diagram illustrating a process in which a neural network device converts weights into block floating point numbers according to some embodiments. [Figure 6A]11 is a diagram illustrating a manner in which weights converted to block floating point are stored according to some embodiments. [Figure 6B] 13 is a diagram illustrating a method of storing weights converted to block floating point according to another embodiment; [Figure 7] 1 is a diagram illustrating a process in which a neural network device converts input activations into block floating point numbers according to some embodiments. [Figure 8] 11 is a diagram illustrating a process of performing floating-point arithmetic according to some embodiments in which an analog crossbar array supports signed inputs and weights. [Figure 9A] 11 is a diagram illustrating a process for performing floating-point arithmetic according to some embodiments in which an analog crossbar array supports signed weights but unsigned inputs. [Figure 9B] 11 is a diagram illustrating a process for performing floating-point arithmetic according to some embodiments in which an analog crossbar array supports signed weights but unsigned inputs. [Figure 10] 11 is a diagram illustrating a process of performing floating-point arithmetic according to some embodiments in which an analog crossbar array supports unsigned inputs and weights. [Figure 11] 11 is a diagram illustrating a process in which a final output of a floating-point operation is output by combining an operation result of an analog crossbar array with an operation result of a digital operation circuit according to some embodiments. [Figure 12] 4 is a flow chart illustrating a method of operation of a neural network device according to some embodiments. [Figure 13] FIG. 1 is a block diagram illustrating a configuration of an electronic system according to some embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0007] The terms used in this embodiment are currently common terms that are used as much as possible while taking into consideration the functions in this embodiment, but they may vary depending on the intentions of engineers in this technical field, precedents, or the emergence of new technologies. In addition, in certain cases, some terms are arbitrarily selected, and in such cases, their meanings are described in detail in the description of the embodiment. Therefore, the terms used in this embodiment should be defined based on the meanings of the terms and the overall content of this embodiment, rather than simply the names of the terms.

[0008] In the description of the present embodiment, when a part is connected to another part, this does not only mean that the part is directly connected to the other part, but also means that the part is electrically connected to the other part through another component in between. In addition, the terms "comprise" and "include" used in the present embodiment should not be interpreted as including all of the various components or steps described in the specification, but should be interpreted as including some of the components or steps not included, or including additional components or steps.

[0009] In addition, terms including ordinal numbers such as "first" or "second" are used in the present specification to describe various components, but the components are not limited to the terms. The terms may be used to distinguish one component from another.

[0010] The following description of the embodiments is not intended to limit the scope of the invention, and anything that a person skilled in the art can easily infer should be interpreted as falling within the scope of the invention. Hereinafter, the embodiments will be described in detail with reference to the accompanying drawings for illustrative purposes only.

[0011] FIG. 1 is a diagram illustrating the architecture of a neural network according to some embodiments.

[0012] Referring to FIG. 1, the neural network 1 is also represented by a mathematical model using nodes and edges. The neural network 1 is also an architecture of a deep neural network (DNN) or n-layer neural networks. The DNN or the n-layer neural network may correspond to a convolutional neural network (CNN), a recurrent neural network (RNN), a deep belief network, a restricted Boltzman machine, or the like. For example, the neural network 1 may be embodied by a convolutional neural network (CNN), but is not limited thereto. The neural network 1 in FIG. 1 may also correspond to some layers of a convolutional neural network. Thus, the neural network 1 may also correspond to a convolutional layer, a pooling layer, a fully connected layer, or the like of a convolutional neural network. However, for the sake of convenience, the following description will be given on the assumption that the neural network 1 corresponds to a convolution layer of a convolution neural network.

[0013] In the convolution layer, the first feature map FM1 also corresponds to an input feature map, and the second feature map FM2 also corresponds to an output feature map. The feature map may mean a data set in which various features of input data are expressed. The first feature map FM1 and the second feature map FM2 are also high-dimensional matrices of two or more dimensions, and have respective activation parameters. When the first feature map FM1 and the second feature map FM2 correspond to, for example, three-dimensional feature maps, the first feature map FM1 and the second feature map FM2 have a width (W) (also called a column), a height (H) (also called a row), and a depth (C). At that time, the depth (C) also corresponds to the number of channels.

[0014] The first feature map FM1 may include multiple input stripes. For example, the first feature map FM1 may include HxW input stripes. An input stripe is a channel-wise input data for one space of the input feature map and may have a size of 1x1xC. For example, the input stripe may include C input activations.

[0015] In the convolution layer, a convolution operation involving the first feature map FM1 and the weight map WM is performed, and as a result, a second feature map FM2 may be generated. The weight map WM may filter the first feature map FM1 and may also be referred to as a weight filter or weight kernel. In one example, the depth of the weight map WM, i.e., the number of channels, is the same as the depth of the first feature map FM1, i.e., the number of channels. The weight map WM is shifted in a manner of crossing the first feature map FM1 with a sliding window. During each shift, each weight included in the weight map WM is multiplied by all feature values ​​in the area overlapped with the first feature map FM1 and added. One channel of the second feature map FM2 may be generated by convolving the first feature map FM1 with the weight map WM.

[0016] Although one weight map WM is shown in FIG. 1, in practice, multiple weight maps may be convolved with the first feature map FM1 to generate multiple channels of the second feature map FM2. Meanwhile, the second feature map FM2 of the convolution layer may also be the input feature map of the next layer. For example, the second feature map FM2 may also be the input feature map of the pooling layer. However, it is not limited thereto.

[0017] FIG. 2 is a diagram illustrating operations performed in a neural network according to some embodiments.

[0018] Referring to FIG. 2, the neural network 2 has a structure including an input layer, a hidden layer, and an output layer, and receives input data (e.g., I 1 and I 2 ) and based on the result of the calculation, output data (e.g., O 1 and O 2 ) can be generated.

[0019] As described above, the neural network 2 may be a DNN or n-layer neural network including two or more hidden layers. For example, as shown in FIG. 2, the neural network 2 may be a DNN including an input layer (Layer 1), two hidden layers (Layer 2 and Layer 3), and an output layer (Layer 4). When the neural network 2 is embodied by a DNN architecture, the neural network 2 may process more complex data sets than a neural network having a single layer, since it includes more layers that can process useful information. Meanwhile, the neural network 2 is illustrated as including four layers, but this is merely an example, and the neural network 2 may include fewer or more layers, or fewer or more channels. That is, the neural network 2 may include layers of various structures different from those illustrated in FIG. 2.

[0020] Each layer included in the neural network 2 may include a number of channels, which may correspond to a number of artificial nodes, also known as neurons, processing elements (PEs), units, or similar terms. For example, as shown in FIG. 2, Layer 1 may include two channels (nodes), and Layer 2 and Layer 3 may each include three channels. However, this is by way of example only, and each layer included in the neural network 2 may include a variety of numbers of channels (nodes).

[0021] The channels included in each layer of the neural network 2 can be connected to each other to process data. For example, one channel can receive data from another channel, perform a calculation, and output the calculation result to yet another channel.

[0022] The input and output of each channel are also called input activation and output activation, respectively. That is, the activation is both the output of one channel and a parameter corresponding to the input of the channel included in the next layer. Meanwhile, each channel can determine its own activation based on the activation and weights received from the channel included in the previous layer. The weights are parameters used to calculate the output activation in each channel and are also values ​​assigned to the connection relationship between channels.

[0023] Each channel may also be processed by a computational unit or processing element that receives inputs and outputs activations, and the inputs and outputs of each channel may be mapped, e.g., where σ is the activation function and w i jk is the weight from the kth channel in the (i-1)th layer to the jth channel in the ith layer, and b i j is the bias of the j-th channel included in the i-th layer, and a i j is the activation of the j-th channel of the i-th layer, then activation a i j can be calculated using the following formula 1:

[0024]

number

[0025]

number

[0026] As described above, in the neural network 2, many data sets are exchanged between multiple interconnected channels and undergo a computation process through layers. In such a computation process, many MAC (multiply-accumulate) operations (or multiply-accumulate operations) are performed, along with many memory access operations to load the activations and weights, which are the operands of the MAC operations, at the appropriate time.

[0027] Meanwhile, a general digital computer uses a von Neumann structure in which the operation unit and the memory are separated and a common data bus is used for data transmission between the two separated blocks. Therefore, in the process of implementing a neural network 2 in which data transfer and calculation are continuously repeated, it takes a lot of time to transmit data and consumes excessive power. To overcome such problems, an in-memory computing circuit has been proposed as an architecture that integrates the memory and the operation unit for performing MAC calculations into one. The in-memory computing circuit will be described in more detail below with reference to FIG. 3.

[0028] FIG. 3 is a diagram illustrating an in-memory computing circuit according to some embodiments.

[0029] 3, the in-memory computing circuit 3 may include an analog crossbar array 30 and an ADC (analog to digital converter) 40. The ADC may also be referred to as an analog to digital converter. However, only components related to the present embodiment are illustrated in the in-memory computing circuit 3 illustrated in FIG. 3. Therefore, it will be apparent to those skilled in the art that the in-memory computing circuit 3 may further include other general-purpose components in addition to the components illustrated in FIG. 3.

[0030] The analog crossbar array 30 may include a number of row lines 310, a number of column lines 320, and a number of memory cells 330. The number of row lines 310 may also be used to receive input data. For example, if the number of row lines 310 is N (N is any natural number), the N row lines may receive a voltage V 1 ,V 2 ,…,V N may be applied. The plurality of column lines 320 may cross the plurality of row lines 310. For example, when the plurality of column lines 320 is M column lines (M is any natural number), the plurality of column lines 320 and the plurality of row lines 310 may cross at NxM crossing points.

[0031] Meanwhile, a plurality of memory cells 330 may be disposed at intersections of the row lines 310 and the column lines 320. Each of the memory cells 330 may be implemented as a non-volatile memory such as a resistive random access memory (ReRAM) or eFlash to store weights, but is not limited thereto. Each of the memory cells 330 may also be implemented as a volatile memory such as a static random access memory (SRAM).

[0032] In the example shown in FIG. 3, the memory cells 330 each have a conductance G corresponding to a weight. 11 ,…,G NM When a voltage corresponding to input activation is applied to each of the row lines 310, a current having a magnitude of I=V×G can be output through each memory cell 330 according to Ohm's law. Since the currents output from the memory cells arranged along one column line join together, the sum of the currents I 1 ,…,I M The sum of currents I 1 ,…,I M may correspond to the result of a MAC operation performed in an analog manner.

[0033] The ADC 40 receives the analog MAC operation result (i.e., the current sum I 1 ,…,I M ) can be converted into a digital signal. The result of the MAC operation converted into a digital signal is output from the ADC 40 and is also used in the subsequent neural network operation process.

[0034] In addition, the in-memory computing circuit 3 as shown in FIG. 3 has the advantages of a lower core calculation complexity, less power consumption, and smaller circuit size compared to a digital computer, but it can only perform fixed point-based calculations and has difficulty performing floating point-based calculations that support a large dynamic range.

[0035] Therefore, the conventional technology uses the in-memory computing circuit 3 only in the process of quantizing the trained neural network after training the neural network based on floating point, converting the trained neural network into a fixed point format, and implementing the quantized neural network. However, according to the conventional technology, accuracy loss occurs in the process of quantizing the neural network, or re-training is required to minimize the accuracy loss. In addition, the neural network implementing a specific application has a very large dynamic range of parameters, making it impossible to quantize while minimizing the accuracy loss.

[0036] According to the present disclosure, a neural network device capable of performing floating-point-based calculations while utilizing an in-memory computing circuit 3 having various advantages may be provided. Hereinafter, a method for performing floating-point calculations in the neural network device according to the present disclosure will be described in detail with reference to the drawings.

[0037] FIG. 4 is a schematic diagram illustrating the overall process of a neural network device performing floating-point arithmetic according to some embodiments.

[0038] 4, the neural network device 4 may include at least one control circuit 410, an in-memory computing circuit 420, and a digital arithmetic circuit 430. However, only components related to the present embodiment are shown in the neural network device 4 illustrated in FIG 4. Therefore, it will be obvious to those skilled in the art that the neural network device 4 may further include other general-purpose components in addition to the components illustrated in FIG 4.

[0039] The at least one control circuit 410 performs an overall function for controlling the neural network device 4. For example, the at least one control circuit 410 can control the operation of the in-memory computing circuit 420 and the digital arithmetic circuit 430. Meanwhile, the at least one control circuit 410 can also be realized by an array of a number of logic gates, or by a combination of a general-purpose microprocessor and a memory in which a program that can be executed by the microprocessor is stored.

[0040] At least one control circuit 410 may determine, for each weight kernel, a first shared exponent representative of the weights included in the weight kernel, and obtain a first block floating point weight including a first fraction adjusted based on the first shared exponent. The manner in which at least one control circuit 410 obtains the block floating point weights is described in further detail below with reference to Figures 5, 6A, and 6B.

[0041] FIG. 5 is a diagram illustrating a process in which a neural network device according to some embodiments converts weights into block floating point numbers.

[0042] Referring to Figure 5, when the height of the weight kernel is R, the width of the weight kernel is Q, the depth of the weight kernel (i.e., the number of channels) is C, and the number of weight kernels is K, the process of converting CRQK weights contained in the weight kernel into block floating point format is shown.

[0043] The neural network device according to the present disclosure performs floating-point calculations with low power consumption by utilizing an in-memory computing circuit (e.g., in-memory computing circuit 420 in FIG. 4) that can only perform fixed-point-based calculations, and it is necessary to separate the portion of the floating-point weight that can perform calculations using the in-memory computing circuit.

[0044] Meanwhile, the weights included in the weight kernel are floating-point format data and may have the same or different exponents depending on the size of the values ​​they indicate, but for efficient calculation, a block floating-point format in which exponents are shared for blocks of a specific size may be used. For example, as shown in FIG. 5, a shared exponent is extracted for each weight kernel, and weights included in one weight kernel may share the same exponent. Meanwhile, the shared exponent is determined to be the maximum value among the existing exponents of the weights in order to express all the weights included in the weight kernel, but is not necessarily limited thereto. As the exponent of the weight is changed from the existing exponent to the shared exponent, the decimal part of the weight is also adjusted through a corresponding shift operation.

[0045] As shown in Fig. 5, when the CRQK weights are converted to a block floating-point format per weight kernel, they are also represented by only CRQK sign bits, K shared exponents, and CRQK fractions. The K shared exponents are input to the digital arithmetic circuit 430 and used for digital arithmetic, and the CRQK fractions are input to the in-memory computing circuit 420 and also used for analog arithmetic. Hereinafter, with reference to Figs. 6A and 6B, a method of storing the shared exponent and fractions for use in digital arithmetic and analog arithmetic, respectively, will be described in detail.

[0046] FIG. 6A is a diagram illustrating a method for storing weights converted to block floating point according to some embodiments, and FIG. 6B is a diagram illustrating a method for storing weights converted to block floating point according to another embodiment.

[0047] On the other hand, the analog crossbar array 610 in FIG. 6A and the analog crossbar array 615 in FIG. 6B each correspond to the analog crossbar array 30 in FIG. 3, and the digital arithmetic circuit 620 in FIG. 6A and the digital arithmetic circuit 625 in FIG. 6B each correspond to the digital arithmetic circuit 430 in FIG. 4, so duplicated explanations will be omitted.

[0048] Referring to FIG. 6A, when weights are converted to have a shared exponent for each weight kernel as described in FIG. 5, a manner in which the shared exponent and decimal part of the weight are preserved for a particular layer is illustrated.

[0049] In one example, W denotes the weights included in the nth layer, the ith input channel, and the kth weight kernel. n,i,k can also be expressed as the following Equation 2.

[0050]

number

[0051] The fractional part of the weight in block floating-point format, F W(n,i,k) are also stored in the analog crossbar array 610. For example, if the analog crossbar array 610 includes M column lines, each of the M column lines corresponds to each of the M weight kernels, and the weights included in the weight kernels are also stored in memory cells arranged along the corresponding column lines. Weights, each consisting of f bits, are stored in units of stripes having a size of N, by the number of kernels (i.e., M).

[0052] On the other hand, the shared exponent E extracted from the weights W(n,k) may be stored separately in the digital arithmetic circuit 620. The digital arithmetic circuit 620 may include a register or memory for storing a digital value corresponding to the shared exponent. The shared exponents, each consisting of e bits, are stored as many times as the number of kernels (i.e., M), so a total of exM bits are stored.

[0053] If the analog crossbar array 610 supports signed weights, then E W(n,k) is stored separately in the storage space of the digital arithmetic circuit 620, and the weight sign bit S W is also stored in the analog crossbar array 610. However, if the analog crossbar array 610 supports weights with no sign, S W E W(n,k) As such, the data is stored in the storage space of the separate digital arithmetic circuit 620 in addition to the analog crossbar array 610 .

[0054] In the above, the case where the weights are converted to have a shared exponent for each weight kernel has been described with reference to FIG. 6A, but this is merely one example. The weights may also be converted to have a shared exponent for each predetermined group, not for each weight kernel. For example, as shown in FIG. 6B, a certain number of columns (i.e., weight kernels) included in the analog crossbar array 615 are grouped into one group, and a shared exponent is extracted for each group. In this case, the number of shared exponents for indicating the overall weights is reduced, and the amount of calculation in the process of converting the output of the analog crossbar array 615 to a final output may be reduced.

[0055] However, if an excessive number of columns are grouped into one group, the accuracy loss increases, but the number of groups can be appropriately set by considering the trade-off between the accuracy loss and the amount of calculation. In one example, the weights correspond to data known in advance by the neural network device, so the neural network device can determine an appropriate number of groups in advance through simulation. However, this is not necessarily limited, and the number of groups can also be determined in real time. In addition to the number of groups, the types of columns included in one group can also be determined.

[0056] Meanwhile, the shared exponent extracted from the weight may be separately stored in the digital arithmetic circuit 625. The digital arithmetic circuit 625 may include a register or memory for storing a digital value corresponding to the shared exponent. The number of the shared exponents stored may be equal to the number of groups (i.e., G).

[0057] In some embodiments, when the analog crossbar array 615 can amplify the sensed current or the input voltage with a predetermined amplification factor, a portion corresponding to the difference in the shared exponents between specific groups can be assigned to the amplification factor of the analog crossbar array 615 to unify the shared exponents between the groups. For example, when the analog crossbar array 615 can amplify the sensed current by 2 and the difference between the shared exponent assigned to the first group and the shared exponent assigned to the second group is 1, the same shared exponent is assigned to the first group and the second group instead of different shared exponents, but the sensed current can be amplified by 2 when obtaining the current sum corresponding to the first group. This can further reduce the number of shared exponents for indicating the overall weight.

[0058] Meanwhile, without being limited to the above examples, it will be easily understood by those skilled in the art that, instead of amplifying the sensed current, the input voltage can be amplified, or both the sensed current and the input voltage can be amplified, and the amplification factor can be set in various ways, such as 4 times or 8 times, instead of 2 times. If the sensed current and the input voltage can be amplified separately, and the amplification factor can be selected from various amplification factors, the number of shared exponents for indicating the overall weight can be further reduced.

[0059] Returning again to Figure 4, the at least one control circuit 410 may determine, for each of a plurality of input stripes included in the input feature map, a second shared exponent representative of the input activations included in the input stripe, and obtain input activations in a second block floating point format including second fractions adjusted based on the second shared exponent. A manner in which the at least one control circuit 410 obtains the block floating point input activations will be described in further detail below with reference to Figure 7.

[0060] FIG. 7 is a diagram illustrating a process in which a neural network device converts input activations into block floating point numbers according to some embodiments.

[0061] Referring to Figure 7, when the height of the input feature map IFM is H, the width of the input feature map IFM is W, and the depth of the input feature map IFM (i.e., the number of channels) is C, the process of converting HWC input activations contained in the input feature map IFM into a block floating-point format is illustrated.

[0062] The input activations included in the input feature map IFM may be converted to block floating-point format on an input stripe basis. For example, as shown in FIG. 7, when the input activations are converted to block floating-point format on an input stripe basis, the input activations included in one input stripe are also represented by only one shared exponent. One shared exponent corresponding to one input stripe is input to the digital arithmetic circuit 430 and used for digital arithmetic, and C fractions corresponding to one input stripe are input to the in-memory computing circuit 420 and also used for analog arithmetic. C sign bits corresponding to one input stripe are input to the digital arithmetic circuit 430 and used for digital arithmetic, or are input to the in-memory computing circuit 420 and also used for analog arithmetic, depending on the characteristics of the crossbar analog array (e.g., whether or not it supports inputs with signs).

[0063] Meanwhile, the block floating-point conversion of the input activations is also performed in real time by inputting the floating-point input feature map IFM. The shared exponent representing one input stripe is determined to be, but is not necessarily limited to, the maximum value among the existing exponents of the input activations included in the input stripe. The shared exponent is also appropriately determined according to a predefined rule.

[0064] The fractional part of the input activations is also adjusted via a corresponding shift operation by changing the exponent of the input activations from the existing exponent to a shared exponent. In one example, for the nth layer and the ith input channel, the input activations X n,i can also be expressed as the following Equation 3.

[0065]

number

[0066] 7, an example in which input activations are converted to block floating-point format in units of input stripes is described, but the present invention is not necessarily limited thereto. For example, when the number of channels C corresponding to one input stripe is greater than the number of row lines ROW of the analog crossbar array included in the in-memory computing circuit 420, one input stripe is C0 and C 1 (C=C 0 +C 1 , where C 0 <ROW、C 1 <ROWである)。

[0067] In that case, the input activations are also converted to block floating-point form on a stripe portion basis, but are not limited thereto, and the input activations may be converted to block floating-point form on an input stripe portion basis as is conventional, and input to the in-memory computing circuitry 420 only on a stripe portion basis.

[0068] 4, the at least one control circuit 410 can utilize weights and input activations in block floating point format to perform neural network operations. For example, the at least one control circuit 410 can utilize in-memory computing circuitry 420 and digital computing circuitry 430 to perform neural network operations.

[0069] The in-memory computing circuit 420 may include an analog crossbar array and an ADC, as described with reference to the in-memory computing circuit 3 of Figure 3. The analog crossbar array may include a number of row lines, a number of column lines intersecting the number of row lines, and a number of memory cells disposed at intersections of the number of row lines and the number of column lines.

[0070] At least one control circuit 410 can store first fractions corresponding to each weight in the weight kernel in memory cells arranged along the column lines corresponding to the weight kernel, and input second fractions corresponding to each input activation in the input stripe to the row lines. The analog crossbar array can then perform a MAC operation on the first fractions (corresponding to the weights) and the second fractions (corresponding to the input activations) in an analog manner, with the results of the MAC operation also being output along the column lines. Meanwhile, the ADC can convert the results of the MAC operation into digital signals to subsequently enable digital operations on the results of the MAC operation in the digital operation circuit 430.

[0071] The digital arithmetic circuit 430 may perform an addition operation on the first shared exponent and the second shared exponent, combine a result of the MAC operation with a result of the addition operation, and output a partial-sum of the floating-point output activations included in the channel of the output feature map. For example, the digital arithmetic circuit 430 may combine a result of the MAC operation corresponding to one input stripe with a result of the addition operation to obtain a partial sum used to calculate the floating-point output activations. When partial sums corresponding to all input stripes included in the input feature map are obtained, the digital arithmetic circuit 430 may use the obtained partial sums to calculate the floating-point output activations included in the channel of the output feature map.

[0072] On the other hand, if the weights are expressed as in Equation 2 above and the input activations are expressed as in Equation 3 above, the partial sum calculated for the kth weight kernel can also be expressed as Equation 4 below.

[0073]

number

[0074]

number

[0075] FIG. 8 is a diagram illustrating a process for performing floating-point arithmetic according to some embodiments in which an analog crossbar array supports signed inputs and weights.

[0076] In some embodiments, if the analog crossbar array supports signed inputs, the at least one control circuit 410 can input sign bits (e.g., IFM signs) of input activations in a second block floating point format along with second fractions (e.g., IFM fractions) onto multiple row lines. Also, if the analog crossbar array supports signed weights, the at least one control circuit 410 can store sign bits of weights in a first block floating point format along with the first fractions in memory cells.

[0077] Since the analog crossbar array supports both signed inputs and weights, simply inputting the sign bits of the input activations and weights together with the fractional part to the analog crossbar array will allow the MAC operation to take into account both the sign of the input activations and the sign of the weights. For example, the output CO of the analog crossbar array corresponding to the k-th weight kernel is k can also be calculated as in Equation 5 below.

[0078]

number

[0079] In some embodiments, since the analog crossbar array supports signed weights, the at least one control circuit 410 can store the sign bit of the first block floating point format weight in a memory cell along with the first fractional part, as described above with reference to Figure 8. This allows the analog crossbar array to store not only weights with positive signs, but also weights with negative signs.

[0080] On the other hand, if the analog crossbar array supports inputs without a sign, the at least one control circuit 410 may obtain a first current sum output along each of the plurality of column lines by preferentially activating only row lines whose sign bits of input activations in the second block floating point format are a first value. Thereafter, the at least one control circuit 410 may obtain a second current sum output along each of the plurality of column lines by activating only row lines whose sign bits of input activations in the second block floating point format are a second value. The first value and the second value may be 0 or 1, and may have different values ​​from each other.

[0081] In that way, the at least one control circuit 410 can separately obtain a first current sum corresponding to positive input activations and a second current sum corresponding to negative input activations over at least two cycles (e.g., cycle 0 and cycle 1). However, depending on the configuration of the in-memory computing circuit 420, the manner in which the first current sum and the second current sum are combined may vary.

[0082] In one example, as shown in FIG 9A, before the first current sum and the second current sum are combined, a digital conversion may be performed first. For example, an ADC may convert the first current sum to a first digital signal and the second current sum to a second digital signal. The time at which the first current sum is converted to the first digital signal and the time at which the second current sum is converted to the second digital signal may be different from each other, but is not limited thereto.

[0083] Meanwhile, in this example, the neural network device may further include a digital accumulator 910. The digital accumulator 910 may combine the first digital signal and the second digital signal to output a digital signal corresponding to the result of the MAC operation. For example, the digital accumulator 910 may combine the first digital signal and the second digital signal by adding the first digital signal corresponding to a positive input activation and subtracting the second digital signal corresponding to a negative input activation.

[0084] In another example, as shown in FIG. 9B, the combination of the first current sum and the second current sum is performed before the digital conversion. In this example, the neural network device may further include an analog accumulator 920 (920). The analog accumulator 920 can output a final current sum by combining the first current sum and the second current sum. For example, the analog accumulator 920 can output a final current sum by adding the first current sum corresponding to a positive input activation and subtracting the second current sum corresponding to a negative input activation. The ADC can convert the final current sum output from the analog accumulator 920 into a digital signal corresponding to the result of the MAC operation.

[0085] 9A shows digital accumulator 910 located outside in-memory computing circuitry 420, and FIG. 9B shows analog accumulator 920 located within in-memory computing circuitry 430, but is not limited thereto. Digital accumulator 910 and analog accumulator 920 may each be located in any suitable location within or outside in-memory computing circuitry 420.

[0086] FIG. 10 is a diagram illustrating a process for performing floating-point arithmetic according to some embodiments in which an analog crossbar array supports unsigned inputs and weights.

[0087] If the analog crossbar array supports weights without signs, at least one control circuit 410 can store first fractional parts corresponding to each weight whose sign bit is a first value in memory cells arranged along a first column line of the analog crossbar array and store first fractional parts corresponding to each weight whose sign bit is a second value in memory cells arranged along a second column line of the analog crossbar array.

[0088] The in-memory computing circuit 420 can output a final current sum by combining the first current sum output along each of the first column lines and the second current sum output along each of the second column lines. As such, the analog crossbar array included in the in-memory computing circuit 420 can also be operated to include two crossbar arrays (i.e., a first crossbar array 1010 that stores positive weights and a second crossbar array 1020 that stores negative weights).

[0089] The in-memory computing circuit 420 can output a final current sum by combining a first current sum output along each of the first column lines and a second current sum output along each of the second column lines using an add / subtract (ADD / SUB) module 1030. For example, the add / subtract (ADD / SUB) module 1030 can output a final current sum by adding the first current sum output from the first crossbar array 1010 and subtracting the second current sum output from the second crossbar array 1020.

[0090] On the other hand, since the analog crossbar array supports unsigned inputs, the at least one control circuit 410 can input positive and negative input activations separately to the in-memory computing circuit 420 over two cycles (e.g., cycle 0 and cycle 1) as first described with reference to Figures 9A and 9B, such that the final current sums corresponding to the positive input activations can be output from the analog crossbar array in cycle 0, and the final current sums corresponding to the negative input activations can be output from the analog crossbar array in cycle 1.

[0091] The accumulator 1040 may combine the final current sums corresponding to the positive input activations and the final current sums corresponding to the negative input activations to obtain fractional partial sums. For example, the accumulator 1040 may add the final current sums corresponding to the positive input activations and subtract the final current sums corresponding to the negative input activations to obtain fractional partial sums.

[0092] 10, an example in which the analog crossbar array supports unsigned weights and inputs has been described, but the analog crossbar array can also support unsigned weights but signed inputs. In that case, since input over two cycles is not required considering the sign of the input activation, all input activations are input in one cycle regardless of the sign, and the configuration of the accumulator 1040 can be omitted.

[0093] In the above, a method for calculating the partial sum of the fractional part according to whether the analog crossbar array included in the in-memory computing circuit 420 supports weights or inputs having a sign has been described in detail with reference to Figures 8 to 10. The partial sum of the fractional part output from the analog crossbar array is input to the digital arithmetic circuit 430 and is also used for calculating the final output. Hereinafter, a method for the digital arithmetic circuit 430 to calculate the final output will be described in detail with reference to Figure 11.

[0094] FIG. 11 is a diagram illustrating a process in which a final output of a floating-point operation is output by combining an operation result of an analog crossbar array with an operation result of a digital operation circuit according to some embodiments.

[0095] The digital arithmetic circuit 430 may obtain a third fractional part by performing a shift operation on the result of the MAC operation output from the analog crossbar array so that the most significant bit becomes 1. The digital arithmetic circuit 430 may include a shift operator 1110 for performing the shift operation. The third fractional part after the shift operation may correspond to a fractional part of the partial sum of the output activations.

[0096] The digital arithmetic circuit 430 may obtain a third exponent by performing a conversion operation of adding or subtracting the number of times a shift operation is performed before an addition result involving a first shared exponent (i.e., a shared exponent of a weight) and a second shared exponent (i.e., a shared exponent of an input activation). The third exponent on which the conversion operation is performed may correspond to an exponent of a partial sum of an output activation. Thus, the digital arithmetic circuit 430 may output a partial sum of a floating-point output activation including a third fraction and a third exponent.

[0097] After the digital arithmetic circuit 430 calculates the partial sums of the output activations corresponding to one input stripe included in the input feature map, the digital arithmetic circuit 430 can also sequentially calculate the partial sums of the output activations corresponding to the remaining input stripes. When the partial sums corresponding to all the input stripes included in the input feature map are calculated, the digital arithmetic circuit 430 can obtain the final output activations based on the calculated partial sums.

[0098] Meanwhile, the neural network device can selectively apply an activation function during or before the process of combining the result of the MAC operation and the result of the addition operation in the digital arithmetic circuit 430 to obtain a floating-point output activation. In one example, when ReLU, which outputs 0 when a negative number is input, is applied as the activation function, the neural network device can determine whether the output activation is a negative number based on the sign bit included in the result of the MAC operation, and if it is determined that the output activation is a negative number, omit the shift operation and conversion operation and output the output activation as 0. This can omit unnecessary operations and operations.

[0099] The floating-point output activations finally output from the digital arithmetic circuit 430 are also used as input activations for the next layer. In the next layer, the above process is repeated, and a forward pass or a backward pass may be performed along the layer of the neural network implemented by the neural network device. Also, by performing the forward pass or the backward pass, learning may be performed for the neural network implemented by the neural network device, or inference may be performed using the neural network.

[0100] The neural network device according to the present disclosure converts input activations and weights into block floating-point format, and then performs operations related to the exponent part, for which accuracy is important, digitally, and performs floating-point operations while minimizing power consumption and accuracy loss by utilizing an existing fixed-point based analog crossbar array for operations related to the fractional part, for which many operations are required.

[0101] FIG. 12 is a flow chart illustrating a method of operation of a neural network device according to some embodiments.

[0102] 12, the operation method of the neural network device is composed of steps that are processed in a time series manner in the neural network device 4 shown in FIG 4. Therefore, even if the content is omitted below, it can be understood that the content described above with respect to FIG 4 to FIG 11 also applies to the operation method of the neural network device of FIG 12.

[0103] In step 1210, the neural network device may determine, for each weight kernel, a first shared exponent representative of the weight included in the weight kernel, and obtain a weight in a first block floating-point format including a first fraction adjusted based on the first shared exponent.

[0104] At step 1220, the neural network device may determine, for each of a plurality of input stripes included in the input feature map, a second shared exponent representative of the input activations included in the input stripe, and obtain input activations in a second block floating point format including a second fraction adjusted based on the second shared exponent.

[0105] In step 1230, the neural network device may use an analog crossbar array to perform a MAC operation on the first fractional part and the second fractional part, and use an ADC to convert the result of the MAC operation into a digital signal.

[0106] The neural network device can store first fractional parts corresponding to each weight included in the weight kernel in memory cells arranged along the column lines corresponding to the weight kernel in a plurality of column lines of the analog crossbar array, and can input second fractional parts corresponding to each input activation included in the input stripe to a plurality of row lines of the analog crossbar array, so that the MAC operation related to the first fractional parts and the second fractional parts is also performed in an analog manner.

[0107] In some embodiments, if the analog crossbar array supports signed inputs, the neural network device can input the sign bit of the input activation in a second block floating point format along with the second fractional part to a number of row lines, so that the analog crossbar array can output a result of the MAC operation that takes into account the sign bit of the input activation.

[0108] However, according to another embodiment, when the analog crossbar array supports inputs without signs, the neural network device can obtain a first current sum output along each of the plurality of column lines by preferentially activating only row lines whose sign bits of input activations in the second block floating point format are a first value. Then, the neural network device can obtain a second current sum output along each of the plurality of column lines by activating only row lines whose sign bits of input activations in the second block floating point format are a second value. The first value and the second value can be 0 or 1, and can have different values ​​from each other.

[0109] In this way, the neural network device can separately obtain a first current sum corresponding to positive input activations and a second current sum corresponding to negative input activations over at least two cycles, whereas the manner in which the first current sum and the second current sum are combined may differ depending on the configuration of the in-memory computing circuit including the analog crossbar array.

[0110] In one example, the neural network device may first perform digital conversion before the first current sum and the second current sum are combined. The neural network device may use an ADC to convert the first current sum to a first digital signal and use an ADC to convert the second current sum to a second digital signal. The neural network device may use a digital accumulator to combine the first digital signal and the second digital signal to output a digital signal corresponding to the result of the MAC operation. In an example where the first value is 0 and the second value is 1, the neural network device may output a digital signal corresponding to the result of the MAC operation by subtracting the second digital signal from the first digital signal.

[0111] In another example, the neural network device may combine the first and second current sums prior to digital conversion. The neural network device may use an analog accumulator to combine the first and second current sums to output a final current sum. In an example where the first value is 0 and the second value is 1, the neural network device may subtract the second current sum from the first current sum to output a final current sum. The neural network device may also use an ADC to convert the final current sum into a digital signal corresponding to the result of the MAC operation.

[0112] Meanwhile, in some embodiments, when the crossbar array supports weights having signs, the neural network device may store the sign bits of the weights in the first block floating point format in memory cells together with the first fractional part, so that the analog crossbar array may output the results of the MAC operation taking into account the sign bits of the weights.

[0113] According to another embodiment, when the analog crossbar array supports weights without signs, the neural network device can store first fractions corresponding to each weight whose sign bit is a first value in memory cells arranged along a first column line of the analog crossbar array, and store first fractions corresponding to each weight whose sign bit is a second value in memory cells arranged along a second column line of the analog crossbar array. The first value and the second value can be 0 or 1 and can have different values ​​from each other.

[0114] The neural network device can output a final current sum by combining the first current sum output along each of the first column lines and the second current sum output along each of the second column lines. As such, the analog crossbar array included in the neural network device can also be operated to include two crossbar arrays (i.e., a first crossbar array in which the sign bits of the weights are a first value, and a second crossbar array in which the sign bits of the weights are a second value).

[0115] At step 1240, the neural network device may utilize digital arithmetic circuitry to perform an addition operation involving the first shared exponent and the second shared exponent, combine a result of the MAC operation with a result of the addition operation, and output a partial sum of the floating-point output activations included in the channel of the output feature map.

[0116] For example, the neural network device may obtain a third fractional part by performing a shift operation on the result of the MAC operation so that the most significant bit becomes 1. The neural network device may obtain a third exponent part by performing a conversion operation of adding or subtracting the number of times the shift operation has been performed on the result of the addition operation. As a result, the neural network device may output a partial sum of the floating-point output activations including the third fractional part and the third exponent part.

[0117] After the neural network device calculates the partial sums of the output activations corresponding to one input stripe included in the input feature map, the neural network device can also sequentially calculate the partial sums of the output activations corresponding to the remaining input stripes. When the partial sums corresponding to all the input stripes included in the input feature map are calculated, the neural network device can obtain the final output activations according to the calculated partial sums.

[0118] Meanwhile, the neural network device can selectively apply an activation function during or before the process of combining the result of the MAC operation and the result of the addition operation to obtain a floating-point output activation. In one example, when ReLU, which outputs 0 when a negative number is input, is applied as an activation function, the neural network device can determine whether the output activation is negative based on the sign bit included in the result of the MAC operation, and if it is determined that the output activation is negative, omit the shift operation and conversion operation and output the output activation as 0. This can omit unnecessary operations and operations.

[0119] In the method of operating the neural network device according to the present disclosure, input activations and weights are converted into a block floating-point format, and then calculations related to the exponent part, for which accuracy is important, are performed digitally, while calculations related to the fractional part, for which many calculations are required, are performed using an existing fixed-point based analog crossbar array, thereby minimizing power consumption and loss of accuracy while performing floating-point calculations.

[0120] 12 are described in order to explain the overall flow of the operation method of the neural network device, and are not necessarily performed in the described order. For example, after step 1210 is performed and steps 1220 to 1240 are performed for one input stripe included in the input feature map, step 1210 is not performed any more. Instead of performing step 1210 further, steps 1220 to 1240 may be repeatedly performed for the remaining input stripes included in the input feature map.

[0121] Also, when floating-point operations for outputting output activations corresponding to other input feature maps are started after performing steps 1220 to 1240 for all input stripes included in one input feature map to output output activations corresponding to the input feature map, the performance of step 1210 may be omitted. For example, in the process of performing floating-point operations for all input feature maps, step 1210 may be performed only once at the beginning. However, this is merely an example, and step 1210 may be performed again at any time when it is necessary to change weights used in the floating-point operations. Also, each of steps 1220 to 1240 may be performed at an appropriate time depending on the situation.

[0122] FIG. 13 is a block diagram illustrating a configuration of an electronic system according to some embodiments.

[0123] 13, the electronic system 1300 can analyze input data in real time based on a neural network, extract useful information, and make a situational decision based on the extracted information, or control the configuration of an electronic device in which the electronic system 1300 is installed. For example, the electronic system 1300 can be applied to robotic devices such as drones and advanced driver assistance systems (ADAS), smart TVs, smartphones, medical devices, mobile devices, image display devices, measuring devices, and Internet of Things (IoT) devices, and can also be installed in at least one of various other electronic devices.

[0124] The electronic system 1300 may include a processor 1310, a random access memory (RAM) 1320, a neural network device 1330, a memory 1340, a sensor module 1350, and a communication module 1360. The electronic system 1300 may further include an input / output module, a security module, a power control device, etc. A part of the hardware configuration of the electronic system 1300 is also mounted on at least one semiconductor chip.

[0125] The processor 1310 controls the overall operation of the electronic system 1300. The processor 1310 may include one processor core (single core) or multiple processor cores (multi-core). The processor 1310 may process or execute programs and / or data stored in the memory 1340. In some embodiments, the processor 1310 may control the function of the neural network device 1330 by executing programs stored in the memory 1340. The processor 1310 may also be embodied as a central processing unit (CPU), a graphics processing unit (GPU), an application processor (AP), etc.

[0126] The RAM 1320 may temporarily store programs, data, or instructions. For example, the programs and / or data stored in the memory 1340 may be temporarily stored in the RAM 1320 under the control or startup code of the processor 1310. The RAM 1320 may be implemented by a memory such as a dynamic random access memory (DRAM) or a static random access memory (SRAM).

[0127] The neural network device 1330 can perform neural network operations based on received input data and generate information signals based on the results of the operations. The neural network may include, but is not limited to, CNN, RNN, FNN, deep belief networks, restricted Boltzman machines, etc. The neural network device 1330 is a dedicated neural network hardware accelerator itself or a device including the same, and corresponds to the above-mentioned neural network device (e.g., the neural network device 4 in FIG. 4).

[0128] The neural network device 1330 can perform floating-point arithmetic while utilizing an in-memory computing circuit that can significantly reduce power consumption, and can minimize accuracy loss by digitally processing exponent-related arithmetic that is sensitive to errors.

[0129] The information signal may include one of various types of recognition signals such as a voice recognition signal, an object recognition signal, a video recognition signal, a biometric recognition signal, etc. For example, the neural network device 1330 may receive frame data included in a video stream as input data and generate a recognition signal related to an object included in an image represented by the frame data from the frame data. However, the present invention is not limited thereto, and depending on the type or function of the electronic device installed in the electronic system 1300, the neural network device 1330 may receive various types of input data and generate a recognition signal based on the input data.

[0130] The memory 1340 is a storage location for storing data, and may store an operating system (OS), various programs, and various data. In one embodiment, the memory 1340 may store intermediate results generated during the process of performing calculations by the neural network device 1330.

[0131] The memory 1340 may be, but is not limited to, a DRAM. The memory 1340 may include at least one of a volatile memory or a non-volatile memory. The non-volatile memory may include a read only memory (ROM), a programmable read only memory (PROM), an erasable programmable read only memory (EPROM), an electrically erasable programmable read only memory (EEPROM), a flash memory, a phase change random access memory (PRAM), a magnetic random access memory (MRAM), a resistive random access memory (ReRAM), a ferroelectric random access memory (FeRAM), etc. The volatile memory may include a DRAM, an SRAM, a synchronous dynamic random access memory (SDRAM), etc. In one embodiment, the memory 1340 may include at least one of a hard disk drive (HDD), a solid static driver (SSD), a CF, an SD, a micro-SD, a mini-SD, an xD, or a memory stick.

[0132] The sensor module 1350 may collect information about the surroundings of an electronic device in which the electronic system 1300 is mounted. The sensor module 1350 may sense or receive a signal (e.g., a video signal, an audio signal, a magnetic signal, a biosignal, a touch signal, etc.) from outside the electronic device and convert the sensed or received signal into data. To this end, the sensor module 1350 may include at least one of various sensing devices, such as a microphone, an image capture device, an image sensor, a LIDAR (light detection and ranging) sensor, an ultrasonic sensor, an infrared sensor, a biosensor, and a touch sensor.

[0133] The sensor module 1350 can provide the converted data as input data to the neural network device 1330. For example, the sensor module 1350 can include an image sensor to capture an external environment of the electronic device, generate a video stream, and provide successive data frames of the video stream in turn as input data to the neural network device 1330. However, without being limited thereto, the sensor module 1350 can provide various types of data to the neural network device 1330.

[0134] The communication module 1360 may include various wired or wireless interfaces capable of communicating with external devices. For example, the communication module 1360 may include a communication interface that can connect to a wired local area network (LAN), a wireless local area network (WLAN) such as wireless fidelity (Wi-Fi), a wireless personal area network (WPAN) such as Bluetooth, a wireless universal serial bus (USB), Zigbee, near field communication (NFC), radio frequency identification (RFID), power line communication (PLC), or a mobile cellular network such as 3rd generation (3G), 4th generation (4G), or long term evolution (LTE).

[0135] The above-described embodiments of the present invention can be created into a program that can be executed by a computer, and can also be realized by a general-purpose digital computer that runs the program using a computer-readable recording medium. The data structures used in the above-described embodiments of the present invention can also be recorded in a computer-readable recording medium through various means. The above-described computer-readable recording medium includes recording media such as magnetic recording media (e.g., ROM, floppy disk, hard disk, etc.) and optically readable media (e.g., CD-ROM (compact disc read only memory), DVD (digital versatile disc), etc.).

[0136] The present invention has been described above with reference to its preferred embodiments. Those skilled in the art will understand that the present invention can be embodied in modified forms without departing from the essential characteristics of the present invention. Therefore, the disclosed embodiments should be considered from an illustrative rather than a restrictive perspective. The scope of the present invention is defined by the claims, not the above description, and all differences within the scope of the equivalents should be interpreted as being included in the present invention. [Explanation of symbols]

[0137] 1,2 Neural Networks 3,420 In-memory computing circuits 30,610,615 Analog Crossbar Array 40 ADC 4,1330 Neural network device 330 memory cells 410 At least one control circuit 430,620,625 Digital arithmetic circuit 910,920 Digital Accumulator 1030 Addition / Subtraction Module 1040 Accumulator 1110 Shift Operator 1300 Electronic Systems 1310 Processor 1320 RAM 1340 Memory 1350 Sensor Module 1360 Communication Module

Claims

1. A method for operating a neural network device that performs floating-point arithmetic, comprising: determining, for each weight kernel, a first shared exponent representative of the weights included in the weight kernel, and obtaining a first block floating-point format weight including a first fraction adjusted based on the first shared exponent; determining, for each of a plurality of input stripes included in the input feature map, a second shared exponent representative of the input activations included in the input stripe, and obtaining input activations in a second block floating point format including a second fraction adjusted based on the second shared exponent; performing a MAC operation on the first fractional part and the second fractional part using an analog crossbar array, and converting a result of the MAC operation into a digital signal using an ADC; performing an addition operation involving the first shared exponent and the second shared exponent using digital arithmetic circuitry, combining a result of the MAC operation and a result of the addition operation to output a partial sum of floating-point output activations included in a channel of an output feature map.

2. The method comprises: storing the first fractional parts corresponding to each of the weights included in the weight kernel in memory cells arranged along a column line corresponding to the weight kernel in a plurality of column lines of the analog crossbar array; 2. The method of claim 1, further comprising: inputting the second fractional portions corresponding to each of the input activations included in the input stripe into a plurality of row lines of the analog crossbar array.

3. If the analog crossbar array supports signed inputs, the inputting step comprises:

3. The method of claim 2, further comprising inputting a sign bit of an input activation in the second block floating point format along with the second fractional portion onto the plurality of row lines.

4. If the analog crossbar array supports unsigned inputs, the method further comprises: obtaining a first sum of currents output along each of the plurality of column lines by activating only row lines having a sign bit of an input activation of the second block floating point format having a first value; 3. The method of claim 2, further comprising: obtaining a second current sum output along each of the plurality of column lines by activating only row lines whose sign bit of the input activation of the second block floating point format is a second value.

5. The method comprises: converting the first sum of currents into a first digital signal using the ADC; converting the second current sum to a second digital signal using the ADC; 5. The method of claim 4, further comprising: combining the first digital signal and the second digital signal using a digital accumulator to output the digital signal corresponding to a result of the MAC operation.

6. The method comprises: combining the first current sum and the second current sum using an analog accumulator to output a final current sum; 5. The method of claim 4, further comprising: utilizing the ADC to convert the final current sum to the digital signal corresponding to a result of the MAC operation.

7. If the analog crossbar array supports signed weights, the storing step comprises:

3. The method of claim 2, further comprising storing a sign bit of the weight of the first block floating-point format in the memory cell along with the first fractional portion.

8. If the analog crossbar array supports weights without signs, the method further comprises: storing first fractional parts corresponding to each weight having a first value of sign bit in memory cells arranged along a first column line of the analog crossbar array; storing first fractional parts corresponding to each of the weights whose sign bits have a second value in memory cells disposed along a second column line of the analog crossbar array; 3. The method of claim 2, further comprising: outputting a final current sum by combining a first current sum output along each of the first column lines and a second current sum output along each of the second column lines.

9. The outputting step includes: performing a shift operation on the result of the MAC operation so that the most significant bit becomes 1 to obtain a third fraction part; obtaining a third exponent by performing a conversion operation of adding or subtracting the number of times the shift operation has been performed to a result of the addition operation; and outputting a partial sum of the floating-point output activations including the third fractional portion and the third exponent portion.

10. The method comprises: determining whether the floating-point output activation is negative based on a sign bit included in a result of the MAC operation; 10. The method of claim 9, further comprising: if the floating-point output activation is determined to be a negative number, then omitting the shift operation and the convert operation and outputting the floating-point output activation to zero.

11. In a neural network device performing floating-point arithmetic, at least one control circuitry for determining, for each weight kernel, a first shared exponent representative of the weights included in the weight kernel, and obtaining weights in a first block floating point format with a first fractional part scaled based on the first shared exponent, and for each of a plurality of input stripes included in the input feature map, determining a second shared exponent representative of input activations included in the input stripe, and obtaining input activations in a second block floating point format with a second fractional part scaled based on the second shared exponent; an in-memory computing circuit including an analog crossbar array that performs a MAC operation on the first fractional part and the second fractional part, and an ADC that converts a result of the MAC operation into a digital signal; a digital arithmetic circuit that performs an addition operation on the first shared exponent and the second shared exponent, combines a result of the MAC operation and a result of the addition operation, and outputs a partial sum of floating-point output activations included in a channel of an output feature map.

12. The analog crossbar array comprises: Multiple lowlines and a plurality of column lines intersecting the plurality of row lines; a plurality of memory cells arranged at intersections of the plurality of row lines and the plurality of column lines; The at least one control circuit comprises:

12. The neural network device of claim 11, wherein the first fractional parts corresponding to each of the weights included in the weight kernel are stored in memory cells arranged along a column line among the plurality of column lines corresponding to the weight kernel, and the second fractional parts corresponding to each of the input activations included in the input stripe are input to the plurality of row lines.

13. If the analog crossbar array supports signed inputs, the at least one control circuitry 13. The neural network device of claim 12, wherein a sign bit of the input activation in the second block floating point format is input to the plurality of row lines along with the second fractional portion.

14. If the analog crossbar array supports unsigned inputs, the at least one control circuitry 13. The neural network device of claim 12, wherein a first current sum is output along each of the plurality of column lines by activating only row lines whose sign bit of the input activation in the second block floating point format is a first value, and a second current sum is output along each of the plurality of column lines by activating only row lines whose sign bit of the input activation in the second block floating point format is a second value.

15. The ADC comprises: converting the first current sum into a first digital signal and converting the second current sum into a second digital signal; The neural network device comprises:

15. The neural network device of claim 14, further comprising a digital accumulator that combines the first digital signal and the second digital signal to output the digital signal corresponding to a result of the MAC operation.

16. The neural network device comprises: an analog accumulator that combines the first current sum and the second current sum to output a final current sum; The ADC comprises:

15. The neural network device of claim 14, further comprising: converting said final current sum into said digital signal corresponding to a result of said MAC operation.

17. If the analog crossbar array supports signed weights, the at least one control circuitry 13. The neural network device according to claim 12, wherein a sign bit of the weight in the first block floating point format is stored in the memory cell together with the first fractional part.

18. If the analog crossbar array supports unsigned weights, the at least one control circuitry storing first fractional parts corresponding to each weight whose sign bit has a first value in memory cells arranged along a first column line of the analog crossbar array, and storing first fractional parts corresponding to each weight whose sign bit has a second value in memory cells arranged along a second column line of the analog crossbar array; The in-memory computing circuit comprises:

13. The neural network device of claim 12, wherein a final current sum is output by combining a first current sum output along each of the first column lines and a second current sum output along each of the second column lines.

19. The digital arithmetic circuit includes:

19. The neural network device of claim 11, further comprising: a shift operation for the result of the MAC operation so that the most significant bit becomes 1 to obtain a third fractional part; a conversion operation for the result of the addition operation to add or subtract the number of times the shift operation has been performed to obtain a third exponent part; and an output of a partial sum of the floating-point output activations including the third fractional part and the third exponent part.

20. The digital arithmetic circuit includes: The neural network device of claim 19, further comprising: determining whether the floating-point output activation is a negative number based on a sign bit included in the result of the MAC operation; and, if the floating-point output activation is determined to be a negative number, skipping the shift operation and the conversion operation and outputting the floating-point output activation to 0.

Citation Information

Patent Citations

  • Block floating point for neural network implementations

    US20180157465A1

  • Digital Architecture Supporting Analog Co-Processor

    US20190205741A1

  • Accelerated quantized multiply-and-add operations

    WO2019183202A1

  • Matrix vector multiplier with a vector register file comprising a multi-port memory

    WO2019204068A1

  • Multiply and accumulate circuit

    WO2020060769A1