Device for implementing a neural network and its operation method
The neural morphic device addresses inefficiencies in memory-centric neural networks by using binary weight values and time-domain binary vectors for convolution operations, achieving reduced model size and computation with maintained accuracy.
Patent Information
- Application Number
- JP2021094973
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-08
- Filing Date
- 2021-06-07
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-06-07
AI Technical Summary
Existing memory-centric neural network technologies face challenges in efficiently processing large amounts of input data in real-time to extract desired information, requiring improved computational efficiency and reduced model size.
A neural morphic device utilizing a crossbar array circuit with binary weight values and time-domain binary vectors for convolution operations, reducing model size and computation through binary weight values and time-axis XNOR operations.
This approach maintains learning performance and classification accuracy similar to multi-bit data while significantly reducing model size and computation requirements.
Smart Images

Figure 0007708381000010 
Figure 0007708381000011 
Figure 0007708381000012
Abstract
Description
Technical Field
[0001] The present invention relates to an apparatus for implementing a neural network and an operating method thereof.
Background Art
[0002] Memory-centric neural network devices refer to a computer science architecture that models the biological brain. With the development of memory-centric neural network technology, research is actively underway to utilize memory-centric neural networks to analyze input data and extract effective information in various electronic systems.
[0003] Therefore, in order to utilize a memory-centric neural network to analyze a large amount of input data in real time and extract desired information, a technology capable of efficiently processing operations is required.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] The problem to be solved by the present invention is to provide an apparatus and a method for generating a chemical structure using a neural network. In addition, a computer-readable recording medium recording a program for causing a computer to execute the above method is provided. The technical problems to be solved are not limited to the above-mentioned technical problems, and other technical problems may exist.
Means for Solving the Problems
[0006] As a technical means for achieving the foregoing technical problem, a first aspect of the present disclosure provides a neural morphic device for implementing a neural network, including a memory storing at least one program, an on-chip memory including a crossbar array circuit, and at least one processor configured to drive the neural network by executing the at least one program. The at least one processor stores binary weight values in synapse circuits included in the crossbar array circuit, obtains an input feature map from the memory, converts the input feature map into a temporal domain binary vector, provides the temporal domain binary vector as an input value of the crossbar array circuit, and outputs an output feature map by performing a convolution operation between the binary weight values and the temporal domain binary vector.
[0007] A second aspect of the present disclosure provides a neural network device for implementing a neural network, including a memory storing at least one program, and at least one processor configured to drive the neural network by executing the at least one program. The at least one processor obtains binary weight values and an input feature map from the memory, converts the input feature map into a temporal domain binary vector, and outputs an output feature map by performing a convolution operation between the binary weight values and the temporal domain binary vector.
[0008] A third aspect of the present disclosure can provide a method in a neural morphic device for implementing a neural network, the method including: storing binary weight values in synaptic circuits included in a crossbar array circuit; obtaining an input feature map from a memory; converting the input feature map into a time-domain binary vector; providing the time-domain binary vector as an input value of the crossbar array circuit; and outputting an output feature map by performing a convolution operation between the binary weight values and the time-domain binary vector.
[0009] A fourth aspect of the present disclosure can provide a method in a neural network device for implementing a neural network, the method including: obtaining binary weight values and an input feature map from a memory; converting the input feature map into a time-domain binary vector; and outputting an output feature map by performing a convolution operation between the binary weight values and the time-domain binary vector.
[0010] A fifth aspect of the present disclosure can provide a computer-readable recording medium recording a program for causing a computer to execute the methods of the third and fourth aspects.
Advantages of the Invention
[0011] According to the above-described problem-solving means of the present disclosure, by using binary weight values and time-domain binary vectors, the model size and the amount of computation can be reduced.
[0012] Also, according to one of the other problem-solving means of the present disclosure, by performing a time-axis XNOR operation between binary weight values and a time-domain binary vector, learning performance and final classification / recognition accuracy similar to those of a neural network using multi-bit data can be ensured.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2A
Figure 2B
Figure 3A
Figure 3B
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 6C
Figure 7A
Figure 7B
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Embodiments for Carrying Out the Invention
[0014] In this specification, the phrases "in some embodiments" or "in an embodiment" described in various places do not necessarily indicate the same embodiments.
[0015] Some embodiments of the present disclosure are also shown by functional block configurations and various processing steps. Some or all of such functional blocks are also embodied by various numbers of hardware and / or software configurations that perform specific functions. For example, the functional blocks of the present disclosure are embodied by one or more microprocessors or by circuit configurations for a predetermined function. Also, for example, the functional blocks of the present disclosure are embodied by various programming languages or scripting languages. The functional blocks are also embodied by algorithms executed by one or more processors. Further, the present disclosure can employ prior art for electronic environment settings, signal processing, and / or data processing. Terms such as "mechanism", "element", "means", and "configuration" are used generically and are not limited to mechanical and physical configurations.
[0016] Also, the connection lines or connection members between the components illustrated in the drawings merely exemplarily show functional connections and / or physical or circuit connections. In an actual device, the connections between components are shown by various functional connections, physical connections, or circuit connections that can be alternative or additional.
[0017] Hereinafter, the present disclosure will be described in detail with reference to the accompanying drawings.
[0018] FIG. 1 is a drawing for explaining a biological neuron and a mathematical model that mimics the operation of the biological neuron.
[0019] The biological neuron means a cell existing in the human nervous system. The biological neuron is one of the basic biological computing individuals. The human brain contains approximately 100 billion biological neurons and approximately 100 trillion connections (interconnects) between the biological neurons.
[0020] Referring to FIG. 1, the biological neuron 10 is a single cell. The biological neuron 10 includes a neuronal cell body containing a cell nucleus and various cell organelles. The various cell organelles include mitochondria, a large number of dendrites radiating from the cell body, and an axon process longitudinally cut by many branched extensions.
[0021] Generally, the axon process performs the function of transmitting signals from one neuron to another, and the dendrites perform the function of receiving signals from other neurons. For example, when different neurons are connected to each other, the signals transmitted through the axon process of a neuron can be received by the dendrites of other neurons. At this time, signals between neurons are transmitted through a specialized connection called a synapse, and various neurons are connected to each other to form a neural network. The neuron that secretes neurotransmitters based on the synapse is called the pre-synaptic neuron, and the neuron that receives the information transmitted through the neurotransmitter is also called the post-synaptic neuron.
[0022] On the other hand, the human brain can learn and remember a huge amount of information by transmitting and processing various signals through a neural network formed by a large number of neurons connected to each other. The huge number of connections between neurons in the human brain is directly related to the massively parallel nature of biological computing, and various attempts have been made to imitate the artificial neural network and efficiently process a huge amount of information. For example, as a computing system designed to implement the artificial neural network at the neuron level, the neural morphic device is being studied.
[0023] On the one hand, the operation of the biological neuron 10 is also replicated in the mathematical model 11. The mathematical model 11 corresponding to the biological neuron 10 includes, as an example of neural morphic computing, a multiplication operation of multiplying synaptic weights for information from a number of neurons, an addition operation (Σ) for the values (ω0x0, ω1x1, ω2x2) multiplied by the synaptic weights, and an operation of applying a characteristic function (b) and an activation function (f) to the addition operation result. A neural morphic computing result can be provided by neural morphic computing. Here, values such as x0, x1, x2, … correspond to axon values, and values such as ω0, ω1, ω2, … correspond to synaptic weights.
[0024] FIGS. 2A and 2B are diagrams for explaining a method of operating a neural morphic device according to an embodiment.
[0025] Referring to FIG. 2A, the neural morphic device may include a crossbar array circuit unit. The crossbar array circuit unit includes a plurality of crossbar array circuits, and each crossbar array circuit is also implemented by RCA (resistive crossbar memory arrays). Specifically, each crossbar array circuit may include an input node 210 corresponding to a presynaptic neuron, a neuron circuit 220 corresponding to a postsynaptic neuron, and a synapse circuit 230 providing a connection between the input node 210 and the neuron circuit 220.
[0026] In one embodiment, the crossbar array circuit of the neural morphic device includes four input nodes 210, four neuron circuits 220, and sixteen synapse circuits 230, but the numbers thereof can be variously changed. When the number of input nodes 210 is N (where N is a natural number of 2 or more) and the number of neuron circuits 220 is M (where M is a natural number of 2 or more and may be the same as or different from N), N*M synapse circuits 230 are also arranged in a matrix.
[0027] Specifically, a wiring 21 connected to the input node 210 and extending in a first direction (e.g., the horizontal direction), and a wiring 22 connected to the neuron circuit 220 and extending in a second direction (e.g., the vertical direction) intersecting the first direction may be provided. Hereinafter, for convenience of explanation, the wiring 21 extending in the first direction is referred to as a row line, and the wiring 22 extending in the second direction is referred to as a column line. A plurality of synapse circuits 230 are arranged at each intersection of the row line 21 and the column line 22, and can connect the corresponding row line 21 and the corresponding column line 22 to each other.
[0028] The input node 210 generates a signal, e.g., a signal corresponding to specific data, and sends it to the row line 21. The neuron circuit 220 can perform a role of receiving and processing a synaptic signal that has passed through the synapse circuit 230 via the column line 22. The input node 210 also corresponds to an axon, and the neuron circuit 220 also corresponds to a neuron. However, whether it is a presynaptic neuron or a postsynaptic neuron is also determined by the relative relationship with other neurons. For example, when the input node 210 receives a synaptic signal in the relationship with other neurons, it can function as a postsynaptic neuron. Similarly, when the neuron circuit 220 sends a signal in the relationship with other neurons, it can function as a presynaptic neuron.
[0029] The connection between the input node 210 and the neuron circuit 220 is also made via the synapse circuit 230. Here, the synapse circuit 230 is an element whose electrical conductance or weight value changes due to an electrical pulse, e.g., a voltage or a current, applied to both ends.
[0030] The synaptic circuit 230 may include, for example, a variable resistance element. The variable resistance element is an element that can switch between different resistance states by a voltage or current applied across its two ends, and can have a variety of substances having a plurality of resistance states, for example, a single film structure or a multi-layer film structure including metal oxides such as transition metal oxides and perovskite-based substances; phase change substances such as chalcogenide-based substances; ferroelectric substances; ferromagnetic substances, etc. The operation in which the variable resistance element and / or the synaptic circuit 230 changes from a high resistance state to a low resistance state can be referred to as a set operation, and the operation in which it changes from a low resistance state to a high resistance state can be referred to as a reset operation.
[0031] Regarding the operation of the neural morphic device, referring to FIG. 2B for explanation, it is as follows. For convenience of explanation, the row wiring 21 is referred to as the first row wiring 21A, the second row wiring 21B, the third row wiring 21C, and the fourth row wiring 21D in order from the top, and the column wiring 22 is referred to as the first column wiring 22A, the second column wiring 22B, the third column wiring 22C, and the fourth column wiring 22D in order from the left.
[0032] Referring to FIG. 2B, in the initial state, all of the synaptic circuits 230 are also in a state where the conductivity is relatively low, that is, a high resistance state. If some of the synaptic circuits 230 are in a low resistance state, an initialization operation to set them to a high resistance state may be additionally required. Each of the synaptic circuits 230 can have a predetermined threshold value required for a change in resistance and / or conductivity. More specifically, if a voltage or current having a magnitude smaller than the predetermined threshold value is applied across both ends of each synaptic circuit 230, the conductivity of the synaptic circuit 230 does not change, while if a voltage or current larger than the predetermined threshold value is applied to the synaptic circuit 230, the conductivity of the synaptic circuit 230 changes.
[0033] In this state, in order to perform an operation of outputting specific data as a result of the specific column wiring 22, an input signal corresponding to the output of the input node 210 and corresponding to the specific data also enters the row wiring 21. At this time, the input signal can also be shown as an application of an electrical pulse to each of the row wirings 21. For example, when an input signal corresponding to the data of "0011" enters the row wiring 21, no electrical pulse is applied to the row wirings 21 corresponding to "0", for example, the first row wiring 21A and the second row wiring 21B, and an electrical pulse can be applied only to the row wirings 21 corresponding to "1", for example, the third row wiring 21C and the fourth row wiring 21D. At this time, the column wiring 22 is also driven by an appropriate voltage or current for output.
[0034] As an example, when the column wiring 22 for outputting specific data is predetermined, the column wiring 22 is driven so that a voltage having a magnitude equal to or higher than the voltage (hereinafter referred to as the set voltage) required when the synaptic circuit 230 located at the intersection with the row wiring 21 corresponding to "1" is in the set operation is applied, and the remaining column wirings 22 are also driven so that the remaining synaptic circuits 230 are applied with a voltage having a magnitude smaller than the set voltage. For example, if the magnitude of the set voltage is V set and the column wiring 22 for outputting the data of "0011" is determined to be the third column wiring 22C, the first synaptic circuit 230A and the second synaptic circuit 230B located at the intersections of the third column wiring 22C with the third row wiring 21C and the fourth row wiring 21C are connected to V set or more voltage is applied, the magnitude of the electrical pulse applied to the third row wiring 21C and the fourth row wiring 21D is also V set or more, and the voltage applied to the third column wiring 22C is also 0V. Thereby, the first synaptic circuit 230A and the second synaptic circuit 230B also enter a low-resistance state. The conductivity of the first synaptic circuit 230A and the second synaptic circuit 230B in the low-resistance state gradually increases as the number of electrical pulses increases. The magnitude and width of the applied electrical pulse are also substantially constant. The remaining synaptic circuits 230 except for the first synaptic circuit 230A and the second synaptic circuit 230B are Vset The voltages applied to the remaining column wirings, namely, the first column wiring 22A, the second column wiring 22B, and the fourth column wiring 22D, are such that a smaller voltage is applied, and the voltage is between 0V and V set and a value therebetween, for example, 1 / 2V set can have a value of. Thereby, the resistance states of the remaining synaptic circuits 230 other than the first synaptic circuit 230A and the second synaptic circuit 230B do not change.
[0035] As another example, a column wiring 22 that outputs specific data is not defined. In such a case, while applying an electrical pulse corresponding to the specific data to the row wiring 21, the current flowing through each column wiring 22 is measured, and the column wiring 22 that first reaches a predetermined critical current, for example, the third column wiring 22C, becomes the column wiring 22 that outputs the specific data.
[0036] By the method described above, different data can be output to different column wirings 22 respectively.
[0037] FIG. 3A and FIG. 3B are diagrams for comparing vector-matrix multiplication and operations performed in a crossbar array according to an embodiment.
[0038] First, referring to FIG. 3A, the convolution operation between an input feature map and weight values is also performed using vector-matrix multiplication. For example, the pixel data of the input feature map is also represented by a matrix X 310, and the weight values are also represented by a matrix W 311. The pixel data of the output feature map is also represented by a matrix Y 312 which is the multiplication operation result value of the matrix X 310 and the matrix W 311.
[0039] Referring to FIG. 3B, a vector multiplication operation is performed using the non-volatile memory elements of the crossbar array. Comparing with FIG. 3A, the pixel data of the input feature map is also received as the input value of the non-volatile memory element, and the input value is also the voltage 320. Also, the weighted value is stored in the synapse of the non-volatile memory element, that is, in the memory cell, and the weighted value stored in the memory cell is also the conductance 321. Therefore, the output value of the non-volatile memory element is also represented by the current 322 which is the multiplication operation result value of the voltage 320 and the conductance 321.
[0040] FIG. 4 is a drawing for explaining an example in which a convolution operation is performed in a neural morphic device according to an embodiment. The neural morphic device is provided with pixels of an input feature map 410, and the crossbar array circuit 400 of the neural morphic device is also implemented by RCA.
[0041] The neural morphic device can receive an input feature map in digital signal form and can use a DAC (digital analog converter) 420 to convert the input feature map into a voltage in analog signal form. In one embodiment, the neural morphic device can use the DAC 420 to convert the pixel value of the input feature map into a voltage and then provide the voltage as the input value 401 of the crossbar array circuit 400.
[0042] In addition, the learned weight values are stored in the crossbar array circuit 400 of the neural morphic device. The weight values are also stored in the memory cells of the crossbar array circuit 400, and the weight values stored in the memory cells are also conductances 402. At this time, the neural morphic device can calculate an output value by performing a vector multiplication operation between the input value 401 and the conductance 402, and the output value is also represented by a current 403. That is, the neural morphic device can output the same result value as the convolution operation result of the input feature map and the weight value by using the crossbar array circuit 400.
[0043] Since the current 403 output from the crossbar array circuit 400 is an analog signal, in order to use the current 403 as an input feature map of another crossbar array circuit, the neural morphic device can utilize an ADC (analog digital converter) 430. The neural morphic device can utilize the ADC 430 to convert the current 403, which is an analog signal, into a digital signal. In one embodiment, the neural morphic device can utilize the ADC 430 to convert the current 403 into a digital signal with the same number of bits as the pixels of the input feature map 410. For example, when the input feature map 410 is data with 4-bit pixels, the neural morphic device can utilize the ADC 430 to convert the current 403 into 4-bit data.
[0044] The neural morphic device can apply an activation function to the digital signal converted by the ADC 430 by utilizing an activation unit 440. As the activation function, a sigmoid function, a Tanh function, and a ReLU (rectified linear unit) function can be utilized, but the activation functions applicable to the digital signal are not limited thereto.
[0045] The digital signal to which the activation function is applied is also used as an input feature map of another crossbar array circuit 450. When the digital signal to which the activation function is applied is used as an input feature map of another crossbar array circuit 450, the above-described process is similarly applied to the other crossbar array circuit 450.
[0046] FIG. 5 is a diagram for explaining operations performed by a neural network according to an embodiment.
[0047] Referring to FIG. 5, the neural network 500 has a structure including an input layer, a hidden layer, and an output layer, performs operations based on received input data (e.g., I1 and I2), and can generate output data (e.g., O1 and O2) based on the execution result.
[0048] For example, as shown in FIG. 5, the neural network 500 may include an input layer (Layer 1), two hidden layers (Layer 2 and Layer 3), and an output layer (Layer 4). Since the neural network 500 includes more layers that can process valid information, the neural network 500 can process a more complex data set than a neural network having a single layer. On the other hand, although the neural network 500 is illustrated as including four layers, this is merely an example, and the neural network 500 may include more or fewer layers or may include more or fewer channels. That is, the neural network 500 may include layers having various structures different from those illustrated in FIG. 5.
[0049] Each layer included in the neural network 500 may include a plurality of channels. The channels also correspond to a plurality of artificial nodes known by terms such as neurons, processing elements (PEs), units, or similar terms. For example, as shown in FIG. 5, Layer 1 may include two channels (nodes), and each of Layer 2 and Layer 3 may include three channels. However, this is merely an example, and each layer included in the neural network 500 may include various numbers of channels (nodes).
[0050] The channels included in each layer of the neural network 500 are connected to each other and can process data. For example, one channel can receive and calculate data from other channels and output the calculation result to still other channels.
[0051] The inputs and outputs of each channel are also referred to as input feature maps and output feature maps, respectively. The input feature map includes a plurality of input activations, and the output feature map may include a plurality of output activations. That is, the feature map or the activation is both the output of one channel at a time and also the parameter corresponding to the input of the channels included in the next layer.
[0052] On the other hand, each channel can determine its own activation based on the activations and weight values received from the channels included in the previous layer. The weight value is a parameter used to calculate the output activation in each channel and is also a value assigned to the connection relationship between channels.
[0053] Each channel is also processed by a computational unit or a processing element that receives an input and outputs an output activation, and the input and output of each channel are mapped. For example, σ is an activation function, and w i j,k is the weight value from the k-th channel included in the (i - 1)-th layer to the j-th channel included in the i-th layer, and b i j is the bias of the j-th channel included in the i-th layer, and a i j When a is the activation of the j-th channel of the i-th layer, the activation a i j is also calculated using Equation 1 as follows.
Equation
[0054] As shown in FIG. 5, the activation of the first channel CH1 of the second layer (Layer 2) is a 2 which is also expressed as a 2 1. Also, a 2 1 can have the value of a 2 1,1 1 = σ(w 1 ×a 2 1,2 ×a 1 2 + b 2 1) according to Equation 1. However, the aforementioned Equation 1 is merely an example for explaining the activation and weight values used for processing data in the neural network 500, and is not limited thereto. The activation is also a value obtained by applying batch normalization and an activation function to the sum of the activations received from the previous layer.
[0055] FIGS. 6A to 6C are diagrams for explaining an example of converting an initial weighted value into a binary weighted value according to an embodiment.
[0056] Referring to FIG. 6A, an input layer 601, an output layer 602, and initial weighted values W 11 , W 12 , …, W 32 , W 33 are illustrated. To each of the three neurons in the input layer 601, three input activations I1, I2, and I3 may correspond, and to each of the three neurons in the output layer 602, three output activations O1, O2, and O3 may correspond. Also, the nth input activation I n and the mth output activation O m are applied with the initial weighted value W nm .
[0057] The initial weighted value 610 in FIG. 6B is a matrix representation of the initial weighted values W 11 , W 12 , …, W 32 , W 33 in FIG. 6A.
[0058] In the learning process of the neural network, the initial weighted value 610 may be determined. In one embodiment, the initial weighted value 610 is also represented by 32-bit floating point.
[0059] The initial weighted value 610 is also converted into a binary weighted value 620. The binary weighted value 620 can have a size of 1 bit. In the present invention, in the inference process of the neural network, by using the binary weighted value 620 instead of the initial weighted value 610, the model size and the amount of computation can be reduced. For example, when converting a 32-bit initial weighted value 610 into a 1-bit binary weighted value 620, the model size can be compressed to about 1 / 32.
[0060] In one embodiment, based on the maximum value and the minimum value of the initial weight value 610, the initial weight value 610 is also converted into a binary weight value 620. In other embodiments, based on the maximum value and the minimum value of the initial weight value that can be input to the neural network, the initial weight value 610 is also converted into a binary weight value 620.
[0061] For example, the maximum value of the initial weight value that can be input to the neural network is 1.00, and the minimum value is also -1.00. When the initial weight value is 0.00 or more, it is also converted into a binary weight value 1, and when the initial weight value is less than 0.00, it is also converted into a binary weight value -1.
[0062] Also, the average value 630 of the absolute values of the initial weight value 610 is multiplied by the binary weight value 620. By multiplying the binary weight value 620 by the average value 630 of the absolute values of the initial weight value 610, when using the binary weight value 620, a result value similar to the case of using the initial weight value 610 can be obtained.
[0063] For example, assuming that there are 1024 neurons in the previous layer and 512 neurons in the current layer, each of the 512 neurons belonging to the current layer will have 1024 initial weight values 610. However, for each neuron, after calculating the average value of the absolute values of the 1024 32-bit floating-point initial weight values 610, the calculation result is multiplied by the binary weight value 620.
[0064] Specifically, the average value of the absolute values of the initial weight value 610 used for calculating the predetermined output activations O1, O2, O3 is also multiplied by the initial weight value 610 used for calculating the predetermined output activations O1, O2, O3.
[0065] For example, referring to FIG. 6A, in the process of calculating the first output activation O1, the initial weight values W 11 , W 21 , and W 31 can be used. The initial weight values W 11 , W2, 1 and W31 Each is converted into binary weighted values W 11 ’, W 21 ’, and W 31 ’, and the binary weighted values W 11 ’, W 21 ’, and W 31 ’ are multiplied by the average value of the absolute values of the initial weighted values W 11 , W 21 , and W 31 .
Number
[0066] In the same manner, the binary weighted values W 12 ’, W 22 ’, and W 32 ’ are multiplied by the average value of the absolute values of the initial weighted values W 12 , W 22 , and W 32 . Also, the binary weighted values W
Number
Number
[0067] Referring to FIG. 6C, the initial weighted value 610, the binary weighted value 620, and the average value 630 of the absolute values of the initial weighted value 610 are shown as specific numerical values. In FIG. 6C, for convenience of explanation, the initial weighted value 610 is expressed in decimal, but it is assumed that the initial weighted value 610 is a 32-bit floating point number.
[0068] In FIG. 6C, it is illustrated that when the initial weighting value is 0.00 or more, it is converted to the binary weighting value 1, and when the initial weighting value is less than 0.00, it is converted to the binary weighting value -1.
[0069] Also, the average value of the absolute values of the initial weighting values W 11 、W 21 、and W 31 is “0.28”, and the average value of the absolute values of the initial weighting values W 12 、W 22 、and W 32 is “0.37”, and it is illustrated that the average value of the absolute values of the initial weighting values W 13 、W 23 、and W 33 is “0.29”.
[0070] FIGS. 7A and 7B are diagrams for explaining an example of converting an input feature map into a temporal domain binary vector according to an embodiment.
[0071] The input feature map is also converted into a plurality of temporal domain binary vectors. The input feature map includes a plurality of input activations, and each of the plurality of input activations is also converted into a temporal domain binary vector.
[0072] The input feature map is also converted into a plurality of temporal domain binary vectors based on the quantization level. In one embodiment, the range between the maximum value and the minimum value of the input activations that can be input to the neural network is also divided into N (N is a natural number) quantization levels. For example, for the quantization level division, a sigmoid function or a tanh function or the like can be used, but it is not limited thereto.
[0073] For example, referring to FIG. 7A, when the quantization level is set to 9 and the maximum and minimum values of the input activation that can be input to the neural network are 1.0 and -1.0 respectively, the quantization level is also divided into "1.0, 0.75, 0.5, 0.25, 0, -0.25, -0.5, -0.75, -1.0".
[0074] On the other hand, in FIG. 7A, although it is shown that the intervals between the quantization levels are set to be the same, the intervals between the quantization levels can also be set non-linearly.
[0075] When the quantization level is set to N, the time-domain binary vector can have (N - 1) elements. For example, referring to FIG. 7A, when the quantization level is set to 9, the time-domain binary vector can have 8 elements t1, t2, …, t7, t8.
[0076] Based on which quantization level among the N quantization levels the input activation belongs to, the input activation is also converted into a time-domain binary vector. For example, when a predetermined input activation has a value of 0.75 or more, the predetermined input activation is also converted into the time-domain binary vector (+1, +1, +1, +1, +1, +1, +1, +1). Also, for example, when a predetermined input activation has a value less than -0.25 and greater than or equal to -0.5, the predetermined input activation is also converted into the time-domain binary vector (+1, +1, +1, -1, -1, -1, -1, -1).
[0077] Referring to FIG. 7B, an example is shown in which each of a plurality of input activations included in the input feature map 710 is converted into a time binary vector. Since the first activation has a value less than 0 and greater than or equal to -0.25, it is also converted into a time domain binary vector (-1, -1, -1, -1, +1, +1, +1, +1), and since the second activation has a value less than 0.5 and greater than or equal to 0.25, it is also converted into a time domain binary vector (-1, -1, +1, +1, +1, +1, +1, +1). Also, since the third activation has a value less than -0.75 and greater than or equal to -1.0, it is also converted into a time domain binary vector (-1, -1, -1, -1, -1, -1, -1, +1), and since the fourth activation has a value greater than or equal to 0.75, it is also converted into a time domain binary vector (+1, +1, +1, +1, +1, +1, +1, +1).
[0078] On the other hand, in the conventional method, when each input activation of each layer of the neural network is converted into a binary value, the information that the input activation has is lost, and information transmission between the layers cannot be performed normally.
[0079] On the other hand, as in the present disclosure, when each input activation of each layer of the neural network is converted into a time domain binary vector, it is possible to approximate the original input activation based on a plurality of binary values.
[0080] FIG. 8 is a diagram for explaining applying a binary weighted value and a time domain binary vector to a batch normalization process according to an embodiment.
[0081] Generally, in a neural network algorithm model, after performing the product of the input activation (the first input value or the output value of the previous layer) and the initial weight value (32-bit floating point), and the sum of the results (MAC: multiply and accumulate), a separate bias value is further added for each neuron. Then, batch normalization is performed for each neuron on the resulting value. After inputting the result into the activation function, the activation function output value is transmitted as the input value of the next layer.
[0082] It can be expressed as in Equation 2. In Equation 2, I n represents the input activation, W nm represents the initial weight value, B m represents the bias value, α m represents the initial scale value of batch normalization, β m represents the bias value of batch normalization, f represents the activation function, O m represents the output activation.
Equation
[0083] Referring to FIG. 8, the input activation I n 810 is also converted into the time-domain binary vector I b n (t)820. The temporal binary vector generator can convert the input activation I n 810 into the time-domain binary vector I b n (t)820.
[0084] As described in FIGS. 7A and 7B, according to the existing quantization level, the input activation I n 810 is also converted into the time-domain binary vector I b n (t)820. On the other hand, the time-domain binary vector Ib n (t)820 The number of elements included in each of them is also determined by the number of quantization levels. For example, when the number of quantization levels is N, the time-domain binary vector I b n (t)820 The number of elements included in each of them is also (N - 1).
[0085] On the other hand, when the input activation I n 810 is converted into the time-domain binary vector I b n (t)820, the operation result using the time-domain binary vector I b n (t)820 is amplified by the number T 850 of elements included in the time-domain binary vector I b n (t)820. Accordingly, when using the time-domain binary vector I b n (t)820, by dividing the operation result by the number T 850 of elements, a result identical to the original MAC operation result can be obtained. The specific explanation thereof is performed using Equations 5 and 6.
[0086] As described in FIGS. 6A - 6C, the initial weight value W nm is also converted into the binary weight value W b nm 830. For example, through the sign function, the initial weight value W nm is also converted into the binary weight value W b nm 830.
[0087] The time-domain binary vector I b n (t)820 and the binary weight value W b nm 830 are subjected to a convolution operation. In one embodiment, the time-domain binary vector I b n (t)820 and the binary weight value W bnm An XNOR operation and an addition operation with 830 are performed.
[0088] Time-domain binary vector I b n (t) 820 and binary weight value W b nm After performing the XNOR operation with 830 and combining all the results, the original multi-bit input activation I n 810 and binary initial weight value W nm will show the same increasing and decreasing pattern as the result of the convolution operation performed with them.
[0089] Time-domain binary vector I b n (t) 820 and binary weight value W b nm The operation with 830 can be expressed as shown in Equation 3 below.
Equation
Equation
[0090] Also, intermediate activation X m 840 is also divided by the number T 850 of elements included in each of the time-domain binary vectors I b n (t) 820. The number of elements included in each of the time-domain binary vectors I b n (t) 820 is also determined by the number of quantization levels. Time-domain binary vector I bn (t)820 is amplified by the number of elements T 850 of the intermediate activation X, so the intermediate activation X m By dividing 840 by the number of elements T 850, the same result as the original MAC operation result can be obtained.
[0091] Intermediate activation X m 840 is multiplied by the average value S m 860 of the absolute value of the initial weight value, and the time-domain binary vector I b n By dividing each by the number of elements T 860 included in (t)820, the output activation O m 870 can be calculated. The output activation O m 870 is also shown as in Equation 5 below.
Equation
[0092] For the neural network algorithm model according to Equation 2, with the binary weight value W b nm 830, the time-domain binary vector I bn (t)820 and the modified scale value α” m When applied, Equation 2 is also expressed as Equation 6 below.
Equation
[0093] In the present disclosure, the binary weight value W b nm Multiply 830 by the average value S of the absolute values of the initial weight values m By multiplying 860, the binary weight value W b nm Even when using 830, the same result value as when using the initial weight value W nm can be obtained. On the other hand, the average value S of the absolute values of the initial weight values m 860 can be included in the batch normalization operation as in the above formula 6 (M m Xα m ), so that no additional model parameters are generated, there is no loss in reducing the model size, and there is also no loss in reducing the amount of computation. That is, it can be confirmed that in formula 6, compared with formula 2, the operation can be performed without adding separate parameters and separate procedures.
[0094] In the present disclosure, the multi-bit input activation I n 810 is quantized to a low number of bits of 2 bits to 3 bits and then converted into a time-domain binary vector I b n (t)820 having a plurality of elements. Also, in the present disclosure, the binary weight value W b nm 830 and the time-domain binary vector I b n (t)820, by performing a time-axis XNOR operation, it is possible to ensure learning performance and final classification / recognition accuracy at the same level as a 32-bit floating-point neural network based on a binary MAC operation.
[0095] On the other hand, the number of elements T 860 can be included in the batch normalization operation as in the above formula 6 (α mWithout further generation of model parameters, there is no loss in reducing the model size, nor is there any loss in reducing the amount of computation. That is, it can be confirmed that in Equation 6, operations can be performed without adding separate parameters and separate procedures when compared with Equation 2.
[0096] For example, for a 32-bit input activation I n 810 is converted to a time-domain binary vector I b n (t)820, the model size can be compressed by about T / 32.
[0097] FIG. 9 is a block diagram of a neural network device using a von Neumann structure according to an embodiment.
[0098] Referring to FIG. 9, the neural network device 900 may include an external input receiving unit 910, a memory 920, a time-domain binary vector generation unit 930, a convolution operation unit 940, and a neural operation unit 950.
[0099] Only the components related to this embodiment are illustrated in the neural network device 900 shown in FIG. 9. Therefore, it will be apparent to those skilled in the art that the neural network device 900 may further include other general-purpose components in addition to the components illustrated in FIG. 9.
[0100] The external input receiving unit 910 can receive neural network model-related information, input image (or audio) data, etc. from the outside. Various information and data received from the external input receiving unit 910 are also stored in the memory 920.
[0101] In one embodiment, the memory 920 can be divided into a first memory for storing input feature maps and a second memory for storing binary weight values, other real-valued parameters, model structure definition variables, and the like. On the other hand, the binary weight values stored in the memory 920 are also values obtained by converting the initial weight values (e.g., 32-bit floating-point numbers) after the learning of the neural network is completed.
[0102] The time-domain binary vector generation unit 930 can receive an input feature map from the memory 920. The time-domain binary vector generation unit 930 can convert the input feature map into a time-domain binary vector. The input feature map includes a plurality of input activations, and the time-domain binary vector generation unit 930 can convert each of the plurality of input activations into a time-domain binary vector.
[0103] Specifically, the time-domain binary vector generation unit 930 can convert the input feature map into a plurality of time-domain binary vectors based on the quantization level. In one embodiment, when the range between the maximum value and the minimum value of the input activations that can be input to the neural network is divided into N (N is a natural number) quantization levels, the time-domain binary vector generation unit 930 can convert the input activation into a time-domain binary vector having (N - 1) elements.
[0104] The convolution operation unit 940 can receive binary weight values from the memory 920. Also, the convolution operation unit 940 can receive a plurality of time-domain binary vectors from the time-domain binary vector generation unit 930.
[0105] The convolution operation unit 940 includes an adder, and the convolution operation unit 940 can perform a convolution operation between the binary weight values and the plurality of time-domain binary vectors.
[0106] The neural operation unit 950 can receive the convolution operation result of the binary weighted value and a plurality of time-domain binary vectors from the convolution operation unit 940. Further, the neural operation unit 950 can receive the correction scale value of batch normalization, the bias value of batch normalization, and the activation function, etc. from the memory 920.
[0107] In the neural operation unit 950, batch normalization and pooling can be performed, and the activation function can be applied, but the operations that can be performed and applied in the neural operation unit 950 are not limited to those.
[0108] On the other hand, the correction scale value of batch normalization can be calculated by multiplying the initial scale value by the average value of the absolute values of the initial weighted values and dividing by the number T of elements included in the time-domain binary vector.
[0109] By performing batch normalization and applying the activation function in the neural operation unit 950, an output feature map can be output. The output feature map may include a plurality of output activations.
[0110] FIG. 10 is a block diagram of a neural morphic device using an in-memory structure according to an embodiment.
[0111] Referring to FIG. 10, the neural morphic device 1000 may include an external input receiving unit 1010, a memory 1020, a time-domain binary vector generation unit 1030, an on-chip memory 1040, and a neural operation unit 1050.
[0112] Only the components related to this embodiment are shown in the neural morphic device 1000 shown in FIG. 10. Therefore, it will be apparent to those skilled in the art that the neural morphic device 1000 may further include other general-purpose components in addition to the components shown in FIG. 10.
[0113] The external input receiving unit 1010 can receive neural network model related information, input image (or audio) data, etc. from the outside. Various information and data received from the external input receiving unit 1010 are also stored in the memory 1020. The memory 1020 can store input feature maps, other real-valued parameters, model structure definition variables, etc. Different from the neural network device 900 in FIG. 9, the binary weight values are stored not in the memory 1020 but in the on-chip memory 1040, and the detailed content will be described later.
[0114] The time-domain binary vector generation unit 1030 can receive an input feature map from the memory 1020. The time-domain binary vector generation unit 1030 can convert the input feature map into a time-domain binary vector. The input feature map includes a plurality of input activations, and the time-domain binary vector generation unit 1030 can convert each of the plurality of input activations into a time-domain binary vector.
[0115] Specifically, the time-domain binary vector generation unit 1030 can convert the input feature map into a plurality of time-domain binary vectors based on the quantization level. In one embodiment, when the range between the maximum value and the minimum value of the input activations that can be input to the neural network is divided into N (N is a natural number) quantization levels, the time-domain binary vector generation unit 1030 can convert the input activation into a time-domain binary vector having (N - 1) elements.
[0116] The on-chip memory 1040 may include an input unit 1041, a crossbar array circuit 1042, and an output unit 1043.
[0117] The crossbar array circuit 1042 may include a plurality of synapse circuits (e.g., variable resistors). Binary weight values are also stored in the plurality of synapse circuits. The binary weight values stored in the plurality of synapse circuits are also values obtained by converting initial weight values (e.g., 32-bit floating-point numbers) after the learning of the neural network is completed.
[0118] The input unit 1041 can receive a plurality of time-domain binary vectors from the time-domain binary vector generation unit 1030.
[0119] When a plurality of time-domain binary vectors are received by the input unit 1041, the crossbar array circuit 1042 can perform a convolution operation between the binary weight values and the plurality of time-domain binary vectors.
[0120] The output unit 1043 can transmit the convolution operation result to the neural operation unit 1050.
[0121] The neural operation unit 1050 can receive the convolution operation result between the binary weight values and the plurality of time-domain binary vectors from the output unit 1043. Also, the neural operation unit 1050 can receive a batch normalization correction scale value, a batch normalization bias value, an activation function, etc. from the memory 1020.
[0122] In the neural operation unit 1050, batch normalization and pooling can be performed, and an activation function can be applied, but the operations that can be performed and applied in the neural operation unit 1050 are not limited to these.
[0123] On the one hand, the modified scale value of batch normalization is calculated by multiplying the initial scale value by the average value of the absolute values of the initial weight values and then dividing by the number T of elements included in the time-domain binary vector.
[0124] Batch normalization is performed in the neural operation unit 1050, and by applying an activation function, an output feature map can be output. The output feature map may include a plurality of output activations.
[0125] FIG. 11 is a flowchart for explaining a method of implementing a neural network in a neural network device according to an embodiment.
[0126] Referring to FIG. 11, in step 1110, the neural network device can obtain binary weight values and an input feature map from a memory.
[0127] In step 1120, the neural network device can convert the input feature map into a time-domain binary vector.
[0128] In one embodiment, the neural network device can convert the input feature map into a time-domain binary vector based on the quantization level related to the input feature map.
[0129] Specifically, the neural network device divides the range between the maximum value and the minimum value that can be input to the neural network into N (N is a natural number) quantization levels, and based on which quantization level among the N quantization levels each activation of the input feature map belongs to, each activation can be converted into a time-domain binary vector.
[0130] On the one hand, the neural network device can divide the range of the maximum value and the minimum value that can be input into the neural network into linear quantization levels or non-linear quantization levels.
[0131] In step 1130, the neural network device can output an output feature map by performing a convolution operation on the binary weight value and the time-domain binary vector.
[0132] The neural network device can output an output feature map by performing batch normalization on the convolution operation result.
[0133] In one embodiment, the neural network device can calculate a corrected scale value by multiplying the initial scale value of batch normalization by the average value of the absolute values of the initial weight values and dividing by the number of elements included in each time-domain binary vector. The neural network device can perform batch normalization based on the corrected scale value.
[0134] The neural network device can perform a multiplication operation of multiplying each bias value applied to the neural network by the initial scale value, and reflect the multiplication operation result in the output feature map.
[0135] Also, the neural network device can perform batch normalization on the convolution operation result, and output an output feature map by applying an activation function to the execution result of the batch normalization.
[0136] FIG. 12 is a flowchart for explaining a method of implementing a neural network in a neural morphic device according to one embodiment.
[0137] Referring to FIG. 12, at stage 1210, the neural morphing device can store binary weight values in the synaptic circuits included in the crossbar array circuit.
[0138] At stage 1220, the neural morphing device can obtain an input feature map from the memory.
[0139] At stage 1230, the neural morphing device can convert the input feature map into a time-domain binary vector.
[0140] In one embodiment, the neural morphing device can convert the input feature map into a time-domain binary vector based on the quantization level related to the input feature map.
[0141] Specifically, the neural morphing device divides the range between the maximum value and the minimum value that can be input to the neural network into N (N is a natural number) quantization levels, and based on which quantization level among the N quantization levels each activation of the input feature map belongs to, each activation can be converted into a time-domain binary vector.
[0142] On the other hand, the neural morphing device can divide the range between the maximum value and the minimum value that can be input to the neural network into linear quantization levels or non-linear quantization levels.
[0143] At stage 1240, the neural morphing device can provide the time-domain binary vector as an input value to the crossbar array circuit.
[0144] At stage 1250, the neural morphing device can output an output feature map by performing a convolution operation on the binary weight value and the time-domain binary vector.
[0145] The neural morphic device can output an output feature map by performing batch normalization on the convolution operation result.
[0146] In one embodiment, the neural morphic device can calculate a corrected scale value by multiplying the initial scale value of batch normalization by the average value of the absolute values of the initial weights and dividing by the number of elements included in each of the time-domain binary vectors. The neural morphic device can perform batch normalization based on the corrected scale value.
[0147] The neural morphic device can perform a multiplication operation of multiplying each bias value applied to the neural network by the initial scale value, and reflect the multiplication operation result in the output feature map.
[0148] Also, the neural morphic device can perform batch normalization on the convolution operation result, and output an output feature map by applying an activation function to the batch normalization execution result.
[0149] FIG. 13 is a block diagram illustrating the hardware configuration of a neural network device according to one embodiment.
[0150] The neural network device 1300 can also be implemented by various devices such as a PC (personal computer), a server device, a mobile device, and an embedded device. As a specific example, it may apply to, but is not limited to, smartphones, tablet devices, AR (augmented reality) devices, IoT (internet of things), autonomous driving automobiles, robotics, medical devices, etc. that perform speech recognition, video recognition, video classification, etc. using neural networks. Furthermore, the neural network device 1300 also corresponds to a dedicated hardware accelerator (HW accelerator) installed in the devices as described above. The neural network device 1300 is also a hardware accelerator such as an NPU (neural processing unit), a TPU (Tensor processing unit), or a Neural Engine, which is a dedicated module for driving the neural network, but is not limited thereto.
[0151] Referring to FIG. 13, the neural network device 1300 includes a processor 1310 and a memory 1320. Only the components related to this embodiment are illustrated in the neural network device 1300 shown in FIG. 9. Therefore, it will be apparent to those skilled in the art that the neural network device 1300 may further include other general-purpose components in addition to the components illustrated in FIG. 13.
[0152] Processor 1310 controls the overall functions for executing neural network device 1300. For example, processor 1310 controls neural network device 1300 generally by executing a program stored in memory 1320 within neural network device 1300. Processor 1310 is implemented by, but not limited to, a CPU (central processing unit), a GPU (graphics processing unit), an AP (application processor), etc. provided within neural network device 1300.
[0153] Memory 1320 is hardware that stores various data processed within neural network device 1300. For example, memory 1320 can store data processed by neural network device 1300 and data to be processed. In addition, memory 1320 can store applications, drivers, etc. driven by neural network device 1300. Memory 1320 may include RAM (random access memory) such as DRAM (dynamic random access memory) and SRAM (static random access memory), ROM (read only memory), EEPROM (electrically erasable programmable read only memory), CD-ROM (compact disc read only memory), Blu-ray (registered trademark), or other optical disk storage, HDD (hard disk drive), SSD (solid static driver), or flash memory.
[0154] Processor 1310 reads from and writes to memory 1320 neural network data, such as image data, feature map data, weight value data, etc., and utilizes the read / written data to execute a neural network. When the neural network is executed, processor 1310 repeatedly performs a convolution operation on an input feature map and weight values in order to generate data related to an output feature map. At that time, depending on various factors such as the number of channels of the input feature map, the number of channels of the weight values, the size of the input feature map, the size of the weight values, and the precision of the values, the amount of computation of the convolution operation is determined.
[0155] The actual neural network driven by neural network device 1300 is also implemented by a more complex architecture. Thereby, processor 1310 will perform a very large number of operations with an operation count ranging from hundreds of millions to tens of billions, and the frequency at which processor 1310 accesses memory 1320 for operations will both increase dramatically. Due to such an operation amount burden, in mobile devices such as smartphones, tablets, and wearable devices with relatively low processing performance, and embedded devices, etc., the processing of the neural network becomes not smooth.
[0156] Processor 1310 can perform operations such as convolution operations, batch normalization operations, pooling operations, activation function operations, etc. In one embodiment, processor 1310 can perform matrix multiplication operations, transformation operations, and transposition operations in order to obtain multi-head self-attention. In the process of obtaining multi-head self-attention, the transformation operation and the transposition operation are also performed after or before the matrix multiplication operation.
[0157] Processor 1310 can obtain the binary weighted values and the input feature map from memory 920, and can convert the input feature map into a time-domain binary vector. Further, processor 1310 can output an output feature map by performing a convolution operation between the binary weighted values and the time-domain binary vector.
[0158] FIG. 14 is a block diagram illustrating the hardware configuration of a neural morphic device according to an embodiment.
[0159] Referring to FIG. 14, neural morphic device 1400 may include processor 1410 and on-chip memory 1420. Only the components related to this embodiment are illustrated in neural morphic device 1400 shown in FIG. 14. Therefore, it will be apparent to those skilled in the art that neural morphic device 1400 may further include other general-purpose components in addition to the components shown in FIG. 14.
[0160] Neural morphic device 1400 is also mounted on digital systems that require low-power neural network driving, such as smartphones, drones, tablet devices, AR devices, IoT devices, autonomous driving automobiles, robotics, and medical devices, but is not limited thereto.
[0161] Neural morphic device 1400 includes a plurality of on-chip memories 1420, and each on-chip memory 1420 is also constituted by a plurality of crossbar array circuits. The crossbar array circuit may include a plurality of presynaptic neurons, a plurality of postsynaptic neurons, and a synapse circuit that provides a connection between each of the plurality of presynaptic neurons and the plurality of postsynaptic neurons, that is, a memory cell. In one embodiment, the crossbar array circuit is also implemented by RCA.
[0162] The external memory 1430 is hardware that stores various data processed by the neural morphic device 1400, and the external memory 1430 can store the data processed by and to be processed by the neural morphic device 1400. Further, the external memory 1430 can store applications, drivers, etc. driven by the neural morphic device 1400. The external memory 1430 may include RAM such as DRAM and SRAM, ROM, EEPROM, CD-ROM, Blu-ray, or other optical disk storage, HDD, SSD, or flash memory.
[0163] The processor 1410 plays a role in controlling the overall functions for driving the neural morphic device 1400. For example, the processor 1410 generally controls the neural morphic device 1400 by executing a program stored in the on-chip memory 1420 within the neural morphic device 1400. The processor 1410 is also embodied by, but not limited to, a CPU, GPU, AP, etc. provided within the neural morphic device 1400. The processor 1410 reads / writes various data from / to the external memory 1430, utilizes the read / written data, and executes the neural morphic device 1400.
[0164] The processor 1410 can generate a plurality of binary feature maps by binarizing the pixel values of the input feature map based on a plurality of threshold values. The processor 1410 can provide the pixel values of the plurality of binary feature maps as input values to the crossbar array circuit unit. The processor 1410 can use a DAC to convert the pixel values into analog signals (voltages). Processor 1410 can store the weighted values applied to the crossbar array circuit unit in the synaptic circuits included in the crossbar array circuit unit. The weighted values stored in the synaptic circuits are also conductances. Further, Processor 1410 can calculate the output value of the crossbar array circuit unit by performing a multiplication operation between the input value and the kernel value stored in the synaptic circuit.
[0165] Processor 1410 can generate the pixel values of the output feature map by merging the output values calculated by the crossbar array circuit unit. On the other hand, since the output values calculated by the crossbar array circuit unit (or the resulting values obtained by multiplying the calculated output values by the weighted values) are in the form of analog signals (currents), Processor 1410 can use an ADC to convert the output values into digital signals. Further, Processor 1410 can apply an activation function to the output values converted into digital signals in the ADC.
[0166] Processor 1410 can store binary weighted values in the synaptic circuits included in the crossbar array circuit and acquire the input feature map from the external memory 1430. Further, Processor 1410 can convert the input feature map into a time-domain binary vector and provide the time-domain binary vector as the input value of the crossbar array circuit. Further, Processor 1410 can output the output feature map by performing a convolution operation between the binary weighted values and the time-domain binary vector.
[0167] This embodiment is also embodied in the form of a recording medium containing computer-executable instructions such as program modules executed by a computer. A computer-readable medium is any available medium that can be accessed by a computer and includes both volatile and non-volatile media, as well as removable and non-removable media. The computer-readable medium may also include either computer storage media and communication media. The computer storage media includes volatile and non-volatile, removable and non-removable media embodied by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. The communication media typically includes modulated data signals such as computer-readable instructions, data structures, program modules, or other data, or other transmission mechanisms, and includes any information transmission media.
[0168] Also, in this specification, a "section" is also a hardware component such as a processor or a circuit, and / or a software component executed by a hardware component such as a processor.
[0169] The above description of this specification is for illustrative purposes, and those skilled in the technical field to which the content of this specification belongs will be able to understand that it can be easily deformed into other specific forms without changing the technical idea and essential features of the present invention. Therefore, the above embodiments should be understood to be illustrative in all respects and not restrictive. For example, each component described as a single type is implemented in a distributed manner, and similarly, components described as distributed are also implemented in a combined form.
[0170] The scope of this embodiment is indicated by the claims rather than the above detailed description, and all changes and modifications derived from the meaning, scope, and equivalent concept of the claims should be construed as being included.
Description of Symbols
[0171] 10 Biological neuron 11 Mathematical model of biological neuron 210 Input node 220 Neural circuit 400,1042 Crossbar array circuit 410 Input feature map 420 DAC 430 ADC 440 Activation unit 500 Neural network 601 Input layer 602 Output layer 810 Input activation 820 Time-domain binary vector 830 Binary weighted value 840 Intermediate activation 850 Number T 860 Average value S m 870 Output activation O M 900,1300 Neural network device 910 External input device 920,1020,1320 Memory 930,1030 Time-domain binary vector generation unit 940 Convolution operation unit 950,1050 Neural operation unit 1000,1400 Neuromorphic device 1010 Input reception unit 1040,1420 On-chip memory 1041 Input section 1043 Output section 1310,1410 Processor 1430 External memory
Claims
1. A neural morphic device that implements a neural network, comprising: a memory storing at least one program; an on-chip memory including a crossbar array circuit; at least one processor that drives the neural network by executing the at least one program. The at least one processor: stores a binary weighted value converted from the initial weight value in a synapse circuit included in the crossbar array circuit based on a maximum value and a minimum value of the initial weight value; acquires an input feature map from the memory; converts each activation of the input feature map into a time-domain binary vector expressed as a sequence of elements including at least one of a positive element and a negative element based on which quantization level among N quantization levels each activation belongs to (N is a natural number); provides the time-domain binary vector as an input value to the crossbar array circuit; outputs an output feature map by performing a convolution operation between the binary weighted value and the time-domain binary vector. A neural morphic device.
2. The at least one processor: outputs an output feature map by performing batch normalization on the convolution operation result. The neural morphic device according to claim 1.
3. The at least one processor: calculates a corrected scale value by multiplying an initial scale value of the batch normalization by an average value of absolute values of the initial weight values and dividing by the number of elements included in each of the time-domain binary vectors; performs the batch normalization based on the corrected scale value. The neural morphic device according to claim 2.
4. The at least one processor: divides a range between a maximum value and a minimum value that can be input to the neural network into N quantization levels (N is a natural number). The neural morphic device according to claim 1.
5. The at least one processor: divides a range between a maximum value and a minimum value that can be input to the neural network into non-linear quantization levels. The neural morphic device according to claim 4.
6. The at least one processor performs a multiplication operation of multiplying each bias value applied to the neural network by the initial scale value, and reflects the multiplication operation result in the output feature map. The neural morphic device according to claim 3.
7. The at least one processor performs the batch normalization on the convolution operation result, and outputs an output feature map by applying an activation function to the execution result of the batch normalization. The neural morphic device according to claim 3.
8. A neural network device for implementing a neural network, comprising: a memory storing at least one program; and at least one processor that drives a neural network by executing the at least one program. The at least one processor acquires, from the memory, a binary weighted value and an input feature map converted from the initial weight value based on the maximum value and the minimum value of the initial weight value, converts each activation of the input feature map into a time-domain binary vector expressed as a sequence of elements including at least one of a positive element and a negative element based on which quantization level among N quantization levels each activation belongs to (N is a natural number), and outputs an output feature map by performing a convolution operation between the binary weighted value and the time-domain binary vector. Neural network device.
9. The at least one processor outputs an output feature map by performing batch normalization on the convolution operation result. The neural network device according to claim 8.
10. The at least one processor calculates a corrected scale value by multiplying the initial scale value of the batch normalization by the average value of the absolute values of the initial weight values and dividing by the number of elements included in each of the time-domain binary vectors, and performs the batch normalization based on the corrected scale value. The neural network device according to claim 9.
11. The at least one processor Divide the range between the maximum value and the minimum value that can be input to the neural network into N quantization levels (N is a natural number). The neural network device according to claim 8.
12. The at least one processor Divides the range between the maximum value and the minimum value that can be input to the neural network into non-linear quantization levels. The neural network device according to claim 11.
13. The at least one processor Performs a multiplication operation of multiplying each bias value applied to the neural network by the initial scale value, Reflects the multiplication operation result in the output feature map. The neural network device according to claim 10.
14. The at least one processor Performs the batch normalization on the convolution operation result, and outputs an output feature map by applying an activation function to the result of performing the batch normalization. The neural network device according to claim 10.
15. In a neural morphic device, a method for implementing a neural network, comprising: Under the control of a processor provided in the neural morphic device, based on the maximum value and the minimum value of the initial weight value, storing a binary weighted value converted from the initial weight value in a synaptic circuit included in a crossbar array circuit; The processor obtains an input feature map from a memory; Under the control of the processor, based on which quantization level among N quantization levels each activation of the input feature map belongs to, converting each activation into a time-domain binary vector expressed as a sequence of elements including at least one of a positive element and a negative element (N is a natural number); Under the control of the processor, providing the time-domain binary vector as an input value of the crossbar array circuit; Under the control of the processor, performing a convolution operation between the binary weighted value and the time-domain binary vector to output an output feature map; A method including the above steps.
16. In a neural network device, a method for implementing a neural network, comprising: Under the control of a processor provided in the neural network device, obtaining from a memory a binary weighted value and an input feature map converted from the initial weight value based on the maximum value and the minimum value of the initial weight value; Under the control of the processor, converting each activation of the input feature map into a time-domain binary vector expressed as a sequence of elements including at least one of a positive element and a negative element based on which quantization level among N quantization levels each activation belongs to (N is a natural number); Under the control of the processor, outputting an output feature map by performing a convolution operation between the binary weighted value and the time-domain binary vector; A method comprising the above.
17. A computer-readable recording medium having a program recorded thereon, wherein when the program is executed by a processor of the computer, the method according to claim 15 is caused to be implemented on the computer; A computer-readable recording medium.
18. A computer-readable recording medium having a program recorded thereon, wherein when the program is executed by a processor of the computer, the method according to claim 16 is caused to be implemented on the computer; A computer-readable recording medium.
Citation Information
Patent Citations
Neural network circuit device, neural network, neural network processing method and neural network executing program
JP2018092377A
Method and device for optimizing and applying multi-layer neural network model, and storage medium
JP2019160319A
Resistive processing unit array, method for forming a resistive processing unit array and method for hysteretic operation - Patents.com
JP2020514886A
Quantized neural network training and inference
US20170286830A1
Binary, ternary and bit serial compute-in-memory circuits
US20190102359A1