Neuromorphic apparatus and methods for implementing neural networks
By using memory and cross-switch matrix array circuitry in a neuromorphic device, the problem of inefficiently processing large amounts of input data in existing technologies is solved, enabling efficient neural network computation and data analysis.
Patent Information
- Application Number
- CN202011409845.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-08
- Filing Date
- 2020-12-04
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2040-12-04
AI Technical Summary
Existing technologies struggle to efficiently process large amounts of input data and extract desired information, necessitating a more efficient computing technique for real-time analysis of neural networks.
A neuromorphic device, including a memory and a cross-switch matrix array circuit, is used to transform the input feature map into a temporal binary vector by storing binary weight values in synaptic circuits, and to perform convolution calculations to output the feature map.
It achieves efficient computational processing, improves the data processing capabilities of neural networks, and can handle complex datasets while maintaining high accuracy.
Smart Images

Figure CN113837371B_ABST
Abstract
Description
[0001] This application claims the benefit of Korean Patent Application No. 10-2020-0069100, filed June 8, 2020, the disclosure of which is incorporated herein in its entirety by reference. TECHNICAL FIELD
[0002] The disclosure relates to a device for implementing a neural network and an operating method of the device. BACKGROUND
[0003] A memory-oriented neural network device represents a computing architecture that models a biological brain. As the memory-oriented neural network technology has developed, research into analyzing input data and extracting effective information by using a memory-oriented neural network in various types of electronic systems has been actively conducted.
[0004] Accordingly, in order to analyze a large amount of input data in real time and extract desired information by using a memory-centered neural network, a technology for efficiently processing computation is required. SUMMARY
[0005] The disclosure provides a device and a method for generating a chemical structure by using a neural network. The disclosure also provides a computer-readable recording medium having recorded thereon a program for executing the method on a computer. The technical objects to be achieved by one or more embodiments are not limited to what has been described above and other technical objects can be inferred from the following embodiments.
[0006] Additional aspects will be set forth in part in the description which follows, and in part will be apparent from the description, or can be learned by practice of the presented disclosure.
[0007] According to an aspect of an embodiment, a neuromorphic device for implementing a neural network includes a memory in which at least one program is stored, an on-chip memory including a crossbar matrix array circuit, and at least one processor configured to drive a neural network by executing the at least one program, wherein the at least one processor is further configured to store binary weight values in synapse circuits included in the crossbar matrix array circuit, obtain an input feature map from the memory, convert the input feature map into a time-domain binary vector, provide the time-domain binary vector as an input value of the crossbar matrix array circuit, and output an output feature map by performing convolution calculation between the binary weight values and the time-domain binary vector.
[0008] According to another aspect of an embodiment, a neural network apparatus for implementing a neural network includes a memory in which at least one program is stored, and at least one processor configured to drive a neural network by executing the at least one program, wherein the at least one processor is further configured to obtain binary weight values and an input feature map from the memory, convert the input feature map into a time-domain binary vector, and output an output feature map by performing a convolution calculation between the binary weight values and the time-domain binary vector.
[0009] According to another aspect of an embodiment, a method of implementing a neural network in a neuromorphic apparatus includes storing binary weight values in synapse circuits included in a crossbar matrix array circuit in the neuromorphic apparatus, obtaining an input feature map from a memory in the neuromorphic apparatus, converting the input feature map into a time-domain binary vector, providing the time-domain binary vector as an input value to the crossbar matrix array circuit, and outputting an output feature map by performing a convolution calculation between the binary weight values and the time-domain binary vector.
[0010] According to another aspect of an embodiment, a method of implementing a neural network in a neuromorphic apparatus includes storing binary weight values in synapse circuits included in a crossbar matrix array circuit in the neuromorphic apparatus, obtaining an input feature map from a memory in the neuromorphic apparatus, converting the input feature map into a time-domain binary vector, providing the time-domain binary vector as an input value to the crossbar matrix array circuit, and outputting an output feature map by performing a convolution calculation between the binary weight values and the time-domain binary vector.
[0011] According to another aspect of an embodiment, there is provided a computer-readable recording medium having recorded thereon a program for implementing the above method on a computer. BRIEF DESCRIPTION OF DRAWINGS
[0012] The above and other aspects, features, and advantages of certain embodiments of the disclosure will be more apparent from the following description taken in conjunction with the accompanying drawings, in which:
[0013] Figure 1 is a diagram for describing a mathematical model of a biological neuron and an operation of simulating the biological neuron.
[0014] Figures 2A-2B is a diagram for describing an operation method of a neuromorphic apparatus according to an embodiment.
[0015] Figures 3A-3B is a diagram for comparing vector matrix multiplication with a calculation performed in a crossbar matrix (or crossbar) array according to an embodiment.
[0016] Figure 4 is a diagram for describing an example of performing a convolution calculation in a neuromorphic apparatus according to an embodiment.
[0017] Figure 5 is a diagram for describing a computation performed in a neural network according to an embodiment.
[0018] Figures 6A-6C is a diagram for describing an example of converting an initial weight value into a binary weight value according to an embodiment.
[0019] Figure 7A and Figure 7B is a diagram for describing an example of converting an input feature map into a time-domain binary vector according to an embodiment.
[0020] Figure 8 is a diagram for describing an application of a binary weight value and a time-domain binary vector to a batch normalization (or batch normalization) process according to an embodiment.
[0021] Figure 9 is a block diagram of a neural network apparatus using a von Neumann structure according to an embodiment.
[0022] Figure 10 is a block diagram of a neural network apparatus using an in-memory structure according to an embodiment.
[0023] Figure 11 is a flowchart of a method of implementing a neural network in a neural network apparatus according to an embodiment.
[0024] Figure 12 is a flowchart of a method of implementing a neural network in a neuromorphic apparatus according to an embodiment.
[0025] Figure 13 is a block diagram illustrating a hardware configuration of a neural network apparatus according to an embodiment.
[0026] Figure 14 is a block diagram illustrating a hardware configuration of a neuromorphic apparatus according to an embodiment. DETAILED DESCRIPTION
[0027] Reference will now be made in detail embodiments, examples of which are illustrated in the accompanying drawings, wherein like reference numerals refer to like elements throughout. In this regard, the present embodiments can have different forms and should not be construed as being limited to the descriptions set forth herein. Accordingly, the embodiments are merely descriptive of aspects and are not intended to limit the aspects. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. Expressions such as "at least one of," when preceding a list of elements, modify the entire list of elements and do not modify the individual elements of the list.
[0028] The phrases "in some embodiments," "in certain embodiments," "in various embodiments," and similar language, as used throughout this specification, can but do not necessarily refer to the same embodiments, unless the context specifically states otherwise. Unless specifically stated otherwise, as apparent from the following
[0029] Some embodiments can be described in terms of functional block components and various processing steps. Such functional blocks can be realized by any number of hardware and / or software components configured to perform the specified functions. For example, the disclosed functional blocks can be implemented with one or more microprocessors or circuitry configured to perform the specified functions. Additionally, the disclosed functional blocks can be implemented with any programming or scripting language. The functional blocks can be implemented in algorithms executing on one or more processors. Additionally, the disclosure can employ any number of conventional techniques for electronics configuration, signal processing and / or data processing. The terms "mechanism," "element," "means" and "composition" can be broad terms that encompass both hardware and software. The disclosed embodiments can be implemented in software and / or firmware. The software can be stored on one or more computer-readable storage media, such as RAM, ROM, EEPROM, flash memory or other memory, CD-ROM, etc. The software can include one or more programs that can be used to program a computer or other electronic device to implement the disclosed embodiments. The programs can be written in any of a number of high level programming languages such as C, C++, Java, Visual Basic, or the like, for use with any operating system or operating environment.
[0030] In addition, the connection lines or connectors shown in the various figures presented can be intended to represent exemplary functional and / or physical and logical connections between various elements. It should be noted that many alternative or additional functional relationships or physical connections or logical connections can be present in actual devices.
[0031] Hereinafter, the present disclosure will be described in detail with reference to the accompanying drawings.
[0032] Figure 1 is a diagram for describing a mathematical model of a biological neuron and an operation of the biological neuron.
[0033] A biological neuron represents a cell in a human nervous system. The biological neuron can be one of basic biological computing entities. The human brain includes about 100 billion biological neurons and about 100 trillion interconnections between the biological neurons.
[0034] The biological neuron 10 is a single cell. The biological neuron 10 includes a neuron cell body including a nucleus and various organelles. The various organelles can include mitochondria, a plurality of dendrites extending from the neuron cell body, and an axon extending from a large number of extensions.
[0035] In general, an axon performs a function of transmitting a signal from a neuron to another neuron, and a dendrite performs a function of receiving a signal from another neuron. For example, when different neurons are connected to each other, a signal transmitted through an axon of a neuron can be received by a dendrite of another neuron. At this time, a signal between neurons is transmitted through a dedicated connection called a synapse, and a plurality of neurons are connected to each other to form a neural network. A neuron that secretes a neurotransmitter based on a synapse can be referred to as a pre-synaptic neuron, and a neuron that receives information transmitted through a neurotransmitter can be referred to as a post-synaptic neuron.
[0036] Meanwhile, a human brain can learn and memorize a large amount of information by transmitting and processing various signals through a neural network formed by a large number of neurons connected to each other. Since a large number of connections between neurons in a human brain is directly related to the massive parallel nature of biological computing, various attempts have been made to efficiently process a large amount of information by artificially simulating a biological neural network. For example, a neuromorphic device is being researched as a computing system designed to implement an artificial neural network at a neuron level.
[0037] Meanwhile, the operation of a biological neuron 10 can be simulated through a mathematical model 11. The mathematical model 11 corresponding to the biological neuron 10 can include the following calculations as an example of neuromorphic computing: a multiplication calculation of multiplying information from a plurality of neurons by a synaptic weight, an addition calculation ∑ for values ω0x0, ω1x1, ω2x2 obtained by multiplying the synaptic weight, and a calculation for applying a feature function b and an activation function f to the result of the addition calculation. A neuromorphic computing result can be provided by neuromorphic computing. Here, values like x0, x1, x2, etc. correspond to axon values, and values like ω0, ω1, ω2, etc. correspond to synaptic weights.
[0038] Figures 2A-2B is a diagram for describing a method of operating a neuromorphic device according to an embodiment.
[0039] Referring to Figure 2A , the neuromorphic device can include a crossbar matrix array circuit unit. The crossbar matrix array circuit unit can include a plurality of crossbar matrix array circuits, each of which can be implemented as a resistive crossbar matrix memory array (RCA). In detail, each crossbar matrix array circuit can include an input node 210 corresponding to a pre-synaptic neuron, a neuron circuit 220 corresponding to a post-synaptic neuron, and a synapse circuit 230 providing a connection between the input node 210 and the neuron circuit 220.
[0040] In one embodiment, the crossbar matrix array circuit of the neuromorphic device includes four input nodes 210, four neuron circuits 220, and sixteen synapse circuits 230, but the number can vary. When the number of input nodes 210 is N (here, N is a natural number equal to or greater than 2) and the number of neuron circuits 220 is M (here, M is a natural number equal to or greater than 2 and can be the same as or different from N), NxM synapse circuits 230 can be arranged in a matrix shape.
[0041] In detail, a line 21 connected to the input nodes 210 and extending in a first direction (for example, a longitude direction) and a line 22 connected to the neuron circuits 220 and extending in a second direction (for example, a latitude direction) crossing the first direction can be provided. Hereinafter, for convenience of explanation, the line 21 extending in the first direction will be referred to as a row line, and the line 22 extending in the second direction will be referred to as a column line. The plurality of synapse circuits 230 can be arranged at respective intersections of the row lines 21 and the column lines 22, thereby connecting the corresponding row line 21 and the corresponding column line 22.
[0042] The input nodes 210 can generate a signal (for example, a signal corresponding to specific data) and transmit the signal to the row line 21, and the neuron circuits 220 can receive a synaptic signal passing through the synapse circuits 230 through the column line 22 and process the synaptic signal. The input nodes 210 can correspond to axons, and the neuron circuits 220 can correspond to neurons. However, whether the neuron is a presynaptic neuron or a postsynaptic neuron can be determined based on the relative relationship with other neurons. For example, when the input node 210 receives a synaptic signal related to another neuron, the neuron circuit 220 can function as a postsynaptic neuron. Similarly, when the neuron circuit 220 transmits a signal related to another neuron, the neuron circuit 220 can function as a presynaptic neuron.
[0043] The connection between the input nodes 210 and the neuron circuits 220 can be established through the synapse circuits 230. Here, the synapse circuit 230 is a device whose conductance or weight changes according to an electrical pulse (for example, voltage or current) applied to both ends thereof.
[0044] The synaptic circuit 230 can include, for example, a variable resistance element. The variable resistance device is a device that can be switched between different resistance states according to a voltage or a current applied to both ends thereof and can have a single layer structure or a multi-layer structure including various materials that can have a plurality of resistance states (for example, metal oxides like transition metal oxides and perovskite-based materials, phase change materials like chalcogenide materials, ferroelectric materials, ferromagnetic materials, etc.). The operation of the variable resistance element and / or the synaptic circuit 230 changing from a high resistance state to a low resistance state can be referred to as a set operation, and the operation of the variable resistance element and / or the synaptic circuit 230 changing from a low resistance state to a high resistance state can be referred to as a reset operation.
[0045] The following will be described with reference to Figure 2B The operation of the neuromorphic device will be described. For ease of explanation, the row lines 21 will be referred to as a first row line 21A, a second row line 21B, a third row line 21C, and a fourth row line 21D in order from the top, and the column lines 22 will be referred to as a first column line 22A, a second column line 22B, a third column line 22C, and a fourth column line 22D in order from the left.
[0046] Referring to Figure 2B In an initial state, all of the synaptic circuits 230 can be in a relatively low conductance state (i.e., a high resistance state). When some of the synaptic circuits 230 are in a low resistance state, an initialization operation for switching them to a high resistance state can be additionally required. Each of the synaptic circuits 230 can have a predetermined threshold for changing resistance and / or conductance. In detail, when a voltage or a current having a magnitude less than the predetermined threshold is applied to both ends of each of the synaptic circuits 230, the conductance of the synaptic circuit 230 can not be changed. On the other hand, when a voltage and / or a current having a magnitude greater than the predetermined threshold is applied to the synaptic circuit 230, the conductance of the synaptic circuit 230 can be changed.
[0047] In this state, in order to perform an operation for outputting a specific data as a result of a specific column line 22, an input signal corresponding to the specific data can be input to the row lines 21 in response to the output of the input node 210. At this time, the input signal can occur as a result of applying an electric pulse to each of the row lines 21. For example, when an input signal corresponding to data "0011" is input through the row lines 21, an electric pulse can not be applied to the row lines 21 (e.g., the first row line 21A and the second row line 21B) corresponding to "0", and an electric pulse can be applied only to the row lines 21 (e.g., the third row line 21C and the fourth row line 21D) corresponding to "1". At this time, the column lines 22 can be driven with an appropriate voltage or current for the output of the column lines 22.
[0048] For example, when the column line 22 that outputs specific data has been determined, the corresponding column line 22 can be driven so that the synapse circuit 230 located at the intersection with the row line 21 corresponding to "1" receives a voltage having a magnitude equal to or greater than a voltage required to perform a set operation (hereinafter, referred to as a set voltage), and the remaining column lines 22 can be driven so that the synapse circuits 230 receive a voltage having a magnitude smaller than the set voltage. For example, when the magnitude of the set voltage is Vsetand the third column line 22C is determined as the column line 22 for outputting data "0011", the magnitude of the electric pulse applied to the third row line 21C and the fourth row line 21D can be equal to or greater than Vsetand the voltage applied to the third column line 22C can be 0 V so that the first synapse circuit 230A and the second synapse circuit 230B located at the intersection between the third column line 22C and the third row line 21C and the fourth row line 21D receive a voltage equal to or greater than Vset. Accordingly, the first synapse circuit 230A and the second synapse circuit 230B can be in a low resistance state. The conductance of the first synapse circuit 230A and the second synapse circuit 230B in the low resistance state can gradually increase as the number of electric pulses increases. The magnitude and width of the electric pulse applied to its third row line 21C and fourth row line 21D can be substantially constant. The voltage applied to the remaining column lines (i.e., the first column line 22A, the second column line 22B, and the fourth column line 22D) can have a value between 0 V and Vset(e.g., 1 / 2 Vset) so that the synapse circuits 230 other than the first synapse circuit 230A and the second synapse circuit 230B receive a voltage smaller than Vset. Accordingly, the resistance state of the remaining synapse circuits 230 other than the first synapse circuit 230A and the second synapse circuit 230B can not be changed.
[0049] In another example, a specific column line 22 can not be designated to output specific data. In this case, the current flowing through each of the column lines 22 can be measured while applying an electric pulse corresponding to specific data to the row lines 21, and the column line 22 (e.g., the third column line 22C) first reaching a predetermined threshold current can be determined as the column line 22 that will output the specific data.
[0050] By the above-described method, different data can be respectively output to different column lines 22.
[0051] Figures 3A-3B is a diagram for comparing vector matrix multiplication with a calculation performed in a crossbar matrix array according to an embodiment.
[0052] First, referring to Figure 3AThe convolution computation between the input feature map and the weight values can be performed by using vector matrix multiplication. For example, the pixel data of the input feature map can be represented as a matrix X 310, and the weight values can be represented as a matrix W 311. The pixel data of the output feature map can be represented as a matrix Y 312, which is the result of the multiplication computation between the matrix X 310 and the matrix W 311.
[0053] Referring to Figure 3B The vector multiplication computation can be performed by using a non-volatile memory device of a crossbar matrix array. In contrast to Figure 3A The pixel data of the input feature map can be received as input values of the non-volatile memory device, which can be voltages 320 (e.g., voltages V1, V2, …, Vm, where m is a positive integer). In addition, the weight values can be stored in synapses (i.e., memory cells) of the non-volatile memory device, which can be conductances 321 (e.g., conductances G1, G2, …, Gn, where n is a positive integer) stored in the memory cells. Accordingly, the output values of the non-volatile memory device can be represented as currents 322 (e.g., currents I1, I2, …, In, where n is a positive integer), which are the result of the multiplication computation between the voltages 320 and the conductances 321. m 11 12 1m 21 22 2m n1 n2 nm m
[0054] Figure 4 is a diagram for describing an example of performing a convolution computation in a neuromorphic device according to an embodiment.
[0055] The neuromorphic device can receive pixels of an input feature map 410, and a crossbar matrix array circuit 400 of the neuromorphic device can be implemented as a resistive crossbar matrix memory array (RCA).
[0056] The neuromorphic device can receive an input feature map in a digital signal form, and convert the input feature map into a voltage in an analog signal form by using a digital-to-analog converter (DAC) 420. In one embodiment, the neuromorphic device can convert pixel values of the input feature map into a voltage by using the DAC 420, and provide the voltage as an input value 401 of the crossbar matrix array circuit 400.
[0057] Further, the learned weight values can be stored in the crossbar matrix array circuit 400 of the neuromorphic device. The weight values can be stored in the memory cells of the crossbar matrix array circuit 400, and the weight values stored in the memory cells can be the conductances 402. At this time, the neuromorphic device can calculate the output values by performing vector multiplication calculation between the input values 401 and the conductances 402, and the output values can be expressed as the currents 403. In other words, the neuromorphic device can output the same result as a result of convolution calculation between the input feature maps and the weight values by using the crossbar matrix array circuit 400.
[0058] Since the currents 403 output from the crossbar matrix array circuit 400 are analog signals, the neuromorphic device can use an analog-to-digital converter (ADC) 430 to use the currents 403 as input feature maps of another crossbar matrix array circuit. The neuromorphic device can use the ADC 430 to convert the currents 403, which are analog signals, into digital signals. In one embodiment, the neuromorphic device can convert the currents 403 into digital signals having the same number of bits as the pixels of the input feature maps 410 by using the ADC 430. For example, when the pixels of the input feature maps 410 are 4-bit data, the neuromorphic device can convert the currents 403 into 4-bit data by using the ADC 430.
[0059] The neuromorphic device can apply an activation function to the digital signals converted by the ADC 430 by using an activation unit 440. A sigmoid function, a Tanh function, and a rectified linear unit (ReLU) function can be used as the activation function, but the activation function that can be applied to the digital signals is not limited thereto.
[0060] The digital signals to which the activation function is applied can be used as input feature maps of another crossbar matrix array circuit 450. When the digital signals to which the activation function is applied are used as the input feature maps of the other crossbar matrix array circuit 450, the above-described processes can be applied to the other crossbar matrix array circuit 450 in the same manner.
[0061] Figure 5 is a diagram for describing a calculation performed in a neural network according to an embodiment.
[0062] Referring to Figure 5 , the neural network 500 can have a structure including an input layer, a hidden layer, and an output layer, perform a calculation based on received input data (e.g., I1 and I2), and generate output data (e.g., O1 and O2) based on a result of performing the calculation.
[0063] For example, as Figure 5As illustrated in FIG. 5, the neural network 500 can include an input layer (layer 1), two hidden layers (layer 2 and layer 3), and an output layer (layer 4). Since the neural network 500 includes more layers capable of processing effective information, the neural network 500 can process more complex data sets than a neural network having a single layer. Meanwhile, although the neural network 500 is illustrated as including four layers, it is merely an example, and the neural network 500 can include fewer or more layers or can include fewer or more channels. In other words, the neural network 500 can include layers of various structures different from the structure illustrated in FIG. 5. Figure 5 As illustrated in FIG. 5, the neural network 500 can include an input layer (layer 1), two hidden layers (layer 2 and layer 3), and an output layer (layer 4). Since the neural network 500 includes more layers capable of processing effective information, the neural network 500 can process more complex data sets than a neural network having a single layer. Meanwhile, although the neural network 500 is illustrated as including four layers, it is merely an example, and the neural network 500 can include fewer or more layers or can include fewer or more channels. In other words, the neural network 500 can include layers of various structures different from the structure illustrated in FIG. 5. Figure 5 As illustrated in FIG. 5, the neural network 500 can include an input layer (layer 1), two hidden layers (layer 2 and layer 3), and an output layer (layer 4). Since the neural network 500 includes more layers capable of processing effective information, the neural network 500 can process more complex data sets than a neural network having a single layer. Meanwhile, although the neural network 500 is illustrated as including four layers, it is merely an example, and the neural network 500 can include fewer or more layers or can include fewer or more channels. In other words, the neural network 500 can include layers of various structures different from the structure illustrated in FIG. 5.
[0064] Each layer included in the neural network 500 can include a plurality of channels. A channel can correspond to a plurality of artificial nodes (referred to as neurons, processing elements (PEs), units, or similar terms). For example, as illustrated in FIG. 5, layer 1 can include two channels (nodes), and layers 2 and 3 can each include three channels. However, this is merely an example, and the layers included in the neural network 500 can each include various numbers of channels (nodes). Figure 5
[0065] The channels included in each layer of the neural network 500 can be connected to each other and process data. For example, one channel can receive data from other channels and perform a calculation, and output a result of the calculation to other channels.
[0066] The input and output of each channel can be referred to as an input feature map and an output feature map. The input feature map can include a plurality of input activations, and the output feature map can include a plurality of output activations. In other words, a feature map or an activation can be an output of one channel and, at the same time, can be a parameter corresponding to an input of a channel included in a next layer. In one example, the input feature map and the output feature map can be generated through the neural network 500 in response to image data being input to the neural network 500, thereby performing image recognition. For example, the input feature map of the input layer (e.g., layer 1) can correspond to input image data, and the input feature map of each intermediate layer (e.g., a hidden layer) can correspond to an output feature map output from a previous layer. In this case, the output layer can output an image recognition result.
[0067] Meanwhile, each channel can determine its own activation based on the activation and weight values received from the channels included in the previous layer. The weight value is a parameter used to calculate the output activation in each channel, and can be a value assigned to a connection relationship between channels.
[0068] Each channel can be processed by a calculation unit or a PE that receives an input and outputs an output activation, and the input and output of each channel can be mapped. For example, when σ is an activation function, is a weight value from the kth channel included in the (i-1)th layer to the jth channel included in the ith layer, is a bias of the jth channel included in the ith layer and is an activation of the jth channel included in the ith layer, the activation may be calculated by using Equation 1 below.
[0069]
Equation 1
[0070]
[0071] As shown in Equation 1, Figure 5 the activation of the first channel CH1 of the second layer 2 can be expressed as In addition, according to Equation 1, may have a value However, Equation 1 described above is only an example for describing the activation and the weight value used to process data in the neural network 500, and the present disclosure is not limited thereto. The activation can be a value obtained by applying batch normalization and an activation function to a sum of activations received from a previous layer.
[0072] Figures 6A-6C is a diagram for describing an example of converting an initial weight value into a binary weight value according to an embodiment.
[0073] Referring to Figure 6A , an input layer 601, an output layer 602, and initial weight values W 11 , W 12 , …, W 32 , and W 33 are shown. Three input activations I1, I2, and I3 can correspond to three neurons of the input layer 601, respectively, and three output activations O1, O2, and O3 can correspond to three neurons of the output layer 602, respectively. In addition, the initial weight values W nm may be applied to the n-th input activation I n and the m-th output activation O m .
[0074] Figure 6B The initial weight values 610 are expressed in the form of a matrix Figure 6A W 11 , W 12 , …, W 32 , and W 33 shown in Equation 1.
[0075] The initial weight values 610 can be determined during a training process of a neural network. In one embodiment, the initial weight values 610 can be expressed as 32-bit floating point numbers.
[0076] The initial weight values 610 can be converted into binary weight values 620. The binary weight values 620 can each have a size of 1 bit. In the disclosure, a model size and an operation count can be reduced by using the binary weight values 620 instead of the initial weight values 610 during inference processing of a neural network. For example, when 32-bit initial weight values 610 are converted into 1-bit binary weight values 620, a model size can be compressed to 1 / 32.
[0077] In one embodiment, the initial weight values 610 can be converted into the binary weight values 620 based on a maximum value and a minimum value of the initial weight values 610. In one embodiment, the initial weight values 610 can be converted into the binary weight values 620 based on a maximum value and a minimum value of initial weight values that can be input to a neural network.
[0078] For example, a maximum value of initial weight values that can be input to a neural network can be 1.00, and a minimum value can be -1.00. When an initial weight value is 0.00 or more, the initial weight value can be converted into a binary weight value 1. When the initial weight value is less than 0.00, the initial weight value can be converted into a binary weight value -1.
[0079] Further, the binary weight values 620 can be multiplied by the average value 630 of the absolute values of the initial weight values 610. Since the binary weight values 620 are multiplied by the average value 630 of the absolute values of the initial weight values 610, a result similar to that of a case where the initial weight values 610 are used can be obtained even when the binary weight values 620 are used.
[0080] For example, when it is assumed that 1024 neurons exist in a previous layer and 512 neurons exist in a current layer, the 512 neurons belonging to the current layer each have 1024 initial weight values 610. Here, after the average value of the absolute values of the 1024 initial weight values 610 that are 32-bit floating-point numbers is calculated for each neuron, the binary weight values 620 can be multiplied by the result of the calculation.
[0081] In detail, the initial weight values 610 used to calculate the predetermined output activations O1, O2, and O3 can be multiplied by the average value 630 of the absolute values of the initial weight values 610 used to calculate the predetermined output activations O1, O2, and O3.
[0082] For example, with reference to Figure 6A , the initial weight values W 11 , W 21 , and W 31 may be used during processing of calculating the first output activation O1. The initial weight values W 11 , W 21 , and W 31 are converted into binary weight values W11 '、W 21 'and W 31 ', binary weight value W 11 '、W 21 'and W 31 'Can be compared with the initial weight value W 11 W 21 and W 31 The average of absolute values Multiply.
[0083] In the same respect, the binary weight value W 12 '、W 22 'and W 32 'Can be compared with the initial weight value W 12 W 22 and W 32 The average of absolute values Multiplication. Furthermore, the binary weight value W 13 '、W 23 'and W 33 'Can be compared with the initial weight value W 13 W 23 and W 33 The average of absolute values Multiply.
[0084] Reference Figure 6C The initial weight value 610, the binary weight value 620, and the average of the absolute values of the initial weight value 610, 630, are shown as specific values. Figure 6C For ease of explanation, the initial weight value 610 is represented in decimal, but it is assumed that the initial weight value 610 is a 32-bit floating-point number.
[0085] Figure 6C The diagram shows how to convert initial weight values equal to or greater than 0.00 into binary weight values of 1, and how to convert initial weight values less than 0.00 into binary weight values of -1.
[0086] also, Figure 6C The initial weight value W is shown. 11 W 21 and W 31 The average absolute value is 0.28, and the initial weight value W 12 W 22 and W 32 The average absolute value is "0.37", and the initial weight value W 13 W 23 and W 33 The average absolute value is "0.29".
[0087] Figure 7A and Figure 7Bis a diagram for describing an example of converting an input feature map into a time-domain binary vector according to an embodiment.
[0088] The input feature map can be converted into a plurality of time-domain binary vectors. The input feature map can include a plurality of input activations, and each of the plurality of input activations can be converted into a time-domain binary vector.
[0089] The input feature map can be converted into a plurality of time-domain binary vectors based on quantization levels. In one embodiment, a range between a maximum value and a minimum value of input activations that can be input to a neural network can be divided into N quantization levels (N is a natural number). For example, a sigmoid function or a tanh function can be used to classify the quantization levels, but the present disclosure is not limited thereto.
[0090] For example, referring to Figure 7A When nine quantization levels are set and the maximum value and the minimum value of input activations that can be input to a neural network are 1.0 and -1.0, respectively, the quantization levels can be "1.0, 0.75, 0.5, 0.25, 0, -0.25, -0.5, -0.75, and -1.0."
[0091] Meanwhile, although Figure 7A It is illustrated that intervals between the quantization levels are set to be the same, but the intervals between the quantization levels can be set in a non-linear manner.
[0092] When N quantization levels are set, a time-domain binary vector can have N-1 elements. For example, referring to Figure 7A When nine quantization levels are set, a time-domain binary vector can have eight elements t1, t2,..., t7, and t8.
[0093] An input activation can be converted into a time-domain binary vector based on a quantization level to which the input activation belongs from among N quantization levels. For example, when a predetermined input activation has a value equal to or greater than 0.75, the predetermined input activation can be converted into a time-domain binary vector "+1, +1, +1, +1, +1, +1, +1, +1". Further, in another example, when a predetermined input activation has a value less than -0.25 and equal to or greater than -0.5, the predetermined input activation can be converted into a time-domain binary vector "+1, +1, +1, -1, -1, -1, -1, -1".
[0094] Referring to Figure 7B, shows an example in which each of a plurality of input activations included in the input feature map 710 is converted into a time-domain binary vector. Since the first activation has a value less than 0 and equal to or greater than -0.25, the first activation can be converted into a time-domain binary vector of "-1, -1, -1, -1, +1, +1, +1, +1". Meanwhile, since the second activation has a value less than 0.5 and equal to or greater than 0.25, the second activation can be converted into a time-domain binary vector of "-1, -1, +1, +1, +1, +1, +1, +1". Also, since the third activation has a value less than -0.75 and equal to or greater than -1.0, the third activation can be converted into a time-domain binary vector of "-1, -1, -1, -1, -1, -1, -1, +1". Meanwhile, since the fourth activation has a value equal to or greater than 0.75, the fourth activation can be converted into a time-domain binary vector of "+1, +1, +1, +1, +1, +1, +1, +1".
[0095] Meanwhile, when each input activation of each layer of a neural network is converted into a binary value in a conventional manner, information carried by the input activation is lost, so the information can not be properly transmitted between layers.
[0096] On the other hand, as in the present disclosure, when each input activation of each layer of a neural network is converted into a time-domain binary vector, the original input activation can be approximated based on a plurality of binary values. Thus, for example, when a neural network is used for image recognition, a relatively high accuracy rate can be achieved.
[0097] Figure 8 is a diagram for describing application of binary weight values and time-domain binary vectors to a batch normalization (or batch normalization) process according to an embodiment.
[0098] Generally, in a neural network algorithm model, after a multiply-accumulate (MAC) calculation for multiplying an input activation (an initial input value or an output value from a previous layer) with an initial weight value (a 32-bit floating point number) and adding a result of the multiplication is performed, a separate bias value is added for each neuron. Then, a result thereof is subjected to batch normalization for each neuron. Next, after a result of the batch normalization is input into an activation function, an output value of the activation function is transmitted as an input value of a next layer.
[0099] The above-described process can be expressed as Equation 2. In Equation 2, I n denotes an input activation, W nm denotes an initial weight value, B m denotes a bias value, α m denotes an initial scale value of batch normalization, β mdenotes a bias value of batch normalization, f denotes an activation function, O m denotes an output activation.
[0100]
Equation 2
[0101]
[0102] Referring to Figure 8 , an input activation I n 810 can be converted into a time-domain binary vector I b n (t) 820. The time-domain binary vector generator can convert the input activation I n 810 into the time-domain binary vector I b n (t) 820.
[0103] As described above with reference to Figures 7A-7B , the input activation I n 810 can be converted into a time-domain binary vector I b n (t) 820 according to a preset quantization level. Meanwhile, the number of elements included in each time-domain binary vector I b n (t) 820 can be determined according to the number of quantization levels. For example, when the number of quantization levels is N, the number of elements included in each time-domain binary vector I b n (t) 820 can be N-1.
[0104] On the other hand, when the input activation I n 810 is converted into the time-domain binary vector I b n (t) 820, a calculation result using the time-domain binary vector I b n (t) 820 can be amplified by the number T 850 of elements included in the time-domain binary vector I b n (t) 820. Accordingly, in the case of using the time-domain binary vector I b n (t) 820, the calculation result can be divided by the number T 850 of elements to obtain the same result as that of the original MAC calculation. Detailed descriptions thereof will be given by using Equation 5 and Equation 6.
[0105] As described above with reference to Figures 6A-6C , the initial weight value W nm can be converted into a binary weight value W b nm830. For example, the initial weight value W nm It can be converted into a binary weight value W through a sign function. b nm 830.
[0106] Time-domain binary vector I b n (t)820 and binary weight value W b nm Convolution computations between 830 and 830 can be performed. In one embodiment, the time-domain binary to I... b n (t)820 and binary weight value W b nm XNOR and addition operations between 830 can be performed.
[0107] Execute the time-domain binary vector I b n (t)820 and binary weight value W b nm After calculating XNOR between 830 and summing the results, the original multi-bit input activation I is performed. n 810 and binary initial weight value W nm The same increasing / decreasing pattern appears in the results of convolution calculations between them.
[0108] Time-domain binary vector I b n (t)820 and binary weight value W b nm The calculation between 830 can be expressed as Equation 3 below.
[0109] Equation 3
[0110]
[0111] As a result of the convolution calculation, the intermediate activation X m 840 can be obtained. Mid-activation X m 840 can be represented as Equation 4 below.
[0112] Equation 4
[0113]
[0114] Intermediate activation X m 840 can be averaged with the absolute values of the initial weights S m Multiply by 860.
[0115] In addition, intermediate activation X m840 can be divided by the binary vector I included in the time domain. b n The number of elements in each of (t)820 is T 850. This includes the time-domain binary vector I. b n The number of elements T in each of (t)820, T850, can be determined based on the number of quantization levels. This is due to the use of a time-domain binary vector I. b n The result of the calculation of (t)820 is amplified by the number of elements T850, thus activating the intermediate X. m 840 can be divided by the number of elements T 850 to obtain the same result as the original MAC calculation.
[0116] When X is activated in the middle m The average of 840 and the absolute values of the initial weights, S m 860 multiplied and divided by the binary vector I included in the time domain b n When the number of elements in each of (t)820 is T = 850, the output is activated O. m 870 is available. Output activation O m 870 can be represented as Equation 5 below.
[0117] Equation 5
[0118] O m =X m ×S m ÷T
[0119] In one embodiment, when batch normalization is performed, the average of the absolute values S of the initial scale value and the initial weight value of the batch normalization is used. m 860 multiplied and divided by the binary vector I included in the time domain b n The number of elements in each of (t)820 is T 850, therefore the scale value α is modified. m It can be obtained.
[0120] When the binary weight value W is determined according to Equation 2 b nm Time-domain binary vector I b n (t) and correction scale value α" m (That is, the modified scale value α) m When applied to a neural network algorithm model, Equation 2 can be expressed as Equation 6 below.
[0121] Equation 6
[0122]
[0123] In the disclosure, by converting an initial weight value W nm expressed as a multi-bit floating point number into a binary weight value W b nm 830, the model size and operation count can be reduced.
[0124] In the disclosure, by multiplying a binary weight value W b nm 830 by an average value S m 860 of absolute values of the initial weight values, even when the binary weight value W b nm 830 is used, a result similar to that of a case where the initial weight value W nm is used can be obtained.
[0125] On the other hand, since the average value S m 860 of absolute values of the initial weight values can be included in the batch normalization calculation (M m × α m ) as shown in Equation 6, no additional model parameters are generated, and thus there is no loss in model size reduction and operation count reduction. In other words, as compared to Equation 2, it can be seen that, in Equation 6, the calculation can be performed without additional parameters and a separate process.
[0126] In the disclosure, a multi-bit input activation I n 810 can be quantized to a low number of bits (e.g., 2 to 3 bits), and a result thereof can be converted into a time-domain binary vector I b n (t) 820 having a plurality of elements. Further, in the disclosure, by performing a time-axis XNOR calculation between a binary weight value W b nm 830 and the time-domain binary vector I b n (t) 820, a learning (or training) performance and a final classification / recognition accuracy level similar to those of a 32-bit floating point neural network based on MAC calculation can be ensured.
[0127] On the other hand, since the number T 850 of elements can be included in the batch normalization calculation (α mIn Equation 6, among the equations of Equation 1, Equation 2, Equation 3, Equation 4, and Equation 5, additional model parameters are not generated, and thus there is no loss in model size reduction and operation count reduction. In other words, as compared with Equation 2, it can be seen that, in Equation 6, the calculation can be performed without additional parameters and a separate process. Thus, for example, when the neural network is used for image recognition, a relatively fast recognition speed can be achieved.
[0128] For example, when a 32-bit input activation I n 810 is converted into a time-domain binary vector I b n (t) 820, the model size can be compressed to T / 32.
[0129] Figure 9 is a block diagram of a neural network apparatus using a von Neumann structure according to an embodiment.
[0130] Referring to Figure 9 , the neural network apparatus 900 can include an external input receiver 910, a memory 920, a time-domain binary vector generator 930, a convolution calculation unit 940, and a neural calculation unit 950.
[0131] In the neural network apparatus 900 shown in Figure 9 , only components related to the present disclosure are shown. Thus, it is obvious to one of ordinary skill in the art that the neural network apparatus 900 can include other general components in addition to the components shown in Figure 9 .
[0132] The external input receiver 910 can receive neural network model-related information, input image (or audio) data, etc. from the outside. The various types of information and data received by the external input receiver 910 can be stored in the memory 920.
[0133] In one embodiment, the memory 920 can be divided into a first memory for storing input feature maps and a second memory for storing binary weight values, other real number parameters, and model structure definition variables. Meanwhile, the binary weight values stored in the memory 920 can be values obtained by converting initial weight values (e.g., 32-bit floating point numbers) for which learning (or training) of the neural network is completed.
[0134] The time-domain binary vector generator 930 can receive the input feature maps from the memory 920. The time-domain binary vector generator 930 can convert the input feature maps into time-domain binary vectors. The input feature maps can include a plurality of input activations, and the time-domain binary vector generator 930 can convert each of the plurality of input activations into a time-domain binary vector.
[0135] In detail, the time-domain binary vector generator 930 can convert the input feature map into a plurality of time-domain binary vectors based on quantization levels. In one embodiment, when a range between a maximum value and a minimum value of input activations that can be input to a neural network can be divided into N quantization levels (N is a natural number), the time-domain binary vector generator 930 can convert the input activations into a time-domain binary vector having N-1 elements.
[0136] The convolution calculation unit 940 can receive the binary weight values from the memory 920. Also, the convolution calculation unit 940 can receive the plurality of time-domain binary vectors from the time-domain binary vector generator 930.
[0137] The convolution calculation unit 940 can include an adder, and the convolution calculation unit 940 can perform a convolution calculation between the binary weight values and the plurality of time-domain binary vectors.
[0138] The neural calculation unit 950 can receive the binary weight values and a result of the convolution calculation between the binary weight values and the plurality of time-domain binary vectors from the convolution calculation unit 940. Also, the neural calculation unit 950 can receive a modified scale value of batch normalization, a bias value of batch normalization, an activation function, etc. from the memory 920.
[0139] Batch normalization and pooling can be performed, and an activation function can be applied in the neural calculation unit 950. However, the calculations that can be performed and applied in the neural calculation unit 950 are not limited thereto.
[0140] Meanwhile, the modified scale value of batch normalization can be obtained by multiplying an initial scale value by an average value of absolute values of initial weight values and dividing the result by the number T of elements included in each time-domain binary vector.
[0141] When batch normalization is performed and an activation function is applied in the neural calculation unit 950, an output feature map can be output. The output feature map can include a plurality of output activations.
[0142] Figure 10 is a block diagram of a neural network apparatus using an in-memory structure according to an embodiment.
[0143] Referring to Figure 10 , the neuromorphic apparatus 1000 can include an external input receiver 1010, a memory 1020, a time-domain binary vector generator 1030, an on-chip memory 1040, and a neural calculation unit 1050.
[0144] In Figure 10Among the neuromorphic device 1000 illustrated in FIG. 10, only components related to the present disclosure are illustrated. Thus, it is obvious to one of ordinary skill in the art that the neuromorphic device 1000 can include other general components in addition to Figure 10 The neuromorphic device 1000 illustrated in FIG. 10 can include other general components in addition to the components illustrated in FIG. 10.
[0145] The external input receiver 1010 can receive, from the outside, neural network model-related information, input image (or audio) data, etc. The various types of information and data received by the external input receiver 1010 can be stored in the memory 1020.
[0146] The memory 1020 can store input feature maps, other real number parameters, model structure definition variables, etc., with respect to the neural network device 900 of FIG. 9. Figure 9 Unlike the neural network device 900 of FIG. 9, binary weight values can be stored in the on-chip memory 1040 instead of the memory 1020, a detailed description of which will be given later.
[0147] The time-domain binary vector generator 1030 can receive input feature maps from the memory 1020. The time-domain binary vector generator 1030 can convert the input feature maps into time-domain binary vectors. The input feature maps can include a plurality of input activations, and the time-domain binary vector generator 1030 can convert each of the plurality of input activations into a time-domain binary vector.
[0148] In detail, the time-domain binary vector generator 1030 can convert the input feature maps into a plurality of time-domain binary vectors based on quantization levels. In one embodiment, when a range between a maximum value and a minimum value of input activations that can be input to a neural network can be divided into N quantization levels (N is a natural number), the time-domain binary vector generator 1030 can convert the input activations into time-domain binary vectors having N-1 elements.
[0149] The on-chip memory 1040 can include an input unit 1041, a crossbar matrix array circuit 1042, and an output unit 1043.
[0150] The crossbar matrix array circuit 1042 can include a plurality of synapse circuits (e.g., variable resistors). Binary weight values can be stored in the plurality of synapse circuits. The binary weight values stored in the plurality of synapse circuits can be values obtained by converting initial weight values (e.g., 32-bit floating point numbers) for which learning (or training) of a neural network is completed.
[0151] The input unit 1041 can receive a plurality of time-domain binary vectors from the time-domain binary vector generator 1030.
[0152] When the input unit 1041 receives the plurality of time-domain binary vectors, the crossbar matrix array circuit 1042 can perform a convolution calculation between the binary weight values and the plurality of time-domain binary vectors.
[0153] The output unit 1043 can transmit a result of the convolution calculation to the neural computing unit 1050.
[0154] The neural computing unit 1050 can receive the binary weight values and the result of the convolution calculation between the binary weight values and the plurality of time-domain binary vectors from the output unit 1043. Also, the neural computing unit 1050 can receive the modified scale value of the batch normalization, the bias value of the batch normalization, the activation function, etc. from the memory 1020.
[0155] The batch normalization and the pooling can be performed, and the activation function can be applied in the neural computing unit 1050. However, the calculations that can be performed and applied in the neural computing unit 1050 are not limited thereto.
[0156] Meanwhile, the modified scale value of the batch normalization can be obtained by multiplying the initial scale value by the average value of the absolute values of the initial weight values and dividing the result by the number T of elements included in each time-domain binary vector.
[0157] When the batch normalization is performed and applied in the activation function neural computing unit 1050, an output feature map can be output. The output feature map can include a plurality of output activations.
[0158] Figure 11 is a flowchart of a method of implementing a neural network in a neural network device according to an embodiment.
[0159] Referring to Figure 11 In operation 1110, the neural network device can obtain the binary weight values and the input feature map from the memory.
[0160] In operation 1120, the neural network device can convert the input feature map into a time-domain binary vector.
[0161] In one embodiment, the neural network device can convert the input feature map into a time-domain binary vector based on quantization levels.
[0162] Specifically, the neural network device can divide a range between a maximum value and a minimum value that can be input to the neural network into N (N is a natural number) quantization levels, and convert respective activations of the input feature map into a time-domain binary vector based on a quantization level to which the respective activations belong among the N quantization levels.
[0163] Meanwhile, the neural network device can divide a range between a maximum value and a minimum value that can be input to the neural network into linear quantization levels or non-linear quantization levels.
[0164] In operation 1130, the neural network device can output the output feature map by performing a convolution calculation between the binary weight values and the time domain binary vectors.
[0165] The neural network device can output the output feature map by performing batch normalization on the result of the convolution calculation.
[0166] In one embodiment, the neural network device can obtain a modified scale value by multiplying an initial scale value of the batch normalization by an average value of absolute values of the initial weight values and dividing the result thereof by the number of elements included in each time domain binary vector. The neural network device can perform the batch normalization based on the modified scale value.
[0167] The neural network device can perform a multiplication calculation for multiplying each bias value applied to the neural network by the initial scale value, and reflect the result of the multiplication calculation to the output feature map.
[0168] The neural network device can output the output feature map by performing batch normalization on the result of the convolution calculation and applying an activation function to the result of the batch normalization.
[0169] Figure 12 is a flowchart of a method of implementing a neural network in a neuromorphic device according to an embodiment.
[0170] Referring to Figure 12 In operation 1210, the neuromorphic device can store the binary weight values in the synapse circuits included in the crossbar matrix array circuit.
[0171] In operation 1220, the neuromorphic device can obtain the input feature map from the memory.
[0172] In operation 1230, the neuromorphic device can convert the input feature map into the time domain binary vector.
[0173] In one embodiment, the neuromorphic device can convert the input feature map into the time domain binary vector based on quantization levels.
[0174] Specifically, the neuromorphic device can divide a range between a maximum value and a minimum value, which can be input to the neural network, into N (N is a natural number) quantization levels, and convert respective activations of the input feature map into the time domain binary vector based on a quantization level to which the respective activations belong among the N quantization levels.
[0175] Meanwhile, the neuromorphic device can divide a range between a maximum value and a minimum value, which can be input to the neural network, into linear quantization levels or non-linear quantization levels.
[0176] In operation 1240, the neuromorphic device can provide the time-domain binary vector as an input value of the crossbar matrix array circuit.
[0177] In operation 1250, the neuromorphic device can output the output feature map by performing a convolution calculation between the binary weight value and the time-domain binary vector.
[0178] The neuromorphic device can output the output feature map by performing batch normalization on the result of the convolution calculation.
[0179] In one embodiment, the neuromorphic device can obtain a modified scale value by multiplying an initial scale value of the batch normalization by an average value of absolute values of the initial weight values and dividing the result thereof by the number of elements included in each time-domain binary vector. The neuromorphic device can perform the batch normalization based on the modified scale value.
[0180] The neuromorphic device can perform a multiplication calculation for multiplying each bias value applied to the neural network by the initial scale value and reflect the result of the multiplication calculation to the output feature map.
[0181] The neuromorphic device can output the output feature map by performing batch normalization on the result of the convolution calculation and applying an activation function to the result of the batch normalization.
[0182] Figure 13 is a block diagram illustrating a hardware configuration of a neuromorphic device according to an embodiment.
[0183] The neuromorphic device 1300 can be implemented as various types of devices such as a personal computer (PC), a server device, a mobile device, and an embedded device. In detail, for example, the neuromorphic device 1300 can be a smart phone, a tablet device, an augmented reality (AR) device, an Internet of Things (IoT) device, an autonomous vehicle, a robot, a medical device, etc. that performs voice recognition, image recognition, image classification, etc. by using a neural network. However, the present disclosure is not limited thereto. Further, the neuromorphic device 1300 can correspond to a dedicated hardware accelerator mounted on a device as described above. The neuromorphic device 1300 can be a hardware accelerator like a neural processing unit (NPU), a tensor processing unit (TPU), or a neural engine that is a dedicated module for driving a neural network, but is not limited thereto.
[0184] Referring to Figure 13 , the neuromorphic device 1300 includes a processor 1310 and a memory 1320. In the neuromorphic device 1300 illustrated in Figure 13 , only components related to the present disclosure are illustrated. Thus, it is apparent to one of ordinary skill in the art that the neuromorphic device 1300 can include other general components in addition to the components illustrated in Figure 13 .
[0185] The processor 1310 serves to control the overall functions for operating the neural network device 1300. For example, the processor 1310 overall controls the neural network device 1300 by executing a program stored in the memory 1320 in the neural network device 1300. The processor 1310 can be implemented to include a central processing unit (CPU), a graphics processing unit (GPU), an application processor (AP), etc., in the neural network device 1300, but is not limited thereto.
[0186] The memory 1320 is a hardware component that stores various types of data processed in the neural network device 1300. For example, the memory 1320 can store data processed in the neural network device 1300 and data to be processed. In addition, the memory 1320 can store an application, a driver, etc., to be executed by the neural network device 1300. The memory 1320 can include a random access memory (RAM) such as a dynamic random access memory (DRAM) and a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a CD-ROM, a Blu-ray or another optical disk storage device, a hard disk drive (HDD), a solid state drive (SSD), or a flash memory.
[0187] The processor 1310 reads neural network data such as image data, feature map data, and weight value data from the memory 1320 / writes neural network data to the memory 1320, and performs a neural network by using the read / written data. When the neural network is performed, the processor 1310 repeatedly performs convolution calculation between input feature maps and weight values to generate data on output feature maps. At this time, the amount of calculation of the convolution calculation can be determined based on various factors such as the number of channels of the input feature maps, the number of channels of the weight values, the size of the input feature maps, the size of the weight values, and the precision of values.
[0188] A real neural network driven by the neural network device 1300 can be implemented with a more complex architecture. Accordingly, the processor 1310 performs a very large operation count ranging from several hundred million to several hundred billion, and thus it is inevitable that the frequency of accessing the memory 1320 by the processor 1310 for calculation sharply increases. Due to the burden of such calculation, the neural network can not be smoothly processed in mobile devices like smartphones, tablet PCs, wearable devices, etc., and embedded devices having relatively low processing power.
[0189] The processor 1310 can perform convolution calculation, batch normalization calculation, pooling calculation, activation function calculation, etc. In one embodiment, the processor 1310 can perform matrix multiplication calculation, conversion calculation, and transpose calculation to obtain multi-head self attention. In the process of obtaining multi-head self attention, the conversion calculation and the transpose calculation can be performed after or before the matrix multiplication calculation.
[0190] The processor 1310 can obtain the binary weight value and the input feature map from the memory 1320, and convert the input feature map into a time domain binary vector. Furthermore, the processor 1310 can output the output feature map by performing convolution calculation between the binary weight value and the time domain binary vector.
[0191] Figure 14 is a block diagram illustrating a hardware configuration of a neuromorphic device according to an embodiment.
[0192] Referring to Figure 14 , the neuromorphic device 1400 can include a processor 1410 and an on-chip memory 1420. In Figure 14 , the neuromorphic device 1400 illustrated in FIG. 14A includes only components related to the present disclosure. Thus, it is obvious to one of ordinary skill in the art that the neuromorphic device 1400 can include other general components in addition to the components illustrated in FIG. 14A. Figure 14
[0193] The neuromorphic device 1400 can be mounted on a digital system requiring a low-power neural network, such as a smart phone, a drone, a tablet device, an augmented reality (AR) device, an Internet of Things (IoT) device, an autonomous vehicle, a robot, a medical device, etc. However, the present disclosure is not limited thereto.
[0194] The neuromorphic device 1400 can include a plurality of on-chip memories 1420, each of which can include a plurality of crossbar matrix array circuits. The crossbar array circuit can include a plurality of presynaptic neurons, a plurality of postsynaptic neurons, and a synapse circuit (i.e., a memory cell) providing a connection between the plurality of presynaptic neurons and the plurality of postsynaptic neurons. In one embodiment, the crossbar matrix array circuit can be implemented as an RCA.
[0195] The external memory 1430 is a hardware component that stores various types of data processed in the neuromorphic device 1400. For example, the external memory 1430 can store data processed in the neuromorphic device 1400 and data to be processed. Also, the external memory 1430 can store an application, a driver, etc. to be executed by the neuromorphic device 1400. The external memory 1430 can include a random access memory (RAM) such as a dynamic random access memory (DRAM) and a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a CD-ROM, a Blu-ray or another optical disk storage device, a hard disk drive (HDD), a solid state drive (SSD), or a flash memory.
[0196] The processor 1410 serves to control the overall functions for operating the neuromorphic device 1400. For example, the processor 1410 overall controls the neuromorphic device 1400 by executing a program stored in the on-chip memory 1420 included in the neuromorphic device 1400. The processor 1410 can be implemented as a central processing unit (CPU), a graphics processing unit (GPU), an application processor (AP), etc. included in the neuromorphic device 1400, but is not limited thereto. The processor 1410 reads / writes various data from / to the external memory 1430 and performs the neuromorphic device 1400 by using the read / written data.
[0197] The processor 1410 can generate a plurality of binary feature maps by binarizing pixel values of an input feature map based on a plurality of threshold values. The processor 1410 can provide the pixel values of the plurality of binary feature maps as input values of the crossbar matrix array circuit. The processor 1410 can convert the pixel values into analog signals (voltages) using a DAC.
[0198] The processor 1410 can store weight values to be applied to the crossbar matrix array circuit unit in the synapse circuit included in the crossbar matrix array circuit unit. The weight values stored in the synapse circuit can be conductances. Also, the processor 1410 can obtain output values of the crossbar matrix array circuit unit by performing a multiplication calculation between the input values and the kernel values stored in the synapse circuit.
[0199] The processor 1410 can generate pixel values of an output feature map by merging the output values calculated by the crossbar matrix array circuit unit. Meanwhile, since the output values calculated by the crossbar matrix array circuit unit (or result values obtained by multiplying the calculated output values with weight values) are in the form of analog signals (currents), the processor 1410 can convert the output values into digital signals by using an ADC. Also, the processor 1410 can apply an activation function to the output values converted into digital signals by the ADC.
[0200] The processor 1410 can store the binary weight values in the synapse circuit including the crossbar matrix array circuit, and obtain the input feature map from the external memory 1630. Further, the processor 1410 can convert the input feature map into a time-domain binary vector, and provide the time-domain binary vector as an input value of the crossbar matrix array circuit. Further, the processor 1410 can output the output feature map by performing convolution calculation between the binary weight values and the time-domain binary vector.
[0201] One or more exemplary embodiments can be implemented by a computer-readable recording medium such as a program module executed by a computer. The computer-readable recording medium can be any available medium that can be accessed by a computer, and examples thereof include all volatile media (e.g., RAM) and non-volatile media (e.g., ROM), and detachable media and non-detachable media. Further, examples of the computer-readable recording medium can include computer storage media and communication media. Examples of the computer storage media include all volatile media and non-volatile media that have been implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, and other data, and detachable media and non-detachable media. The communication media typically include computer-readable instructions, data structures, program modules, other data modulated data signals, or additional transmission mechanisms, and examples thereof include any information transmission media.
[0202] Further, in the present specification, a "unit" can be a hardware component like a processor or a circuit and / or a software component executed by a hardware configuration like a processor.
[0203] While the inventive concept has been particularly shown and described with reference to exemplary embodiments thereof, it will be understood by those of ordinary skill in the art that various changes in form and details can be made therein without departing from the spirit and scope of the inventive concept as defined by the following claims. Accordingly, it will be understood that the above-described exemplary embodiments are not limiting of the scope of the inventive concept. For example, each component described in a single type can be executed in a distributed manner, and the distributed components described can be executed in an integrated form.
[0204] According to the above-described embodiments disclosed, by using the binary weight values and the time-domain binary vector, it is possible to reduce the model size and the operation count.
[0205] Further, according to another embodiment of the present disclosure, by performing the time-axis XNOR calculation between the binary weight values and the time-domain binary vector, a learning (or training) performance and a final classification / recognition accuracy level similar to those of a neural network using multi-bit data can be ensured.
[0206] It is to be understood that the embodiments described herein are to be considered in a descriptive sense only and not for purposes of limitation. Descriptions of features or aspects within each embodiment should typically be considered as being applicable to similar features or aspects within other embodiments. While one or more embodiments have been described with reference to the attached drawings, it will be apparent to those of ordinary skill in the art that various changes in form and details can be made therein without departing from the spirit and scope as defined by the following claims.
Claims
1. A neuromorphic device for implementing a neural network, the neuromorphic device comprising: a memory storing at least one program; an on-chip memory including a crossbar matrix array circuit; and at least one processor configured to drive a neural network by executing the at least one program, wherein the at least one processor is further configured to: store binary weight values in synapse circuits included in the crossbar matrix array circuit, wherein initial weight values allowed to be input to the neural network are converted to the binary weight values having values of +1 or -1 based on a maximum value and a minimum value of the initial weight values; obtain an input feature map from the memory; convert the input feature map to a time-domain binary vector; provide the time-domain binary vector as an input value to the crossbar matrix array circuit; and output an output feature map by performing a convolution calculation between the binary weight values and the time-domain binary vector using the crossbar matrix array circuit, wherein the at least one processor is configured to divide a range between a maximum value and a minimum value of input activations allowed to be input to the neural network into N quantization levels, and convert each of a plurality of input activations of the input feature map to a time-domain binary vector having N-1 elements each of +1 or -1 based on a quantization level to which each of the plurality of input activations of the input feature map belongs, wherein N is a positive integer. The at least one processor is further configured to output the output feature map by performing batch normalization on a result of the convolution calculation.
2. The neuromorphic device of claim 1, wherein, The at least one processor is further configured to:
3. The neuromorphic device of claim 2, wherein, calculate a modified scale value by multiplying an initial scale value of the batch normalization by an average value of absolute values of the initial weight values and dividing a result of the multiplication by a number of elements included in each of the time-domain binary vectors; and perform the batch normalization based on the modified scale value. The at least one processor is further configured to divide the range between the maximum value and the minimum value allowed to be input to the neural network into non-linear quantization levels. The at least one processor is further configured to:
4. The neuromorphic device of claim 1, wherein, perform a multiplication calculation for multiplying each bias value applied to the neural network by the initial scale value; and 5. The neuromorphic device of claim 3 or 4, wherein, reflect a result of the multiplication calculation to the output feature map. The at least one processor is further configured to output the output feature map by performing the batch normalization on a result of the convolution calculation and applying an activation function to a result of the batch normalization. 7.A method of implementing a neural network in a neuromorphic device, the method comprising:
6. The neuromorphic device of claim 3 or 4, wherein, storing binary weight values in synapse circuits included in a crossbar matrix array circuit in the neuromorphic device, wherein initial weight values allowed to be input to the neural network are converted to the binary weight values having values of +1 or -1 based on a maximum value and a minimum value of the initial weight values; obtaining an input feature map from a memory in the neuromorphic device; converting the input feature map to a time-domain binary vector; providing the time-domain binary vector as an input value to the crossbar matrix array circuit; and outputting an output feature map by performing a convolution calculation between the binary weight values and the time-domain binary vector using the crossbar matrix array circuit, wherein the converting the input feature map into the time-domain binary vector includes dividing a range between a maximum value and a minimum value of input activations allowed to be input to the neural network into N quantization levels, and converting each of a plurality of input activations of the input feature map into the time-domain binary vector having N-1 elements each of which is +1 or -1 based on a quantization level to which each of the plurality of input activations of the input feature map belongs, wherein N is a positive integer.
8. The method of claim 7, wherein, The outputting the output feature map includes outputting the output feature map by performing batch normalization on a result of the convolution calculation.
9. The method of claim 8, wherein, The performing the batch normalization includes: calculating a modified scale value by multiplying an initial scale value of the batch normalization by an average value of absolute values of initial weight values and dividing a result of the multiplication by a number of elements included in each time-domain binary vector; and performing the batch normalization based on the modified scale value.
10. The method of claim 7, wherein, The dividing includes dividing a range between a maximum value and a minimum value allowed to be input to the neural network into non-linear quantization levels.
11. The method of claim 9 or 10, wherein, The method further includes: performing a multiplication calculation for multiplying each bias value applied to the neural network by the initial scale value; and reflecting a result of the multiplication calculation to the output feature map.
12. The method of claim 9 or 10, wherein, The outputting the output feature map by performing the batch normalization on the result of the convolution calculation includes outputting the output feature map by performing the batch normalization on the result of the convolution calculation and applying an activation function to a result of the batch normalization. 13.A computer-readable recording medium having recorded thereon a program for executing the method according to any one of claims 7 to 12 on a computer.
Citation Information
Patent Citations
Heat-transfer type air purification apparatus for window
KR1020200069100A