In-memory computing (IMC) device and method of operating an IMC device
Through the IMC device that performs multiplication and accumulation operations in memory, the problems of large amount of computing and high power consumption in neural network computing are solved, and efficient neural network computing mode switching is realized, which improves computing performance and reduces power consumption.
Patent Information
- Application Number
- CN202411517046.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-10-29
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, when implementing neural network computing, there is a problem of large amount of calculation and high power consumption, especially when performing simple operations such as multiplication and accumulation operations, data movement becomes a bottleneck in performance and power.
Using an in-memory computing (IMC) device, the mode switching between pulsed neural network (SNN) and artificial neural network (NN) is implemented by performing multiplication and accumulation operations in memory, combining cross-switch arrays and post-arithm circuits, and the calculation mode is optimized to improve efficiency.
By reducing data movement and improving power efficiency, the IMC device improves computing performance and reduces power consumption when performing neural network computing, and is suitable for a variety of electronic devices.
Smart Images

Figure CN120235201A_ABST
Abstract
Description
[0001] This application claims the benefit of Korean Patent Application No. 10-2023-0194423, filed on Dec. 28, 2023, with the Korean Intellectual Property Office, the entire disclosure of which is incorporated herein by reference for all purposes. Technical Field
[0002] The following description relates to in-memory computing (IMC) devices and methods of operating IMC devices. Background Art
[0003] In many application fields, various types of neural networks trained using machine learning and / or deep learning can be used to provide high performance in terms of, for example, accuracy, speed, and / or energy efficiency. Algorithms for machine learning that implement neural networks may require high computational amounts, but the operations for the calculations of the algorithms may be relatively simple operations (such as multiply-accumulate (MAC) operations that calculate the dot product of two vectors and accumulate the result values of the dot product). Uncomplicated operations such as MAC operations can be implemented through in-memory computing (IMC). Summary of the Invention
[0004] This Summary of the Invention is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to help determine the scope of the claimed subject matter.
[0005] In one general aspect, an in-memory computing (IMC) macro has an operation mode that can alternate between a first mode and a second mode, and the IMC macro includes: an input control circuit configured to be able to generate a signal in which a predefined mode is applied to an input signal, and be able to send a previous operation result as feedback, and the input control circuit is executed according to which mode the operation mode is in; a crossbar array including memory cells, the memory cells including additional rows that process and store the previous operation result of the feedback, and columns including adder trees corresponding to the memory cells; and a post-arithmetic circuit configured to be able to execute a first operation corresponding to a spiking neural network (SNN) and a second operation corresponding to an artificial neural network (ANN), wherein which of the first operation and the second operation is executed depends on which mode the operation mode is in.
[0006] The memory cells may include rows that store weights corresponding to the input signal, and wherein the adder tree is configured to: add a first operation result between the input signal and the weights and a second operation result between the predefined mode and the previous operation result.
[0007] The input signal may include: a pulse signal for an SNN or a feature map for an NN.
[0008] The IMC macro can be configured to set the operation mode to a first mode for the SNN or a second mode for the NN according to a command sent from the host.
[0009] The first mode can be used for the SNN, and the second mode can be used for the NN. The input control circuit can also be configured to set a predefined mode to 1 based on the operation mode being in the first mode, and to set the predefined mode to a mode corresponding to the number of bits of the input signal based on the operation mode being in the second mode.
[0010] The previous operation result can include the previous membrane potential value of the SNN, and the input control circuit can include an additional input port configured to send the processed value of the previous membrane potential value or the bias value of each column in the plurality of columns to an additional row of each in the memory unit according to which mode the operation mode is in.
[0011] The first mode can be used for the SNN. The processed value of the previous membrane potential value can be the arithmetic negation of the previous membrane potential value fed back from the post-arithmetic circuit, and the additional input port can be configured to send the processed value to an additional row of each in the memory unit based on the operation mode being in the first mode.
[0012] The second mode can be used for the NN, and the additional input port can be configured to send the bias value of each column in the plurality of columns to an additional row of each in the memory unit based on the operation mode being in the second mode.
[0013] The additional row can be configured to store the processed previous membrane potential value based on the operation mode being in the first mode, and to store the bias value of each column in the plurality of columns based on the operation mode being in the second mode.
[0014] The crossbar switch array can be configured to store the result of adding (i) a first multiplication operation result and (ii) a second multiplication operation result by an adder tree. The first multiplication operation result is obtained by adding each product between the input signal and the weights stored in the memory unit, and the second multiplication operation result is obtained by multiplying the predefined mode by the value stored in the additional row.
[0015] The post-arithmetic circuit can include: a first shifter configured to adjust the operation result of the adder tree by a right shift operation based on the operation mode being in the first mode; a second shifter configured to adjust the value stored in the accumulator by a left shift operation based on the operation mode being in the second mode; and an accumulator.
[0016] The post arithmetic circuit may be configured to, based on the operation mode being in the first mode: perform a right shift operation on the operation result of adding the pulse signal and the weight by an adder tree through a first shifter; transmit the membrane potential value stored in an additional row through a second shifter; and store the result of the right shift operation and the transmitted membrane potential value in an accumulator.
[0017] The post arithmetic circuit may be configured to, based on the operation mode being in the second mode: transmit the result of adding (i) a first multiplication operation between an input signal and a weight stored in a memory cell and (ii) a second multiplication operation between a value of a predefined pattern and a bias value of each of the plurality of columns stored in an additional row to an accumulator through a first shifter; perform a left shift operation on the operation result corresponding to the input signal serially applied bit by bit in the accumulator through a second shifter; and accumulate the result of the left shift operation by the accumulator to generate a multi-bit value.
[0018] The adder tree may be configured to, for each operation, for each of the plurality of columns: simultaneously perform (i) a first multiplication operation between an input signal and a weight stored in a memory cell, and (ii) a second multiplication operation between the weight and a previous operation result.
[0019] The IMC macro may be integrated in at least one of the following devices: a mobile device, a mobile computing device, a mobile phone, a smart phone, a personal digital assistant (PDA), a fixed location terminal, a tablet computer, a computer, a wearable device, a laptop computer, a server, a music player, a video player, an entertainment unit, a navigation device, a communication device, a global positioning system (GPS) device, a television (TV), a tuner, a satellite radio, a song player, a digital video player, a digital video disc (DVD) player, a vehicle, a component of a vehicle, an avionics system, a drone, a multi-axis aircraft, and a medical device.
[0020] In another general aspect, there is a method of operating an in-memory computing (IMC) macro having an operating mode that can alternate between a first mode and a second mode, and the method includes: sending, depending on which mode the operating mode is in, a result of applying a predefined mode to an input signal or a previous membrane potential value as feedback; storing weights corresponding to the input signal in a row of a memory cell and processing and storing the previous membrane potential value of the feedback in an additional row of the memory cell; adding (i) a first operation result between the input signal and the weights and (ii) a second operation result between the predefined mode and the previous membrane potential value of the feedback by an adder tree; and selectively performing a first operation corresponding to a spiking neural network (SNN) or a second operation corresponding to an artificial neural network (ANN), where which operation is performed depends on which mode the operating mode is in.
[0021] The step of sending may include: sending, depending on which mode the operating mode is in, a processed value of the previous membrane potential value or a bias value of each of the plurality of columns to an additional row of each of the memory cells.
[0022] The first mode may be used for an SNN, and the second mode may be used for an NN, and the step of sending to the additional row may include: based on the operating mode being in the first mode, updating the additional row by sending the processed value to an additional row of each of the memory cells, where the processed value is obtained by arithmetically inverting the previous membrane potential value; and based on the operating mode being in the second mode, updating the additional row by sending the bias value of each of the plurality of columns of the memory cell to an additional row of each of the memory cells.
[0023] The first mode may be used for an SNN, and the step of selectively performing may include: based on the operating mode being in the first mode, performing a right shift operation on an operation result added by the adder tree between a pulse signal and the weights by a first shifter; passing, by a second shifter, a membrane potential value stored in the additional row; and storing the result of the right shift operation and the passed membrane potential value in an accumulator.
[0024] The second mode may be used for an NN, and the step of selectively performing may include: based on the operating mode being in the second mode: passing, by a first shifter, a result of adding (i) a first multiplication operation between the input signal and the weights stored in the memory cell and (ii) a second multiplication operation between a value of the predefined mode and a bias value of each of the plurality of columns stored in the additional row to the accumulator; performing a left shift operation on the operation result corresponding to the input signal applied bit-serial in the accumulator by a second shifter; and accumulating the result of the left shift operation by the accumulator to generate a multi-bit.
[0025] In another general aspect, a in-memory computing (IMC) device is provided. The IMC device has an operating mode that alternates between a first mode and a second mode, and the IMC device includes: an input control circuit configured to apply a predefined pattern to an input signal and / or send a processed value of a previous operation result as feedback according to the operating mode; a crossbar array including a plurality of columns, each of the plurality of columns including a memory cell and an adder tree, and the memory cell including an additional row for storing a processed value of a previous operation result or a bias value; and a plurality of post-arithmetic circuits corresponding to the plurality of columns, and for each of the plurality of columns, a corresponding post-arithmetic circuit is configured to: perform a first operation corresponding to a spiking neural network (SNN) or a second operation corresponding to a neural network (NN) different from the SNN according to the operating mode.
[0026] For each of the plurality of columns, a corresponding memory cell may include a row for storing weights corresponding to the input signal, and a corresponding adder tree may be configured to add the result of a first multiplication operation between the input signal and the weights and the result of a second multiplication operation between the predefined pattern and the processed value of the previous operation result or the bias value according to the operating mode.
[0027] The input signal may include a spike signal for the SNN or a feature map for the NN.
[0028] According to a command sent from a host, the operating mode may be set to the first mode for the SNN or the second mode for the NN.
[0029] The first mode may be used for the SNN and the second mode may be used for the NN, and the input control circuit may further be configured to: based on the operating mode being in the first mode, set the predefined pattern to 1; and based on the operating mode being in the second mode, set the predefined pattern to a pattern corresponding to the number of bits of the input signal.
[0030] The previous operation result may include a previous membrane potential value of the SNN, and the input control circuit may include an additional input port configured to: according to the operating mode, for each of the plurality of columns, send a processed value of the previous membrane potential value or the bias value to a corresponding additional row of a corresponding memory cell.
[0031] The first mode may be used for the SNN, and the additional input port may be configured to: based on the operating mode being in the first mode, for each of the plurality of columns, send a processed value of the previous membrane potential value to a corresponding additional row of a corresponding memory cell, where the processed value of the previous membrane potential value is the arithmetic negation of the previous membrane potential value fed back from a corresponding post-arithmetic circuit.
[0032] The second mode can be used for the NN, and the additional input port can be configured to: based on the operation mode being in the second mode, for each of the plurality of columns, send a bias value to a corresponding additional row of a corresponding memory cell.
[0033] For each of the plurality of columns, a corresponding additional row can be configured to: based on the operation mode being in the first mode, store a processed value of a previous membrane potential value; and based on the operation mode being in the second mode, store a bias value.
[0034] The crossbar array can be configured to: for each of the plurality of columns, store a result of adding a result of a first multiplication operation and a result of a second multiplication operation by a corresponding adder tree, the result of the first multiplication operation being obtained by adding respective products between an input signal and weights stored in corresponding memory cells, the result of the second multiplication operation being obtained by multiplying a predefined pattern by a value stored in an additional row.
[0035] Each column of the plurality of post-arithmetic circuits can include: a first shifter configured to, based on the operation mode being in the first mode, adjust an operation result of a corresponding adder tree by a right shift operation; a second shifter configured to, based on the operation mode being in the second mode, adjust a value stored in an accumulator by a left shift operation; and an accumulator.
[0036] For each of the plurality of columns, a corresponding post-arithmetic circuit can be configured, based on the operation mode being in the first mode: by the first shifter, perform a right shift operation on an operation result added by a corresponding adder tree between a pulse signal and weights; by the second shifter, transfer a membrane potential value stored in the accumulator; and store the result of the right shift operation and the transferred membrane potential value in the accumulator.
[0037] For each of the plurality of columns, a corresponding post-arithmetic circuit can be configured, based on the operation mode being in the second mode: by the first shifter, transfer a result of adding a result of a first multiplication operation between an input signal and weights stored in a corresponding memory cell and a result of a second multiplication operation between a predefined pattern and a bias value stored in a corresponding additional row to the accumulator; by the second shifter, perform a left shift operation on an operation result corresponding to the input signal serially applied bit by bit in the accumulator; and accumulate the result of the left shift operation by the accumulator to generate a multi-bit result.
[0038] For each of the plurality of columns, a corresponding adder tree can be configured to: simultaneously perform a first multiplication operation between an input signal and weights stored in a corresponding memory cell, and a second multiplication operation between a predefined pattern and a processed value or a bias value of a previous operation result.
[0039] The IMC device may be integrated in at least one of the following devices: a mobile device, a mobile computing device, a mobile phone, a smart phone, a personal digital assistant (PDA), a fixed-position terminal, a tablet computer, a computer, a wearable device, a laptop computer, a server, a music player, a video player, an entertainment unit, a navigation device, a communication device, a global positioning system (GPS) device, a television (TV), a tuner, a satellite radio device, a song player, a digital video player, a digital video disc (DVD) player, a vehicle, a component of a vehicle, an avionics system, a drone, a multi-axis aircraft, and a medical device.
[0040] In another general aspect, a method of operating a in-memory computing (IMC) device is provided, the IMC device having an operating mode that alternates between a first mode and a second mode, and the method includes: applying a predefined mode to an input signal and / or sending a processed value of a previous membrane potential value as feedback according to the operating mode; storing weights corresponding to the input signal in a row of a memory cell, and storing a bias value of a column corresponding to the memory cell or the processed value of the previous membrane potential value in an additional row of the memory cell; adding a result of a first multiplication operation between the input signal and the weights and a result of a second multiplication operation between the predefined mode and the processed value of the previous membrane potential value or the bias value through an adder tree corresponding to the memory cell; and selectively performing a first operation corresponding to a spiking neural network (SNN) or a second operation corresponding to a neural network (NN) different from the SNN according to the operating mode.
[0041] The step of sending may include: sending the processed value of the previous membrane potential value or the bias value to an additional row of the memory cell according to the operating mode.
[0042] The first mode may be used for the SNN and the second mode may be used for the NN, and the step of sending to the additional row of the memory cell may include: updating the additional row of the memory cell by sending the processed value of the previous membrane potential value to the additional row of the memory cell based on the operating mode being in the first mode, wherein the processed value of the previous membrane potential value is obtained by arithmetically inverting the previous membrane potential value; and updating the additional row of the memory cell by sending the bias value to the additional row of the memory cell based on the operating mode being in the second mode.
[0043] The first mode can be used for the SNN, and the selectively executed steps may include: based on the operation mode being in the first mode, performing a right shift operation on the operation result of adding the pulse signal and the weight by the adder tree by a first shifter; passing the membrane potential value stored in the accumulator by a second shifter; and storing the result of the right shift operation and the passed membrane potential value in the accumulator.
[0044] The second mode can be used for the NN, and the selectively executed steps may include: based on the operation mode being in the second mode, passing the result of adding the result of the first multiplication operation between the input signal and the weight stored in the memory cell and the result of the second multiplication operation between the predefined pattern and the bias value stored in the additional row to the accumulator by a first shifter; performing a left shift operation on the operation result corresponding to the input signal serially applied to the bits in the accumulator by a second shifter; and accumulating the result of the left shift operation by the accumulator to generate a multi-bit result.
[0045] According to the following detailed description, drawings, and claims, other features and aspects will be clear. Description of the Drawings
[0046] Figure 1A An example embodiment of an in-memory computing (IMC) system that performs multiplication-accumulation (MAC) operations of a neural network according to one or more example embodiments is shown.
[0047] Figure 1B An example structure of a neural network according to one or more example embodiments is shown.
[0048] Figure 1C An example operation of a spiking neural network (SNN) and a neural network (NN) according to one or more example embodiments is shown.
[0049] Figure 2 An example IMC macro according to one or more example embodiments is shown.
[0050] Figure 3 An example structure and operation of an IMC macro according to one or more example embodiments are shown.
[0051] Figure 4A and Figure 4B An example operation of a post-arithmetic circuit in which the operation depends on the operation mode according to one or more example embodiments is shown.
[0052] Figure 5 An example flow of the operation of an IMC macro according to one or more example embodiments is shown.
[0053] Figure 6Shows an example process of the operation of an IMC macro according to one or more example embodiments.
[0054] Figure 7 Shows an example electronic system including an IMC macro according to one or more example embodiments.
[0055] Throughout the drawings and the detailed description, unless otherwise described or provided, the same or similar reference numerals can be understood to represent the same or similar elements, features, and structures. The drawings may not be to scale, and for clarity, illustration, and convenience, the relative sizes, proportions, and depictions of elements in the drawings may be exaggerated. Detailed Description
[0056] The following detailed description is provided to assist the reader in obtaining a comprehensive understanding of the methods, devices, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, devices, and / or systems described herein will be apparent after understanding the disclosure of this application. For example, the order of operations described herein is merely an example and is not limited to those set forth herein, but may be changed as will be apparent after understanding the disclosure of this application, except for operations that must occur in a specific order. Additionally, descriptions of known features may be omitted for greater clarity and conciseness after understanding the disclosure of this application.
[0057] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. Instead, the examples described herein have been provided only to illustrate some of the many possible ways of implementing the methods, devices, and / or systems described herein that will be apparent after understanding the disclosure of this application.
[0058] The terms used herein are for describing various examples only and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more of them. As a non-limiting example, the terms "comprising," "including," and "having" specify the presence of the stated features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.
[0059] Throughout the specification, when an element or component is described as being "connected to", "coupled to", or "joined to" another element or component, the element or component can be directly "connected to", "coupled to", or "joined to" the other element or component, or one or more other elements or components can reasonably be present therebetween. When an element or component is described as being "directly connected to", "directly coupled to", or "directly joined to" another element or component, no other elements or components can be present therebetween. Similarly, phrases such as "between" and "immediately between" and "adjacent to" and "immediately adjacent to" can be interpreted as previously described.
[0060] Although terms such as "first", "second", and "third", or A, B, (a), (b), etc. may be used herein to describe various elements, components, regions, layers, or sections, these elements, components, regions, layers, or sections should not be limited by these terms. Each of these terms is not used to define, for example, the nature, order, or sequence of the corresponding element, component, region, layer, or section, but only to distinguish the corresponding element, component, region, layer, or section from other elements, components, regions, layers, or sections. Thus, a first element, first component, first region, first layer, or first section as referred to in the examples described herein can also be referred to as a second element, second component, second region, second layer, or second section without departing from the teachings of the examples.
[0061] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains and based on an understanding of the disclosure of this application. Terms such as those defined in a general dictionary should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the disclosure of this application, and should not be interpreted in an idealized or overly formal sense, unless explicitly so defined herein. The use of the term "can" with respect to an example or embodiment herein (e.g., with respect to what an example or embodiment can include or achieve) means that there is at least one example or embodiment that includes or achieves such a feature, and not all examples are so limited.
[0062] Figure 1A An example implementation of an in-memory computing (IMC) system that performs multiply-accumulate (MAC) operations of a neural network according to one or more example embodiments is shown. Referring to Figure 1A , an example structure of the IMC system 100 is shown.
[0063] In a computing device using a von Neumann architecture, due to the frequent data movement between the arithmetic unit part (e.g., the main processor) and the memory part, there may be limitations in performance and power. The exchange of data between the arithmetic unit part and the memory part usually becomes a bottleneck, where the exchange of data cannot keep up with the pace of computing operations. IMC is a computing architecture for directly performing computing operations (e.g., MAC operations) on the data in the memory that stores the data. IMC can be provided to overcome such limitations in performance and power. Since the operations are executed inside the memory, one or a limited number of basic operations can be executed instead of various operations. IMC can reduce the frequency of data movement between the processor 120 and the memory device 110, and can increase power efficiency. For most IMC devices, the data affected by the operations remains stored in the IMC device before, during, and after the operations. In addition, although the IMC device can perform in-memory operations (logical operations / mathematical operations), the IMC device can also act as a memory device. For example, the IMC device can work in a typical manner of a memory device (e.g., having a similar interface, addressing scheme, etc.).
[0064] When the host (e.g., the processor 120) that incorporates or controls the IMC system 100 inputs data (to be computed) into the memory device 110, the memory device 110 can perform operations (or computations) on the data by itself. The processor 120 can read the result of the operations from the memory device 110. Therefore, the data movement or data transmission during such a computing process can be minimized.
[0065] For example, the IMC system 100 can perform MAC operations that are frequently used in artificial intelligence (AI) algorithms and various other types of operations.
[0066] The neural network 130 can be an overall model in which nodes form a network through connections therebetween. The neural network 130 can have the ability to solve problems by changing the strength / weight of the connections through learning. The neural network 130 can include one or more layers containing nodes, and each layer is connected to another layer. The nodes in the neural network 130 can include a combination of weights or biases. The way the neural network 130 infers (predicts) results from any input can be changed by changing the weights of the nodes through learning. As Figure 1A shown, for a given layer containing a given node, the computing operations between the layers in the neural network 130 can include a MAC operation of adding the results of "multiplying each of the input values of the given node by the corresponding weight of the given node". For example, referring to Figure 1A , in the neural network 130, four nodes i0, i1, i2, and i3 in the i-th layer can be connected to three nodes o0, o1, and o2 in the o-th layer, and the weights between the node o0 and the nodes i0, i1, i2, and i3 can be w00 , w 10 , w 20 and w 30 , where, o0 = i0 × w 00 + i1 × w 10 + i2 × w 20 + i3 × w 30 . The neural network 130 can be a deep neural network (DNN). As a non-limiting example, the neural network 130 can be / include a spiking neural network (SNN) and a neural network (NN) (such as, a convolutional neural network (CNN), a recurrent neural network (RNN), a perceptron, a multi-layer perceptron, a feed-forward (FF) network, a radial basis function (RBF) network, a deep feed-forward (DFF) network, a long short-term memory (LSTM), a gated recurrent unit (GRU), an autoencoder (AE), a variational autoencoder (VAE), a denoising autoencoder (DAE), a sparse autoencoder (SAE), a Markov chain (MC), a Hopfield network (HN), a Boltzmann machine (BM), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a deep convolutional network (DCN), a deconvolution network (DN), a deep convolutional inverse graphics network (DC-IGN), a generative adversarial network (GAN), a liquid state machine (LSM), an extreme learning machine (ELM), an echo state network (ESN), a deep residual network (DRN), a differentiable neural computer (DNC), a neural Turing machine (NTM), a capsule network (CN), a Kohonen network (KN), or an attention network (AN)) or a combination of a spiking neural network (SNN) and a neural network (NN).
[0067] For example, a MAC operation that repeats multiplication and addition operations can be represented by Equation 1.
[0068] Equation 1
[0069] In Equation 1, the value of the node in the (n + 1)-th layer can be calculated by applying an appropriate activation function f() to the sum of the product of the value of the node in the n-th (previous) layer and the weights mapped to it. For example, can represent the value of the j-th node in the (n + 1)-th layer, can represent the weight between the i-th node in the n-th layer and the j-th node in the (n + 1)-th layer, and can represent the value of the i-th node in the n-th layer, where the parameter n can represent the index of the layer in the neural network, the parameter m can represent the number of nodes in the n-th layer, the parameter i can represent the index of the node in the n-th layer, and the parameter j can represent the index of the node in the (n + 1)-th layer. The MAC operation can be performed by applying the remaining data to store the input xn,i or weight w n,i,j is executed by a memory.
[0070] In one example embodiment, the memory device 110 of the IMC system 100 may perform the above-described MAC operations and / or vector matrix multiplication (VMM) operations. The memory device 110 may include an IMC macro that performs MAC operations and / or VMM operations. The memory device 110 may also be referred to as a "memory array" or an "IMC device".
[0071] In addition to performing MAC operations and / or VMM operations, the memory device 110 may be used as a memory for storing data, and the memory device 110 may be used to drive algorithms including multiplication operations. (Although in some cases, operand inputs may be input to the memory device 110) The memory device 110 may directly perform logical / mathematical operations within the memory without data movement or transfer, thereby reducing data movement or transfer while improving power efficiency.
[0072] Figure 1B illustrates an example structure of a neural network according to one or more example embodiments. Referring to Figure 1B , an example structure of a neural network 130 in which the SNN 140 and the NN 150 are combined is shown. Although a spiking neural network is technically a type of neural network, the NN 150 may be considered a non-SNN neural network.
[0073] Since the SNN 140 and the NN 150 (which is not a spiking NN) are combined, the neural network 130 may also be referred to as a "hybrid neural network" 130.
[0074] The hybrid neural network 130 may provide both the characteristics of the SNN 140 (e.g., low operations and low power) and the characteristics of the NN 150 (e.g., high performance), and thus may correspond to a new network structure in which the SNN 140 and the NN 150 are combined in a hybrid manner.
[0075] The memory device 110 may perform MAC operations and / or VMM operations with low power and may be designed to be used in a deep learning-based NN. Generally, in order to improve power efficiency, it may be desirable to configure the memory device 110 having a greater depth in the input channel direction with a wide adder tree, thereby allowing more efficient processing of VMM operations for computing large weight matrices.
[0076] In an example embodiment, the static random access memory (SRAM) IMC macro (of the hybrid neural network 130) has a structure that uses low power to compute an input signal applied bit-serially and multiple-bit weights. For the SNN 140, the SRAM IMC can store membrane potential and control leakage voltage, and at the same time, for the NN 150, the SRAM IMC can support the VMM operation between an input signal (input operand) and a weight matrix (stored operand).
[0077] Refer to Figure 1C Describe the operations of the SNN 140 and the operations of the NN 150. In addition, refer to Figures 2 to 6 Describe the structure and operations of the IMC macro for supporting the operations of both the SNN 140 and the NN 150.
[0078] Figure 1C Illustrate example operations of the SNN and example operations of the NN according to one or more example embodiments. Refer to Figure 1C ,, what is shown are the operations of the SNN 140 and the operations of the NN 150.
[0079] The SNN 140 can be configured such that the concept of time is included in the interaction between nodes, and the interaction is generally referred to as a spike. In the SNN 140, the internal state of a node (i.e., a neuron) can be changed by time information and spike signals sent from other nodes. For example, refer to Figure 1C ,, a node of the SNN 140 can receive inputs x1, x2, x3, and x4 from other nodes, and weights W 1,i , W i,i , W k,i and W j,i can correspond to the inputs x1, x2, x3, and x4 respectively, and the internal state of the node can be based on the inputs x1, x2, x3, and x4 and the weights W 1,i , W i,i , W k,i and W j,i and be changed. When the changed internal state of the node satisfies a specific condition due to an incoming spike, the node can generate its own spike.
[0080] For example, when a first node (first neuron) is upstream of a second node (second neuron) and is connected to the second node (second neuron), the first node can send information to the second node. For example, (as Figure 1CAs shown in [Figure X], the first node can sequentially send pulse signals to the second node three times along the time axis. In this case, whenever the second neuron receives a pulse signal (e.g., the first pulse signal and the second pulse signal) from the first node, the action potential of the second neuron can rise to a specific value, and this rise gradually decays due to leakage current (until reinforced by another pulse).
[0081] When the third pulse signal is sent to the second node and the action potential of the second node thus exceeds the threshold voltage u th , an output pulse signal "1" can be generated from the second node. In combination with the generation of the output pulse signal, the value of the action potential of the second node can be set to "0". Even when multiple nodes are connected (e.g., when there are multiple first nodes connected to the second node), the foregoing process can be similarly applied.
[0082] For understanding, an analogy with biological neurons can be helpful. The inside and outside of the cell body of a neuron can be separated by a cell membrane (cell wall), which can have a membrane potential specific to the cell. The cell membrane can be modeled as the leaky integrate-and-fire (LIF) neuron model shown in the upper half of Figure 1C [Figure X].
[0083] The LIF neuron model can model the following rules of a neuron.
[0084] (i)The LIF neuron model can calculate the sum of the pulses of the presynaptic neuron (the first node). In this case, the pulses of the presynaptic neuron can be considered as electric power from the outside and can correspond to the power source of the neuron.
[0085] (ii)The LIF neuron model can generate an output pulse signal when the membrane potential U exceeds its threshold voltage u th and can be initialized. A neuron can store sodium ions inside the neuron through the action potential sent to the presynaptic neuron. Such a characteristic can be modeled as a capacitor C that temporarily stores electric power. In addition, the membrane potential can have a voltage that increases due to the action potential and returns to the reset voltage over time when ions escape through the (cell) membrane, and this behavior can be modeled as a resistor R.
[0086] Based on the foregoing, the LIF neuron model can be implemented as a resistor-capacitor (RC) circuit. The cell membrane can be represented as the capacitor C of the RC circuit, and the potential difference between the two ends of the battery can be represented as the membrane potential. When an external current I is input to the RC circuit, the capacitor C corresponding to the battery can be charged. In this case, the presynaptic neuron can receive an input pulse signal, and when the action potential (membrane potential) of the postsynaptic neuron exceeds the threshold voltage u as the input pulse signals are accumulated thWhen this occurs, the postsynaptic neuron can generate an output pulse. The postsynaptic neuron that generates the output pulse can recover after experiencing a refractory period. The "refractory period" described herein can be a period during which the initialization / reset state is briefly maintained immediately after the pulse generation.
[0087] (iii) In the LIF neuron model, the voltage of the membrane potential can leak continuously (i.e., the membrane can gradually dissipate the voltage).
[0088] Turning to NN 150, NN 150 can have a network structure in which nodes are connected by links with weights (each link having its own weight). NN 150 can include, for example, an input layer that receives input signals, an output layer that outputs the results processed by the hidden layer, and a hidden layer that is provided between the input layer and the output layer and is generally not exposed to the outside (this is not a strict requirement). There can be multiple hidden layers. The input data received by the input layer can be processed by the hidden layer and then output through the output layer.
[0089] The input nodes included in the input layer can send the input data as it is to the hidden layer without any special operations, and thus the input nodes can correspond to the input values themselves. The nodes in the hidden layer and the output layer can perform specific operations on the received input data.
[0090] The nodes in the layers other than the input layer can receive their input values through links / connections, calculate the weighted sum, and generate an output signal by applying an activation function (etc.) to the weighted sum. (In the case of the output layer nodes) the output signal can be the final output value, or the output signal can be the input value of another node. In this case, the activation function (etc.) can determine whether the node is activated. When the weighted sum is greater than or equal to the threshold of the activation function, the node can be activated, and when the weighted sum is less than the threshold, the node can not be activated.
[0091] The weighted sum (e.g., Y) can be the multiplication operation and repeated addition operation between the inputs (e.g., I1, I k and I j ) and the corresponding weights (e.g., W 1,i , W k,i and W j,i ), and can also be referred to as "MAC operation". Since the MAC operation is performed using a memory that combines the computing operation function, the circuit that performs the MAC operation can also be referred to as an IMC circuit.
[0092] Figure 2 An example IMC macro according to one or more example embodiments is shown. Refer to Figure 2, the IMC macro 200 may include an input control circuit 210, a crossbar array 230, and a post-arithmetic circuit 250, the details of which will become clear as other figures are discussed.
[0093] The IMC macro 200 may be configured based on, for example, SRAM and may perform digital-based MAC operations and / or VMM operations. In terms of SRAM, the cells storing bits may have SRAM characteristics.
[0094] The IMC macro 200 may operate as a hybrid network in which SNN and NN are combined. For example, the IMC macro 200 may set the operation mode to a first mode for SNN or a second mode for NN according to a command sent from a host. The host may repeatedly switch the IMC macro 200 back and forth between these two modes (which may also be referred to as network modes). The "first mode" may represent an operation mode in which "the operation of SNN (rather than NN) is active", and the "second mode" may represent an operation mode in which "the operation of NN (rather than SNN) is active".
[0095] According to the operation mode, the input control circuit 210 may (i) apply a predefined pattern to the input signal (e.g., the input signal used as an input operand for MAC / VMM operations), and / or (ii) send (feedback) the processed value of the previous operation result of the crossbar array 230 to the crossbar array 230. The input signal may be, for example, a pulse signal for SNN or a feature map for NN. The previous operation result may be, for example, the previous membrane potential value of SNN.
[0096] When the operation mode is set to the first mode for SNN, the input control circuit 210 may set the predefined pattern to "1". When the operation mode is set to the second mode for NN, the input control circuit 210 may set the predefined pattern to a pattern or bias value corresponding to the number of bits of the input signal. For example, as a non-limiting example, when the number of bits of the input signal is 4, the predefined pattern may be "0001".
[0097] The input control circuit 210 may include (also shown in Figure 3 ), an additional input port 215, which is configured to send (i) the processed value of the previous membrane potential value or (ii) the bias value of each column in a set of columns to an additional row of the corresponding memory cell 231 according to the operation mode (e.g., Figure 3Additional rows 310-1, 310-2, ……, and 310-M). The bias value for each of the set of columns can be, for example, a static bias value that can vary between columns. Since the additional rows and additional input ports were not found in previous IMC devices (and the other components mentioned herein can also be new), the additional rows and additional input ports are referred to as "additional".
[0098] When the operation mode is in the first mode for the SNN, the additional input port 215 can receive the processed value (e.g., -U1(t), -U2(t), ……, -U m (t)) obtained by multiplying the previous membrane potential values 305 (e.g., U1(t), U2(t), ……, U m (t)) fed back from the post arithmetic circuit 250 to the input control circuit 210 by -1 from the input control circuit 210 and send it to the additional rows of the corresponding memory cells 231. When the operation mode is in the second mode for the NN, the additional input port 215 can send the bias value of the corresponding column to the corresponding additional rows in the memory cells 231. For example, when the operation mode is in the first mode, the additional rows can store the processed previous membrane potential values. And when the operation mode is in the second mode, the additional rows can store the bias values of the corresponding columns.
[0099] In one example embodiment, the IMC macro 200 can include memory cells 231 configured in the form of a crossbar array 230. The memory cells 231 can include word lines, memory cells (i.e., bit cells), and bit lines. The word lines can be used to receive input data or input signals of a neural network (e.g., Figure 1A the neural network 130 in). For example, when there are N word lines, the values corresponding to the input signals of the neural network can be applied to the N word lines.
[0100] For example, the crossbar array 230 can perform multiplication operations (e.g., VMM operations) between a single vector and a matrix in a number of cycles, and this operation can be used in both the SNN (configured to perform pulse-based signal processing) and the NN, where the NN can be a discrete domain (non-pulse) neural network (e.g., CNN, RNN, or LSTM, just listing some examples of digital neural network architectures).
[0101] The crossbar array 230 can include memory cells 231 and an adder tree 235 corresponding to the memory cells 231, and each memory cell 231 includes at least one additional row that stores the result of processing the feedback of the previous operation result (e.g., the previous membrane potential value) or the bias value of the corresponding column.
[0102] As a non-limiting example, the memory cell 231 may include at least one of a diode, a transistor (e.g., a metal-oxide-semiconductor field-effect transistor (MOSFET)), a SRAM bit cell, and a resistive memory. Hereinafter, the SRAM memory cell will be used as an example to describe the memory cell 231, but the example need not be limited thereto.
[0103] The memory cell 231 may include rows storing weights (e.g., stored operands) corresponding to input signals (e.g., input operands). The memory cell 231 may be, for example, a SRAM memory array. The adder tree 235 may be, for example, a digital adder tree. Although the examples described herein represent weights stored in the memory cell 231, the examples and embodiments described herein are not limited to any particular type of application or data.
[0104] The crossbar array 230 may receive input signals or input data (input operands) applied serially by bits, perform a multiplication operation between the multi-bit weights stored in the memory cell 231 and a one-bit input signal, and may add the results of the multiplication operation through the adder tree 235. The result of the addition performed by the adder tree 235 obtained in each cycle may be output as a final operation result through the accumulator 255 of the post-arithmetic circuit 250.
[0105] The adder tree 235 may add “(i) the result of a first multiplication operation between an input signal and a weight” and “(ii) the result of a second multiplication operation between a predefined pattern and a processed previous operation result or a bias value”. In each operation, for each in a column, the adder tree 235 may simultaneously perform “(i) a first multiplication operation between an input signal and a weight stored in the memory cell 231” and “(ii) a second multiplication operation between a predefined pattern and a processed previous operation result or a bias value”.
[0106] The crossbar array 230 may store the result of adding, by the adder tree 235, “(i) a first multiplication operation result obtained by adding respective products (multiplications) between the weights stored in the memory cell 231 and an input signal” and “(ii) a second multiplication operation result obtained by multiplying a predefined pattern by a value stored in an additional row”. In this case, the respective products between the input signal and the weights may be added within a column. In addition, a pattern input (as an input of a predefined pattern) may be multiplied by data stored in an additional row, and the result of the multiplication may also be added in the same column. Subsequently, the above two results may be added again by the adder tree 235 within the same column. Briefly, the foregoing process may involve multiplication of {input, pattern} by {weights, additional row}, and after each multiplication, the results may be added in the adder tree 235.
[0107] The post arithmetic circuit 250 may selectively perform a first operation corresponding to the SNN or a second operation corresponding to the NN according to the operation mode.
[0108] The post arithmetic circuit 250 may include, for example, a first shifter 251, a second shifter 253, and an accumulator 255. When the operation mode is in the first mode, the first shifter 251 may adjust the operation result of the adder tree 235 through a right shift operation. When the operation mode is the second mode, the second shifter 253 may adjust the value stored in the accumulator 255 through a left shift operation.
[0109] When the operation mode is in the first mode, the accumulator 255 may store the membrane potential value. When the operation mode is in the second mode, the accumulator 255 may convert the bit-serial calculation result into a multi-bit calculation result.
[0110] For example, when the operation mode is in the first mode, the post arithmetic circuit 250 may perform a right shift operation on the operation result of the operation between the pulse signal added by the adder tree 235 and the weight through the first shifter 251. The post arithmetic circuit 250 may bypass the membrane potential value stored in the accumulator 255 through the second shifter 253 (that is, the second shifter 253 may not perform any operation on the membrane potential value stored in the accumulator 255 and may transfer the membrane potential value). The post arithmetic circuit 250 may store the result of the right shift operation and the bypassed membrane potential value in the accumulator 255.
[0111] Optionally, when the operation mode is in the second mode, the post arithmetic circuit 250 may bypass the result through the first shifter 251 and enter the accumulator 255, where the result is the result of adding "(i) the first multiplication operation between the weight stored in the memory cell 231 and the input signal" and "(ii) the second multiplication operation between the predefined pattern and the bias value stored in the additional row". The post arithmetic circuit 250 may perform a left shift operation on the operation result of the accumulator 255 corresponding to the input signal applied bit-serial through the second shifter 253. The post arithmetic circuit 250 may generate a multi-bit result by accumulating the result of the left shift operation via the accumulator 255.
[0112] The IMC macro 200 can be integrated into at least one device (e.g., a mobile device, a mobile computing device, a mobile phone, a smart phone, a personal digital assistant (PDA), a fixed - location terminal, a tablet computer, a computer, a wearable device, a laptop computer, a server, a music player, a video player, an entertainment unit, a navigation device, a communication device, a global positioning system (GPS) device, a television (TV), a tuner, a satellite radio device, a song player, a digital video player, a digital video disc (DVD) player, a vehicle, a part of a vehicle, an avionics system, a drone, a multi - axis aircraft, or a medical device).
[0113] In one example embodiment, using the IMC macro 200 can allow for selectively performing operations for SNNs and operations for NNs without additional hardware configuration, thereby improving the operational efficiency of SNNs and NNs in terms of power, hardware, and / or performance.
[0114] As a non - limiting example, the IMC macro 200 can be implemented as a neural network device, an IMC circuit, or a MAC arithmetic circuit and / or a MAC arithmetic device.
[0115] Figure 3 Shows an example structure and operation of an IMC macro according to one or more example embodiments. Referring to Figure 3 , shown is an example structure of the SRAM IMC macro 300.
[0116] The input control circuit 210 can receive an external input signal 301 (e.g., X1, X2, ……, X N ). For example, when the operation mode is in the first mode, the input signal 301 can be in the form of a pulse signal of a node in the previous layer. When the operation mode is in the second mode, the input signal 301 can be in the form of a feature map (non - pulse signal).
[0117] The input control circuit 210 can include an additional input port 215 that receives a static bias value or a previous membrane potential value. The input control circuit 210 can apply a predefined pattern (indicated as PP in some figures) or an idle counter value sent through the additional input port 215 to the input signal 301 to generate a result, and send the generated result as an input to the memory unit 231. The memory unit 231 can include bit cells corresponding to memory banks and an arithmetic circuit that outputs a signal corresponding to the arithmetic result, and the arithmetic result corresponds to each of the bit cells. The memory unit 231 can be, for example, a unit of an SRAM memory array.
[0118] In this case, the predefined pattern can be hard - wired into the hardware of the SRAM IMC macro 300 or can be set in a register before run - time.
[0119] The SRAM IMC macro 300 may include, for each column, an arithmetic module 320 as a crossbar array 230. The arithmetic module 320 includes memory cells 231 that implement multiplication and an adder tree 235 that adds all the arithmetic results and outputs the result of the addition.
[0120] For example, the crossbar array 230 may include M columns (e.g., column 1, column 2, ……, column M) that receive external inputs and q additional rows 310. Although the crossbar array 230 may have multiple additional rows 310, an example case where the number of additional rows 310 is one (i.e., q = 1) is described below, which is the most basic structure.
[0121] When the operation mode is in the first mode, the crossbar array 230 may have N rows that receive N corresponding input signals 301 (e.g., signals of presynaptic neurons) as inputs. The crossbar array 230 may also have an additional row 310 that receives A (here, "A" is variable) additional input signals 303 corresponding to a predefined pattern. In this case, the arithmetic results of the N input signals 301 and the A additional input signals 303 may be added by an adder tree 235 having a length of N + A.
[0122] For a previous arithmetic result (e.g., U(t), where t represents time), the input control circuit 210 may apply arithmetic negation to it (forming, for example, -U(t)), and may store the thus processed previous arithmetic result (e.g., -U(t)) in the additional row 310. For convenience, the processed previous arithmetic result -U(t) may be a result reflecting the leakage voltage generated from the SNN.
[0123] Each of the memory cells 231 in the crossbar array 230 may include at least one additional row 310 that stores the processed previous arithmetic value -U(t). As noted, the processed previous arithmetic value -U(t) is obtained by the input control circuit 210 performing arithmetic negation on the arithmetic result (e.g., the membrane potential value U(t)) received from the post-arithmetic circuit 250. For example, when the operation mode is in the first mode, the input control circuit 210 performs arithmetic negation on the arithmetic result received from the post-arithmetic circuit 250, and then may directly write the arithmetically negated arithmetic result onto the additional rows 310 (e.g., 310-1, 310-2, ……, and 310-M) of the corresponding memory cells 231. When the operation mode is in the second mode, the input control circuit 210 may store a fixed bias value in the additional row 310. The bias value stored in the additional row 310 may be added to subsequent arithmetic results later.
[0124] Each of the memory cells 231 in the crossbar array 230 may have an additional row 310 that stores -U(t+Δt) to which a previous operation result was applied (i.e., the data stored in the additional row 310 may be updated to -U(t+Δt)) for efficient computation of the SNN. The additional row 310 may be directly updated within the SRAM IMC macro 300.
[0125] The SRAM IMC macro 300 may add the corresponding outputs of the memory cells 231 via the adder tree 235 and output the final operation result via the post-arithmetic circuit 250.
[0126] The SRAM IMC macro 300 may store weights in the SRAM memory cells 231 and then apply the input signal 301 to perform an operation. Depending on whether the operation mode is in the first mode or the second mode, the SRAM IMC macro 300 may perform an operation by combining the input signal and a predefined mode.
[0127] For example, when the operation mode is in the first mode, the SRAM IMC macro 300 may perform a multiplication operation between (i) a pulse signal (which is the input signal) and (ii) the weights stored in the memory cells 231; the multiplication operation may be performed in the operation module 320 for each column. The SRAM IMC macro 300 may add the results of the multiplication operation via the adder tree 235 and send the result of the addition to the post-arithmetic circuit 250.
[0128] For example, when the operation mode is in the second mode, the SRAM IMC macro 300 may add via the adder tree 235 "the result of (i) applying a bias value stored in a predefined mode (e.g., PP(b)×B m ) to (ii) the multiplication operation result between the feature map value X(b) as the input signal 301 and the weights W m stored in the memory cells 231 (e.g., X(b)×W m ). The result of the addition may be sent to the post-arithmetic circuit 250. In this case, the bias value may vary for each column.
[0129] The adder tree 235 can add the multiplication operation results corresponding to the SRAM memory cells 231 respectively, and send the result of the addition to the post-arithmetic circuit 250. The post-arithmetic circuit 250 can perform an addition operation by performing bit-shifting on the addition operation result of the corresponding digit-by-digit numbers according to the operation mode. For example, when the operation mode is in the second mode, the post-arithmetic circuit 250 can combine "(i) the addition operation result of the subsequent digit-by-digit numbers" with "(ii) the addition operation result after bit-shifting", and accumulate the multiplication operation results digit by digit, and thus output a multi-bit result corresponding to the final MAC operation result.
[0130] In the case where the input control circuit 210 receives single-bit input data (such as a pulse signal), bit-shifting may not be required, and thus the post-arithmetic circuit 250 can directly output the addition operation result of the adder tree 235, or optionally store the addition operation result of the adder tree 235 in an output register (not shown). The final addition operation result (such as the MAC operation result) stored in the output register can be read by, for example, a processor (such as the processor 710 in Figure 7 and used for other calculation operations.
[0131] The post-arithmetic circuit 250 can finally combine the operation results output from each column to output the combined result as the MAC operation result.
[0132] The post-arithmetic circuit 250 can support both SNN and NN. The post-arithmetic circuit 250 can send the operation result U(t+Δt) to the input control circuit 210 to allow the operation result U(t+Δt) to be converted to -U(t+Δt) in the input control circuit 210, and can allow the input control circuit 210 to directly write -U(t+Δt) into the additional row 310.
[0133] In each operation, for each column, the adder tree 235 in the SRAM IMC macro 300 can simultaneously add "(i) the operation result between N input signals (such as pulse signals) and N weights stored in the memory cell 231" and "(ii) the operation result of the previous membrane potential value U(t+Δt)".
[0134] The post-arithmetic circuit 250 can use Figure 4A and Figure 4B the two shifters (such as the first shifter 251 and the second shifter 253) shown in according to the first mode and the second mode.
[0135] For example, when the operation mode is in the first mode, the post-arithmetic circuit 250 may send the operation result of the adder tree 235 to the first shifter 251, and accumulate the result of the right-shift operation performed by the first shifter 251 in the accumulator 255.
[0136] When the operation mode is in the second mode, the post-arithmetic circuit 250 may send the operation result of the adder tree 235 to the second shifter 253, and accumulate the result of the left-shift operation performed by the second shifter 253 in the accumulator 255.
[0137] In one example embodiment, when using a single SRAM IMC macro 300, the SRAM IMC macro 300 may selectively operate the SNN and NN, and thus may improve the overall system power efficiency. In addition, effectively operating a hybrid neural network including the SNN and NN may contribute to effectively configuring a large-scale SNN system.
[0138] Figure 4A and Figure 4B illustrates an example operation of a post-arithmetic circuit in which the operation depends on the operation mode according to one or more example embodiments. In one example embodiment, the post-arithmetic circuit may include: a first shifter 251 configured to adjust the addition result obtained through the adder tree included in the arithmetic module 320 of the SRAM IMC macro 300; a second shifter 253 configured to adjust the value stored in the accumulator 255; and an accumulator 255. In Figure 4A and Figure 4B , the arithmetic module 320 corresponds to any m-th column (e.g., column m) in (the M columns).
[0139] For example, when the operation mode is in the first mode, the first shifter 251 may adjust the operation result of the adder tree through a right-shift operation to apply a value of dt / tau in the form of 2 -t (where t is a natural number greater than 1, i.e., t > 1). ).
[0140] When the operation mode is in the second mode, the second shifter 253 may adjust the value stored in the accumulator 255 to the value multiplied by 2 through a bitwise left-shift operation u (where u is a natural number greater than 1, i.e., u > 1).
[0141] Referring to Figure 4A , FIG. 400 illustrates the operation of the post-arithmetic circuit 250 executed when the operation mode of the IMC macro is in the first mode.
[0142] When the operation mode is in the first mode for the SNN, the post-arithmetic circuit 250 may receive, from the arithmetic module 320 of the SRAM IMC macro 300, the result of adding "(i) the processed previous membrane potential value -U(t)" and "(ii) the result of the operation between the pulse signal 401(I in (t)) and the weight R (i.e., RI in (t)) in the adder tree (e.g., -U(t)+RI in (t)).
[0143] The post-arithmetic circuit 250 may send the output -U(t)+RI in (t) of the arithmetic module 320 to the first shifter 251, and the first shifter 251 may send the result obtained by performing a right shift operation on -U(t)+RI in (t) (e.g., ) to the accumulator 255. In this case, dt / tau ( ) may correspond to the time constant.
[0144] In the first mode, the second shifter 253 may not perform a shift operation, but simply bypass the previous membrane potential value U(t) stored in the accumulator 255 (transfer the previous membrane potential value U(t) stored in the accumulator 255).
[0145] The accumulator 255 may add the previous membrane potential value U(t) bypassed from the second shifter 253 to the result of the right shift operation sent from the first shifter 251 and send the updated membrane potential value U(t+Δt) to the arithmetic module 320 through the input control circuit.
[0146] In this way, the operation result of the IMC macro may ultimately be sent to the accumulator 255 (also indicated as "Accum" in the figure), and the value sent to the accumulator 255 may be sent to the arithmetic module 320 through the input control circuit together with the control signal for write-back.
[0147] The input control circuit may arithmetically invert the updated membrane potential value U(t+Δt) to -U(t+Δt) and store the latter in the additional row 310. In this case, each of the M columns may have the value of -U(t+Δt). Since the IMC macro performs row-by-row writing, the IMC macro may simultaneously write the value of -U(t+Δt) into the additional row 310 of the M columns. In the first mode, the accumulator 255 may store the membrane potential value.
[0148] Referring to Figure 4B , FIG. 410 shows the operation of the post-arithmetic circuit 250 performed when the operation mode of the IMC macro is in the second mode.
[0149] When the operation mode is in the second mode for NN, the input feature map 403 may be input, and the post-arithmetic circuit 250 may receive the bias value B stored in the append row 310 in the adder tree from the operation module 320 of the SRAM IMC macro 300. m ” and “(ii) Input signal X(b) and weight W m The result of adding the operation results between the two (for example, X(b)×W m +PP(b)×B m ). In this case, b represents the number of bits, and X(b) may represent the b-th input bit. In this case, the bits may be numbered in reverse order from the most significant bit (MSB) to the least significant bit (LSB). PP(b) may correspond to the input bit of the b-th pattern starting from the MSB.
[0150] When PP(b) is accumulated to multiple bits, the result of adding the bias value becomes X×W m +B m In this case, and as described above, m is the column index.
[0151] In the second mode, for each column, the additional row 310 may store the bias value B m , and the operation module 320 may apply a predefined pattern (PP) value to perform an operation (such as, for example, Y m =X×W m +B m ×PP). In this case, Y m Indicates the operation result of the mth column. For example, W m represents the weight corresponding to the mth column in the N×M weight matrix W. X represents the input vector consisting of {X1, X2, ..., Xn}. The PP value can be selected and set by the user. In this case, when the PP value is "1", the operation result of the operation module 320 can be Y m =X×W m +B m .
[0152] In the second mode, the first shifter 251 may shift the weight W (i) stored in the memory cell to m The result of the first multiplication operation with the input signal X (eg, X×W m )” and “(ii) the offset value B stored in the additional row 310 m The result of the second multiplication operation with the PP value (e.g., B m ×PP)” (for example, X×W m +B m ×PP or X×W m +B m) bypass (the result of the first multiplication operation between “(i) the weight W stored in the memory cell m and the input signal X (e.g., X×W m )” and “(ii) the bias value B stored in the additional row 310 m and the result of the second multiplication operation between the PP value (e.g., B m ×PP)” (e.g., X×W m +B m ×PP or X×W m +B m ) and transfer it to the accumulator 255. In this case, the second shifter 253 can perform a left shift operation on the operation result corresponding to the input signal X serially applied bit - by - bit to the accumulator 255 (e.g., X×W m +B m ).
[0153] In the second mode, the accumulator 255 can perform the function of converting the serially - applied bit - by - bit calculation result (e.g., the result of the left shift operation) into a multi - bit calculation result.
[0154] Figure 5 An example flow showing the operation of the IMC macro according to one or more example embodiments.
[0155] Referring to Figure 5 , the IMC macro of the example embodiment can selectively perform the first operation corresponding to the SNN or the second operation corresponding to the NN by performing operations 505 to 560 described below.
[0156] In operation 505, the IMC macro can perform initialization based on the operation mode (such as, for example, setting the operation mode, the predefined mode (PP), and / or the shifter of the post - arithmetic circuit). The operation mode can be determined by a command from an external device (such as a host).
[0157] In operation 510, the IMC macro can store the weights in the rows of the memory cell and can also store information in the additional row. For example, when the operation mode is in the first mode, the IMC macro can store zero (“0”) in the additional row. When the operation mode is in the second mode, the IMC macro can store the bias value in the additional row. In this case, the bias value can be selectively used and can be different for each column.
[0158] In operation 515, the IMC macro can determine (e.g., by checking the value of the register) whether the set operation mode is in the first mode for the SNN.
[0159] In operation 520, when it has been determined in operation 515 that the operation mode is in the first mode, for each column, the IMC macro may apply an input pulse signal to the rows of the memory cells and apply a PP value to the additional rows to calculate RI in (t - U(t).
[0160] In operation 525, the IMC macro may perform an operation (e.g., ) through the post - arithmetic circuit.
[0161] In operation 530, the IMC macro may calculate the updated membrane potential value through an accumulator.
[0162] In operation 535, the IMC macro may send the updated membrane potential value U(t + Δt) calculated by the accumulator, and this value may be sent to the additional rows together with the control signal for write - back through the input control circuit. In this case, when the updated membrane potential value U(t + Δt) is greater than the threshold, the value of the additional row of the corresponding column may be "0". In addition, when the updated membrane potential value U(t + Δt) is less than or equal to the threshold, the value of the additional row of the corresponding column may be - U(t + Δt).
[0163] In operation 540, the IMC macro may determine whether the input data is the last one. When it is determined in operation 540 that the input data is not the last one, the IMC macro may execute operation 520 again.
[0164] When it is determined in operation 540 that the input data is the last one, the IMC macro may return to the "start" point or end the operation.
[0165] In operation 545, when it has been determined in operation 515 that the operation mode is not in the first mode (i.e., the operation mode is in the second mode for NN), the IMC macro may determine whether to use a bias value. When it has been determined not to use a bias value, the IMC macro may apply all zeros ("0") as the input mode value (i.e., the predefined mode), so that the multiplication operation result is forced to be "0".
[0166] In operation 550, when it has been determined in operation 545 not to use a bias value, the IMC macro may perform an operation in which the bias value is not reflected for the N - bit input data (or input signal) of each column (e.g., Y=(Y << 1)+W×X[j]).
[0167] In operation 555, when it has been determined in operation 545 to use a bias value, the IMC macro may perform an operation in which the bias value is reflected for the N - bit input data (or input signal) of each column (e.g., Y=(Y << 1)+W×X[j]+B[j]).
[0168] In operation 560, the IMC macro can determine whether the input data is the last one. When it is determined in operation 560 that the input data is not the last one, the IMC macro can execute operation 545 again.
[0169] When it is determined in operation 560 that the input data is the last one, the IMC macro can return to the "start" point or end the operation.
[0170] Figure 6 An example flow diagram showing the operations of the IMC macro according to one or more example embodiments is provided. Referring Figure 6 to, the IMC macro of the example embodiment can selectively perform a first operation corresponding to the SNN or a second operation corresponding to the NN by executing operations 610 to 640 described below.
[0171] In operation 610, according to the operation mode, the IMC macro can send the result of applying a predefined pattern to the input signal and / or the previous membrane potential value of the feedback. The input signal can be a pulse signal for the SNN or a feature map for the NN. The previous membrane potential value of the feedback can include the previous membrane potential value of the SNN. According to the operation mode, the IMC macro can send the processed value or the bias value of the previous membrane potential value to an additional row of the corresponding memory cell for each in the column.
[0172] For example, when the operation mode is in the first mode for the SNN, the IMC macro can update the additional row by sending the processed value obtained by multiplying the previous membrane potential value by -1 (or otherwise arithmetically inverting the value) to the additional row of the corresponding memory cell. Optionally, when the operation mode is in the second mode for the NN, the IMC macro can update the additional row by sending the bias value to the additional row of the corresponding memory cell.
[0173] In operation 620, the IMC macro can store the weights corresponding to the input signal in a row of the memory cell and can store the processed value or the bias value of the previous membrane potential value in an additional row of the memory cell.
[0174] In operation 630, the IMC macro can add "the result of the first multiplication operation between (i) the input signal and the weights" and "the result of the second multiplication operation between (ii) the predefined pattern and the processed value or the bias value of the previous membrane potential value" through an adder tree.
[0175] In operation 640, the IMC macro can selectively perform a first operation corresponding to the SNN or a second operation corresponding to the NN according to the operation mode.
[0176] For example, when the operation mode is in the first mode for the SNN, the IMC macro can perform a right shift operation on the operation result between the pulse signal and the weight through the first shifter, and the operation results are added by the adder tree. The IMC macro can bypass the membrane potential value stored in the accumulator (transfer the membrane potential value stored in the accumulator) through the second shifter, and store the result of the right shift operation and the bypassed membrane potential value in the accumulator. Optionally, when the operation mode is in the second mode for the NN, the IMC macro can bypass the result of adding "(i) the first multiplication operation between the weight stored in the memory cell and the input signal" and "(ii) the second multiplication operation between the value of the predefined pattern and the bias value of the corresponding column stored in the additional row" (transfer the result of adding "(i) the first multiplication operation between the weight stored in the memory cell and the input signal" and "(ii) the second multiplication operation between the value of the predefined pattern and the bias value of the corresponding column stored in the additional row") to the accumulator through the first shifter. The IMC macro can perform a left shift operation on the operation result of the accumulator (the operation result corresponding to the input signal applied bit-serially) through the second shifter. The IMC macro can generate a multi-bit result by accumulating the result of the left shift operation via the accumulator.
[0177] Figure 7 FIG. shows an example electronic system including an IMC macro according to one or more example embodiments. Referring to Figure 7 , the electronic system 700 of the example embodiment can analyze input data in real time based on a neural network (e.g., Figure 1A the neural network 130 in ) to extract valid information, and can determine the situation of the electronic device on which the electronic system 700 is installed or can control the components of the electronic device on which the electronic system 700 is installed based on the extracted information. As a non-limiting example, the electronic system 700 can be installed on at least one of an unmanned aerial vehicle, a robotic device (such as an advanced driver assistance system (ADAS)), a vehicle, a smart TV, a smart phone, a medical device, a mobile device, an image display device, an instrument device, an Internet of Things (IoT) device, and other types of electronic devices.
[0178] The electronic system 700 can include a processor 710, a random access memory (RAM) 720, a neural network device 730, a memory 740, a sensor module 750, and a transmit / receive module 760. The electronic system 700 can also include an input / output module, a security module, a power control device, etc. Some of the hardware components of the electronic system 700 can be installed on at least one semiconductor chip.
[0179] The processor 710 may control the overall operation of the electronic system 700. The processor 710 may include a single processor core (e.g., a single core) of any type of processor (including the examples mentioned herein) or may include multiple processor cores of varying types (e.g., a multi-core). Although the term "processor" (e.g., processor 710) is used in the singular in some places, the term means "one or more processors". The processor 710 may process or execute programs and / or data stored in the memory 740. In some example embodiments, the processor 710 may execute a program stored in the memory 740 to control the functions of the neural network device 730. The processor 710 may be implemented as a central processing unit (CPU), a graphics processing unit (GPU), an application processor (AP), etc.
[0180] The RAM 720 may temporarily store programs, data, or instructions. For example, the programs and / or data stored in the memory 740 may be temporarily stored in the RAM 720 in response to control or startup code from the processor 710. The RAM 720 may be implemented as a memory (such as, for example, dynamic RAM (DRAM) or static RAM (SRAM)).
[0181] The neural network device 730 may perform computational operations of a neural network based on received input data and may generate various information signals based on the results of the performed computational operations. As a non-limiting example, the neural network may include a CNN, an RNN, a fuzzy neural network (FNN), a deep belief network (DBN), a restricted Boltzmann machine (RMB), etc. The neural network device 730 may be, for example, a hardware accelerator dedicated to a neural network itself and / or a device including a hardware accelerator.
[0182] For example, the neural network device 730 may correspond to any one of the IMC macros described above (e.g., Figure 2 the IMC macro 200 in Figure 3 and / or the IMC macro 300 in
[0183] The term "information signal" as used herein may include one of various types of identification signals (such as, for example, voice recognition signals, object recognition signals, image recognition signals, biometric information recognition signals, etc.). For example, the neural network device 730 may receive frame data included in a video stream as input data and may generate an identification signal for an object included in the image represented by the frame data from the frame data. The neural network device 730 may receive various types of input data according to the type or function of the electronic device on which the electronic system 700 is installed and may generate an identification signal based on the input data.
[0184] The memory 740, as a storage location for storing data, may store an operating system (OS), various programs, and various data. In one exemplary embodiment, the memory 740 may store intermediate results generated during the processing of performing the computing operations of the neural network device 730.
[0185] The memory 740 may include at least one of a volatile memory and a non-volatile memory (but not including the signal itself). As a non-limiting example, the non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, etc. As a non-limiting example, the volatile memory may include DRAM, SRAM, synchronous DRAM (SDRAM), phase change memory (PCM) RAM (PRAM), magnetic RAM (MRAM), resistive RAM (RRAM), and / or ferroelectric RAM (FRAM). According to an example, the memory 740 may include at least one of a hard disk drive (HDD), a solid state drive (SSD), a compact flash (CF) card, a secure digital (SD) card, a micro SD card, a mini SD card, an extreme digital (Xd) picture card, and a memory stick.
[0186] The sensor module 750 may collect information around the electronic device on which the electronic system 700 is installed. The sensor module 750 may sense or receive signals from outside the electronic system 700 (such as, for example, image signals, voice signals, magnetic signals, biological signals, touch signals, etc.) and convert the sensed or received signals into data. The sensor module 750 may include at least one of various sensing devices (such as, for example, microphones, imaging devices, image sensors, light detection and ranging (LIDAR) sensors, ultrasonic sensors, infrared sensors, biological sensors, and touch sensors).
[0187] The sensor module 750 may provide the data obtained through conversion as input data to the neural network device 730. For example, the sensor module 750 may include an image sensor, and may generate a video stream by capturing an image of the external environment of the electronic system 700, and provide consecutive data frames of the video stream as input data to the neural network device 730. However, the sensor module 750 is not limited thereto, and may provide various types of data to the neural network device 730.
[0188] The transmit / receive module 760 may include various types of wired interfaces or wireless interfaces configured to communicate with an external device. For example, the transmit / receive module 760 may include a local area network (LAN), a wireless LAN (WLAN) (such as, Wi-Fi), a wireless personal area network (WPAN) (such as, Bluetooth), a wireless universal serial bus (USB), ZigBee, near field communication (NFC), radio frequency identification (RFID), power line communication (PLC), a mobile cellular network (such as, 3G, 4G, and long term evolution (LTE)), etc. accessible communication interfaces.
[0189] The examples described herein may be implemented using hardware components, software components, and / or combinations thereof. The processing device may be implemented using one or more general-purpose computers or special-purpose computers (such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of responding and executing instructions in a defined manner). The processing device may run an operating system (OS) and one or more software applications running on the OS. The processing device may also access, store, manipulate, process, and create data in response to the execution of software. For simplicity, the description of the processing device is used in the singular form; however, those skilled in the art will understand that the processing device may include multiple processing elements and various types of processing elements. For example, the processing device may include multiple processors or a processor and a controller. Additionally, different processing configurations (such as, parallel processors) are feasible.
[0190] Herein, with respect to Figures 1A to 7The described computing devices, vehicles, electronic devices, processors, memories, sensors, vehicle / operation function hardware, ADAS systems, displays, information output systems and hardware, storage devices, and other devices, apparatuses, units, modules, and components are implemented by or represent hardware components. Examples of hardware components that can be used to perform the operations described in this application include, where appropriate: controllers, sensors, generators, drivers, memories, comparators, arithmetic logic units, adders, subtractors, multipliers, dividers, integrators, and any other electronic components configured to perform the operations described in this application. In other examples, one or more of the hardware components that perform the operations described in this application are implemented by computing hardware (e.g., by one or more processors or computers). A processor or computer can be implemented by one or more processing elements (such as an array of logic gates, a controller and an arithmetic logic unit, a digital signal processor, a microcomputer, a programmable logic controller, a field programmable gate array, a programmable logic array, a microprocessor, or any other device or combination of devices configured to respond and execute instructions in a defined manner to achieve a desired result). In one example, a processor or computer includes or is connected to one or more memories that store instructions or software executed by the processor or computer. The hardware components implemented by the processor or computer can execute instructions or software (such as an operating system (OS) and one or more software applications running on the OS) to perform the operations described in this application. The hardware components can also access, manipulate, process, create, and store data in response to the execution of the instructions or software. For simplicity, the singular terms "processor" or "computer" may be used in the description of the examples described in this application, but in other examples, multiple processors or computers may be used, or a processor or computer may include multiple processing elements, or multiple types of processing elements, or both. For example, a single hardware component or two or more hardware components can be implemented by a single processor, or two or more processors, or a processor and a controller. One or more hardware components can be implemented by one or more processors, or a processor and a controller, and one or more other hardware components can be implemented by one or more other processors, or additional processors and additional controllers. One or more processors, or a processor and a controller, can implement a single hardware component or two or more hardware components. The hardware components can have any one or more of different processing configurations, examples of different processing configurations including: single processor, independent processor, parallel processor, single instruction single data (SISD) multiprocessing, single instruction multiple data (SIMD) multiprocessing, multiple instruction single data (MISD) multiprocessing, and multiple instruction multiple data (MIMD) multiprocessing.
[0191] Performing the operations described in this application in Figures 1A to 7The method shown is performed by computing hardware (e.g., by one or more processors or computers), which is implemented to execute instructions or software as described above to perform the operations performed by the method described in this application. For example, a single operation or two or more operations can be performed by a single processor, or two or more processors, or a processor and a controller. One or more operations can be performed by one or more processors, or a processor and a controller, and one or more other operations can be performed by one or more other processors, or additional processors and additional controllers. One or more processors, or a processor and a controller, can perform a single operation or two or more operations.
[0192] Instructions or software for controlling computing hardware (e.g., one or more processors or computers) to implement the hardware components and perform the method as described above can be written as a computer program, code segment, instruction, or any combination thereof to individually or jointly direct or configure one or more processors or computers to operate as a machine or a special-purpose computer to perform the operations performed by the hardware components and method as described above. In one example, the instructions or software include machine code (such as machine code generated by a compiler) directly executable by one or more processors or computers. In another example, the instructions or software include high-level code executable by one or more processors or computers using an interpreter. The instructions or software can be written in any programming language based on the block diagrams and flowcharts shown in the figures and the corresponding descriptions herein, which disclose algorithms for performing the operations performed by the hardware components and method as described above.
[0193] Instructions or software for controlling computing hardware (e.g., one or more processors or computers) to implement the hardware components and perform the methods as described above, along with any associated data, data files, and data structures, can be recorded, stored, or fixed in or on one or more non-transitory computer-readable storage media. Examples of non-transitory computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-RLTH, BD-Re, Blu-ray or optical disc storage, hard disk drive (HDD), solid state drive (SSD), card-type memory (such as, multimedia card or micro card (e.g., Secure Digital (SD) or Extreme Digital (XD))), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store instructions or software and any associated data, data files, and data structures in a non-transitory manner and provide the instructions or software and any associated data, data files, and data structures to one or more processors or computers such that the one or more processors or computers can execute the instructions. In one example, the instructions or software and any associated data, data files, and data structures are distributed across a networked computer system such that the instructions and software and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.
[0194] Although the present disclosure includes specific examples, it will be apparent after understanding the disclosure of this application that various changes in form and detail can be made in these examples without departing from the spirit and scope of the claims and their equivalents. The examples described herein will be considered to be merely descriptive and not for purposes of limitation. The description of each feature or aspect in an example will be considered to be applicable to similar features or aspects in other examples. Appropriate results can be achieved if the described techniques are performed in a different order, and / or if the components in the described system, architecture, device, or circuit are combined in a different manner, and / or replaced or supplemented by other components or their equivalents.
[0195] Accordingly, in addition to the foregoing disclosure, the scope of the disclosure may also be defined by the claims and their equivalents, and all variations within the scope of the claims and their equivalents will be construed as being included in the disclosure.
Claims
1. An in-memory computing device having an operating mode that alternates between a first mode and a second mode, the in-memory computing device comprising: An input control circuit configured to: apply a predefined pattern to an input signal and / or send a processed value of a previous operation result as feedback, depending on the operation mode; a crossbar array including a plurality of columns, each of the plurality of columns including a memory cell and an adder tree, and the memory cell including an additional row storing a processed value of a previous operation result or a bias value; as well as A plurality of post-arithmetic circuits correspond to the plurality of columns, and for each of the plurality of columns, the corresponding post-arithmetic circuit is configured to: perform a first operation corresponding to a spiking neural network or a second operation corresponding to a neural network other than the spiking neural network according to an operation mode.
2. The in-memory computing device of claim 1, wherein: For each of the plurality of columns, the corresponding memory unit includes a row storing a weight corresponding to an input signal, and The corresponding adder tree is configured to add the result of a first multiplication operation between the input signal and the weight and the result of a second multiplication operation between the predefined pattern and the processed value of the previous operation result or the bias value, depending on the operation mode.
3. The in-memory computing device of claim 1 , wherein: Input signals include: Spiking signals for spiking neural networks or feature maps for neural networks.
4. The in-memory computing device of claim 1, wherein: According to a command sent from the host, the operation mode is set to the first mode for the spiking neural network or the second mode for the neural network.
5. The in-memory computing device of claim 1, wherein: The first mode is for a spiking neural network and the second mode is for a neural network, and wherein the input control circuit is further configured to: Based on the operation mode being in the first mode, setting the predefined mode to 1; and Based on the operation mode being in the second mode, the predefined mode is set to a mode corresponding to the number of bits of the input signal.
6. The in-memory computing device of claim 1, wherein: The previous operation results include previous membrane potential values of the spiking neural network, and wherein the input control circuit includes an additional input port, which is configured to: according to the operation mode, for each of the multiple columns, send a processed value of the previous membrane potential value or a bias value to a corresponding additional row of the corresponding memory cell.
7. The in-memory computing device of claim 6, wherein: The first mode is for a spiking neural network and wherein the additional input port is configured as: Based on the operation mode being in the first mode, for each of the multiple columns, a processed value of a previous membrane potential value is sent to a corresponding additional row of a corresponding memory cell, wherein the processed value of the previous membrane potential value is the arithmetically inverted previous membrane potential value fed back from the corresponding post-arithmetic circuit.
8. The in-memory computing device of claim 6, wherein: The second mode is for a neural network and wherein the additional input port is configured as: Based on the operating mode being in the second mode, for each column of the plurality of columns, bias values are sent to corresponding additional rows of corresponding memory cells.
9. The in-memory computing device of claim 6, wherein: For each column in the plurality of columns, a corresponding additional row is configured as: storing a processed value of a previous membrane potential value based on the operating mode being in the first mode; and Based on the operating mode being in the second mode, the offset value is stored.
10. The in-memory computing device of claim 9, wherein: The crossbar array is configured as: For each of the plurality of columns, a result of adding a result of a first multiplication operation obtained by adding respective products between an input signal and weights stored in corresponding memory cells and a result of a second multiplication operation is stored, wherein the result of the first multiplication operation is obtained by multiplying a predefined pattern with a value stored in an additional row.
11. The in-memory computing device of claim 1 , wherein: Each column of the plurality of post-arithmetic circuits comprises: A first shifter is configured to: adjust the operation result of the corresponding adder tree by a right shift operation based on the operation mode being in the first mode; a second shifter configured to adjust the value stored in the accumulator by a left shift operation based on the operation mode being in the second mode; and accumulator.
12. The in-memory computing device of claim 11, wherein: For each column of the plurality of columns, a corresponding post-arithmetic circuit is configured to: Based on the operation mode being in the first mode: By means of a first shifter, a right shift operation is performed on the operation result between the pulse signal and the weight added by the corresponding adder tree; passing the membrane potential value stored in the accumulator through the second shifter; and The result of the right shift operation and the transferred membrane potential value are stored in the accumulator.
13. The in-memory computing device of claim 11, wherein: For each column of the plurality of columns, the corresponding post-arithmetic circuit is configured to: be in a second mode based on the operation mode: Passing, through the first shifter, a result of adding a result of a first multiplication operation between the input signal and the weight stored in the corresponding memory cell and a result of a second multiplication operation between the predefined pattern and the bias value stored in the corresponding additional row to the accumulator; By means of the second shifter, a left shift operation is performed on an operation result in the accumulator corresponding to the input signal applied bit-serially; and The results of the left shift operation are accumulated through the accumulator to produce a multi-bit result.
14. The in-memory computing device of claim 1, wherein: For each column of the plurality of columns, a corresponding adder tree is configured as: A first multiplication operation between an input signal and a weight stored in a corresponding memory unit and a second multiplication operation between a predefined pattern and a processed value or a bias value of a previous operation result are performed simultaneously.
15. The in-memory computing device of claim 1, wherein: The in-memory computing device is integrated in at least one of the following: mobile devices, mobile computing devices, mobile phones, smart phones, personal digital assistants, fixed location terminals, tablet computers, computers, wearable devices, laptop computers, servers, music players, video players, entertainment units, navigation devices, communication devices, global positioning system devices, televisions, tuners, satellite radio devices, song players, digital video players, digital video disc players, vehicles, components of vehicles, avionics systems, drones, multicopters, and medical devices.
16. A method of operating an in-memory computing device, the in-memory computing device having an operating mode that alternates between a first mode and a second mode, the method comprising: Depending on the operation mode, applying a predefined pattern to an input signal and / or sending a processed value of a previous membrane potential value as feedback; storing weights corresponding to input signals in rows of memory cells, and storing bias values of columns corresponding to the memory cells or processed values of the previous membrane potential values in additional rows of memory cells; adding, through an adder tree corresponding to the memory unit, a result of a first multiplication operation between the input signal and the weight and a result of a second multiplication operation between the predefined pattern and the processed value of the previous membrane potential value or the bias value; as well as According to the operation mode, a first operation corresponding to a spiking neural network or a second operation corresponding to a neural network other than the spiking neural network is selectively performed.
17. The method according to claim 16, wherein: The steps of sending include: Depending on the operating mode, the processed value of the previous membrane potential value or the bias value is sent to an additional row of memory cells.
18. The method according to claim 17, wherein: The first mode is for a spiking neural network and the second mode is for a neural network, and wherein the step of sending to the additional row of memory cells comprises: updating the additional row of memory cells by sending a processed value of the previous membrane potential value to the additional row of memory cells based on the operating mode being in the first mode, wherein the processed value of the previous membrane potential value is obtained by arithmetically negating the previous membrane potential value; and Based on the operating mode being in the second mode, additional rows of memory cells are updated by sending the bias value to the additional rows of memory cells.
19. The method according to claim 16, wherein: The first mode is for a spiking neural network, and wherein the steps selectively performed include: Based on the operation mode being in the first mode, Performing a right shift operation on the operation result between the pulse signal and the weight added by the adder tree through the first shifter; passing the membrane potential value stored in the accumulator through the second shifter; and The result of the right shift operation and the transferred membrane potential value are stored in the accumulator.
20. The method according to claim 16, wherein: The second mode is for a neural network, and wherein the steps selectively performed include: Based on the operating mode in the second mode: passing, through the first shifter, a result of adding a result of a first multiplication operation between the input signal and the weight stored in the memory unit and a result of a second multiplication operation between the predefined pattern and the bias value stored in the additional row to an accumulator; performing a left shift operation on an operation result corresponding to an input signal applied bit-serially in the accumulator by a second shifter; and The results of the left shift operation are accumulated through the accumulator to produce a multi-bit result.