Neural network device and its operating method

The neural network device addresses precision and size challenges by using a converter and cell array to perform higher bit operations within edge computing devices, achieving efficient and cost-effective calculations.

JP7910788B2Active Publication Date: 2026-08-25PEBBLE SQUARE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024202661
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2024-07-31
Filing Date
2024-11-20
Publication Date
2026-08-25
Estimated Expiration
2044-11-20

AI Technical Summary

Technical Problem

Existing neural network devices face challenges in performing precise operations due to size constraints when implemented in edge computing devices, particularly with high-resolution memory cells and digital-to-analog converters, leading to increased costs and device size.

Method used

A neural network device comprising a digital-to-analog converter, a cell array with memory cells, and an analog-to-digital converter, along with a processor, that allows for higher bit operations by converting digital inputs into analog signals, performing calculations within the cell array, and converting back to digital outputs, thereby overcoming resolution limitations.

Benefits of technology

Enables higher bit operations and output generation beyond the resolution of limited converters and memory cells, reducing device size and cost while maintaining precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007910788000001
    Figure 0007910788000001
  • Figure 0007910788000002
    Figure 0007910788000002
  • Figure 0007910788000003
    Figure 0007910788000003
Patent Text Reader

Abstract

To provide a neural network device that performs an operation on data having a higher number of bits than a limited bit resolution of a digital-to-analog converter, and an operation method thereof.SOLUTION: The neural network device includes a digital-analog converter that converts a digital input into an analog input of either a voltage or a current, a cell array 2 that includes a plurality of memory cells that are arranged in a plurality of bit lines and a plurality of word lines and store neural network weights, performs an operation on the analog input input via the word lines, and outputs an analog output of either a current or a voltage via the bit lines, an analog-digital converter that converts the analog output into a digital output, and at least one processor that is electrically connected to the digital-analog converter and the analog-digital converter and executes control on the digital input and the digital output.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a neural network device and an operating method thereof, and more particularly, to a method for performing more precise bit operations using a neural network device having a relatively small bit resolution.

Background Art

[0002] An artificial neural network imitates a biological neural network, and they can be learned by a large number of input data and are used to estimate or approximate results that are difficult to derive by general techniques. An artificial neural network includes interconnected neuron layers that exchange signals, and synapses have weights determined based on learning or experience.

[0003] On the other hand, a CIM (Computing in memory) device that performs analog operations processes data inside the memory, so data movement between the memory and the processor is minimized and the operation speed is improved. However, there is a problem that precise operations are difficult due to size constraints when realized in an edge computing device.

[0004] The above-described background art is technical information that the inventor possessed for deriving the present invention or acquired in the process of deriving the present invention, and is not necessarily known art publicly disclosed to the general public before the filing of the present invention.

Summary of the Invention

Problems to be Solved by the Invention

[0005] The purpose of this disclosure is to provide a neural network device and a method of operating the same. The problems that this disclosure seeks to solve are not limited to the technical problems mentioned above, and other technical problems not mentioned can be clearly understood by a person ordinary in the art from the description of the invention, and will be understood more clearly from the embodiments of this disclosure. Furthermore, it will be found that the problems and advantages that this disclosure seeks to solve can be achieved by the means and combinations thereof set forth in the claims. [Means for solving the problem]

[0006] As a means to solve the above technical problems, a first aspect of the present disclosure provides a neural network device comprising: a digital-to-analog converter that converts a digital input into an analog input of either voltage or current; a cell array that includes a plurality of memory cells arranged in a plurality of bit lines and a plurality of word lines and storing weights of a neural network, and which performs operations on the analog input input via the word lines and outputs an analog output of either current or voltage via the bit lines; an analog-to-digital converter that converts the analog output into a digital output; and at least one processor that is electrically connected to the digital-to-analog converter and the analog-to-digital converter and performs control on the digital input and the digital output, wherein the at least one processor inputs one or more digital inputs, including at least a portion of the input signal, to the digital-to-analog converter based on the number of bits of the input signal and the DAC bit resolution of the digital-to-analog converter, and generates the output signal using at least one of the digital outputs corresponding to the output of the bit lines based on the number of bits of the input signal and the cell bit resolution of the plurality of memory cells.

[0007] A second aspect of the present disclosure can provide a method for operating a neural network device, comprising the steps of: generating one or more digital inputs including at least a portion of an input signal based on the number of bits of an input signal and the DAC bit resolution of a digital-to-analog converter; obtaining one or more digital outputs corresponding to the one or more digital inputs using a cell array including a plurality of memory cells that store the weights of a neural network; and generating an output signal using at least one of the digital outputs based on the number of bits of an input signal and the cell bit resolution of the plurality of memory cells.

[0008] A third aspect of this disclosure can provide a computer-readable recording medium that stores a program for performing the method of the second aspect on a computer.

[0009] In addition, other methods, apparatuses, and computer-readable recording media containing programs for implementing the present invention can be provided.

[0010] Other aspects, features, and advantages not mentioned above will become clear from the following drawings, claims, and detailed description of the invention. [Effects of the Invention]

[0011] According to the means for solving the problems of this disclosure described above, it is possible to perform operations on data having a higher number of bits than the bit resolution of a limited digital-to-analog converter.

[0012] Furthermore, according to the means for solving the problems of this disclosure, it is possible to generate an output having a higher number of bits than the bit resolution of a limited memory cell.

[0013] The effects of the embodiments are not limited to those mentioned above, and other effects not mentioned will be clearly understood by those ordinary skill in the art from the description of the present invention. [Brief explanation of the drawing]

[0014] [Figure 1] This is a diagram illustrating the implementation of a neural network system according to one embodiment. [Figure 2] This is an illustrative diagram illustrating a comparison between a von Neumann architecture and a CIM (computing in memory) architecture according to one embodiment of the present invention. [Figure 3] This is an illustrative diagram illustrating a comparison between a von Neumann architecture and a CIM (computing in memory) architecture according to one embodiment of the present invention. [Figure 4] This figure shows a neural network device according to one embodiment of the present invention. [Figure 5A] This is a diagram illustrating the operation method of a cell array according to one embodiment. [Figure 5B] This is a diagram illustrating the operation method of a cell array according to one embodiment. [Figure 6A] This figure illustrates a comparison between the product of a matrix and a vector and an operation performed on a cell array, according to one embodiment. [Figure 6B] This figure illustrates a comparison between the product of a matrix and a vector and an operation performed on a cell array, according to one embodiment. [Figure 7] This figure illustrates an example in which a convolution operation is performed in a cell array according to one embodiment. [Figure 8] This diagram illustrates the operation of a neural network device according to one embodiment of the present invention. [Figure 9] This is a diagram illustrating a digital input using one embodiment of the present invention. [Figure 10] This diagram illustrates a method for generating an output signal according to one embodiment of the present invention. [Figure 11] This figure illustrates a digital input according to another embodiment of the present invention. [Figure 12] FIG. for explaining a method of generating an output signal according to another embodiment of the present invention. [Figure 13] FIG. for explaining an operation method of a neural network device according to another embodiment of the present invention. [Figure 14] FIG. for explaining a method of generating an output signal according to another embodiment of the present invention. [Figure 15] FIG. for explaining a method of generating an output signal according to another embodiment of the present invention. [Figure 16] A flowchart of an operation method of a neural network device according to an embodiment of the present invention. [Figure 17] A block diagram of a neural network device according to another embodiment of the present invention.

BEST MODE FOR CARRYING OUT THE INVENTION

[0015] When it is determined that a specific description of the known technology related to explaining the present invention obscures the gist of the present invention, the detailed description thereof can be omitted, and unless otherwise defined, all terms used in this specification have the same meaning as generally understood by those having ordinary knowledge in the technical field to which the present invention belongs.

[0016] Phrases such as "according to an embodiment", "relating to an embodiment", or "by implementing an embodiment" in this specification do not necessarily refer to the same embodiment.

[0017] Embodiments can be modified in various ways and can have various forms, so some embodiments are shown in the drawings and described in detail. However, this is not intended to limit the embodiments to a specific disclosed form, and should be understood to include all modifications, equivalents, or alternatives included in the spirit and technical scope of the embodiments. The terms used in the specification are merely used to explain the embodiments and are not intended to limit the embodiments.

[0018] The terminology used in these embodiments has been selected to the greatest extent possible from commonly used terms, taking into account the functions of these embodiments. However, this may change depending on the intentions of engineers in the technical field to which the embodiments belong, case law, the emergence of new technologies, etc. In some cases, the applicant has arbitrarily selected terms, in which case their meaning will be described in detail in the relevant sections. Therefore, the terminology used in these embodiments should not be merely names of terms, but should be defined based on the meaning of the terms and the context of the embodiments as a whole.

[0019] Some embodiments of this disclosure can be represented by functional block configurations and various processing steps. Some or all of such functional blocks can be implemented by various numbers of hardware and / or software configurations that perform a particular function. For example, a functional block of this disclosure may be implemented by one or more microprocessors, or by a circuit configuration for a given function.

[0020] Furthermore, for example, the functional blocks of this disclosure can be implemented in various programming or scripting languages. Functional blocks can also be implemented in algorithms that run on one or more processors. In addition, this disclosure can employ prior art for electronic environment configuration, signal processing, and / or data processing.

[0021] Terms such as “database,” “element,” “means,” and “configuration” can be used broadly and are not limited to mechanical and physical configurations. Furthermore, terms such as “part” and “module” as described in the specification mean a unit that processes at least one function or operation, which may be implemented in hardware or software, or in combination of hardware and software.

[0022] Furthermore, the connecting lines or members shown in the drawings between components are merely illustrative examples of functional and / or physical or circuit connections. In actual devices, connections between components may be indicated by a variety of alternative or added functional, physical, or circuit connections.

[0023] Furthermore, while ordinal terms such as "first" or "second" used herein may be used to describe various components, the components should not be limited by these terms. The terms are used solely for the purpose of distinguishing one component from another.

[0024] Furthermore, some components in the drawings may be shown with slightly exaggerated sizes or proportions. Additionally, components shown in one drawing may not be shown in other drawings.

[0025] Throughout the specification, “Embodiments” are any classifications that facilitate the description of the invention in this disclosure, and each embodiment does not have to be mutually exclusive. For example, a configuration disclosed in one embodiment may be applied to and / or implemented in other embodiments, and may be modified and applied and / or implemented without departing from the scope of this disclosure.

[0026] Furthermore, the terms used in this disclosure are for illustrative purposes only and are not intended to limit these embodiments. Unless otherwise specified, singular terms in this disclosure also include plural terms.

[0027] The embodiments of this disclosure will be described in detail below with reference to the accompanying drawings, so as to be easily implemented by a person skilled in the art. However, the embodiments of this disclosure can be implemented in a variety of different forms and are not limited to the embodiments described herein.

[0028] The present invention will be described in detail below with reference to the drawings.

[0029] Figure 1 is a diagram illustrating the implementation of a neural network system according to one embodiment.

[0030] Referring to Figure 1, we can see the trained neural network 10 and the device 20 on which the neural network 10 is implemented.

[0031] The training of the neural network 10 means that the weights of each layer of the neural network 10 have been determined by a large amount of training data. If the weights resulting from the training of the neural network 10 are stored in a central cloud server, then a cloud computing device using the neural network 10 can communicate with the central cloud server to send input values ​​to the neural network 10 and receive output values. In this case, even if the neural network 10 is very complex or large-scale, the output values ​​can be used without problems by the cloud computing device.

[0032] However, if device 20 is an edge computing device that processes data from itself without communicating with a central cloud server, the weights of the neural network 10 determined by learning are stored in the actual hardware, device 20, specifically in the memory cells that make up the cell array of device 20. In this case, device 20 may be a neuromorphic chip.

[0033] Neuromorphic chips are hardware that mimics the structure of the human brain by creating circuits that imitate the morphology of neurons. In other words, a neuromorphic chip is a computer chip that mimics the structure of the nervous system. Because neuromorphic chips consist only of the circuits necessary for neural network computation, they can achieve hundreds of times greater advantages in terms of power, area, and speed. Neuromorphic chips mimic the way the brain works by configuring structures that connect neurons and synapses in parallel, and can conserve energy by connecting and disconnecting when not processing data. For example, conventional computers with a von Neumann architecture process data sequentially as it is input, making them excellent for executing precisely crafted programs, but they have problems such as power consumption limitations and low efficiency in pattern recognition and real-time recognition. On the other hand, neuromorphic chips use analog operation where various states gradually change, rather than digital data such as 0s and 1s. In other words, the parallel-configured artificial neurons operate in an event-driven manner without clock operation. Therefore, they can efficiently process non-routine characters, speech, and images that are difficult for conventional computers to intuitively recognize.

[0034] In one embodiment, when input data such as images, sounds, or electromagnetic waves are input to a neuromorphic chip, the input data can be used to output predetermined output data through calculations within the neuromorphic chip. In this case, the data input to the neuromorphic chip is not limited to the images, sounds, or electromagnetic waves mentioned above, but can include various forms of data such as video and text.

[0035] One embodiment of a neuromorphic device can be realized using an Edge AI Chip. Edge AI refers to a technology that executes AI algorithms on hardware devices using edge computing based on data generated by the system. AI processing is mainly performed in cloud-based data centers that require enormous computing capacity and are highly dependent on servers. On the other hand, using Edge AI, AI algorithm calculations are performed locally, reducing dependence on the cloud (server), thereby reducing communication costs, and protecting privacy by not sending sensitive personal information to the cloud. Therefore, by configuring a neuromorphic device with an Edge AI Chip, not only are costs reduced and security improved, but calculations are processed immediately within the same hardware, resulting in a highly responsive system.

[0036] On the other hand, in the neural network 10, the state values ​​(States) of each weight can be very diverse (for example, 128 states), and the memory cells of the cell array implemented in the device 20 are formed of multi-bit (for example, 8-bit) memory cells that can store the state values ​​of the weights. On the other hand, in order to perform calculations on the data input from the device 20, the cell array of the neural network must be composed of memory cells that have state values ​​greater than or equal to the resolution of the input data, i.e., the number of bits of the input data, and the digital-to-analog converter of the device 20 must also have a resolution greater than or equal to the number of bits of the input data.

[0037] However, realizing a neural network with high-resolution memory cells can incur excessive costs, and high-resolution digital-to-analog converters occupy a large area, potentially unnecessarily increasing the size of the device 20. Therefore, considering the size and cost of the device 20, a high-precision (i.e., a method for processing data with a large number of bits) computation method is required, even if the device 20 is composed of low-resolution components.

[0038] In this specification, "cell bit resolution" is expressed in bits as the number of distinct state values ​​that a single memory cell can represent. For example, a cell bit resolution of 7 bits for a memory cell may mean that the memory cell can store any of 128 distinct state values.

[0039] In the following description, the apparatus 20 according to one embodiment of the present invention, i.e., the neural network apparatus, may be the neuromorphic apparatus described above. That is, the neuromorphic apparatus described above can function as a neural network apparatus according to one embodiment of the present invention.

[0040] Figures 2 and 3 are illustrative diagrams illustrating a comparison between a von Neumann architecture and a CIM (computing in memory) architecture according to one embodiment of the present invention.

[0041] Referring to Figure 2, the von Neumann architecture is a computer architecture proposed by John von Neumann, and is a stored-program computer architecture consisting of a typical three-tier architecture of main memory, central processing unit, and input / output device.

[0042] The von Neumann architecture has the advantage of greatly increasing versatility because when switching from computing to other tasks, only the software (programs) needs to be changed without having to rearrange the hardware (wires, etc.). However, because it consists of sequentially executing an enumerated set of instructions, each of which modifies a value in a specific memory location, it causes serious problems in the design of high-speed computers. This is known as the von Neumann bottleneck.

[0043] To solve the von Neumann bottleneck, alternatives have been proposed, such as the Harvard architecture, which divides memory into areas for storing instructions and areas for storing data; the CIM architecture, which performs not only data storage but also data computation from memory; and neuromorphic computing, which uses an artificial neural network type integrated circuit that mimics the brain structure of higher animals, forming a large number of units that integrate computation and memory functions and connecting them in parallel like a network, and then operating each unit in an event-driven manner.

[0044] Referring to Figure 3, it can be seen that the CIM architecture consists of a processor and memory with computing capabilities.

[0045] Unlike conventional von Neumann architectures, where all data in memory is moved to the processor for computation, the CIM architecture performs calculations in memory when a processor instruction is received, and only transfers the result data to the processor. This avoids the movement of large amounts of data, effectively resolving the aforementioned von Neumann bottleneck. It also has the advantage of significantly lower power consumption.

[0046] A neural network device according to one embodiment of the present invention can perform calculations using only on-chip memory without using external memory. For example, the neural network can perform calculations without memory updates during input signal processing by performing calculations for each layer on a CIM basis using only on-chip memory, without using external memory (e.g., off-chip memory). Specifically, the neural network device can perform CIM-based calculations with each memory cell and processor directly connected.

[0047] However, CIM-based AI chips perform calculations directly within internal memory without exchanging data with external memory, eliminating the bottleneck caused by data movement between conventional memory and computing units. This allows CIM-based AI chips to fundamentally solve memory bandwidth problems. Furthermore, this structure offers the advantages of reduced power consumption and minimized heat generation. The cell array of a neural network device according to one embodiment of the present invention can be configured with multi-bit feasible memory to maximize the computation of such a CIM architecture. For example, the neural network device can be configured with 7 bits (128 analog memory states) of feasible memory. By configuring the neural network with a large capacity, unlike typical CIM chips which suffer from heat generation and performance degradation, it is possible to process vast amounts of data with low power consumption and high performance even during prolonged use.

[0048] On the other hand, on-chip memory can be implemented using a cell array. That is, a cell array can receive instructions from a processor and perform calculations, and CIM calculations can be achieved by integrating the memory cells of the cell array into on-chip memory. For example, a processor can receive an input signal and drive a neural network device trained on predetermined training data to obtain an output signal.

[0049] Figure 4 shows a neural network device according to one embodiment of the present invention.

[0050] Neural network devices can be implemented on various types of devices, including personal computers (PCs), server devices, mobile devices, and embedded devices. Specific examples include, but are not limited to, smartphones, tablet devices, augmented reality (AR) devices, internet of things (IoT) devices, autonomous vehicles, robotics, and medical devices that utilize neural networks for speech recognition, image recognition, and image classification. Furthermore, neural network devices can be compatible with dedicated hardware accelerators (HW accelerators) installed in the aforementioned devices. These hardware accelerators may include, but are not limited to, NPUs (neural processing units), TPUs (Tensor Processing Units), and Neural Engines, which are dedicated modules for driving neural networks.

[0051] The neural network device may include a digital-to-analog converter 1, a cell array 2, an analog-to-digital converter 3, and a processor 4. The neural network device shown in Figure 4 only shows components relevant to this embodiment, and it will be obvious to those of the art that the neural network device may further include other general-purpose components in addition to those shown in Figure 4.

[0052] A neural network device according to one embodiment may include a digital-to-analog converter 1.

[0053] A digital-to-analog converter 1 according to one embodiment can convert an input signal having a digital value into an analog signal. For example, the analog signal may be a voltage or a current. That is, the digital-to-analog converter 1 can convert a digital input into an analog input of either voltage or current. As an example, the digital-to-analog converter 1 can receive a digital voltage composed of multiple bits, convert it into an analog voltage corresponding to the number of bit lines, and apply the analog voltage to multiple bit lines.

[0054] A neural network device according to one embodiment may include a cell array 2 that includes a plurality of memory cells arranged on a plurality of bit lines and a plurality of word lines.

[0055] In one embodiment, multiple word lines of the cell array 2 are connected to a digital-to-analog converter 1, and the digital input can be converted to an analog input from the digital-to-analog converter 1.

[0056] As described above, multiple memory cells can store the weights of a neural network. For example, when an analog input is input through each of the multiple word lines of cell array 2, a MAC (multiply and accumulate) operation is performed with the neural network weights stored in the multiple memory cells, and an analog output can be output through each of the multiple bit lines. In this case, similar to the analog input, the analog output can be either a current or a voltage signal.

[0057] A neural network device according to one embodiment may include an analog-to-digital converter 3.

[0058] According to one embodiment, the analog-to-digital converter 3 is connected to multiple bit lines of the cell array 2 and can receive analog outputs.

[0059] An analog-to-digital converter 3 according to one embodiment can convert an analog output into a digital output having a digital value. That is, the analog-to-digital converter 3 can convert either an analog output of voltage or current into a digital output. As an example, the analog-to-digital converter 3 can receive analog voltages output from multiple bit lines and convert them into a digital output having a predetermined number of bits.

[0060] In one embodiment, the processor 4 is electrically connected to the digital-to-analog converter 1 and the analog-to-digital converter 3, and can perform control over digital inputs and digital outputs. Specifically, the processor 4 can control digital inputs based on input signals or control output signals based on digital outputs.

[0061] Figures 5A and 5B are diagrams illustrating the operation method of a cell array according to one embodiment.

[0062] Referring to Figure 5A, the cell array can include multiple memory cells 530. Here, each memory cell 530 may be an element whose electrical conductivity or weight changes depending on the electrical pulses applied to its terminals, such as voltage or current. For example, each memory cell 530 may be an RCA (Resistive Crossbar Memory Array), or a multi-level memory such as ReRAM (Resistive RAM), FeRAM (Ferroelectric RAM), PRAM (Phase-change RAM), MRAM (Magnetic RAM), or NAND / NOR flash memory.

[0063] In one embodiment, the cell array can provide wiring 512 extending in a first direction (e.g., horizontal direction) and wiring 522 extending in a second direction (e.g., vertical direction) intersecting the first direction. For the sake of explanation, the wiring 512 extending in the first direction will be referred to as a row line, and the wiring 522 extending in the second direction will be referred to as a column line. Multiple memory cells 530 are arranged at each intersection of the row lines 512 and the column lines 522, and the corresponding row lines 512 and the corresponding column lines 522 can be connected.

[0064] The memory cell 530 can be realized to have various characteristics, such as exhibiting analog behavior in which there is no abrupt change in resistance during set and reset operations, and the conductivity changes gradually according to the number of electrical pulses input. Specifically, the processor of the neural network device can apply a rapidly changing voltage to the memory cell 530. This makes it possible to gradually change the resistance value or weight of the memory cell 530.

[0065] The operation of the above cell array can be explained as follows with reference to Figure 5B. For the sake of explanation, the row wirings 512 can be referred to from top to bottom as the first row wiring 512A, the second row wiring 512B, the third row wiring 512C, and the fourth row wiring 512D, and the column wirings 522 can be referred to from left to right as the first column wiring 522A, the second column wiring 522B, the third column wiring 522C, and the fourth column wiring 522D.

[0066] Referring to Figure 5B, in the initial state, all of the multiple memory cells 530 may be in a state of relatively low conductivity, i.e., a high-resistance state. If at least some of the multiple memory cells 530 are in a low-resistance state, further initialization operations may be required to bring them into a high-resistance state. Each of the multiple memory cells 530 may have a predetermined threshold required for a change in resistance and / or conductivity. More specifically, when a voltage or current smaller than the predetermined threshold is applied across each memory cell 530, the conductivity of the memory cell 530 remains unchanged, while when a voltage or current greater than the predetermined threshold is applied to the memory cell 530, the conductivity of the memory cell 530 may change.

[0067] In this state, an input signal corresponding to the specific data (or an analog input obtained by converting the input signal) can be entered into the row wiring 512 in order to perform the operation of outputting specific data as the result of a specific column wiring 522. For example, the input signal can appear as the application of an electrical pulse to each of the row wirings 512. Furthermore, the column wirings 522 can be driven with an appropriate voltage or current for output.

[0068] For the sake of explanation, the following will use a single-bit (1-bit) operation as an example. In one example, if a column wiring 522 that outputs specific data has already been defined, this column wiring 522 can be driven so that the memory cell 530 located at the intersection with the row wiring 512 corresponding to "1" is supplied with a voltage greater than or equal to the voltage required during set operation (hereinafter referred to as the set voltage), and the remaining column wirings 522 can be driven so that the remaining memory cells 530 are supplied with a voltage less than the set voltage. For example, if the magnitude of the set voltage is Vset, and the column wiring 522 that outputs the data "0011" is defined as the third column wiring 522C, then the size of the electrical pulses applied to the third and fourth row wirings 512C and 512D can be greater than or equal to Vset so that the first and second memory cells 530A and 530B located at the intersection of the third column wiring 522C and the third and fourth row wirings 512C and 512D are supplied with a voltage greater than or equal to Vset, and the voltage applied to the third column wiring 522C can be 0V. Therefore, the first and second memory cells 530A and 530B can be in a low-resistance state. The conductivity of the first and second memory cells 530A and 530B in the low-resistance state can gradually increase as the number of electrical pulses increases. The size and width of the applied electrical pulses can be substantially constant. The voltage applied to the remaining column wiring, i.e., the first, second and fourth column wirings 522A, 522B, and 522D, can have a value between 0V and Vset, for example, a value of 1 / 2Vset, so that the remaining memory cells 530 other than the first and second memory cells 530A and 530B are subjected to a voltage smaller than Vset. Therefore, the resistance state of the remaining memory cells 530 other than the first and second memory cells 530A and 530B does not have to change.

[0069] As another example, it is not necessary to define a column wiring 522 that outputs specific data. In this case, by applying an electrical pulse corresponding to the specific data to the row wiring 512 and measuring the current flowing through each of the column wirings 522, the column wiring 522 that first reaches a predetermined threshold current, for example, the third column wiring 522C, can become the column wiring 522 that outputs this specific data.

[0070] In this way, different data can be output to different column wirings 522.

[0071] On the other hand, the row wiring 512 of the cell array described above can represent word lines, and the column wiring 522 of the cell array can represent bit lines.

[0072] Figures 6A and 6B are diagrams illustrating a comparison between the product of matrices and vectors and operations performed on a cell array, according to one embodiment.

[0073] First, referring to Figure 6A, the convolution operation between the input data and the kernel can be performed using matrix-vector multiplication. For example, the input data can be represented by matrix X610, and the weight values ​​can be represented as the kernel by matrix W611. The output data can be represented by matrix Y612, which is the result of multiplying matrix X610 and matrix W611.

[0074] Referring to Figure 6B, a vector multiplication operation can be performed using multiple memory cells in a cell array. In comparison to Figure 6A, the input data may be received as the input value of a memory cell, and the input value may be a voltage of 620. Furthermore, the weight values ​​may be stored in the core synapses, i.e., memory cells, and the weight values ​​stored in the memory cells may be conductances of 621. Therefore, the output value of the memory cell can be represented by a current of 622, which is the result of the multiplication operation between voltage 620 and conductance 621.

[0075] Figure 7 illustrates an example in which a convolution operation is performed in a cell array according to one embodiment.

[0076] In one embodiment, the neural network device can receive an input signal 710. In this case, the input signal 710 may be a digital input having a digital value. The input signal 710 can be converted to an analog input 701 via a digital-to-analog converter 720. Furthermore, the converted analog input 701 can be input to multiple word lines of a core 700, which is implemented as at least part of a cell array.

[0077] Furthermore, the core 700 can store learned kernel values ​​in multiple memory cells. For example, kernel values ​​stored in multiple memory cells may be conductances 702. In this case, the cell array can calculate an output value by performing a vector multiplication operation between the analog input 701 and the conductance 702, and the output value can be represented as an analog output 703 (e.g., a current value).

[0078] Since the analog output 703 (e.g., current) output from core 700 is an analog signal, the analog output 703 can be converted to a digital input via analog-to-digital converter 730 for use as input data for other cores 750 in the cell array. The cell array can use the analog-to-digital converter 730 to convert the analog output 703 to a digital signal. In one embodiment, the neural network device can use the analog-to-digital converter 730 to convert the analog output 703 to a digital signal having the same bit resolution as the number of bits of the input signal 710. For example, if the number of bits of the input signal 710 is 1 bit resolution, the neural network device can use the analog-to-digital converter 730 to convert the analog output 703 to a 1 bit resolution digital signal.

[0079] The neural network device can use the activation unit 740 to apply an activation function to the digital signal converted by the analog-to-digital converter 730. While sigmoid, tanh, and ReLU (Rectified Linear Unit) functions can be used as activation functions, the activation functions applicable to the digital signal are not limited to these. The digital signal to which the activation function has been applied can then be used as an input value for other cores 750. When the digital signal to which the activation function has been applied is used as an input value for other cores 750, the process described above can be applied similarly to the other cores 750.

[0080] On the other hand, core 700 and the other cores 750 are not physically separated, and it can be said that the weight values ​​of the memory cells included in the cell array are changed according to the weight and / or bias values ​​of each core 700, 750.

[0081] On the other hand, the number of bits in the input signal 710 can have various bit resolution values, such as 1 bit, 4 bits, and 8 bits. In this case, the number of bits in the input signal 710 may be higher than the resolution of the memory cells contained in the cell array 700. In this case, calculations can be performed by setting the number of bits in the input signal 710 to be less than or equal to the resolution of the memory cells contained in the cell array 700, but this presents a problem in that higher precision calculations cannot be achieved. Therefore, a calculation method to solve this problem will be explained using Figure 8 and subsequent figures.

[0082] Figure 8 is a diagram illustrating the operation method of a neural network device according to one embodiment of the present invention.

[0083] Figure 8 illustrates the environment in which the processor of a neural network device generates output signals based on input signals. That is, Figure 8 shows only the components (e.g., cell array 2) necessary to explain the execution process of the processor, but is not limited to these and may include other components that have been omitted.

[0084] On the other hand, the processor may be a component of the control logic (not shown) included in the neural network device. Alternatively, the processor may be a component provided separately from the control logic (not shown).

[0085] In one embodiment, the processor can generate one or more digital inputs 811, 812 based on the input signal 800. For example, the processor can input one or more digital inputs, including at least a portion of the input signal 800, to the digital-to-analog converter (not shown) based on the number of bits of the input signal 800 and the bit resolution of the digital-to-analog converter (not shown) (hereinafter referred to as "DAC bit resolution").

[0086] For example, the processor may input the input signal 800 as digital inputs 811, 812 to a digital-to-analog converter (not shown) based on the bit depth and DAC bit resolution of the input signal 800, or it may generate a plurality of digital inputs 811, 812 that include at least a portion of the input signal 800 based on the bit depth and DAC bit resolution of the input signal 800, and input the generated plurality of digital inputs 811, 812 to a digital-to-analog converter (not shown).

[0087] Figure 9 is a diagram illustrating a digital input according to one embodiment of the present invention.

[0088] Referring to Figure 9, the processor can input two or more digital inputs, including at least a portion of the input signal 900, to the digital-to-analog converter in response to the number of bits in the input signal 900 exceeding the DAC bit resolution. For example, in response to the input signal 900 having 16 bits and the DAC bit resolution being 8 bits, the processor can input two or more digital inputs, including at least a portion of the input signal 900, to the digital-to-analog converter.

[0089] In one embodiment, two or more digital inputs may include an upper bit sequence 931 corresponding to the upper bit 921 of the input signal 900 and a lower bit sequence 932 corresponding to the lower bit 922 of the input signal 900. For example, if the input signal 900 is 16-bit data, the upper bit sequence 931 may be a bit sequence corresponding to the upper 8 bits of the input signal 900, and the lower bit sequence 932 may be a bit sequence corresponding to the lower 8 bits of the input signal 900. As another example, if the input signal 900 is 16-bit data, the upper bit sequence 931 may be a bit sequence corresponding to the upper 10 bits of the input signal 900, and the lower bit sequence 932 may be a bit sequence corresponding to the lower 6 bits of the input signal 900. However, in this case, as will be described later, the DAC bit resolution must be 10 bits or more.

[0090] On the other hand, the multiple bits that constitute the upper bit 921 of the input signal 900 have values ​​that are, on average, positioned higher than the multiple bits that constitute the lower bit 922 within the input signal 900. In one embodiment, the upper bit 921 of the input signal 900 is a bit sequence from the first bit positioned higher than the reference bit 901 to the reference bit 901, and the lower bit 922 may be a bit sequence from any of the multiple bits of the upper bit 921 and the bit 902 following the reference bit 901 to the second bit positioned lower than the bit 902 following the reference bit 901.

[0091] In one embodiment, if the DAC bit resolution is n and the most significant bit of the input signal 900 is the first bit, then the reference bit 901 may be the nth bit of the input signal 900. In another embodiment, the reference bit 901 may be either the most significant bit of the input signal 900 or the nth bit.

[0092] In one embodiment, the upper bit sequence 931 may be obtained by shifting the upper bit 921 of the input signal 900 to the right by a first bit length 910 such that the least significant bit (LSB) of the upper bit sequence 931 aligns with the least significant bit of the input signal 900.

[0093] Here, the least significant bit refers to the bit located at the lowest position in the bit sequence. For example, the least significant bit of the bit sequence (1, 1, 1, 1, 1, 1, 1, 0) can be 0. Furthermore, the alignment of the least significant bit of the upper bit sequence 931 with the least significant bit of the input signal 900 can mean removing the remaining bits of the input signal 900, except for the upper bit 921, so that the least significant bit of the upper bit sequence 931 can be in the same position as the least significant bit of the input signal 900.

[0094] In one embodiment, the upper bit sequence 931 is the bit sequence from the most significant bit (MSB) of the input signal 900 to the reference bit 901, and the lower bit sequence 932 can be the bit sequence from any of the upper bits 921 of the input signal 900 and the next bit 902 after the reference bit 901 to the least significant bit of the input signal 900.

[0095] Here, the most significant bit refers to the bit located at the highest position in the bit sequence. For example, in the bit sequence (1, 0, 0, 0, 0, 0, 0), the most significant bit can be 1.

[0096] For example, if the lower bit sequence 932 is the bit sequence from the bit 902 following the reference bit 901 to the least significant bit of the input signal 900, then the first bit length 910 can be the number of bits in the lower bit sequence 932. That is, if the lower bit sequence 932 is the bit sequence from the bit 902 following the reference bit 901 to the least significant bit of the input signal 900, then the upper bit sequence 931 and the lower bit sequence 932 are generated by dividing the input signal 900 without any overlapping bits, so the first bit length 910 can be the number of bits in the lower bit sequence 932. Figure 9 shows the case where the input signal 900 is divided in half to generate the upper bit sequence 931 and the lower bit sequence 932, but the number of bits in the upper bit sequence 931 and the lower bit sequence 932 may differ depending on the position of the reference bit 901.

[0097] For example, if the input signal 900 has 16 bits (1, 1, 0, 1, 0, 0, 0, 1, 1, 1, 1, 1, 0, 0, 1, 0) and the reference bit is 1 in the 9th position, then the upper bit sequence 931 is (1, 1, 1, 1, 0, 0, 0, 1) obtained by removing the remaining bits (1, 1, 1, 1, 0, 0, 1, 0) from the upper bit portion of the input signal 900, and the lower bit sequence 932 can be the bit sequence from 1, the bit following the reference bit, to the least significant bit (1, 1, 1, 1, 0, 0, 1, 0).

[0098] On the other hand, the number of bits of the digital input may be less than or equal to the DAC bit resolution. That is, according to one embodiment of the present invention, even if an input signal with a higher number of bits than the DAC bit resolution is input, the processor generates two or more digital inputs (e.g., 931, 932) with a number of bits less than or equal to the DAC bit resolution, thereby enabling input to the digital-to-analog converter.

[0099] Returning to Figure 8, one or more digital inputs 811, 812 can be input to multiple word lines 821 of the cell array 2 via a digital-to-analog converter (not shown). One or more digital inputs 811, 812 input to multiple word lines 821 of the cell array 2 can be output to an analog output through calculations with neural network weights stored in multiple memory cells 823. This analog output can be output via the bit line 822 of the cell array 2. That is, the processor can obtain digital outputs 831, 832 corresponding to the output of the bit line 822. At this time, digital outputs 831, 832 corresponding to digital inputs 811, 812 can be obtained, respectively. For example, digital output 831 corresponding to digital input 811 can be obtained, and digital output 832 corresponding to digital input 812 can be obtained sequentially.

[0100] For example, the first analog output, which is output first via bit line 822, may be converted to digital output 831 via an analog-to-digital converter, and the second analog output, which is output subsequently via bit line 822, may be converted to digital output 832 via an analog-to-digital converter. The processor can then sequentially acquire the digital outputs 831 and 832 output via the analog-to-digital converters.

[0101] In one embodiment, when digital input 811 is the upper bit sequence according to the above embodiment and digital input 812 is the lower bit sequence according to the above embodiment, the processor inputs the upper bit sequence 811 and the lower bit sequence 812 to a digital-to-analog converter (not shown) and receives a digital output 831 corresponding to the upper bit sequence 811 and a digital output 832 corresponding to the lower bit sequence 812, which are output via an analog-to-digital converter (not shown). In this case, in the following description, the digital output 831 corresponding to the upper bit sequence 811 can be defined as the upper bit output, and the digital output 832 corresponding to the lower bit sequence 812 can be defined as the lower bit output.

[0102] In one embodiment, the processor can generate an output signal 840 using at least one of the digital outputs 831, 832 corresponding to the output of the bit line 822, based on the number of bits of the input signal 800 to be acquired and the bit resolution of the plurality of memory cells 823 (hereinafter referred to as "cell bit resolution").

[0103] Figure 10 is a diagram illustrating a method for generating an output signal according to one embodiment of the present invention.

[0104] Referring to Figure 10, the upper bit output 1021 and lower bit output 1022 according to one embodiment are shown. In one embodiment, the processor can shift the upper bit output 1021 to the left by a first bit length of 1010. For example, if the upper bit output 1021 is (1, 1, 0, 1, 0, 0, 0, 1), the upper bit output 1031, shifted to the left by a first bit length of 1010, may be (1, 1, 0, 1, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0). Here, the lower 8 bits "0" can represent the actual value 0, but can also be replaced with meaningless values ​​or null values.

[0105] In one embodiment, the processor can generate an output signal 1000 based on the moved upper bit output 1031 and lower bit output 1032. For example, if the moved upper bit output 1031 is (1, 1, 0, 1, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0) and the lower bit output 1032 is (1, 1, 1, 1, 0, 0, 1, 0), the processor can combine the upper bit output 1031 and the lower bit output 1032 to generate an output signal 1000 of (1, 1, 0, 1, 0, 0, 0, 1, 1, 1, 1, 1, 0, 0, 1, 0). That is, the output signal 1000 can be generated by embedding the lower bit output 1032, which corresponds to the first bit length 1010, into the lower bit of the upper bit output 1031, which corresponds to the first bit length 1010.

[0106] Returning to Figure 9, as an example, if the lower bit sequence 932 is the bit sequence from the bit 902 following the reference bit 901 to the least significant bit of the input signal 900, then the first bit length 910 can be the number of bits in the lower bit sequence 932. That is, if the lower bit sequence 932 is the bit sequence from the bit 902 following the reference bit 901 to the least significant bit of the input signal 900, then the upper bit sequence 931 and the lower bit sequence 932 divide the input signal 900 without any overlapping bits, so the first bit length 910 can be the number of bits in the lower bit sequence 932.

[0107] Figure 9 shows the case where the input signal 900 is split in half to generate the upper bit sequence 931 and the lower bit sequence 932. The number of bits in the upper bit sequence 931 and the lower bit sequence 932 may vary depending on the position of the reference bit 901.

[0108] Figure 11 is a diagram illustrating a digital input according to another embodiment of the present invention.

[0109] In one embodiment, if the lower bit sequence 1132 is a bit sequence from any of the multiple bits of the upper bit 1121 of the input signal 1100 to the least significant bit of the input signal 1100, the first bit length 1110 can be the number of bits in the lower bit sequence 1132 minus the number of overlapping bits 1111 in the upper bit sequence 1131 and the lower bit sequence 1132.

[0110] For example, in order for the upper bit sequence 1131 and the lower bit sequence 1132 not to overlap, the processor must determine the bit sequence from the bit 1102 following the reference bit 1101 to the least significant bit of the input signal 1100 as the lower bit sequence 1132. Conversely, the processor may determine that the upper bit sequence 1131 and the lower bit sequence 1132 overlap.

[0111] For example, if the upper bit sequence 1131 is the bit sequence from the most significant bit of the input signal 1100 to the reference bit 1101, and the lower bit sequence 1132 is the bit sequence from any of the multiple bits of the upper bit 1121 of the input signal 1100 to the least significant bit of the input signal 1100, then the upper bit sequence 1131 and the lower bit sequence 1132 can have overlapping bits 1111. Specifically, as shown in Figure 11, if the lower bit sequence 1132 is the bit sequence from the bit before the reference bit 1101 to the least significant bit of the input signal 1100, then the upper bit sequence 1131 and the lower bit sequence 1132 can have two overlapping bits 1111.

[0112] Figure 12 is a diagram illustrating a method for generating an output signal according to another embodiment of the present invention.

[0113] Referring to Figure 12, an embodiment is shown in which, when the upper bit sequence and lower bit sequence have overlapping bits as in Figure 11, the processor generates an output signal 1200 based on the upper bit output 1221 corresponding to the upper bit sequence and the lower bit output 1222 corresponding to the lower bit sequence.

[0114] In one embodiment, the processor can shift the upper bit output 1221 corresponding to the upper bit sequence to the left by a first bit length 1210. As described above, the first bit length 1210 may be the number of bits in the lower bit sequence minus the number of overlapping bits in the upper and lower bit sequences.

[0115] In one embodiment, the processor can generate an output signal 1200 based on the moved upper bit output 1231 and lower bit output 1232. The specific method by which the processor generates the output signal 1200 based on the moved upper bit output 1231 and lower bit output 1232 is the same as in Figure 10, however there may be problems with how the processor handles the overlapping bit 1211 when combining the upper bit output 1221 and lower bit output 1222 in that embodiment.

[0116] In one embodiment, the overlapping bits 1211 of the upper bit output 1221 and the overlapping bits 1211 of the lower bit output 1222 do not have to have the same value. As a result of the cell array calculation on specific data, a predetermined error in the lower bits tends to be ignored. However, in the case of the upper bit output 1221, it becomes the upper bit of the output signal 1200 when generating the output signal 1200, so it is necessary to preserve the lower bits of the upper bit output 1221. Therefore, as explained with reference to Figure 11, the processor generates the upper bit example and the lower bit example so that there are overlapping bits between the upper bit example and the lower bit example, and the lower bits of the upper bit output 1221 can be preserved as described above by replacing the overlapping bits 1211 of the upper bit output 1221 corresponding to the upper bit sequence with the value of the overlapping bits 1211 of the lower bit output 1222 corresponding to the lower bit sequence.

[0117] However, the method for handling duplicate bits 1211 is not limited to this; the output signal 1200 can also be determined based on the difference between the duplicate bit 1211 of the upper bit output 1221 and the duplicate bit 1211 of the lower bit output 1222. For example, the corresponding portion of the output signal 1200 can be determined by the average value of the duplicate bit 1211 of the upper bit output 1221 and the duplicate bit 1211 of the lower bit output 1222.

[0118] Figure 13 is a diagram illustrating the operation method of a neural network device according to another embodiment of the present invention.

[0119] Figure 13 illustrates the environment in which the processor of a neural network device generates an output signal 1350 based on an input signal 1300. That is, Figure 13 shows only the components (e.g., cell array 2) necessary to explain the execution process of the processor, but is not limited to these and may include other components that have been omitted.

[0120] In one embodiment, the processor can generate an output signal 1350 by combining two or more combinations of any two digital outputs 1331 and 1332 generated based on the input signal 1300. Specifically, in response to the number of bits in the input signal 1300 exceeding the cell bit resolution, the processor can generate an output signal 1350 by combining two or more combinations of any two digital outputs 1331 and 1332 corresponding to the outputs of bit lines 1310A, 1310B, 1320A, and 1320B. That is, in response to the number of bits in the input signal 1300 and the DAC bit resolution being 16 bits and the cell bit resolution being 8 bits, the processor can generate an output signal 1350 by combining two or more combinations of any two digital outputs 1331 and 1332 corresponding to the outputs of bit lines 1310A, 1310B, 1320A, and 1320B.

[0121] On the other hand, the number of bits of digital outputs 1331 and 1332 may be the same as or lower than the cell bit resolution. In other words, according to one embodiment of the present invention, the processor can combine two or more combinations of any two digital outputs 1331 and 1332, each having a number of bits less than or equal to the cell bit resolution, to generate an output signal with a number of bits higher than the cell bit resolution.

[0122] In one embodiment, the cell array 2 may consist of a pair of first bit lines 1310A to which a first memory cell storing weights corresponding to the higher bits of the output signal 1350 is connected, and second bit lines 1310B to which a second memory cell storing weights corresponding to the lower bits of the output signal 1350 is connected. Furthermore, in one embodiment, the digital outputs 1331 and 1332 corresponding to the outputs of bit lines 1310A and 1310B may include a higher bit output 1331 corresponding to the output of the first bit line 1310A and a lower bit output 1332 corresponding to the output of the second bit line 1310B. Thus, the processor can generate the output signal 1350 based on the digital output 1331 output by the first bit line 1310A and the digital output 1332 output by the second bit line 1310B.

[0123] Figure 14 is a diagram illustrating a method for generating an output signal according to another embodiment of the present invention.

[0124] Referring to Figure 14, the upper bit output 1421 and lower bit output 1422 according to one embodiment are shown. In one embodiment, the processor can shift the upper bit output 1421 to the left by a second bit length 1410. For example, in an embodiment in which the processor shifts the upper bit output 1421 to the left by a second bit length 1410, the same method as in the embodiment in which the processor shifts the upper bit output to the left by a first bit length can be applied, as described above with reference to Figure 10.

[0125] Specifically, the processor can shift the upper bit output 1421 to the left by a second bit length of 1410 so that its most significant bit aligns with the most significant bit of the output signal 1400, and then generate the output signal 1400 based on the shifted upper bit output 1431 and lower bit output 1432. For example, the processor can combine the shifted upper bit output 1431 and lower bit output 1432 to generate the output signal 1400.

[0126] Returning to Figure 13, in one embodiment, the first bit line 1310A stores the upper bit weights from the most significant bit of the output signal 1350 to the reference bit (not shown), and the second bit line 1310B can store the lower bit weights from any of the multiple upper bits of the output signal 1350 and the bit following the reference bit (not shown) to the least significant bit of the output signal 1350. Therefore, the output of the first bit line 1310A may be an upper bit output (e.g., digital output 1331), and the output of the second bit line 1310B may be a lower bit output (e.g., digital output 1332).

[0127] Referring to Figure 14, in one embodiment, if the second bit line stores the lower bit weights from the bit following the reference bit (not shown) to the least significant bit of the output signal 1400, the second bit length 1410 could be the number of bits in the lower bit output 1422. In this regard, the principle described above can be directly applied via Figure 10.

[0128] Figure 15 is a diagram illustrating a method for generating an output signal according to another embodiment of the present invention.

[0129] In one embodiment, if the second bit line stores the lower bit weights from any of the multiple upper bits of the output signal 1500 down to the least significant bit of the output signal 1500, the second bit length 1510 may be the number of bits in the lower bit output 1522 minus the number of overlapping bits 1511 in the upper bit output 1521 and the lower bit output 1522.

[0130] In one embodiment, the processor can calculate the result of a calculation between the upper bit sequence and the weights of the neural network stored in multiple memory cells of the cell array (hereinafter referred to as the "correct output"). The processor can define the difference between the correct output and the upper bit output 1521 as the residual error and reflect the residual error in the weight of the second bit line that outputs the lower bit output 1522. That is, the processor can save the duplicate bits 1511 of the upper bit output 1521 by replacing the duplicate bits 1511 of the upper bit output 1521 with the duplicate bits 1511 of the lower bit output 1522.

[0131] The method by which the processor combines the moved upper bit output 1531 and lower bit output 1532 to generate the output signal 1500 can be achieved by applying the same principle as the method described above, as shown in Figure 14, and the details are omitted.

[0132] Figure 16 is a flowchart illustrating the operation method of a neural network device according to one embodiment of the present invention.

[0133] Referring to Figure 16, in step 1610, the neural network device can generate one or more digital inputs that include at least a portion of the input signal, based on the number of bits in the input signal and the DAC bit resolution of the digital-to-analog converter.

[0134] In one embodiment, the neural network device can generate two or more digital inputs, including at least a portion of the input signal, and input them to the digital-to-analog converter in response to the number of bits in the input signal exceeding the DAC bit resolution of the digital-to-analog converter.

[0135] In one embodiment, the number of bits of the digital input may be the same as or higher than the DAC bit resolution.

[0136] In one embodiment, two or more digital inputs may include a high-order bit sequence corresponding to the high-order bits of the input signal and a low-order bit sequence corresponding to the low-order bits of the input signal.

[0137] In one embodiment, the upper bit sequence may be obtained by shifting the upper bits of the input signal to the right by a first bit length such that the least significant bit of the upper bit sequence aligns with the least significant bit of the input signal.

[0138] In one embodiment, the neural network device can input the upper bit sequence and the lower bit sequence to a digital-to-analog converter.

[0139] In one embodiment, the upper bit sequence is the bit sequence from the most significant bit of the input signal to the reference bit, and the lower bit sequence can be the bit sequence from any of the multiple bits of the upper bits of the input signal and the bit following the reference bit to the least significant bit of the input signal.

[0140] In one embodiment, if the lower bit sequence is a bit sequence from the bit following the reference bit to the least significant bit of the input signal, the first bit length may be the number of bits in the lower bit sequence.

[0141] In one embodiment, if the lower bit sequence is a bit sequence from any of the multiple upper bits of the input signal to the least significant bit of the input signal, the first bit length can be the number of bits in the lower bit sequence minus the number of overlapping bits between the upper and lower bit sequences.

[0142] In step 1620, the neural network device can obtain one or more digital outputs corresponding to one or more digital inputs using a cell array containing multiple memory cells that store the weights of the neural network.

[0143] In step 1630, the neural network device can generate an output signal using at least one of the digital outputs based on the number of bits in the input signal and the cell bit resolution of the multiple memory cells.

[0144] In one embodiment, the neural network device receives the upper bit output corresponding to the upper bit sequence and the lower bit output corresponding to the lower bit sequence output via an analog-to-digital converter, shifts the upper bit output to the left by a first bit length, and generates an output signal based on the shifted upper bit output and lower bit output.

[0145] In one embodiment, the neural network device can generate an output signal by combining two or more combinations of digital outputs corresponding to the output of a bit line, in response to the number of bits in the input signal exceeding the cell bit resolution of multiple memory cells. In this case, the number of bits in the digital output may be the same as or lower than the cell bit resolution.

[0146] In one embodiment, the cell array of the neural network device may consist of a pair of first bit lines that store weights corresponding to the higher bits of the output signal and second bit lines that store weights corresponding to the lower bits of the output signal.

[0147] In one embodiment, the digital output corresponding to the output of a bit line may include a higher bit output corresponding to the output of a first bit line and a lower bit output corresponding to the output of a second bit line.

[0148] In one embodiment, the neural network device can shift the upper bit output to the left by a length of two bits so that the most significant bit of the upper bit output aligns with the most significant bit of the output signal, and generate an output signal based on the shifted upper bit output and lower bit output.

[0149] In one embodiment, the first bit line can store the upper bit weights from the most significant bit of the output signal to the reference bit.

[0150] In one embodiment, the second bit line can store the lower bit weights from any of the higher bits of the output signal and the bit following the reference bit down to the least significant bit of the output signal.

[0151] In one embodiment, if the second bit line stores the lower bit weights from the bit following the reference bit to the least significant bit of the output signal, the length of the second bit may be equal to the number of bits in the lower bit output.

[0152] In one embodiment, if the second bit line stores the lower bit weights from any of the multiple upper bits of the output signal down to the least significant bit of the output signal, the second bit length can be the number of bits in the lower bit output minus the number of overlapping bits in the upper bit output and the lower bit output.

[0153] Figure 17 is a block diagram of a neural network device according to another embodiment of the present invention.

[0154] Referring to Figure 17, the neural network device (hereinafter referred to as "device") 1700 may include a communication unit 1710, a processor 1720, and a DB 1730. Only components relevant to the embodiment are shown in the device 1700 in Figure 17. Therefore, it will be understood by an ordinary person of the art that other general-purpose components may be further included in addition to the components shown in Figure 17.

[0155] The communication unit 1710 may include one or more components that enable wired / wireless communication with an external server or external device. For example, the communication unit 1710 may include at least one of a short-range communication unit (not shown), a mobile communication unit (not shown), and a broadcast receiving unit (not shown). In one embodiment, the communication unit 1710 may use at least one communication protocol from among Serial Peripheral Interface (SPI) and Universal Asynchronous Transceiver (UART). Furthermore, in one embodiment, the communication unit 1710 may communicate with sensors, external memory, and external control devices.

[0156] DB1730 is hardware that stores various data processed within the device 1700, and can store programs for processing and controlling the processor 1720.

[0157] The DB1730 can include RAM (random access memory) such as DRAM (dynamic random access memory) and SRAM (static random access memory), ROM (read-only memory), EEPROM (electrically erasable programmable read-only memory), CD-ROM, Blu-ray or other optical disc storage devices, HDD (hard disk drive), SSD (solid state drive), or flash memory.

[0158] The processor 1720 controls the overall operation of the device 1700. For example, the processor 1720 can control the input unit (not shown), display (not shown), communication unit 1710, DB1730, etc., by executing a program stored in DB1730. The processor 1720 can control the operation of the device 1700 by executing a program stored in DB1730.

[0159] The processor 1720 can control at least some of the operation of the components of the device 1700 described above in Figures 1 to 16.

[0160] The processor 1720 can be implemented using at least one of the following: ASICs (application-specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, microcontrollers, microprocessors, or other electrical units for functional execution.

[0161] In one embodiment, the device 1700 may be a server. The server can be implemented as a computer device or a group of computer devices that communicate over a network and provide instructions, code, files, content, services, etc. For example, the server may receive input signals and generate output signals.

[0162] On the other hand, embodiments of the present invention can be realized in the form of a computer program that can be executed on a computer via various components, and such a computer program can be recorded on a computer-readable medium. In this case, the medium may include magnetic media such as hard disks, floppy disks and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical recording media such as floptical disks, and hardware devices specially configured to store and execute program instructions, such as ROM, RAM, and flash memory.

[0163] On the other hand, the computer program may be specifically designed and configured for the present invention, or it may be publicly known and available to those skilled in the field of computer software. Examples of computer programs may include not only machine code, such as that produced by a compiler, but also high-level language code that can be executed by a computer using an interpreter or the like.

[0164] According to one embodiment, the methods according to various embodiments of the present disclosure may be provided in a computer program product. The computer program product may be traded as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or online (e.g., by download or upload) via an application store (e.g., Play Store®), or directly between two user devices. In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily generated in a device-readable storage medium such as the memory of the manufacturer's server, the application store's server, or an intermediary server.

[0165] Unless otherwise stated, the steps constituting the method according to the present invention may be performed in any order. The present invention is not necessarily limited to the order in which the steps are described. In the present invention, the use of all examples or exemplary terms is solely for the purpose of illustrating the invention in detail, and the scope of the present invention is not limited by such examples or exemplary terms unless otherwise limited by the claims. Furthermore, those skilled in the art will understand that various modifications, combinations, and changes may be added to the scope of the claims or their equivalents, and that design conditions and factors may be comprised of these.

[0166] Therefore, the concept of the present invention should not be limited to the embodiments described above. Not only the claims described later, but also all scopes equivalent to or modified from these claims fall within the scope of the concept of the present invention.

Claims

1. A digital-to-analog converter that converts a digital input into either a voltage or current analog input, A cell array arranged in multiple bit lines and multiple word lines, including multiple memory cells that store the weights of a neural network, which performs operations on the analog input input via the word lines and outputs an analog output of either current or voltage via the bit lines, An analog-to-digital converter that converts the aforementioned analog output to a digital output, The digital-to-analog converter and the analog-to-digital converter are electrically connected to at least one processor which performs control over the digital input and the digital output, The aforementioned at least one processor is In response to the number of bits in the input signal exceeding the DAC bit resolution of the digital-to-analog converter, two or more digital inputs, including at least a portion of the input signal, are input to the digital-to-analog converter. An output signal is generated using at least one of the digital outputs corresponding to the output of the bit line. The number of bits for each of the two or more digital inputs is less than or equal to the DAC bit resolution. The two or more digital inputs include a higher bit sequence corresponding to the higher bits of the input signal and a lower bit sequence corresponding to the lower bits of the input signal. The aforementioned upper bit sequence is obtained by shifting the upper bits of the input signal to the right by a first bit length such that the least significant bit (LSB) of the upper bit sequence aligns with the least significant bit of the input signal. In the case where the lower bit sequence is a bit sequence from any of the multiple upper bits of the input signal to the least significant bit of the input signal, the first bit length is the number of bits in the lower bit sequence minus the number of overlapping bits in the upper bit sequence and the lower bit sequence, in a neural network device.

2. The aforementioned at least one processor is The upper bit sequence and the lower bit sequence are input to the digital-to-analog converter. The upper bit output corresponding to the upper bit sequence and the lower bit output corresponding to the lower bit sequence output via the analog-to-digital converter are received. The output of the higher bits is shifted to the left by the length of the first bit, The neural network device according to claim 1, which generates the output signal based on the moved upper bit output and the lower bit output.

3. The aforementioned upper bit sequence is This is the bit sequence from the most significant bit (MSB) to the reference bit of the aforementioned input signal. The lower bit sequence is, The neural network device according to claim 1, wherein the bit sequence is from any of the multiple upper bits of the input signal and the bit following the reference bit to the least significant bit of the input signal.

4. The neural network device according to claim 3, wherein the lower bit sequence is a bit sequence from the bit following the reference bit to the least significant bit of the input signal, and the first bit length is the number of bits in the lower bit sequence.

5. A digital-to-analog converter that converts a digital input to an analog input of either voltage or current, A cell array arranged in multiple bit lines and multiple word lines, including multiple memory cells that store the weights of a neural network, which performs operations on the analog input input via the word lines and outputs an analog output of either current or voltage via the bit lines, An analog-to-digital converter that converts the aforementioned analog output to a digital output, The digital-to-analog converter and the analog-to-digital converter are electrically connected to at least one processor which performs control over the digital input and the digital output, wherein the at least one processor One or more digital inputs, including at least a portion of the input signal, are input to the digital-to-analog converter. In response to the number of bits in the input signal exceeding the cell bit resolution of the plurality of memory cells, an output signal is generated by combining any two or more combinations of digital outputs corresponding to the output of the bit line. The cell bit resolution is expressed in bits as the number of distinguishable state values ​​that each memory cell can store. The number of bits in each of the aforementioned digital outputs is less than or equal to the cell bit resolution. The digital output corresponding to the output of the bit line includes an upper bit output corresponding to the output of the first bit line to which a first memory cell storing weights corresponding to the upper bits of the output signal is connected, and a lower bit output corresponding to the output of the second bit line to which a second memory cell storing weights corresponding to the lower bits of the output signal is connected. Generating the output signal includes shifting the upper bit output to the left by a second bit length so that the most significant bit (MSB) of the upper bit output aligns with the most significant bit of the output signal, and generating the output signal based on the shifted upper bit output and the lower bit output. The first memory cell stores the upper bit weights from the most significant bit of the output signal to the reference bit, The second memory cell stores the weights of the lower bits from any of the multiple upper bits of the output signal and the bit following the reference bit down to the least significant bit (LSB) of the output signal. A neural network device in which, when the second memory cell stores the lower bit weights from any of the multiple upper bits of the output signal to the least significant bit of the output signal, the second bit length is the number of bits of the lower bit output minus the number of overlapping bits between the upper bit output and the lower bit output.

6. The cell array is The neural network device according to claim 5, wherein the first bit line and the second bit line are configured as a pair.

7. The neural network device according to claim 5, wherein the second memory cell stores the lower bit weights from the bit following the reference bit to the least significant bit of the output signal, and the second bit length is the number of bits in the lower bit output.

8. The steps include generating two or more digital inputs, including at least a portion of the input signal, in response to the number of bits in the input signal exceeding the DAC bit resolution of the digital-to-analog converter, A step of obtaining one or more digital outputs corresponding to the two or more digital inputs using a cell array containing multiple memory cells that store the weights of a neural network, The step of generating an output signal using at least one of the aforementioned digital outputs, The number of bits for each of the two or more digital inputs is less than or equal to the DAC bit resolution. The two or more digital inputs include a higher bit sequence corresponding to the higher bits of the input signal and a lower bit sequence corresponding to the lower bits of the input signal. The aforementioned upper bit sequence is obtained by shifting the upper bits of the input signal to the right by a first bit length such that the least significant bit (LSB) of the upper bit sequence aligns with the least significant bit of the input signal. A method for operating a neural network device, wherein, if the lower bit sequence is a bit sequence from any of the multiple upper bits of the input signal to the least significant bit of the input signal, the first bit length is the number of bits in the lower bit sequence minus the number of overlapping bits in the upper bit sequence and the lower bit sequence.

9. A step of generating one or more digital inputs that include at least a portion of an input signal, A cell array arranged in multiple bit lines and multiple word lines, including multiple memory cells that store the weights of a neural network, which performs calculations on analog inputs input via the word lines and outputs an analog output of either current or voltage via the bit lines; a digital-to-analog converter that converts one or more digital inputs into analog inputs; and an analog-to-digital converter that converts the analog outputs into digital outputs, to obtain two or more digital outputs corresponding to one or more digital inputs. The step includes generating an output signal by combining two or more combinations of the two or more digital outputs in response to the number of bits in the input signal exceeding the cell bit resolution of the plurality of memory cells, The cell bit resolution is expressed in bits as the number of distinguishable state values ​​that each memory cell can store. The number of bits for each of the two or more digital outputs is less than or equal to the cell bit resolution. The two or more digital outputs include a high-bit output corresponding to the output of a first bit line to which a first memory cell storing weights corresponding to the high-order bits of the output signal is connected, and a low-bit output corresponding to the output of a second bit line to which a second memory cell storing weights corresponding to the low-order bits of the output signal is connected. The step of generating the output signal is: The process includes the steps of shifting the most significant bit (MSB) of the upper bit output to the left by a second bit length so that it aligns with the most significant bit of the output signal, and generating the output signal based on the shifted upper bit output and the lower bit output. The first memory cell stores the upper bit weights from the most significant bit of the output signal to the reference bit, The second memory cell stores the weights of the lower bits from any of the multiple upper bits of the output signal and the bit following the reference bit to the least significant bit (LSB) of the output signal. A method for operating a neural network device, wherein, when the second memory cell stores the lower bit weights from any of the multiple upper bits of the output signal to the least significant bit of the output signal, the second bit length is the value obtained by subtracting the number of overlapping bits between the upper bit output and the lower bit output from the number of bits of the lower bit output.

10. A computer-readable recording medium that stores a program for causing a computer to perform the method according to claim 8 or 9.

Citation Information

Patent Citations

  • Neuromorphic device implementing neural network and operation method of the same

    KR102627460B1

  • Reconfigurable input precision in-memory computing

    US20210326110A1