Copy data in a memory system with artificial intelligence mode
Patent Information
- Application Number
- KR1020247040571
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-08-29
- Filing Date
- 2020-08-27
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2040-08-27
Smart Images

Figure 112024135371033-PAT00001_ABST
Abstract
Description
Technology Field
[0001] The present disclosure generally relates to a memory device, and more specifically, to an apparatus and method for copying data in a memory system having an artificial intelligence mode. Background Technology
[0002] Memory devices are typically provided inside computers or other electronic devices, in semiconductors or integrated circuits. There are various types of memory, including volatile and non-volatile memory. Volatile memory may require power to retain its data and includes random-access memory (RAM), dynamic random access memory (DRAM), and synchronous dynamic random access memory (SDRAM). Non-volatile memory can provide permanent data by retaining stored data when power is not supplied and may include read-only memory (ROM), electrically erasable programmable ROM (EEPROM), erasable programmable ROM (EPROM), and resistance variable memory such as phase change random access memory (PCRAM), resistive random access memory (RRAM), and magnetoresistive random access memory (MRAM).
[0003] Memory is also utilized as a storage for volatile and non-volatile data for a wide range of electronic applications. Non-volatile memory can be used, for example, in personal computers, portable memory sticks, digital cameras, cellular telephones, portable music players such as MP3 players, movie players, and other electronic devices. Memory cells can be arranged into arrays, and arrays are used in memory devices.
[0004] (Patent Document 1) U.S. Patent Application Publication US 2019 / 0205737 Brief explanation of the drawing
[0005] FIG. 1a is a block diagram of a device of a computing system type including a memory device having an artificial intelligence (AI) accelerator according to a plurality of embodiments of the present disclosure. FIG. 1b is a block diagram of a device of a computing system type including a memory system having a memory device having an AI accelerator according to a plurality of embodiments of the present disclosure. FIG. 2 is a block diagram of a plurality of registers of a memory device having an artificial intelligence (AI) accelerator according to a plurality of embodiments of the present disclosure. FIGS. 3a and 3b are block diagrams of a plurality of bits of a plurality of registers of a memory device having an AI accelerator according to a plurality of embodiments of the present disclosure. FIG. 4 is a block diagram of a plurality of blocks of a memory device having an artificial intelligence (AI) accelerator according to a plurality of embodiments of the present disclosure. FIG. 5 is a flowchart illustrating an exemplary artificial intelligence process in a memory device having an artificial intelligence (AI) accelerator according to a plurality of embodiments of the present disclosure. FIG. 6 is a flowchart illustrating an exemplary method of copying data according to a plurality of embodiments of the present disclosure. Specific details for implementing the invention
[0006] The present disclosure includes an apparatus and a method for copying data in a memory system having an artificial intelligence (AI) mode. An exemplary apparatus may include a command indicating that the apparatus operates in an AI mode, a command to perform an AI operation using an AI accelerator based on the state of a plurality of registers, and a command to copy data between memory devices performing the AI operation. The data copied between memory devices may be neural network data, activation function data, bias data, input data and / or output data related to an AI operation. The AI accelerator may include hardware, software, and / or firmware configured to perform operations associated with the AI operation (e.g., a logic operation among other operations). The hardware may include a circuit composed of an adder and / or multiplier for performing operations such as a logic operation associated with the AI operation.
[0007] The memory device may contain data stored in the error of a memory cell used by an AI accelerator to perform AI operations. Input data, along with data defining a neural network such as neuron data, activation function data, and / or bias value data, may be stored in the memory device, copied between memory devices, and used to perform AI operations. Additionally, the memory device may include a temporary block for storing a portion of the results of the AI operation and an output block for storing the results of the AI operation. The host may issue a read command to the output block, and the results of the output block may be sent to the host to complete the execution of a command requesting that the AI operation be performed.
[0008] The host and / or controller of the memory system may issue a command to copy input and / or output data between memory devices performing AI operations. For example, the memory system may copy input data from a first memory device to a second memory device so that the first memory device can use the input data in a first AI operation and the second memory device can use the input data in a second AI operation. The input data copied from the first memory device to the second memory device may be received from the host. The first memory device performing the first AI operation and the second memory device performing the second AI operation may include identical or different neural network data, activation function data, and / or bias data. The results of the first and second AI operations may be reported to the controller and / or host and compared with each other.
[0009] The host and / or controller of the memory system may issue a command to copy neural network data, activation function data, and bias data between memory devices performing AI operations. The memory system may copy neural network data, activation function data, and / or bias data from the first memory device to the second memory device so that the first memory device can use the neural network data, activation function data, and / or bias data in the first AI operation, and the second memory device can use the neural network data, activation function data, and / or bias data in the second AI operation. The results of the first and second AI operations may be reported to the host and / or controller of the memory system and compared with each other.
[0010] Each memory device of the memory system can send input data and neuron data to the AI accelerator, and the AI accelerator can perform AI operations on the input data and neuron data. The memory device can store the result of the AI operation in the memory device's temporary block. The memory device can apply bias value data to the AI accelerator and send the result of the temporary block. The AI accelerator can perform AI operations on the result of the temporary block using the bias value data. The memory device can store the result of the AI operation in the memory device's temporary block. The memory device can transmit the result of the temporary block and activation function data to the AI accelerator. The AI accelerator can perform AI operations on the result of the temporary block and / or activation function data. The memory device can store the result of the AI operation in the memory device's output block.
[0011] An AI accelerator can reduce latency and power consumption related to AI operations compared to AI operations performed on a host. AI operations performed on a host utilize data exchanged between a memory device and the host, which adds to the latency and power consumption of the AI operations. Although AI operations performed according to an embodiment of the present disclosure may be performed on a memory device using an AI accelerator and a memory array, data is not transmitted from the memory device during the performance of the AI operations.
[0012] In the following detailed description of the present disclosure, reference is made to the accompanying drawings, which constitute part of this document, and the methods by which a plurality of embodiments of the present disclosure may be implemented are illustrated. These embodiments are sufficiently described to enable a person skilled in the art to implement the embodiments of the present disclosure, and it should be understood that other embodiments may be utilized and that process, electrical and / or structural changes may be made without departing from the scope of the present disclosure. As used herein, the designator "N" indicates that a specified plurality of specific features may be included with the plurality of embodiments of the present disclosure.
[0013] As used herein, "plural" may indicate one or more of such things. For example, a plurality of memory devices may indicate one or more memory devices. Additionally, a designator such as "N" used herein, particularly in relation to reference numbers in the drawings, indicates that a specified plurality of specific features may be included in a plurality of embodiments of the present disclosure.
[0014] The drawings herein follow a numbering rule in which the first digit corresponds to the figure number and the remaining digits identify elements or components of the drawings. Similar configurations or components between different drawings may be identified by the use of similar numbers. As is understood, elements illustrated in the various embodiments herein may be added, exchanged, and / or removed to provide a plurality of additional embodiments of the present disclosure. Furthermore, the proportions and relative sizes of elements provided in the drawings are intended to illustrate various embodiments of the present disclosure and should not be used in a limiting sense.
[0015] FIG. 1a is a block diagram of a device in the form of a computing system (100) including a memory device (120) according to a plurality of embodiments of the present disclosure. The memory device (120), memory array (125-1, ..., 125-N), memory controller (122) and / or AI accelerator (124) used herein may also be considered as separate "devices".
[0016] As illustrated in FIG. 1a, a host (102) may be coupled to a memory device (120). The host (102) may be a laptop computer, a personal computer, a digital camera, a digital recording and playback device, a mobile phone, a PDA, a memory card reader, an interface hub, or another host system, and may include a memory access device, for example, a processor. A person skilled in the art will understand that "processor" may refer to one or more processors, such as a parallel processing system, multiple coprocessors, etc.
[0017] The host (102) includes a host controller (108) for communicating with the memory device (120). The host controller (108) may send commands to the memory device (120). The host controller (108) may communicate with the memory device (120), the memory controller (122) of the memory device (120), and / or the AI accelerator (124) of the memory device (120) to perform AI operations, data reading, data writing, and / or data deletion, among other operations. AI operations may include machine learning or neural network operations, which may include training operations or inference operations, or both. In some examples, each memory device (120) may represent a layer within a neural network or a deep neural network (e.g., a network with three or more hidden layers). Alternatively, each memory device (120) may be or may include a node of a neural network, and a layer of the neural network may be composed of multiple memory devices or parts of some memory devices (120). The memory device (120) can store weights (or models) for AI operations of the memory array (125).
[0018] A physical host interface may provide an interface for transmitting control, address, data, and other signals between a memory device (120) and a host (102) having a compatible receptor for the physical host interface. Signals may be communicated between the host (102) and the memory device (120) on a plurality of buses, such as a data bus and / or an address bus, for example.
[0019] The memory device (120) may include a controller (120), an accelerator (124), and a memory array (125-1,...,125-N). The memory device (120) may be, among other types of devices, a low-power double data rate dynamic random access memory such as an LPDDR5 device, and / or a graphics double data rate dynamic random access memory such as a GDDR6 device. The memory array (125-1,...,125-N) may include a plurality of memory cells such as volatile memory cells (e.g., among other types of volatile memory cells, a DRAM memory cell) and / or non-volatile memory cells (e.g., among other types of non-volatile memory cells, an RRAM memory cell). The memory device (120) may read and / or write data to the memory array (125-1,...,125-N). The memory array (125-1,...,125-N) can store data used during AI operations performed in the memory device (120). The memory array (125-1,...,125-N) can store inputs, outputs, weight matrices and bias information of the neural network, and / or activation function information used by the AI accelerator to perform AI operations in the memory device (120).
[0020] The AI accelerator (124) of the host controller (108), memory controller (122), and / or memory device (120) may include control circuits such as hardware, firmware, and / or software, for example. In one or more embodiments, the host controller (108), memory controller (122), and / or AI accelerator (124) may be an application-specific integrated circuit (ASIC) coupled to a printed circuit board that includes a physical interface. Additionally, the memory controller (122) of the memory device (120) may include a register (130). The register (130) may be programmed to provide information to the AI accelerator to perform AI operations. The register (130) may include any number of registers. The register (130) may be written to and / or read by the host (102), memory controller (122), and / or AI accelerator (124). Register (130) may provide input, output, neural network, and / or active function information for the AI accelerator (124). Register (130) may include a mode register (131) for selecting an operation mode for the memory device (120). The AI mode operation may be selected by writing a word to the register (131), which prohibits access to registers regarding normal operations of the memory device (120) and allows access to registers regarding AI operations, such as 0xAA and / or 0x2AA. Additionally, the AI mode operation may be selected using a signature that utilizes an encryption algorithm authenticated by a key stored in the memory device (120). Register (130) may also be located in the memory array (125-1,..., 125-N) and may be accessible by the controller (122).
[0021] The AI accelerator (124) may include hardware (126) and / or software / firmware (128) for performing AI operations. The hardware (126) may include an adder / multiplier (126) to perform logical operations regarding AI operations. The memory controller (122) and / or accelerator (124) may receive commands from the host (102) to perform AI operations. The memory device (120) may perform the AI operations requested by the command of the host (102) using the AI accelerator (124), data of the memory array (125-1,...,125-N), and information of the register (130). The memory device may report information of the AI operations, such as results and / or error information, back to the host (120), for example. The AI operations performed by the AI accelerator (124) may be performed without using external processing resources.
[0022] The memory array (125-1,...,125-N) may provide main memory for the memory system or may be used as additional memory or storage throughout the memory system. Each memory array (125-1,...,125-N) may include multiple memory blocks of memory cells. The blocks of memory cells may be used to store data used during AI operations performed by the memory device (120). The memory array (125-1,...,125-N) may include, for example, DRAM memory cells. The embodiments do not limit the specific type of memory device. For example, the memory device may include, among others, RAM, ROM, DRAM, SDRAM, PCRAM, RRAM, 3D XPoint, and flash memory.
[0023] For example, the memory device (120) may perform one or more inference steps or AI operations including such steps. The memory array (125) may be a layer of a neural network or each individual node, and the memory device (120) may be a layer; or the memory device (120) may be a node within a larger network. Additionally or alternatively, the memory array (125) may store data or weights, or both, to be utilized (e.g., summed) within the node. Each node (e.g., memory array (125)) may combine the input of data read from the same or different cells of the memory array (125) with the weights read from the cells of the memory array (125). The combination of weights and data may be summed within the hardware (126) or within the periphery of the memory array (125) using, for example, an adder / multiplier (127). In such cases, the summation result may be passed to an active function that is represented or instantiated within the hardware (126) or at the periphery of the memory array (125). The result may be passed to another memory device (120) or used within an AI accelerator (124) (e.g., software / firmware (128)) to train a network containing the memory device (120) or to make decisions.
[0024] A network using a memory device (120) can be capable of supervised or unsupervised learning or can be used for this purpose. It can be combined with other learning or training systems. In some cases, a training network or model is used or called with the memory device (120), and the operations of the memory device (120) are mainly or exclusively associated with inference.
[0025] The embodiment of FIG. 1a may include additional circuitry not illustrated so as not to obscure the embodiments of the present disclosure. For example, the memory device (120) may include an address circuit for latching an address signal provided through an I / O connection via an I / O circuit. The address signal may be received and decoded by a row decoder or a column decoder to access the memory array (125-1,...,125-N). A person skilled in the art will understand that the number of address input connections may depend on the density and architecture of the memory array (125-1,...,125-N).
[0026] FIG. 1b is a block diagram of a device of the type of computing system comprising a memory system having a memory device having an artificial intelligence (AI) accelerator according to a plurality of embodiments of the present disclosure. As used herein, the memory device (120-1, 120-2, 120-3, and 120-X), the controller (10) and / or the memory system (104) may also be separately considered as a “device”.
[0027] As illustrated in FIG. 1b, a host (102) may be coupled to a memory system (104). The host (102) may be a laptop computer, a personal computer, a digital camera, a digital recording and playback device, a mobile phone, a PDA, a memory card reader, an interface hub, among other host systems, and may include a memory access device such as a processor, for example. A person skilled in the art should understand that "processor" may refer to one or more processors, such as a parallel processing system, multiple coprocessors, etc.
[0028] The host (102) includes a host controller (108) that communicates with the memory system (104). The host controller (108) can send commands to the memory system (104). The memory system (104) may include the controller (104) and memory devices (120-1, 120-2, 120-3, and 120-X). The memory devices (120-1, 120-2, 120-3, and 120-X) may be the memory devices (120) described above with respect to FIG. 1a and may include an AI accelerator having hardware, software and / or firmware to perform AI operations. The host controller (108) can communicate with the controller (105) and / or memory devices (120-1, 120-2, 120-3, and 120-X) to perform AI operations, data reading, data writing, and / or data deletion among other operations. The physical host interface may provide an interface for transmitting control, address, data, and other signals between the memory system (104) and the host (102) having a compatible acceptor for the physical host interface. Signals may be communicated between the host (102) and the memory system (104) on a plurality of buses, such as a data bus and / or an address bus, for example.
[0029] The memory system (104) may include a controller (105) coupled to memory devices (120-1, 120-2, 120-3, and 120-X) via a bus (121). The bus (121) may be configured so that the full bandwidth of the bus (121) can be consumed when some or all of the memory devices of the memory system are in operation. For example, two of the four memory devices (120-1, 120-2, 120-3, and 120-X) shown in FIG. 1b may be configured to operate while utilizing the full bandwidth of the bus (121). For example, the controller (105) may send a command to a selection line (117) that can select memory devices (120-1 and 120-3) for operation during a specific period, such as the same time. The controller (105) may send a command to a selection line (119) that can select memory devices (120-2 and 120-X) for operation during a specific period such as the same time. In a plurality of embodiments, the controller (105) may be configured to send a command to the selection lines (117 and 119) to select any combination of memory devices (120-1, 120-2, 120-3, and 120-X).
[0030] In a plurality of embodiments, the command of the select line (117) may be used to select memory devices (120-1 and 120-3), and the command of the select line (119) may be used to select memory devices (120-2 and 120-X). The selected memory devices may be used during the execution of the AI operation. Data regarding the AI operation may be copied and / or transmitted between the selected memory devices (120-1, 120-2, 120-3 and 120-X) on the bus (121). For example, a first part of the AI operation may be performed in the memory device (120-1), and the output of the first part of the AI operation may be copied from the bus (121) to the memory device (120-3). The output of the first part of the AI operation may be used by the memory device (120-3) as an input for the second part of the AI operation. Additionally, neural network data, activation function data, and / or bias data regarding AI computation can be copied between memory devices (120-1, 120-2, 120-3, and 120-X) of the bus (121).
[0031] FIG. 2 is a block diagram of a plurality of registers of a memory device having an artificial intelligence (AI) accelerator according to a plurality of embodiments of the present disclosure. A register (230) may be an AI register and may include input information, output information, neural network information, and / or activation function information between other types of information for use by an AI accelerator, a controller, and / or a memory array of a memory device (e.g., the AI accelerator (124), memory controller (122), and / or memory array (125-1,..., 125-N)) of FIG. 1. The register may be read and / or written based on instructions from a host, an AI accelerator, and / or controller (e.g., the host (102), AI accelerator (124), and memory controller (122)) of FIG. 1.
[0032] Register (232-0) can define parameters regarding the AI mode of the memory device. Bits of register (232-0) can start an AI operation, restart an AI operation, indicate that the contents of the register are valid, clear the contents of the register, and / or terminate the AI mode.
[0033] Registers (232-1, 232-2, 232-3, 232-4, and 232-5) may define the size of the input used in the AI operation, the number of inputs used in the AI operation, and the start and end addresses of the inputs used in the AI operation. Registers (232-7, 232-8, 232-9, 232-10, and 232-11) may define the size of the output used in the AI operation, the number of outputs used in the AI operation, and the start and end addresses of the outputs used in the AI operation.
[0034] Registers (232-12) can be used to enable the use of input banks, neuron banks, output banks, bias banks, activation functions, and temporary banks used during AI operations.
[0035] Registers (232-13, 232-14, 232-15, 232-16, 232-17, 232-18, 232-19, 232-20, 232-21, 232-22, 232-23, 232-24, and 232-25) can be used to define the neural network used during AI computation. Registers (232-13, 232-14, 232-15, 232-16, 232-17, 232-18, 232-19, 232-20, 232-21, 232-22, 232-23, 232-24, and 232-25) can define the size, number, and location of neurons and / or layers of the neural network used during AI computation.
[0036] Registers (232-26) can enable the debug / hold mode of the AI accelerator and allow the output to be observed at a layer of the AI operation. Registers (232-26) can indicate that the active state may be applied during the AI operation and that the AI operation may advance in the AI operation (e.g., perform the next step in the AI operation). Registers (232-26) can indicate that a temporary block containing the output of a layer is valid. The data in the temporary block may be modified by the host and / or controller of the memory device, and thus the modified data may be used in the AI operation as the AI operation advances. Registers (232-27, 232-28, and 232-29) can define the layer to which the debug / hold mode stops the AI operation, changes the contents of the neural network, and / or observes the output of the layer.
[0037] Registers (232-30, 232-31, 232-32, and 232-33) can define the size of the temporary bank used in the AI operation and the start and end addresses of the temporary bank used in the AI operation. Register (232-30) can define the start and end addresses of the first temporary bank used in the AI operation, and register (232-33) can define the start and end addresses of the first temporary bank used in the AI operation. Registers (232-31, and 232-32) can define the size of the temporary bank used in the AI operation.
[0038] Registers (232-34, 232-35, 232-36, 232-37, 232-38, and 232-39) may relate to activation functions used in AI operations. Register (232-34) may enable the use of activation function blocks, activation functions for each neuron, activation functions for each layer, and external activation functions. Register (232-35) may define the start and end addresses of the location of the activation function. Registers (232-36, 232-37, 232-38, and 232-39) may define the resolution of the input (e.g., x-axis) and output (e.g., y-axis) of the activation function and / or custom-defined activation function.
[0039] Registers (232-40, 232-41, 232-42, 232-43, and 232-44) can define the size of the bias value used in the AI operation, the number of bias values used in the AI operation, and the start and end addresses of the bias value used in the AI operation.
[0040] Registers (232-45) can provide state information for AI computation and information for debug / hold mode. Registers (232-45) can enable debug / hold mode, indicate that the AI accelerator is performing AI computation, indicate that the full capability of the AI accelerator can be used, indicate that only the metric computation of the AI computation can be performed, and / or indicate that the AI computation can proceed to the next neuron and / or layer.
[0041] Registers (232-46) may provide error information regarding AI operations. Registers (232-46) may indicate that there was an error in the sequence of AI operations, an error in the algorithm of the AI operations, an error in a page of data that the ECC could not correct, and / or an error in a page of data that the ECC could correct.
[0042] Registers (232-47) may represent an activation function to be used for AI operations. Registers (232-47) may indicate that one of a plurality of predefined activation functions may be used for AI operations and / or that a custom activation function located in the block may be used for AI operations.
[0043] Registers (232-48, 232-49, and 232-50) may represent neurons and / or layers where AI operations are executed. In the case where an error occurs during the AI operation, registers (232-48, 232-49, and 232-50) represent the neuron and / or layer where the error occurred.
[0044] FIGS. 3a and 3b are block diagrams of a plurality of bits in a plurality of registers of a memory device having an artificial intelligence (AI) accelerator according to a plurality of embodiments of the present disclosure. Each register (332-0,..., 332-50) includes a plurality of bits, bits (334-0, 334-1, 334-2, 334-3, 334-4, 334-5, 334-6, and 334-7) to represent information regarding performing an AI operation.
[0045] Register (332-0) may define parameters regarding the AI mode of the memory device. Bits (334-5) of register (332-0) may be read / write bits and, when programmed to 1b, may indicate that the elaboration of the AI operation can be restarted from the beginning (360). Bits (334-5) of register (332-0) may be reset to 0b when the AI operation is restarted. Bits (334-4) of register (332-0) may be read / write bits and, when programmed to 1b, may indicate that the elaboration of the AI operation can be started (361). Bits (334-4) of register (332-0) may be reset to 0b when the AI operation is started.
[0046] Bit (334-3) of register (332-0) may be a read / write bit and may indicate that the contents of the AI register are valid (362) when programmed to 1b and invalid when programmed to 0b. Bit (334-2) of register (332-0) may be a read / write bit and may indicate that the contents of the AI register are deleted (363) when programmed to 1b. Bit (334-1) of register (332-0) may be a read-only bit and may indicate that the AI accelerator is performing an AI operation and is in use (363) when programmed to 1b. Bit (334-0) of register (332-0) may be a write-only bit and may indicate that the memory device exits AI mode (365) when programmed to 1b.
[0047] Registers (332-1, 332-2, 332-3, 332-4, and 332-5) may define the size of the input used in the AI operation, the number of inputs used in the AI operation, and the start and end addresses of the inputs used in the AI operation. Bits (334-0, 334-1, 334-2, 334-3, 334-4, 334-5, 334-6, and 334-7) of registers (332-1 and 332-2) may define the size of the input (366) used in the AI operation. The size of the input may represent the width of the input in terms of the input type and / or the number of bits, such as floating-point, integer, and / or double among other types. Bits (334-0, 334-1, 334-2, 334-3, 334-4, 334-5, 334-6, and 334-7) of registers (332-3 and 332-4) may indicate the number of inputs (367) used in the AI operation. Bits (334-4, 334-5, 334-6, and 334-7) of register (332-5) may indicate the starting address (368) of a block in the memory array of inputs used in the AI operation. Bits (334-0, 334-1, 334-2, and 334-3) of register (332-5) may indicate the ending address (369) of a block in the memory array of inputs used in the AI operation. If the start address (368) and the end address (369) are the same address, only one input block is displayed for the AI operation.
[0048] Registers (332-7, 332-8, 332-9, 332-10, and 332-11) may define the size of the output of the AI operation, the number of outputs of the AI operation, and the start and end addresses of the outputs of the AI operation. Bits (334-0, 334-1, 334-2, 334-3, 334-4, 334-5, 334-6, and 334-7) of registers (332-7 and 332-8) may define the size (370) of the output used in the AI operation. The size of the output may represent the width of the output in terms of the number of bits and the output type, such as floating-point, integer, and / or double, among other types. Bits (334-0, 334-1, 334-2, 334-3, 334-4, 334-5, 334-6, and 334-7) of registers (332-9 and 332-10) may represent the number of outputs (371) used in the AI operation. Bits (334-4, 334-5, 334-6, and 334-7) of register (332-11) may represent the starting address (372) of a block in the memory array of outputs used in the AI operation. Bits (334-0, 334-1, 334-2, and 334-3) of register (332-11) may represent the ending address (373) of a block in the memory array of outputs used in the AI operation. If the start address (372) and the end address (373) are the same address, only one output block is displayed for the AI operation.
[0049] Registers (332-12) can be used to enable the use of input banks, neuron banks, output banks, bias banks, activation functions, and temporary banks used during AI operations. Bit (334-0) of register (332-12) can activate the input bank (380), bit (334-1) of register (332-12) can activate the neural network bank (379), bit (334-2) of register (332-12) can activate the output bank (378), bit (334-3) of register (332-12) can activate the bias bank (377), bit (334-4) of register (332-12) can activate the activation function bank (376), and bits (334-5 and 334-6) of register (332-12) can activate the first temporary bank (375) and the second temporary bank (374).
[0050] Registers (332-13, 332-14, 332-15, 332-16, 332-17, 332-18, 332-19, 332-20, 332-21, 332-22, 332-23, 332-24, and 332-25) can be used to define the neural network used during the AI operation. Bits (334-0, 334-1, 334-2, 334-3, 334-4, 334-5, 334-6, and 334-7) of registers (332-13 and 332-14) can define the number of rows (381) of the matrix used in the AI operation. The bits (334-0, 334-1, 334-2, 334-3, 334-4, 334-5, 334-6, and 334-7) of the registers (332-15 and 332-16) can define the number of columns (382) of the matrix used in the AI operation.
[0051] The bits (334-0, 334-1, 334-2, 334-3, 334-4, 334-5, 334-6, and 334-7) of registers (332-17 and 332-18) may define the size (383) of the neurons used in the AI operation. The size of the neurons may represent the width of the neurons in terms of the type of input, such as floating-point, integer, and / or double among other types, and / or the number of bits. The bits (334-0, 334-1, 334-2, 334-3, 334-4, 334-5, 334-6, and 334-7) of registers (332-19, 332-20, and 322-21) may represent the number (384) of the neurons in the neural network used in the AI operation. Bits (334-4, 334-5, 334-6, and 334-7) of register (332-22) may represent the starting address (385) of a block in the memory array of neurons used in the AI operation. Bits (334-0, 334-1, 334-2, and 334-3) of register (332-5) may represent the ending address (386) of a block in the memory array of neurons used in the AI operation. If the starting address (385) and the ending address (386) are the same address, only one neuron block is represented for the AI operation. The bits (334-0, 334-1, 334-2, 334-3, 334-4, 334-5, 334-6, and 334-7) of the registers (332-23, 332-24, and 322-25) may represent the number of layers (387) of the neural network used in the AI operation.
[0052] Register (332-26) can enable the debug / hold mode and output of the AI accelerator so that they can be observed at the layer of the AI operation. Bit (334-0) of register (332-26) can indicate that the AI accelerator is in debug / hold mode and that an active function can be applied (391) during the AI operation. Bit (334-1) of register (332-26) can indicate that the AI operation can proceed (390) (e.g., to perform the next step in the AI operation). Bits (334-2) and (334-3) of register (232-26) can indicate that the temporary block where the output of the layer is located is valid (388 and 389). Since the data in the temporary block can be modified by the controller and / or host of the memory device, the modified data can be used in the AI operation as the AI operation proceeds.
[0053] Bits (334-0, 334-1, 334-2, 334-3, 334-4, 334-5, 334-6, and 334-7) of registers (332-27, 332-28, and 332-29) can define the layer to observe the output of when the debug / hold mode stops the AI operation (392).
[0054] Registers (332-30, 332-31, 332-32, and 332-33) can define the size of the temporary bank used in the AI operation and the start and end addresses of the temporary bank used in the AI operation. Bits (334-4, 334-5, 334-6, and 334-7) of register (332-30) can define the start address (393) of the first temporary bank used in the AI operation. Bits (334-0, 334-1, 334-2, and 334-3) of register (332-30) can define the end address (394) of the first temporary bank used in the AI operation. The bits (334-0, 334-1, 334-2, 334-3, 334-4, 334-5, 334-6, and 334-7) of registers (332-31 and 332-32) may define the size (395) of the temporary bank used for AI operations. The size of the temporary bank may represent the width of the temporary bank in terms of the type of input and / or the number of bits, such as floating-point, integer, and / or double among other types. The bits (334-4, 334-5, 334-6, and 334-7) of register (332-33) may define the starting address (396) of the second temporary bank used for AI operations. The bits (334-0, 334-1, 334-2, and 334-3) of register (332-34) can define the exit address (397) of the second temporary bank used in the AI operation.
[0055] Registers (332-34, 332-35, 332-36, 332-37, 332-38, and 332-39) may be associated with activation functions used in AI operations. Bit (334-0) of register (332-34) may enable the use of activation function blocks (3101). Bit (334-1) of register (332-34) may maintain AI in neurons (3100) and enable the use of activation functions for each neuron. Bit (334-2) of register (332-34) may maintain AI in layers (399) and enable the use of activation functions for each layer. Bit (334-3) of register (332-34) may enable the use of external activation functions (398).
[0056] Bits (334-4, 334-5, 334-6, and 334-7) of register (332-35) can define the starting address (3102) of the active function bank used for AI operations. Bits (334-0, 334-1, 334-2, and 334-3) of register (332-35) can define the ending address (3103) of the active function bank used for AI operations. Bits (334-0, 334-1, 334-2, 334-3, 334-4, 334-5, 334-6, and 334-7) of registers (332-36 and 332-37) can define the resolution (3104) of the input (e.g., x-axis) of the active function. The bits (334-0, 334-1, 334-2, 334-3, 334-4, 334-5, 334-6, and 334-7) of the registers (332-38 and 332-39) can define the output (e.g., y-axis) (3105) and / or resolution of the activation function for a given x-axis value of the custom activation function.
[0057] Registers (332-40, 332-41, 332-42, 332-43, and 332-44) may define the size of the bias value used in the AI operation, the number of bias values used in the AI operation, and the start and end addresses of the bias values used in the AI operation. Bits (334-0, 334-1, 334-2, 334-3, 334-4, 334-5, 334-6, and 334-7) of registers (332-40 and 332-41) may define the size (3106) of the bias value used in the AI operation. The size of the bias value may represent the width of the bias value in terms of the type of bias value, such as floating-point, integer, and / or double among other types, and / or the number of bits. Bits (334-0, 334-1, 334-2, 334-3, 334-4, 334-5, 334-6, and 334-7) of registers (332-42 and 332-43) may represent the number (3107) of bias values used in the AI operation. Bits (334-4, 334-5, 334-6, and 334-7) of register (332-44) may represent the starting address (3108) of a block in the memory array of bias values used in the AI operation. Bits (334-0, 334-1, 334-2, and 334-3) of register (332-44) may represent the ending address (3109) of a block in the memory array of bias values used in the AI operation. If the start address (3108) and the end address (3109) are the same address, only one bias value block is displayed for the AI operation.
[0058] Registers (332-45) may provide status information for AI computation and information for debug / hold mode. Bit (334-0) of registers (332-45) may enable debug / hold mode (3114). Bit (334-1) of registers may indicate that the AI accelerator is in use (3113) and is performing AI computation. Bit (334-2) of registers (332-45) may indicate that the AI accelerator is turned on (3112) and / or that the full capability of the AI accelerator can be used. Bit (334-3) of registers (332-45) may indicate that the matrix computation (3111) of the AI computation should be performed. Bit (334-4) of registers (332-45) may indicate that the AI computation can proceed (3110) and can proceed to the next neuron and / or layer.
[0059] Registers (332-46) can provide error information regarding AI operations. Bit (334-3) of register (332-46) can indicate that there was an error (3115) in the sequence of AI operations. Bit (334-2) of register (332-46) can indicate that there was an error (3116) in the algorithm of AI operations. Bit (334-1) of register (332-46) can indicate that there was an error (3117) in a page of data that the ECC could not correct. Bit (334-0) of register (332-46) can indicate that there was an error (3118) in a page of data that the ECC could correct.
[0060] Registers (332-47) may indicate an activation function used in an AI operation. Bits (334-0, 334-1, 334-2, 334-3, 334-4, 334-5, and 334-6) of registers (332-47) may indicate that one of a plurality of predefined activation functions (3120) may be used in an AI operation. Bit (334-7) of registers (332-47) may indicate that a custom activation function (3119) located in a block may be used in an AI operation.
[0061] Registers (332-48, 332-49, and 332-50) may represent neurons and / or layers executing AI operations. Bits (334-0, 334-1, 334-2, 334-3, 334-4, 334-5, 334-6, and 334-7) of registers (332-48, 332-49, and 332-50) may represent the addresses of neurons and / or layers executing AI operations. If an error occurs during an AI operation, registers (332-48, 332-49, and 332-50) may represent neurons and / or layers where the error occurred.
[0062] FIG. 4 is a block diagram of a plurality of blocks of a memory device having an artificial intelligence (AI) accelerator according to a plurality of embodiments of the present disclosure. An input block (440) is a block of a memory array in which input data is stored. The data in the input block (440) can be used as an input for an AI operation. The address of the input block (440) can be indicated in a register (5) (e.g., register (232-5) in FIG. 2 and register (332-5) in FIG. 3a). Since the embodiment may have a plurality of input blocks, it is not limited to a single input block. The data in the input block (440) can be sent from a host to the memory device. The data may be accompanied by a command indicating that an AI operation can be performed in the memory device using the data.
[0063] The output block (420) is a block of a memory array in which output data of an AI operation is stored. The data in the output block (442) can be used to store the output of the AI operation and send it to a host. The address of the output block (442) can be indicated in a register (11) (e.g., register (232-11) in FIG. 2 and register (332-11) in FIG. 3a). Since the embodiment may have multiple output blocks, it is not limited to a single output block.
[0064] Data in the output block (442) may be sent to the host upon completion and / or suspension of the AI operation. Temporary blocks (444-1 and 444-2) may be blocks of a memory array where data is temporarily stored while the AI operation is being performed. Data may be stored in the temporary blocks (444-1 and 444-2) while the AI operation iterates through the neurons and layers of the neural network used for the AI operation. The addresses of the temporary blocks (448) may be indicated in registers (30 and 33) (e.g., registers (232-30 and 232-33) in FIG. 2 and registers (332-30 and 332-33) in FIG. 3b). Since the embodiment may have multiple temporary blocks, it is not limited to two temporary blocks.
[0065] The active function block (446) is a convex part of a memory array in which the active functions of the AI operation are stored. The active function block (446) may store custom active functions and / or predefined active functions generated by the host and / or AI accelerator. The address of the active function block (448) may be indicated in a register (35) (e.g., registers (232-35) in FIG. 2 and registers (332-35) in FIG. 3b). Since the embodiment may have multiple active function blocks, it is not limited to a single active function block.
[0066] The bias value block (448) is a block of memory array in which the bias value of the AI operation is stored. The address of the bias value block (448) may be indicated in a register (44) (e.g., registers (232-44) in FIG. 2 and registers (332-44) in FIG. 3b). Since the embodiment may have multiple bias value blocks, it is not limited to a single bias value block.
[0067] The neural network blocks (450-1, 450-2, 450-3, 450-4, 450-5, 450-6, 450-7, 450-8, 450-9, and 450-10) are blocks of a memory array in which the neural network of the AI operation is stored. The neural network blocks (450-1, 450-2, 450-3, 450-4, 450-5, 450-6, 450-7, 450-8, 450-9, and 450-10) can store information about the neurons and layers used in the AI operation. The addresses of the neural network blocks (450-1, 450-2, 450-3, 450-4, 450-5, 450-6, 450-7, 450-8, 450-9, and 450-10) can be displayed in register (22) (e.g., register (232-22) in FIG. 2 and register (332-22) in FIG. 3a).
[0068] FIG. 5 is a flowchart illustrating an exemplary artificial intelligence process of a memory device having an artificial intelligence (AI) accelerator according to a plurality of embodiments of the present disclosure. In response to initiating an AI operation, the AI accelerator may write input data (540) and neural network data (550) to an input and a neural network block, respectively. The AI accelerator may perform an AI operation using the input data (540) and neural network data (550). The results may be stored in temporary banks (544-1 and 544-2). Temporary banks (544-1 and 544-2) may be used to store data while performing matrix calculations, adding bias data, and / or applying an activation function during the AI operation.
[0069] The AI accelerator receives partial result and bias value data (548) of an AI operation stored in temporary banks (544-1 and 544-2) and can perform an AI operation using partial result and bias value data (548) of an AI operation. The result can be stored in temporary banks (544-1 and 544-2).
[0070] The AI accelerator receives partial results of AI operations and activation function data (546) stored in temporary banks (544-1 and 544-2) and can perform AI operations using partial results of AI operations and activation function data (546). The results can be stored in an output bank (542).
[0071] FIG. 6 is a flowchart illustrating an exemplary method of copying data according to a plurality of embodiments of the present disclosure. The method described in FIG. 6 may be performed by a memory system including a memory device, such as the memory device (120) illustrated in FIG. 1a and 1b, for example.
[0072] In block (6150), the method may include the step of receiving a command including a first address and a second address, wherein the first address identifies a first memory device storing data for a training or inference operation, and the second address identifies a second memory device as a target for the data. The method may include the step of copying data from the first memory device to the second memory device of the memory system. For example, the first memory device may copy an output block to an input block of the second memory device. The host and / or controller may format the data to be stored in the second memory device and used for AI operations.
[0073] In block (6152), the method may include the step of transferring data from the first device to the second device in response to a command and based at least partially on the first address and the second address of the command. The method may include the step of enabling the second memory device to perform artificial intelligence (AI) operations by programming a register for the second memory device to enter an artificial intelligence (AI) mode.
[0074] In block (6154), the method may include the step of performing a training or inference operation in a second memory device using data from a first device. The method may include the step of performing an artificial intelligence (AI) operation in the second memory device using data copied from the first memory device. The method may include the step of copying data between memory devices combined together. For example, if the density of the neural network is too high to be stored in a single memory device, the input, output, and / or temporary blocks may be copied between memory devices to perform AI operations on the neural network.
[0075] Although specific embodiments have been illustrated and described herein, those skilled in the art will understand that a configuration calculated to achieve the same result may replace the specific embodiments illustrated. This disclosure is intended to cover adaptations or variations of various embodiments of this disclosure. The foregoing description should be understood as illustrative rather than limiting. Combinations of the foregoing embodiments and other embodiments not specifically described herein will be apparent to those skilled in the art upon reviewing the foregoing description. The scope of the various embodiments of this disclosure includes other applications in which the structures and methods are used. Accordingly, the scope of the various embodiments of this disclosure should be determined by reference to the appended claims, together with the full scope of equivalents to which such claims are granted.
[0076] In the foregoing detailed description, various features are grouped together into a single embodiment to simplify the disclosure. This method of disclosure should not be interpreted as reflecting an intention that the disclosed embodiments of the present disclosure must use more features than explicitly cited in each claim. Rather, as indicated by the following claim, the inventive content is less than all features of the single disclosed embodiment. Accordingly, the following claim is incorporated into the detailed description, and each claim exists as a separate embodiment in itself.
Claims
Claim 1 A device for performing AI operations in memory devices, comprising: a controller; and a plurality of memory devices coupled to the controller, wherein each of the plurality of memory devices is configured as part of a neural network and comprises a plurality of memory arrays, and the controller: provides input information associated with data stored in a first memory device among the plurality of memory devices for a training operation, an inference operation, or both; receives a command including a first address identifying the first memory device as a target for the stored data and a second address identifying the second memory device among the plurality of memory devices; transmits data from the first memory device to the second memory device in response to the command and the input information; performs the training operation, the inference operation, or both in the first memory device to generate first results using the stored data based on the command; and performs the training operation, the inference operation, or both in the second memory device to generate second results using the transmitted data based on the command; A device for performing AI operations in memory devices configured to compare the first results and the second results in the controller. Claim 2 An apparatus for performing AI operations in memory devices according to claim 1, wherein the input information comprises a set of data that provides information associated with the stored data. Claim 3 A device for performing AI operations in memory devices according to claim 1, wherein the instruction instructs the second memory device to receive the stored data and to write the stored data to the memory array of the second memory device. Claim 4 The device for performing AI operations in memory devices, wherein the command is issued by a host device according to claim 1. Claim 5 A device for performing AI operations in memory devices, wherein the input information is stored in the second memory device according to claim 1. Claim 6 A device for performing AI operations in memory devices, wherein the input information is stored in a host device according to claim 1. Claim 7 The device for performing AI operations in memory devices, wherein the controller configured to transfer the stored data from the first memory device to the second memory device comprises copying the input information for the training operation, the inference operation, or both. Claim 8 A system for performing AI operations in memory devices, comprising: a controller; and a plurality of memory devices coupled to the controller, wherein each of the plurality of memory devices is configured as part of a neural network and comprises a plurality of memory arrays, and the controller: receives a first command to execute a first training or inference operation in a first memory device among the plurality of memory devices; executes the first training or inference operation in the first memory device to generate first results based on the first command; receives a second command to copy data from the first memory device to a second memory device among the plurality of memory devices and execute a second training or inference operation in the second memory device, wherein the second command includes a first address and a second address, wherein the first address identifies the first memory device storing data associated with the first training or inference operation, and the second address identifies the second memory device as a target for the data. A system for performing AI operations in memory devices, comprising input information including a set of data providing information associated with the copied data, the first command, and the second command, wherein data is transmitted from the first memory device to the second memory device in response to the second command, and the second training operation or the inference operation is performed in the second memory device using the copied data to generate second results based on the second command, and the first results and the second results are compared with comparison results transmitted to a host device. Claim 9 A system for performing AI operations in memory devices, wherein the copied data includes data associated with the first training or inference operation, in accordance with claim 8. Claim 10 A system for performing AI operations in memory devices according to claim 8, wherein neural network data is copied from the first memory device to the second memory device. Claim 11 A system for performing AI operations in memory devices according to claim 10, wherein the neural network data is used by the first memory device to perform the first training or inference operation and is used by the second memory device to perform the second training or inference operation. Claim 12 A system for performing AI operations in memory devices, wherein, in claim 8, the first command and the second command are issued by a host device, and the input information is stored in the host device. Claim 13 A system for performing AI operations in memory devices according to claim 8, wherein the first command and the second command are issued by a host device, and the input information is stored in the second memory device. Claim 14 A system for performing AI operations in memory devices according to claim 8, wherein activation function data is copied from the first memory device to the second memory device. Claim 15 A method for performing AI operations in memory devices, comprising: receiving input information associated with data stored for a training operation, an inference operation, or both, in a first memory device among a plurality of memory devices; executing the training operation, the inference operation, or both, in the first memory device to generate first results using the stored data; receiving a command from a host device to the first memory device, the command including a first address and a second address, wherein the first address identifies the first memory device storing the data for the training operation, the inference operation, or both, and the second address identifies the second memory device among the plurality of memory devices as a target for the stored data; transmitting the stored data from the first memory device to the second memory device in response to the command and the input information, wherein the stored data includes the input information for the training operation, the inference operation, or both; executing the training operation, the inference operation, or both, in the second memory device to generate second results using the transmitted stored data; and a comparison of the first results and the second results transmitted to the host device. A method for performing AI operations on memory devices, comprising a step of comparing with results. Claim 16 A method for performing AI operations in memory devices according to claim 15, wherein the step of transferring the stored data from the first memory device to the second memory device comprises the step of copying neural network data for the training operation, the inference operation, or both. Claim 17 A method for performing AI operations in memory devices according to claim 15, wherein the step of receiving the command includes receiving a command to select the first memory device and the second memory device to transmit data on a bus shared by the plurality of memory devices. Claim 18 A method for performing AI operations in memory devices according to claim 15, wherein the step of receiving the command includes receiving a command to activate the second memory device to enter an AI mode to perform artificial intelligence (AI) operations. Claim 19 A method for performing AI operations in memory devices according to claim 15, further comprising the step of receiving another instruction including the first address and the third address, wherein the first address identifies the first memory device storing data for the training operation or the inference operation, and the third address identifies the third memory device among the plurality of memory devices as a target for the stored data, and in response to the first address, copying the stored data from the first memory device to the third memory device. Claim 20 A method for performing AI operations in memory devices according to claim 19, further comprising the step of performing other training operations or inference operations on the third memory device using the data copied from the first memory device.
Citation Information
Patent Citations
Data access between computing nodes
US20170344283A1
Memory device performing parallel arithmetic processing and memory module including the same
US20190146788A1
Architectural enhancements for computing systems having artificial intelligence logic disposed locally to memory
US20190251034A1