Memory system and method of operating the same

By adopting a pooled memory structure of multiple sub-memory systems in the memory system, partitioning and depooling table data, the problem of insufficient bandwidth and memory capacity in the deep learning recommendation system service is solved, and efficient and scalable data processing and storage is achieved.

CN114647372BActive Publication Date: 2025-05-13SK HYNIX INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111160008.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-18
Filing Date
2021-09-30
Publication Date
2025-05-13
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

When existing memory systems perform recommendation system services based on deep learning, they can easily lead to bandwidth problems and insufficient memory capacity, making it difficult to efficiently process massive data.

Method used

The pooled memory structure of multiple sub-memory systems is adopted, and by partitioning and de-pooling the embedded tables, the distributed storage and processing of data is realized, the dependence on the host device is reduced, and the memory capacity utilization and data processing efficiency are improved.

Benefits of technology

It realizes that sufficient memory capacity is provided without increasing bandwidth requirements, supports high-performance training and inference operations, and improves the efficiency and scalability of the recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114647372B_ABST
    Figure CN114647372B_ABST
Patent Text Reader

Abstract

The present disclosure provides a memory system and an operation method thereof. The memory system includes: a plurality of memory devices configured to store partial data strips obtained by partitioning an embedding table, the embedding table including vector information strips about items of an acquired learning model; and a memory controller configured to obtain a data strip corresponding to the query from each of the plurality of memory devices among the partial data strips in response to a query received from a host, perform a pooling operation for generating embedded data using the obtained data strips, and provide the embedded data to the host.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the priority of Korean Patent Application No. 10-2020-0178442 filed in the Korean Intellectual Property Office on December 18, 2020, which is incorporated herein by reference in its entirety. Technical Field

[0003] Various embodiments of the present disclosure relate generally to an electronic device, and more particularly, to a memory controller and an operating method thereof. Background Art

[0004] A memory system generally stores data in response to control of a host device such as a computer or a smart phone. A memory system includes a memory device that stores data and a memory controller that controls the memory device. A memory device is generally classified as a volatile memory device or a nonvolatile memory device.

[0005] A volatile memory device stores data only when power is supplied thereto, and loses the stored data when power is not supplied. Examples of volatile memory devices include static random access memory (SRAM) and dynamic random access memory (DRAM).

[0006] Nonvolatile memory devices retain stored data even when power is not supplied. Examples of nonvolatile memory devices include read only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), and flash memory. Summary of the invention

[0007] Various embodiments of the present disclosure relate to a plurality of improved sub-memory systems that perform near data processing (NDP) and a pooled memory system including the sub-memory systems.

[0008] According to an embodiment of the present disclosure, a memory system may include: a plurality of memory devices configured to store partial data strips obtained by partitioning an embedding table, the embedding table including vector information strips about items of an acquired learning model; and a memory controller configured to obtain, in response to a query received from a host, a data strip corresponding to the query from each of the plurality of memory devices among the partial data strips, perform a pooling operation for generating embedded data using the obtained data strips, and provide the embedded data to the host, wherein the query includes a request for the pooling operation, an address at which the host memory receives the embedded data, and a physical address corresponding to any one of the plurality of memory devices.

[0009] According to an embodiment of the present disclosure, a memory system may include: a plurality of memory devices configured to store partial data strips obtained by partitioning an embedding table, the embedding table including vector information strips about items of an acquired learning model; and a memory controller configured to respond to a query and embedded data received from a host, use the embedded data to perform an unpooling operation for generating partitioned data strips, and control the plurality of memory devices to update partial data corresponding to the partitioned data strips stored in each of the plurality of memory devices.

[0010] According to an embodiment of the present disclosure, a pooled memory system may include a host device and multiple sub-memory systems, wherein the multiple sub-memory systems are configured to store partial data strips obtained by partitioning an embedding table, the embedding table including vector information strips about items of an acquired learning model, wherein the host device is configured to: broadcast a first query to the multiple sub-memory systems, and control each of the multiple sub-memory systems to obtain a data strip corresponding to the first query among the partial data strips by using the first query, use the data strips that have been obtained to perform a pooling operation for generating embedded data, and provide the embedded data to the host device.

[0011] According to an embodiment of the present disclosure, a pooled memory system may include multiple sub-memory systems, each of the multiple sub-memory systems including: a memory device configured to store partial data strips obtained by partitioning an embedding table, the embedding table including vector information strips about items of an acquired learning model; and a memory controller configured to perform a pooling operation for generating the embedded data strips and a de-pooling operation for partitioning the training data, the pooled memory system also including a host device configured to generate training data using the embedded data strips received from the multiple sub-memory systems, and to control the multiple sub-memory systems to learn the training data. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a block diagram illustrating a pooled memory system according to an embodiment of the present disclosure;

[0013] Figure 2 is a block diagram illustrating a sub-memory system according to an embodiment of the present disclosure;

[0014] Figure 3 is a diagram showing an embedding table according to an embodiment of the present disclosure;

[0015] Figure 4 is a diagram illustrating a method of storing a portion of a data strip according to an embodiment of the present disclosure;

[0016] Figure 5is a diagram illustrating a method of storing a portion of a data strip according to an embodiment of the present disclosure;

[0017] Figure 6 is a diagram illustrating a method of storing a portion of a data strip according to an embodiment of the present disclosure;

[0018] Figure 7 is a diagram illustrating a lookup operation and a pooling operation according to an embodiment of the present disclosure;

[0019] Figure 8 is a diagram illustrating an inference operation according to an embodiment of the present disclosure;

[0020] Fig. 9 is a diagram illustrating a training operation according to an embodiment of the present disclosure;

[0021] Fig.10 is a diagram illustrating a communication packet of a host device and a sub-memory system according to an embodiment of the present disclosure;

[0022] Fig.11 is a diagram illustrating an inference operation and a training operation according to an embodiment of the present disclosure; and

[0023] Fig.12 is a block diagram showing a configuration of a memory controller according to another embodiment of the present disclosure. DETAILED DESCRIPTION

[0024] The specific structural or functional descriptions illustrating the embodiments according to the concepts disclosed in this specification are only used to describe the embodiments according to the concepts. The embodiments according to the concepts can be implemented in various forms, but the description is not limited to the embodiments described in this specification.

[0025] Various modifications and changes can be applied to the embodiments according to the concepts, so that the embodiments will be shown in the drawings and described in the specification. However, the embodiments according to the concepts of the present disclosure are not to be construed as being limited to a specific disclosure, but include all variations, equivalents or alternative forms that do not depart from the spirit and technical scope of the present disclosure. In some embodiments, well-known processes, device structures and techniques will not be described in detail to avoid blurring the present disclosure. It is intended to more clearly disclose the gist of the present disclosure by omitting unnecessary descriptions.

[0026] Hereinafter, various embodiments of the present disclosure will be described with reference to the accompanying drawings to describe the present disclosure in detail.

[0027] Figure 1 is a block diagram showing a pooled memory system 10000 according to an embodiment of the present disclosure.

[0028] Reference Figure 1, the pooled memory system 10000 may include a plurality of sub-memory systems 1000 and a host device 2000 .

[0029] The pooled storage system 10000 may be a device that provides a recommendation system service. The recommendation system service may refer to a service that provides personalized information to users by filtering information. The pooled storage system 10000 may obtain a user information profile by querying the user's personal information, interests, preferences, etc., and may recommend or provide information and items that are suitable for the user's preference information based on the obtained user information profile.

[0030] Recommendation system services can be implemented as deep learning-based algorithms for high-performance training and inference to effectively use or provide a large and increasing amount of service data. However, since deep learning-based recommendation systems are mainly based on host devices to perform embedded operations, etc., bandwidth issues may occur and insufficient memory capacity may be caused due to the need for massive service data.

[0031] According to an embodiment of the present disclosure, the pooled memory system 10000 can provide a recommendation system service with sufficient memory capacity and no bandwidth problem, that is, the pooled memory system 10000 can ensure sufficient memory capacity through a pooled memory structure capable of increasing memory capacity, and can be implemented to perform near data processing (NDP) to solve bandwidth problems. More specifically, the pooled memory system 10000 can be implemented as a pooled memory structure including a plurality of sub-memory systems 1000, in which a plurality of memory devices are connected in parallel to each other. In addition, the pooled memory system 10000 can control the plurality of sub-memory systems 1000 to perform simple operations related to the embedded operation.

[0032] The pooled memory system 10000 may be implemented in a form in which a plurality of sub-memory systems 1000 and a host device 2000 included therein are interconnected. The plurality of sub-memory systems 1000 may perform reasoning operations and training operations in response to the control of the host device 2000. The reasoning operation may be associated with a read operation of the memory system, and the training operation may be associated with a write operation of the memory system. The reasoning operation and the training operation will be described in more detail later.

[0033] In order to perform reasoning operations and training operations, multiple sub-memory systems 1000 may store data about the learning model that has been acquired. In the present disclosure, multiple sub-memory systems 1000 have acquired and stored therein an embedding table or data about the learning model. Through reasoning operations, multiple sub-memory systems 1000 may provide at least a portion of the stored embedding table to the host device 2000. Through training operations, multiple sub-memory systems 1000 may update at least a portion of the stored embedding table based on training data provided from the host device 2000. The host device 2000 may receive at least a portion of the embedding table from multiple sub-memory systems 1000 through reasoning operations. The host device 2000 may update the received portion to generate training data. The host device 2000 may provide training data to multiple sub-memory systems 1000 to update the embedding tables stored in the multiple sub-memory systems 1000 through training operations.

[0034] More specifically, multiple sub-memory systems 1000 may store vector information about items of a learning model that has been acquired. A learning model may refer to an artificial intelligence (AI) model, and the AI ​​model may be constructed by training. The meaning of constructing by training may refer to constructing a predefined set of operating rules or AI models by using a learning algorithm to train a basic AI model using multiple training data to perform the desired features. The above training may be performed in the pooled memory system 10000 according to the present disclosure or in an external server and / or system. Examples of learning algorithms include supervised learning algorithms, unsupervised learning algorithms, semi-supervised learning algorithms, or reinforcement learning algorithms, and in particular, the learning algorithm may refer to a supervised learning algorithm.

[0035] The learning model may include multiple network nodes that mimic neurons of a human neural network and have weights. Multiple network nodes may form respective connection relationships to simulate the synaptic activity of neurons to exchange signals through synapses. For example, the AI ​​model may include a neural network model or a deep learning model developed from a neural network model. In a deep learning model, multiple network nodes may be located at different depths (or layers) and exchange data according to convolutional connection relationships. For example, models such as the following may be used as data recognition models: deep neural network (DNN) models, recurrent neural network (RNN) models, or bidirectional recurrent deep neural network (BRDNN) models. However, examples of data recognition models are not limited thereto.

[0036] According to an embodiment of the present disclosure, functions related to the recommendation system service of the pooled memory system 10000 may be performed according to the control of the host device 2000. More specifically, the host device 2000 may include a host processor and a host memory. The host processor may be a general-purpose processor such as a central processing unit (CPU), an application processor (AP), or a digital signal processor (DSP), a graphics-specific processor such as a graphics processing unit (GPU) or a visual processing unit (VPU), or an AI-specific processor such as a neural processing unit (NPU). The host memory may store an operating system or application for providing a recommendation system service.

[0037] The host device 2000 may broadcast a query to the plurality of sub-memory systems 1000. By using the query, the host device 2000 may control the plurality of sub-memory systems 1000 to obtain a data strip corresponding to the query from each of the plurality of sub-memory systems 1000 among the partial data strips. The host device 2000 may control the plurality of sub-memory systems 1000 to perform a pooling operation for generating embedded data using the obtained data strips, and provide the embedded data to the host device 2000. The query may include a request for the pooling operation, an address of a host memory receiving the embedded data, and a physical address corresponding to any one of the plurality of memory devices.

[0038] The host device 2000 can generate training data for updating the embedded table based on data received from each of the multiple sub-memory systems. As will be described later, the training data can also be represented as embedded data, which is broadcast from the host device 2000 to each of the sub-memory systems and is used by each of the sub-memory systems to update the partial data stored in the sub-memory system, which is part of the embedded table. In addition, the host device 2000 can broadcast the training data as well as the query. By using the query broadcasted together with the training data, the host device 2000 can control the multiple sub-memory systems 1000 to perform a depooling operation for generating partitioned data strips by partitioning the training data, and update the partial data corresponding to the partitioned data strips in each of the sub-memory systems. The query may include a request for the depooling operation, an address of the host memory corresponding to the embedded data, and a physical address corresponding to any one of the multiple memory devices.

[0039] Figure 2 is a block diagram illustrating a sub-memory system according to an embodiment of the present disclosure.

[0040] Reference Figure 2 , shows a sub-memory system 1000a among the multiple sub-memory systems 1000.

[0041] The sub-memory system 1000a may be implemented as one of various types of memory systems according to a host interface corresponding to a communication method with the host device 2000. For example, the sub-memory system 1000a may be implemented as one of various types of storage devices such as a solid state drive (SSD), a multimedia card (MMC), an eMMC, an RS-MMC, and a micro-MMC, a secure digital (SD) card, a mini SD card, and a micro SD card, a universal serial bus (USB) memory system, a universal flash memory (UFS) device, a personal computer memory card international association (PCMCIA) card type memory system, a peripheral component interconnect (PCI) card type memory system, a high-speed PCI (PCI-E) card type memory device, a compact flash (CF) card, a smart media card, and a memory stick.

[0042] The sub-memory system 1000a may be implemented as one of various package types. For example, the sub-memory system 1000a may be implemented as one of various package types such as: package on package (POP), system-in-package (SIP), system on chip (SOC), multi-chip package (MCP), chip on board (COB), wafer-level manufacturing package (WFP), and wafer-level stacked package (WSP).

[0043] The sub memory system 1000 a may include a memory device 100 and a memory controller 200 .

[0044] The memory device 100 may store data or use the stored data. More specifically, the memory device 100 may operate in response to the control of the memory controller 200. The memory device 100 may include a plurality of banks storing data. Each of the plurality of banks may include a memory cell array having a plurality of memory cells.

[0045] The memory device 100 may be a volatile random access memory such as a dynamic random access memory (DRAM), SDRAM, DDR SDRAM, DDR2 SDRAM, DDR3 SDRAM, LPDDR SDRAM, LPDDR2 SDRAM, and LPDDR3 SDRAM. As an example, in the context of the following description, the memory device 100 is a DRAM.

[0046] The memory device 100 may receive a command and an address from the memory controller 200. The memory device 100 may be configured to access an area selected by a received address in a memory cell array. Accessing a selected area may refer to performing an operation corresponding to a received command on the selected area. For example, the memory device 100 may perform a write operation (or a programming operation), a read operation, and an erase operation. A write operation may be an operation in which the memory device 100 writes data to an area selected by an address. A read operation may refer to an operation in which the memory device 100 reads data from an area selected by an address. An erase operation may refer to an operation in which the memory device 100 erases data from an area selected by an address.

[0047] According to an embodiment of the present disclosure, the memory device 100 may include an embedding table storage device 110 storing an embedding table. The embedding table may be a table including vector information about items of the acquired learning model. More specifically, the memory device 100 may store partial data strips obtained by partitioning the embedding table. The partial data strips may be data obtained by partitioning the embedding table in units of dimensions of vector information, and the memory device 100 may store partial data corresponding to at least one dimension. To avoid repeated description, the following reference is made to Figures 3 to 6 Description Description of the embedded table and some of the data.

[0048] The memory controller 200 may control general operations of the sub memory system 1000 a .

[0049] When power is supplied to the sub-memory system 1000a, the memory controller 200 may run firmware (FW). The firmware (FW) may include: a host interface layer (HIL) that receives a request input from the host device 2000 or outputs a response to the host device 2000; a translation layer (TL) that manages operations between an interface of the host device 2000 and an interface of the memory device 100; and a memory interface layer (MIL) that provides a command to the memory device 100 or receives a response from the memory device 100.

[0050] The memory controller 200 may control the memory device 100 to perform a write operation, a read operation, or an erase operation in response to a request from the host device 2000. During a write operation, the memory controller 200 may provide a write command, a bank address, and data to the memory device 100. During a read operation, the memory controller 200 may provide a read command and a bank address to the memory device 100. During an erase operation, the memory controller 200 may provide an erase command and a bank address to the memory device 100.

[0051] According to an embodiment of the present disclosure, the memory controller 200 may include a read operation control component 210 , a weight update component 220 , and an operation component 230 .

[0052] The read operation control component 210 can identify a physical address corresponding to the query from the query broadcast by the host device 2000, and control the memory device 100 to read the data stripe corresponding to the physical address. In response to the query received from the host device 2000, the read operation control component 210 of the memory controller 200 can obtain the data stripe corresponding to the query from the memory device 100 among the partial data stripes.

[0053] The weight update component 220 may control the memory device 100 to update the weight of the partial data stripe stored in the memory device 100 based on the data received from the host device 2000. More specifically, the weight update component 220 may control the memory device 100 to update the partial data stripe corresponding to the partition data obtained by partitioning the embedded data. The weight update component 220 may be a configuration for controlling the memory device 100 to store the partition data in the memory device 100.

[0054] The operation component 230 may perform a pooling operation for generating embedded data or a de-pooling operation for generating partitioned data stripes.

[0055] More specifically, in response to a query received from the host device 2000, the operation component 230 can obtain a data strip corresponding to the query from the memory device 100. In addition, the operation component 230 can use the data corresponding to the query to perform a pooling operation for generating embedded data. More specifically, the operation component 230 can generate embedded data by a pooling operation for compressing the data strip corresponding to the query. The above-mentioned pooling operation can be used to perform element-by-element operations on the vector information strips of the data strips. For example, the elements of the vector information strips can be added by a pooling operation. The pooling operation can be a process for integrating multiple vectors into a single vector. The pooled memory system 10000 can reduce the size of the data by a pooling operation, and can reduce the amount of data transmission to the host device 2000. In addition, the operation component 230 can perform a de-pooling operation for generating partitioned data strips in response to the query received from the host device 2000 and the embedded data. More specifically, the operation component 230 can perform a de-pooling operation for partitioning the embedded data into multiple partitioned data strips.

[0056] Figure 3 is a diagram illustrating an embedding table according to an embodiment of the present disclosure.

[0057] Reference Figure 3 , shows the embedded tables stored in multiple sub-memory systems 1000.

[0058] The embedding table may include vector information strips about the items of the learning model that have been acquired. The learning model may refer to an artificial intelligence (AI) model, and the AI ​​model may have features built by training. More specifically, the embedding table may be a collection of category data that can be classified into categories and embedding vector pairs corresponding to the category data, and each of the items of the learning model may be category data. Category data in natural language form may be digitized in the form of vectors having similarities to each other using an embedding algorithm. For example, a vector may be a set of numbers consisting of several integers or floating point numbers, such as "(3, 5)" or "(0.1, -0.5, 2, 1.2)". In other words, the embedding table may refer to a set of category data strips of a learning model constructed to classify data into categories and vector information strips of category data. When vector values ​​(such as the slope of an embedding vector or the form of an embedding vector) are more similar to each other, the corresponding words are more similar in semantics.

[0059] According to an embodiment of the present disclosure, the embedding table may have a three-dimensional structure. More specifically, referring to Figure 3 , the embedding table can have a three-dimensional structure including "table size", "number of features" and "dimensionality". "Table size" can refer to the number of items included in a category. For example, when the category labeled <movie name> includes "Harry Potter" and "Shrek", the number of items is two (2). "Number of features" can refer to the number of categories. For example, when the embedding table includes <movie name> and <movie genre>, the number of categories is two (2). "Dimension" can refer to the dimension of the embedding vector. In other words, "dimension" can refer to the number of numbers included in the vector. For example, since a group of "(0.1, -0.5, 2, 1.2)" includes four (4) numbers, the group can be represented as four (4) dimensions.

[0060] Figure 4 is a diagram illustrating a method of storing a partial data strip according to an embodiment of the present disclosure.

[0061] Reference Figure 4 , showing embedded tables stored in multiple sub-memory systems respectively.

[0062] Each of the embedding tables may include a partial data strip. More specifically, each of the partial data strips may be obtained by partitioning the embedding table in units of dimensions of the vector information strip. Each of the plurality of sub-memory systems may store partial data corresponding to at least one dimension. By setting the embedding tables in parallel as much as possible in the pooled memory system 10000, operations related to embedding may be performed without degrading performance in the plurality of sub-memory systems 1000.

[0063] Figure 5 is a diagram illustrating a method of storing a partial data strip according to an embodiment of the present disclosure.

[0064] Reference Figure 5 , showing a portion of data stripes stored in each of the plurality of sub-memory systems 1000. Figure 5 As shown, each of the multiple sub-memory systems 1000 can store partial data strips obtained by partitioning the embedding table in units of "k" dimensions. For example, the first sub-memory system can store partial data strips obtained by partitioning the embedding table in "k" dimensions from the "nth" dimension to the "(n+k-1)" dimension. The second sub-memory system can store partial data strips obtained by partitioning the embedding table in the dimension of vector information of "k" dimensions from the "(n+k)" dimension to the "(n+2k-1)" dimension. Through the above method, the pooled memory system 10000 can store the embedding table in multiple sub-memory systems 1000.

[0065] according to Figure 5 In the embodiment shown, each of the plurality of sub-memory systems 1000 stores partial data strips obtained by the same number of dimensions. However, the embodiment is not limited thereto. According to another embodiment, each of the plurality of sub-memory systems 1000 may be implemented to store partial data strips obtained by partitioning the embedded table by any number of dimensions.

[0066] Figure 6 is a diagram illustrating a method of storing a partial data strip according to an embodiment of the present disclosure.

[0067] Reference Figure 6 , showing a portion of a data strip stored in each of a plurality of memory devices.

[0068] As mentioned above Figure 5 As described, each of the sub-memory systems can store partial data stripes obtained by partitioning the embedding table in units of dimensions. More specifically, each of the sub-memory systems can store partial data stripes obtained by partitioning the embedding table in units of "k" dimensions.

[0069] In addition, the sub-memory system may store the partial data stripes in a plurality of memory devices.The sub-memory system may include a plurality of memory devices, and each of the plurality of memory devices may store partial data corresponding to at least one dimension.

[0070] For example, the i-th sub-memory system may include first to n-th memory devices MD1 to MDn, and each of the first to n-th memory devices MD1 to MDn may store partial data corresponding to at least one dimension.

[0071] Figure 7 is a diagram illustrating a lookup operation and a pooling operation according to an embodiment of the present disclosure.

[0072] Reference Figure 7 , showing a lookup operation for searching a vector information piece in an embedding table and a pooling operation for compressing the vector information piece.

[0073] When receiving a query from the host device 2000 , each of the plurality of sub memory systems 1000 coupled to the host device 2000 may perform a lookup operation for obtaining data corresponding to the query.

[0074] More specifically, the search operation may be performed in parallel in each of the plurality of sub-memory systems 1000. Each of the sub-memory systems 1000 may include a plurality of memory devices 100, and each of the plurality of memory devices 100 may perform a search operation for obtaining data corresponding to a query from the stored partial data. In other words, since the plurality of memory devices 100 perform the search operation in parallel, the plurality of sub-memory systems 1000 may perform the search operation simultaneously.

[0075] In addition, each of the plurality of sub memory systems 1000 may perform a pooling operation for generating embedded data by using a data stripe obtained through a lookup operation.

[0076] More specifically, a pooling operation may be performed in parallel in each of the plurality of sub-memory systems 1000. Each of the sub-memory systems 1000 may include a memory controller 200, and the memory controller 200 may perform a pooling operation for generating embedded data by compressing data strips obtained by a lookup operation. According to an embodiment of the present disclosure, the memory controller 200 may perform element-by-element operations on vector information strips of data strips obtained from the plurality of memory devices 100.

[0077] More specifically, the memory controller 200 may perform element-by-element operations on pieces of vector information to generate a single piece of vector information. The pooled memory system 10000 may perform a pooling operation in each of the plurality of sub-memory systems 1000 and reduce the amount of data transmission from the plurality of sub-memory systems 1000 to the host device 2000 .

[0078] Figure 8 is a diagram illustrating an inference operation according to an embodiment of the present disclosure.

[0079] Reference Figure 8, showing an inference operation. In the inference operation, each of the plurality of sub-memory systems 1000 performs a pooling operation for compressing a vector information piece obtained from an embedding table, and provides an embedding data piece generated by the pooling operation to the host device 2000.

[0080] More specifically, each of the sub-memory systems may store partial data strips obtained by partitioning the embedding table in a plurality of memory devices 100. The embedding table may be a table including vector information strips about items of the acquired learning model. A plurality of memory devices 100 may store partial data strips obtained by partitioning the embedding table in units of dimensions of the vector information strips. For example, each of the partial data strips may be a partial data strip obtained by partitioning the embedding table in units of 16 to 1024 dimensions. The memory controller 200 included in each of the sub-memory systems may obtain data strips corresponding to queries received from the host device 2000 from each of the plurality of memory devices 100 among the partial data strips.

[0081] The memory controller 200 included in each of the sub-memory systems can perform a pooling operation for generating embedded data by using the obtained data strips. For example, the memory controller 200 can obtain 1 to 80 pieces of data from a plurality of memory devices 100, and generate embedded data by performing element-by-element operations on the vector information strips of the obtained data strips. The above-mentioned pooling operation can be used to perform element-by-element operations on the vector information strips of the data strips. For example, the elements of the vector information strips can be added by the pooling operation. The pooling operation can be a process for integrating multiple read vectors into a single vector.

[0082] The memory controller 200 included in each of the sub-memory systems may provide the generated embedded data to the host device 2000. More specifically, the memory controller 200 may generate a batch by accumulating embedded data stripes. In addition, the memory controller 200 may provide the generated batch to the host device 2000. The batch may be a group of embedded data stripes or set data, and one batch may include 128 to 1024 embedded data stripes.

[0083] Fig. 9 is a diagram illustrating a training operation according to an embodiment of the present disclosure.

[0084] Reference Fig. 9 , showing a training operation. In the training operation, each of the plurality of sub-memory systems 1000 performs a depooling operation for partitioning a gradient received from the host device 2000, and updates an embedding table using the partitioned gradient. The gradient is data for updating the embedding table, and may refer to embedding data including a weight.

[0085] Each of the sub-memory systems may store partial data strips obtained by partitioning the embedding table in a plurality of memory devices 100. The embedding table may include vector information strips about items of the acquired learning model. The plurality of memory devices 100 may store partial data strips obtained by partitioning the embedding table in units of the dimensions of the vector information strips. For example, each of the partial data strips may be a partial data strip obtained by partitioning the embedding table in units of 16 to 1024 dimensions.

[0086] The memory controller 200 included in each of the sub-memory systems may receive a broadcast query requesting an update of the embedded table and a broadcast gradient (i.e., embedded data including weights) from the host device 2000. The broadcast gradient may be received from the host device 2000 in batches by each of the sub-memory systems. In addition, the memory controller 200 may use the gradient received from the host device 2000 to perform a depooling operation for generating partitioned data strips. For example, the memory controller 200 may receive a gradient of a broadcast batch (i.e., embedded data including weights) from the host device 2000. In addition, the memory controller 200 may use the batch gradient broadcasted from the host device 2000 to generate partitioned data strips. The memory controller 200 may perform a depooling operation for partitioning the embedded data into partitioned data strips to easily update the embedded table. The above-mentioned depooling operation may be used to perform element-by-element operations on the vector information strips of the data strips.

[0087] In addition, the memory controller 200 included in each of the sub memory systems may control the plurality of memory devices 100 to update partial data corresponding to the partition stripe in each of the plurality of memory devices 100 .

[0088] The pooled memory system 10000 can broadcast a query requesting an update and a gradient (i.e., embedded data including weights) to multiple sub-memory systems 1000, and can control each of the sub-memory systems to update a portion of the data. The pooled memory system 10000 can use the query and embedded data to train the learning model stored in each of the sub-memory systems.

[0089] Fig.10 is a diagram illustrating a communication packet of the host device 2000 and the sub memory system 1000 according to an embodiment of the present disclosure.

[0090] Reference Fig.10, shows a communication packet between the host device 2000 and the plurality of sub-memory systems 1000. More specifically, the first communication packet 11 may be a message transmitted from the host device 2000 to the plurality of sub-memory systems 1000. The first communication packet 11 may include a "task ID", an "operation code", a "source address", a "source size", and a "destination address".

[0091] The first communication packet 11 may consist of a total of 91 bits, and 4 bits may be allocated to a “task ID”. The “task ID” may indicate an operation state of the host device 2000. For example, the “task ID” may indicate whether the operation of the host device 2000 is running or terminated. The host device 2000 may readjust the operation of the sub memory system using the “task ID”.

[0092] 3 bits may be allocated to the "Opcode", and the "Opcode" may include data for distinguishing a plurality of embedding operations from each other. More specifically, the host device 2000 may use the "Opcode" to distinguish between initialization, inference operation, and training operation of the embedding table.

[0093] 32 bits may be allocated to the "source address", and the "source address" may include data about the source address of the query or gradient. More specifically, the host device 2000 may include data about each of the gradient addresses acquired from the host memory in the sub-memory system into the "source address". The gradient is data for updating the embedding table and may refer to embedding data including a weight.

[0094] 20 bits may be allocated to "source size", and "source size" may include data on the size of the query or gradient that each of the sub-memory systems obtains from the host memory. "Target address" may include an address in the host memory where the internal operation results of each of the sub-memory systems will be stored.

[0095] The host device 2000 may communicate with the plurality of sub memory systems 1000 using the first communication packet 11. In addition, when the plurality of sub memory systems 1000 receive the first communication packet 11 from the host device 2000, the plurality of sub memory systems 1000 may transmit a second communication packet 12 as a response message. The second communication packet 12 may include a "task ID" and an "operation code".

[0096] Fig.11 is a diagram illustrating an inference operation and a training operation according to an embodiment of the present disclosure.

[0097] Reference Fig.11, shows a pooled memory system 10000 that includes a plurality of sub-memory systems 1000 and a host device 2000 and performs an inference operation and a training operation.

[0098] The sub memory system 1000 may store partial data strips obtained by partitioning an embedding table including vector information strips about items of a learning model that have been acquired in the memory device 100 .

[0099] The host device 2000 may control the sub-memory system 1000 to perform a pooling operation for generating an embedded data strip. The above-mentioned pooling operation may be used to perform element-by-element operations on the vector information strips of the data strip to generate a single vector information strip. The sub-memory system 1000 may provide the generated embedded data strip to the host device 2000.

[0100] The host device 2000 may generate training data using the embedded data strips received from the sub-memory system 1000. More specifically, the host device 2000 may compare the features of the embedded data strips with each other. The embedded data strips may include embedded vectors, and the host device 2000 may compare vector values ​​of the embedded vectors, such as slope, form, or size, with each other. When the vector values ​​(such as the slope of the embedded vectors or the form of the embedded vectors) are more similar to each other, the corresponding words are more similar to each other in terms of semantics. The host device 2000 may calculate the scores of the embedded data strips received from the sub-memory system 1000, and generate training data based on the calculated scores. The training data may be used to update the embedding table, and may refer to embedded data with corrected weights.

[0101] The host device 2000 may control the sub memory system 1000 to learn the training data. The sub memory system 1000 may perform a depooling operation for partitioning the training data received from the host device 2000 and update the weight of the embedding table.

[0102] Fig.12 is a block diagram showing a configuration of a memory controller 1300 according to another embodiment of the present disclosure.

[0103] Reference Fig.12 , the memory controller 1300 may include a processor 1310 , a RAM 1320 , an error correction code (ECC) circuit 1330 , a ROM 1360 , a host interface 1370 , and a memory interface 1380 . Fig.12 The memory controller 1300 shown may be Figure 1 and Figure 2 An embodiment of a memory controller 200 is shown.

[0104] The processor 1310 may communicate with the host device 2000 using the host interface 1370, and perform logic operations to control the operation of the memory controller 1300. For example, the processor 1310 may load a write command, a data file or a data structure, perform various operations, or generate commands and addresses in response to a request received from the host device 2000 or an external device. For example, the processor 1310 may generate various commands for a write operation, a read operation, an erase operation, a suspend operation, and a parameter setting operation.

[0105] The processor 1310 may perform a function of a translation layer (TL). The processor 1310 may recognize a physical address provided by the host device 2000.

[0106] The processor 1310 may generate a command without a request from the host device 2000. For example, the processor 1310 may generate a command for a background operation such as an operation for a refresh operation of the memory device 100.

[0107] The RAM 1320 may be used as a buffer memory, an operation memory, or a cache memory of the processor 1310. The RAM 1320 may store codes and commands executed by the processor 1310. The RAM 1320 may store data processed by the processor 1310. When the RAM 1320 is implemented, the RAM 1320 may include a static RAM (SRAM) or a dynamic RAM (DRAM).

[0108] The ECC circuit 1330 may detect and correct errors during a write operation or a read operation. More specifically, the ECC circuit 1330 may perform an error correction operation according to an error correction code (ECC). The ECC circuit 1330 may perform ECC encoding based on data to be written to the memory device 100. The data on which the ECC encoding has been performed may be transmitted to the memory device 100 through the memory interface 1380. In addition, the ECC circuit 1330 may perform ECC decoding on data received from the memory device 100 through the memory interface 1380.

[0109] The ROM 1360 may be used as a storage unit to store various types of information used for the operation of the memory controller 1300. The ROM 1360 may be controlled by the processor 1310.

[0110] The host interface 1370 may include a protocol for exchanging data between the host device 2000 and the memory controller 1300. More specifically, the host interface 1370 may communicate with the host device 2000 through one or more different interface protocols such as a universal serial bus (USB) protocol, a multimedia card (MMC) protocol, a peripheral component interconnect (PCI) protocol, a high-speed PCI (PCI-E) protocol, an advanced technology attachment (ATA) protocol, a serial ATA protocol, a parallel ATA protocol, a small computer system interface (SCSI) protocol, an enhanced small disk interface (ESDI) protocol, an integrated drive electronics (IDE) protocol, and a dedicated protocol.

[0111] The memory interface 1380 may communicate with the memory device 100 using a communication protocol according to the control of the processor 1310. More specifically, the memory interface 1380 may communicate commands, addresses, and data with the memory device 100 through a channel.

[0112] According to embodiments of the present disclosure, a plurality of improved sub-memory systems that perform near data processing (NDP) and a pooled memory system including the sub-memory systems may be provided.

Claims

1. A memory system, comprising: a plurality of memory devices storing partial data strips obtained by partitioning an embedding table including vector information strips about terms of the acquired learning model; as well as a memory controller, in response to a query received from a host, obtaining a data stripe corresponding to the query from among the partial data stripes from each of the plurality of memory devices, performing a pooling operation for generating embedded data using the obtained data stripes, and providing the embedded data to the host, wherein the query includes a request for the pooling operation, an address of a host memory receiving the embedded data, and a physical address corresponding to any one of the plurality of memory devices, wherein the embedding table categorizes the items according to categories and includes pieces of vector information digitized based on similarities between the categorized items; and Each of the partial data pieces is data obtained by partitioning the embedding table in units of dimensions of the vector information piece.

2. The memory system of claim 1 , wherein the memory controller comprises: a read operation control component that identifies a physical address corresponding to the query and controls each of the plurality of memory devices to read a data stripe corresponding to the physical address; as well as An operation component performs the pooling operation for generating the embedded data by compressing the obtained data strips. 3 . The memory system according to claim 2 , wherein the operation component performs an element-by-element sum operation on vector information of the already obtained data stripes to generate the embedded data. 4 . The memory system according to claim 1 , wherein the embedded data is an embedded vector including vector information of the data strip that has been obtained.

5. The memory system according to claim 1, Each of the plurality of memory devices stores partial data corresponding to at least one dimension partitioned from the embedded table.

6. A memory system comprising: a plurality of memory devices storing partial data strips obtained by partitioning an embedding table including vector information strips about terms of the acquired learning model; as well as a memory controller that, in response to a query and embedded data received from a host, performs a depooling operation for generating a partition data stripe using the embedded data, and controls the plurality of memory devices to update a portion of data corresponding to the partition data stripe stored in each of the plurality of memory devices, wherein the embedding table categorizes the items according to categories and includes pieces of vector information digitized based on similarities between the categorized items; and Each of the partial data pieces is data obtained by partitioning the embedding table in units of dimensions of the vector information piece.

7. The memory system of claim 6, wherein the memory controller comprises: An operation component for performing the depooling operation for partitioning the embedded data into partitioned data stripes; as well as A weight updating component controls the plurality of memory devices to update a weight of a portion of data corresponding to a partition data stripe stored in each of the plurality of memory devices. 8 . The memory system of claim 7 , wherein the operation component performs an element-by-element sum operation on the vector information of the embedded data.

9. The memory system of claim 6, wherein the query comprises a request for the depooling operation, an address of a host memory corresponding to the embedded data, and a physical address corresponding to any one of the plurality of memory devices.

10. The memory system according to claim 6, Each of the plurality of memory devices stores partial data corresponding to at least one dimension partitioned from the embedded table.

11. A pooled memory system, comprising: Host device; as well as a plurality of sub-memory systems storing partial data strips obtained by partitioning an embedding table including vector information strips about items of the acquired learning model, wherein the host device: broadcasting a first query to the plurality of sub-memory systems; and controlling each of the plurality of sub-memory systems to obtain a data strip corresponding to the first query among the partial data stripes by using the first query, performing a pooling operation for generating embedded data using the obtained data stripes, and providing the embedded data to the host device, wherein the embedding table categorizes the items according to categories and includes pieces of vector information digitized based on similarities between the categorized items; and Each of the partial data pieces is data obtained by partitioning the embedding table in units of dimensions of the vector information piece.

12. The pooled memory system according to claim 11, wherein each of the plurality of sub-memory systems comprises: a plurality of memory devices storing the partial data pieces obtained by partitioning the embedding table including vector information pieces about items of the acquired learning model; a read operation control component that identifies a physical address corresponding to the first query and controls each of the plurality of memory devices to read a data stripe corresponding to the physical address; as well as An operation component performs the pooling operation for generating the embedded data by compressing the obtained data strips. 13 . The pooled memory system according to claim 12 , wherein the operation component performs element-by-element operations on the vector information of the acquired data stripes.

14. The pooled memory system of claim 12, wherein the host device controls the plurality of sub-memory systems by using the first query, the first query comprising a request for the pooling operation, an address at which the host memory device receives the embedded data, and physical addresses of the plurality of memory devices.

15. The pooled memory system according to claim 11, Each of the plurality of sub-memory systems stores partial data corresponding to at least one dimension partitioned from the embedded table. 16 . The pooled memory system of claim 11 , wherein the host device generates training data for updating the embedding table based on embedded data received from each of the plurality of sub-memory systems.

17. The pooled memory system according to claim 16, wherein the host device broadcasts a second query and the training data to the plurality of sub-memory systems, and Each of the multiple sub-memory systems responds to the second query and the training data received from the host device, uses the training data to perform a depooling operation for generating partitioned data strips, and updates partial data corresponding to the partitioned data strips stored in each of the multiple sub-memory systems.

18. The pooled memory system according to claim 17, wherein each of the plurality of sub-memory systems comprises: An operation component, performing the depooling operation for partitioning the training data into partition data strips; as well as The weight updating component updates the weights of some data stripes corresponding to the partitioned data stripes.

19. The pooled memory system of claim 17, wherein the host device controls the plurality of sub-memory systems by using the second query, the second query comprising a request for the depooling operation, an address of the host memory device corresponding to the training data, and a physical address corresponding to any one of the plurality of sub-memory systems.

20. A pooled memory system, comprising: A plurality of sub-memory systems, each of the plurality of sub-memory systems comprising: a memory device storing partial data pieces obtained by partitioning an embedding table including vector information pieces about items of the acquired learning model; and a memory controller that performs a pooling operation for generating embedded data stripes and a de-pooling operation for partitioning training data; and a host device that generates the training data using the embedded data strips received from the plurality of sub-memory systems and controls the plurality of sub-memory systems to learn the training data, wherein the embedding table categorizes the items according to categories and includes pieces of vector information digitized based on similarities between the categorized items; and Each of the partial data pieces is data obtained by partitioning the embedding table in units of dimensions of the vector information piece.

Citation Information

Patent Citations

  • Method for identifying virtual machine in local area network

    CN112068926A

  • Object localization within a semantic domain

    US20190050648A1