Memory controller and operating method thereof
By identifying and sequentially inputting vectors to the computing unit through the memory controller, the problem of low computing performance of the memory controller is solved, achieving more efficient computing and power saving.
Patent Information
- Application Number
- CN202411903289.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-16
- Filing Date
- 2024-12-23
- Publication Date
- 2025-10-24
AI Technical Summary
Existing memory controllers access memory too frequently during computation, resulting in low computational performance and high power consumption.
The memory controller identifies the information of multiple vectors required for calculations on multiple rows of the matrix and controls multiple computing units to sequentially input these vectors to perform calculations, reducing the number of information reading operations.
This reduces the time and power consumption required for the memory controller to perform calculations, thereby improving computing performance.
Smart Images

Figure CN120832464A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments relate to a memory controller and a method of operating the same. Background Art
[0002] Memory, such as dynamic random access memory (DRAM), can store data used for computation. Because computations using large amounts of data are necessary, the time required for a memory controller to read the data stored in the memory and perform computations using the read data can be lengthy. Therefore, it is desirable to develop a method for operating a memory controller that improves the computational performance of the memory controller when performing computations using the data stored in the memory. Summary of the Invention
[0003] One aspect provides a memory controller and an operating method thereof that improves computing performance by minimizing the number of times the memory controller is accessed.
[0004] The technical problems to be solved by the present disclosure are not limited to the technical problems described above, and other technical problems can be inferred from the following exemplary embodiments.
[0005] A memory controller is provided herein, comprising: a plurality of computing units, wherein the plurality of computing units are configured to perform computations related to matrix multiplication; and a controller configured to: identify first information indicating a plurality of vectors required to perform the computations on a plurality of rows of a matrix, identify second information indicating at least one row to which each of the plurality of vectors among the plurality of rows corresponds, and control the plurality of computing units to perform the computations on the plurality of rows by sequentially inputting the plurality of vectors into the plurality of computing units based on the first information and the second information.
[0006] This article also provides an operating method of a memory controller, which includes multiple computing units and a controller, and the multiple computing units are configured to perform calculations related to matrix multiplication. The operating method includes: identifying first information indicating multiple vectors required to perform the calculations on multiple rows of a matrix, identifying second information indicating at least one row corresponding to each of the multiple vectors among the multiple rows, and controlling the multiple computing units to perform calculations on the multiple rows by sequentially inputting the multiple vectors into the multiple computing units based on the first information and the second information.
[0007] Also provided herein is a non-transitory computer-readable recording medium having a program for executing the operation method on a computer.
[0008] Also provided herein is a memory system including a host system including a memory device controller, and a plurality of memory controllers operating according to commands received from the memory device controller, wherein the plurality of memory controllers includes a first memory controller including a plurality of computing units configured to perform a computation related to matrix multiplication, and a controller configured to identify first information indicating a plurality of vectors required for the computation with respect to a plurality of rows of a matrix, identify second information indicating at least one row to which each of the plurality of vectors among the plurality of rows corresponds, and control the plurality of computing units to perform the computation with respect to the plurality of rows by sequentially inputting the plurality of vectors into the plurality of computing units based on the first information and the second information.
[0009] Additional aspects will be set forth in part in the description which follows, and in part will become apparent to those skilled in the art upon examination of the description, or can be learned by practice of the disclosure.
[0010] According to example embodiments, by the memory controller sequentially inputting the plurality of vectors into the plurality of computing units based on the first information and the second information, and by the memory controller controlling the plurality of computing units to perform the computation with respect to the plurality of rows, the number of operations to read information about the plurality of vectors can be minimized. Accordingly, the time taken for the memory controller to perform the computation can be reduced, and the amount of power consumed by the memory controller to perform the computation with respect to the plurality of rows of the matrix can also be reduced.
[0011] Effects of the disclosure are not limited to what has been described above, and other effects will be apparent to those skilled in the art from the following description.
[0012] No description in the present application should be interpreted as implying any particular element, step, or function is an essential element that is necessarily included in the claim scope. The scope of the patent subject matter is defined only by the claims. Also, unless the word "means" is followed by the word "for" in a claim, no claim element is intended to be means-plus-function or step-plus-function. Any claim element that does not include the word "means" or "step" followed by the word "for" is not intended to be a means or step-plus-function element. The use of the terms "including," "containing," "comprising," "having," or "including" in a claim element does not in and of itself require that any particular element be present. Any use of the terms "either," "or," or "and" in a claim element does not in and of itself require that all options or alternatives be present and exclusive. BRIEF DESCRIPTION OF DRAWINGS
[0013] These and / or other aspects, features, and advantages will become apparent and more readily appreciated from the following description, taken in conjunction with the accompanying drawings, in which: Figure 1 FIG. 1 is a diagram illustrating a memory controller according to an example embodiment; Figure 2 FIG. 2 is a diagram illustrating a memory controller according to an example embodiment in more detail; Figure 3A FIG. 3 is a diagram illustrating a memory system including a memory device and a host system according to an example embodiment; Figure 3B FIG. 4 is a diagram illustrating a memory device according to an example embodiment; Figure 4 FIG. 5 is a flowchart for explaining an operation method of a first controller according to an example embodiment; Figure 5 FIG. 6 is a diagram illustrating a method of performing a calculation related to matrix multiplication of a plurality of rows of a sparse matrix and a target matrix according to an example embodiment; Figure 6 FIG. 7 is a diagram for explaining a number of times a second controller transmits a command including address information to a memory when performing a calculation related to matrix multiplication according to an example embodiment; Figure 7 FIG. 8 is a diagram illustrating a method of selecting a preset number of a plurality of rows among rows of a sparse matrix; Figure 8 FIG. 9 is a flowchart of an operation method of a memory controller according to an example embodiment. DETAILED DESCRIPTION
[0014] The terms used in the example embodiments are selected from general terms that are widely used at present, taking into account the functions in the disclosure, if possible. However, the terms can vary according to the intention of those skilled in the art or precedents, appearance of new technologies, etc. Also, in some cases, there are terms arbitrarily selected by the applicant, and in these cases, the meanings will be described in the corresponding description. Therefore, the terms used in the disclosure should be defined based on the meanings of the terms and the content of the disclosure, not based on simple names of the terms.
[0015] Throughout the specification, when a component is described as "including or including" a component, unless otherwise specified, it does not exclude other components but can further include other components. Also, the terms "… unit", "… group", and "… module" and the like described in the specification are terms that refer to a unit that processes at least one function or operation, which can be implemented as hardware, software, or a combination thereof.
[0016] Hereinafter, example embodiments of the disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art to which the disclosure pertains can easily implement them. However, the disclosure can be implemented in various different forms, and is not limited to the example embodiments described herein.
[0017] Hereinafter, example embodiments will be described in detail with reference to the accompanying drawings.
[0018] Figure 1 FIG. 1 is a diagram illustrating a memory controller according to an example embodiment.
[0019] Referring to Figure 1 According to an example embodiment, a memory controller 100 can include a plurality of computing units 110 and a controller 120. Here, the memory controller 100 can be an apparatus for controlling performance of various computations using information read from a memory. The memory controller 100 can be located close to the memory, and can be an apparatus for controlling performance of computations using information read from the memory. In view of this, the memory controller 100 can be referred to as a near data processing (NDP) module or a near memory processing (NMP) module. Here, the memory can include a DRAM, but is not limited thereto. More specifically, the memory controller 100 can be an apparatus for controlling various operations related to matrix multiplication, including an operation of reading information about vectors from the memory and an operation of performing computations related to matrix multiplication using the read information about vectors.
[0020] According to an example embodiment, computations related to matrix multiplication performed in the memory controller 100 can be computations related to sparse matrix multiplication. Here, a sparse matrix can denote a matrix in which most of the values of the elements of the matrix are zero. For example, only 10% or less of the elements can be non-zero. In view of this, among the elements of the sparse matrix, an element whose value is 0 can be referred to as a zero element. In contrast, among the elements of the sparse matrix, an element whose value is not 0 can be referred to as a non-zero element. A target matrix that is a target of matrix multiplication of the sparse matrix can be a dense matrix. Here, a dense matrix can denote a matrix in which most of the values of the elements of the matrix are non-zero. In other words, the target matrix can have a much larger capacity than the sparse matrix, and thus it is efficient in terms of storage space management to store the target matrix in the memory rather than in a buffer (not shown) in the memory controller 100. In view of this, at least some of the rows of the target matrix can be identified as a plurality of vectors required for computations related to matrix multiplication, and the memory controller 100 can sequentially perform an operation to read the plurality of vectors identified as the vectors required for the computations related to matrix multiplication.
[0021] The plurality of computing units 110 can be computing units for performing a computation based on various information including information read from the memory. More specifically, each of the plurality of computing units 110 can be a computing unit for performing a computation related to matrix multiplication based on information about any one vector among the plurality of vectors read from the memory. In an example embodiment, each of the plurality of computing units 110 can be a multiply and accumulation (MAC) unit, but is not limited thereto. In another example embodiment, each of the plurality of computing units 110 can include an arithmetic and logical unit (ALU) and a computing unit for a fused multiply-add (FMA).
[0022] According to an example embodiment, the controller 120 can control the overall operation of the memory controller 100. More specifically, the controller 120 can control the overall operation of a computation related to matrix multiplication of the memory controller 100.
[0023] According to an example embodiment, the controller 120 can identify information required for a computation related to matrix multiplication. More specifically, the controller 120 can identify first information about a plurality of vectors required for a computation about a plurality of rows of a matrix, and second information about at least one row corresponding to each of the plurality of vectors. In an example embodiment, the plurality of vectors can be identified based on columns of non-zero elements included in the plurality of rows. More specifically, among rows of a target matrix, a row in which multiplication with non-zero elements included in a plurality of rows of a sparse matrix is performed can be identified as a vector required for a computation related to matrix multiplication. In another example embodiment, among the plurality of vectors, at least one row corresponding to a first vector can be identified based on a row of a non-zero element included in a first column corresponding to the first vector.
[0024] According to an example embodiment, by inputting information related to a computation into the plurality of computing units 110, the controller 120 can control the plurality of computing units 110 to perform a computation on the plurality of rows. For example, by sequentially inputting the plurality of vectors into the plurality of computing units 110 based on the first information and the second information, the controller 120 can control the plurality of computing units 110 to perform a computation on the plurality of rows. The plurality of computing units 110 can perform a computation on the plurality of rows based on input information including information about the plurality of vectors.
[0025] Figure 2 is a diagram illustrating a memory controller according to an example embodiment in more detail.
[0026] More specifically, Figure 2is a diagram illustrating an entire calculation process related to matrix multiplication, including operations in which the controller 120 controls the plurality of calculation units 110 to perform calculations, and the controller 120 includes a first controller 140 and a second controller 150.
[0027] Referring to Figure 2 , the memory controller 100 can include the plurality of calculation units 110, the controller 120, the buffer 130, and the interface 160. The controller 120 can include the first controller 140 and the second controller 150, and the first controller 140 can include a matrix buffer 141, an extractor 142, a first queue 143, and a second queue 144. Here, the first controller 140 can be located close to the memory, and can perform operations to control the plurality of calculation units 110 to perform calculations. In this regard, the first controller 140 can be referred to as an NDP controller. In an example embodiment, when the memory is a DRAM, the second controller 150 can perform operations to control the DRAM. Here, the second controller 150 can be referred to as a DRAM controller. The extractor can be implemented by a CPU, a processor that executes instructions from the memory, custom hardware, or a combination.
[0028] The sparse matrix can be stored in the matrix buffer 141. More specifically, most elements of the sparse matrix have zero values, and thus a compressed sparse matrix can be stored in the matrix buffer 141. For example, the compressed sparse matrix can be a sparse matrix compressed in a compressed sparse row (CSR) format based on row compression or in a compressed sparse column (CSC) format based on column compression, but the compressed sparse matrix is not limited thereto. Hereinafter, for convenience of explanation, descriptions will be made based on an uncompressed sparse matrix. Further, only a plurality of rows of the sparse matrix can be stored in the matrix buffer 141. The plurality of rows stored in the matrix buffer 141 can be determined by a microcontroller (not shown) of a host system, which will be described in detail with reference to FIG. 3.
[0029] According to an example embodiment, based on the sparse matrix stored in the matrix buffer 141, the extractor 142 can identify first information about a plurality of vectors required for calculations with respect to the plurality of rows and second information about at least one row corresponding to each of the plurality of vectors among the plurality of rows. More specifically, by sequentially searching for non-zero elements included in the plurality of rows for each column according to a predetermined first order, the extractor 142 can perform operations to identify the plurality of vectors and at least one row corresponding to each of the plurality of vectors. Here, sequentially searching for each column according to the predetermined first order can be sequentially searching from a first column to a last column, but the sequential search is not limited thereto.
[0030] The information recognized by the extractor 142 can be split and stored separately into the first queue 143 and the second queue 144. According to an example embodiment, the first information can be stored in the first queue 143. The first information can include address information of the plurality of vectors stored in the memory. In other words, the first information can include memory address information for accessing each of the plurality of vectors. For example, the vectors of the target matrix stored in the memory can be distinguished only by information about the row of the target matrix to which each of the plurality of vectors corresponds, and thus the address information stored in the first queue 143 can be index information indicating the row of the target matrix corresponding to the vector. In an example embodiment, the index of the vector corresponding to the i-th row of the target matrix can be i-1, but is not limited thereto. According to another example embodiment, the second information can be stored in the second queue 144. The second information can include index information of the row to which each of the plurality of vectors corresponds among the plurality of rows. Among the plurality of rows, when the first vector corresponds to the first row and the second row, the second information related to the first row can include first index information of the first row and second index information of the second row. The index indicating the i-th row of the sparse matrix can be i-1, but is not limited thereto.
[0031] According to an example embodiment, the first queue 143 and the second queue 144 can have a first-in first-out (FIFO) data structure. In view of this, the first information related to the vector first recognized by the extractor 142 can be first stored in the first queue 143. In other words, the first information related to the plurality of vectors can be sequentially stored in the first queue 143 according to a second order with respect to the plurality of vectors. Also, the second information related to at least one row to which each of the plurality of vectors corresponds can be sequentially stored in the second queue 144 according to the second order with respect to the plurality of vectors. Here, the second order of the plurality of vectors can be the same as the order in which the plurality of vectors are recognized by the extractor 142, the order in which the plurality of vectors are stored in the first queue 143, the order in which the plurality of vectors are transmitted to the second controller 150, and the order in which the second controller 150 reads the plurality of vectors from the memory. For example, the second order with respect to the plurality of vectors can be ascending order of the plurality of rows of the target matrix corresponding to the plurality of vectors, but the second order is not limited thereto. Also, after the calculation of the plurality of rows is completed, the information stored in the first queue 143 and the second queue 144 can be deleted.
[0032] According to an example embodiment, the first controller 140 can sequentially transmit the first information and the second information to the second controller 150.
[0033] More specifically, the first queue 143 and the second queue 144 have a FIFO data structure, and thus the first controller 140 can sequentially transmit the first information and the second information stored in the first queue 143 and the second queue 144 to the second controller 150. In the second order regarding the plurality of vectors, among the plurality of vectors, any one vector can be included only once without overlapping, and thus the number of times that the first information and the second information are transmitted to the second controller 150 can be equal to the number of the plurality of vectors. In other words, the operation of the second controller 150 reading one vector of the plurality of vectors from the memory is performed only once, and thus each of the plurality of vectors can be referred to as a unique vector.
[0034] According to an example embodiment, the second controller 150 can be a controller for controlling the memory. More specifically, the second controller 150 can be a controller for performing an operation of accessing data stored in the memory. In this regard, the second controller 150 can perform an overall operation to control the memory, including reading data stored in the memory and writing data to the memory. For example, the second controller 150 can sequentially transmit a command including address information to the memory and sequentially receive information regarding the plurality of vectors from the memory. More specifically, based on the second order of the plurality of vectors, in response to receiving the first information and the second information regarding any one vector of the plurality of vectors from the first controller 140, the second controller 150 can transmit a command including address information regarding the corresponding vector to the memory. Here, the command can include reading information regarding any one vector stored in the memory based on the address information and returning the information regarding the vector to the second controller 150. In this regard, the address information included in the command can include information about the location of a first element of the vector corresponding to the first information in the memory and information regarding the vector size. Thus, the second controller 150 can receive information about the vector corresponding to the command.
[0035] According to an example embodiment, the second controller 150 can control the memory through the interface 160. Here, the interface 160 can be a device that allows interaction between the memory controller 100 and an external device. More specifically, the interface 160 can be an interface that allows interaction between the memory controller 100 and an external device such as a memory or a memory device controller. For example, the interface 160 can be a double data rate (DDR) physical layer, but is not limited thereto. The interface 160 can be connected to a command pin 161 for transmitting and receiving a command to and from the external device, and a data pin 162 for transmitting and receiving data to and from the external device. For example, the command pin 161 can correspond to DDR. Command / Address (C / A), and the data pin 162 can correspond to DDR. Data Queue (DQ) as a DDR data line, but the command pin 161 and the data pin 162 are not limited thereto. In view of this, a command transmitted from the second controller 150 can be transmitted to the memory through the command pin 161, and information related to a vector transmitted from the memory can be transmitted to the second controller 150 through the data pin 162. Hereinafter, a process of performing a calculation on a first vector read from the memory among a plurality of vectors is described.
[0036] According to an example embodiment, the plurality of calculation units 110 can include two or more calculation units to perform a calculation based on various information including information read from the memory. In view of this, among a plurality of vectors, some vectors can correspond to two or more rows, and thus in order to process a calculation on the corresponding vectors in parallel, the plurality of calculation units 110 can include two or more calculation units. Referring to Figure 2 , the plurality of calculation units 110 can include a first calculation unit 111, a second calculation unit 112, and a third calculation unit 113, which are three calculation units, but the disclosure is not limited thereto.
[0037] According to an example embodiment, the second information can include information about a first row and a second row to which the first vector corresponds. A calculation for the first row and a calculation for the second row can be performed in parallel.
[0038] In the first computing unit 111, a computation for the first row can be performed. In this regard, the controller 120 can control an operation of inputting information about the first vector, which is read out by the second controller 150, into the first computing unit 111, and an operation of inputting information about the first non-zero element stored in the matrix buffer 141 and information about the first row stored in the second queue 144 into the first computing unit 111. More specifically, by inputting the information about the first vector, the information about the first non-zero element, and the information about the first row into the first computing unit 111 at the same time, the controller 120 can control the first computing unit 111 to perform a computation for the first row. Here, the first non-zero element can be an element in a column corresponding to the first row of the sparse matrix and the first vector. The first computing unit 111 can perform scalar multiplication based on the first vector and the first non-zero element.
[0039] When the computation for the first row is performed in the first computing unit 111, a computation for a second row can be performed in the second computing unit 112 in parallel. In this regard, the controller 120 can control an operation of inputting information about the first vector, which is read out by the second controller 150, into the second computing unit 112, and an operation of inputting information about a second non-zero element stored in the matrix buffer 141 and information about the second row stored in the second queue 144 into the second computing unit 112. More specifically, by inputting the information about the first vector, the information about the second non-zero element, and the information about the second row into the second computing unit 112 at the same time, the controller 120 can control the second computing unit 112 to perform a computation for the second row. Here, the second non-zero element can be an element in a column corresponding to the second row of the sparse matrix and the first vector. The second computing unit 112 can perform scalar multiplication based on the first vector and the second non-zero element.
[0040] Although Figure 2 Although not shown in FIG. 1, the number of rows corresponding to a certain vector can be greater than the number of computing units. For example, if the number of rows corresponding to a certain vector is 4, computations for three rows can be first performed in the first computing unit 111, the second computing unit 112, and the third computing unit 113. Thereafter, a computation for the remaining one row can be performed in any one of the first computing unit 111, the second computing unit 112, and the third computing unit 113.
[0041] Further, even Figure 2The controller 120 can also control an operation of inputting information about a vector to be read out by the second controller 150 into the plurality of computing units 110 and an operation of inputting information about a row stored in the second queue 144 into the plurality of computing units 110, according to an example embodiment, although not shown. In view of this, when the values of non-zero elements of a sparse matrix are all 1, a result of performing scalar multiplication based on a vector and a non-zero element can be the same as the non-zero element, and thus, in order to improve computational efficiency related to matrix multiplication, the controller 120 can not perform an operation of inputting any one of non-zero elements included in a sparse matrix.
[0042] According to an example embodiment, the buffer 130 can store a result vector corresponding to each row of a sparse matrix. In view of this, when a computation for a particular row is performed in any one of the plurality of computing units 110, a result vector corresponding to the particular row can be updated based on a computation in the computing unit corresponding to the particular row. For example, a result vector corresponding to a first row can be updated according to scalar multiplication based on a first vector and a first non-zero element. More specifically, a vector according to scalar multiplication based on the first vector and the first non-zero element and a vector according to scalar addition between an existing result vector corresponding to the first row can be calculated as a result vector corresponding to the updated first row. Similarly, a result vector corresponding to a second row can be updated according to scalar multiplication based on the first vector and a second non-zero element. More specifically, a vector according to scalar multiplication based on the first vector and the second non-zero element and a vector according to scalar addition between an existing result vector can be calculated as a result vector corresponding to the updated second row.
[0043] Figure 2 It is shown that the controller 120 controls the plurality of computing units 110 to perform a computation for a first vector, but the present disclosure is not limited thereto. According to an example embodiment, the second order with respect to the plurality of vectors can include an order in which the first controller 140 transmits first information about the first vector to the second controller 150 and then the first controller 140 transmits first information about a second vector to the second controller 150. Accordingly, the controller 120 can control the plurality of computing units 110 to perform a computation based on the first vector and then perform a computation based on the second vector. Further, the second order with respect to the plurality of vectors can include an order in which the first controller 140 finally transmits first information about a third vector to the second controller 150. Here, when the plurality of computing units 110 perform a computation based on the third vector, the computation for the plurality of rows can be terminated.
[0044] Figure 2It is only shown that the first controller 140 sends both the first information and the second information to the second controller 150, but the present disclosure is not limited thereto. For example, the first controller 140 may send only the first information related to the first vector to the second controller 150. In view of this, since the second controller 150 does not receive the second information, the information about the first vector may be sent to all computing units in the plurality of computing units 110. Therefore, the first controller 140 may control the calculation to be performed only in the computing unit corresponding to the second information about the first vector among the plurality of computing units 110. In view of this, the first controller 140 may send information related to the first non-zero element and the first row to the first computing unit 111, send information related to the second non-zero element and the second row to the second computing unit 112, and send the zero element to the third computing unit 113. The calculations in the first computing unit 111 and the second computing unit 112 may be the same as those in the first computing unit 111 and the second computing unit 112. Figure 3A . On the other hand, the third computing unit 113 may perform scalar multiplication based on the zero element and the first vector, and as a result of the scalar multiplication, a vector having zero elements may be calculated. The calculation in the third computing unit 113 may not be reflected in the update of the result vector stored in the buffer 130.
[0045] Figure 3A is a diagram illustrating a memory system including a memory device and a host system according to example embodiments.
[0046] The memory system 10 may be a system for processing calculations related to matrix multiplication in parallel using a memory device equipped with multiple memories. The memory system 10 may include multiple memory devices and a host system 310 for controlling the multiple memory devices. Here, the memory device may be a memory module composed of multiple memory controllers for performing various calculations using multiple memories and information read out from the multiple memories. According to an example embodiment, the memory device may be a dual in-line memory module (DIMM). Referring to Figure 3A , the plurality of memory devices may include a first memory device 300 and a second memory device 301. In an example embodiment, the memory may be a DRAM, but is not limited thereto. In another example embodiment, the memory may be a volatile memory such as a cache memory, a register, and a static random access memory (SRAM).
[0047] The host system 310 can be a system for controlling a memory device equipped with a plurality of memories to perform a calculation in parallel. More specifically, the host system 310 can include a central processing unit (CPU), a controller, or an application specific integrated circuit (ASIC). Also, for example, the host system 310 can include a memory chip such as a DRAM, an SRAM, a phase change RAM (PRAM), a magnetoresistive RAM (MRAM), a ferroelectric RAM (FeRAM), and a resistive RAM (RRAM).
[0048] Referring to Figure 2 , the host system 310 can include a memory device controller 311. In order to minimize the time required to perform a calculation related to a matrix multiplication using high-capacity data stored in a DRAM, the memory device controller 311 can control a calculation related to a matrix multiplication to be performed separately in a plurality of memory controllers included in each of a plurality of memory devices. More specifically, the memory device controller 311 can control a calculation to be performed, which is related to a matrix multiplication between a plurality of rows corresponding to each of a plurality of memory controllers and a target matrix. Here, the plurality of rows can be a set number of rows among rows of a sparse matrix. Here, the set number can be based on the size of the buffer 130 of the memory controller 100. In view of this, the larger the size of the buffer 130 of the memory controller 100, the more result vectors can be stored, and thus the set number can be set larger as the size of the buffer 130 in the memory controller 100 is larger.
[0049] According to an example embodiment, a plurality of rows corresponding to each of a plurality of memory controllers can be determined so that non-zero elements included in the plurality of rows are placed in the same column in the largest number. The memory device controller 311 can transmit a command including information about the determined plurality of rows to the memory controller of each of the plurality of memory devices. The memory controller 100 can receive the command from the memory device controller 311 through the interface 160. Figure 7 Only one interface 160 is illustrated, but the present disclosure is not limited thereto. For example, an interface to interact with the memory device controller 311 and an interface to interact with a DRAM can be separately included in the memory controller 100. It will be described in detail with reference to Figure 3B a specific method of selecting a set number of a plurality of rows among rows of a sparse matrix.
[0050] The first memory device 300 can include a first memory controller 321, a second memory controller 322, a first memory bank 330 including a first DRAM 331, and a second memory bank 340 including a second DRAM 341. Here, each of the first memory controller 321 and the second memory controller 322 can correspond to the memory controller 100. In light of this, in the first memory controller 321, the calculation regarding the plurality of rows included in the command received from the memory device controller 311 can be performed. Further, the memory bank can be a block for distinguishing a plurality of DRAMs included in the memory device. In light of this, each of the plurality of memory controllers included in the memory device can read the information regarding the vector from the DRAM included in the corresponding memory bank. For example, the first memory controller 321 can transmit a command to the first DRAM 331 included in the first memory bank 330 and receive the information regarding the vector from the first DRAM 331. Similarly, the second memory controller 322 can transmit a command to the second DRAM 341 included in the second memory bank 340 and receive the information regarding the vector from the second DRAM 340.
[0051] Figure 3B is a diagram illustrating a memory device according to an example embodiment.
[0052] Referring to Figure 3B The first memory device 300 can include a plurality of memory chips mounted on a board 350. The first memory device 300 can further include a buffer chip 320 that provides command and address information to the plurality of memory chips and an input / output pad 360 disposed at one end of the board 350. The input / output pad 360 can be connected to a data input / output path of each of the plurality of memory chips. In light of this, the input / output pad 360 can correspond to the command pin 161 and the data pin 162.
[0053] Each of the plurality of memory chips including the memory chip 332 and the memory chip 342 can include the memory of the first memory device 300. For example, the plurality of memory chips including the memory chip 332 and the memory chip 342 can include the first DRAM 331 and the second DRAM 341. Further, the buffer chip 320 can provide command and address information to the plurality of memory chips included in the first memory device 300. For example, the buffer chip 320 can include the first memory controller 321 that provides command and address information to the first DRAM 331 and the second memory controller 322 that provides command and address information to the second DRAM 341.
[0054] Figure 4The first memory device 300 is shown to include eight memory chips, but this is merely an example and the present disclosure is not limited thereto. The number of memory chips may vary depending on the data storage capacity of the first memory device 300 or each memory chip. For example, the first memory device 300 may include 16 memory chips. When a memory device including eight memory chips and a memory device including 16 memory chips have the same data storage capacity, the data storage capacity of the eight memory chips may be twice that of the data storage capacity of the 16 memory chips. Furthermore, the number of data input / output paths connected to each memory chip of the memory device including eight memory chips may be twice the number of data input / output paths connected to each memory chip of the memory device including 16 memory chips.
[0055] Figure 4 is a flowchart for explaining an operating method of a first controller according to an example embodiment.
[0056] For operation Figure 1 Each operation of the method of the first controller 140 may be performed by the first controller 140 as described above, and thus reference will be omitted. Figure 5 3. It is obviously understood that within the scope of the present disclosure clearly understood by those skilled in the art to which the exemplary embodiments of the present disclosure pertain, each operation of the method of operating the first controller 140 may be partially changed and / or replaced, or some orders between the operations may be changed.
[0057] In operation S410 , the first controller 140 may recognize that the first queue 143 and the second queue 144 are empty.
[0058] According to example embodiments, in response to completion of a calculation related to matrix multiplication, information stored in the first queue 143 and the second queue 144 may be deleted. In other words, the first queue 143 and the second queue 144 being empty may indicate that there is no calculation related to matrix multiplication currently being executed in the memory controller 100 and that the previously executed calculation has been completed.
[0059] In operation S420 , the first controller 140 may identify a plurality of vectors and at least one row to which each of the plurality of vectors corresponds among the plurality of rows.
[0060] According to an example embodiment, the first controller 140 can extract first information about a plurality of vectors required for the calculation of a plurality of rows of a matrix, and second information about at least one row to which each of the plurality of vectors among the plurality of vectors corresponds. More specifically, by sequentially searching for a non-zero element included in the plurality of rows for each column according to a first order that is preset, the first controller 140 can extract the plurality of vectors and at least one row to which each of the plurality of vectors corresponds. Here, the extraction of the plurality of vectors and at least one row to which each of the plurality of vectors corresponds can be sequentially extracting the plurality of vectors and at least one row to which each of the plurality of vectors corresponds based on a second order of the plurality of vectors.
[0061] In operation S430, the first controller 140 can store the first information about the plurality of vectors in the first queue 143 and store the second information about at least one row in the second queue 144.
[0062] According to an example embodiment, when all first information about all vectors required for the calculation of a plurality of rows of a matrix is extracted, the first controller 140 can store the first information about the plurality of vectors in the first queue 143 and store the second information about at least one row in the second queue 144. The first queue 143 and the second queue 144 can have a FIFO data structure, and thus the first information identified in operation S420 can be sequentially stored in the first queue 143 based on a second order, and the second information identified in operation S420 can be sequentially stored in the second queue 144 based on the second order, which is an order in which the plurality of vectors are identified.
[0063] In operation S440, the first controller 140 can sequentially transmit the first information and the second information to the second controller 150.
[0064] According to an example embodiment, the first queue 143 and the second queue 144 have a FIFO data structure, and thus the first controller 140 can sequentially transmit the first information and the second information to the second controller 150 based on a second order, which is an order in which the first information and the second information are stored in the first queue 143 and the second queue 144. For example, the first controller 140 can transmit the first information about a first vector, which is stored first in the first queue 143 and the second queue 144, and the second information about at least one row corresponding to the first vector to the second controller 150. Thereafter, the first controller 140 can transmit the first information about a second vector, which is identified after the first vector, and the second information about at least one row corresponding to the second vector to the second controller 150. The first controller 140 can repeat the operation of transmitting the first information and the second information until all the first information about the plurality of vectors is transmitted to the second controller 150. In this regard, the total number of times that the first information and the second information are transmitted to the second controller 150 can be equal to the number of the plurality of vectors. When all the first information about the plurality of vectors is transmitted to the second controller 150, the calculation about the plurality of rows of the sparse matrix is completed, and thus all the information stored in the first queue 143 and the second queue 144 can be deleted.
[0065] Figure 5 FIG. 1 is a diagram illustrating a method of performing a calculation related to matrix multiplication based on a plurality of rows of a sparse matrix and a target matrix according to an example embodiment.
[0066] According to an example embodiment, the sparse matrix 500 can be a matrix in which most of the values of the matrix elements are zero. According to an example embodiment, the sparse matrix 500 can be a sparse matrix used in a calculation related to a graph neural network (GNN). Here, the GNN can be an artificial neural network for analyzing and modeling a graph structure. The graph can be a data structure composed of nodes and edges representing a relationship between two nodes. In this regard, the i-th row and the i-th column of the sparse matrix corresponding to the graph can correspond to the i-th node, and the elements corresponding to (i, j) and (j, i) in the sparse matrix corresponding to the graph can correspond to an edge between the i-th node and the j-th node. For example, when a first node and a second node are connected, an element corresponding to the first node and the second node in the sparse matrix can be a non-zero element, and when the first node is not connected to the second node, an element corresponding to the first node and the second node in the sparse matrix can be a zero element. Due to the nature of the graph, the sparse matrix used in the GNN-related calculation can be a symmetric matrix, and the non-zero elements included in the sparse matrix used in the GNN-related calculation can be 1. The sparse matrix 500 can be a sparse matrix used in a GNN-related calculation, and can be a sparse matrix used in a calculation related to a deep learning recommendation model (DLRM), but the sparse matrix 500 is not limited thereto.
[0067] According to an example embodiment, each row of the sparse matrix 500 can include non-zero elements. More specifically, each row of the sparse matrix 500 can include the same number of five non-zero elements, but is not limited thereto. Referring to FIG. 5, a first element, a fourth element, a sixth element, a tenth element, and an eleventh element of a first row of the sparse matrix 500 can be non-zero elements, and other elements in the first row can be zero elements. Figure 5
[0068] According to an example embodiment, the target matrix 520 can be a matrix composed of vectors that are objects of matrix multiplication with the sparse matrix 500. As described above, the rows of the sparse matrix 500 can correspond to at least some vectors of the target matrix 520. More specifically, the vectors of the target matrix 520 corresponding to the rows of the sparse matrix 500 can be identified based on columns corresponding to the non-zero elements included in the rows of the sparse matrix 500. Referring to FIG. 5, the vectors of the target matrix 520 corresponding to the first row of the sparse matrix 500 can be identified based on the columns corresponding to the non-zero elements included in the first row of the sparse matrix 500. Figure 6 The vectors of the target matrix 520 corresponding to the first row of the sparse matrix 500 can include a vector 521, a vector 522, a vector 523, a vector 524, and a vector 525 of the target matrix 520. The vector 521 of the target matrix 520 can correspond to the non-zero element 511 that is the first element of the first row of the sparse matrix 500, the vector 522 of the target matrix 520 can correspond to the fourth element of the first row of the sparse matrix 500, the vector 523 of the target matrix 520 can correspond to the sixth element of the first row of the sparse matrix 500, the vector 524 of the target matrix 520 can correspond to the tenth element of the first row of the sparse matrix 500, and the vector 525 of the target matrix 520 can correspond to the eleventh element of the first row of the sparse matrix 500.
[0069] According to an example embodiment, the result matrix 530 can be a matrix that is a calculation result related to matrix multiplication based on the sparse matrix 500 and the target matrix 520. More specifically, a result vector corresponding to an i-th row of the result matrix 530 can be calculated based on non-zero elements included in the i-th row of the sparse matrix 500 and vectors of the target matrix 520 corresponding to the i-th row. For example, a result vector 531 corresponding to a first row of the result matrix 530 can be calculated based on non-zero elements included in the first row of the sparse matrix 500 and vectors of the target matrix 520 corresponding to the first row of the sparse matrix 500, a result vector 532 corresponding to a second row of the result matrix 530 can be calculated based on non-zero elements included in the second row of the sparse matrix 500 and vectors of the target matrix 520 corresponding to the second row of the sparse matrix 500, and a result vector 533 corresponding to a third row of the result matrix 530 can be calculated based on non-zero elements included in the third row of the sparse matrix 500 and vectors of the target matrix 520 corresponding to the third row of the sparse matrix 500.
[0070] According to an example embodiment, the memory device controller 311 can control the calculations related to the multiplication of the sparse matrix 500 to be separately performed in the first memory device 300 and the second memory device 301. More specifically, the memory device controller 311 can control the first memory device 300 to perform the calculations of the first to sixth rows of the sparse matrix 500, and the memory device controller 311 can control the second memory device 301 to perform the calculations of the seventh to twelfth rows of the sparse matrix 500. The calculations of the first to third rows can be performed in the first memory controller 321 of the first memory device 300, and the calculations of the fourth to sixth rows can be performed in the second memory controller 322 of the first memory device 300. In other words, the calculations related to the multiplication of the sparse matrix 500 can be separately performed in the plurality of memory controllers included in each of the plurality of memory devices. Hereinafter, a specific example embodiment for identifying the plurality of vectors and at least one row to which each of the plurality of vectors corresponds required for the calculations of the first to third rows of the sparse matrix 500 is described. In view of this, the plurality of rows 510 in the sparse matrix 500 can include the first to third rows of the sparse matrix 500.
[0071] According to an example embodiment, the extractor 142 can identify the plurality of vectors and at least one row to which each of the plurality of vectors corresponds by sequentially searching for the non-zero elements included in the plurality of rows for each column according to a predetermined first order. More specifically, the extractor 142 sequentially searches for the non-zero elements included in the plurality of rows from the first row to the last row, and when a column contains a non-zero element, sequentially performs an operation of identifying first information about a vector corresponding to the column of the non-zero element and second information about a row of the non-zero element, an operation of determining whether a next column contains a non-zero element, and when the column does not contain a non-zero element, performs an operation of determining whether a next column contains a non-zero element.
[0072] For example, the extractor 142 can identify that the first column of the plurality of rows 510 includes two non-zero elements, which are non-zero element 511 and non-zero element 512. In view of this, a vector corresponding to the first column can be identified as a vector required for the calculation of the plurality of rows, and the first row corresponding to the non-zero element 511 and the third row corresponding to the non-zero element 512 can be identified as rows corresponding to the vector corresponding to the first column. The extractor 142, which completes the identification, can perform an operation of determining whether the next column, i.e., the second column, contains a non-zero element. As similar operations are repeatedly performed, the extractor 142 can identify that the eighth column of the sparse matrix 500 does not contain a non-zero element. In other words, a vector corresponding to the eighth column can be identified as an unnecessary vector for the calculation of the plurality of rows. Subsequently, the extractor 142 can identify that the next column, i.e., the ninth column, contains a non-zero element 513. In view of this, a vector corresponding to the ninth column can be identified as a vector required for the calculation of the plurality of rows, and the second row corresponding to the non-zero element 513 can be identified as a row corresponding to the vector corresponding to the ninth column. Finally, the extractor 142 can identify that the twelfth column includes two non-zero elements, which are non-zero element 514 and non-zero element 515. In view of this, a vector corresponding to the twelfth column can be identified as a vector required for the calculation of the plurality of rows, and the second row corresponding to the non-zero element 514 and the third row corresponding to the non-zero element 515 can be identified as rows corresponding to the vector corresponding to the twelfth column. Accordingly, a plurality of vectors can be identified based on the second order for the plurality of vectors without overlapping.
[0073] Figure 5 is a diagram for explaining the number of times the second controller transmits a command including address information to a memory when performing a calculation related to matrix multiplication according to an example embodiment.
[0074] The order 602 can represent an order in which the second controller 150 transmits a command including address information to a memory when performing a calculation related to matrix multiplication with respect to the plurality of rows 510 according to an example embodiment. Figure 6 As described above, the indices included in the order 601 and the order 602 can correspond to each row of the target matrix 520. For example, the i-th row of the target matrix 520 can correspond to the index i-1.
[0075] With respect to the order 602, according to an example embodiment, the first controller 140 can perform an operation to identify a plurality of vectors required for the calculation of the plurality of rows 510 without overlapping according to the second order of the plurality of vectors. Accordingly, information about the plurality of vectors can be sequentially stored in the first queue 143 according to the second order. In addition, the first controller 140 can sequentially transmit first information about the vectors stored in the first queue 143 to the second controller 150, and the second controller 150 can sequentially transmit a command including the first information to the memory. Referring toFigure 5 In other words, the number of times that the command is sequentially transmitted to the memory based on the sequence 602 can be 11, which is the number of columns corresponding to the non-zero elements included in the plurality of rows 510. Also, it can take as much time as T4 to sequentially transmit the command to the memory based on the sequence 602.
[0076] On the other hand, when performing a calculation related to matrix multiplication based on each of the plurality of rows 510 and the target matrix 520, the number of times that the command is transmitted to the memory can be 15 times. More specifically, a command including address information of five vectors corresponding to a first row of the sparse matrix 500 can be transmitted to the memory from T0 to T1, a command including address information of five vectors corresponding to a second row of the sparse matrix 500 can be transmitted to the memory from T1 to T2, and a command including address information of five vectors corresponding to a third row can be transmitted to the memory from T2 to T3. In other words, the number of times that the command is sequentially transmitted to the memory based on the sequence 602 can be less than 15, which is the number of times that the command is transmitted to the memory based on the sequence 601.
[0077] In other words, according to Figure 5 the example embodiment, when performing a calculation related to matrix multiplication on the plurality of rows 510, the command including address information of each vector is transmitted to the memory only once, and thus the number of accesses to the memory can be reduced. In particular, although Figure 7 Although it is shown that the size of the sparse matrix 500 is 12x12 and the size of the target matrix 520 is 12x5, the size of the vector included in the target matrix 520 can be very large. For example, the size of the vector included in the target matrix 520 can be very large, exceeding 1K. Thus, when performing a calculation related to matrix multiplication using a high-capacity vector, the time required to read a specific vector can significantly become longer. In view of this, as the number of accesses to the memory is reduced, the time required to read a specific vector can be greatly reduced.
[0078] Figure 6 is a diagram illustrating a method of selecting a preset number of a plurality of rows among rows of a sparse matrix.
[0079] According to the example embodiment, the number of times that a command including address information is transmitted to the memory when sequentially performing a calculation related to matrix multiplication based on each row and a target matrix is less than the number of times that a command including address information is transmitted to the memory when performing a calculation according to Figure 6The difference between the number of times the commands including address information are transmitted to the memory in relation to the matrix multiplication of the example embodiment can be based on the number of times the non-zero elements included in the plurality of rows overlap in the same column. More specifically, the more the non-zero elements included in the plurality of rows overlap in the same column, the greater the difference in the count that can be calculated. Referring to Figure 7 The difference between the number of times the commands are sequentially transmitted to the memory based on the order 601 and the number of times the commands are sequentially transmitted to the memory based on the order 602 can be 4, i.e., the number of times the non-zero elements included in the plurality of rows 510 overlap in the same column is 4.
[0080] In other words, when the plurality of rows are configured so that the non-zero elements among the rows of the sparse matrix are placed in the same column in the maximum number, the number of operations in which the second controller 150 reads information related to the vector from the memory can be minimized. In light of this, the commands transmitted to the plurality of memory devices by the memory device controller 311 can include information related to the plurality of rows corresponding to each of the plurality of memory controllers. Here, the plurality of rows can be determined by the memory device controller 311 so that the non-zero elements included in a set number of rows of the matrix are placed in the same column in the maximum number. However, determining the plurality of rows corresponding to each of the plurality of memory controllers can also take time. In other words, determining the plurality of rows corresponding to each of the plurality of memory controllers in a short time while ensuring that many non-zero elements included in the plurality of rows are placed in the same column can be efficient in reducing the total time required for the calculation related to the matrix multiplication.
[0081] According to an example embodiment, when the sparse matrix is a symmetric matrix, the memory device controller 311 can more efficiently determine the plurality of rows by using the characteristics of the symmetric matrix. In an example embodiment, the sparse matrix used in the GNN-related calculation can be a symmetric matrix. In light of this, if the i-th node and the j-th node are connected to each other, the elements corresponding to (i,j) and (j,i) among the elements of the sparse matrix can be non-zero elements. Conversely, if the i-th node and the j-th node are not connected to each other, the elements corresponding to (i,j) and (j,i) among the elements of the sparse matrix can be zero elements. In addition, among the elements of the sparse matrix, the elements located on the diagonal line can be non-zero elements.
[0082] According to an example embodiment, when the plurality of rows include a first row and at least one second row, the at least one second row can be determined based on the columns corresponding to the non-zero elements included in the first row. In other words, the at least one row can be determined heuristically based on the columns corresponding to the non-zero elements included in the first row. In light of this, referring to Figure 7The row 710 of the sparse matrix 700, which is a symmetric matrix, can include a non-zero element 711 corresponding to the first column, a non-zero element 712 corresponding to the fourth column, a non-zero element 713 corresponding to the sixth column, a non-zero element 714 corresponding to the tenth column, and a non-zero element 715 corresponding to the eleventh column. When the plurality of rows consists of a set number of rows, for example, three rows, the at least one second row can be determined based on any two elements among the non-zero element 712, the non-zero element 713, the non-zero element 714, and the non-zero element 715. Figure 6 It is shown that the at least one second row is determined based on the fourth column corresponding to the non-zero element 712 and the sixth column corresponding to the non-zero element 713. In view of this, the row corresponding to the non-zero element 712 can be the fourth row 720 of the sparse matrix 700, and the row corresponding to the non-zero element 713 can be the sixth row 730 of the sparse matrix 700.
[0083] When performing the matrix multiplication-related calculation according to the example embodiment of operation S820 with respect to the plurality of rows 740 including the row 710, the fourth row 720, and the sixth row 730, the number of times that the second controller 150 sequentially sends a command to the memory can be 8, which is the number of columns corresponding to the non-zero elements included in the plurality of rows 740. In other words, by determining the plurality of rows using the characteristics of the symmetric matrix, the number of operations of reading information about the vector from the memory can be significantly reduced. Accordingly, the time required for the memory controller 100 to perform the calculation can also be greatly reduced. Figure 8
[0084] Figure 8 is a flowchart illustrating an operation method of a memory controller according to an example embodiment.
[0085] Figures 1 to 7 Each operation of the operation method of operation S820 can be performed by the memory controller 100 as described above, and thus a description overlapping with the description of operation S820 will be omitted. Here, the memory controller 100 can include a plurality of calculation units 110 performing a calculation related to matrix multiplication, and a controller 120. The controller 120 can include a first controller 140 and a second controller 150 configured to perform an operation to access data stored in the memory. The first controller 140 can include a matrix buffer 141 storing a plurality of rows, an extractor 142 configured to identify a plurality of vectors and at least one row corresponding to each of the plurality of vectors, a first queue 143 storing first information, and a second queue 144 storing second information. In a case where the first queue 14 and the second queue 144 are empty, the extractor 142 is configured to identify the plurality of vectors related to operation S810 and at least one row corresponding to each of the plurality of vectors related to operation S820.
[0086] In operation S810, the memory controller 100 can identify first information indicating a plurality of vectors related to a calculation for a plurality of rows of a matrix. The first information can include address information in which each of the plurality of vectors is stored in the memory. The plurality of vectors can be identified based on columns corresponding to non-zero elements included in the plurality of rows.
[0087] According to an example embodiment, the plurality of rows is determined so that non-zero elements included in a set number of rows of the matrix are placed in the same column in a maximum number. The set number is determined based on a size of the buffer 130 included in the memory controller 100. Alternatively, in a case where the plurality of rows includes a first row and at least one second row, the at least one second row is determined based on columns corresponding to non-zero elements included in the first row.
[0088] In operation S820, the memory controller 100 can identify second information indicating at least one row corresponding to each of a plurality of vectors among the plurality of rows. The second information can include index information of at least one row corresponding to each of the plurality of vectors among the plurality of rows. At least one first row corresponding to a first vector among the plurality of vectors can be identified based on a row of a first non-zero element included in a first column corresponding to the first vector.
[0089] According to an example embodiment, the memory controller 100 can identify the plurality of vectors and at least one row corresponding to each of the plurality of vectors by sequentially searching for non-zero elements included in the plurality of rows of the matrix for each column in a predetermined order.
[0090] According to an example embodiment, the first controller 140 can sequentially transmit the address information and the index information to the second controller 150. The second controller 150 can sequentially transmit a command including the address information to the memory. Here, the memory can include a DRAM, but is not limited thereto. The number of times of sequentially transmitting the command to the memory can be less than the number of non-zero elements included in the plurality of rows. More specifically, the number of times of sequentially transmitting the command to the memory is the number of columns corresponding to the non-zero elements included in the plurality of rows. The second controller 150 can sequentially receive information related to the plurality of vectors from the memory and transmit the information related to the plurality of vectors to the plurality of computing units 110.
[0091] In operation S830, the memory controller 100 can control the plurality of computing units 110 to perform a calculation on the plurality of rows by sequentially inputting the plurality of vectors into the plurality of computing units 110 based on the first information and the second information. More specifically, by sequentially inputting the plurality of vectors and non-zero elements into the plurality of computing units 110 based on the first information, the second information, and the non-zero elements included in the plurality of rows, the memory controller 100 can control the plurality of computing units 110 to perform a calculation on the plurality of rows.
[0092] According to an example embodiment, the second information includes information about a first row and a second row among the plurality of rows to which a first vector among the plurality of vectors corresponds. The controller 120 is further configured to input the first vector to i) a first computing unit 111 corresponding to the first row, the plurality of computing units 110 including the first computing unit 111, and ii) a second computing unit 112 corresponding to the second row, the plurality of computing units 110 including the second computing unit 112, control the first computing unit 111 to perform a first computation of the first row, and control the second computing unit 112 to perform a second computation of the second row.
[0093] According to an example embodiment, the buffer 130 included in the memory controller 100 can store a result vector related to the computation, and in the case where the second information includes information about a first row among the plurality of rows to which a first vector among the plurality of vectors corresponds, a result vector corresponding to the first row is updated based on a computation in a first computing unit 111 corresponding to the first row among the plurality of computing units.
[0094] According to an example embodiment, in the case where the order in which the plurality of vectors are sequentially input includes inputting a second vector after inputting a first vector, the controller 120 can be further configured to control the plurality of computing units 110 to perform a third computation based on the first vector and then perform a fourth computation based on the second vector.
[0095] According to an example embodiment, the memory controller 100 can include an interface 160. The interface 160 can receive a command from a host system 310 including a memory device controller 311 to perform a computation.
[0096] The memory controller 100 or terminal according to the above-described example embodiments can include a processor, a memory for storing and running program data, a permanent storage device such as a disk drive, and / or a user interface device such as a communication port, a touch panel, a key, and / or a button for communicating with an external device. The method implemented as a software module or algorithm can be stored in a computer-readable recording medium as computer-readable code or program instructions executable on a processor. Here, the computer-readable recording medium includes a magnetic storage medium (e.g., a ROM, a RAM, a floppy disk, and a hard disk), and an optically readable medium (e.g., a CD-ROM and a DVD). The computer-readable recording medium can be distributed among computer systems connected to a network, so that the computer-readable code can be stored and executed in a distributed manner. The medium can be read by a computer, stored in a memory, and run on a processor.
[0097] Example embodiments can be represented by functional block elements and various processing steps. The functional blocks can be realized in any number of hardware and / or software configurations, including integrated circuit configurations such as memory, processing, logic, and / or look-up tables that can be controlled by one or more microprocessors or other control devices to perform various functions. Similarly, like these elements can be implemented as software programming or software elements, example embodiments can be implemented in programming or scripting languages such as C, C++, Java, assembly, and / or others, including various algorithms implemented as combinations of data structures, procedures, routines, or other programming constructs. Functional aspects can be implemented in algorithms running on one or more processors. Moreover, example embodiments can employ existing technologies for electronic environment setup, signal processing, and / or data processing. Terms such as "mechanism," "element," "means," and "configuration" can be broadly used, not limited to mechanical and physical elements. These terms can include the meaning of a series of software routines associated with a processor or the like.
[0098] The above example embodiments are merely examples, and other embodiments can be implemented within the scope of the claims described later.
Claims
1. A memory controller comprising: a plurality of computing units, wherein the plurality of computing units are configured to perform a computation related to a matrix multiplication; and a controller configured to: identify first information indicating a plurality of vectors required for the computation for a plurality of rows of a matrix, identify second information indicating at least one row to which each of the plurality of vectors among the plurality of rows corresponds, and control the plurality of computing units to perform the computation for the plurality of rows by sequentially inputting the plurality of vectors into the plurality of computing units based on the first information and the second information.
2. The memory controller of claim 1, wherein, the second information includes information about a first row and a second row among the plurality of rows to which a first vector among the plurality of vectors corresponds, and wherein the controller is further configured to: input a first vector to i) a first computing unit corresponding to the first row, the plurality of computing units including the first computing unit, and ii) a second computing unit corresponding to the second row, the plurality of computing units including the second computing unit, control the first computing unit to perform a first computation of the first row, and control the second computing unit to perform a second computation of the second row.
3. The memory controller of claim 1, wherein, the controller comprises: a first controller including a matrix buffer storing the plurality of rows, an extractor configured to identify the plurality of vectors and the at least one row to which each of the plurality of vectors corresponds, a first queue storing the first information, and a second queue storing the second information, wherein the first information includes address information in which each of the plurality of vectors is stored in a memory, and wherein the second information includes index information of the at least one row to which each of the plurality of vectors among the plurality of rows corresponds.
4. The memory controller of claim 3, wherein, the extractor is configured to identify the plurality of vectors and the at least one row to which each of the plurality of vectors corresponds by sequentially searching for a non-zero element included in the plurality of rows for each column in a predetermined order.
5. The memory controller of claim 3, wherein, the controller further comprises a second controller configured to perform an operation to access data stored in the memory, wherein the first controller sequentially transmits the address information and the index information to the second controller, and wherein the second controller sequentially transmits a command including the address information to the memory, sequentially receives information about the plurality of vectors from the memory, and transmits information about the plurality of vectors to the plurality of computing units.
6. The memory controller of claim 5, wherein, a number of times of sequentially transmitting the command to the memory is less than a number of non-zero elements included in the plurality of rows.
7. The memory controller of claim 5, wherein, a number of times of sequentially transmitting the command to the memory is a number of columns corresponding to the non-zero elements included in the plurality of rows.
8. The memory controller of claim 1, wherein, the plurality of vectors are identified based on the columns corresponding to the non-zero elements included in the plurality of rows, and wherein at least one first row corresponding to a first vector among the plurality of vectors is identified based on a row including a first non-zero element included in a first column corresponding to the first vector.
9. The memory controller of claim 1, wherein, The controller is further configured to control the plurality of computing units to perform the computation for the plurality of rows by sequentially inputting the plurality of vectors and the non-zero elements into the plurality of computing units based on the first information, the second information, and the non-zero elements included in the plurality of rows. 10.The memory controller of claim 1, the memory controller comprising a buffer storing a result vector related to the computation, wherein, in case that the second information includes information about the first row corresponding to a first vector among the plurality of vectors among the plurality of rows, updating a result vector corresponding to the first row based on the computation in a first computing unit among the plurality of computing units corresponding to the first row.
11. The memory controller of claim 1, wherein the order in which the plurality of vectors are sequentially inputted includes inputting a second vector after a first vector, and wherein the controller is further configured to control the plurality of computing units to perform a third computation based on the first vector and then perform a fourth computation based on the second vector.
12. The memory controller of claim 3, wherein, The extractor is configured to identify the plurality of vectors and the at least one row corresponding to each of the plurality of vectors in case that the first queue and the second queue are empty.
13. The memory controller of claim 1, the memory controller further comprising an interface, wherein, The interface receives a command from a host system including a memory device controller to perform the computation.
14. The memory controller of claim 1, wherein, The plurality of rows are determined so that non-zero elements included in a set number of rows of the matrix are placed in the same column in a maximum number.
15. The memory controller of claim 1, wherein, The plurality of rows include the first row and at least one second row, wherein the at least one second row is determined based on a column corresponding to a non-zero element included in the first row.
16. The memory controller of claim 14, further comprising a buffer, wherein, The buffer is configured to store a result vector related to the computation, wherein the set number is determined based on a size of the buffer.
17. The memory controller of claim 3, wherein, The memory includes a dynamic random access memory (DRAM). 18.An operating method of a memory controller including a controller and a plurality of computing units configured to perform a computation related to matrix multiplication, the operating method comprising: identifying first information indicating a plurality of vectors required for the computation for a plurality of rows of a matrix; identifying second information indicating at least one row corresponding to each of the plurality of vectors among the plurality of rows; and controlling the plurality of computing units to perform the computation for the plurality of rows by sequentially inputting the plurality of vectors into the plurality of computing units based on the first information and the second information. 19.A non-transitory computer-readable recording medium having a program for executing the operation method of claim 18 on a computer.
20. A memory system comprising: a host system including a memory device controller; and a plurality of memory controllers operating according to commands received from the memory device controller, wherein the plurality of memory controllers includes a first memory controller, wherein the first memory controller includes: a plurality of computing units configured to perform computations related to matrix multiplication; and a controller configured to: identify first information indicating a plurality of vectors required for the computations for a plurality of rows of a matrix, identify second information indicating at least one row to which each of the plurality of vectors among the plurality of rows corresponds, and control the plurality of computing units to perform computations for the plurality of rows by sequentially inputting the plurality of vectors into the plurality of computing units based on the first information and the second information.