Data caching method and host system

By using the data caching method in artificial intelligence applications, by reading and generating the key matrix and value matrix of parts and writing them into the physical units of multiple channels, the problem of slow reading and writing speed of rewritable non-volatile memory modules in artificial intelligence applications is solved, and more efficient data processing is achieved.

CN120123262APending Publication Date: 2025-06-10PHISON ELECTRONICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510184431.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In artificial intelligence applications, how to improve the read and write speed of rewritable nonvolatile memory modules in frequent data writing and reading.

Method used

A data cache method is proposed, by obtaining the current character in the inference stage, reading the previous key matrix and value matrix from the rewritable non-volatile memory module, generating the current key matrix and value matrix according to the current character and the previous matrix, and writing it in part to the physical unit of multiple channels.

Benefits of technology

Through partial writing and reading, the read and write speed of data cache is improved and the performance of artificial intelligence algorithms in rewritable non-volatile memory modules is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123262A_ABST
    Figure CN120123262A_ABST
Patent Text Reader

Abstract

The invention provides a data caching method and a host system. In one embodiment, a host system includes a rewritable non-volatile memory module that includes a plurality of lanes. The data caching method comprises the following steps: acquiring a current substitute symbol in a reasoning stage; reading a previous key matrix and a previous value matrix from the rewritable non-volatile memory module; generating a current key matrix according to the current generation symbol and the previous key matrix; generating a current value matrix according to the current substitute and the previous value matrix; writing a plurality of parts of the current key matrix into different channels; and writing a plurality of portions of the current value matrix to different channels. Therefore, the caching speed can be increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a data caching method, a memory storage device, and a memory control circuit unit for an artificial intelligence algorithm, and more particularly to a data caching method and a host system. Background Art

[0002] The growth of portable electronic devices such as mobile phones and laptop computers has been very rapid in recent years, leading to a sharp increase in consumers' demand for storage media. Since rewritable non-volatile memory modules (e.g., flash memory) have characteristics such as data non-volatility, power saving, small size, and no mechanical structure, they are very suitable for being built into various portable electronic devices exemplified above.

[0003] In recent years, with the development of artificial intelligence, there have been more and more related applications. Artificial intelligence involves a large amount of calculations and is accompanied by frequent data writing and reading. When using a rewritable non-volatile memory module to complete artificial intelligence-related calculations, how to improve the reading and writing speed is an issue of concern to those skilled in the art. Summary of the Invention

[0004] To solve the above problems, the present disclosure proposes a data caching method and a host system.

[0005] The present disclosure proposes a data caching method for a rewritable non-volatile memory module, which includes a plurality of channels. Each channel includes a plurality of physical units. The data caching method includes: obtaining a current token in the inference stage; reading a previous key matrix and a previous value matrix from the rewritable non-volatile memory module; generating a current key matrix according to the current token and the previous key matrix; generating a current value matrix according to the current token and the previous value matrix; writing a part of the current key matrix into a physical unit in one of the channels, and storing another part of the current key matrix into a physical unit in another channel; and writing a part of the current value matrix into a physical unit in one of the channels, and storing another part of the current value matrix into a physical unit in another channel.

[0006] In an embodiment of the present disclosure, the above-mentioned part of the current key matrix, another part of the current key matrix, part of the current value matrix, and another part of the current value matrix are written into the corresponding physical units in a single-level storage cell programming mode.

[0007] In an embodiment of the present disclosure, the above-mentioned part of the current key matrix belongs to a first token, and another part of the current key matrix belongs to a second token, and the first token is different from the second token.

[0008] In one embodiment of the present disclosure, a part of the current key matrix described above belongs to a first feature, and another part of the current key matrix belongs to a second feature, and the first feature is different from the second feature.

[0009] In one embodiment of the present disclosure, the above data caching method further includes: calculating at least one query vector, at least one key vector, and at least one value vector according to a current token; multiplying the query vector by the current key matrix to obtain a temporary vector; and multiplying the temporary vector by the current value matrix to obtain an attention output. The step of generating the current key matrix according to the current token and the previous key matrix includes: combining the key vector with the previous key matrix to generate the current key matrix. The step of generating the current value matrix according to the current token and the previous value matrix includes: combining the value vector with the previous value matrix to generate the current value matrix.

[0010] In one embodiment of the present disclosure, the above query vector includes a first query vector and a second query vector, the key vector includes a first key vector and a second key vector, the value vector includes a first value vector and a second value vector, the previous key matrix includes a first previous key matrix and a second previous key matrix, and the previous value matrix includes a first previous value matrix and a second previous value matrix. The step of combining the key vector with the previous key matrix to generate the current key matrix includes: combining the first key vector with the first previous key matrix to generate a first current key matrix, and combining the second key vector with the second previous key matrix to generate a second current key matrix. The step of combining the value vector with the previous value matrix to generate at least one current value matrix includes: combining the first value vector with the first previous value matrix to generate a first current value matrix, and combining the second value vector with the second previous value matrix to generate a second current value matrix. A part of the current key matrix described above belongs to a first current key matrix, and another part of the current key matrix belongs to a second current key matrix, where a part of the current value matrix belongs to a first current value matrix, and another part of the current value matrix belongs to a second current value matrix.

[0011] In one embodiment of the present disclosure, the above current token is input to the first layer of a neural network, the neural network further includes a second layer, and the data caching method further includes: writing a part of the current key matrix corresponding to the second layer and a part of the current key matrix corresponding to the first layer into physical units in the same channel.

[0012] In one embodiment of the present disclosure, the above data caching method further includes: obtaining a next token; reading a part of the current key matrix and another part of the current key matrix in parallel from a channel; and reading a part of the current value matrix and another part of the current value matrix in parallel from the channel.

[0013] In an embodiment of the present disclosure, the steps of calculating the query vector, key vector, and value vector according to the current token include: multiplying the current token by a query weight matrix to obtain the query vector; multiplying the current token by a key weight matrix to obtain the key vector; and multiplying the current token by a value weight matrix to obtain the value vector.

[0014] In another aspect, an embodiment of the present invention provides a host system including a memory storage device and a processor. The memory storage device includes a rewritable non-volatile memory module, which includes a plurality of channels, and each channel includes a plurality of physical units. The processor is electrically connected to the memory storage device and is configured to execute a plurality of steps: obtaining a current token in the inference stage; reading a previous key matrix and a previous value matrix from the rewritable non-volatile memory module; generating a current key matrix according to the current token and the previous key matrix; and generating a current value matrix according to the current token and the previous value matrix. The memory storage device is configured to write a part of the current key matrix into the physical units in one of the channels, store another part of the current key matrix into the physical units in another channel, write a part of the current value matrix into the physical units in one of the channels, and store another part of the current value matrix into the physical units in another channel.

[0015] In an embodiment of the present disclosure, the above-mentioned memory storage device is configured to program the part of the current key matrix, another part of the current key matrix, the part of the current value matrix, and another part of the current value matrix into the corresponding physical units in a single-level storage unit programming mode.

[0016] In an embodiment of the present disclosure, the above-mentioned current token is input to the first layer of a neural network, and the neural network further includes a second layer. The memory storage device is configured to write the part of the current key matrix corresponding to the second layer and the part of the current key matrix corresponding to the first layer into the physical units in the same channel.

[0017] In an embodiment of the present disclosure, the above-mentioned host system is further configured to obtain the next token. The memory storage device reads the part of the current key matrix and another part of the current key matrix in parallel from the channels, and reads the part of the current value matrix and another part of the current value matrix in parallel from the channels.

[0018] To make the above features and advantages of the present invention more obvious and understandable, the following specific embodiments are given and described in detail in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a schematic diagram of a host system and an input / output (I / O) device according to an exemplary embodiment of the present invention;

[0020] Figure 2Schematic diagrams of a host system, a memory storage device, and an I / O device shown according to exemplary embodiments of the present invention;

[0021] Figure 3 Schematic diagram of a memory storage device shown according to exemplary embodiments of the present invention;

[0022] Figure 4 Schematic diagram showing multiple channels according to an embodiment;

[0023] Figure 5 Schematic diagram showing a neural network in the inference phase according to an embodiment;

[0024] Figure 6 Schematic diagram showing the calculation of self-attention according to an embodiment;

[0025] Figure 7 Schematic diagram showing the calculation of self-attention according to an embodiment;

[0026] Figure 8 Schematic diagram showing matrix operations regarding attention according to an embodiment;

[0027] Figure 9 Schematic diagram showing the operation when applying a KV cache according to an embodiment;

[0028] Figure 10 Schematic diagram showing caches of multiple layers according to an embodiment;

[0029] Figure 11 Schematic diagram showing multiple heads according to an embodiment;

[0030] Figure 12 Flowchart showing a data caching method according to an embodiment. Detailed implementation manners

[0031] Some embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings. For the reference numerals of the components cited in the following description, when the same reference numerals appear in different drawings, they will be regarded as the same or similar components. These embodiments are only a part of the present invention and do not disclose all the implementable ways of the present invention. More precisely, these embodiments are only examples of the systems and methods within the scope of the patent application of the present invention.

[0032] Regarding the "first", "second", etc. used herein, they do not particularly refer to the meaning of order or sequence, but are only used to distinguish components or operations described with the same technical terms.

[0033] Generally, a memory storage device (also referred to as a memory storage system) includes a rewritable non-volatile memory module and a controller (also referred to as control circuitry). The memory storage device can be used with a host system so that the host system can write data to the memory storage device or read data from the memory storage device.

[0034] Figure 1 FIG. 4 is a schematic diagram of a host system and an input / output (I / O) device according to an exemplary embodiment of the present invention. Figure 2 FIG. 6 is a schematic diagram of a host system, a memory storage device, and an I / O device according to an exemplary embodiment of the present invention.

[0035] Please refer to Figure 1 and Figure 2 FIG. 13, the host system 11 is a computer system, which can be a desktop computer, a server, a distributed system, a laptop computer, etc., and the present invention is not limited thereto. The host system 11 may include a processor 111, a random access memory (RAM) 112, a read only memory (ROM) 113, and a data transmission interface 114. The processor 111, the random access memory 112, the read only memory 113, and the data transmission interface 114 may be coupled to a system bus 110. The processor 111 may be a graphic processing unit (GPU), a tensor processing unit (TPU), a neural processing unit (NPU), a central processing unit, etc. In some embodiments, a memory is also provided in the processor 111.

[0036] In an exemplary embodiment, the processor 111 may be coupled to the memory storage device 10 through the data transmission interface 114. For example, the processor 111 may store data to the memory storage device 10 or read data from the memory storage device 10 via the data transmission interface 114. In addition, the host system 11 may be coupled to the I / O device 12 through the system bus 110. For example, the host system 11 may transmit an output signal to the I / O device 12 or receive an input signal from the I / O device 12 via the system bus 110. In other embodiments, the processor 111 may also be electrically connected to the memory storage device 10 through a dedicated data transmission interface 114 instead of through the system bus 110.

[0037] In an exemplary embodiment, the processor 111, the random access memory 112, the read-only memory 113, and the data transmission interface 114 may be disposed on the motherboard 20 of the host system 11. The number of data transmission interfaces 114 may be one or more. Through the data transmission interface 114, the motherboard 20 may be coupled to the memory storage device 10 via wired or wireless means.

[0038] In an exemplary embodiment, the memory storage device 10 may be, for example, a USB flash drive 201, a memory card 202, or a solid state drive (SSD) 203. In some embodiments, the memory storage device 10 may be disposed outside the host system 11 as a wireless memory storage device 204. The wireless memory storage device 204 may be, for example, a near field communication (NFC) memory storage device, a wireless local area network (WiFi) memory storage device, a Bluetooth memory storage device, or a low energy Bluetooth memory storage device (e.g., iBeacon), etc., which are memory storage devices based on various wireless communication technologies. In addition, the motherboard 20 may also be coupled to various I / O devices such as a global positioning system (GPS) module 205, a network interface card 206, a wireless transmission device 207, a keyboard 208, a screen 209, a speaker 210, etc. via the system bus 110. For example, in an exemplary embodiment, the motherboard 20 may access the wireless memory storage device 204 via the wireless transmission device 207.

[0039] Figure 3 is a schematic diagram of a memory storage device shown in an exemplary embodiment of the present invention. Please refer to Figure 3 , the memory storage device 10 includes a connection interface unit 31, a memory control circuit unit 32, and a rewritable non-volatile memory module 33.

[0040] The connection interface unit 31 is used to couple to the processor 111. The memory storage device 10 can communicate with the processor 111 via the connection interface unit 31. In an exemplary embodiment, the connection interface unit 31 is compatible with the Peripheral Component Interconnect Express (PCI Express) standard. In an exemplary embodiment, the connection interface unit 31 can also be compliant with the Serial Advanced Technology Attachment (SATA) standard, the Parallel Advanced Technology Attachment (PATA) standard, the Institute of Electrical and Electronic Engineers (IEEE) 1394 standard, the Universal Serial Bus (USB) standard, the SD interface standard, the Ultra High Speed-I (UHS-I) interface standard, the Ultra High Speed-II (UHS-II) interface standard, the Memory Stick (MS) interface standard, the MCP interface standard, the MMC interface standard, the eMMC interface standard, the Universal Flash Storage (UFS) interface standard, the eMCP interface standard, the CF interface standard, the Integrated Device Electronics (IDE) standard, or other suitable standards. The connection interface unit 31 can be packaged in a chip with the memory control circuit unit 32, or the connection interface unit 31 is disposed outside a chip including the memory control circuit unit 32.

[0041] The memory control circuit unit 32 is coupled to the connection interface unit 31 and the rewritable non-volatile memory module 33. The memory control circuit unit 32 is used to execute multiple logic gates or control instructions implemented in hardware form or firmware form and perform operations such as writing, reading, and erasing data in the rewritable non-volatile memory module 33 according to the instructions of the processor 111.

[0042] The rewritable non-volatile memory module 33 is used to store data written by the processor 111. The rewritable non-volatile memory module 33 may include a single-level cell (SLC) NAND flash memory module (i.e., a flash memory module in which 1 bit can be stored in one memory cell), a multi-level cell (MLC) NAND flash memory module (i.e., a flash memory module in which 2 bits can be stored in one memory cell), a triple-level cell (TLC) NAND flash memory module (i.e., a flash memory module in which 3 bits can be stored in one memory cell), a quad-level cell (QLC) NAND flash memory module (i.e., a flash memory module in which 4 bits can be stored in one memory cell), other flash memory modules, or other memory modules with the same characteristics.

[0043] Each memory cell in the rewritable non-volatile memory module 33 stores one or more bits by a change in voltage (hereinafter also referred to as threshold voltage). Specifically, there is a charge trapping layer between the control gate and the channel of each memory cell. By applying a write voltage to the control gate, the amount of electrons in the charge trapping layer can be changed, thereby changing the threshold voltage of the memory cell. This operation of changing the threshold voltage of the memory cell is also referred to as "writing data into the memory cell" or "programming the memory cell". As the threshold voltage changes, each memory cell in the rewritable non-volatile memory module 33 has multiple storage states. By applying a read voltage, it can be determined which storage state a memory cell belongs to, thereby obtaining one or more bits stored in this memory cell.

[0044] In an exemplary embodiment, the memory cells of the rewritable non-volatile memory module 33 may form multiple physical programming units, and these physical programming units may form multiple physical erasure units. Specifically, the memory cells on the same word line may form one or more physical programming units. If each memory cell can store more than 2 bits, the physical programming units on the same word line can be at least classified into lower physical programming units and upper physical programming units. For example, the least significant bit (LSB) of a memory cell belongs to the lower physical programming unit, and the most significant bit (MSB) of a memory cell belongs to the upper physical programming unit. Generally, in an MLC NAND flash memory, the write speed of the lower physical programming unit is greater than that of the upper physical programming unit, and / or the reliability of the lower physical programming unit is higher than that of the upper physical programming unit.

[0045] In some embodiments, the processor 111 or the memory control circuit unit 32 may determine the programming mode for writing data to the rewritable non-volatile memory module 33. When using the single-level cell programming mode, data is written to the next physical programming unit. When using the multi-level (including second-level, third-level, fourth-level, etc.) cell programming mode, data is written to at least the next physical programming unit and the upper physical programming unit.

[0046] Figure 4 is a schematic diagram showing multiple channels according to an embodiment. Please refer to Figure 4 The rewritable non-volatile memory module 33 includes multiple channels 401 to 404, and each channel has multiple physical units. In an exemplary embodiment, a physical unit refers to a physical address or a physical programming unit. In an exemplary embodiment, a physical unit may also be composed of multiple consecutive or non-consecutive physical addresses. In an exemplary embodiment, a physical unit may also refer to a virtual block (VB). A virtual block may include multiple physical addresses or multiple physical programming units. In an exemplary embodiment, a virtual block may include one or more physical erasure units. Each of the channels 401 to 404 operates independently. For example, each channel belongs to an independent chip and is controlled by a specific chip enable (CE) signal. The memory control circuit unit 32 can write data to the channels 401 to 404 in parallel or read data from the channels 401 to 404 in parallel.

[0047] In this embodiment, the processor 111 executes the inference stage of a neural network, and some data is generated during the inference stage and cached in the rewritable non-volatile memory module 33. In particular, according to the operation, architecture, and data characteristics of the neural network, different data can be written to different channels 401 to 404 in parallel, thereby improving the read and write speeds. Regarding the following writing and reading of data, the processor 111 transmits a write instruction or a read instruction to the memory control circuit unit 32, and the memory control circuit unit 32 reads data from the rewritable non-volatile memory module 33 or writes data to the rewritable non-volatile memory module 33.

[0048] Figure 5 is a schematic diagram showing a neural network in the inference stage according to an embodiment. Please refer to Figure 5, the neural network 500 includes a first layer 501, a second layer 502, …, up to an Nth layer 503, where N is a positive integer. The input of the neural network 500 is a plurality of tokens 511 to 513. Each token can be text data, voice data, image data, or other types of data, and the present invention is not limited thereto. Each token is a vector generated through encoding or other processing. Here, the operation of the neural network 500 at different time points is unfolded. Specifically, the neural network 500 outputs a next token 514, and the next token 514 is fed back as the input of the neural network 500 to generate a next token 515. Similarly, the next token 515 is fed back as the output of the neural network 500 to generate a next token 516. Such a mechanism is also called autoregressive (AT).

[0049] Each of the first layer 501 to the Nth layer 503 is also called a decoder, and these decoders have the same architecture. The output of the first layer 501 can also be called a token, an embedding, or a feature vector. The output of the first layer 501 is transmitted to the second layer 502, and the output of the second layer 502 is transmitted to the next layer until it is transmitted to the Nth layer 503. In Figure 5 the architecture of the second layer 502 is enlarged as an example.

[0050] The operation of the second layer 502 includes steps 521 to 526. Generally speaking, in step 521, projection is performed to obtain a query matrix Q, a key matrix K, and a value matrix V. In step 522, an attention score is calculated based on the query matrix Q and the key matrix K. In step 523, normalization is performed on these attention scores, and this normalization is used to change the numerical range. For example, soft-max can be used, and in other embodiments, sparsemax can also be used, etc. Then in step 524, the normalized attention scores are multiplied by the value matrix V and summed to calculate an attention output. In step 525, this attention output can be projected again and added to the current token, and this is also called a residual layer. Finally, in step 526, the output of this layer is obtained through a fully connected layer. Among them, steps 522 to 524 are also called the calculation of self-attention.

[0051] Figure 6 is a schematic diagram showing the calculation of self-attention according to an embodiment. Please refer to Figure 5 and Figure 6 , assuming there are three tokens α 1 , α 2 , α 3 , where the token α 3 is also called the current token.

[0052] Multiply the token α 1 by a key weight matrix W K to obtain a key vector k 1 , multiply the token α 1 by a value weight matrix W V to obtain a value vector v 1 . Similarly, multiply the token α 2 by the key weight matrix W K to obtain a key vector k 2 , multiply the token α 2 by the value weight matrix W V to obtain a value vector v 2 . Multiply the token α 3 by the key weight matrix W K to obtain a key vector k 3 , multiply the token α 3 by the value weight matrix W V to obtain a value vector v 3 . In addition, multiply the token α 3 by a query weight matrix W Q to obtain a query vector q 3 .

[0053] Next, calculate the inner product of the query vector q 3 and the key vector k 1 to obtain an attention score, denoted here as β 1 . Similarly, calculate the inner product of the query vector q 3 and the key vector k 2 to obtain the attention score β 2 , calculate the inner product of the query vector q 3 and the key vector k 3 to obtain the attention score β 3 . Next, normalize the attention scores β 1 , β 2 , and β 3 . This normalization is, for example, soft-max, but the present invention is not limited thereto.

[0054] Then, multiply the normalized attention score β 1 by the value vector v 1 , multiply the normalized attention score β 2 by the value vector v 2 , multiply the normalized attention score β 3 by the value vector v 3 , and finally sum up these products to obtain an attention output, expressed as the following mathematical formula 1, where o 3 is a vector, called the attention output.

[0055] [Mathematical formula 1]

[0056]

[0057] The calculated attention output o 3 corresponds to the current token α 3 . For the token α 1 、α 2 , the corresponding attention output o 1 、o 2 can also be calculated. As Figure 7 shown, taking the token α 2 as an example, multiplying the token α 2 by the query weight matrix W Q can obtain the query vector q 2 , and the remaining operations are the same as Figure 6 . In particular, the query vector q 2 will not be multiplied by the key vector k 3 because the token α 2 occurs before the token α 3 , and there is no information about the token α 2 when calculating the attention output for the token α 3 . In other words, the query vector of each token will only be multiplied by the key vector of the tokens that occurred before.

[0058] Because for each token α 1 、α 2 、α 3 , the corresponding attention output will be calculated. Here, the tokens α 1 、α 2 、α 3 can be formed into a matrix X, and the above calculations are represented by the matrix. Figure 8 is a schematic diagram showing the matrix operations regarding attention according to an embodiment. Please refer to Figure 6 and Figure 8 . Each row of the matrix X represents a token. In this example, there are L tokens in total, and the length of each token (as a vector) is D, where L and D are positive integers. On the other hand, the sizes of the query weight matrix W Q , the key weight matrix W K , and the value weight matrix W V are all D×d, where d is a positive integer.

[0059] Each token α 1 、α 2 、α 3 will be multiplied by the query weight matrix W QMultiply. After being represented as matrices, these calculations are equivalent to multiplying matrix X and the query weight matrix W Q and obtaining a query matrix Q. Similarly, each token α 1 、α 2 、α 3 will be multiplied with the key weight matrix W K After being represented as matrices, these calculations are equivalent to multiplying matrix X and the query weight matrix W K and obtaining a key matrix K. Each token α 1 、α 2 、α 3 will be multiplied with the value weight matrix W V After being represented as matrices, these calculations are equivalent to multiplying matrix X and the query weight matrix W V and obtaining a value matrix V. Among them, the sizes of the query matrix Q, the key matrix K, and the value matrix V are all D×L, where each row corresponds to a token and each column corresponds to a feature. In another perspective, the query matrix Q contains L query vectors, such as the above-mentioned query vectors q 1 、q 2 、q 3 . Similarly, the key matrix K contains L key vectors, such as the above-mentioned key vectors k 1 、k 2 、k 3 . The value matrix V contains L value vectors, such as the above-mentioned value vectors v 1 、v 2 、v 3 .

[0060] The multiplication of the query vectors q 1 、q 2 、q 3 and the key vectors k 1 、k 2 、k 3 After being represented as matrices, it is equivalent to multiplying the query matrix Q and the transpose of the key matrix K, as shown in Figure 8 Matrix 800 in it, whose size is L×L, and each element is an attention score. However, there are multiple elements in matrix 800 as "×", which means that this element does not need to be calculated because the query vector of each token will only be multiplied with the key vectors of the previous tokens. Next, normalization is performed row by row, that is, normalization (such as soft-max) is performed on the elements in the same row of matrix 800, and thus matrix S can be obtained.

[0061] Next, the attention scores will be multiplied with the value vectors v 1 、v 2 、v 3Multiply them and then calculate the sum. After being represented as a matrix, such a calculation is equivalent to multiplying matrix S and query matrix V to obtain output matrix O h , which has a size of L×d. In other words, output matrix O h contains L attention outputs, such as containing the above-mentioned attention output o 1 、o 2 、o 3 , which respectively correspond to tokens α 1 、α2、α3.

[0062] When multiple tokens are input into the neural network, some calculations will be repeated. For example, when processing token α 1 , the key vector k 1 and value vector v 1 must be calculated. When processing α 2 , the key vector k 1 and value vector v 1 also need to be calculated. Therefore, after processing token α 1 , the key vector k 1 and value vector v 1 can be written into the rewritable non-volatile memory module 33. When processing token α 2 , the key vector k 1 and value vector v 1 can be read from the rewritable non-volatile memory module 33. Such an approach is called KV caching.

[0063] Figure 9 is a schematic diagram showing the operation when applying KV caching according to an embodiment. Please refer to Figure 9 , first obtain the current token 901, and this current token 901 is the Figure 8 last row of matrix X in Q . Next, multiply the current token by the query weight matrix W K to obtain the query vector 902. Then multiply the current token 901 by the key weight matrix W K to obtain the key vector 903, and this key vector 903 is the Figure 8 last row of the key matrix K in Figure 8 . Similarly, multiply the current token 901 by the value weight matrix W V to obtain the value vector 904, and this value vector 904 is the Figure 8 last row of the value matrix V in

[0064] Then, read the previous key matrix 911 from the rewritable non-volatile memory module 33. Combine the key vector 903 and the previous key matrix 911 to generate the current key matrix 912. Specifically, the key vector 903 can be transposed and then appended to the last column of the previous key matrix 911 to generate the current key matrix 912. The current key matrix 912 is the same as Figure 8The transpose of the middle key matrix K. Next, multiply the query vector 902 by the current key matrix 912 to obtain a temporary vector 920, where each element in the temporary vector 920 is an attention score. After normalizing (e.g., soft-max) all the elements of the temporary vector 920, a temporary vector 921 can be obtained.

[0065] On the other hand, read the previous value matrix 931 from the rewritable non-volatile memory module 33. Combine the value vector 904 with the previous value matrix 931 to generate the current value matrix 932. Specifically, append the value vector 904 to the last row of the previous value matrix 931 to generate the current value matrix 932. This current value matrix 932 is the same as Figure 8 the value matrix V in h .

[0066] After processing the current token, write the current key matrix 912 and the current value matrix 932 to the rewritable non-volatile memory module 33. When processing the next token, read the current key matrix 912 and the current value matrix 932 from the rewritable non-volatile memory module 33. It can be seen that the rewritable non-volatile memory module 33 will be frequently written to and read from. To improve the performance of the KV cache, multiple channels in the rewritable non-volatile memory module 33 can be used for parallel processing. Specifically, a part of the current key matrix 912 can be written to the physical units in one channel, and another part of the current key matrix 912 can be written to the physical units in another channel. In this way, multiple parts of the current key matrix 912 can be written to multiple channels in parallel, thereby increasing the writing speed. Similarly, a part of the current value matrix 932 can be written to one channel, and another part of the current value matrix 932 can be written to another channel, which can also increase the writing speed. In addition, when processing the next token, multiple parts of the current key matrix 912 can be read from multiple channels in parallel, and multiple parts of the current value matrix 932 can be read from multiple channels in parallel, which can increase the reading speed.

[0067] The current key matrix 912 has two dimensions, namely the token dimension (corresponding to L tokens) and the feature dimension (corresponding to d features). In some embodiments, the portion of the current key matrix 912 belonging to a first token can be written to one channel, and the portion belonging to another second token can be written to another channel, where the first token is different from the second token. In some embodiments, different features can also be written to different channels. For example, the portion of the current key matrix 912 belonging to a first feature can be written to one channel, and the portion belonging to another second feature can be written to another channel, where the first feature is different from the second feature. Similarly, the current value matrix 932 also has two dimensions, namely the token dimension and the feature dimension. In some embodiments, the portion of the current value matrix 932 belonging to a first token can be written to one channel, and the portion belonging to another second token can be written to another channel. In some embodiments, the portion of the current value matrix 932 belonging to a first feature can be written to one channel, and the portion belonging to another second feature can be written to another channel.

[0068] In some embodiments, multiple portions of the current key matrix 912 and multiple portions of the current value matrix 932 are written to the physical units in the rewritable non-volatile memory module 33 in a single-level cell programming mode. In other words, each memory cell in the written physical unit stores only one bit. Since the write and read speeds of the single-level cell programming mode are relatively fast and the number of erasable times is relatively large (compared with the multi-level cell programming mode), the single-level cell programming mode is more suitable for the frequent write / read operations of the above-mentioned KV cache.

[0069] Please refer to Figure 5 , the above operations are related to the KV cache of a certain layer in the neural network 500. When processing another layer in the neural network 500, since different layers must be processed sequentially (cannot be processed in parallel), the data of different layers can be written to the same channel. Figure 10 is a cache schematic diagram showing multiple layers according to an embodiment. Please refer to Figure 10, taking the current key matrices 912-1 and 912-2 as examples here, where the current key matrix 912-1 belongs to the first layer and the current key matrix 912-2 belongs to the second layer. Here, multiple parts corresponding to different tokens in the current key matrix are written into different channels. For example, the part of the current key matrix 912-1 belonging to the first token is written into channel 401, the part belonging to the second token is written into channel 402, the part belonging to the third token is written into channel 403, and the part belonging to the fourth token is written into channel 404. On the other hand, the part of the current key matrix 912-2 belonging to the first token is written into channel 401, the part belonging to the second token is written into channel 402, the part belonging to the third token is written into channel 403, and the part belonging to the fourth token is written into channel 404. When the number of channels is greater than or equal to the number of tokens, the part of the current key matrix 912-2 belonging to the fifth token can be written into the fifth channel. When the number of channels is less than the number of tokens, the part of the current key matrix 912-2 belonging to the fifth token can be written into the first channel. Figure 9 The current key matrix in

[0070] can be replaced with the current value matrix, and the rest of the operations are the same, which will not be elaborated here. Figure 6 and Figure 7 for calculation.

[0071] Figure 11 is a schematic diagram showing multiple heads according to an embodiment. Please refer to Figure 11 , for simplicity, taking two heads as an example here to illustrate. After the token α 1 is projected, it can generate the key vector k 1 and the value vector v 1 . The key vector k 1 can be further projected (multiplied by two weight matrices respectively) to obtain the key vector k 1,1 and the key vector k 1,2 . Similarly, the value vector v 1 can be further projected to obtain the value vector v 1,1 and the value vector v 1,2 . Among them, the key vector k 1,1 and the value vector v1,1 belongs to the first head, while the key vector k 1,2 and the value vector v 1,2 belong to the second head.

[0072] For the token α 2 , after generating the query vector q 2 , the key vector k 2 , and the value vector v 2 , the query vector q 2 can generate the query vector q 2,1 , q 2,2 through further projection. The key vector k 2 can generate the key vector k 2,1 and the key vector k 2,2 through further projection. The value vector v 2 can generate the value vector v 2,1 and the value vector v 2,2 . Among them, the query vector q 2,1 , the key vector k 2,1 , and the value vector v 2,1 belong to the first head, and the query vector q 2,2 , the key vector k 2,2 and the value vector v 2,2 belong to the second head.

[0073] The query vector q 2,1 will be multiplied by the key vector k 1,1 to obtain an attention score, and this attention score will be multiplied by the value vector v 1,1 . The query vector q 2,1 will also be multiplied by the key vector k 2,1 to obtain an attention score, and this attention score will be multiplied by the value vector v 2,1 . The sum of these two products will result in a vector as the first input for the projection in step 1110.

[0074] The query vector q 2,2 will be multiplied by the key vector k 1,2 to obtain an attention score, and this attention score will be multiplied by the value vector v 1,2 . The query vector q 2,2 will also be multiplied by the key vector k 2,2 to obtain an attention score, and this attention score will be multiplied by the value vector v 2,2 . The sum of these two products will result in a vector as the second input for the projection in step 1110.

[0075] In step 1110, the two input vectors will be concatenated and then multiplied by a matrix to obtain the attention output.

[0076] For each head, the KV cache is handled the same as for a single head. Specifically, the key vectors k 1,1 and the key vector k 2,1 will form a first key matrix and be stored in the rewritable non-volatile memory module 33. When processing the next token, this first key matrix is referred to as the first previous key matrix. On the other hand, the key vectors k 1,2 and the key vector k 2,2 will form a second key matrix and be stored in the rewritable non-volatile memory module 33. When processing the next token, this second key matrix is referred to as the second previous key matrix. In some embodiments, the key matrices belonging to different heads can be written in parallel to different channels, so the previous key matrices can be read in parallel from different channels. Subsequent calculations can refer to Figure 9 For the next token, the first previous key matrix will be combined with the first key vector belonging to the first head to generate a first current key matrix. In addition, the second previous key matrix and the second key vector belonging to the second head are combined to generate a second current key matrix. After processing the next token, the first current key matrix and the second current key matrix can be written to different channels.

[0077] Similarly, the value vectors v 1,1 and the value vector v 2,1 will form a first value matrix and be stored in the rewritable non-volatile memory module 33. When processing the next token, this first value matrix is referred to as the first previous value matrix. On the other hand, the value vectors v 1,2 and the value vector v 2,2 will form a second value matrix and be stored in the rewritable non-volatile memory module 33. When processing the next token, this second value matrix is referred to as the second previous value matrix. In some embodiments, the value matrices belonging to different heads can be written in parallel to different channels, so the previous value matrices can be read in parallel from different channels. Subsequent calculations can refer to Figure 9 For the next token, the first previous value matrix will be combined with the first value vector belonging to the first head to generate a first current value matrix. In addition, the second previous value matrix and the second value vector belonging to the second head are combined to generate a second current value matrix. After processing the next token, the first current value matrix and the second current value matrix can be written to different channels.

[0078] In view of the above-described multiple embodiments, after generating at least one current key matrix, a part of the current key matrix can be written to one channel, and another part can be written to another channel, and these parts can be different headers, tokens, or features. Similarly, after generating at least one current value matrix, a part of the current value matrix can be written to one channel, and another part can be written to another channel, and these parts can be different headers, tokens, or features. When processing different layers of a neural network, the tokens, features, or headers in different layers can be written to the same channel.

[0079] In some embodiments, the processor 111 can transmit the arrangement information of the current key matrix and the current value matrix to the memory control circuit unit 32. This arrangement information can be used to calculate which header, token, and feature the data in each logical address belongs to. For example, the arrangement information is used to indicate that the first header is transmitted first and then the second header, and the first token in the same header is transmitted first and then the second token. In this way, the memory control circuit unit 32 will first receive the data belonging to the first header, the first token, and the first feature, and then receive the data belonging to the first header, the first token, and the second feature, and so on. The memory control circuit unit 32 can write different parts of the current key matrix and the current value matrix to different channels according to this arrangement information.

[0080] Figure 12 is a flowchart showing a data caching method according to an embodiment. Please refer to Figure 12 , in step 1201, obtain the current token in the inference stage. In step 1202, read the previous key matrix and the previous value matrix from the rewritable non-volatile memory module. In step 1203, generate the current key matrix according to the current token and the previous key matrix. For example, first calculate the query vector, key vector, and value vector according to the current token, and then combine the key vector and the previous key matrix to generate the current key matrix. In step 1204, generate the current value matrix according to the current token and the previous value matrix. For example, combine the value vector and the previous value matrix to generate the current value matrix. In step 1205, multiply the query vector and the current key matrix to obtain a temporary vector. In step 1206, multiply the temporary vector and the current value matrix to obtain the attention output. In step 1207, write a part of the current key matrix to the physical unit in one channel, and store another part of the current key matrix to the physical unit in another channel. In step 1208, write a part of the current value matrix to the physical unit in one channel, and store another part of the current value matrix to the physical unit in another channel. Figure 12 The steps in have been described in detail above and will not be repeated here. It should be noted that Figure 12 The steps in can be implemented as multiple pieces of program code or circuits, and the present invention is not limited thereto. In addition,Figure 12 The method can be used in combination with the above embodiments or alone. In other words, Figure 12 Other steps can also be added between the steps of Figure 12 . In some embodiments, steps 1205 and 1206 can also be omitted. The present invention does not limit

[0081] the order of the steps in

[0082] . For example, step 1203 and step 1204 can be executed simultaneously, or steps 1205 and 1206 can be executed after steps 1207 and 1208. When performing KV caching, data writing and reading need to be performed frequently, and the above-disclosed technology can improve the writing and reading speeds. Although the present invention has been disclosed as above with embodiments, it is not intended to limit the present invention. Any person skilled in the art within the technical field, without departing from the spirit and scope of the present invention, can make some modifications and refinements. Therefore, the protection scope of the present invention shall be subject to what is defined by the claims.

Claims

1. A data caching method, characterized in that: For a rewritable non-volatile memory module, wherein the rewritable non-volatile memory module includes a plurality of channels, each of the plurality of channels includes a plurality of physical units, and the data caching method includes: Get the current token in the inference phase; reading a previous key matrix and a previous value matrix from the rewritable non-volatile memory module; generating a current key matrix according to the current token and the previous key matrix; generating a current value matrix according to the current token and the previous value matrix; Writing a portion of the current key matrix to the plurality of physical cells in one of the plurality of channels and storing another portion of the current key matrix to the plurality of physical cells in another one of the plurality of channels; and A portion of the current value matrix is ​​written to the plurality of physical units in one of the plurality of channels, and another portion of the current value matrix is ​​stored to the plurality of physical units in another one of the plurality of channels.

2. The data caching method according to claim 1, characterized in that: Wherein the portion of the current key matrix, the other portion of the current key matrix, the portion of the current value matrix, and the other portion of the current value matrix are written to the corresponding multiple physical units in a single-level storage unit programming mode.

3. The data caching method according to claim 1, characterized in that: The portion of the current key matrix belongs to a first symbol, the other portion of the current key matrix belongs to a second symbol, and the first symbol is different from the second symbol.

4. The data caching method according to claim 1, characterized in that: The part of the current key matrix belongs to a first feature, the other part of the current key matrix belongs to a second feature, and the first feature is different from the second feature.

5. The data caching method according to claim 1, characterized in that: Also includes: Calculate a query vector, a key vector, and a value vector according to the current token; multiplying the query vector and the current key matrix to obtain a temporary vector; as well as Multiply the temporary vector and the current value matrix to obtain the attention output, The step of generating a current key matrix according to the current token and the previous key matrix comprises: combining the key vector and the previous key matrix to produce the current key matrix, The step of generating the current value matrix according to the current token and the previous value matrix comprises: The value vector and the previous value matrix are combined to produce the current value matrix.

6. The data caching method according to claim 5, characterized in that: wherein the query vector includes a first query vector and a second query vector, the key vector includes a first key vector and a second key vector, the value vector includes a first value vector and a second value vector, the previous key matrix includes a first previous key matrix and a second previous key matrix, and the previous value matrix includes a first previous value matrix and a second previous value matrix, The step of combining the key vector and the previous key matrix to generate the current key matrix comprises: combining the first key vector and the first previous key matrix to generate a first current key matrix, combining the second key vector and the second previous key matrix to generate a second current key matrix, The step of combining the value vector and the previous value matrix to generate a current value matrix comprises: combining the first value vector and the first previous value matrix to produce a first current value matrix, combining the second value vector and the second previous value matrix to produce a second current value matrix, wherein the part of the current key matrix belongs to the first current key matrix, and the other part of the current key matrix belongs to the second current key matrix, Wherein the part of the current value matrix belongs to the first current value matrix, and the other part of the current value matrix belongs to the second current value matrix.

7. The data caching method according to claim 5, characterized in that: The step of calculating the query vector, the key vector, and the value vector according to the current token includes: Multiplying the current token by a query weight matrix to obtain the query vector; Multiplying the current token by a key weight matrix to obtain the key vector; as well as The current token is multiplied by a value weight matrix to obtain the value vector.

8. The data caching method according to claim 1, characterized in that: The current token is input to the first layer of the neural network, the neural network further comprises a second layer, and the data caching method further comprises: The portion of the current key matrix corresponding to the second layer and the portion of the current key matrix corresponding to the first layer are written to the plurality of physical cells in a same one of the plurality of channels.

9. The data caching method according to claim 1, characterized in that: Also includes: Get the next generation of symbols; reading the portion of the current key matrix and the another portion of the current key matrix in parallel from the plurality of channels; as well as The portion of the current value matrix and the other portion of the current value matrix are read in parallel from the plurality of channels.

10. A host system, characterized in that: Include: A memory storage device comprising a rewritable non-volatile memory module, wherein the rewritable non-volatile memory module comprises a plurality of channels, each of the plurality of channels comprising a plurality of physical cells; as well as A processor, electrically connected to the memory storage device, is configured to execute a plurality of steps: Get the current token during the inference phase; reading a previous key matrix and a previous value matrix from the rewritable non-volatile memory module; generating a current key matrix according to the current token and the previous key matrix; as well as generating a current value matrix according to the current token and the previous value matrix, The memory storage device is used to write a portion of the current key matrix to the multiple physical units in one of the multiple channels, store another portion of the current key matrix to the multiple physical units in another of the multiple channels, write a portion of the current value matrix to the multiple physical units in one of the multiple channels, and store another portion of the current value matrix to the multiple physical units in another of the multiple channels.

11. The host system according to claim 10, characterized in that: The memory storage device is used to write the portion of the current key matrix, the other portion of the current key matrix, the portion of the current value matrix, and the other portion of the current value matrix to the corresponding multiple physical units in a single-level storage unit programming mode.

12. The host system according to claim 10, characterized in that: The portion of the current key matrix belongs to a first symbol, the other portion of the current key matrix belongs to a second symbol, and the first symbol is different from the second symbol.

13. The host system according to claim 10, characterized in that: The part of the current key matrix belongs to a first feature, the other part of the current key matrix belongs to a second feature, and the first feature is different from the second feature.

14. The host system according to claim 10, characterized in that: The multiple steps also include: Calculate a query vector, a key vector, and a value vector according to the current token; multiplying the query vector and the current key matrix to obtain a temporary vector; and Multiply the temporary vector and the current value matrix to obtain the attention output, The step of generating a current key matrix according to the current token and the previous key matrix comprises: combining the key vector and the previous key matrix to produce the current key matrix, The step of generating the current value matrix according to the current token and the previous value matrix comprises: The value vector and the previous value matrix are combined to produce the current value matrix.

15. The host system according to claim 14, characterized in that: wherein the query vector includes a first query vector and a second query vector, the key vector includes a first key vector and a second key vector, the value vector includes a first value vector and a second value vector, the previous key matrix includes a first previous key matrix and a second previous key matrix, and the previous value matrix includes a first previous value matrix and a second previous value matrix, The step of combining the key vector and the previous key matrix to generate the current key matrix comprises: combining the first key vector and the first previous key matrix to generate a first current key matrix, combining the second key vector and the second previous key matrix to generate a second current key matrix, The step of combining the value vector and the previous value matrix to generate a current value matrix comprises: combining the first value vector and the first previous value matrix to produce a first current value matrix, combining the second value vector and the second previous value matrix to produce a second current value matrix, wherein the part of the current key matrix belongs to the first current key matrix, and the other part of the current key matrix belongs to the second current key matrix, Wherein the part of the current value matrix belongs to the first current value matrix, and the other part of the current value matrix belongs to the second current value matrix.

16. The host system according to claim 14, characterized in that: The step of calculating the query vector, the key vector, and the value vector according to the current token includes: Multiplying the current token by a query weight matrix to obtain the query vector; Multiplying the current token by a key weight matrix to obtain the key vector; as well as The current token is multiplied by a value weight matrix to obtain the value vector.

17. The host system according to claim 10, characterized in that: wherein the current token is input to a first layer of a neural network, the neural network further comprising a second layer, The memory storage device is used to write the portion of the current key matrix corresponding to the second layer and the portion of the current key matrix corresponding to the first layer to the multiple physical units in the same one of the multiple channels.

18. The host system according to claim 10, characterized in that: The host system is also used to obtain the next generation symbol, Wherein the memory storage device reads the portion of the current key matrix and the other portion of the current key matrix from the multiple channels in parallel, and reads the portion of the current value matrix and the other portion of the current value matrix from the multiple channels in parallel.