Data caching method and host system

By dividing the storage device into cached and shared blocks, and using the access speed differences between different storage devices to dynamically adjust the number of blocks, the problem of memory limitations of large language models is solved, and the inference efficiency and processing capabilities are improved.

CN120371730APending Publication Date: 2025-07-25XIAMEN JIAXIN ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510451644.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

As the tag lengths supported by large language models increase, the memory demand for key-value caches grows rapidly, resulting in memory limiting the tag lengths supported by the model, affecting inference efficiency.

Method used

The storage device is divided into cache blocks and shared blocks. The cache blocks are used to store some key-value data, and the shared blocks are used to store key-value data of other layers. The number of cached and shared blocks is dynamically adjusted to optimize storage and reading efficiency using the differences in access speeds of different storage devices.

Benefits of technology

By optimizing the storage solution, the mark length that the neural network can process is increased, the inference efficiency is improved, and the memory requirement is reduced, and the processing capability of the model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371730A_ABST
    Figure CN120371730A_ABST
Patent Text Reader

Abstract

The invention provides a data caching method and a host system. The method comprises the following steps: inputting a plurality of marks into a neural network, and obtaining key value data corresponding to a plurality of layers; dividing the first storage device into a cache block and a shared block; storing the first key value data corresponding to all layers to a cache block; storing the second key value data corresponding to the first layer to a shared block; storing second key value data corresponding to at least part of other layers except the first layer in the key value data to a second storage device; and reading the first key value data corresponding to the first layer from the cache block and reading the second key value data corresponding to the first layer from the shared block when an inference program about the first layer is performed. Therefore, the neural network can process more marks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a data caching method and a host system applied to key-value caching. Background Art

[0002] With the rapid development of artificial intelligence technology, especially in the field of natural language processing, large language models have been widely applied to various tasks, such as machine translation, dialogue generation, etc. These models usually adopt the self-attention mechanism to enhance the ability to process long texts. These models often need to perform key-value caching (KV cache) during the inference process to avoid repeated calculations and improve the inference efficiency. Key-value caching refers to temporarily storing the keys and values generated for past tokens in a memory for quick query during subsequent token generation to avoid repeated calculations. However, as the token length supported by the model increases, the memory requirement for key-value caching also grows rapidly, and when the memory is insufficient, it will limit the token length supported by the model. Summary of the Invention

[0003] To solve the above problems, the present disclosure proposes a data caching method and a host system.

[0004] The present disclosure proposes a data caching method applicable to a first storage device and a second storage device. The data caching method includes: inputting a plurality of tokens into a neural network and obtaining a plurality of key-value data corresponding to a plurality of layers; dividing the first storage device into a cache block and a shared block; respectively storing the first key-value data corresponding to all layers in the key-value data into the cache blocks corresponding to their respective layers; storing the second key-value data corresponding to the first layer in the key-value data into the shared block; storing the second key-value data corresponding to at least some other layers except the first layer in the key-value data into the second storage device; and during the inference process regarding the first layer, reading the first key-value data corresponding to the first layer from the cache block and reading the second key-value data corresponding to the first layer from the shared block.

[0005] In an embodiment of the present disclosure, the access speed of the above-mentioned first storage device is greater than the access speed of the second storage device, and the size of all key-value data is greater than the capacity of the first storage device.

[0006] In an embodiment of the present disclosure, the above data caching method further includes: after executing the inference program of the first layer, moving the second key-value data corresponding to the new tag from the shared block to the second storage device; reading the second key-value data corresponding to the second layer from the second storage device to the shared block; and when performing the inference program for the second layer, reading the first key-value data corresponding to the second layer from the cache block and reading the second key-value data corresponding to the second layer from the shared block.

[0007] In an embodiment of the present disclosure, the above data caching method further includes: obtaining the tag length supported by the neural network; and determining the number of cache blocks and the number of shared blocks according to the tag length.

[0008] In an embodiment of the present disclosure, the step of determining the number of cache blocks and the number of shared blocks according to the tag length includes: obtaining the total number of blocks of the first storage device; setting the sum of the number of cache blocks and the number of shared blocks to be the same as the total number of blocks; dividing the tag length by the block size to obtain the block length; and calculating the minimum value of the number of shared blocks under a condition that the value obtained by multiplying the number of shared blocks by the number of layers and adding the number of cache blocks is greater than or equal to the block length.

[0009] In an embodiment of the present disclosure, the above data caching method further includes: gradually increasing the tag length as the neural network continuously generates a plurality of output tags, where the block length is greater than the total number of blocks.

[0010] In an embodiment of the present disclosure, the above data caching method further includes: generating a plurality of output tags by the neural network, connecting the output tags to the tags to obtain a plurality of historical tags; obtaining a first prefix of the historical tags; calculating a first hash value of the first prefix; and storing the first hash value and the key-value data corresponding to the first prefix in the second storage device.

[0011] In an embodiment of the present disclosure, the above data caching method further includes: obtaining a second prefix of the historical tags, where the second prefix covers the first prefix; calculating a second hash value of the second prefix; and storing the second hash value and the key-value data corresponding to the second prefix in the second storage device.

[0012] In an embodiment of the present disclosure, the step of storing the first hash value and the key-value data corresponding to the first prefix into the second storage device includes: determining whether the first hash value hits the first previous hash value of the third storage device, where the access speed of the third storage device is greater than that of the second storage device, and the access speed of the third storage device is less than that of the first storage device; if the first hash value does not hit the first previous hash value, determining whether the third storage device is full; if the third storage device is not full, writing the first hash value and the key-value data corresponding to the first prefix into the third storage device; if the third storage device is full, determining whether the second storage device is full; if the second storage device is full, clearing the data of the second storage device; if the second storage device is not full, selecting a previous key-value data in the third storage device; determining whether a hash value of the previous key-value data hits the second previous hash value of the second storage device; and if the hash value of the previous key-value data does not hit the second previous hash value, moving the previous key-value data to the second storage device.

[0013] In an embodiment of the present disclosure, the data caching method further includes: when performing an inference program for one of the layers, the first key-value data of M layers is stored in the cache block, and the second key-value data of N layers is stored in the shared block. The cache block and the shared block are distributed on the same memory chip, M and N are positive integers, and M is greater than N.

[0014] In an embodiment of the present disclosure, obtaining the key-value data includes: obtaining the marked prefix; calculating the hash value of the prefix; determining whether the hash value hits a previous hash value of the second storage device; and if the hash value hits the previous hash value, reading the key-value data corresponding to the prefix from the second storage device.

[0015] In another aspect, an embodiment of the present invention provides a host system, including a processor and a second storage device, and the processor includes a first storage device. The processor is electrically connected to the second storage device and is configured to execute the above data caching method.

[0016] To make the above features and advantages of the present invention more obvious and understandable, specific embodiments are hereinafter given and described in detail in conjunction with the accompanying drawings as follows. Description of the Drawings

[0017] Figure 1 is a schematic diagram of a host system and an input / output (I / O) device drawn according to an exemplary embodiment of the present invention;

[0018] Figure 2 is a schematic diagram of a host system, a storage device, and an I / O device drawn according to an exemplary embodiment of the present invention;

[0019] Figure 3A schematic diagram of a storage device drawn according to an exemplary embodiment of the present invention;

[0020] Figure 4 A schematic diagram of a memory hierarchy drawn according to an embodiment;

[0021] Figure 5 A schematic diagram of a neural network in the inference stage drawn according to an embodiment;

[0022] Figure 6 A schematic diagram of dividing a storage device into multiple blocks according to an embodiment;

[0023] Figure 7 A schematic diagram of data when performing the inference program of the first layer drawn according to an embodiment;

[0024] Figure 8 A schematic diagram of data when performing the inference program of the second layer drawn according to an embodiment;

[0025] Figure 9 A schematic diagram of merging key-value data of cache blocks and shared blocks drawn according to an embodiment;

[0026] Figure 10 A schematic diagram of a conversation record drawn according to an embodiment;

[0027] Figure 11 A schematic diagram of another conversation record drawn according to an embodiment;

[0028] Figure 12 A flowchart of reading key-value data drawn according to an embodiment;

[0029] Figure 13 A flowchart of writing key-value data drawn according to an embodiment;

[0030] Figure 14 A flowchart of a data caching method drawn according to an embodiment. Detailed implementation manners

[0031] Some embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings. For the reference numerals of the components cited in the following description, when the same reference numerals appear in different drawings, they will be regarded as the same or similar components. These embodiments are only a part of the present invention and do not disclose all the implementable ways of the present invention. More precisely, these embodiments are only examples of the systems and methods in the claims of the present invention.

[0032] Regarding the use of "first", "second", etc. in this article, it does not specifically refer to the meaning of order or sequence, but is only used to distinguish components or operations described with the same technical terms.

[0033] Figure 1A schematic diagram of a host system and an input / output (I / O) device drawn according to an exemplary embodiment of the present invention. Figure 2 A schematic diagram of a host system, a storage device, and an I / O device drawn according to an exemplary embodiment of the present invention.

[0034] Please refer to Figure 1 and Figure 2 , the host system 11 is a computer system, which can be a desktop computer, a server, a distributed system, a laptop computer, etc., and the present invention is not limited thereto. The host system 11 may include a processor 111, a storage device 112, a read only memory (ROM) 113, and a data transmission interface 114. The storage device 112 is, for example, a random access memory (RAM).

[0035] The processor 111, the storage device 112, the read only memory 113, and the data transmission interface 114 may be coupled to a system bus 110. In some embodiments, the storage device 112 may also be connected to the processor 111 through a dedicated transmission interface instead of through the bus 110. The processor 111 may be a graphic processing unit (GPU), a tensor processing unit (TPU), a neural processing unit (NPR), etc. The processor 111 has a storage device 120, which is, for example, a Video RAM (VRAM).

[0036] In an exemplary embodiment, the processor 111 may be coupled to the storage device 10 through the data transmission interface 114. For example, the processor 111 may store data to or read data from the storage device 10 via the data transmission interface 114. In addition, the host system 11 may be coupled to the I / O device 12 through the system bus 110. For example, the host system 11 may transmit an output signal to the I / O device 12 or receive an input signal from the I / O device 12 via the system bus 110. In other embodiments, the processor 111 may also be electrically connected to the storage device 10 through a dedicated data transmission interface 114 instead of through the system bus 110.

[0037] In an exemplary embodiment, the processor 111, the storage device 112, the read only memory 113, and the data transmission interface 114 may be disposed on the motherboard 20 of the host system 11. The number of the data transmission interfaces 114 may be one or more. Through the data transmission interface 114, the motherboard 20 may be coupled to the storage device 10 in a wired or wireless manner.

[0038] In an exemplary embodiment, the storage device 10 may be, for example, a USB flash drive 201, a memory card 202, or a solid state drive (SSD) 203. In some embodiments, the storage device 10 may be disposed outside the host system 11 as a wireless storage device 204. The wireless storage device 204 may be, for example, a near field communication (NFC) storage device, a wireless fidelity (WiFi) storage device, a Bluetooth storage device, or a low energy Bluetooth storage device (e.g., iBeacon), etc., which are storage devices based on various wireless communication technologies. In addition, the motherboard 20 may also be coupled to various I / O devices such as a global positioning system (GPS) module 205, a network interface card 206, a wireless transmission device 207, a keyboard 208, a screen 209, a speaker 210, etc. through the system bus 110. For example, in an exemplary embodiment, the motherboard 20 may access the wireless storage device 204 through the wireless transmission device 207.

[0039] Figure 3 is a schematic diagram of a storage device shown according to an exemplary embodiment of the present invention. Please refer to Figure 3 , the storage device 10 includes a connection interface unit 31, a memory control circuit unit 32, and a rewritable non-volatile memory module 33.

[0040] The connection interface unit 31 is used to couple to the processor 111. The storage device 10 can communicate with the processor 111 via the connection interface unit 31. In an exemplary embodiment, the connection interface unit 31 is compatible with the Peripheral Component Interconnect Express (PCI Express) standard. In an exemplary embodiment, the connection interface unit 31 can also be compliant with the Serial Advanced Technology Attachment (SATA) standard, the Parallel Advanced Technology Attachment (PATA) standard, the Institute of Electrical and Electronic Engineers (IEEE) 1394 standard, the Universal Serial Bus (USB) standard, the SD interface standard, the Ultra High Speed-I (UHS-I) interface standard, the Ultra High Speed-II (UHS-II) interface standard, the Memory Stick (MS) interface standard, the MCP interface standard, the MMC interface standard, the eMMC interface standard, the Universal Flash Storage (UFS) interface standard, the eMCP interface standard, the CF interface standard, the Integrated Device Electronics (IDE) standard, or other suitable standards. The connection interface unit 31 can be encapsulated in a single chip with the memory control circuit unit 32, or the connection interface unit 31 is disposed outside a chip containing the memory control circuit unit 32.

[0041] The memory control circuit unit 32 is coupled to the connection interface unit 31 and the rewritable non-volatile memory module 33. The memory control circuit unit 32 is used to execute multiple logic gates or control instructions implemented in hardware form or firmware form and perform operations such as data writing, reading, and erasing in the rewritable non-volatile memory module 33 according to the instructions of the processor 111.

[0042] The rewritable non-volatile memory module 33 is used to store data written by the processor 111. The rewritable non-volatile memory module 33 may include a single-level cell (SLC) NAND flash memory module (i.e., a flash memory module in which 1 bit can be stored in one memory cell), a multi-level cell (MLC) NAND flash memory module (i.e., a flash memory module in which 2 bits can be stored in one memory cell), a triple-level cell (TLC) NAND flash memory module (i.e., a flash memory module in which 3 bits can be stored in one memory cell), a quad-level cell (QLC) NAND flash memory module (i.e., a flash memory module in which 4 bits can be stored in one memory cell), other flash memory modules, or other memory modules with the same characteristics.

[0043] Each memory cell in the rewritable non-volatile memory module 33 stores one or more bits by a change in voltage (hereinafter also referred to as the threshold voltage). Specifically, there is a charge trapping layer between the control gate and the channel of each memory cell. By applying a write voltage to the control gate, the amount of electrons in the charge trapping layer can be changed, thereby changing the threshold voltage of the memory cell. This operation of changing the threshold voltage of the memory cell is also referred to as "writing data to the memory cell" or "programming the memory cell". As the threshold voltage changes, each memory cell in the rewritable non-volatile memory module 33 has multiple storage states. By applying a read voltage, it can be determined which storage state a memory cell belongs to, and thus one or more bits stored in this memory cell can be obtained.

[0044] In an exemplary embodiment, the memory cells of the rewritable non-volatile memory module 33 can form a plurality of physical programming units, and these physical programming units can form a plurality of physical erasure units. Specifically, the memory cells on the same word line can form one or more physical programming units. If each memory cell can store more than 2 bits, the physical programming units on the same word line can be at least classified into lower physical programming units and upper physical programming units. For example, the least significant bit (LSB) of a memory cell belongs to the lower physical programming unit, and the most significant bit (MSB) of a memory cell belongs to the upper physical programming unit. Generally speaking, in MLC NAND flash memory, the write speed of the lower physical programming unit is greater than that of the upper physical programming unit, and / or the reliability of the lower physical programming unit is higher than that of the upper physical programming unit.

[0045] In some embodiments, the processor 111 or the memory control circuit unit 32 can determine the programming mode for writing data to the rewritable non-volatile memory module 33. When using the single-level memory cell programming mode, the data is written to the lower physical programming unit. When using the multi-level (including second-level, third-level, fourth-level, etc.) memory cell programming mode, the data is written to at least the lower physical programming unit and the upper physical programming unit.

[0046] Figure 4 is a schematic diagram showing a memory hierarchy according to an embodiment. Please refer to Figure 4 , in this embodiment, the access speeds of the storage devices 120, 112, and 10 are different. The access speed of the storage device 120 is greater than that of the storage device 112, and the access speed of the storage device 112 is greater than that of the storage device 10. In some embodiments, the capacity of the storage device 120 is less than that of the storage device 112, and the capacity of the storage device 112 is less than that of the storage device 10, but the present invention is not limited thereto. A data caching method will be proposed here. The storage device 120 is used to store the data required by a neural network during the inference stage. However, since the capacity of the storage device 120 may not be sufficient, some data will be stored in the storage device 112 or the storage device 10. In this embodiment, a total of three storage devices 120, 112, and 10 are used, but more or fewer storage devices can also be used in other embodiments. For example, in some embodiments, the storage device 112 can be omitted, and the data that cannot be stored in the storage device 120 will be stored in the storage device 10.

[0047] Figure 5 is a schematic diagram showing a neural network during the inference stage according to an embodiment. Please refer to Figure 5, the neural network 500 includes a first layer 501, a second layer 502, …, up to an Nth layer 503, where N is a positive integer. The input to the neural network 500 is a plurality of tokens 511-513. Each token can be text data, speech data, image data, or other types of data, and the present invention is not limited thereto. Each token is a vector generated through encoding or other processing. Here, the operation of the neural network 500 at different time points is unfolded. Specifically, the neural network 500 outputs a next token 514, and the next token 514 is fed back as the input to the neural network 500 and generates a next token 515. Similarly, the next token 515 is fed back as the output of the neural network 500 and generates a next token 516. Such a mechanism is also called autoregressive (AT).

[0048] Each of the first layer 501 to the Nth layer 503 is also called a decoder, and these decoders have the same architecture. The output of the first layer 501 can also be called a token, an embedding, or a feature vector. The output of the first layer 501 is transmitted to the second layer 502, and the output of the second layer 502 is transmitted to the next layer until it is transmitted to the Nth layer 503. In Figure 5 the architecture of the second layer 502 is enlarged as an example.

[0049] The operation of the second layer 502 includes steps 521-526. Generally speaking, in step 521, projection is performed to obtain a query matrix Q, a key matrix K, and a value matrix V. In step 522, the query matrix Q and the key matrix K are multiplied to calculate a plurality of attention scores. In step 523, a transformation is performed on these attention scores, and this transformation is, for example, soft-max. In other embodiments, sparsemax, etc. can also be used. Then in step 524, the transformed attention scores are multiplied by the value matrix V and summed to calculate an attention output. In step 525, this attention output can be projected again and added to the input token, which is also called a residual layer. Finally, in step 526, the output of this layer is obtained through a fully connected layer. Among them, steps 522-524 are also called the calculation of self-attention.

[0050] Here, the key matrix K and the value matrix V are collectively called key-value data. The inference stage can generally be divided into two stages: prefill and decode. During prefill, the key-value data of all tokens 511-513 are calculated, and these key-value data are stored in the storage device. During decode, the key-value data are read from the storage device to avoid repeated calculation.

[0051] Figure 6It is a schematic diagram of dividing a storage device into multiple blocks according to an embodiment. Please refer to Figure 6 , where the memory space of the storage device 120 is divided into multiple blocks B1 to B10, and each block is a memory space in the vertical direction. Additionally, assume that the neural network 500 has a total of 6 layers L1 to L6, and these layers correspond to the memory space in the horizontal direction. After inputting multiple tokens into the neural network 500, the key-value data corresponding to these layers and tokens will be calculated first. The size of the key-value data corresponding to each token is 2 * 6 * D, where 2 represents keys and values, 6 represents the number of layers L1 to L6, and D represents the vector length. Figure 6 Each space in Figure 6 can store the key-value data corresponding to one or more tokens. For simplicity, it is assumed here that each space can store the key-value data corresponding to one token. Therefore, the entire storage device 120 can store the key-value data corresponding to 10 tokens and 6 layers. For example,

[0052] the first space in block B1 can store the key-value data of the first token in the first layer L1, the second space in block B1 can store the key-value data of the first token in the second layer L2, the first space in block B2 can store the key-value data of the second token in the first layer L1, and so on. Figure 6In the medium, the second key-value data KV9 to KV17 of the first layer L1 are stored in the shared block 620, where the second key-value data KV9 to KV17 respectively correspond to the 9th to 17th markers. On the other hand, the second key-value data of at least some other layers L2 to L6 except the first layer L1 are stored in another storage device (for example, storage device 112 or storage device 10). For example, here the second key-value data of layers L2 to L6 can be stored in storage device 112; or a part of the second key-value data of layers L2 to L6 is stored in storage device 112, and another part of the second key-value data of layers L2 to L6 is stored in storage device 10; or when there is still idle space in the shared block 620, a part of the second key-value data of layers L2 to L6 can also be stored in the shared block 620, and the remaining second key-value data is stored in storage device 112.

[0053] Figure 7 It is a data schematic diagram showing the inference procedure of the first layer according to an embodiment. Please refer to Figure 6 and Figure 7 , when performing the inference procedure of the first layer L1, the first key-value data KV1 to KV8 corresponding to the first layer L1 can be read from the cache block 610, and the second key-value data KV9 to KV17 corresponding to the first layer L1 can be read from the shared block 620. Here, the second key-value data KV9 to KV17 are shown at the position of the first layer L1, but the actual storage position is as Figure 6 shown. During the inference procedure, a new marker (i.e., the 18th marker) is also generated. For example, Figure 5 when the next marker 514 of is input to the neural network, the corresponding key-value data KV18 is also generated, and this key-value data KV18 is also stored in the shared block 620. After the inference procedure of the first layer L1 is completed, the second key-value data KV9 to KV18 in the shared block 620 are moved to the storage device 112.

[0054] Figure 8 It is a data schematic diagram showing the inference procedure of the second layer according to an embodiment. When performing the inference procedure of the second layer L2, the second key-value data KV9 to KV17 corresponding to the second layer L2 are read from the storage device 112 to the shared block 620. When performing the inference procedure of the second layer L2, the first key-value data KV1 to KV8 corresponding to the second layer L2 can be read from the cache block 610, and the second key-value data KV9 to KV17 corresponding to the second layer L2 can be read from the shared block 620. Similarly, the key-value data KV18 of a new marker is generated during the inference procedure of the second layer L2. After the inference procedure of the second layer L2 is completed, the key-value data KV18 corresponding to the second layer L2 is written into the storage device 112.

[0055] In other words, when performing the inference procedure for a certain layer, the first key-value data is read from the cache block 610, and the second key-value data is read from the shared block 620. After the inference procedure is completed, the key-value data (including the newly marked key-value data) in the shared block 620 is written to the storage device 112. The storage device 120 could originally only handle a tag length of 10, but in this way, the tag length that can be processed can be increased to 20.

[0056] From another perspective, when performing the inference procedure for a certain layer, the cache block 610 stores the first key-value data of the Mth layer, and the shared block 620 stores the second key-value data of the Nth layer. M and N are positive integers, and M is greater than N. In this example, M = 6 and N = 1, but in other embodiments, it can also be M = 6 and N = 2. The present invention does not limit the values of the positive integers M and N. On the other hand, the cache block 610 and the shared block 620 are distributed on the same memory chip.

[0057] In some embodiments, the number of shared blocks 620 can be determined dynamically. Specifically, the tag length supported by the neural network can be obtained first, and then the number of cache blocks 610 and the number of shared blocks 620 can be determined based on this tag length. When the tag length is larger, more shared blocks 620 are needed to store the tags of the same layer; when the tag length is smaller, fewer shared blocks 620 can be designed to reduce the read and write operations on the storage device 112.

[0058] Assume the number of cache blocks 610 is X, and the number of shared blocks 620 is Y. The total number of blocks of the storage device 120 is represented as Z. In some embodiments, X and Y can be determined according to the following mathematical formula.

[0059] [Mathematical formula 1]

[0060] X + Y = Z

[0061] [Mathematical formula 2]

[0062]

[0063] where N L is the number of all layers in the neural network. L M is the tag length supported by the neural network. S B is the block size, indicating how many tags each block can store. In this example, S B = 1. L M / S Bis the block length, representing how many blocks the neural network can support. In Mathematical Formula 1, it is set that the sum of the number X of cache blocks and the number Y of shared blocks is the same as the total number Z of blocks. In Mathematical Formula 2, it is to calculate the minimum value of Y under a condition that the value obtained by multiplying the number Y of shared blocks by the number N of layers L plus the number X of cache blocks must be greater than or equal to the block length L M / S B . Taking the Figure 8 example, X = 8, Y = 2, and Z = 10 can be substituted into Mathematical Formula 1 and Mathematical Formula 2. Mathematical Formula 2 can be rewritten as 8+(2×6)≥20, where Y = 2 is the minimum value that meets the condition. Such a value can support a longer token length while avoiding too many reads and writes to the storage device 112

[0064] In some embodiments, the token length L can also be increased progressively M . The token length supported by a language model may be several thousand or tens of thousands. If the token length L M is set to such a large value, too many shared blocks will be generated. Therefore, as the neural network continuously generates output tokens, the token length can be gradually increased when the shared blocks are insufficient. For example, the token length L M can be first set to 100, increased to 200 next time, and gradually increased to the supported upper limit. At this time, the calculated block length L M / S B will be greater than the total number Z of blocks, but due to the shared blocks, the storage device 120 can still store so many tokens

[0065] Please refer to Figure 6 , in this example, the key-value data KV9~KV17 are stored interleaved in block B9 and block B10. When writing the key-value data in the shared block 620 to the storage device 112 or the storage device 10, the key-value data in the same block will be stored in the same file. Therefore, the first file stores the key-value data KV9, KV11, KV13, KV15, KV17, and the second file stores the key-value data KV10, KV12, KV14, KV16. Each layer corresponds to two files, so a total of 2*6 = 12 files will be generated. However, in other embodiments, the key-value data of the same layer can also be stored in the same file, or each piece of key-value data can be stored in a separate file, and the present invention is not limited thereto. On the other hand, the key-value data KV9~KV17 can be first stored in block B9 in sequence and then stored in block B10, and the present invention does not limit the arrangement of the key-value data KV9~KV17

[0066] Figure 9It is a schematic diagram showing the merging of key-value data of cache blocks and shared blocks according to an embodiment. Please refer to Figure 9 , in this example, the storage device 120 has a total of 8 blocks B1 to B8, where blocks B1 to B7 are cache blocks, and block B8 is a shared block. In this example, a total of 6 files (such as files 901 to 903) are generated to store the second key-value data corresponding to 6 layers L1 to L6 respectively. In particular, in this embodiment, the key-value data in the cache block can also be stored in the storage device 112 or the storage device 10. In such an example, the key-value data of the same layer (including the first key-value data and the second key-value data) all belong to the same file. Since the first key-value data of the same layer is arranged horizontally in the storage device 120, but the second key-value data is arranged vertically, format conversion is required. In other words, it is necessary to first read the file 901 corresponding to the first layer L1 from the storage device 112 or the storage device 10, convert the second key-value data in this file 901 into a horizontal arrangement and continue it after the first key-value data of the first layer L1 to form the file 911, and then write the file 911 into the storage device 112 or the storage device 10. Similarly, for the second layer L2, first read the file 902, convert the second key-value data in this file 902 into a horizontal arrangement and continue it after the first key-value data of the second layer L2 to form the file 912, and then write the file 912 into the storage device 112 or the storage device 10. Therefore, a total of 6 files (including files 911 to 913) are generated to store all the key-value data corresponding to all layers L1 to L6.

[0067] After storing both the first key-value data and the second key-value in the storage device 112 or the storage device 10, if the same content input appears subsequently, the corresponding key-value data can be directly read from the storage device 112 or the storage device 10 to the storage device 120, thereby reducing the calculation time. In some embodiments, the hash values of the corresponding tags can also be calculated, and these hash values are also stored in the storage device 112 or the storage device 10, and the corresponding key-value data is found by comparing the hash values subsequently.

[0068] Figure 10It is a schematic diagram showing a conversation record according to an embodiment. In the prior art, only the key-value data of the current conversation is temporarily stored, but the key-value data of all conversation records is used for inference, so the inference speed of the model will become slower and slower. In this example, a long-term key-value cache is to be established. For example, user 1020 and chatbot 1030 are having a conversation. The first sentence 1021 input by user 1020 is "What is the longest river in the world?", and this sentence 1021 will be converted into tokens 1001~1004, and the corresponding key-value data will also be calculated. The chatbot 1030 will use the key-value data of sentence 1021 during inference. The chatbot 1030 will generate multiple output tokens 1011~1014, and these output tokens 1011~1014 will be converted into sentence 1031, the content of which is "The longest river in the world is the Nile River". These output tokens 1011~1014 will be connected after tokens 1001~1004 and are collectively called multiple historical tokens. Next, multiple prefixes of these historical tokens will be obtained, the hash value of the prefix will be calculated, and finally the hash value and the corresponding key-value data will be stored in the storage device 112. For example, token 1001 forms the first prefix, so the hash value 1041 of token 1001 and the corresponding key-value data will be written into the storage device 112. Tokens 1001 and 1002 form the second prefix, so the hash value 1042 of tokens 1001 and 1002 and the corresponding key-value data will be written into the storage device 112. Tokens 1001~1003 form the third prefix, so the hash value 1043 of tokens 1001~1003 and the corresponding key-value data will be written into the storage device 112. The second prefix covers the first prefix, and the third prefix covers the second prefix. From another perspective, the first prefix is a subset of the second prefix, and the second prefix is a subset of the third prefix. The above process is repeated continuously until the hash values of all tokens 1001~1004 and 1011~1014 are calculated, that is, in Figure 10 this example, there are 10 prefixes in total, and the hash value of each prefix and the corresponding key-value data will be written into the storage device 112.

[0069] Next, user 1020 inputs sentence 1022. Key-value data of sentences 1021 and 1031 are required for inference. At this time, the hash value after combining sentences 1021 and 1031 can be calculated, and this hash value will hit (hit) the hash value of storage device 112 (referred to as the previous hash value), so the key-value data corresponding to sentences 1021 and 1031 can be read from storage device 112. In this way, when user 1020 inputs sentence 1022, there is no need to repeatedly calculate the key-value data of sentences 1021 and 1031. This approach will save more and more time as the conversation progresses because the conversation record gets longer and more key-value data needs to be calculated.

[0070] In this embodiment, the hash value of a prefix is calculated, so a sentence does not need to be exactly the same as the previous sentence to find the matching key-value data. For example, Figure 11 is a schematic diagram showing another conversation record according to an embodiment. Please refer to Figure 10 and Figure 11 User 1120 is different from user 1020, or they are the same but Figure 10 and Figure 11 belong to different dialog boxes. User 1120 inputs sentence 1121, and this sentence 1121 has a common prefix with sentence 1021, that is, "What is the longest". Sentence 1121 will be converted into multiple tokens 1101 to 1104. Next, multiple prefixes are obtained from these tokens 1101 to 1104, the hash values of these prefixes are calculated, and it is determined whether these hash values hit the previous hash value in storage device 112. For example, token 1101 forms the first prefix; tokens 1101 and 1102 form the second prefix; tokens 1101 to 1103 form the third prefix; tokens 1101 to 1104 form the fourth prefix. In this example, the hash value of tokens 1101 and 1102 (corresponding to "What is the longest") will hit the previous hash value of storage device 112, so the key-value data 1150 corresponding to the second prefix can be read from storage device 112. In this way, there is no need to repeatedly calculate the key-value data of tokens 1101 and 1102.

[0071] In some embodiments, Figure 10 and Figure 11The above approach can be applied to Retrieval Augmented Generation (RAG). On the one hand, in some applications, users may frequently input the same questions. For example, in financial customer service, questions like "What should I do if my credit card is lost?" are often asked, or the same symptoms are inquired about in medical customer service. On the other hand, if chatbot 1030 retrieves a fixed database, the same reference data will be obtained for the same question, and the key-value data of these reference data will also be stored in storage device 112. Therefore, calculations can be reduced. In other words, the above historical tags can also include tags of the text obtained after retrieving the database.

[0072] Figure 12 is a flowchart showing the reading of key-value data according to an embodiment. Please refer to Figure 4 and Figure 12 , in this example, there are a total of three storage devices 120, 112, and 10. When reading or writing key-value data, it is sequentially determined whether the data has been stored in these three storage devices 120, 112, and 10. For a certain prefix, the hash value is first calculated. In step 1201, it is determined whether this hash value hits the previous hash value stored in storage device 120. If so, in step 1202, the corresponding key-value data is read from storage device 120. If the result of step 1201 is no, in step 1203, it is determined whether the hash value hits the previous hash value stored in storage device 112. If the result of step 1203 is yes, step 1204 is executed to read the corresponding key-value data from storage device 112. If the result of step 1203 is no, step 1205 is executed to determine whether the hash value hits the previous hash value stored in storage device 10. If the result of step 1205 is yes, step 1206 is executed to read the corresponding key-value data from storage device 10. If the result of step 1205 is no, step 1207 is executed to calculate the key-value data. After calculating the key-value data of all tags, in step 1208, a decoding program is performed by the neural network.

[0073] After the neural network performs the decoding program, multiple historical tags are obtained (refer to Figure 10 ), and the hash value of a prefix is calculated. Next, please refer to Figure 13 , Figure 13is a flowchart for writing key-value data according to an embodiment. In step 1301, it is determined whether this hash value hits a previous hash value in the storage device 112. If so, it means that the corresponding key-value data has already been stored in the storage device 112, so this process can be ended. If the result of step 1301 is no, it is determined in step 1302 whether the storage device 112 is full. If the storage device 112 is not full, the hash value and the corresponding key-value data are written to the storage device 112 in step 1303. If the storage device 112 is full, it is determined in step 1304 whether the storage device 10 is full. If the result of step 1304 is yes, a part of the data in the storage device 10 is cleared, and the process returns to step 1304. If the storage device 10 is not full, the previous key-value data in the storage device 112 is selected in step 1306. For example, the key-value data that has not been used for the longest time can be selected. In step 1307, it is determined whether the hash value of the previous key-value data hits a previous hash value in the storage device 10. If the result of step 1307 is yes, it means that the corresponding key-value data has already been stored in the storage device 10, so the process can be ended. If the result of step 1307 is no, the previous key-value data is moved to the storage device 10 in step 1308, and then the process returns to step 1302.

[0074] Figure 14 is a flowchart showing a data caching method according to an embodiment. Please refer to Figure 14 , in step 1401, a plurality of tokens are input into a neural network, and a plurality of key-value data corresponding to a plurality of layers are obtained. In some embodiments, the key-value data of all layers can be directly calculated; in other embodiments, part or all of the key-value data can also be obtained from a storage device by calculating the hash value of a prefix. In step 1402, the first storage device is divided into a cache block and a shared block. In step 1403, the first key-value data corresponding to all layers in the key-value data are respectively stored in the cache blocks corresponding to these layers. In step 1404, the second key-value data corresponding to the first layer in the key-value data are stored in the shared block. In step 1405, the second key-value data corresponding to at least some other layers except the first layer in the key-value data are stored in a second storage device. In step 1406, when performing an inference procedure regarding the first layer, the first key-value data corresponding to the first layer are read from the cache block, and the second key-value data corresponding to the first layer are read from the shared block.

[0075] In the above data caching method and host system, the problem that more tokens cannot be processed due to insufficient capacity of the storage device can be solved. In addition, the above means also propose long-term key-value caching, thereby reducing repeated calculations.

[0076] Although the present invention has been disclosed above by way of examples, it is not intended to limit the present invention. Any person skilled in the relevant technical field may make some modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be subject to what is defined by the claims.

Claims

1. A data caching method, characterized in that, Applicable to a first storage device and a second storage device, the data caching method includes: Inputting a plurality of tokens into a neural network and obtaining a plurality of key-value data corresponding to a plurality of layers; Dividing the first storage device into a plurality of cache blocks and a shared block; Respectively storing first key-value data corresponding to the plurality of layers among the plurality of key-value data into the cache blocks corresponding to the respective layers among the plurality of cache blocks; Storing second key-value data corresponding to the first layer among the plurality of layers in the plurality of key-value data into the shared block; Storing second key-value data corresponding to at least some other layers among the plurality of layers except the first layer in the plurality of key-value data into the second storage device; And When performing an inference procedure regarding the first layer, reading the first key-value data corresponding to the first layer from the plurality of cache blocks and reading the second key-value data corresponding to the first layer from the shared block.

2. The data caching method according to claim 1, wherein Wherein the access speed of the first storage device is greater than the access speed of the second storage device, and the size of the plurality of key-value data is greater than the capacity of the first storage device.

3. The data caching method according to claim 1, wherein Further includes: After executing the inference procedure of the first layer, moving the second key-value data corresponding to the new token from the shared block to the second storage device; Reading the second key-value data corresponding to the second layer among the plurality of layers from the second storage device to the shared block; And When performing an inference procedure regarding the second layer, reading the first key-value data corresponding to the second layer from the plurality of cache blocks and reading the second key-value data corresponding to the second layer from the shared block.

4. The data caching method according to claim 1, wherein Further includes: Obtaining the token length supported by the neural network; And Determining the number of the plurality of cache blocks and the number of the shared block according to the token length.

5. The data caching method according to claim 4, wherein Wherein the step of determining the number of the plurality of cache blocks and the number of the shared block according to the token length includes: Obtaining the total number of blocks of the first storage device; Setting the sum of the number of the plurality of cache blocks and the number of the shared block to be the same as the total number of blocks; Dividing the token length by the block size to obtain a block length; and Calculating the minimum value of the number of the shared block under the condition that the value obtained by multiplying the number of the shared block by the number of the plurality of layers and adding the number of the plurality of cache blocks is greater than or equal to the block length.

6. The data caching method according to claim 5, wherein Further includes: Gradually increasing the token length as the neural network continuously generates a plurality of output tokens, wherein the block length is greater than the total number of blocks.

7. The data caching method according to claim 1, wherein Further includes: Generating a plurality of output tokens by the neural network, connecting the plurality of output tokens to the plurality of tokens to obtain a plurality of historical tokens; Obtaining a first prefix of the plurality of historical tokens; Calculating a first hash value of the first prefix; And Storing the first hash value and the key-value data corresponding to the first prefix into the second storage device.

8. The data caching method according to claim 7, wherein, Further includes: Obtain a second prefix of the multiple historical tags, where the second prefix encompasses the first prefix; Calculate a second hash value of the second prefix; And Store the second hash value and the key-value data corresponding to the second prefix in the second storage device.

9. The data caching method according to claim 7, wherein Where the step of storing the first hash value and the key-value data corresponding to the first prefix in the second storage device includes: Determine whether the first hash value hits a first previous hash value of a third storage device, where the access speed of the third storage device is greater than that of the second storage device, and the access speed of the third storage device is less than that of the first storage device; If the first hash value does not hit the first previous hash value, determine whether the third storage device is full; If the third storage device is not full, write the first hash value and the key-value data corresponding to the first prefix into the third storage device; If the third storage device is full, determine whether the second storage device is full; If the second storage device is full, clear the data in the second storage device; If the second storage device is not full, select previous key-value data in the third storage device; Determine whether the hash value of the previous key-value data hits a second previous hash value of the second storage device; and If the hash value of the previous key-value data does not hit the second previous hash value, move the previous key-value data to the second storage device.

10. The data caching method according to claim 1, wherein Where obtaining the multiple key-value data corresponding to the multiple layers includes: Obtain a prefix of the multiple tags; Calculate a hash value of the prefix; Determine whether the hash value hits a previous hash value stored in the second storage device; and If the hash value hits the previous hash value, read the key-value data corresponding to the prefix from the second storage device.

11. The data caching method according to claim 1, wherein Further includes: When performing an inference program for one of the multiple layers, first key-value data of M layers is stored in the multiple cache blocks, and second key-value data of N layers is stored in the shared block, where the multiple cache blocks and the shared block are distributed on the same memory chip, M and N are positive integers, and M is greater than N.

12. A host system, characterized in that, Includes: A processor, including a first storage device; and A second storage device; and Where the processor is electrically connected to the second storage device and is configured to execute multiple steps: Input multiple tags into a neural network and obtain multiple key-value data corresponding to multiple layers; Divide the first storage device into multiple cache blocks and a shared block; Store the first key-value data corresponding to the multiple layers in the multiple key-value data into the cache blocks corresponding to the respective layers in the multiple cache blocks; Store the second key-value data corresponding to the first layer of the multiple layers in the multiple key-value data into the shared block; Store the second key-value data corresponding to at least some other layers except the first layer of the multiple layers in the multiple key-value data into the second storage device; And When performing the inference procedure for the first layer, read the first key-value data corresponding to the first layer from the plurality of cache blocks, and read the second key-value data corresponding to the first layer from the shared block.

13. The host system according to claim 12, wherein Wherein the access speed of the first storage device is greater than the access speed of the second storage device, and the size of the plurality of key-value data is greater than the capacity of the first storage device.

14. The host system according to claim 12, characterized in that, Wherein the plurality of steps further include: After executing the inference procedure for the first layer, move the second key-value data corresponding to the new tag from the shared block to the second storage device; Read the second key-value data corresponding to the second layer in the plurality of layers from the second storage device to the shared block; and When performing the inference procedure for the second layer, read the first key-value data corresponding to the second layer from the plurality of cache blocks, and read the second key-value data corresponding to the second layer from the shared block.

15. The host system according to claim 12, wherein Wherein the plurality of steps further include: Obtain the tag length supported by the neural network; and Determine the number of the plurality of cache blocks and the number of the shared blocks according to the tag length.

16. The host system according to claim 15, wherein, Wherein the step of determining the number of the plurality of cache blocks and the number of the shared blocks according to the tag length includes: Obtain the total number of blocks of the first storage device; Set the sum of the number of the plurality of cache blocks and the number of the shared blocks to be the same as the total number of blocks; Divide the tag length by the block size to obtain the block length; and Calculate the minimum value of the number of the shared blocks under the condition that the value obtained by multiplying the number of the shared blocks by the number of the plurality of layers and then adding the number of the plurality of cache blocks is greater than or equal to the block length.

17. The host system according to claim 16, wherein, Wherein the plurality of steps further include: Gradually increase the tag length as the neural network continuously generates a plurality of output tags, wherein the block length is greater than the total number of blocks.

18. The host system according to claim 12, characterized in that, Wherein the plurality of steps further include: Generate a plurality of output tags by the neural network, and connect the plurality of output tags to the plurality of tags to obtain a plurality of historical tags; Obtain a first prefix of the plurality of historical tags; Calculate a first hash value of the first prefix; and Store the first hash value and the key-value data corresponding to the first prefix in the second storage device.

19. The host system according to claim 18, wherein Wherein the plurality of steps further include: Obtain a second prefix of the plurality of historical tags, wherein the second prefix covers the first prefix; Calculate a second hash value of the second prefix; and Store the second hash value and the key-value data corresponding to the second prefix in the second storage device.

20. The host system according to claim 18, characterized in that, Further includes a third storage device, wherein the step of storing the first hash value and the key-value data corresponding to the first prefix in the second storage device includes: Determine whether the first hash value hits the first previous hash value of the third storage device, where the access speed of the third storage device is greater than that of the second storage device and less than that of the first storage device; If the first hash value does not hit the first previous hash value, determine whether the third storage device is full; If the third storage device is not full, write the first hash value and the key-value data corresponding to the first prefix to the third storage device; If the third storage device is full, determine whether the second storage device is full; If the second storage device is full, clear the data of the second storage device; If the second storage device is not full, select the previous key-value data in the third storage device; Determine whether the hash value of the previous key-value data hits the second previous hash value of the second storage device; and If the hash value of the previous key-value data does not hit the second previous hash value, move the previous key-value data to the second storage device.

21. The host system according to claim 12, characterized in that, Where obtaining the multiple key-value data corresponding to the multiple layers includes: Obtain the prefixes of the multiple tags; Calculate the hash value of the prefix; Determine whether the hash value hits the previous hash value stored in the second storage device; and If the hash value hits the previous hash value, read the key-value data corresponding to the prefix from the second storage device.

22. The host system according to claim 12, wherein, Where the multiple steps further include: When performing an inference program for one of the multiple layers, the first key-value data of M layers is stored in the multiple cache blocks, and the second key-value data of N layers is stored in the shared block, where the multiple cache blocks and the shared block are distributed on the same memory chip, M and N are positive integers, and M is greater than N.