Ring buffer storage method and ring buffer storage system
By adopting a ring buffer storage method in a large language model (LLM), the cache tensor buffer is updated according to the number of input tags, which solves the problem of large storage overhead in the prior art and realizes efficient cache storage management.
Patent Information
- Application Number
- CN202411552859.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-03
- Filing Date
- 2024-11-01
- Publication Date
- 2025-05-06
AI Technical Summary
Large language model (LLM) requires larger storage areas during inference. The existing ring buffer storage method is difficult to effectively manage cache buffers, resulting in large storage overhead.
By adopting a ring buffer storage method in a large language model (LLM), the start storage address of the cache tensor buffer is moved according to the number of input tags of the LLM, an updated cache tensor buffer matrix is formed, and additional space segments are added to the last row of the matrix to optimize cache storage.
Reduces the number of additional space segments used by the cache tensor buffer, avoids unnecessary storage area replication operations, maintains consistent latency, and improves the utilization efficiency of cache storage areas.
Smart Images

Figure CN119938556A_ABST
Abstract
Description
[Technical field]
[0001] The present application relates to data storage, and in particular to a ring buffer storage method and a ring buffer storage system. [Background technology]
[0002] Large language models (LLM), especially those using transformer decoders, usually require large storage areas because they rely on information from previous tokens to predict information from subsequent tokens. To speed up the inference process, common optimization techniques include implementing key (K) cache / value (V) cache.
[0003] During LLM reasoning, new K / V values are generated and written to the K / V cache buffer. In order to effectively manage this cache buffer, a ring buffer mechanism can be used. The K / V values of the current model are stored in a ring buffer. Ideally, the size of the ring buffer should be at least twice the size of the model input cache to allow the model to fully utilize the cache without causing the ring buffer to be reset. For example, a model that can access the previous 20 tokens will require a ring buffer with storage space for at least 40 tokens.
[0004] Therefore, given the storage requirements of LLM, it is crucial to develop a ring buffer that can minimize the storage overhead. [Summary of the invention]
[0005] In an embodiment of the present invention, a ring buffer storage method is disclosed. The ring buffer storage method includes: generating data of a first output according to Q input tokens of a large language model (LLM); and writing the data of the first output into the last Q column vectors of an updated first cache tensor buffer matrix, wherein the starting storage address of the first cache tensor buffer is moved based on the number of input tokens Q of the LLM to update the first cache tensor buffer, wherein the first cache tensor buffer forms a first cache tensor buffer matrix, the updated first cache tensor buffer forms an updated first cache tensor buffer matrix, the first cache tensor buffer matrix includes a plurality of first spatial segments, each row of the first cache tensor buffer matrix includes C spatial segments, C is a cache size, wherein the plurality of first spatial segments have continuous storage addresses, the starting address of each row of the first cache tensor buffer matrix is continuous with the ending address of the previous row, and the last row of the updated first cache tensor buffer matrix includes at least one overhead spatial segment.
[0006] In another embodiment of the present invention, a ring buffer storage method is disclosed. The ring buffer storage method includes: generating data of multiple outputs according to Q input tokens of a large language model (LLM), wherein each output corresponds to a cache tensor buffer, multiple outputs correspond to multiple cache tensor buffers, and multiple cache tensor buffers have continuous addresses and form a concatenated cache tensor buffer, and the concatenated cache tensor buffers form a cache tensor buffer matrix; the cache tensor buffer matrix includes multiple space segments, each row of the cache tensor buffer matrix includes C space segments, C is the cache size, and the starting address of each row of the cache tensor buffer matrix is continuous with the ending address of the previous row; writing the data of the multiple outputs into the last Q column vectors of the updated cache tensor buffer matrix, wherein the starting storage address of each cache tensor buffer is moved according to the number of input tokens Q of the LLM to update the concatenated cache tensor buffer, the updated concatenated cache tensor buffer forms an updated cache tensor buffer matrix, and the last row of the updated cache tensor buffer matrix includes at least one additional space segment.
[0007] In another embodiment of the present invention, a ring buffer storage system is disclosed. The ring buffer storage system includes a ring buffer and a processor. The processor generates data of a first output according to Q input tokens of a large language model (LLM), and writes the data of the first output into the last Q column vectors of an updated first cache tensor buffer matrix, wherein the starting storage address of the first cache tensor buffer is moved based on the number of input tokens Q of the LLM to update the first cache tensor buffer, wherein the first cache tensor buffer forms a first cache tensor buffer matrix, the updated first cache tensor buffer forms an updated first cache tensor buffer matrix, the first cache tensor buffer matrix includes a plurality of first space segments, each row of the first cache tensor buffer matrix includes C space segments, C is a cache size, wherein the plurality of first space segments have continuous storage addresses, the starting address of each row of the first cache tensor buffer matrix is continuous with the ending address of the previous row, and the last row of the updated first cache tensor buffer matrix includes at least one additional space segment.
[0008] The ring buffer storage method and the ring buffer storage system of the present application can reduce the number of additional space segments used by the cache tensor buffer / connected cache tensor buffer.
Brief Description of the Drawings
[0009] Figure 1A A block diagram of a ring buffer storage system for a large language model according to an embodiment of the present invention is shown.
[0010] Figure 1BAnother block diagram of a ring buffer storage system for a large language model according to an embodiment of the present invention is shown.
[0011] Figure 2 A schematic diagram of a ring buffer storage method for a large language model according to an embodiment of the present invention is shown.
[0012] Figure 3 Shows Figure 1A The ring buffer storage system in the first storage mode stores a first state of a K / V cache tensor buffer matrix.
[0013] Figure 4 Shows Figure 1A The ring buffer storage system in the second state of the K / V cache tensor buffer matrix in the first storage mode.
[0014] Figure 5 Shows Figure 1A The third state of the K / V cache tensor buffer matrix of the ring buffer storage system in the first storage mode.
[0015] Figure 6 Shows Figure 1A The ring buffer storage system in the second storage mode is connected to the cache tensor buffer matrix of the first state.
[0016] Figure 7 Shows Figure 1A The ring buffer storage system in the second storage mode connects the cache tensor buffer matrix to a second state.
[0017] Figure 8 Shows Figure 1A The ring buffer storage system in the second storage mode is connected to the cache tensor buffer matrix of the third state.
[0018] Fig. 9 A first state of a connection cache tensor buffer matrix in a third storage mode is shown.
[0019] Fig.10 A second state of the connection cache tensor buffer matrix in the third storage mode is shown.
[0020] Fig.11 A third state of the connection cache tensor buffer matrix in a third storage mode is shown.
[0021] Fig.12 Shows Figure 1A or Figure 1B A flowchart of a ring buffer storage method executed by a ring buffer storage system in the embodiment.
[0022] Fig.13 Shows Figure 1A or Figure 1B A flowchart of a ring buffer storage method executed by a ring buffer storage system in the embodiment. [Specific implementation method]
[0023] In the autoregressive system of Generative Pre-trained Transformer (GPT) and other transformer-based architectures, at least one token from an input sequence can first be converted into a hidden state associated with the at least one token and containing basic information about the at least one token. The hidden state is then processed by multiple transformer layers of the system. In the architecture of the disclosed system, each transformer layer can include an attention mechanism or a self-attention mechanism for updating the hidden state of the input token. This multi-layer processing ensures that the final output is based on a comprehensive and detailed understanding of the entire input sequence, resulting in more accurate and more contextual results.
[0024] Figure 1A A block diagram of a ring buffer storage system 100 for a large language model (LLM) according to an embodiment of the present invention is shown. Figure 1B Another block diagram of a ring buffer storage system 100 for LLM according to an embodiment of the present invention is shown. The ring buffer storage system 100 can be used to store key data and value data of input tags of LLM in an autoregressive mechanism, such as for transformer-based LLM. LLM can be an artificial intelligence (AI) that can process and generate human language. For example, a neural network-based LLM can perform an attention mechanism or a self-attention mechanism. In Figure 1, the ring buffer storage system 100 includes a ring buffer 10 and a processor 12. The processor 12 is connected to the ring buffer 10. The LLM can be software and runs on the processor 12. In one embodiment, the processor 12 may include a processor 121 and a processor 122, such as Figure 1B As shown, processor 121 and processor 122 are independent. Figure 1B In another embodiment, the LLM may be executed on the processor 121. The processor 122 may execute a software program to determine the read / write storage address of the ring buffer 10. The LLM may perform read / write operations based on the read / write storage address of the ring buffer 10. In another embodiment, the processor 122 is a processor such as Figure 1AAs shown. The LLM may run on the processor 12. The LLM may determine the read / write storage address of the ring buffer 10 and perform a read / write operation according to the read / write storage address of the ring buffer 10. In some embodiments, the ring buffer 10 may be located in a DRAM. Alternatively, the ring buffer 10 may also be integrated in the processor where the LLM is located.
[0025] The processor 122 or LLM may request a first cache tensor buffer in the ring buffer and obtain a starting storage address of the first cache tensor buffer, wherein the first cache tensor buffer includes multiple space segments, which form a first cache tensor buffer matrix. These multiple space segments have continuous storage addresses. Each row of the first cache tensor buffer matrix includes C space segments. C is a cache size. The starting address of each row of the first cache tensor buffer matrix is continuous with the ending address of the previous row. The LLM may generate data for the first output according to the Q input tags of the LLM. The processor 122 or LLM may move the starting storage address of the first cache tensor buffer according to the number Q of the input tags of the LLM to update the first cache tensor buffer. The updated first cache tensor buffer forms an updated first cache tensor buffer matrix, wherein the end of the last row of the updated first cache tensor buffer matrix includes at least one additional space segment, wherein the number of additional space segments may be equal to the number of input tags of the LLM, and the number of additional space segments may be Q. The LLM may write the first output directly or indirectly to the last Q column vectors of the updated first cache tensor buffer matrix. In the ring buffer storage system 100, the address space of the Q additional space segments is equal to Q×S, where S is the stride size of a space segment. In the following embodiments and drawings, the size of a space segment is equal to the stride size S.
[0026] The number Q of additional space segments is less than the cache size C. Q, C and S are positive integers. The details of executing the ring buffer storage method and the ring buffer storage system 100 will be described below.
[0027] Figure 2 The schematic diagram of the ring buffer storage system 100 executing the ring buffer storage method is shown. The ring buffer storage method includes steps S101 to S105.
[0028] Step S101: Decoder layer K in LLM receives the hidden state of the input token from decoder layer K-1. Decoder layer K reads the content of K cache tensor buffer according to the starting storage address of K cache tensor buffer, and reads the content of V cache tensor buffer according to the starting storage address of V cache tensor buffer.
[0029] The LLM includes multiple decoder layers K. One decoder layer outputs the hidden state of the input token to the next decoder layer. Each decoder layer has its own K cache tensor buffer and V cache tensor buffer. The starting storage address of the K cache tensor buffer and the starting storage address of the V cache tensor buffer in this step can be calculated by the processor 122. Alternatively, the starting storage address of the K cache tensor buffer and the starting storage address of the V cache tensor buffer can be calculated by the LLM.
[0030] The K-cache tensor buffer includes multiple spatial segments that form a K-cache tensor buffer matrix. These multiple spatial segments have consecutive storage addresses. Each row of the K-cache tensor buffer matrix includes C spatial segments, where C is the cache size. The starting address of each row of the K-cache tensor buffer matrix is consecutive to the ending address of the previous row. The V-cache tensor buffer includes multiple spatial segments that form a V-cache tensor buffer matrix, and the structure of the V-cache tensor buffer matrix is similar to that of the K-cache tensor buffer matrix.
[0031] Step S102: The decoder layer K determines the updated hidden state of the input token and the key data and numerical data of the input token based on the received hidden state of the input token and the contents of the K cache tensor buffer and the V cache tensor buffer.
[0032] Step S103: Move the starting storage address of the K cache tensor buffer according to the number Q of input tokens of the LLM to update the K cache tensor buffer. Move the starting storage address of the V cache tensor buffer according to the number Q of input tokens of the LLM to update the V cache tensor buffer.
[0033] Step S103 may be performed by the processor 122. Alternatively, step S103 may be performed by the LLM. The updated K cache tensor buffer forms an updated K cache tensor buffer matrix, wherein the end of the last row of the updated K cache tensor buffer matrix includes Q additional spatial segments. The updated V cache tensor buffer forms an updated V cache tensor buffer matrix, wherein the end of the last row of the updated V cache tensor buffer matrix includes Q additional spatial segments.
[0034] Step S104: Decoder layer K writes the key data of the input token into the last Q column vectors of the updated K cache tensor buffer matrix, and writes the numerical data of the input token into the last Q column vectors of the updated V cache tensor buffer matrix. Decoder layer K outputs the updated hidden state of the input token to decoder layer K+1.
[0035] Alternatively, the decoder layer K may write the key data of the input tokens into a buffer with consecutive addresses. The data segments of the key data are copied from the buffer to the last Q column vectors of the updated K cache tensor buffer matrix. The decoder layer K may write the numerical data of the input tokens into a buffer with consecutive addresses. The data segments of the numerical data are copied from the buffer to the last Q column vectors of the updated V cache tensor buffer matrix.
[0036] Figure 3 The first write state of the K / V cache tensor buffer matrix in the first storage mode of the ring buffer storage system 100 is shown. As described above, the K cache tensor buffer forms a K cache tensor buffer matrix M, and the V cache tensor buffer forms a V cache tensor buffer matrix N. To avoid ambiguity, the K / V cache tensor buffer matrix is referred to as the first cache tensor buffer matrix M (for key cache tensors) and the second cache tensor buffer matrix N (for value cache tensors). The dimension of the first cache tensor buffer matrix M is (R1, C×S). The dimension of the second cache tensor buffer matrix is (R2, C×S). R1 is the row dimension of the first cache tensor buffer matrix M. R2 is the row dimension of the second cache tensor buffer matrix N. R1 and R2 are positive integers. C is the cache size. S is the step size, and S is the step size of a segment. In one embodiment, the row dimensions R1 and R2, and the step size S are determined according to the architecture of the LLM. The cache size C may be determined according to a use case scenario and is less than or equal to the longest tag length (number) supported by the LLM. The first cache tensor buffer matrix M includes a plurality of spatial segments M11 to M16, M21 to M26, and M31 to M36. The plurality of spatial segments M11 to M16, M21 to M26, and M31 to M36 have consecutive storage addresses. For example, the storage addresses of the spatial segments M11 to M16, M21 to M26, and M31 to M36 may be represented as Table T1:
[0037] Storage address #1 to #6 #7 to #12 #13 to #18 Space segment M11 to M16 M21 to M26 M31 to M36
[0038] Table T1
[0039] Similarly, the second cache tensor buffer matrix N includes a plurality of space segments N11 to N16, N21 to N26, and N31 to N36. The plurality of space segments N11 to N16, N21 to N26, and N31 to N36 have continuous storage addresses. Since the second cache tensor buffer matrix N also includes "continuous" space segments, the description of its storage address allocation is omitted here.
[0040] The processor 122 or LLM allocates a predetermined number m' of space segments as additional space segments to be appended to the end of the first cache tensor buffer matrix. In some embodiments, the predetermined number m' may be at least C. Similarly, the processor 122 or LLM allocates a predetermined number n' of space segments as additional space segments to be appended to the end of the second cache tensor buffer matrix. In some embodiments, the predetermined number n' may be at least C. Figure 3 As shown, the additional space segments attached to the first cache tensor buffer matrix include A1, B1, C1, D1, E1 and F1, and the additional space segments attached to the second cache tensor buffer matrix include A2, B2, C2, D2, E2 and F2. Here, each additional space segment occupies an address space equal to S. The first cache tensor buffer matrix M can be regarded as an "empty" matrix. Similarly, the second cache tensor buffer matrix N can be regarded as an "empty" matrix.
[0041] Assume that there are two input tokens processed by a large language model (LLM), for example, the two input tokens are "Nice" and "to". After the output key data of the two input tokens are generated from the LLM, the output key data of the two input tokens are directly or indirectly written to the last Q=2 column vectors of the first cache tensor buffer matrix M in the ring buffer 10. If the hardware running the LLM is capable of directly writing the output key data of the two input tokens to the last Q=2 column vectors of the first cache tensor buffer matrix M, then the output key data of the two input tokens are directly written to the last Q=2 column vectors of the first cache tensor buffer matrix M. If the hardware running the LLM does not have such a capability, the output key data of the two input tokens are written to a continuous buffer, and then the output key data of the two input tokens are copied from the continuous buffer to the last Q=2 column vectors of the first cache tensor buffer matrix segment by segment. For example, the data of data segment K1 is written to the "empty" space segment M15. The data of data segment K2 is written to the "empty" space segment M16. The data of data segment K3 is written to the "empty" space segment M25. The data of data segment K4 is written into the "empty" space segment M26. The data of data segment K5 is written into the "empty" space segment M35. The data of data segment K6 is written into the "empty" space segment M36. Similarly, after generating the output value (Value) data of the two input tags, the output value data of the two input tags are directly or indirectly written into the last Q=2 column vectors of the second cache tensor buffer matrix N. For example, the data of data segment V1 is written into the "empty" space segment N15. The data of data segment V2 is written into the "empty" space segment N16. The data of data segment V3 is written into the "empty" space segment N25. The data of data segment V4 is written into the "empty" space segment N26. The data of data segment V5 is written into the "empty" space segment N35. The data of data segment V6 is written into the "empty" space segment N36.
[0042] Figure 4 1 shows a second write state of the K / V cache tensor buffer matrix in the first storage mode of the ring buffer storage system 100. In the previous state, the data of the data segments K1 to K6 are written to the last Q=2 column vectors of the first cache tensor buffer matrix M. The data of the data segments V1 to V6 are written to the last Q=2 column vectors of the second cache tensor buffer matrix N. Then, the LLM further processes two input tags, for example, the two input tags are "meet" and "you".
[0043] Then, the starting storage address of the first cache tensor buffer is shifted according to the number of input tags of the LLM to update the first cache tensor buffer. The updated first cache tensor buffer forms an updated first cache tensor buffer matrix M. The updated first cache tensor buffer matrix M includes additional spatial segments A1 and B1 at the end of its last row. Similarly, the starting storage address of the second cache tensor buffer is shifted according to the number of input tags of the LLM to update the second cache tensor buffer. The updated second cache tensor buffer forms an updated second cache tensor buffer matrix N. The updated second cache tensor buffer matrix N includes additional spatial segments A2 and B2 at the end of its last row.
[0044] For example, the original first cache tensor buffer matrix M can be shown in an expanded form as shown in Table T2.
[0045]
[0046] Table T2
[0047] For example, when the number of input tokens is 2, the starting storage address of the first cache tensor buffer is moved by increasing the starting storage address of the first cache tensor buffer by a first offset storage address equal to 2×S. The first cache tensor buffer matrix M may be updated as shown in Table T4.
[0048]
[0049] Table T4
[0050] After the first cache tensor buffer matrix M is updated, its first row includes spatial segments M13 to M16, and M21 to M22, the second row includes spatial segments M23 to M26 and M31 to M32, and the third row includes spatial segments M33 to M36, A1, and B1. Figure 4, the last Q=2 column vectors of the first cache tensor buffer matrix M include "empty" space segments {M21 and M22, M31 and M32, A1 and B1}. It should be understood that the last Q=2 column vector space segments {M21 and M22, M31 and M32, A1 and B1} of the updated first cache tensor buffer matrix M are discontinuous. Therefore, the last Q=2 column vectors of the first cache tensor buffer matrix M can be used to cache the data of the data segment of the output key data. The data of each data segment of the output key data is directly or indirectly written to the corresponding space segment of the last Q=2 column vectors of the updated first cache tensor buffer matrix M based on its storage address. For example, the data of data segment K7 is written to the "empty" space segment M21. The data of data segment K8 is written to the "empty" space segment M22. The data of data segment K9 is written to the "empty" space segment M31. The data of data segment K10 is written to the "empty" space segment M32. The data of the data segment K11 is written into the "empty" space segment A1. The data of the data segment K12 is written into the "empty" space segment B1.
[0051] Similarly, when the number of input tags is 2, the starting storage address of the second cache tensor buffer is moved by increasing the starting storage address of the second cache tensor buffer by an offset storage address equal to 2×S. Therefore, the second cache tensor buffer matrix N can be updated. After the second cache tensor buffer matrix N is updated, its first row includes spatial segments N13 to N16 and N21 to N22, the second row includes spatial segments N23 to N26 and N31 to N32, and the third row includes spatial segments N33 to N36, A2, and B2. Figure 4 , the last Q=2 column vectors of the second cache tensor buffer matrix N include "empty" space segments {N21 and N22, N31 and N32, A2 and B2}. It should be understood that the last Q=2 column vector space segments {N21 and N22, N31 and N32, A2 and B2} of the updated second cache tensor buffer matrix N are discontinuous. Therefore, the last Q=2 column vectors of the second cache tensor buffer matrix N can be used to cache data segments of output numerical data. The data of each data segment of the output numerical data is written directly or indirectly to the corresponding space segment of the last Q=2 column vectors of the updated second cache tensor buffer matrix N based on its storage address. For example, the data of data segment V7 is written to the "empty" space segment N21. The data of data segment V8 is written to the "empty" space segment N22. The data of data segment V9 is written to the "empty" space segment N31. The data of data segment V10 is written to the "empty" space segment N32. The data of data segment V11 is written into the "empty" space segment A2. The data of data segment V12 is written into the "empty" space segment B2.
[0052] Figure 5The third state of the K / V cache tensor buffer matrix in the first storage mode of the ring buffer storage system 100 is shown. Here, the LLM further processes an input tag, for example, the input tag is "!". Then, the starting storage address of the first cache tensor buffer can be moved according to the number of input tags to cache the output key data generated from the LLM. The updated first cache tensor buffer forms an updated first cache tensor buffer matrix M. The updated first cache tensor buffer matrix M includes an additional space segment C1 at the end of its last row. The starting storage address of the second cache tensor buffer can be moved according to the number of input tags to cache the output numerical data generated from the LLM. The updated second cache tensor buffer forms an updated second cache tensor buffer matrix N. The updated second cache tensor buffer matrix N includes an additional space segment C2 at the end of its last row.
[0053] For example, when the number of input tokens is 1, the starting storage address of the first cache tensor buffer is moved by increasing the starting storage address of the first cache tensor buffer by an offset storage address equal to 1×S. Therefore, the first cache tensor buffer matrix M can be updated. In this way, the first cache tensor buffer matrix M can be shown in an expanded form in Table T5.
[0054]
[0055] Table T5
[0056] After the first cache tensor buffer matrix M is updated, its first row includes spatial segments M14 to M16 and M21 to M23, its second row includes spatial segments M24 to M26 and M31 to M33, and its third row includes M34 to M36, A1, B1, and C1. Figure 5 , the last Q=1 column vectors of the first cache tensor buffer matrix M include "empty" space segments {M23, M33 and C1}. It should be understood that the space segments {M23, M33 and C1} of the last Q=1 column vectors of the updated first cache tensor buffer matrix M are discontinuous. The last Q=1 column vectors of the first cache tensor buffer matrix M can be used to cache data segments of output key data. The data of each data segment of the output key data is written directly or indirectly to the corresponding space segment of the last Q=1 column vector of the updated first cache tensor buffer matrix M based on its storage address. For example, the data of data segment K13 is written to the "empty" space segment M23. The data of data segment K14 is written to the "empty" space segment M33. The data of data segment K15 is written to the "empty" space segment C1.
[0057] Similarly, when the number of input tags is 1, the starting storage address of the second cache tensor buffer is moved by increasing the starting storage address of the second cache tensor buffer by an offset storage address equal to 1×S. As a result, the second cache tensor buffer matrix N is updated. After the second cache tensor buffer matrix N is updated, its first row includes spatial segments N14 to N16, and N21 to N23, the second row includes spatial segments N24 to N26 and N31 to N33, and the third row includes N34 to N36, A2, B2, and C2. Figure 5 , the last Q=1 column vectors of the second cache tensor buffer matrix N include "empty" space segments {N23, N33 and C2}. It should be understood that the space segments {N23, N33 and C2} of the last Q=1 column vectors of the updated second cache tensor buffer matrix N are discontinuous. Therefore, the last Q=1 column vectors of the second cache tensor buffer matrix N can be used to cache the data of the data segments of the output numerical data. The data of each data segment of the output numerical data is written directly or indirectly to the corresponding segment of the last Q=1 column vectors of the updated second cache tensor buffer matrix N based on its storage address. For example, the data of data segment V13 is written to the "empty" space segment N23. The data of data segment V14 is written to the "empty" space segment N33. The data of data segment V15 is written to the "empty" space segment C2.
[0058] Therefore, compared with the solution of using twice the model input cache size to avoid resetting the ring buffer, the ring buffer storage method of the present application only requires fewer additional storage segments and does not need to reset the ring buffer.
[0059] Figure 6The first write state of the concatenated cache tensor buffer matrix in the second storage mode of the ring buffer storage system 100 is shown. In this embodiment, different cache tensor buffers are connected to generate a connected cache tensor buffer. Each cache tensor buffer forms a cache tensor buffer matrix, and the connected cache tensor buffers form a connected cache tensor buffer matrix. In this embodiment, the buffer cache efficiency can be further improved by combining different cache tensor buffer matrices. As described above, the dimension of the first cache tensor buffer matrix M is (R1, C×S). The dimension of the second cache tensor buffer matrix N is (R2, C×S). The row dimension R1 and the row dimension R2 may be different. The first cache tensor buffer matrix M and the second cache tensor buffer matrix N may be connected to generate a connected cache tensor buffer matrix F. The connected cache tensor buffer matrix F includes a plurality of spatial segments M11 to M16, M21 to M26, M31 to M36, N11 to N16, N21 to N26, and N31 to N36. The space segments M11 to M16, M21 to M26, M31 to M36, N11 to N16, N21 to N26, and N31 to N36 have consecutive storage addresses. For example, the storage addresses of the space segments M11 to M16, M21 to M26, M31 to M36, N11 to N16, N21 to N26, and N31 to N36 may be as shown in Table T6:
[0060] Storage address #1 to #6 #7 to #12 #13 to #18 Space segment M11 to M16 M21 to M26 M31 to M36 Storage address #19 to #24 #25 to #30 #31 to #36 Space segment N11 to N16 N21 to N26 N31 to N36
[0061] Table T6
[0062] The processor 122 or LLM allocates a predetermined number L' of spatial segments as additional spatial segments to be attached to the connected cache tensor buffer matrix. In some embodiments, the predetermined number L' may be at least C. Figure 6 As shown, the additional space segments attached to the connected cache tensor buffer matrix include A, B, C, D, E, and F. Here, each additional space segment occupies an address space equal to S. The connected cache tensor buffer matrix F can be regarded as an "empty" matrix.
[0063] Assume that there are two input tags processed by LLM, for example, the two input tags are "Nice" and "to". After the output key data of the two tags are generated from LLM, the output key data of the two tags can be written directly or indirectly to the corresponding space segments of the connected cache tensor buffer matrix F. The methods of direct writing and indirect writing are similar to the previous embodiments and are not described in detail here for the sake of brevity. For example, the data of data segment K1 is written to the "empty" space segment M15. The data of data segment K2 is written to the "empty" space segment M16. The data of data segment K3 is written to the "empty" space segment M25. The data of data segment K4 is written to the "empty" space segment M26. The data of data segment K5 is written to the "empty" space segment M35. The data of data segment K6 is written to the "empty" space segment M36. Similarly, after the output numerical data of the two tags are generated from LLM, the output numerical data of the two tags can be written directly or indirectly to the corresponding space segments of the connected cache tensor buffer matrix F. For example, the data of data segment V1 is written to the "empty" space segment N15. The data of data segment V2 is written to the "empty" space segment N16. The data of data segment V3 is written to the "empty" space segment N25. The data of data segment V4 is written to the "empty" space segment N26. The data of data segment V5 is written to the "empty" space segment N35. The data of data segment V6 is written to the "empty" space segment N36.
[0064] Figure 7 1 shows a second write state of the connected cache tensor buffer matrix F in the second storage mode of the ring buffer storage system 100. In the previous state, the data segments K1 to K6 of the output key data and the data segments V1 to V6 of the output value data are written into the last Q=2 column vectors of the connected cache tensor buffer matrix F. Then, the LLM further processes two input tags, for example, the two input tags are "meet" and "you". The number of input tags Q is 2.
[0065] The original concatenated cache tensor buffer matrix F can be shown in expanded form in Table T7.
[0066]
[0067] Table T7
[0068] Then, the starting storage address of the first cache tensor buffer is moved by increasing the starting storage address of the first cache tensor buffer by an offset storage address equal to 2×S. As a result, the first cache tensor buffer matrix M is updated. The starting storage address of the second cache tensor buffer is moved by increasing the starting storage address of the second cache tensor buffer by an offset storage address equal to 2×S. As a result, the second cache tensor buffer matrix N is updated. The address space size of the first cache tensor buffer matrix is R1×C×S, and the address space size of the second cache tensor buffer matrix is R2×C×S. Therefore, the connected cache tensor buffer matrix F can be updated as shown in Table T8.
[0069]
[0070] Table T8
[0071] After the connection cache tensor buffer matrix F is updated, its first row includes spatial segments M13 to M16, and M21 to M22, the second row includes spatial segments M23 to M26 and M31 to M32, the third row includes spatial segments M33 to M36 and N11 to N12, the fourth row includes spatial segments N13 to N16 and N21 to N22, the fifth row includes spatial segments N23 to N26 and N31 to N32, and the sixth row includes spatial segments N33 to N36, A and B. Figure 7, the last Q=2 column vectors of the connected cache tensor buffer matrix F include "empty" space segments {M21 to M22, M31 to M32, N11 to N12, N21 to N22, N31 to N32, A and B}. It should be understood that the last Q=2 column vector space segments {M21 to M22, M31 to M32, N11 to N12, N21 to N22, N31 to N32, A and B} of the connected cache tensor buffer matrix F can be used to cache data of data segments of output key data and output numerical data. The data of each segment of output key data and output numerical data is written directly or indirectly to the corresponding segment of the connected cache tensor buffer matrix F based on its storage address. For example, the data of data segment K7 is written to the "empty" space segment M21. The data of data segment K8 is written to the "empty" space segment M22. The data of data segment K9 is written to the "empty" space segment M31. The data of data segment K10 is written to the "empty" space segment M32. The data of data segment K11 is written to the "empty" space segment N11. The data of data segment K12 is written to the "empty" space segment N12. The data of data segment V7 is written to the "empty" space segment N21. The data of data segment V8 is written to the "empty" space segment N22. The data of data segment V9 is written to the "empty" space segment N31. The data of data segment V10 is written to the "empty" space segment N32. The data of data segment V11 is written to the "empty" space segment A. The data of data segment V12 is written to the "empty" space segment B.
[0072] Figure 8 The third write state of the connected cache tensor buffer matrix F in the second storage mode of the ring buffer storage system 100 is shown. Here, the LLM further processes an input tag, for example, the input tag is "!". The number of input tags Q is 1.
[0073] Then, the starting storage address of the first cache tensor buffer is moved by increasing the starting storage address of the first cache tensor buffer by the offset storage address equal to 1×S. As a result, the first cache tensor buffer matrix M is updated. The starting storage address of the second cache tensor buffer is moved by increasing the starting storage address of the second cache tensor buffer by the offset storage address equal to 1×S. As a result, the second cache tensor buffer matrix N is updated. Therefore, the connected cache tensor buffer matrix F can be updated as shown in Table T9.
[0074]
[0075] Table T9
[0076] Here, the first row of the updated connection cache tensor buffer matrix F includes space segments M14 to M16 and M21 to M23, the second row includes space segments M24 to M26 and M31 to M33, the third row includes space segments M33 to M36 and N11 to N13, the fourth row includes space segments N14 to N16 and N21 to N23, the fifth row includes space segments N24 to N26 and N31 to N33, and the sixth row includes space segments N34 to N36, A, B, and C. Figure 8 , the last Q=1 column vectors of the connection cache tensor buffer matrix F include "empty" space segments {M23, M33, N13, N23, N33 and C}. It should be understood that the space segments {M23, M33, N13, N23, N33 and C} of the last Q=1 column vectors of the updated connection cache tensor buffer matrix F are discontinuous. Therefore, the last Q=1 column vectors of the connection cache tensor buffer matrix F can be used to cache the data of the data segments of the output key data and the output numerical data. The data of each segment of the output key data is directly or indirectly written to the corresponding space segment of the connection cache tensor buffer matrix F based on its storage address. The data of each segment of the output numerical data is directly or indirectly written to the corresponding space segment of the connection cache tensor buffer matrix F based on its storage address. For example, the data of data segment K13 is written to the "empty" space segment M23. The data of data segment K14 is written to the "empty" space segment M33. The data of data segment K15 is written into the "empty" space segment N13. The data of data segment V13 is written into the "empty" space segment N23. The data of data segment V14 is written into the "empty" space segment N33. The data of data segment V15 is written into the "empty" space segment C.
[0077] In the above Figure 6-8 In , each cache tensor buffer forms a cache tensor buffer matrix, and each cache tensor buffer matrix includes a plurality of rows. Figure 6-8 The method in can also be applied to the case where the cache tensor buffer matrix consists of only one row.
[0078] In the ring buffer storage system 100, any software / hardware or technical modifications fall within the scope of the present invention. For example, if a single large continuous storage space is impractical or undesirable, B continuous storage blocks can be introduced to divide the large continuous storage space. Assuming there are B storage blocks, and assuming that the sliding window can move at most Q marks without triggering a reset of the ring buffer, then the total additional storage size of the buffer is at most equal to B×Q×S. B is a positive integer greater than or equal to 2. For example, in the aforementioned Figure 3-Figure 5 In the embodiment of the present invention, the first cache tensor buffer and the second cache tensor buffer are separate storage blocks. The ends of the first cache tensor buffer and the second cache tensor buffer are attached with corresponding additional space segments.
[0079] In the third storage mode, different cache tensor buffers are connected to generate a connected cache tensor buffer. Different cache tensor buffers have consecutive addresses. The processor 122 or LLM allocates a predetermined number L' of space segments as additional space segments to be attached to the end of the connected cache tensor buffer. In some embodiments, the predetermined number L' may be at least C. The connected cache tensor buffers form a cache tensor buffer matrix. In this embodiment, each cache tensor buffer may be a row of the cache tensor buffer matrix.
[0080] The third storage mode can be applied to the following scenario. LLM includes multiple decoder layers. Each decoder layer has a corresponding K cache tensor buffer and a V cache tensor buffer. Each decoder layer outputs the corresponding key data of the input tag to the corresponding K cache tensor buffer. Multiple K cache tensor buffers are connected to generate a connected K cache tensor buffer. The connected K cache tensor buffers form a cache tensor buffer matrix. Each decoder layer outputs the corresponding numerical data of the input tag to the corresponding V cache tensor buffer. Multiple V cache tensor buffers are connected to generate a connected V cache tensor buffer. The connected V cache tensor buffers form a cache tensor buffer matrix. The following Figure 9-11 It will be described using an example of storing multiple output key data into multiple K-cache tensor buffers.
[0081] Fig. 9 1 shows a first write state of a cache tensor buffer matrix in a third storage mode of the ring buffer storage system 100. In this embodiment, 3 K cache tensor buffers are shown. Each K cache tensor buffer may be a row of the cache tensor buffer matrix.
[0082] Assume that the LLM processes two input tags, for example, the two input tags are "Nice" and "to". After the LLM generates the first output key data of the two input tags, the first output key data of the two input tags are written directly or indirectly to the first K cache tensor buffer. For example, the data of data segment K1 is written to the "empty" space segment M19. The data of data segment K2 is written to the "empty" space segment M20. After the LLM generates the second output key data of the two input tags, the second output key data of the two input tags are written directly or indirectly to the second K cache tensor buffer. For example, the data of data segment K1' is written to the "empty" space segment M29. The data of data segment K2 is written to the "empty" space segment M30. After the LLM generates the third output key data of the two input tags, the third output key data of the two input tags are written directly or indirectly to the third K cache tensor buffer. For example, the data of data segment K1" is written into the "empty" space segment M39. The data of data segment K2" is written into the "empty" space segment M40. In this embodiment, the data of the first output key data is written into continuous space segments. The data of the second output key data is written into continuous space segments. The data of the third output key data is written into continuous space segments.
[0083] Fig.10 1 shows a second write state of the cache tensor buffer matrix in the third storage mode of the ring buffer storage system 100. Here, the LLM further processes two input tags, for example, the two input tags are "meet" and "you".
[0084] Then, the starting storage address of the first K cache tensor buffer is updated according to the number of input tags of the LLM. The starting storage address of the second K cache tensor buffer is updated according to the number of input tags of the LLM. The starting storage address of the third K cache tensor buffer is updated according to the number of input tags of the LLM. For example, when the number of input tags is 2, the starting storage address of each K cache tensor buffer is moved by increasing the first offset storage address equal to 2×S. As a result, the cache tensor buffer matrix F' can be updated as shown in Table T10.
[0085]
[0086] Table T10
[0087] Each segment of the first, second and third output key data is written directly or indirectly to the corresponding segment of the cache tensor buffer matrix F' based on its storage address. For example, the data of data segment K3 is written into the "empty" space segment M21. The data of data segment K4 is written into the "empty" space segment M22. The data of data segment K3' is written into the "empty" space segment M31. The data of data segment K4' is written into the "empty" space segment M32. The data of data segment K3" is written into the "empty" space segment A. The data of data segment K4" is written into the "empty" space segment B.
[0088] Fig.11 1 shows a third write state of the cache tensor buffer matrix in the third storage mode of the ring buffer storage system 100. Here, the LLM further processes one input tag, for example, the input tag is "!". At this time, the number of input tags is 1.
[0089] Then, the starting storage addresses of the first, second, and third K cache tensor buffers are updated according to the number of input tokens of the LLM. For example, when the number of input tokens is 1, the starting storage address of each K cache tensor buffer is moved by increasing the first offset storage address equal to 1×S. As a result, the cache tensor buffer matrix F' can be updated as shown in Table T11.
[0090]
[0091] Table T11
[0092] Each segment of the first, second and third output key data is written directly or indirectly to the corresponding segment of the cache tensor buffer matrix F' based on its storage address. For example, the data of data segment K5 is written to the "empty" space segment M23. The data of data segment K5' is written to the "empty" space segment M33. The data of data segment K5" is written to the "empty" space segment C.
[0093] Fig.12 The flowchart of the ring buffer storage method executed by the ring buffer storage system 100 is shown. The ring buffer storage method includes steps S1201 to S1203. Any hardware / software or technical modification belongs to the scope of the present invention. Fig.12 The method shown in can be executed Figure 1A , Figure 1B , Figure 2-Figure 11 The corresponding scheme in . Steps S1201 to S1203 are as follows.
[0094] Step S1201: Generate first output data according to Q input tags of LLM.
[0095] Step S1202: According to the number of input tags Q of the LLM, the starting storage address of the first cache tensor buffer is moved to update the first cache tensor buffer. The first cache tensor buffer forms a first cache tensor buffer matrix. The updated first cache tensor buffer forms an updated first cache tensor buffer matrix. The first cache tensor buffer matrix includes multiple first space segments, and each row of the first cache tensor buffer matrix includes C space segments, where C is the cache size. Multiple first space segments have continuous storage addresses. The starting address of each row of the first cache tensor buffer matrix is continuous with the ending address of the previous row of the first cache tensor buffer matrix. The last row of the updated first cache tensor buffer matrix includes at least one additional space segment. Among them, the end of the last row of the updated first cache tensor buffer matrix includes at least one additional space segment, and the number of additional space segments can be Q.
[0096] Wherein, "cache tensor buffer" is a buffer for storing cache tensors. Any reasonable hardware or technical transformation belongs to the scope of the present invention. For example, the data of the cache tensor buffer or the cache tensor buffer matrix can be in tensor format. In addition, the tensor format can include an array format, a tuple format, or other signal formats. "
[0097] In some embodiments, the method further includes generating data of a second output according to the Q input tags of the LLM; and writing the data of the second output into the last Q column vectors of the updated second cache tensor buffer matrix; wherein the starting storage address of the second cache tensor buffer is moved based on the number Q of input tags of the LLM to update the second cache tensor buffer. Wherein, the second cache tensor buffer forms a second cache tensor buffer matrix, and the updated second cache tensor buffer forms the updated second cache tensor buffer matrix, wherein the second cache tensor buffer matrix includes a plurality of second spatial segments, the plurality of second spatial segments have continuous storage addresses, the starting address of each row of the second cache tensor buffer matrix is continuous with the ending address of the previous row of the second cache tensor buffer matrix, and the starting storage address of the first cache tensor buffer matrix is immediately adjacent to the ending storage address of the second cache tensor buffer matrix. The first cache tensor buffer matrix is continuous with the second cache tensor buffer matrix.
[0098] Wherein, the number of additional spatial segments is the number of input tags Q of the LLM, the address space of the Q additional spatial segments is equal to Q×S, the number of additional spatial segments Q is less than the cache size C, Q and S are positive integers, the dimension of the first cache tensor buffer matrix is (R1, C×S), R1 is the row dimension, and R1 is a positive integer. Wherein, in some embodiments, before the starting storage address of the first cache tensor buffer is moved for the first time, the address space of the additional spatial segments attached to the first cache tensor buffer is at least C×S. In one embodiment, before the starting storage address of the first cache tensor buffer is moved for the first time, the number of additional spatial segments attached to the first cache tensor buffer is at least C. Wherein, S is the step size of the spatial segment.
[0099] In some embodiments, the starting storage address of the first cache tensor buffer is moved by increasing a first offset storage address equal to Q×S.
[0100] In some embodiments, the spatial segments of the last Q column vectors of the updated first cache tensor buffer matrix are discontinuous, and the data of each segment of the first output is written directly or indirectly to the corresponding segment of the last Q column vectors of the updated first cache tensor buffer matrix according to its storage address.
[0101] In some embodiments, the method further includes generating data of a second output according to the Q input tags of the LLM; and writing the data of the second output into the last Q column vectors of the updated second cache tensor buffer matrix; wherein the starting storage address of the second cache tensor buffer is moved by the number Q of input tags of the LLM to update the second cache tensor buffer, wherein the second cache tensor buffer forms a second cache tensor buffer matrix, and the updated second cache tensor buffer forms the updated second cache tensor buffer matrix, wherein the second cache tensor buffer matrix includes a plurality of second spatial segments, the plurality of second spatial segments have continuous storage addresses, the starting address of each row of the second cache tensor buffer matrix is continuous with the ending address of the previous row of the second cache tensor buffer matrix, and the starting storage address of the first cache tensor buffer matrix is immediately adjacent to the ending storage address of the second cache tensor buffer matrix. The dimension of the second cache tensor buffer matrix is (R2, C×S), R2 is the row dimension of the second cache tensor buffer matrix, and R2 is a positive integer. Wherein, the spatial segments of the last Q column vectors of the updated second cache tensor buffer matrix are discontinuous, and the data of each segment of the second output is directly or indirectly written to the corresponding segment of the last Q column vectors of the updated second cache tensor buffer matrix according to its storage address. Wherein, the starting storage address of the second cache tensor buffer is moved by adding a second offset storage address equal to Q×S. Wherein, in some embodiments, the offset storage address to which the starting storage address of the second cache tensor buffer is moved is the same as the offset storage address to which the starting storage address of the first cache tensor buffer is moved.
[0102] In some embodiments, the contents of the first cache tensor buffer are read before the starting storage address of the first cache tensor buffer is moved, wherein the data of the first output is generated based on the contents of the first cache tensor buffer and Q input tags of the LLM.
[0103] Step S1203: Write the first output data into the last Q column vectors of the updated first cache tensor buffer matrix.
[0104] The details of steps S1201 to S1203 have been described above, so they will not be repeated here. In the ring buffer storage system 100, since the cache tensor buffer is in matrix form, appending Q additional space segments to the matrix can provide Q*R space segments for writing, where R is the row dimension of the matrix. With this arrangement, the total number of additional space segments is reduced. Instead of using twice the model input cache size to avoid resetting the ring buffer, the ring buffer only requires a small number of additional space segments without resetting the ring buffer. Therefore, since the number of additional space segments is small enough, the capacity of the ring buffer can be expanded to enable it to handle the theoretical limit of the model. A ring buffer with sufficient capacity eliminates the need to perform storage area copying to reset the cache to the top. In addition, by avoiding storage area copying when the ring buffer is reset, the ring buffer storage system can maintain a consistent delay each time the LLM updates the ring buffer.
[0105] Fig.13 The flowchart of the ring buffer storage method executed by the ring buffer storage system 100 is shown. The ring buffer storage method includes steps S1301 to S1303. Any hardware / software or technical modification belongs to the scope of the present invention. Fig.13 The method shown in can be executed Figure 1A , Figure 1B , Figure 2 , Figure 6-Figure 11 Steps S1301 to S1303 are as follows.
[0106] Step S1301: Generate data for multiple outputs according to Q input tokens of a large language model (LLM). Each output corresponds to a cache tensor buffer, and multiple outputs correspond to multiple cache tensor buffers. These cache tensor buffers have consecutive addresses and form a connected cache tensor buffer. The connected cache tensor buffers form a cache tensor buffer matrix. The cache tensor buffer matrix includes multiple space segments, each row of the cache tensor buffer matrix includes C space segments, C is the cache size, and the starting address of each row of the cache tensor buffer matrix is continuous with the ending address of the previous row of the cache tensor buffer matrix.
[0107] Step S1302: The starting storage address of each cache tensor buffer is moved by the input tag quantity Q based on the LLM to update the connected cache tensor buffer, wherein the updated connected cache tensor buffer forms an updated cache tensor buffer matrix, and the last row of the updated cache tensor buffer matrix includes at least one additional space segment.
[0108] The end of the last row of the updated cache tensor buffer matrix includes at least one additional spatial segment, and the number of the additional spatial segments can be Q.
[0109] In some embodiments, the number of the additional space segments is the number Q of input tags of the LLM, the address space of the Q additional space segments is equal to Q×S, the number Q of the additional space segments is less than the cache size, and Q and S are positive integers. In some embodiments, before the starting storage address of the connected cache tensor buffer is first moved, the address space of the additional space segment attached to the connected cache tensor buffer is at least C×S. Where S is the stride of the space segment.
[0110] In some embodiments, the starting storage address of each cache tensor buffer is shifted by increasing the offset storage address equal to Q×S.
[0111] In some embodiments, the spatial segments of the last Q column vectors of the updated cache tensor buffer matrix are discontinuous, and the data of each segment of each output is written directly or indirectly to the corresponding segment of the last Q column vectors of the updated cache tensor buffer matrix based on its storage address.
[0112] In some embodiments, each cache tensor buffer is a row of the cache tensor buffer matrix.
[0113] Optionally, in some embodiments, the contents of each cache tensor buffer are read before the starting storage address of the cache tensor buffer is moved, wherein multiple output data are generated based on the contents of the multiple cache tensor buffers and the Q input tags of the LLM.
[0114] Step S1303: write the multiple output data into the last Q column vectors of the updated cache tensor buffer matrix.
[0115] In general, the present invention discloses a ring buffer storage method and a ring buffer storage system for efficiently managing cache storage areas in a large language model (LLM). The ring buffer storage system can minimize the number of additional space segments of the storage area by strategically using the ring buffer (cache tensor buffer) in matrix form and optimizing the segment configuration, so that less additional space is required for the storage area. By avoiding unnecessary storage area copy operations, the ring buffer storage system can maintain a consistent delay each time the LLM updates the ring buffer. In addition, the ring buffer storage system can optimize the utilization of the cache storage area, thereby improving overall performance and efficiency. Since the row dimension, cache size and stride are adjustable, the ring buffer storage system can adapt to different LLM architectures and cache requirements through these adjustable parameters, and can be expanded to adapt to larger models and increased data volumes.
[0116] Those skilled in the art will immediately appreciate that many modifications and variations of the apparatus and methods are possible while retaining the teachings of the present invention. Therefore, the above disclosure should be limited only within the limits of the appended claims.
Claims
1. A storage method for a ring buffer, characterized in that: include: Generate data for the first output based on Q input tokens of the large language model LLM; as well as Writing the first output data into the last Q column vectors of the updated first cache tensor buffer matrix; wherein a starting storage address of the first cache tensor buffer is moved by a number Q of input tags based on the LLM to update the first cache tensor buffer; Wherein, the first cache tensor buffer forms a first cache tensor buffer matrix, the updated first cache tensor buffer forms an updated first cache tensor buffer matrix, the first cache tensor buffer matrix includes multiple space segments, each row of the first cache tensor buffer matrix includes C space segments, C is the cache size, wherein the multiple space segments have continuous storage addresses, the starting address of each row of the first cache tensor buffer matrix is continuous with the ending address of the previous row, and the last row of the updated first cache tensor buffer matrix includes at least one additional space segment.
2. The method according to claim 1, characterized in that The number of additional spatial segments is the number Q of input tags of the LLM, the address space of the Q additional spatial segments is equal to Q×S, the number Q of additional spatial segments is less than the cache size, Q and S are positive integers, the dimension of the first cache tensor buffer matrix is (R1, C×S), R1 is the row dimension, and R1 is a positive integer, where S is the step size of the spatial segment.
3. The method according to claim 2, characterized in that Before the starting storage address of the first cache tensor buffer is moved for the first time, the address space of the additional space segment appended to the end of the first cache tensor buffer is at least C×S.
4. The method according to claim 2, characterized in that The starting storage address of the first cache tensor buffer is shifted by increasing the first offset storage address equal to Q×S.
5. The method according to claim 1, characterized in that The spatial segments of the last Q column vectors of the updated first cache tensor buffer matrix are discontinuous, and the data of each segment of the first output is written directly or indirectly to the corresponding segment of the last Q column vectors of the updated first cache tensor buffer matrix based on its storage address.
6. The method according to claim 1, characterized in that Further including: Generate data of a second output according to the Q input tags of the LLM; as well as Writing the second output data into the last Q column vectors of the updated second cache tensor buffer matrix; wherein the starting storage address of the second cache tensor buffer is moved by the number Q of input tags based on the LLM to update the second cache tensor buffer; The second cache tensor buffer forms a second cache tensor buffer matrix, and the updated second cache tensor buffer forms the updated second cache tensor buffer matrix, wherein the second cache tensor buffer matrix includes multiple spatial segments, the multiple spatial segments of the second cache tensor buffer matrix have continuous storage addresses, the starting address of each row of the second cache tensor buffer matrix is continuous with the ending address of the previous row, and the starting storage address of the first cache tensor buffer matrix is adjacent to the ending storage address of the second cache tensor buffer matrix.
7. The method according to claim 6, characterized in that The dimension of the second cache tensor buffer matrix is (R2, C×S), R2 is the row dimension of the second cache tensor buffer matrix, and R2 is a positive integer; the starting storage address of the second cache tensor buffer is moved by increasing the second offset storage address equal to Q×S.
8. The method according to claim 7, characterized in that The spatial segments of the last Q column vectors of the updated second cache tensor buffer matrix are discontinuous, and the data of each segment of the second output is written directly or indirectly to the corresponding segment of the last Q column vectors of the updated second cache tensor buffer matrix based on its storage address.
9. The method according to claim 1, characterized in that Further including: reading the contents of the first cache tensor buffer before the starting storage address of the first cache tensor buffer is moved, The data of the first output is generated based on the content of the first cache tensor buffer and the Q input tags of the LLM.
10. A ring buffer storage method, characterized in that: include: Generate data of multiple outputs according to Q input tokens of a large language model LLM, wherein each output corresponds to a cache tensor buffer, the multiple outputs correspond to multiple cache tensor buffers, and the multiple cache tensor buffers have consecutive addresses and form a connected cache tensor buffer, and the connected cache tensor buffers form a cache tensor buffer matrix; the cache tensor buffer matrix includes multiple space segments, each row of the cache tensor buffer matrix includes C space segments, C is a cache size, and the starting address of each row of the cache tensor buffer matrix is continuous with the ending address of the previous row; and Writing the multiple output data into the last Q column vectors of the updated cache tensor buffer matrix; Wherein a starting storage address of each cache tensor buffer is moved by the number Q of input tags based on the LLM to update the connected cache tensor buffers, wherein the updated connected cache tensor buffers form an updated cache tensor buffer matrix, and a last row of the updated cache tensor buffer matrix includes at least one additional space segment.
11. The method according to claim 10, characterized in that The number of the additional space segments is the number Q of input tags of the LLM, the address space of the Q additional space segments is equal to Q×S, the number Q of the additional space segments is less than the cache size, Q and S are positive integers, where S is the step size of the space segment.
12. The method according to claim 10, characterized in that The starting storage address of each cache tensor buffer is moved by increasing the offset storage address by an amount equal to Q×S, where S is the stride of the spatial segment.
13. The method according to claim 10, characterized in that The spatial segments of the last Q column vectors of the updated cache tensor buffer matrix are discontinuous, and the data of each segment of each output is written directly or indirectly to the corresponding segment of the last Q column vectors of the updated cache tensor buffer matrix based on its storage address.
14. The method according to claim 10, characterized in that Each cache tensor buffer is a row of the cache tensor buffer matrix.
15. A ring buffer storage system, characterized in that: include: Ring buffer; as well as processor; The processor generates data for a first output based on Q input tokens of a large language model LLM, and writes the data into the last Q column vectors of an updated first cache tensor buffer matrix, wherein a starting storage address of the first cache tensor buffer is moved based on the number Q of input tokens of the LLM to update the first cache tensor buffer, wherein the first cache tensor buffer forms a first cache tensor buffer matrix, and the updated first cache tensor buffer forms the updated first cache tensor buffer matrix, the first cache tensor buffer matrix includes multiple space segments, each row of the first cache tensor buffer matrix includes C space segments, C is the cache size, wherein the multiple space segments have continuous storage addresses, the starting address of each row of the first cache tensor buffer matrix is continuous with the end address of the previous row of the first cache tensor buffer matrix, and the last row of the updated first cache tensor buffer matrix includes at least one additional space segment.
16. The system of claim 15, wherein: The number of additional spatial segments is equal to the number Q of input tags of the LLM, the address space of the Q additional spatial segments is equal to Q×S, the number Q of the additional spatial segments is less than the cache size, Q and S are positive integers, the size of the first cache tensor buffer matrix is (R1, C×S), R1 is the row dimension, and R1 is a positive integer, and before the starting storage address of the first cache tensor buffer is moved for the first time, the address space of the additional spatial segment attached to the end of the first cache tensor buffer is at least C×S, where S is the stride of the spatial segment.
17. The system of claim 16, wherein: The starting storage address of the first cache tensor buffer is moved by increasing the first offset storage address equal to Q×S, where S is the stride of the spatial segment.
18. The system of claim 15, wherein: The spatial segments of the last Q column vectors of the updated first cache tensor buffer matrix are discontinuous, and the data of each segment of the first output is written directly or indirectly to the corresponding segment of the last Q column vectors of the updated first cache tensor buffer matrix based on its storage address.
19. The system of claim 15, wherein: The processor generates data of a second output according to Q input tags of the LLM, and writes the data of the second output into the last Q column vectors of an updated second cache tensor buffer matrix, wherein the starting storage address of the second cache tensor buffer is moved based on the number of input tags Q of the LLM to update the second cache tensor buffer, wherein the second cache tensor buffer forms a second cache tensor buffer matrix, and the updated second cache tensor buffer forms an updated second cache tensor buffer matrix, wherein the second cache tensor buffer matrix includes multiple spatial segments, and the multiple spatial segments of the second cache tensor buffer matrix have continuous storage addresses, the starting address of each row of the second cache tensor buffer matrix is continuous with the ending address of the previous row of the second cache tensor buffer matrix, and the starting storage address of the first cache tensor buffer matrix is immediately adjacent to the ending storage address of the second cache tensor buffer matrix.
20. The system of claim 11, wherein: The method reads the contents of the first cache tensor buffer before the starting storage address of the first cache tensor buffer is moved, wherein the data of the first output is generated based on the contents of the first cache tensor buffer and Q input tags of the LLM.