Computing system, multi-round session reasoning method and device and computing device cluster

By preloading the historical KV cache in the host memory and loading it synchronously during accelerator calculation, the recomputation problem caused by limited HBM capacity is solved, and the efficiency and resource utilization of multiple rounds of session reasoning are improved.

CN120406814APending Publication Date: 2025-08-01HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410223741.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-31
Filing Date
2024-02-28
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Generative big model based on Transformer architecture In multiple rounds of session inference, the limited capacity of HBM leads to re-computation of KV cache, resulting in reduced inference efficiency and waste of computing resources.

Method used

By preloading the historical KV cache in the host's memory and synchronously loading the required KV cache when the accelerator performs calculations, the capacity of the external storage device is greater than that of HBM, avoiding recalculation and improving the hit rate of the historical KV cache.

Benefits of technology

It improves the efficiency of multiple rounds of session reasoning, saves computing resources, hides the time overhead of accelerator accessing external storage devices, and improves the synchronization and efficiency of the computing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120406814A_ABST
    Figure CN120406814A_ABST
Patent Text Reader

Abstract

The invention provides a computing system, a multi-round session reasoning method and device and a computing device cluster, and belongs to the technical field of cloud computing. In the system, an external storage device stores a historical key value cache of a completed session, and the capacity of the external storage device is far greater than the capacity of an HBM in an accelerator, so that the hit rate of the historical key value cache can be improved, and key value cache re-calculation is avoided; moreover, when the accelerator processes the session, the host can preload the historical key value cache of the session to be processed in the task queue from the external storage device to the memory of the host, and when the accelerator carries out the ith layer calculation on the session, the accelerator can preload the historical key value cache required by the (i + 1) th layer calculation from the memory of the host, the calculation process and the data loading process are synchronously performed, and the calculation process does not need to wait for data loading completion, so that the time overhead of accessing the external storage device by the accelerator can be hidden, and the efficiency of multi-round session reasoning is improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of a Chinese patent application with the application number 202410137574.7 and the invention title "A Method, Device and Other Equipment for Data Processing" filed on January 31, 2024, the entire content of which is incorporated herein by reference. Technical Field

[0002] This application relates to the field of cloud computing technology, and particularly relates to a computing system, a multi-round conversation reasoning method, a device and a computing device cluster. Background Art

[0003] The generative large model based on the Transformer architecture can be applied to the scenario of multi-round conversation reasoning. In a single round of conversation, the generative large model can process the input token sequence token[1:s] (i.e., the question prompt), infer the (s + 1)-th token token[s + 1] to obtain a new token sequence token[1:s + 1], then input token[1:s + 1] into the model to infer token[s + 2], and so on until a complete token sequence is obtained or an inference termination symbol appears. The generative large model outputs the inferred token sequence, that is, outputs the answer, thus completing a single round of conversation, where s is an integer greater than 1. During the process of inferring a token, the generative large model generates corresponding intermediate data key and value when processing any input token, and the generated intermediate data is also used in the process of inferring subsequent tokens. In addition, during multi-round conversations, each round of conversation uses the tokens input and generated in the previous few rounds, and thus also uses the intermediate data corresponding to the tokens input and generated in the previous few rounds. Therefore, in order to save computing resources and improve inference efficiency, the intermediate data corresponding to the tokens input and generated in the previous few rounds is stored in the high-bandwidth memory (HBM) inside the accelerator, that is, the HBM inside the accelerator is used to store the historical key-value cache (KV cache) of the conversation, so that the historical KV cache of the previous few rounds can be directly reused for inference during subsequent conversations without having to perform recalculation.

[0004] However, in the above method, relying on the HBM to store the KV cache, and the capacity of the HBM is limited. When facing the problem of insufficient available storage space in the HBM, the accelerator will delete the KV cache stored in the HBM according to business requirements, resulting in the inability to reuse the historical KV cache in subsequent conversations. The generative large model faces a huge recalculation overhead for the historical KV cache, still causing a reduction in inference efficiency and waste of computing resources. Summary of the Invention

[0005] The embodiments of the present application provide a computing system, a multi-round conversation reasoning method, a device, and a computing device cluster, which can avoid recalculating the KV cache and improve the efficiency of multi-round conversation reasoning. The technical solution is as follows.

[0006] In a first aspect, a computing system is provided. The computing system includes a host, an accelerator of the host, and an external storage device;

[0007] Among them, the external storage device is used to store the historical KV cache of the completed conversations; the host is used to preload the historical KV cache of the conversations to be processed in the task queue from the external storage device into the memory of the host while the accelerator processes the conversations, and the accelerator is used to preload the historical KV cache required for the (i + 1)-th layer calculation from the memory of the host into the accelerator while performing the i-th layer calculation of the generative large model for the conversations.

[0008] In the above system, since the capacity of the external storage device is much larger than the capacity of the HBM in the accelerator, the external storage device can store a large amount of historical KV cache, thereby increasing the hit rate of the historical KV cache, avoiding recalculating the KV cache, improving the efficiency of multi-round conversation reasoning, and saving computing resources; and since the calculation process and the data loading process are carried out synchronously, the calculation process does not need to wait for the data loading to be completed, thereby being able to hide the time overhead of the accelerator accessing the external storage device and improving the efficiency of multi-round conversation reasoning.

[0009] In some embodiments, the historical KV cache of the input of the first conversation at the (i + 1)-th layer has nothing to do with the positional encoding of the input of the first conversation.

[0010] Here, "has nothing to do with" means that the positional encoding of the input of the first conversation has no influence on the historical KV cache of the first conversation.

[0011] In some embodiments, the number of tokens included in the input of the first conversation is greater than or equal to the number of tokens that the context window of the generative large model can accommodate;

[0012] The accelerator is used to: when performing the i-th layer calculation for the first conversation, load the historical KV cache of the first sub-input at the (i + 1)-th layer from the memory, where the first sub-input is a sub-input of the input of the first conversation, and the number of tokens included in the first sub-input is less than the number of tokens that the context window can accommodate.

[0013] In the above system, the stored KV cache is decoupled from the positional encoding. When the context window overflows, the change in the positional encoding of the input of the generative large model does not cause the stored KV cache to become invalid. Therefore, the accelerator can reuse the historical KV cache of the cropped input, thus avoiding the recalculation of the KV cache due to the overflow of the context window, improving the inference efficiency, and saving computing resources.

[0014] In some embodiments, the accelerator is further configured to:

[0015] Embed the positional encoding of the first sub-input into the historical KV cache of the first sub-input at the (i + 1)-th layer; perform the (i + 1)-th layer calculation on the first session based on the input of the first session and the historical KV cache of the first sub-input at the (i + 1)-th layer after embedding the positional encoding.

[0016] Wherein, the positional encoding is relative positional encoding (RPE).

[0017] In some embodiments, the accelerator includes a first buffer and a second buffer. The first buffer is used to store the historical KV cache of the first session, and the second buffer is used to store the historical KV cache of the input of the second session at the first layer in the task queue. The second session is the first pending session after the first session in the task queue;

[0018] The accelerator is configured to: when performing the i-th layer calculation on the first session, load the historical KV cache of the input of the first session at the (i + 1)-th layer from the memory to the first buffer; when performing the N-th layer calculation on the first session, load the historical KV cache of the input of the second session at the first layer from the memory to the second buffer.

[0019] In the above system, the HBM of the accelerator includes a first buffer and a second buffer. The first buffer is used to store the KV cache of the session currently being processed by the accelerator, and the second buffer is used to store the KV cache of the session to be processed by the accelerator. While the accelerator processes the first session based on the KV cache in the first buffer, it loads the KV cache required for the first-layer calculation of the second session into the second buffer. When the accelerator finishes processing the first session, it processes the second session based on the KV cache in the second buffer. At this time, the second buffer becomes the execution buffer. While processing the second session, the accelerator releases the first buffer through an asynchronous thread, enabling the computing thread of the accelerator to immediately start processing the second session after finishing processing the first session without waiting for the release of the first buffer to complete, thereby hiding the time gap caused by loading the data required for the next session between the two sessions and improving the inference efficiency.

[0020] In some embodiments, the accelerator is further configured to: when performing the (i + 1)-th layer calculation on the first session, write the KV cache generated during the i-th layer calculation of the first session into the memory.

[0021] In the above system, while the accelerator performs the (i + 1)-th layer calculation, writing the KV cache generated during the i-th layer calculation back to the host's memory can hide the time gap caused by writing back the data generated during the previous session between two adjacent rounds of session inferences, and can improve the inference efficiency.

[0022] In some embodiments, the accelerator includes a third buffer, and the third buffer is used to store the KV cache generated by the generative large model in processing the first session;

[0023] The accelerator of the host is configured to: if when finishing the N-th layer calculation of the first session, there is KV cache in the KV cache generated by the generative large model in processing the first session that has not been written into the memory, write the unwritten KV cache into the third buffer; when processing the second session in the task queue through the generative large model, write the KV cache in the third buffer into the memory.

[0024] In the above system, the HBM of the accelerator further includes a third buffer. After the accelerator finishes processing the first session, it writes the KV cache that was generated during the processing of the first session and has not been written back to the memory from the first buffer to the third buffer. Since the data copy efficiency between the first buffer and the third buffer is relatively high, after the accelerator finishes processing the first session, it can quickly release the first buffer, thereby being able to load the KV cache required for subsequent processing into the first buffer without waiting for all the KV cache generated during the processing of the first session to be written back to the memory before releasing the first buffer, hiding the time gap caused by waiting for the data of the previous session to be written back between two sessions and improving the inference efficiency.

[0025] In some embodiments, the host is configured to: if the historical KV cache of the first pending session in the task queue does not exist in the memory, load the historical KV cache of the first pending session from the external storage device into the memory, where the first pending session is the first number of pending sessions at the head of the task queue.

[0026] In the above system, the host can preload the historical KV cache of the pending sessions in the task queue from the external storage device into the host's memory while the accelerator is processing the sessions. Since the calculation process and the data loading process are synchronized, the calculation process does not need to wait for the data loading to complete, thereby being able to hide the time overhead of the accelerator accessing the external storage device and improving the efficiency of multi-round session inference.

[0027] In some embodiments, the number of session rounds included in the first pending session is determined based on the capacity of the memory.

[0028] In some embodiments, the host is configured to: if the available storage space in the memory is sufficient, load the historical KV cache of the first pending session from the external storage device into the memory; if the available storage space in the memory is insufficient, write the historical KV cache of the second pending session in the task queue from the memory to the external storage device and load the historical KV cache of the first pending session from the external storage device into the memory, where the second pending session is the session in the task queue other than the first pending session.

[0029] In some embodiments, the host is further configured to: if the available storage space in the external storage device is insufficient, delete the historical KV cache of the third pending session from the external storage device, where the third pending session is the second number of pending sessions at the end of the task queue.

[0030] Second aspect, a multi-round conversation reasoning method is provided, which is applied to a computing system. The computing system includes a host, an accelerator of the host, and an external storage device. The external storage device is used to store the historical key-value cache (KV cache) of the conversation. The host is used to load the historical KV cache of the conversation to be processed in the task queue from the external storage device into the memory of the host. The method includes:

[0031] Through the accelerator, based on the task queue, the generative large model is used to reason about the first conversation in the task queue. The first conversation is one round of the multi-round conversation. The generative large model includes N layers. Among them, when the accelerator performs the i-th layer calculation on the first conversation, the historical KV cache of the input of the first conversation in the (i + 1)-th layer is loaded from the memory to the accelerator, where i is an integer greater than or equal to 1 and less than N.

[0032] In some embodiments, the historical KV cache of the input of the first conversation in the (i + 1)-th layer is independent of the positional encoding of the input of the first conversation.

[0033] In some embodiments, the number of tokens included in the input of the first conversation is greater than or equal to the number of tokens that can be accommodated by the context window of the generative large model;

[0034] When performing the i-th layer calculation on the first conversation, loading the historical KV cache of the input of the first conversation in the (i + 1)-th layer from the memory includes: when performing the i-th layer calculation on the first conversation, loading the historical KV cache of the first sub-input in the (i + 1)-th layer from the memory. The first sub-input is a sub-input of the input of the first conversation, and the number of tokens included in the first sub-input is less than the number of tokens that can be accommodated by the context window.

[0035] In some embodiments, processing the first conversation in the task queue by the generative large model includes:

[0036] Embedding the positional encoding of the first sub-input into the historical KV cache of the first sub-input in the (i + 1)-th layer; based on the input of the first conversation and the historical KV cache of the first sub-input in the (i + 1)-th layer after embedding the positional encoding, performing the (i + 1)-th layer calculation on the first conversation.

[0037] In some embodiments, the accelerator includes a first buffer and a second buffer. The first buffer is used to store the historical KV cache of the first conversation, and the second buffer is used to store the historical KV cache of the input of the second conversation in the task queue in the first layer. The second conversation is the first conversation to be processed after the first conversation in the task queue;

[0038] When performing the i-th layer calculation on the first session, load the historical KV cache of the input of the first session at the (i + 1)-th layer from the memory, including: when the accelerator performs the i-th layer calculation on the first session, load the historical KV cache of the input of the first session at the (i + 1)-th layer from the memory to the first buffer through the accelerator.

[0039] In some embodiments, the method further includes:

[0040] When the accelerator performs the N-th layer calculation on the first session, load the historical KV cache of the input of the second session at the first layer from the memory to the second buffer through the accelerator.

[0041] In some embodiments, the method further includes:

[0042] When the accelerator performs the (i + 1)-th layer calculation on the first session, write the KV cache generated by performing the i-th layer calculation on the first session into the memory through the accelerator.

[0043] In some embodiments, the accelerator includes a third buffer for storing the KV cache generated by the generative large model processing the first session;

[0044] When performing the (i + 1)-th layer calculation on the first session, writing the KV cache generated by performing the i-th layer calculation on the first session into the memory includes: if there is still KV cache in the KV cache generated by the generative large model processing the first session that has not been written into the memory when the accelerator completes the N-th layer calculation on the first session, write the unwritten KV cache into the third buffer through the accelerator; when the accelerator processes the second session in the task queue through the generative large model, write the KV cache in the third buffer into the memory through the accelerator.

[0045] In some embodiments, the method further includes: if the historical KV cache of the first pending session in the task queue does not exist in the memory, load the historical KV cache of the first pending session from the external storage device to the memory through the host, where the first pending session is the first number of pending sessions at the head of the task queue.

[0046] In some embodiments, the number of conversation rounds included in the first pending session is determined based on the capacity of the memory.

[0047] In some embodiments, if the historical KV cache of the first pending session in the task queue does not exist in the memory, the host loads the historical KV cache of the first pending session from the external storage device into the memory, including:

[0048] If the available storage space in the memory is sufficient, the host loads the historical KV cache of the first pending session from the external storage device into the memory; if the available storage space in the memory is insufficient, the host writes the historical KV cache of the second pending session in the task queue from the memory to the external storage device, and loads the historical KV cache of the first pending session from the external storage device into the memory, where the second pending session is a session other than the first pending session in the task queue.

[0049] In some embodiments, the method further includes:

[0050] If the available storage space in the external storage device is insufficient, the host deletes the historical KV cache of the third pending session from the external storage device, where the third pending session is the second number of pending sessions at the end of the task queue.

[0051] In a third aspect, a multi-round session inference method is provided, which is applied to an accelerator of a host in a computing system. The computing system further includes a host and an external storage device. The external storage device is used to store the historical key-value cache (KV cache) of sessions, and the host is used to load the historical KV cache of the pending sessions in the task queue from the external storage device into the memory of the host. The method includes:

[0052] Based on the task queue, the generative large model infers the first session in the task queue. The first session is one round of the multi-round session. The generative large model includes N layers. When performing the i-th layer calculation on the first session, the historical KV cache of the input of the first session at the (i + 1)-th layer is loaded from the memory, where i is an integer greater than or equal to 1 and less than N.

[0053] In some embodiments, the historical KV cache of the input of the first session at the (i + 1)-th layer is independent of the positional encoding of the input of the first session.

[0054] In some embodiments, the number of tokens included in the input of the first session is greater than or equal to the number of tokens that can be accommodated by the context window of the generative large model;

[0055] When performing the i-th layer calculation on the first session, load the historical KV cache of the input of the first session at the (i + 1)-th layer from the memory, including: when performing the i-th layer calculation on the first session, load the historical KV cache of the first sub-input at the (i + 1)-th layer from the memory, where the first sub-input is a sub-input of the input of the first session and the number of tokens included in the first sub-input is less than the number of tokens that the context window can accommodate.

[0056] In some embodiments, processing the first session in the task queue by the generative large model includes:

[0057] Embed the position encoding of the first sub-input into the historical KV cache of the first sub-input at the (i + 1)-th layer; based on the input of the first session and the historical KV cache of the first sub-input with the embedded position encoding at the (i + 1)-th layer, perform the (i + 1)-th layer calculation on the first session.

[0058] In some embodiments, the accelerator includes a first buffer and a second buffer. The first buffer is used to store the historical KV cache of the first session, and the second buffer is used to store the historical KV cache of the input of the second session in the task queue at the first layer, where the second session is the first session to be processed after the first session in the task queue;

[0059] When performing the i-th layer calculation on the first session, loading the historical KV cache of the input of the first session at the (i + 1)-th layer from the memory includes: when performing the i-th layer calculation on the first session, loading the historical KV cache of the input of the first session at the (i + 1)-th layer from the memory to the first buffer.

[0060] In some embodiments, the method further includes: when performing the N-th layer calculation on the first session, loading the historical KV cache of the input of the second session at the first layer from the memory to the second buffer.

[0061] In some embodiments, the method further includes: when performing the (i + 1)-th layer calculation on the first session, writing the KV cache generated by performing the i-th layer calculation on the first session into the memory.

[0062] In some embodiments, the accelerator includes a third buffer, and the third buffer is used to store the KV cache generated by the generative large model processing the first session;

[0063] When performing the (i + 1)-th layer calculation on the first session, the KV cache generated by performing the i-th layer calculation on the first session is written into the memory, including: when the N-th layer calculation on the first session is completed, if there is a KV cache in the KV cache generated by the generative large model processing the first session that has not been written into the memory, then the KV cache that has not been written into the memory is written into the third buffer; when processing the second session in the task queue by the generative large model, the KV cache in the third buffer is written into the memory.

[0064] In a fourth aspect, a multi-round session inference method is provided, which is applied to a host in a computing system. The computing system further includes an accelerator of the host and an external storage device. The external storage device is used to store the historical key-value cache KV cache of the session, and the accelerator is used to perform inference on the first session in the task queue through a generative large model based on the task queue. The first session is one round of a multi-round session. The generative large model includes N layers. Among them, when performing the i-th layer calculation on the first session, the historical KV cache of the input of the first session at the (i + 1)-th layer is loaded from the memory, where i is an integer greater than or equal to 1 and less than N;

[0065] The method includes:

[0066] Loading the historical KV cache of the session to be processed in the task queue from the external storage device into the memory of the host; storing the historical KV cache of the session to be processed.

[0067] In some embodiments, loading the historical KV cache of the session to be processed in the task queue from the external storage device into the memory of the host includes:

[0068] If the historical KV cache of the first session to be processed in the task queue does not exist in the memory, then the historical KV cache of the first session to be processed is loaded from the external storage device into the memory. The first session to be processed is the first number of sessions to be processed at the head of the task queue.

[0069] In some embodiments, the number of session rounds included in the first session to be processed is determined based on the capacity of the memory.

[0070] In some embodiments, if the historical KV cache of the first session to be processed in the task queue does not exist in the memory, then the host loads the historical KV cache of the first session to be processed from the external storage device into the memory, including:

[0071] If the available storage space of the memory is sufficient, load the historical KV cache of the first session to be processed from the external storage device into the memory; if the available storage space of the memory is insufficient, write the historical KV cache of the second session to be processed in the task queue from the memory to the external storage device, and load the historical KV cache of the first session to be processed from the external storage device into the memory, where the second session to be processed is the session in the task queue other than the first session to be processed.

[0072] In some embodiments, the method further includes:

[0073] If the available storage space of the external storage device is insufficient, delete the historical KV cache of the third session to be processed from the external storage device, where the third session to be processed is the second number of sessions to be processed at the end of the task queue.

[0074] In a fifth aspect, there is provided a multi-round session inference device, which is applied to an accelerator of a host in a computing system. The device includes at least one functional module, and the at least one functional module is configured to execute the multi-round session inference method provided in the foregoing third aspect or any possible implementation manner of the third aspect.

[0075] In a sixth aspect, there is provided a multi-round session inference device, which is applied to a host in a computing system. The device includes at least one functional module, and the at least one functional module is configured to execute the multi-round session inference method provided in the foregoing fourth aspect or any possible implementation manner of the fourth aspect.

[0076] In a seventh aspect, there is provided an accelerator, which includes a computing core and a memory. The accelerator is configured to execute the multi-round session inference method provided in the foregoing third aspect or any possible implementation manner of the third aspect. The memory is used to store computing data, and the computing core is used to perform computing operations on the computing data stored in the memory.

[0077] In an eighth aspect, there is provided a host, which includes a processor and a memory. The host is configured to execute the multi-round session inference method provided in the foregoing fourth aspect or any possible implementation manner of the fourth aspect.

[0078] In a ninth aspect, there is provided a computing device cluster, which includes at least one computing device. Each computing device includes a host, an accelerator of the host, and an external storage device is located in at least one computing device. The external storage device is used to store the historical KV cache of the session, and the host is used to load the historical KV cache of the session to be processed in the task queue from the external storage device into the memory of the host;

[0079] The host in the computing device cluster is used to execute the multi-round session reasoning method provided in the foregoing fourth aspect or any possible implementation manner of the fourth aspect;

[0080] The accelerator in the computing device cluster is used to execute the multi-round session reasoning method provided in the foregoing third aspect or any possible implementation manner of the third aspect.

[0081] In a tenth aspect, there is provided a computer program product including instructions, which, when run on a computing device cluster, cause the computing device cluster to execute the multi-round session reasoning method provided in the foregoing third aspect or any possible implementation manner of the third aspect, or execute the multi-round session reasoning method provided in the foregoing fourth aspect or any possible implementation manner of the fourth aspect.

[0082] In an eleventh aspect, there is provided a computer-readable storage medium including computer program instructions, which, when executed by a computing device cluster, cause the computing device cluster to execute the multi-round session reasoning method provided in the foregoing third aspect or any possible implementation manner of the third aspect, or execute the multi-round session reasoning method provided in the foregoing fourth aspect or any possible implementation manner of the fourth aspect.

[0083] Based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Figure 1 is a schematic structural diagram of a computing system provided by an embodiment of the present application;

[0085] Figure 2 is a schematic functional diagram of a computing system provided by an embodiment of the present application;

[0086] Figure 3 is a flowchart of a multi-round session reasoning method provided by an embodiment of the present application;

[0087] Figure 4 is a schematic diagram of directly trimming the KV cache after the context window overflows provided by an embodiment of the present application;

[0088] Figure 5 is a schematic diagram of performing reasoning based on a KV cache decoupled from positional encoding provided by an embodiment of the present application;

[0089] Figure 6 is a schematic diagram of parallelizing the calculation process and the data loading process in a multi-round session reasoning method provided by an embodiment of the present application;

[0090] Figure 7It is a schematic diagram of the parallel computing process and data write-back process in a multi-round conversation inference method provided by an embodiment of the present application;

[0091] Figure 8 It is a schematic diagram of the data migration process between the memory of the host and the external storage device during the multi-round conversation inference process provided by an embodiment of the present application;

[0092] Figure 9 It is a schematic diagram of the structure of a multi-round conversation inference device provided by an embodiment of the present application;

[0093] Figure 10 It is a schematic diagram of the structure of a multi-round conversation inference device provided by an embodiment of the present application;

[0094] Figure 11 It is a schematic diagram of the structure of a computing device provided by an embodiment of the present application;

[0095] Figure 12 It is a schematic diagram of a computing device cluster provided by an embodiment of the present application;

[0096] Figure 13 It is a schematic diagram of a possible implementation manner of a computing device cluster provided by an embodiment of the present application. Detailed implementation manners

[0097] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0098] First, the present application relates to the application of a generative large model based on the Transformer architecture in multi-round conversation inference. To facilitate the understanding of the content of the embodiments of the present application, the following explains several technical terms related to the embodiments of the present application.

[0099] Token: The smallest semantic unit represented by a vector. For example, the smallest semantic unit can be a character, a Chinese character, a word, or a phrase, etc. A sequence of tokens forms the input to the generative large model. In a multi-round conversation reasoning scenario, the sequence of tokens input to the generative large model is also the problem prompt of the conversation. The generative large model performs reasoning on the input sequence of tokens. Through multiple iterations, the subsequent tokens of the problem prompt are obtained one by one until a complete subsequent token sequence is obtained or a reasoning termination symbol appears. The generative large model outputs the generated token sequence, and the generated token sequence is also the answer to the problem prompt. For example, if the sequence of tokens input to the generative large model is X[1:s], where s is an integer greater than 1, the generative large model performs the first iteration based on the token sequence X[1:s] to obtain the (s + 1)-th token token[s + 1]; the generative large model performs the second iteration based on X[1:s] and token[s + 1] to obtain the (s + 2)-th token token[s + 2]; and so on until a complete token sequence is obtained or a reasoning termination symbol appears. The generative large model outputs the inferred token sequence, which is also the output answer, thus completing one round of conversation reasoning.

[0100] Generative large model based on the Transformer architecture: It includes N Transformer layers, where N is an integer greater than 1. Each Transformer layer includes a self-attention mechanism and a feedforward neural network (FFN).

[0101] One round of conversation reasoning: In each iteration, the N Transformer layers of the generative large model sequentially process the sequence of tokens input in this iteration to obtain a token. The process of processing the sequence of tokens by the i-th layer of the generative large model includes: projecting each token in the input of this iteration respectively with the model parameters corresponding to the self-attention mechanism in the i-th layer to generate the intermediate data key and value corresponding to each token; performing a non-linear transformation on the key and value. The non-linear transformation can be, for example, a residual connection and normalization, and passing the result of the non-linear transformation to the feedforward neural network in the i-th layer; the feedforward neural network in the i-th layer processes the result of the non-linear transformation to obtain the processing result of the i-th layer, and passes the processing result of the i-th layer to the (i + 1)-th layer for processing. Here, i is an integer greater than or equal to 1 and less than N. The above process of processing the sequence of tokens by the i-th layer of the generative large model to obtain the processing result of the i-th layer is also the calculation of the i-th layer.

[0102] Multi-round Conversation Reasoning: To maintain the semantic consistency of the conversation context and ensure an accurate understanding of the input in the current conversation, when generating an answer for the current conversation, the generative large model will refer to the inputs and outputs of the historical conversations with the same object. That is, the input for one round of conversation in multi-round conversations is: the sequence of tokens input in the historical conversation, the sequence of tokens output in the historical conversation, and the sequence of tokens input by the object in this round of conversation.

[0103] Key-Value Cache (KV cache): In one iteration of the generative large model (i.e., the process of generating one token), when the generative large model processes any token, corresponding intermediate data keys and values will be generated, and the generated intermediate data will also be used in subsequent iterations. In addition, during multi-round conversations, each round of conversation will use the sequences of tokens input and output in the historical conversation, and thus will also use the intermediate data corresponding to the tokens input and output in the historical conversation. Therefore, to save computing resources and improve inference efficiency, the intermediate data corresponding to the tokens input and output in the historical conversation is stored to form the historical KV cache of the conversation, so that the historical KV cache of the conversation can be directly reused for inference during conversation reasoning without having to recalculate.

[0104] Prefill Phase: This prefill phase is also the first iteration in the process of one round of conversation reasoning. In the prefill phase, the generative large model performs the first iteration based on the sequence of tokens X[1:s] input in this round of conversation to obtain token[s + 1], and stores the keys and values corresponding to each token in X[1:s] obtained in this iteration to form KV cache[1:s].

[0105] Decoding Phase: This decoding phase is also the second iteration to the last iteration in the process of one round of conversation reasoning. In any iteration of the decoding phase, the generative large model calculates the keys and values corresponding to the token obtained in the previous iteration, reads the stored historical KV cache, and based on the keys and values calculated in this iteration and the historical KV cache, obtains the token obtained in this iteration, and stores the keys and values calculated in this iteration, that is, updates the historical KV cache.

[0106] Context Window: The number of tokens that the generative large model can process simultaneously.

[0107] Accelerator Stream: An abstract unit for the accelerator to execute parallel computing tasks.

[0108] Execution buffer: The area in high-bandwidth memory (HBM) used to store the KV Cache that the accelerator inference computing stream can currently process.

[0109] The above introduced several technical terms related to the embodiments of the present application. Next, the implementation environment of the embodiments of the present application will be introduced.

[0110] Figure 1 It is a schematic structural diagram of a computing system provided by the embodiments of the present application. As Figure 1 shown, the computing system includes a host 101, an accelerator 102 of the host, and an external storage device 103. Among them, the host 101, the accelerator 102, and the external storage device 103 communicate with each other through a wired network or a wireless network.

[0111] Among them, the computing system can be deployed on a computing device cluster, and the computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0112] Among them, the external storage device 103 is, for example, a solid state disk (SSD) or a hard disk drive (HDD), or is, for example, an object storage service (OBS), an elastic volume service (EVS), or a scalable file service (SFS) and other cloud storage services. The embodiments of the present application do not limit the type of the external storage device 103. Schematically, in a multi-level storage system, the external storage device 103 is used to store the historical KV cache of the session.

[0113] Among them, the host 101 is, for example, the host of a computing device, or for example, the main computing device in a computing device cluster for controlling and managing other computing devices. The memory of the host 101 is, for example, dynamic random access memory (DRAM). Schematically, a task scheduler runs on the host 101, and the task scheduler maintains a task queue, which is used to indicate the pending sessions of the accelerator; the memory of the host 101 is used to store the historical KV cache of the pending sessions in the task queue; the host 101 is used to load the historical KV cache of the pending sessions in the task queue from the external storage device 101 into the memory of the host 101 based on the task queue.

[0114] Among them, the accelerator 102 is, for example, a general central processing unit (CPU), a graphics processing unit (GPU), a switching module processor unit (SMPU), a network processor unit (NPU), a microprocessor, or for example, one or more integrated circuits for implementing the solution of this application. For example, an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. The embodiments of this application do not limit the type of the accelerator. Schematically, the HBM in the accelerator 102 is used to store the historical KV cache of the first session. The accelerator 102 can perform inference on the first session in the task queue through a generative large model, that is, perform the calculation of the i-th layer of the generative large model, so as to generate an answer to the first session. Among them, the first session is one round of a multi-round session. When the accelerator 102 performs the i-th layer calculation, it reuses the historical KV cache of the input of the first session in the i-th layer in the HBM of the accelerator 102, and when performing the i-th layer calculation, it loads the historical KV cache of the first session in the (i + 1)-th layer from the memory into the HBM.

[0115] In some embodiments, each computing device in a computing device cluster includes a host 101, an accelerator 102, and an external storage device 103, that is, the storage device configured by each computing device is used as the external storage device 103; in other embodiments, each computing device in the computing device cluster includes the host 101 and the accelerator 102, and at least one computing device with storage capabilities in the computing cluster acts as the external storage device 103. The embodiments of the present application do not limit this.

[0116] Exemplarily, Figure 2 is a functional schematic diagram of a computing system provided by an embodiment of the present application. As Figure 2 shown, the computing system includes a control system and a multi-level storage system. The multi-level storage system includes the memory of the host 101, the HBM in the accelerator 102, and the external storage device 103. Among them, the control system runs on the host 101 and the accelerator 102. The control system includes a task scheduler and a key-value cache management unit. The task scheduler maintains a task queue, and the task queue is used to indicate the sessions to be processed by the accelerator; the key-value cache management unit is used to control the data migration between the memory of the host 101 and the external storage device 103. Specifically, the key-value cache management subunit includes a key-value cache pulling subunit and a key-value cache placing subunit. The key-value cache pulling subunit is used to pre-pull the historical KV cache of the sessions to be processed in the task queue into the memory of the host 101, and to load the KV cache in the memory of the host 101 into the HBM of the accelerator; the key-value cache placing subunit is used to evict the KV cache from the memory of the host 101 to the external storage device 103 when the available storage space in the memory of the host 101 is insufficient, and to store (write back) the KV cache generated by the accelerator processing the session into the memory of the host. It should be noted that Figure 2 the division of the functional modules of the computing system shown is only exemplary, and the embodiments of the present application do not limit the division method of the functional modules of the computing system.

[0117] In some embodiments, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is usually a Transmission Control Protocol / Internet Protocol (TCP / IP) network and an RDMA network in a data center network, such as an RDMA over Converged Ethernet (RoCE) network based on aggregated Ethernet, an InfiniBand (IB) network, etc., and this is not limited. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.

[0118] The following introduces a multi-round conversation reasoning method provided by an embodiment of the present application. This method is applied to the above computing system. In this method, the external storage device stores the historical KV cache of the completed conversations. Since the capacity of the external storage device is much larger than the capacity of the HBM in the accelerator, the external storage device can save a large amount of historical KV cache, thereby increasing the hit rate of the historical KV cache, avoiding the recalculation of the KV cache, improving the efficiency of multi-round conversation reasoning, and saving computing resources. At the same time, while the accelerator is processing a conversation, the host can pre-load the historical KV cache of the conversation to be processed in the task queue from the external storage device into the memory of the host. While the accelerator is performing the i-th layer calculation of the generative large model for a conversation, it can pre-load the historical KV cache required for the (i + 1)-th layer calculation from the memory of the host into the accelerator. Since the calculation process and the data loading process are synchronized, the calculation process does not need to wait for the data loading to complete, thus hiding the time overhead of the accelerator accessing the external storage device and improving the efficiency of multi-round conversation reasoning.

[0119] The above method involves the data migration process between the memory of the accelerator and the host during the multi-round conversation reasoning process, as well as the data migration process between the memory of the host and the external storage device. The following will introduce these two data migration processes separately.

[0120] First, the data migration process between the memory of the accelerator and the host is introduced. Figure 3 is a flowchart of a multi-round conversation reasoning method provided by an embodiment of the present application. As Figure 3 shown, this method is applied to a computing system, which includes a host, an accelerator of the host, and an external storage device. This method includes the following steps 301 to 308.

[0121] Step 301: The accelerator loads the historical KV cache of the input of the first conversation at the first layer of the generative large model from the memory of the host into the first buffer of the accelerator. The generative large model includes N layers, and N is an integer greater than 1.

[0122] Among them, the first session is one round in a multi-round session initiated by a first object. The input of the first session includes the token sequence input in the historical session of the first session, the token sequence output in the historical session, and the token sequence input by the object in the first session. Among them, multi-round sessions initiated by the same object have the same session window identifier. The historical session of the first session is also the completed session that has the same session window identifier as the first session. For example, the first object initiates three rounds of sessions. The first round of session and the second round of session are completed. The first session is the third round of session. Among them, the token sequence input in the first round of session is Q1, the token sequence output is A1, the token sequence input in the second round of session is Q2, the token sequence output is A2, and the token sequence input in the third round of session is Q3. Then the input of the first session is [Q1 A1 Q2 A2 Q3].

[0123] Among them, the historical KV cache of the first session at layer 1 is also the KV cache associated with layer 1 in the historical KV cache of the first session. This first buffer is the area in the HBM of the accelerator used to store the KV cache that the accelerator can currently process, that is, the execution buffer of the accelerator.

[0124] Among them, before the accelerator executes the layer 1 calculation of the first iteration of the first session, the accelerator starts a data reading thread. Through this data reading thread, it loads the historical KV cache of the input of the first session at layer 1 from the memory of the host to this first buffer. In the above method, before the layer 1 calculation occurs, by means of the data reading thread, the data required for this layer of calculation is made ready in the first buffer of the accelerator, which can avoid the time gap caused by loading data from the memory of the host, thereby improving the inference efficiency.

[0125] Among them, the historical KV cache is a multi-dimensional vector with variable length. The historical KV cache can be stored at different granularities. For example, in some embodiments, the historical KV cache is stored at the granularity of one round of conversation, that is, the intermediate data keys and values obtained by processing multiple conversations through the generative large model are stored as multiple KV caches according to the conversation, and each KV cache has a conversation identifier. For another example, in some other embodiments, the historical KV cache is stored at the granularity of one layer of the generative large model, that is, the intermediate data keys and values obtained through multiple layers of the generative large model are stored as multiple KV caches according to the layer, and each KV cache has a conversation identifier and a layer identifier. For yet another example, in some other embodiments, the historical KV cache is stored at the granularity of a group of multi-round conversations, that is, the intermediate data keys and values obtained by processing multiple rounds of conversations initiated by the same object through the generative large model are stored as one KV cache. It should be noted that the above description of the storage granularity of the historical KV cache is only exemplary, and the embodiments of the present application do not limit the storage granularity of the KV cache.

[0126] Among them, the historical KV cache of the input of the first conversation in the first layer has nothing to do with the position encoding of the input of the first conversation. Here, "nothing to do with" means that the position encoding of the input of the first conversation has no influence on the historical KV cache of the first conversation.

[0127] In some embodiments, the number of tokens included in the input of the first conversation is greater than the number of tokens that the context window of the generative large model can accommodate; the accelerator loads the historical KV cache of the first sub-input in the first layer from the memory of the host, and the first sub-input is a sub-input of the input of the first conversation, and the number of tokens included in the first sub-input is less than the number of tokens that the context window can accommodate. This process can be understood as that when the input of the first conversation causes the context window to overflow, the input of the first conversation is clipped to discard the tokens at the front of the input of the first conversation, and the first sub-input of the first conversation is used as the input of the generative large model. Correspondingly, the accelerator loads the historical KV cache corresponding to the first sub-input from the memory of the host, that is, directly clips the KV cache after the context window overflows. The following is an example of Figure 4 the above loading process. Figure 4 is a schematic diagram of directly clipping the KV cache after the context window overflows provided by the embodiments of the present application, as Figure 4As shown, the historical KV cache of the input of the first session at layer 1 is stored in the memory of the host, and the position encoding of the input of the first session is [0:2048]; when the input of the first session causes a context window overflow, the accelerator loads the historical KV cache of the first sub-input of the first session at layer 1 from the memory of the host; the position encoding of the first sub-input is [0:1536], and the accelerator embeds the position encoding of the first sub-input into the historical KV cache of the first sub-input at layer 1, and performs subsequent inference based on the historical KV cache after embedding the position encoding. Among them, the accelerator embedding the position encoding into the historical KV cache means embedding the position encoding into the key in the historical KV cache.

[0128] In the above embodiment, the stored KV cache is decoupled from the position encoding. When the context window overflows, the change of the position encoding of the input of the generative large model will not cause the stored KV cache to become invalid. Therefore, the accelerator can reuse the historical KV cache of the cropped input, thus avoiding the recalculation of the KV cache caused by the context window overflow, improving the inference efficiency and saving computing resources.

[0129] It should be noted that the historical KV cache of the first session is stored in an external storage device. Before the accelerator performs inference on the first session through the generative large model, the host pre-loads the historical KV cache of the first session from the external storage device into the memory of the host, and then the accelerator loads the historical KV cache of the input of the first session at layer 1 from the memory of the host. The data migration process between the external storage device and the memory of the host will be introduced in subsequent embodiments and will not be elaborated here.

[0130] Step 302: The accelerator performs the first-layer calculation of the first iteration of the first session through the generative large model based on the input of the first session. At the same time, the accelerator loads the historical KV cache of the input of the first session at layer 2 from the memory of the host into the first buffer of the accelerator.

[0131] Among them, the accelerator starts a computing thread, and through this computing thread, performs the first-layer calculation of the first iteration of the first session.

[0132] Among them, when the number of tokens included in the input of the first session is less than the number of tokens that can be accommodated by the context window of the generative large model, that is, when the context window does not overflow, the process of the accelerator performing the first-layer calculation of the first iteration of the first session through the generative large model includes: projecting the input token sequence of the first session respectively with the model parameters corresponding to the self-attention mechanism in the first layer to generate intermediate data keys and values corresponding to each token in the input token sequence of the first session; updating the historical KV cache of the input of the first session in the first layer based on the keys and values of the input token sequence of the first session in the first layer; embedding the positional encoding of the input of the first session into the updated historical KV cache, and the positional encoding is relative positional encoding (RPE); performing a non-linear transformation on the keys and values in the historical KV cache embedded with the positional encoding, and the non-linear transformation includes residual connection and normalization, and passing the result of the non-linear transformation to the feed-forward neural network in the first layer; processing the result of the non-linear transformation through the feed-forward neural network in the first layer to obtain the calculation result of the first layer. Among them, the accelerator updating the historical KV cache of the input of the first session in the first layer based on the keys and values of the input token sequence of the first session in the first layer means that every time an accelerator calculation thread calculates an intermediate data key or value, the intermediate data is appended to the historical KV cache of the input of the first session in the first buffer.

[0133] The following is an example to illustrate Figure 5 the above-mentioned first-layer calculation process. Figure 5 FIG. is a schematic diagram of inference based on a KV cache decoupled from positional encoding provided by an embodiment of the present application. As Figure 5 shown, the accelerator projects the input token sequence (input) of the first session respectively with the model parameters (W k , W v ) corresponding to the self-attention mechanism in the first layer to generate intermediate data keys and values corresponding to each token in the input token sequence of the first session; the accelerator updates the historical KV cache; the accelerator embeds the positional encoding of the input of the first session into the key in the updated historical KV cache, and performs subsequent inference based on the KV cache embedded with the positional encoding.

[0134] Among them, when the number of tokens included in the input of the first session is greater than or equal to the number of tokens that can be accommodated by the context window of the generative large model, that is, when the context window overflows, the accelerator performs the first-layer calculation of the first iteration of the first session through the generative large model based on the first sub-input of the first session. The calculation process is the same as that in the case where the context window does not overflow. The difference is that when the context window overflows, the input of the generative large model and the historical KV cache loaded by the accelerator from the host's memory are cropped. The same principles will not be elaborated further.

[0135] Among them, the process of the accelerator loading the historical KV cache of the input of the first session at the second layer from the host's memory to the first buffer is the same as the process of loading the historical KV cache at the first layer in step 301, and will not be elaborated further.

[0136] Step 303: The accelerator performs the second-layer calculation of the first iteration of the first session. At the same time, the accelerator loads the historical KV cache of the input of the first session at the third layer from the host's memory to the first buffer, and writes the KV cache obtained from the first-layer calculation of the first iteration into the host's memory.

[0137] Among them, the accelerator performs the second-layer calculation of the first iteration of the first session through the calculation thread. At the same time, the accelerator loads the historical KV cache of the input of the first session at the third layer from the host's memory through the data reading thread, and the accelerator writes the KV cache obtained from the first-layer calculation of the first iteration from the first buffer into the host's memory through the data write-back thread.

[0138] Among them, the accelerator performs the second-layer calculation of the first iteration of the first session through the calculation thread based on the calculation result of the first-layer calculation of the first iteration of the first session. The calculation process of the second layer is the same as that of the first layer, and will not be elaborated further.

[0139] Among them, the process of the accelerator loading the historical KV cache of the input of the first session at the third layer from the host's memory to the first buffer through the data reading thread is the same as the process of loading the historical KV cache at the first layer in step 301, and will not be elaborated further.

[0140] In step 303 above, when the computing thread of the accelerator is performing the calculation of a certain layer, the data reading thread of the accelerator simultaneously starts to load the KV cache of the subsequent layer into the first buffer. After the calculation of the current layer is completed, the computing thread can immediately start the calculation of the subsequent layer based on the loaded data, hiding the time gap caused by loading the data required for the next layer between adjacent layers, thereby improving the efficiency of session inference. Moreover, when the computing thread of the accelerator is performing the calculation of a certain layer, the data write-back thread of the accelerator simultaneously writes back the KV cache generated by the previous layer to the memory of the host. Compared with writing back the KV cache generated by the current iteration to the host memory after one iteration is completed, it can hide the time gap caused by writing back the data generated by the previous iteration between adjacent iterations, and can improve the efficiency of inference.

[0141] Step 304: The accelerator performs the calculation of the 3rd layer to the Nth layer of the 1st iteration of the first session in the same way as in step 303, and infers the 1st token in the output of the first session.

[0142] The above steps 301 to 304 introduce the process of the 1st iteration of the first session. This 1st iteration is also the pre-filling stage in the inference process of the first session. After this pre-filling stage, the 1st token in the output of the first session is inferred. Next, the decoding stage in the inference process of the first session is introduced through step 305. Among them, this decoding stage includes at least one iteration, and each iteration infers one token in the output of the first session.

[0143] Step 305: The accelerator performs the 2nd iteration to the Mth iteration of the first session in the same way as in the above steps 301 to 304, and sequentially infers the 2nd token to the Mth token in the output of the first session. M is equal to the total number of iterations of the first session, and M is an integer greater than or equal to 2.

[0144] Among them, any iteration from the 2nd iteration to the Mth iteration is the same as the 1st iteration. The difference is that in the 1st iteration, each layer of calculation generates the intermediate data key and value corresponding to each token in the token sequence (the question prompt of the first session) of the input of the first session, while in any iteration from the 2nd iteration to the Mth iteration, each layer of calculation generates the intermediate data key and value corresponding to the token obtained in the previous iteration. The same points will not be elaborated.

[0145] Step 306: While the accelerator is performing the calculation of the Nth layer of the Mth iteration of the first session, the accelerator loads the historical KV cache of the input of the second session in the 1st layer in the task queue from the memory of the host to the second buffer of the accelerator.

[0146] Among them, the task queue is used to indicate the sessions to be processed. The second session is the first session to be processed in the task queue. The second buffer is the area in the HBM of the accelerator for storing the KV cache to be processed by the accelerator, that is, for storing the historical KV cache of the input of the second session at the first layer.

[0147] Among them, the accelerator loads the historical KV cache of the input of the second session at the first layer from the memory of the host to the second buffer through a data reading thread.

[0148] In the above method, the HBM of the accelerator includes a first buffer and a second buffer. The first buffer is used to store the KV cache of the session currently being processed by the accelerator, and the second buffer is used to store the KV cache of the session to be processed by the accelerator. While the accelerator processes the first session based on the KV cache in the first buffer, it loads the KV cache required for the first-layer calculation of the second session into the second buffer. When the accelerator finishes processing the first session, it processes the second session based on the KV cache in the second buffer. At this time, the second buffer becomes the execution buffer. While processing the second session, the accelerator releases the first buffer through an asynchronous thread, so that the computing thread of the accelerator can immediately start processing the second session after finishing processing the first session without waiting for the release of the first buffer to complete, thereby hiding the time gap caused by loading the data required for the next session between the two sessions and improving the inference efficiency.

[0149] It should be noted that the above step 306 is an optional step. In some embodiments, the accelerator only includes a first buffer; after the accelerator finishes the Nth-layer calculation of the Mth iteration of the first session, it releases the first buffer and then loads the historical KV cache of the input of the second session at the first layer into the first buffer. The embodiments of the present application do not limit this. In the above embodiments, all of the HBM of the accelerator is used as the first buffer, which can increase the capacity of the first buffer. When the amount of historical KV cache data of the session is large, it can accommodate the complete historical KV cache of the session and avoid the situation where the later-loaded historical KV cache overwrites the earlier-loaded historical KV cache due to insufficient capacity, thereby improving the hit rate of the historical KV cache in the first buffer and avoiding the repeated loading of the historical KV cache.

[0150] Step 307: If, when the accelerator completes the Nth layer calculation of the Mth iteration of the first session, there is KV cache in the KV cache generated by the generative large model for processing the first session that has not been written to the memory, the accelerator writes the KV cache that has not been written to the memory from the first buffer to the third buffer of the accelerator.

[0151] Among them, the KV cache generated by the generative large model for processing the first session is stored in the first buffer. When the accelerator performs the calculation of the (i + 1)th layer through the calculation thread, the data write-back thread of the accelerator writes the KV cache generated by the ith layer calculation from the first buffer back to the memory of the host, where i is an integer greater than or equal to 1 and less than N. The third buffer is the area in the HBM of the accelerator used to store the KV cache that has not been written back.

[0152] In the above method, the HBM of the accelerator further includes a third buffer. After the accelerator completes the processing of the first session, it writes the KV cache generated during the processing of the first session and not yet written back to the memory from the first buffer to the third buffer. Since the data copy efficiency between the first buffer and the third buffer is relatively high, after the accelerator completes the processing of the first session, the accelerator can quickly write the KV cache in the first buffer that has not been written back to the memory into the third buffer, and then can quickly release the first buffer, so as to be able to load the KV cache required for subsequent processing into the first buffer without waiting for all the KV cache generated by processing the first session to be written back to the memory before releasing the first buffer, hiding the time gap caused by waiting for the data write-back of the previous session between two sessions and improving the inference efficiency.

[0153] It should be noted that the above step 307 is an optional step. In some embodiments, the accelerator only includes a first buffer. After the accelerator completes the Nth layer calculation of the Mth iteration of the first session, it waits for all the KV cache generated by processing the first session to be written from the first buffer to the memory of the host, and then releases the first buffer. In the above embodiments, all of the HBM of the accelerator is used as the first buffer, which can increase the capacity of the first buffer. When the historical KV cache data volume of the session is large, it can accommodate the complete historical KV cache of the session, and can avoid the situation where the later-loaded historical KV cache overwrites the earlier-loaded historical KV cache due to insufficient capacity, thereby improving the hit rate of the historical KV cache in the first buffer and avoiding the repeated loading of the historical KV cache.

[0154] Step 308: The accelerator processes the second session in the same manner as steps 301 to 307 above. Meanwhile, the KV cache in the third buffer of the accelerator is written into the memory of the host.

[0155] Among them, in this step 308, the process of the accelerator writing the KV cache in the third buffer into the memory of the host is the same as the process in step 303 where the accelerator writes the KV cache calculated in the first layer of the first iteration into the memory of the host, and will not be elaborated here.

[0156] In the above method, the accelerator can preload the historical KV cache required for the (i + 1)-th layer calculation from the memory of the host into the accelerator while performing the i-th layer calculation of the generative large model for the session. Since the calculation process and the data loading process are synchronized, the calculation process does not need to wait for the data loading to complete, thus being able to hide the time overhead of the accelerator accessing the external storage device and improving the efficiency of multi-round session inference; further, while the accelerator is performing the (i + 1)-th layer calculation, the KV cache generated in the i-th layer calculation is written back to the memory of the host. Compared with writing the KV cache generated in this round of session into the memory of the host after a round of session inference is completed, it can hide the time gap caused by writing back the data generated in the previous round of session between two adjacent rounds of sessions, and can improve the inference efficiency.

[0157] Next, through Figure 6 and Figure 7 the processes shown in steps 301 to 308 above are exemplarily illustrated. Figure 6 is a schematic diagram of the parallelism between the calculation process and the data loading process in a multi-round session inference method provided by an embodiment of the present application. As Figure 6 shown, the generative large model includes three layers (L1, L2, and L3). After the previous session is processed, the accelerator processes the current session through this generative large model. Among them, before the calculation of the L1 layer occurs, the KV cache required for this layer is already ready in the first buffer in the HBM of the accelerator. Among them, the data loading thread of the accelerator, that is, the KV cache reading stream, loads the KV cache required for the L1 layer calculation into the first buffer; the calculation thread of the accelerator, that is, the execution stream, starts to perform the calculation of the L1 layer; while the calculation thread of the accelerator is performing the calculation of the L1 layer, the data loading thread simultaneously starts to load the KV cache required for the L2 layer calculation into the first buffer, and so on, so that the time for the accelerator to load the KV cache from the memory of the host overlaps with the calculation time, thereby being able to hide the time overhead of loading the KV cache from the memory of the host and significantly improving the inference efficiency.

[0158] Figure 7This is a schematic diagram showing the parallelism between the calculation process and the data write-back process in a multi-round conversation inference method provided by an embodiment of the present application. As Figure 7 shown, the generative large model includes three layers (L1, L2, and L3). In the pre-fill stage, the computing threads of the accelerator, that is, the execution streams, perform the calculation of layer L1. After the computing threads of the accelerator complete the calculation of layer L1, the computing threads of the accelerator perform the calculation of layer L2. At the same time, the data write-back threads of the accelerator, that is, the KV cache write-back streams, write the KV cache generated by the calculation of layer L1 from the first buffer back to the memory of the host. After the computing threads of the accelerator complete the calculation of layer L2, the computing threads of the accelerator perform the calculation of layer L3. At the same time, the data write-back threads of the accelerator write the KV cache generated by the calculation of layer L2 from the first buffer back to the memory of the host, and so on. The data write-back process in the decoding stage is the same as that in the pre-fill stage and will not be elaborated here. In Figure 7 the data write-back process shown, the time when the accelerator writes the generated KV cache back to the memory of the host overlaps with the calculation time, so as to hide the time overhead of writing the generated KV cache back to the memory of the host and significantly improve the inference efficiency.

[0159] The data migration process between the accelerator and the memory of the host is introduced above. Next, the data migration process between the memory of the host and the external storage device will be introduced.

[0160] Among them, the data migration process between the memory of the host and the external storage device includes: if the historical KV cache of the first pending conversation in the task queue does not exist in the memory of the host, the host loads the historical KV cache of the first pending conversation from the external storage device to the memory of the host, and the first pending conversation is the first number of pending conversations at the head of the task queue.

[0161] Among them, the number of conversation rounds included in the first pending conversation is determined based on the capacity of the memory of the host, that is, the first number is determined based on the capacity of the memory of the host. For example, the first number = the available storage space of the memory of the host ÷ the average data volume of the historical KV cache of the conversation. It should be noted that for the historical KV cache of the first conversation being processed by the accelerator, while the accelerator is performing the calculation, it writes the newly generated KV cache back to the memory of the host to update the historical KV cache of the first conversation in the memory of the host. The space occupied by the historical KV cache of the first conversation in the memory of the host belongs to system occupancy, and this space cannot be released until the first conversation is processed. Therefore, the available storage space of the memory of the host = the capacity of the memory of the host - the capacity of the space occupied by the system.

[0162] In some embodiments, if the available storage space in the host's memory is sufficient, the host loads the historical KV cache of the first session to be processed from an external storage device into the host's memory; if the available storage space in the host's memory is insufficient, the host writes the historical KV cache of the second session to be processed in the task queue from the host's memory to the external storage device, and loads the historical KV cache of the first session to be processed from the external storage device into the host's memory, where the second session to be processed is a session other than the first session to be processed in the task queue.

[0163] Those skilled in the art can determine how to judge whether the available storage space in the host's memory is sufficient according to actual needs. For example, in some embodiments, if the capacity of the available storage space in the host's memory is greater than or equal to a target threshold, it means that the available storage space in the host is sufficient; if the capacity of the available storage space in the host's memory is less than the target threshold, it means that the available storage space in the host is insufficient. For another example, in some other embodiments, if the capacity of the available storage space in the host's memory is greater than or equal to the data volume of the historical KV cache of the first session to be processed, it means that the available storage space in the host is sufficient; if the capacity of the available storage space in the host's memory is less than the data volume of the historical KV cache of the first session to be processed, it means that the available storage space in the host is insufficient. It should be noted that the above description of the method for judging whether the available storage space in the host's memory is sufficient is only exemplary, and the embodiments of the present application do not limit this judgment method.

[0164] In some embodiments, outside the available storage space in the host's memory, the host's memory further includes a fourth buffer, which is used for: storing the historical KV cache loaded from the external storage device when the available storage space in the host's memory is insufficient. Through this embodiment, the host's memory includes a fourth buffer, ensuring that there is enough migration space when migrating the KV cache from the external storage device to the host's memory. Thus, when the available storage space in the host is insufficient, the host can write the pre-loaded KV cache from the external storage device into this fourth buffer without waiting to evict the KV cache in the host's memory to the external storage device, avoiding the blockage caused by waiting for the host's memory to evict the KV cache when the available storage space in the host is insufficient.

[0165] In some embodiments, if the available storage space in the external storage device is insufficient, the historical KV cache of the third session to be processed in the task queue is deleted from the external storage device, where the third session to be processed is the second number of sessions to be processed at the end of the task queue. Among them, the second number can be determined according to the actual situation. For example, the second number is 1 or 2, etc., and the embodiments of the present application do not limit this.

[0166] The following is an example of the data migration process between the memory of the host and the external storage device through Figure 8 to illustrate. Figure 8 FIG. is a schematic diagram of the data migration process between the memory of the host and the external storage device during a multi-round session inference process provided by an embodiment of the present application. As Figure 8 shown, the accelerator is processing Job1, which is also the first session; the task queue includes 8 pending sessions, which are Job2 to Job9 in sequence from the head to the tail of the queue; the memory of the host includes a fourth buffer ( Figure 8 the area indicated by "buf" in), and the historical KV caches of Job1, Job2, and Job4 are stored in the memory of the host, which are KV1, KV2, and KV4 respectively; the external storage device is a disk (disks), in which the historical KV caches of Job9, Job8, Job7, and Job3 are stored, which are KV9, KV8, KV7, and KV3 respectively. The task scheduler in the host maintains a prefetching window, which is used to determine the first number of pending sessions at the head of the task queue, that is, the first pending session. Figure 8 In, the size of the prefetching window is 2, indicating that the first pending session includes two rounds of sessions, that is, the first number is 2. As Figure 8 shown, the first pending session is Job2 and Job3; the task scheduler in the host also maintains an eviction window. When the available storage space of the external storage device is insufficient, the sessions indicated by the eviction window will be exempted from being deleted from the external storage device. Figure 8 In, the size of the eviction window is 6. As Figure 8 shown, in the first pending session (Job2 and Job3), the historical KV cache (KV2) of Job2 exists in the memory of the host, the historical KV cache (KV3) of Job3 does not exist in the memory of the host, and the available storage space of the memory of the host is insufficient. Then the host loads KV3 from the disk into the fourth buffer in the memory of the host, and writes the historical KV cache of the second pending session in the task queue from the memory of the host to the disk, where the historical KV cache of the second pending session exists in the memory of the host and the second pending session is not the first pending session. The fact that the second pending session is not the first pending session means that the second pending session is outside the prefetching window. Figure 8Among them, the second session to be processed is Job4, and the host writes the historical KV cache (KV4) of job4 from the host's memory to the disk. Since the available storage space on the disk is insufficient, the host deletes the historical KV cache of the third session to be processed from the disk. The third session to be processed is the second number of sessions at the end of the task queue outside the elimination exemption window. Figure 8 Among them, if the value of the second number is 1, then the third session to be processed is Job9, and the host deletes the historical KV cache (KV 9) of Job9 from the disk. It should be noted that Figure 8 This is only an example of the data migration process between the host's memory and the external storage device, and does not limit this application.

[0167] In the above method, the external storage device stores the historical KV cache of the completed sessions. Since the capacity of the external storage device is much larger than the capacity of the HBM in the accelerator, the external storage device can store a large amount of historical KV caches, thereby improving the hit rate of the historical KV cache, avoiding KV cache recalculation, improving the efficiency of multi-round session inference, and saving computing resources. At the same time, while the accelerator is processing the session, the host can preload the historical KV cache of the sessions to be processed in the task queue from the external storage device into the host's memory. Since the calculation process and the data loading process are synchronized, the calculation process does not need to wait for the data loading to complete, thereby hiding the time overhead of the accelerator accessing the external storage device and improving the efficiency of multi-round session inference.

[0168] It should be noted that, for the convenience of introducing the multi-round session inference method provided by the embodiments of this application above, two embodiments are used to separately introduce the data migration process between the accelerator and the host's memory and the data migration process between the host's memory and the external storage device. However, this does not mean that the above two embodiments are independent and occur separately. That is, in the multi-round session inference method provided by the embodiments of this application, while data migration occurs between the accelerator and the host's memory, data migration can also occur between the host's memory and the external storage device.

[0169] Figure 9 This is a schematic structural diagram of a multi-round session inference device provided by an embodiment of this application, which is applied to an accelerator of a host in a computing system. The computing system further includes a host and an external storage device. The external storage device is used to store the historical key-value cache KV cache of the sessions. The host is used to load the historical KV cache of the sessions to be processed in the task queue from the external storage device into the memory of the host. The device includes a processing module 901 and a loading module 902.

[0170] The processing module 901 is configured to perform inference on the first session in the task queue through a generative large model based on the task queue. The first session is one round in a multi-round session, and the generative large model includes N layers.

[0171] The loading module 902 is configured to, when performing the i-th layer calculation on the first session, load the historical KV cache of the input of the first session at the (i + 1)-th layer from the memory, where i is an integer greater than or equal to 1 and less than N.

[0172] In some embodiments, the historical KV cache of the input of the first session at the (i + 1)-th layer is independent of the positional encoding of the input of the first session.

[0173] In some embodiments, the number of tokens included in the input of the first session is greater than or equal to the number of tokens that can be accommodated by the context window of the generative large model.

[0174] The loading module 902 is configured to:

[0175] When performing the i-th layer calculation on the first session, load the historical KV cache of the first sub-input at the (i + 1)-th layer from the memory. The first sub-input is a sub-input of the input of the first session, and the number of tokens included in the first sub-input is less than the number of tokens that can be accommodated by the context window.

[0176] In some embodiments, the processing module 901 includes:

[0177] An embedding unit configured to embed the positional encoding of the first sub-input into the historical KV cache of the first sub-input at the (i + 1)-th layer.

[0178] A calculation unit configured to perform the (i + 1)-th layer calculation on the first session based on the input of the first session and the historical KV cache of the first sub-input with the embedded positional encoding at the (i + 1)-th layer.

[0179] In some embodiments, the accelerator includes a first buffer and a second buffer. The first buffer is used to store the historical KV cache of the first session, and the second buffer is used to store the historical KV cache of the input of the second session in the task queue at the first layer. The second session is the first session to be processed after the first session in the task queue.

[0180] The loading module 902 is configured to:

[0181] When performing the i-th layer calculation on the first session, load the historical KV cache of the input of the first session at the (i + 1)-th layer from the memory to the first buffer.

[0182] In some embodiments, the loading module 902 is further configured to:

[0183] When performing the Nth layer calculation on the first session, load the historical KV cache of the input of the second session at the first layer from the memory to the second buffer.

[0184] In some embodiments, the apparatus further includes:

[0185] A writing module, configured to write the KV cache generated by performing the ith layer calculation on the first session into the memory when performing the (i + 1)th layer calculation on the first session.

[0186] In some embodiments, the accelerator includes a third buffer, and the third buffer is used to store the KV cache generated by the generative large model processing the first session;

[0187] The writing module is configured to:

[0188] If there is KV cache that has not been written into the memory in the KV cache generated by the generative large model processing the first session when the Nth layer calculation of the first session is completed, write the KV cache that has not been written into the memory into the third buffer;

[0189] When processing the second session in the task queue through the generative large model, write the KV cache in the third buffer into the memory.

[0190] Figure 10 It is a schematic structural diagram of a multi-round session inference apparatus provided by an embodiment of the present application, which is applied to a host in a computing system. The computing system further includes an accelerator of the host and an external storage device. The external storage device is used to store the historical key-value cache KV cache of the session. The accelerator is used to perform inference on the first session in the task queue through a generative large model based on the task queue. The first session is one round of a multi-round session. The generative large model includes N layers. Among them, when performing the ith layer calculation on the first session, load the historical KV cache of the input of the first session at the (i + 1)th layer from the memory, where i is an integer greater than or equal to 1 and less than N; the apparatus includes a loading module 1001 and a storage module 1002.

[0191] The loading module 1001 is configured to load the historical KV cache of the session to be processed in the task queue from the external storage device to the memory of the host;

[0192] The storage module 1002 is configured to store the historical KV cache of the session to be processed.

[0193] In some embodiments, the loading module 1001 includes:

[0194] A loading unit, configured to load the historical KV cache of the first session to be processed in the task queue from the external storage device into the memory if the historical KV cache of the first session to be processed in the task queue does not exist in the memory, where the first session to be processed is the first number of sessions to be processed at the head of the task queue.

[0195] In some embodiments, the number of session rounds included in the first session to be processed is determined based on the capacity of the memory.

[0196] In some embodiments, the loading unit is configured to:

[0197] If the available storage space in the memory is sufficient, load the historical KV cache of the first session to be processed from the external storage device into the memory;

[0198] If the available storage space in the memory is insufficient, write the historical KV cache of the second session to be processed in the task queue from the memory to the external storage device, and load the historical KV cache of the first session to be processed from the external storage device into the memory, where the second session to be processed is a session in the task queue other than the first session to be processed.

[0199] In some embodiments, the apparatus further includes a deletion module, configured to:

[0200] If the available storage space in the external storage device is insufficient, delete the historical KV cache of the third session to be processed from the external storage device, where the third session to be processed is the second number of sessions to be processed at the tail of the task queue.

[0201] Among them, the processing module 901, the loading module 902, the loading module 1001, and the storage module 1002 can all be implemented by software or by hardware. Exemplarily, next, taking the processing module 901 as an example, the implementation manner of the processing module 901 will be introduced. Similarly, the implementation manners of the loading module 902, the loading module 1001, and the storage module 1002 can refer to the implementation manner of the processing module 901.

[0202] As an example of a software functional unit, the processing module 901 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the processing module 901 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running the code may be distributed in the same availability zone (AZ) or in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Usually, one region may include multiple AZs.

[0203] Similarly, the multiple hosts / virtual machines / containers for running the code may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, one VPC is set up within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC, and the interconnection between VPCs is achieved through the communication gateway.

[0204] As an example of a hardware functional unit, the processing module 901 may include at least one computing device, such as a server, etc. Alternatively, the processing module 901 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The above PLD may be implemented using a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0205] The multiple computing devices included in the processing module 901 can be distributed in the same region or in different regions. The multiple computing devices included in the processing module 901 can be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the processing module 901 can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices can be any combination of computing devices such as servers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), and generic array logic (GALs).

[0206] It should be noted that in other embodiments, the processing module 901 can be used to execute any step in the multi-round session inference method, the loading module 902 can be used to execute any step in the multi-round session inference method, the loading module 1001 can be used to execute any step in the multi-round session inference method, and the storage module 1002 can be used to execute any step in the multi-round session inference method. The steps to be implemented by the processing module 901, the loading module 902, the loading module 1001, and the storage module 1002 can be specified as needed. By separately implementing different steps of the multi-round session inference method executed by the accelerator through the processing module 901 and the loading module 902, the entire function of the multi-round session inference device as shown in Figure 9 is achieved; or, the entire function of the multi-round session inference device as shown in Figure 10 is achieved by separately implementing different steps of the multi-round session inference method executed by the host through the loading module 1001 and the storage module 1002.

[0207] This application also provides a computing device 1100. Figure 11 It is a schematic structural diagram of a computing device provided by an embodiment of this application. As shown in Figure 11 , the computing device 1100 includes: a bus 1101, a processor 1102, a memory 1103, and a communication interface 1104. The processor 1102, the memory 1103, and the communication interface 1104 communicate with each other through the bus 1101. The computing device 1100 can be a computing device or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1100.

[0208] The bus 1101 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 11The bus 1101 may include a path for transmitting information between various components of the computing device 1100 (eg, the memory 1103, the processor 1102, and the communication interface 1104).

[0209] The processor 1102 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0210] The memory 1103 may include a volatile memory, such as a random access memory (RAM). The memory 1103 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0211] Memory 1103 stores executable program code. Processor 1102 executes this executable program code to implement the functions of processing module 901 and loading module 902, respectively, thereby implementing the steps of the multi-turn conversational reasoning method performed by the accelerator. Alternatively, it implements the functions of loading module 1001 and storage module 1002, thereby implementing the steps of the multi-turn conversational reasoning method performed by the host. In other words, memory 1103 stores instructions for executing the multi-turn conversational reasoning method. Figure 11 The memory 1103 is merely shown to store program codes for implementing the functions of the aforementioned processing module 901 and loading module 902 .

[0212] The communication interface 1104 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1100 and other devices or a communication network.

[0213] An embodiment of the present application also provides a computing device cluster. Figure 12 is a schematic diagram of a computing device cluster provided in an embodiment of the present application, such as Figure 12As shown, the computing device cluster includes at least one computing device 1100. Instructions for executing the multi-round session inference method may be stored in the memory 1103 of one or more of the computing devices 1100 in the computing device cluster.

[0214] In some possible implementation manners, partial instructions for executing the multi-round session inference method may also be separately stored in the memory 1103 of one or more of the computing devices 1100 in the computing device cluster. In other words, a combination of one or more computing devices 1100 may jointly execute the instructions for executing the multi-round session inference method.

[0215] It should be noted that the memories 1103 in different computing devices 1100 in the computing device cluster may store different instructions, which are respectively used to execute partial functions of the multi-round session inference device. That is, the instructions stored in the memories 1103 of different computing devices 1100 may implement the functions of one or more of the foregoing processing module 901, loading module 902, loading module 1001, and storage module 1002.

[0216] It should be understood that Figure 12 the functions of the computing device 1100 shown in

[0217] In some possible implementation manners, one or more computing devices in the computing device cluster may be connected through a network. Among them, the network may be a wide area network or a local area network, etc. Figure 13 A possible implementation manner is shown. Figure 13 FIG. is a schematic diagram of a possible implementation manner of a computing device cluster provided by an embodiment of the present application. As Figure 13 shown, the computing device 1100A and the computing device 1100B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation manner, instructions for implementing the function of the processing module 901 are stored in the memory 1103 of the computing device 1100A. Figure 13 Taking the example that instructions for implementing the function of the processing module 901 are stored in the memory 1103 of the computing device 1100A in Figure 13 Taking the example that instructions for implementing the function of the loading module 902 are stored in the memory 1103 of the computing device 1100B in

[0218] The embodiment of the present application also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster may be similarly referred to Figure 12 or Figure 13Connection methods. The difference is that the memory 1103 in one or more computing devices 1100 in the computing device cluster may store the same instructions for executing the multi-round session inference method.

[0219] In some possible implementation manners, the memory 1103 of one or more computing devices 1100 in the computing device cluster may also separately store partial instructions for executing the multi-round session inference method. In other words, the combination of one or more computing devices 1100 may jointly execute the instructions for executing the multi-round session inference method.

[0220] It should be noted that the memories 1103 in different computing devices 1100 in the computing device cluster may store different instructions, which are respectively used to execute partial functions of the multi-round session inference device. That is, the instructions stored in the memories 1103 of different computing devices 1100 may implement the functions of one or more of the foregoing processing module 901, loading module 902, loading module 1001, and storage module 1002.

[0221] An embodiment of the present application provides an accelerator, which includes a computing core and a memory. The accelerator is used to execute the steps of the multi-round session inference method executed by the accelerator in any possible implementation manner provided in the foregoing method embodiments. The memory is used to store calculation data, and the computing core is used to perform calculation operations on the calculation data stored in the memory.

[0222] An embodiment of the present application provides a host, which includes a processor and a memory. The host is used to execute the steps of the multi-round session inference method executed by the host in any possible implementation manner provided in the foregoing method embodiments.

[0223] An embodiment of the present application further provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, at least one computing device is caused to execute the multi-round session inference method.

[0224] An embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium may be any available medium that a computing device can store or a data storage device such as a data center including one or more available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc. The computer-readable storage medium includes instructions, and the instructions instruct the computing device to execute the multi-round session inference method.

[0225] It should be noted that the information involved in this application (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the sessions and storage spaces involved in this application are obtained under full authorization.

[0226] Those of ordinary skill in the art can realize that, in combination with the method steps and units described in the embodiments disclosed herein, they can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the steps and components of the embodiments have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0227] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0228] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the unit is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can also be electrical, mechanical, or other forms of connection.

[0229] The unit described as a separated component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of this application.

[0230] In addition, in each embodiment of the present application, each unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software unit.

[0231] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computing device (which can be a personal computer, a server, or a computing device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc., which can store program codes.

[0232] In the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor are the quantity and execution order limited. It should also be understood that although the following description uses terms such as first and second to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of various examples, the first session can be referred to as the second session, and similarly, the second session can be referred to as the first session. Both the first session and the second session can be node sessions, and in some cases, they can be separate and different sessions.

[0233] In the present application, the meaning of the term "at least one" refers to one or more, and the meaning of the term "multiple" refers to two or more. In this text, the terms "system" and "network" are often used interchangeably.

[0234] It should also be understood that the term "if" can be interpreted as meaning "when" ("when" or "upon") or "in response to a determination" or "in response to a detection". Similarly, depending on the context, the phrase "if a determination is made..." or "if [the stated condition or event] is detected" can be interpreted as meaning "when a determination is made..." or "in response to a determination..." or "when [the stated condition or event] is detected" or "in response to a detection of [the stated condition or event]".

[0235] The above description is only a specific implementation of the present application. However, the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present application, and these modifications or substitutions should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

[0236] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer program instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0237] The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer program instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid-state drive).

[0238] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.

[0239] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent substitutions for some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A computing system, characterized in that, The computing system includes a host, an accelerator of the host, and an external storage device; wherein, the external storage device is used to store the historical key-value cache (KV cache) of the session; the host is used to load the historical KV cache of the session to be processed in the task queue from the external storage device into the memory of the host; the accelerator is used to perform inference on a first session in the task queue based on the task queue through a generative large model, the first session is one round of a multi-round session, and the generative large model includes N layers, wherein, when performing the i-th layer calculation on the first session, load the historical KV cache of the input of the first session at the (i + 1)-th layer from the memory, where i is an integer greater than or equal to 1 and less than N.

2. The computing system according to claim 1, wherein The historical KV cache of the input of the first session at the (i + 1)-th layer has nothing to do with the positional encoding of the input of the first session.

3. The computing system according to claim 2, wherein The number of tokens included in the input of the first session is greater than or equal to the number of tokens that can be accommodated by the context window of the generative large model; The accelerator is used for: when performing the i-th layer calculation on the first session, load the historical KV cache of a first sub-input at the (i + 1)-th layer from the memory, the first sub-input is a sub-input of the input of the first session, and the number of tokens included in the first sub-input is less than the number of tokens that can be accommodated by the context window.

4. The computing system according to claim 3, wherein The accelerator is further used for: embed the positional encoding of the first sub-input into the historical KV cache of the first sub-input at the (i + 1)-th layer; perform the (i + 1)-th layer calculation on the first session based on the input of the first session and the historical KV cache of the first sub-input with the embedded positional encoding.

5. The computing system according to claim 1, wherein The accelerator includes a first buffer and a second buffer. The first buffer is used to store the historical KV cache of the first session, and the second buffer is used to store the historical KV cache of the input of a second session in the task queue at the first layer. The second session is the first session to be processed after the first session in the task queue; The accelerator is used for: when performing the i-th layer calculation on the first session, load the historical KV cache of the input of the first session at the (i + 1)-th layer from the memory to the first buffer; when performing the N-th layer calculation on the first session, load the historical KV cache of the input of the second session at the first layer from the memory to the second buffer.

6. The computing system according to claim 1, wherein The accelerator is further used for: when performing the (i + 1)-th layer calculation on the first session, write the KV cache generated by performing the i-th layer calculation on the first session into the memory.

7. The computing system according to claim 6, wherein The accelerator includes a third buffer, and the third buffer is used to store the KV cache generated by the generative large model processing the first session; The accelerator of the host is used for: When the N - th layer calculation of the first session is completed, if there is a KV cache in the KV cache generated by the generative large - model for processing the first session that has not been written to the memory, write the KV cache that has not been written to the memory to the third buffer; When processing the second session in the task queue through the generative large - model, write the KV cache in the third buffer to the memory.

8. The computing system according to claim 1, wherein The host is used for: If the historical KV cache of the first pending session in the task queue does not exist in the memory, load the historical KV cache of the first pending session from the external storage device to the memory, where the first pending session is the first number of pending sessions at the head of the task queue.

9. The computing system according to claim 8, wherein The number of conversation rounds included in the first pending session is determined based on the capacity of the memory.

10. The computing system according to claim 8 or 9, characterized in that, The host is used for: If the available storage space in the memory is sufficient, load the historical KV cache of the first pending session from the external storage device to the memory; If the available storage space in the memory is insufficient, write the historical KV cache of the second pending session in the task queue from the memory to the external storage device, and load the historical KV cache of the first pending session from the external storage device to the memory, where the second pending session is the session in the task queue other than the first pending session.

11. The computing system according to claim 1, wherein The host is also used for: If the available storage space in the external storage device is insufficient, delete the historical KV cache of the third pending session from the external storage device, where the third pending session is the second number of pending sessions at the tail of the task queue.

12. A multi-round conversation reasoning method, characterized in that, Applied to a computing system, the computing system includes a host, an accelerator of the host, and an external storage device. The external storage device is used to store the historical key - value cache (KV cache) of sessions. The host is used to load the historical KV cache of the pending sessions in the task queue from the external storage device to the memory of the host. The method includes: Through the accelerator, based on the task queue, perform inference on the first session in the task queue through a generative large - model. The first session is one round of a multi - round session. The generative large - model includes N layers, where When performing the i - th layer calculation on the first session through the accelerator, load the historical KV cache of the input of the first session at the (i + 1) - th layer from the memory to the accelerator, where i is an integer greater than or equal to 1 and less than N.

13. A multi-round conversation reasoning method, characterized in that, Applied to the accelerator of the host in a computing system. The computing system also includes a host and an external storage device. The external storage device is used to store the historical key - value cache (KV cache) of sessions. The host is used to load the historical KV cache of the pending sessions in the task queue from the external storage device to the memory of the host. The method includes: Based on the task queue, perform inference on the first session in the task queue through a generative large model. The first session is one round of a multi-round session. The generative large model includes N layers, where, When performing the i-th layer calculation on the first session, load the historical KV cache of the input of the first session at the (i + 1)-th layer from the memory, where i is an integer greater than or equal to 1 and less than N.

14. The method according to claim 13, wherein The historical KV cache of the input of the first session at the (i + 1)-th layer has nothing to do with the positional encoding of the input of the first session.

15. The method according to claim 14, wherein The number of tokens included in the input of the first session is greater than or equal to the number of tokens that the context window of the generative large model can accommodate; The step of loading the historical KV cache of the input of the first session at the (i + 1)-th layer from the memory when performing the i-th layer calculation on the first session includes: When performing the i-th layer calculation on the first session, load the historical KV cache of the first sub-input at the (i + 1)-th layer from the memory. The first sub-input is a sub-input of the input of the first session, and the number of tokens included in the first sub-input is less than the number of tokens that the context window can accommodate.

16. The method according to claim 15, characterized in that, The step of processing the first session in the task queue through the generative large model includes: Embed the positional encoding of the first sub-input into the historical KV cache of the first sub-input at the (i + 1)-th layer; Based on the input of the first session and the historical KV cache of the first sub-input at the (i + 1)-th layer with the embedded positional encoding, perform the (i + 1)-th layer calculation on the first session.

17. The method according to claim 13, wherein The accelerator includes a first buffer and a second buffer. The first buffer is used to store the historical KV cache of the first session. The second buffer is used to store the historical KV cache of the input of the second session in the task queue at the first layer. The second session is the first session to be processed after the first session in the task queue; The step of loading the historical KV cache of the input of the first session at the (i + 1)-th layer from the memory when performing the i-th layer calculation on the first session includes: When performing the i-th layer calculation on the first session, load the historical KV cache of the input of the first session at the (i + 1)-th layer from the memory to the first buffer.

18. The method according to claim 17, wherein The method further includes: When performing the N-th layer calculation on the first session, load the historical KV cache of the input of the second session at the first layer from the memory to the second buffer.

19. The method according to claim 13, wherein The method further includes: When performing the (i + 1)-th layer calculation on the first session, write the KV cache generated by performing the i-th layer calculation on the first session into the memory.

20. The method according to claim 19, characterized in that The accelerator includes a third buffer, and the third buffer is used to store the KV cache generated by the generative large model processing the first session; The step of writing the KV cache generated by performing the i-th layer calculation on the first session into the memory when performing the (i + 1)-th layer calculation on the first session includes: When the KV cache generated by the generative large model during the processing of the first session exists in the KV cache that has not been written to the memory when the Nth layer calculation of the first session is completed, write the KV cache that has not been written to the memory to the third buffer; When processing the second session in the task queue through the generative large model, write the KV cache in the third buffer to the memory.

21. A multi-round conversation reasoning method, characterized in that, Applied to the host in a computing system, the computing system further includes an accelerator of the host and an external storage device for storing the historical key-value cache KVcache of the session. The accelerator is used to perform inference on the first session in the task queue through the generative large model based on the task queue. The first session is one round of a multi-round session. The generative large model includes N layers. Among them, when performing the ith layer calculation on the first session, load the historical KV cache of the input of the first session at the (i + 1)th layer from the memory of the host, where i is an integer greater than or equal to 1 and less than N; The method includes: Load the historical KV cache of the session to be processed in the task queue from the external storage device to the memory of the host; Store the historical KV cache of the session to be processed.

22. The method according to claim 21, wherein The loading the historical KV cache of the session to be processed in the task queue from the external storage device to the memory of the host includes: If the historical KV cache of the first session to be processed in the task queue does not exist in the memory, load the historical KV cache of the first session to be processed from the external storage device to the memory. The first session to be processed is the first number of sessions to be processed at the head of the task queue.

23. The method according to claim 22, wherein The number of session rounds included in the first session to be processed is determined based on the capacity of the memory.

24. The method according to claim 22 or 23, characterized in that, The if the historical KV cache of the first session to be processed in the task queue does not exist in the memory, then the host loads the historical KV cache of the first session to be processed from the external storage device to the memory includes: If the available storage space in the memory is sufficient, load the historical KV cache of the first session to be processed from the external storage device to the memory; If the available storage space in the memory is insufficient, write the historical KVcache of the second session to be processed in the task queue from the memory to the external storage device, and load the historical KV cache of the first session to be processed from the external storage device to the memory. The second session to be processed is the session other than the first session to be processed in the task queue.

25. The method according to any one of claims 21 to 24, characterized in that, The method further includes: If the available storage space in the external storage device is insufficient, delete the historical KV cache of the third session to be processed from the external storage device. The third session to be processed is the second number of sessions to be processed at the end of the task queue.

26. A multi-round conversation reasoning device, characterized in that, An accelerator for a host in a computing system, the computing system further including a host and an external storage device, the external storage device being used to store the historical key-value cache KVcache of a session, the host being used to load the historical KV cache of the session to be processed in a task queue from the external storage device into the memory of the host, the apparatus including: A processing module, configured to perform inference on a first session in the task queue through a generative large model based on the task queue, the first session being one round of a multi-round session, and the generative large model including N layers; A loading module, configured to, when performing the i-th layer calculation on the first session, load the historical KV cache of the input of the first session at the (i + 1)-th layer from the memory, where i is an integer greater than or equal to 1 and less than N.

27. A multi-round conversation reasoning device, characterized in that A host for a computing system, the computing system further including an accelerator of the host and an external storage device, the external storage device being used to store the historical key-value cache KVcache of a session, the accelerator being used to perform inference on a first session in the task queue through a generative large model based on the task queue, the first session being one round of a multi-round session, and the generative large model including N layers, wherein, when performing the i-th layer calculation on the first session, the historical KV cache of the input of the first session at the (i + 1)-th layer is loaded from the memory of the host, where i is an integer greater than or equal to 1 and less than N; The apparatus includes: A loading module, which loads the historical KV cache of the session to be processed in the task queue from the external storage device into the memory of the host.

28. An accelerator, characterized in that, The accelerator includes a computing core and a memory, the accelerator is used to execute any one of the multi-round session inference methods in claims 13 to 20, the memory is used to store calculation data, and the computing core is used to perform calculation operations on the calculation data stored in the memory.

29. A host, characterized in that, The host includes a processor and a memory, and the host is used to execute any one of the multi-round session inference methods in claims 21 to 25.

30. A cluster of computing devices, characterized in that, The computing device cluster includes at least one computing device, each computing device including a host, an accelerator of the host, and the external storage device is located in at least one computing device, the external storage device being used to store the historical KV cache of a session, and the host being used to load the historical KV cache of the session to be processed in the task queue from the external storage device into the memory of the host; The host in the computing device cluster is used to execute program code to implement the steps performed by the host in any one of the multi-round session inference methods as described in claims 13 to 20 above; The accelerator in the computing device cluster is used to execute program code to implement the steps performed by the accelerator in any one of the multi-round session inference methods as described in claims 21 to 25 above.

31. A computer program product comprising instructions, characterized in that, When the instruction is run by the computing device cluster, the computing device cluster is caused to execute any one of the multi-round session inference methods in claims 13 to 20, or claims 21 to 25.

32. A computer-readable storage medium, characterized in that, including computer program instructions which, when executed by a cluster of computing devices, cause the cluster of computing devices to execute the multi-round session inference method according to any one of claims 13 to 20 or claims 21 to 25.

Citation Information

Cited By

  • Automatic software debugging method, system and device based on large language model

    CN121579328A

  • Computing system, multi-round session inference method, apparatus and computing device cluster

    WO2025161395A1