Inference method and system, electronic equipment and storage medium

By adopting a three-level cache structure and refined communication strategy in the inference system, cache management and communication efficiency are optimized, the problem of inferred inference in the existing technology is solved, and more efficient data transmission and inference calculation are achieved.

CN119990307APending Publication Date: 2025-05-13SHANGHAI DUOYUN ROAD TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411917100.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing inference implementation solutions still have shortcomings in improving inference efficiency and resource utilization, especially when dealing with large-scale artificial intelligence models, cache management and data transmission overhead is high, resulting in inferred inference efficiency.

Method used

The third-level cache structure is adopted, including the first cache, the second cache and the third cache, and the inference computing process is used to optimize cache management and communication efficiency. Matching is performed through the first cache, the second cache performs data transmission, and the third cache performs inference calculations, and the result is stored in the first cache for multiplexing.

Benefits of technology

It reduces the data transmission overhead and time cost during the inference process, improves inference efficiency, and optimizes cache management, reducing system performance bottlenecks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990307A_ABST
    Figure CN119990307A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of artificial intelligence, and discloses an inference method and system, electronic equipment and a storage medium. The reasoning method comprises the steps that a user request is received, and the user request carries input information; matching in a first cache according to the input information; and based on the second cache, transmitting the matched token sequence, the KV parameters of the token sequence and the input information to a third cache, so that a reasoning model performs reasoning calculation based on the third cache, and transmitting the token sequence obtained by reasoning and the KV parameters of the token sequence to the first cache for storage. The data transmission overhead and cost in the reasoning process are at least reduced, so that the reasoning efficiency is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a reasoning method and system, an electronic device, and a storage medium. Background Art

[0002] As the scale of artificial intelligence (AI) models continues to expand, the reasoning process consumes more and more computing resources, and the need to improve the reasoning efficiency and resource utilization of AI models has become more urgent. For example, in order to improve reasoning efficiency and resource utilization, a technical solution of PD (prefill-decode) separation is proposed. In the computationally intensive prefill stage of the reasoning process, the AI ​​model processes all user inputs and calculates the corresponding key-value cache (KV Cache) and the first token; in the memory-intensive decode stage of the reasoning process, new tokens are generated one by one in sequence through iteration based on the calculated KV Cache and the first token, and only one token is calculated each time the memory is accessed. Based on PD separation, other reasoning implementation solutions are further provided.

[0003] However, the reasoning efficiency of existing reasoning implementation solutions still needs to be further improved to meet higher user requirements. Summary of the invention

[0004] The embodiments of the present application provide a reasoning method and system, an electronic device, and a storage medium, which are at least helpful in reducing the data transmission overhead and cost in the reasoning process to further improve the reasoning efficiency.

[0005] According to some embodiments of the present application, a first aspect of the embodiments of the present application provides an inference method, including: receiving a user request, wherein the user request carries input information; matching in a first cache based on the input information; based on the second cache, passing the matched token sequence and its KV parameters, and the input information to a third cache, so that the inference model performs inference calculations based on the third cache, and passing the inferred token sequence and its KV parameters to the first cache for storage.

[0006] According to some embodiments of the present application, a second aspect of the embodiments of the present application also provides an inference system, including: a prefill cluster, a decode cluster, a first cache, a second cache and a third cache, wherein the prefill cluster, the decode cluster, the first cache, the second cache and the third cache are used to implement the inference method as described in the first aspect.

[0007] According to some embodiments of the present application, a third aspect of the embodiments of the present application also provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the reasoning method described in the first aspect.

[0008] According to some embodiments of the present application, a fourth aspect of the embodiments of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the reasoning method described in the first aspect.

[0009] The technical solution provided in the embodiments of the present application has at least the following advantages:

[0010] In the process of implementing reasoning in response to user requests, firstly, the token sequence and its KV parameters stored in the first cache are matched, and then data is transferred between the first cache and the third cache for inference calculation through the second cache, so as to transfer the matched token sequence and its KV parameters and input information to the third cache, and transfer the token sequence and its KV parameters obtained by reasoning to the first cache for storage. While completing the reasoning, the token sequence and its KV parameters obtained by the current reasoning are stored for reuse in subsequent responses to user requests. In this way, through the coordination of the three-level storage structure of the first cache, the second cache and the third cache and the reasoning process, the speed of memory access in the process of reasoning implementation is optimized, so that the data transmission overhead and time cost in the reasoning process are reduced, and further, the reasoning efficiency is also improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0012] Figure 1 This is the process of the reasoning method provided in the embodiment of this application Figure 1 ;

[0013] Figure 2 This is the process of the reasoning method provided in the embodiment of this application Figure 2 ;

[0014] Figure 3 is a schematic diagram of the structure of the reasoning system provided in the embodiment of the present application;

[0015] Figure 4It is a partial structural diagram of the reasoning system provided in the embodiments of the present application;

[0016] Figure 5 It is a schematic diagram of data flow during the application of the reasoning system and reasoning method provided in the embodiments of the present application;

[0017] Figure 6 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0018] It can be seen from the background technology that the reasoning efficiency of existing reasoning implementation solutions still needs to be further improved.

[0019] After analysis, it is found that the room for improving the inference efficiency in the existing inference schemes mainly lies in: there is room for improvement in the application of cache in the inference process. Specifically, with the increasing scale and data volume of artificial intelligence models, the problems of high video memory usage and large computational delay in the calculation process of large artificial intelligence models such as transformers are becoming increasingly prominent. Among them, most artificial intelligence models are composed of multiple layers of encoders and decoders, each layer of which contains a self-attention mechanism and a feedforward neural network, which makes them show obvious serial characteristics when performing inference calculations. Therefore, when processing long sequence data, artificial intelligence models need to manage a large number of weight matrices and intermediate results, which poses a severe challenge to cache management. Cache management will involve frequent data transmission and cache access, which will not only increase the pressure on video memory, but also cause significant data overhead and delay. At the same time, in the scenario of PD separation, frequent data exchange between nodes such as prefill nodes and decode nodes may cause network bottlenecks, increase latency and affect the overall system performance. In the existing inference implementation scheme, the graphics processing unit (GPU) usually uses its parallel processing capabilities to store and process the cache scheme of the intermediate calculation process to cooperate with the inference calculation process, which can significantly improve the calculation efficiency and reduce the required calculation time. However, due to the independent management of the inference instance cache, the cache hit rate is low and the cache effect is poor. Taking the large-scale virtual language model (vLLM) reasoning framework as an example, it uses a method of pre-allocating a large amount of memory to manage the cache on the basis of the PD separation architecture, and uses a multi-card communication method to synchronize data, which not only leads to a waste of video memory, but also generates a large amount of communication overhead. In application scenarios involving large-scale language models (LLMs), the storage consumption of weights and KV Cache often becomes the main bottleneck of system cache management. This phenomenon is particularly significant when it is necessary to process context-related dynamic generation tasks. For example, the length of the generated sequence may be uncertain, which may be very short or abnormally long. Therefore, a fixed cache space cannot be pre-allocated to store these intermediate calculation results during the inference process, otherwise it may cause cache waste or insufficient problems. At this time, significant communication overhead will be generated during the transmission of data between nodes. This communication burden is particularly prominent in hardware environments that do not support high-speed interconnection technologies (such as NVLink). For example, when using GPUs such as NVIDIA RTX4090 that do not support NVLink, the overall performance will be significantly restricted, thus forming a performance bottleneck. In this case, data exchange delays and bandwidth limitations become key factors affecting the efficiency of distributed reasoning, posing severe challenges to the real-time performance and throughput of the system.It can be seen that reducing cross-device data transmission overhead and time costs is the key to improving the reasoning efficiency of artificial intelligence models. At the same time, ensuring that model performance is not affected and avoiding resource waste or system instability caused by improper memory allocation are also important influencing factors.

[0020] Based on this, an inference method and system, an electronic device, and a storage medium are provided in the embodiments of the present application. During the inference process, a three-level cache structure consisting of a first cache, a second cache, and a third cache is provided to cooperate with the inference calculation process from receiving a user request to completing a response. By carefully designing communication strategies and scheduling strategies to optimize cache management and communication efficiency, the system performance is improved while minimizing the overhead caused by network communication, and the normal inference process of the model is not affected.

[0021] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below in conjunction with the accompanying drawings. However, it will be appreciated by those skilled in the art that in the embodiments of the present application, many technical details are provided in order to enable the reader to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical scheme claimed in the present application can be implemented.

[0022] The division of the following embodiments is for the convenience of description and shall not constitute any limitation on the specific implementation of the present application. The embodiments may be combined with each other and referenced to each other without contradiction.

[0023] The first aspect of the embodiment of the present application provides a reasoning method, which is applied to the corresponding system. The three-level cache structure formed by the first cache, the second cache and the third cache is applied in conjunction with the reasoning process of the artificial intelligence model, which optimizes the speed of memory access during the reasoning implementation process, reduces the data transmission overhead and time cost during the reasoning process, and further improves the reasoning efficiency. In order to facilitate those skilled in the art to better understand the reasoning method provided in the embodiment of the present application, it will be described below.

[0024] In some embodiments, Figure 1 As shown, the process of the reasoning method may include the following steps:

[0025] Step 101: receiving a user request, wherein the user request carries input information.

[0026] Step 102: Perform a match in the first cache according to the input information.

[0027] Step 103, based on the second cache, the matched token sequence and its KV parameters and input information are passed to the third cache, so that the inference model performs inference calculation based on the third cache, and the inferred token sequence and its KV parameters are passed to the first cache for storage.

[0028] In this way, in the process of implementing reasoning in response to user requests, firstly, the token sequence and its KV parameters stored in the first cache are matched, and then data is transferred between the first cache and the third cache for inference calculation through the second cache, so as to transfer the matched token sequence and its KV parameters and input information to the third cache, and transfer the token sequence and its KV parameters obtained by reasoning to the first cache for storage. While completing the reasoning, the token sequence and its KV parameters obtained by the current reasoning are stored for reuse in subsequent responses to user requests. In this way, through the coordination of the three-level storage structure of the first cache, the second cache and the third cache and the reasoning process, the speed of memory access in the reasoning implementation process is optimized, so that the data transmission overhead and time cost in the reasoning process are reduced, and further, the reasoning efficiency is also improved.

[0029] For ease of understanding, the following Figure 1 The steps shown are explained.

[0030] In step 101, a user request is received, wherein the user request carries input information. Among them, the embodiment of the present application does not limit the type of user request, which can be any request that can carry corresponding input information. For example, the user request can be a Hypertext Transfer Protocol (HTTP) request. For example, the user request can also be a Hypertext Transfer Protocol Secure (HTTPS) request to enhance the security of the request, etc., which are not listed here one by one. And the embodiment of the present application does not limit the input information, which can be any information that can trigger the artificial intelligence model to perform reasoning. For example, in a dialogue task, the input information can be "What is the weather like today"; for example, in image processing characters, the input information can be a picture to be processed, etc., which are not listed here one by one.

[0031] In step 102, a match is performed in the first cache according to the input information. The cache stores the token sequence and KV parameters corresponding to the user request to which the response has been completed (i.e., "KV Cache"). At this time, for the currently received user request, the input information it carries can also be converted into a corresponding token sequence. Therefore, the token sequence can be matched in the cache to obtain the token sequence stored in the cache and obtain its corresponding KV parameters. Of course, in some cases, it is also possible to directly match the input information carried by the user request to which the response has been completed and stored in the cache with the input information carried by the currently received user request, and further obtain the token sequence and KV parameters corresponding to the matched input information, etc., which will not be described one by one here.

[0032] It should be noted that the "match" in the embodiments of the present application is not a match of the complete input information or the complete token sequence, but a partial match. Taking the matching of input information as an example, assuming that the input information of the current user request is "ABCDE", if there is an input information "ABCHUIRHIOGVNERFD" carried by the user request that has been responded to, then it can be considered that the current user request matches "ABC", then the matched token sequence and its KV parameters, that is, the token sequence and its KV parameters corresponding to "ABC", will not be repeated here and in the future. Among them, the input sequence listed above only uses letters to represent the input information as an example, which is mainly for ease of understanding and has no other meaning.

[0033] It should also be noted that the token sequence matching involved in the embodiments of the present application mainly refers to the matching object involving the token sequence, rather than necessarily directly matching the token sequence itself. For example, when providing a corresponding identifier such as an ID number for a token, the identifier sequence corresponding to each token in the token sequence, i.e., tokenids, etc., can be directly used for matching, which will not be described one by one here.

[0034] In step 103, based on the second cache, the matched token sequence and its KV parameters and input information are passed to the third cache, so that the inference model performs inference calculation based on the third cache, and the inferred token sequence and its KV parameters are passed to the first cache for storage. It can be understood that after the inferred token sequence and its KV parameters (i.e., KVCache and its related data) are stored in the first cache, the subsequently received user requests can support the aforementioned matching steps in the first cache, thereby paving the way for subsequent user requests to reuse the data of the user requests that have been responded to, reduce the amount of calculation, and improve user efficiency.

[0035] It should be noted that the embodiments of the present application do not limit the specific process of how the second cache cooperates between the first cache and the third cache to implement reasoning. As long as the second cache acts as an intermediary between the first cache and the third cache to provide storage resources for the reasoning process, in some embodiments, the second cache can further cooperate with the prefill nodes and the decode nodes in terms of different focuses on computing-intensive and storage-intensive, support more flexible and diverse storage deployment solutions, and support more flexible and efficient reasoning calculations.

[0036] It is understandable that the prefill stage and the decode stage corresponding to the prefill node and the decode node have different focuses. Based on the second cache, the implementation process of the prefill stage and the implementation process of the decode stage can be differentiated, so that the prefill stage and the decode stage can be more fully decoupled and separated, so as to further make full use of relevant resources based on the differences between the prefill stage and the decode stage, improve resource utilization, and reduce data transmission overhead and time cost.

[0037] For example, based on the second cache, the matched token sequence, its KV parameters, and input information are passed to the third cache, so that the inference model performs inference calculation based on the third cache, and the inferred token sequence and its KV parameters are passed to the first cache for storage. This can be achieved in the following way: based on the first path, the matched token sequence, its KV parameters, and input information are passed to the third cache for prefill calculation, and the initial token and KV parameters obtained by the prefill calculation are stored in the first cache, wherein the first path is established based on the second cache; based on the second path, the initial token and KV parameters obtained by the prefill calculation in the first cache are passed to the third cache for decode calculation, and the inferred token sequence and its KV parameters obtained by the decode calculation are stored in the first cache, wherein the second path is established based on the second cache, and the first path is different from the second path.

[0038] It should be noted that the embodiments of the present application do not limit the specific paths of the first path and the second path, as long as they are formed based on the second cache and the first path and the second path are different (forming path differences to achieve decoupling of the prefill stage and the decode stage). For ease of understanding, the first path and the second path will be explained below from the perspective of specific step implementation, but this does not mean that the embodiments of the present application must or can only be implemented in the following manner and its corresponding path.

[0039] In some examples, based on the first path, the matched token sequence and its KV parameters and input information are passed to the third cache for prefill calculation, and the initial token and KV parameters calculated by prefill are stored in the first cache. This can be achieved by: storing the matched token sequence and its KV parameters and input information in the second cache; calling the prefill instance, based on the third cache, reading the matched token sequence and its KV parameters and input information from the second cache, and performing prefill calculation based on the matched token sequence and its KV parameters and input information; storing the initial token and KV parameters calculated by prefill in the first cache. In other words, the first path is: first cache-second cache-third cache-first cache.

[0040] In some examples, based on the second path, the initial token and KV parameters calculated by prefill in the first cache are passed to the third cache for decode calculation, and the inferred token sequence and its KV parameters calculated by decode are stored in the first cache. This can be achieved by: calling the decode instance, based on the third cache, reading the initial token and KV parameters calculated by prefill from the first cache, and performing decode calculation according to the initial token and KV parameters calculated by prefill; storing the inferred token sequence and its KV parameters calculated by decode in the second cache; synchronizing the inferred token sequence and its KV parameters updated on the second cache to the first cache. In other words, the second path is: first cache-third cache-second cache-first cache.

[0041] It is understandable that the above is only an example. As long as different cache call paths are set for the prefill stage and the decode stage, the prefill stage and the decode stage can be further decoupled in the storage, so as to achieve the purpose of more flexibly supporting the deployment scheme of inference models and storage.

[0042] It can also be understood that, in some embodiments, since the decode stage involves multiple token iterations, these iterations can be further decomposed into a combination of several single token iteration processes, so as to achieve more adequate resource scheduling based on the disassembled decode tasks and further improve the reasoning efficiency.

[0043] For example, based on the above embodiments, in some embodiments, the decode calculation based on the initial token and KV parameters calculated by prefill can be implemented in the following way: in the decode calculation round before the inferred token sequence and its KV parameters are calculated, each time a decode instance is called, the data obtained by the previous round of decode calculation is read from the first cache based on the third cache, and a round of decode calculation is performed, and the data obtained by the current round of decode calculation is stored in the first cache through the second cache; wherein the data obtained by the i-th round of decode calculation includes: the initial token and KV parameters obtained by the prefill calculation, and i tokens obtained by the first i rounds of decode calculation.

[0044] For ease of understanding, the above implementation scheme will be described from another perspective below. That is, based on the third cache, reading the data obtained from the previous round of decoding calculation from the first cache, and performing a round of decoding calculation can be implemented as follows: calling the decode instance, based on the third cache, reading the initial token and KV parameters obtained by prefill calculation from the first cache, and performing decoding calculation according to the initial token and KV parameters obtained by prefill calculation, can be implemented as follows: after the i-th round of decoding calculation is completed and the i+1th token is obtained, calling a decode instance in the decode cluster, based on the third cache corresponding to the currently called decode instance, reading the initial token and KV parameters obtained by prefill calculation from the second cache, and, the first i rounds of decoding calculations obtain i tokens, and according to the initial token and KV parameters obtained by prefill calculation, and, the first i rounds of decoding calculations obtain i tokens, and performing the i+1th round of decoding calculations, until the decoding calculation reaches the iteration stop condition, i=0.

[0045] Accordingly, in some examples, storing the data obtained by the current round of decode calculation in the first cache through the second cache can be achieved in the following way: after each round of decode calculation is completed, all tokens currently obtained by decode calculation, as well as the initial token and KV parameters obtained by prefill calculation are stored in the second cache; all tokens currently obtained by decode calculation stored in the second cache, as well as the initial token and KV parameters obtained by prefill calculation are synchronized to the first cache.

[0046] That is to say, only one token iteration task is assigned to a single decoding example each time, that is, the complete decoding reasoning is decomposed into multiple decoding calculation iterations to be distributed to different decoding instances, which is conducive to making full use of resources to support computationally intensive decoding calculations.

[0047] It can also be understood that, in some embodiments, since the prefill stage and the decode stage corresponding to the prefill node and the decode node have different focuses, different caches can be designed for the prefill stage and the decode stage to adapt to their corresponding data interaction requirements.

[0048] For example, the first cache includes a prefill cache pool and a decode cache pool. Matching is performed in the first cache according to the input information. This can be achieved in the following manner: matching is performed in the prefill cache pool according to the input information.

[0049] At the same time, based on the prefill cache pool and the decode cache pool, different cache interaction methods can also be involved in the subsequent prefill stage and decode stage. In some examples, the initial token and KV parameters calculated by prefill can be stored in the prefill cache pool and the decode cache pool; and, after each round of decode calculation, all tokens currently obtained by decode calculation, as well as the initial token and KV parameters calculated by prefill are stored in the second cache and synchronized to the decode cache pool. In other words, the prefill cache pool and the decode cache pool are used as independent individuals to correspond to the storage of the first cache involved in the prefill stage and the storage of the first cache involved in the decode stage.

[0050] In this way, by further decoupling the corresponding storage into two cache pools, namely the prefill cache pool and the decode cache pool, based on the differences between the prefill stage and the decode stage in the first cache, the prefill stage and the decode stage are also decoupled in the storage, making the concept of separation from PD more adaptable, thereby being able to more flexibly support the deployment scheme of inference models and storage.

[0051] In addition, in order to better support matching, the token sequence and its KV parameters obtained by inference can also be passed to the first cache for storage, which can be achieved in the following way: after the inference is completed, all tokens currently calculated by decode stored in the second cache, as well as the initial token and KV parameters calculated by prefill are stored in the prefill cache pool. In other words, the prefill cache pool and the decode cache pool are synchronized for the final data.

[0052] For ease of understanding, the following will combine the above-mentioned decoupling of the prefill cache pool and the decode cache pool, as well as the decoupling of the cache paths in the prefill stage and the decode stage for explanation.

[0053] After matching in the prefill cache pool according to the input information, the matched token sequence, its KV parameters, and input information are passed to the third cache based on the second cache in the following manner, so that the inference model performs inference calculation based on the third cache, and the inferred token sequence and its KV parameters are passed to the first cache for storage: Based on the second cache, the matched token sequence, its KV parameters, and input information are passed to the prefill instance to perform prefill calculation based on the third cache corresponding to the prefill instance; the initial token and KV parameters calculated by prefill are stored in the prefill cache pool and the decode cache pool; based on the third cache, the initial token and KV parameters calculated by prefill are read from the decode cache pool; after each round of decode calculation is completed, all tokens currently calculated by decode, as well as the initial token and KV parameters calculated by prefill are stored in the second cache and synchronized to the decode cache pool; after the inference is completed, all tokens currently calculated by decode, as well as the initial token and KV parameters calculated by prefill stored in the second cache are stored in the prefill cache pool.

[0054] Of course, the above embodiments or examples are mainly provided for ease of understanding. In some embodiments, other methods can be used to further cooperate with the PD separation architecture. For example, the resources used by the third cache to support the corresponding prefill stage and decode stage calculations are further decoupled to adapt to the reasoning requirements of different stages, thereby further improving the reasoning efficiency and reducing the data transmission overhead and time cost of reasoning.

[0055] In some embodiments, Figure 2 As shown, the process of the reasoning method may also include the following steps:

[0056] Step 201: receiving a user request, wherein the user request carries input information.

[0057] Step 202: Perform a match in the first cache according to the input information.

[0058] Step 203, based on the second cache, the matched token sequence and its KV parameters and input information are passed to the third cache, so that the inference model performs inference calculation based on the third cache, and the inferred token sequence and its KV parameters are passed to the first cache for storage.

[0059] Step 204: manage the token sequence and its KV parameters stored in the first cache according to the matching rate of the user request in the first cache.

[0060] In this way, on the basis of the above-mentioned embodiment, by guiding the management of the cache through matching, the first cache can be more adapted to the function of matching the multiplexed data, which is beneficial to improving the matching efficiency and reducing the management difficulty.

[0061] For ease of understanding, the following Figure 2 The steps shown are explained below. Figure 2 Steps 201 to 203 in the illustrated embodiment are Figure 1 Steps 101 to 103 shown in the illustrated embodiment are substantially the same, and the main difference is that: Figure 2 In the illustrated embodiment, step 204 is further executed after step 203, and steps 201 to 203 are not described in detail herein.

[0062] In step 204, the token sequence and its KV parameters stored in the first cache are managed according to the matching rate of the user request in the first cache. Among them, the embodiment of the present application does not limit the triggering conditions of management. It can be understood that the management of the token sequence and its KV parameters stored in the first cache can be a periodic trigger, for example, a management operation of the first cache is performed once a week or 100 hours. Of course, the management of the token sequence and its KV parameters stored in the first cache can also be based on preset conditions. Trigger, for example, whenever the redundancy ratio of the storage space of the first cache reaches a first preset value, or whenever it is not matched within a preset time and the frequency of being matched in the past response to user requests is less than a second preset value, etc., a management operation of the first cache is performed, etc., which will not be listed one by one here.

[0063] It should be noted that the management of the token sequence and its KV parameters stored in the first cache can be expansion, reduction, data deletion, etc., among which the update is mainly based on the token sequence and its KV parameters corresponding to the user request completed by the current response. Of course, in some cases, the update can also be guided based on the matching results, for example, indicating whether to retain the results of the previous update, etc., which will not be elaborated here.

[0064] The step division of the above methods is only for the purpose of clear description. When implemented, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this patent; adding insignificant modifications to the algorithm or process or introducing insignificant designs without changing the core design of the algorithm and process are all within the scope of protection of this patent.

[0065] On the other hand, the embodiment of the present application also provides a reasoning system, such as Figure 3 As shown, it includes: a prefill cluster, a decode cluster, a first cache, a second cache and a third cache.

[0066] Among them, the prefill cluster, the decode cluster, the first cache, the second cache and the third cache are used to implement the reasoning method described in any of the above embodiments.

[0067] In some embodiments, Figure 4 As shown, a third cache and at least one prefill instance in the prefill cluster and at least two decode instances in the decode cluster are deployed on the same GPU host.

[0068] In some embodiments, several GPU hosts correspond to the same first cache.

[0069] In some embodiments, the first cache is disposed in a central processing unit (CPU).

[0070] It is not difficult to find that this embodiment is a system embodiment corresponding to the method embodiment, and this embodiment can be implemented in conjunction with the method embodiment. The relevant technical details mentioned in the method embodiment are still valid in this embodiment, and in order to reduce repetition, they are not repeated here. Accordingly, the relevant technical details mentioned in this embodiment can also be applied in the method embodiment.

[0071] It is worth mentioning that all modules involved in this embodiment are logic modules. In practical applications, a logic unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, in order to highlight the innovative part of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed by this application, but this does not mean that there are no other units in this embodiment.

[0072] To facilitate those skilled in the art to better understand the reasoning method and system provided in the above embodiments, Figure 5 The reasoning system and its data flow process are explained.

[0073] like Figure 5 As shown, the inference system shows a CPU and two GPUs, wherein a first cache (further including a prefill cache pool and a decode cache pool) is deployed on the CPU, a GPU is set on a device directly connected to a user in a network, a third cache is deployed on a GPU, and a decode instance in a prefill cluster and several decode instances in a decode cluster are also deployed on the GPU. Therefore, based on the first cache and the second cache of the GPU, the memory access speed is the fastest, which can significantly improve the access efficiency. The prefill instance and decode instance implemented based on the third cache can support fast access to frequently accessed model parameters and data; the first cache is set as the central node to store intermediate calculation results, and data reuse is performed to improve the calculation speed, reduce repeated calculations, and further reduce overhead. It should be noted that Figure 5 It is provided primarily for ease of understanding. Figure 5 The complete prefill cluster, decode cluster, and other GPUs where the second cache may exist are not shown.

[0074] And, in Figure 5 The inference system shown also deploys a corresponding inference scheduler to schedule the prefill cluster and the decode cluster, the first cache, the second cache, and the third cache to cooperate in implementing the inference method described in the above embodiment. Specifically, the inference scheduler can be responsible for scheduling resource allocation coordination and data transmission between inference instances of each cache and inference cluster (prefill cluster, decode cluster), and / or responsible for data loading and storage. In the inference initialization phase, the most commonly used model parameters are loaded into the third cache to increase data access speed. Among them, for ease of implementation, in some cases, the inference scheduler sets a lightweight consistency protocol to ensure data consistency of the third cache between different nodes.

[0075] Based on the above description, the inference system forms a multi-level centralized root architecture from the inference cluster, inference scheduler layer, inference instance layer, GPU host layer and first cache upward in sequence, so that the system can be flexibly deployed according to different application scenarios and deployment requirements.

[0076] When performing inference, the process is as follows:

[0077] 1. After receiving a user request, the server in the network triggers the inference scheduler to initialize the structure in the system and allocate resources through the strategy.

[0078] 2. The inference scheduler transfers the user request to the CPU, and the user request first enters the prefill cache pool (also known as the "prefill kv catch block manager", the KV cache management block of the prefill stage) in the first storage of the CPU memory space. After matching part of the token sequence and its KV parameters in the prefill cache pool, the matched token sequence and its KV parameters are written into the second cache together with the unmatched part of the user request (i.e., the inputs composed of matches and no matches in the second cache). Among them, the prefill cache pool can use a prefix tree radixcache (i.e., the prefix tree uses a radix tree structure, wherein the update of the radix tree structure has a node splitting characteristic, therefore, the prefill cache pool will continuously split the token sequence and its KV parameters as the number of user requests responded to increases, thereby generating a smaller, corresponding token sequence shorter stage, making it easier to successfully match and more reusable) to provide the token sequence and its KV parameters inferred from the historical responses to user requests.

[0079] 3. The inference scheduler selects the corresponding prefill instance based on the scheduling strategy, and based on the selected prefill instance, writes the input in the second cache to the third cache of the GPU where the prefill instance is located, and calls the third cache to perform the corresponding prefill calculation to obtain the complete KV parameters and initial token (collectively referred to as "prefixkv", that is, the initial token and kv in the figure) corresponding to the user request. The prefix kv includes the matched kv (also called "match kv", that is, the KV parameters obtained by the previous match and reused in the current prefill calculation) and the unmatched kv (also called "no matchkv", that is, the one calculated by the current prefill instance).

[0080] 4. The inference scheduler schedules the third cache to write the prefix kv into the prefill cache pool and the decode cache pool (also called "decode kv catch block manager", the KV cache management block in the decode stage) of the first cache respectively.

[0081] 5. The inference scheduler selects the corresponding decode instance based on the scheduling strategy, and reads the initial token and kv from the decode cache pool of the first cache based on the selected decode instance and writes them into the third cache, and performs the corresponding decode calculation based on the corresponding decode instance to obtain the next token, that is, the token 1kv generated by the decode calculation (also called "generate token 1kv", where the number "1" is used to represent the first token calculated), and writes the initial token and the token 1kv generated by the kv+decode calculation into the second cache, and at the same time schedules the decode cache pool to read and store the initial token and the token 1kv generated by the kv+decode calculation from the second cache.

[0082] 6. The inference scheduler continues to select the corresponding decode instance based on the scheduling strategy, and reads the initial token and token 1kv generated by kv+decode calculation from the decode cache pool of the first cache based on the selected decode instance and writes them into the third cache, and performs the corresponding decode calculation based on the corresponding decode instance to obtain the next token, that is, token 2kv generated by decode calculation (also called "generate token 2kv"), and writes the initial token and token 1kv generated by kv+decode calculation + token 2kv generated by decode calculation into the second cache, and schedules the decode cache pool to read and store the initial token and token 1kv generated by kv+decode calculation + token 2kv generated by decode calculation from the second cache. In this way, the next token is continuously calculated until the inference corresponding to the current user request is completed.

[0083] 7. While writing the initial token and token 1kv generated by kv+decode calculation+token 2kv+…+token s kv generated by decode calculation in the second cache into the decode cache, where s is the number of tokens that need to be calculated in the decode stage before the inference is completed; the inference scheduler also schedules the initial token and token 1kv generated by kv+decode calculation+token 2kv+…+token s kv generated by decode calculation in the second cache to be written into the prefill cache. In addition, after obtaining the initial token and token 1kv+token 2kv+…+token s kv generated by kv+decode calculation, the decode cache pool also directly returns it as output to the user corresponding to the request.

[0084] In the above process of responding to user requests, the various structures in the system are also maintained. For example, detection alarm strategies are set for each prefill instance and decode instance to better schedule them, such as real-time monitoring of cache hit rate, and dynamic adjustment of cache content according to access frequency. When making adjustments, data such as weights in artificial intelligence model reasoning can be stored in the first cache, and updated according to the intelligent elimination algorithm when it is not needed or used infrequently, so as to improve the cache utilization rate and speed of use. For example, the lowest priority data is integrated with storage configuration and asynchronous synchronization mechanism such as LRU according to the intelligent elimination algorithm to improve the cache capacity and concurrency performance of the cluster. The communication between the above structures is based on the underlying NCCL (NVIDIA Collective Communications Library) protocol.

[0085] That is to say, the inference system is a distributed system that runs in layers, in which the largest whole is the entire inference cluster. The inference cluster consists of multiple inference instances, which are distributed on the GPU host entity. The GPU host entity has hierarchical caches, which are layered according to the speed of memory access and the size of storage. There are three layers in total, from bottom to top: the third cache (L1 Cache) is usually the fastest and smallest cache, directly close to the computing core; the second cache (L2Cache) is larger and slightly slower, usually the cache between the processor and the main memory; the first cache (DRAM and host memory) is larger and slower, usually a cache shared by multiple processing cores. Under the architecture centered on the first cache, the storage structure and top-level configuration of the three-layer cache are designed, and the central first cache is set in the host memory to maximize the use of the memory in the inference node. And design data sharing mechanism and transmission mechanism to transmit the intermediate amount in the large model inference process to the central node and share it through a high-speed network. Storing KV Cache related data in the second cache and accessing it through a high-speed network can effectively reduce the utilization of video memory and transmission delay. Therefore, data will be sent to different graphics cards according to the scheduling of the relevant scheduling program, and all data will be stored in a small cluster. This distributed solution can easily meet the needs of ultra-large service scenarios and greatly reduce costs. As mentioned above, data of the same type will be centrally stored according to program configuration, and combined with hash and other search algorithms, the overhead of data query will be greatly reduced.

[0086] In other words, providing a distributed multi-level cache structure with high adaptability based on PD separation solves the problems of distributed reasoning scenarios: slow inter-card communication, high-frequency memory access, high overhead, low memory utilization, and low cache reuse rate and hit rate of distributed reasoning clusters. Specifically, by establishing a multi-level cache structure, from bottom to top, the third cache (L1 Cache) is used to quickly access hot data, with the lowest latency, fastest speed, and closest to the core; the second cache (L2 Cache) is used to store intermediate results such as intermediate values ​​and intermediate variables of KV calculations, and its speed is slower than the first cache; the first cache (L3Cache) is used as a global cache for long-term storage of infrequently used data. This multi-level cache design is planned according to data and memory access priority, which can more effectively improve the cache hit rate, manage memory resources, and improve data access speed. It will no longer incur high operating costs, nor will it affect the real-time performance and response speed of the model due to resource limitations. Finally, the KV vector key-value pairs and intermediate variables are stored in the cache and high-speed communication area through a centralized system, thereby optimizing the memory access speed, meeting the needs of ultra-large parameter calculation scenarios, and solving the communication overhead and synchronization overhead in distributed reasoning scenarios. It can be effectively used to optimize large-scale reasoning clusters, and can also be modified and iterated according to different hardware architectures and application scenarios. For example, the present invention can be applied to distributed reasoning cluster solutions for multi-modality, diffusion models, etc.

[0087] Another aspect of the present application embodiment further provides an electronic device, such as Figure 6 As shown, it includes: at least one processor 601; and a memory 602 that is communicatively connected to the at least one processor 601; wherein the memory 602 stores instructions that can be executed by the at least one processor 601, and the instructions are executed by the at least one processor 601 so that the at least one processor 601 can execute the reasoning method described in any of the above method embodiments.

[0088] The memory 602 and the processor 601 are connected in a bus manner, and the bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors 601 and the memory 602 together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and are therefore not further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be one element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on a transmission medium. The data processed by the processor 601 is transmitted on a wireless medium via an antenna, and further, the antenna also receives data and transmits the data to the processor 601.

[0089] The processor 601 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management and other control functions. The memory 602 can be used to store data used by the processor 601 when performing operations.

[0090] Another aspect of the present application embodiment further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the above method embodiment is implemented.

[0091] That is, those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including a number of instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0092] Those skilled in the art will appreciate that the above embodiments are specific embodiments for implementing the present application, and in actual applications, various changes may be made thereto in form and detail without departing from the spirit and scope of the present application.

Claims

1. A reasoning method, characterized in that: include: receiving a user request, wherein the user request carries input information; According to the input information, matching is performed in the first cache; Based on the second cache, the matched token sequence and its KV parameters and the input information are passed to the third cache, so that the inference model performs inference calculation based on the third cache, and the inferred token sequence and its KV parameters are passed to the first cache for storage.

2. The inference method according to claim 1, characterized in that: Based on the second cache, the matched token sequence and its KV parameter and the input information are passed to the third cache, so that the inference model performs inference calculation based on the third cache, and the inferenced token sequence and its KV parameter are passed to the first cache for storage, including: Based on the first path, passing the matched token sequence and its KV parameter and the input information to the third cache for prefill calculation, and storing the initial token and KV parameter obtained by the prefill calculation into the first cache, wherein the first path is established based on the second cache; Based on the second path, the initial token and KV parameters obtained by prefill calculation in the first cache are passed to the third cache for decode calculation, and the inferred token sequence and its KV parameters obtained by decode calculation are stored in the first cache, wherein the second path is established based on the second cache, and the first path is different from the second path.

3. The inference method according to claim 2, characterized in that: Based on the first path, the matched token sequence and its KV parameter and the input information are passed to the third cache for prefill calculation, and the initial token and KV parameter obtained by the prefill calculation are stored in the first cache, including: Storing the matched token sequence and its KV parameters and the input information into the second cache; Calling the prefill instance, based on the third cache, reading the matched token sequence and its KV parameter, and the input information from the second cache, and performing prefill calculation according to the matched token sequence and its KV parameter, and the input information; The initial token and KV parameters calculated by prefill are stored in the first cache.

4. The inference method according to claim 2, characterized in that: Based on the second path, the initial token and KV parameters calculated by prefill in the first cache are passed to the third cache for decode calculation, and the inferred token sequence and KV parameters obtained by the decode calculation are stored in the first cache, including: Call the decode instance, read the initial token and KV parameters calculated by prefill from the first cache based on the third cache, and perform decode calculation according to the initial token and KV parameters calculated by prefill; The inferred token sequence and its KV parameters calculated by decode are stored in the second cache; The inferred token sequence and its KV parameters updated in the second cache are synchronized to the first cache.

5. The inference method according to claim 4, characterized in that: The decoding calculation is performed based on the initial token and KV parameters calculated by prefill, including: In a decode calculation round before calculating the inferred token sequence and its KV parameters, each time a decode instance is called, data obtained by the previous round of decode calculation is read from the first cache based on the third cache, and a round of decode calculation is performed, and data obtained by the current round of decode calculation is stored in the first cache through the second cache; The data obtained by the i-th round of decoding calculation includes: the initial token and KV parameters obtained by prefill calculation, and the i tokens obtained by the first i-th round of decoding calculation.

6. The inference method according to claim 5, characterized in that: The step of reading data obtained by a previous round of decoding calculation from the first cache based on the third cache and performing a round of decoding calculation includes: After the i-th round of decoding calculation is completed and the i+1th token is obtained, a decode instance in the decode cluster is called, and based on the third cache corresponding to the currently called decode instance, the initial token and KV parameters calculated by prefill are read from the second cache, as well as the i tokens obtained by the first i rounds of decoding calculation, and the i+1th round of decoding calculation is performed according to the initial token and KV parameters calculated by prefill, as well as the i tokens obtained by the first i rounds of decoding calculation, until the decoding calculation reaches the iteration stop condition, where i is a natural number.

7. The inference method according to claim 5, characterized in that: The storing the data obtained by the current round of decoding calculation into the first cache through the second cache includes: After each round of decode calculation is completed, all tokens currently calculated by decode, as well as the initial token and KV parameters calculated by prefill are stored in the second cache; All tokens currently calculated by decode and stored in the second cache, as well as the initial token and KV parameters calculated by prefill are synchronized to the first cache.

8. The inference method according to any one of claims 1 to 7, characterized in that: The first cache includes a prefill cache pool and a decode cache pool, and matching in the first cache according to the input information includes: According to the input information, matching is performed in the prefill cache pool; The method further comprises: The initial token and KV parameters calculated by prefill are stored in the prefill buffer pool and the decode buffer pool; After each round of decode calculation is completed, all tokens currently obtained through decode calculation, as well as the initial token and KV parameters obtained by prefill calculation are stored in the second cache and synchronized to the decode cache pool; The step of transferring the inferred token sequence and its KV parameters to the first cache for storage includes: After the inference is completed, all tokens currently calculated by decode and stored in the second cache, as well as the initial token and KV parameters calculated by prefill are stored in the prefill cache pool.

9. The inference method according to any one of claims 1 to 7, characterized in that: The method further comprises: The token sequence and the KV parameters thereof stored in the first cache are managed according to the matching rate of the user request in the first cache.

10. An inference system, characterized in that include: A prefill cluster, a decode cluster, a first cache, a second cache and a third cache, wherein the prefill cluster, the decode cluster, the first cache, the second cache and the third cache are used to implement the inference method as described in any one of claims 1 to 9.

11. The inference system according to claim 10, characterized in that: The third cache and at least one prefill instance in the prefill cluster and at least two decode instances in the decode cluster are deployed on the same GPU host, and / or several GPU hosts correspond to the same first cache.

12. An electronic device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the reasoning method according to any one of claims 1 to 9.

13. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the inference method according to any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Data processing system, method, device, medium and program product

    CN120315894A

  • Storage optimization method and device, electronic equipment and storage medium

    CN120371221A

  • Model reasoning method, electronic equipment and storage medium

    CN120450057A