Large language model reasoning KV prefetching method and system based on intelligent network card

By using the RDMA technology of smart network cards and GPU direct network card technology in the large language model inference system, an efficient KV cache data transmission path is built, and KV cache prefetch window is opened in the GPU, and the cache allocation strategy is dynamically optimized, which solves the problem of excessive demand for KV cache memory, and efficient KV cache prefetching and reuse is achieved, improving system performance and resource utilization efficiency.

CN120123112APending Publication Date: 2025-06-10CHINA UNIV OF GEOSCIENCES (WUHAN) +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510028110.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the inference process of large language models, the video memory demand of KV cache often exceeds the video memory usage of the model itself, resulting in high pressure on GPU video memory. Especially in the case of long sequences or concurrent requests, how to optimize the efficiency of KV cache usage under limited GPU video memory resources has become an important technical challenge.

Method used

By utilizing the RDMA technology provided by the smart network card and the GPU direct network card technology, an efficient KV cache data transmission path is built, and a KV cache prefetch window is opened in the GPU, a KV prefetch scheduling module is designed, and a cache allocation strategy is dynamically optimized to achieve efficient prefetch and multiplexing of KV caches.

Benefits of technology

It significantly reduces the transmission delay of KV cache, improves system performance, makes full use of GPU computing power and bandwidth resources, optimizes the efficiency of video memory resources, and solves the problem of high GPU memory pressure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123112A_ABST
    Figure CN120123112A_ABST
Patent Text Reader

Abstract

The invention discloses a large language model reasoning KV prefetching method and system based on an intelligent network card, and relates to the field of machine learning, and the large language model reasoning KV prefetching method based on the intelligent network card mainly comprises the steps: migrating a historical KV cache to a remote storage device, setting a KV prefetching window in a local GPU, and storing the KV prefetching window in the remote storage device; according to the current waiting model reasoning task information, determining the next historical KV cache called into the KV prefetching window, and directly fetching the historical KV cache on the remote storage equipment to the KV prefetching window; and reusing the historical KV cache of the KV prefetching window to execute the reasoning task, updating the KV cache data and the corresponding state, and updating the KV cache data and the corresponding state to the remote storage device. By means of the large language model reasoning KV prefetching method and system based on the intelligent network card, efficient utilization of video memory resources and improvement of system performance can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning, and more specifically, to a large language model inference KV pre-fetching method and system based on an intelligent network card. Background Art

[0002] With the rapid development of large language models (LLM) in recent years, large language models have been widely used in various industries. With their powerful natural language processing (NLP) capabilities, these models have played an important role in text generation, machine translation, sentiment analysis, question-answering systems, and other fields. However, with the increasing popularity of these large models, the huge number of inference requests has brought significant resource consumption pressure. How to improve model inference efficiency and reduce computing and storage costs has become an important issue.

[0003] In the reasoning process of large language models, the parallel computing capability of GPU plays a vital role. In the reasoning of large language models, especially deep neural networks like Transformer structures, a large number of matrix multiplications, additions, and activation function calculations are involved. Each computing core of the GPU can perform these operations independently, so a large amount of data can be processed in a short time, which greatly speeds up the reasoning speed. However, with the continuous expansion of the scale of LLM, especially the increase in the number of model parameters and layers, the GPU video memory required by the model has also increased significantly. Most current large language models use autoregressive decoding strategies for text generation. In this process, the model generates each token by iteratively, and uses the previously generated token as context input to predict the next token. Every time a new token is generated, one of the key tasks of the model is to calculate the attention mechanism. As the core of the current mainstream Transformer architecture, the attention mechanism aims to assign a weight to each input token to measure its relevance to other tokens. In autoregressive generation, each time a new token is generated, the model must recalculate the attention weights of all historical tokens, which results in each new token generation step having to access and recalculate the intermediate results of all historical tokens, especially when generating long texts, the amount of calculation increases exponentially.

[0004] A recognized solution is to cache the historical calculation results of the self-attention layer. In this way, the model can directly extract the already calculated key and value tensors from the cache without having to recalculate them each time. This method avoids repeated calculations and significantly reduces the computational overhead. However, the use of KV cache incurs additional GPU memory overhead, which increases with the length of the conversation request. In current practice (the application of multi-turn conversations has led to a substantial increase in context requirements), the memory requirement of the KV cache even exceeds the memory occupancy of the model itself in some cases. This problem has become a bottleneck restricting the application of large-scale models under the condition of limited existing GPU memory, especially when long sequences or concurrent requests need to be supported. Therefore, how to optimize the usage efficiency of the KV cache under limited GPU memory resources and balance computational performance and memory overhead is one of the important technical challenges that need to be solved urgently.

[0005] To solve this problem, an effective solution is to expand the memory capacity of the GPU and use an external dedicated storage server to solve the problem of the overly large historical KV cache, which can significantly reduce the GPU memory pressure. However, the main challenge faced by this solution is the transmission latency of the historical KV cache. Since in practical applications, the LLM faces randomly arriving requests and the LLM does not know the specific context KV cache required for the next request, when expanding the storage space outside the GPU, the inference task must go through two steps: 1. Retrieve the context KV cache information related to the current request and transfer it into the computing unit GPU, 2. Reuse this KV cache and execute the inference task. Obviously, although using an external storage server logically expands the memory space for the GPU to store KV data, the cost is an additional data transfer step. During the data transfer process, the GPU is actually idle, wasting precious computing power resources. And if the bandwidth and latency of the transmission path are relatively high, especially under high load conditions, the transmission rate may not meet the real-time requirements. At this time, the transmission delay of the KV cache may be greater than the delay of recalculating the KV, which is not worth the loss.

[0006] In response to this, some intelligent network cards (with more functions such as built-in hardware accelerators in addition to data transmission, such as DPU (Data Processing Unit)) provide Remote Direct Memory Access (RDMA) technology to support efficient data transmission. Specifically, RDMA is a technology that allows a computer to directly access the memory of a remote computer without going through CPU processing, thus significantly reducing network communication latency and CPU load. Obviously, combining the characteristics of intelligent network cards can effectively relieve the data transmission pressure between the external storage and the GPU.

[0007] However, one limitation of RDMA technology is that it only supports direct memory-to-memory transfers. This means that when KV cache data is transferred from a remote storage server to a local GPU, it must first be transferred from the memory of the storage server to local memory and then further transferred to the GPU memory for computing. Frequent data transfers and multi-level relays may lead to a reduction in resource utilization. Additionally, although RDMA greatly compresses the latency of data transfer, it does not completely solve the problem of GPU computing power being idle due to data waiting.

[0008] In summary, in order to achieve efficient inference for LLMs in a network environment, it is necessary to reasonably balance the limited video memory space of the GPU and the transmission cost of KV cache data, and make full use of the GPU computing power. Therefore, an efficient KV transmission system is needed to achieve low-cost KV cache reuse for LLMs, so as to meet the dual requirements of high-performance and resource utilization for efficient LLM inference. Summary of the Invention

[0009] The purpose of the present invention is to provide a method and system for KV prefetching for large language model inference based on a smart network card, which can achieve efficient utilization of video memory resources and improvement of system performance.

[0010] The present invention provides a method for KV prefetching for large language model inference based on a smart network card, including the following steps: S1: According to a request, use a central processing unit to obtain a waiting queue; the waiting queue includes the request; S2: According to the waiting queue, use a data processing unit and historical KV caches stored in a remote storage device to obtain a registry; S3: According to the waiting queue and the registry, reorder the task requests in the waiting queue, use a graphics processing unit to obtain a KV prefetch window, use a data processing unit to obtain the historical KV cache location; according to the historical KV cache location recorded in the registry and the size of the KV prefetch window in the graphics processing unit, use the data processing unit, remote storage device, technology of a GPU direct connection network card, and remote direct memory access method to perform KV prefetching; S4: According to the smart network card, graphics processing unit, data processing unit, and remote storage device, obtain a data path; use the remote direct memory access method and GPU direct connection network card technology to transfer the target KV cache data in the remote storage device to the KV prefetch window; S5: Use the central processing unit to submit the next inference request from the waiting queue, use the graphics processing unit to execute the inference task according to the KV prefetch window, and use the graphics processing unit to obtain updated KV cache data; S6: Transmit the updated KV cache data to the remote storage device to obtain an updated historical KV cache.

[0011] Further, the above registry includes entries, and each entry includes a dialogue window ID, a historical KV cache size, a KV cache storage address, a KV cache status, and a dirty bit; the dialogue window ID is used to uniquely identify a dialogue window, facilitating the management and retrieval of the corresponding KV cache; the historical KV cache size is used to record the total size of the historical KV cache of the corresponding dialogue window; the KV cache storage address is used to indicate the storage location of the KV cache on the remote storage device for use during prefetching, facilitating efficient data transfer via RDMA; the KV cache status is used to record and identify the current status of the KV cache, and the current status of the KV cache includes whether it has been prefetched to the GPU and whether prefetching is in progress; the dirty bit is used to indicate whether the KV cache has been updated on the GPU but not yet written back to the remote storage device.

[0012] Further, based on the historical KV cache location recorded in the registry and the KV prefetch window size in the graphics processor, using the technologies of the data processing unit, remote storage device, GPU direct connection network card, and the remote direct memory access method for KV prefetching, it includes a transmission stage; in the transmission stage, according to the GPU network card direct connection and RDMA technology, using a first-in-first-out prefetch scheduling strategy, the historical KV information of the dialogue window to which the request in the waiting queue belongs is transferred from the remote storage device to the graphics processor.

[0013] Preferably, based on the historical KV cache location recorded in the registry and the KV prefetch window size in the graphics processor, using the technologies of the data processing unit, remote storage device, GPU direct connection network card, and the remote direct memory access method for KV prefetching, it includes a transmission stage; in the transmission stage, according to the GPU network card direct connection and RDMA technology, using a dynamic scheduling strategy, the historical KV information of the dialogue window to which the request in the waiting queue belongs is transferred from the remote storage device to the graphics processor; the dynamic scheduling strategy gives priority to processing jobs with small data volumes and long calculation times.

[0014] Further, the above historical KV cache includes a KV cache index table, which is used to record the location and size of the KV cache data for subsequent retrieval; the KV prefetch window includes a KV data block index, which is used to ensure that the LLM request can correctly access the required historical KV cache.

[0015] The present invention also provides a large language model inference KV prefetching system based on a smart network card, including the following modules: a request queue module configured to obtain a waiting queue according to a request by using a central processing unit; the waiting queue includes requests; a registry maintenance module configured to obtain a registry according to the waiting queue by using historical KV caches stored in a data processing unit and a remote storage device; a KV prefetching module configured to reorder task requests in the waiting queue according to the waiting queue and the registry, obtain a KV prefetching window by using a graphics processing unit, and obtain historical KV cache positions by using a data processing unit; according to the historical KV cache positions recorded in the registry and the KV prefetching window size in the graphics processing unit, perform KV prefetching by using the techniques of a data processing unit, a remote storage device, a GPU direct connection network card, and a remote direct memory access method; a data path module configured to obtain a data path according to the smart network card, the graphics processing unit, the data processing unit, and the remote storage device; use the remote direct memory access method and the GPU direct connection network card technique to transfer target KV cache data in the remote storage device to the KV prefetching window; a request iteration module configured to submit the next inference request from the waiting queue by using the central processing unit, execute an inference task according to the KV prefetching window by using the graphics processing unit, and obtain updated KV cache data by using the graphics processing unit; a KV cache update module configured to transmit the updated KV cache data to the remote storage device to obtain an updated historical KV cache.

[0016] Further, the above-mentioned registry includes entries, and the entries include a dialogue window ID, a historical KV cache size, a KV cache storage address, a KV cache status, and a dirty bit; the dialogue window ID is used to uniquely identify a dialogue window, facilitating the management and retrieval of the corresponding KV cache; the historical KV cache size is used to record the total size of the historical KV cache of the corresponding dialogue window; the KV cache storage address is used to indicate the storage location of the KV cache on the remote storage device for use during prefetching, facilitating efficient data transmission through RDMA; the KV cache status is used to record and identify the current status of the KV cache, and the current status of the KV cache includes whether it has been prefetched to the GPU and whether it is being prefetched; the dirty bit is used to indicate whether the KV cache has been updated on the GPU but not yet written back to the remote storage device.

[0017] Further, the above-mentioned performing KV prefetching by using the techniques of a data processing unit, a remote storage device, a GPU direct connection network card, and a remote direct memory access method according to the historical KV cache positions recorded in the registry and the KV prefetching window size in the graphics processing unit includes a transmission stage; in the transmission stage, according to the GPU network card direct connection and RDMA techniques, use a first-in-first-out prefetching scheduling strategy to transfer the historical KV information of the dialogue window to which the request in the waiting queue belongs from the remote storage device to the graphics processing unit.

[0018] Preferably, according to the historical KV cache location recorded in the registry and the KV prefetch window size in the graphics processor, the KV prefetch is performed using the technologies of the data processing unit, remote storage device, GPU direct connection network card, and the remote direct memory access method, including the transmission stage; in the transmission stage, according to the GPU network card direct connection and RDMA technology, using the dynamic scheduling strategy, the historical KV information of the conversation window to which the request in the waiting queue belongs is transmitted from the remote storage device to the graphics processor; the dynamic scheduling strategy gives priority to processing jobs with small data volume and long calculation time.

[0019] Furthermore, the above historical KV cache includes a KV cache index table, which is used to record the location and size of the KV cache data for subsequent retrieval; the KV prefetch window includes a KV data block index, which is used to ensure that the LLM request can correctly access the required historical KV cache.

[0020] Implementing the large language model inference KV prefetch method and system based on the intelligent network card provided by the present invention has the following beneficial effects: The present invention utilizes the RDMA (Remote Direct Memory Access) technology provided by the DPU and the GPU and network card direct connection technology that bypasses the CPU to construct an efficient KV cache data transmission path, significantly reducing the impact of transmission latency on the model performance. At the same time, to make full use of the system computing power and bandwidth resources, the present invention opens a KV cache prefetch window in the GPU and designs a KV prefetch scheduling module. This scheduling module maintains a registry for recording the historical request information in the same conversation window and monitors the newly added pending inference requests in the work queue. When the new request belongs to a certain historical conversation window in the registry, the KV prefetch scheduling module will dynamically optimize the cache allocation strategy according to the historical context KV cache of the window and the capacity of the prefetch window, aiming to avoid waste of video memory caused by excessive prefetch or insufficient prefetch that causes the request to wait for cache loading during execution, thereby achieving efficient utilization of video memory resources and improvement of system performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings: Figure 1 is a flowchart of the large language model inference KV prefetch method based on the intelligent network card provided by the present invention; Figure 2 is an architecture diagram of the large language model inference KV prefetch system based on the intelligent network card provided by the present invention; Figure 3 is a schematic diagram of the KV prefetch strategy provided by the present invention; Figure 4 is a comparison diagram of the KV transmission system latency implemented by different data paths provided by the present invention; Figure 5 It is a delay comparison diagram of the KV prefetch system provided by the present invention. Detailed implementation manners

[0022] For a clearer understanding of the technical features, objectives, and effects of the present invention, the detailed implementation manners of the present invention will now be described in detail with reference to the accompanying drawings.

[0023] Figure 1 It shows a schematic diagram of the KV prefetch method for large language model inference based on a smart network card in this embodiment. In this embodiment, the KV prefetch method for large language model inference based on a smart network card includes the following steps: S1: According to the request, use the central processing unit to obtain a waiting queue; the waiting queue includes the request; As an exemplary embodiment, in step S1, first, the work queue on the central processing unit (CPU) will maintain all the arriving requests and wait for the large language model (LLM) to be idle before transmitting tasks to the graphics processing unit (GPU); S2: According to the waiting queue, use the historical KV cache stored in the data processing unit and the remote storage device to obtain a registry; As an exemplary embodiment, in step S2, the KV prefetch module on the data processing unit (DPU) will monitor the work queue in the CPU in real time and query the relevant information of the waiting requests (such as the subordinate relationship with historical requests, the dependent KV cache information, etc.); In an exemplary embodiment, the registry includes entries, and the entries include a dialogue window ID, a historical KV cache size, a KV cache storage address, a KV cache status, and a dirty bit; the dialogue window ID is used to uniquely identify a dialogue window to facilitate the management and retrieval of the corresponding KV cache; the historical KV cache size is used to record the total size of the historical KV cache of the corresponding dialogue window; the KV cache storage address is used to indicate the storage location of the KV cache on the remote storage device for use during prefetching to facilitate efficient data transmission through RDMA; the KV cache status is used to record and identify the current state of the KV cache, and the current state of the KV cache includes whether it has been prefetched to the GPU and whether it is being prefetched; the dirty bit is used to indicate whether the KV cache has been updated on the GPU but has not been written back to the remote storage device; S3: Reorder the task requests in the waiting queue according to the waiting queue and the registry, obtain the KV prefetch window using the graphics processing unit, and obtain the historical KV cache location using the data processing unit; according to the historical KV cache location recorded in the registry and the KV prefetch window size in the graphics processing unit, perform KV prefetch using the technologies of the data processing unit, remote storage device, GPU direct connection network card, and remote direct memory access method; In an exemplary embodiment, according to the historical KV cache location recorded in the registry and the KV prefetch window size in the graphics processing unit, perform KV prefetch using the technologies of the data processing unit, remote storage device, GPU direct connection network card, and remote direct memory access method, including a transmission stage; in the transmission stage, according to the GPU network card direct connection and RDMA technology, use the first-in-first-out prefetch scheduling strategy to transfer the historical KV information of the dialogue window to which the request in the waiting queue belongs from the remote storage device to the graphics processing unit; In a preferred embodiment, according to the historical KV cache location recorded in the registry and the KV prefetch window size in the graphics processing unit, perform KV prefetch using the technologies of the data processing unit, remote storage device, GPU direct connection network card, and remote direct memory access method, including a transmission stage; in the transmission stage, according to the GPU network card direct connection and RDMA technology, use the dynamic scheduling strategy to transfer the historical KV information of the dialogue window to which the request in the waiting queue belongs from the remote storage device to the graphics processing unit; the dynamic scheduling strategy gives priority to processing jobs with small data volume and long calculation time; As an exemplary embodiment, in step S3, according to the information of these waiting requests, the KV prefetch module will determine the next KV cache to be transferred into the KV prefetch window and notify the data reading and writing module to establish an RDMA connection between the remote storage device and the local; S4: Obtain the data path according to the intelligent network card, graphics processing unit, data processing unit, and remote storage device; use the remote direct memory access method and GPU direct connection network card technology to transfer the target KV cache data in the remote storage device to the KV prefetch window; In an exemplary embodiment, the historical KV cache includes a KV cache index table, which is used to record the location and size of the KV cache data for subsequent retrieval; the KV prefetch window includes a KV data block index, which is used to ensure that the LLM request can correctly access the required historical KV cache; As an exemplary embodiment, in step S4, perform data transmission, and directly fetch the KV cache on the remote storage device into the KV prefetch window of the local GPU through the data path directly connected by the DPU network card and the GPU; S5: Use the central processing unit to submit the next inference request from the waiting queue, use the graphics processing unit to execute the inference task according to the KV prefetch window, and use the graphics processing unit to obtain updated KV cache data; S6: Transmit the updated KV cache data to the remote storage device to obtain the updated historical KV cache.

[0024] This embodiment provides a large language model inference KV prefetch system based on a smart network card, including the following modules: a request queue module configured to: according to the request, use the central processing unit to obtain a waiting queue; the waiting queue includes requests; a registry maintenance module configured to: according to the waiting queue, use the historical KV cache stored in the data processing unit and the remote storage device to obtain a registry; a KV prefetch module configured to: according to the waiting queue and the registry, reorder the task requests in the waiting queue, use the graphics processing unit to obtain a KV prefetch window, and use the data processing unit to obtain the historical KV cache location; according to the historical KV cache location recorded in the registry and the KV prefetch window size in the graphics processing unit, use the data processing unit, the remote storage device, the technology of the GPU direct connection network card, and the remote direct memory access method to perform KV prefetch; a data path module configured to: according to the smart network card, the graphics processing unit, the data processing unit, and the remote storage device, obtain a data path; use the remote direct memory access method and the GPU direct connection network card technology to transmit the target KV cache data in the remote storage device to the KV prefetch window; a request iteration module configured to: use the central processing unit to submit the next inference request from the waiting queue, use the graphics processing unit to execute the inference task according to the KV prefetch window, and use the graphics processing unit to obtain updated KV cache data; a KV cache update module configured to: transmit the updated KV cache data to the remote storage device to obtain the updated historical KV cache; Further, the above registry includes entries, and the entries include a dialogue window ID, a historical KV cache size, a KV cache storage address, a KV cache status, and a dirty bit; the dialogue window ID is used to uniquely identify a dialogue window, facilitating the management and retrieval of the corresponding KV cache; the historical KV cache size is used to record the total size of the historical KV cache of the corresponding dialogue window; the KV cache storage address is used to indicate the storage location of the KV cache on the remote storage device for use during prefetching, facilitating efficient data transmission through RDMA; the KV cache status is used to record and identify the current state of the KV cache, and the current state of the KV cache includes whether it has been prefetched to the GPU and whether it is being prefetched; the dirty bit is used to indicate whether the KV cache has been updated on the GPU but has not been written back to the remote storage device.

[0025] Further, according to the historical KV cache location recorded in the registry and the KV prefetch window size in the graphics processor, the KV prefetch is performed using the technologies of the data processing unit, remote storage device, GPU direct-connected network card, and the remote direct memory access method, including the transmission stage; in the transmission stage, according to the GPU network card direct connection and RDMA technology, using the first-in-first-out prefetch scheduling strategy, the historical KV information of the dialogue window to which the requests in the waiting queue belong is transmitted from the remote storage device to the graphics processor.

[0026] Preferably, according to the historical KV cache location recorded in the registry and the KV prefetch window size in the graphics processor, the KV prefetch is performed using the technologies of the data processing unit, remote storage device, GPU direct-connected network card, and the remote direct memory access method, including the transmission stage; in the transmission stage, according to the GPU network card direct connection and RDMA technology, using the dynamic scheduling strategy, the historical KV information of the dialogue window to which the requests in the waiting queue belong is transmitted from the remote storage device to the graphics processor; the dynamic scheduling strategy preferentially processes jobs with small data volume and long calculation time.

[0027] Further, the above historical KV cache includes a KV cache index table, which is used to record the location and size of the KV cache data for subsequent retrieval; the KV prefetch window includes a KV data block index, which is used to ensure that the LLM request can correctly access the required historical KV cache.

[0028] As Figure 2 shown is a schematic diagram of the large language model inference KV prefetch system architecture based on an intelligent network card. In an exemplary embodiment, the large language model inference KV prefetch system based on an intelligent network card includes: 1. DPU (Data Processing Unit): The DPU is a dedicated data processor designed to offload, accelerate, and isolate data processing tasks to improve overall computing and communication efficiency; it strips these complex but core-computation-independent operations from the host CPU by taking over data-intensive tasks such as network transmission; the DPU supports Remote Direct Memory Access (RDMA) technology, allowing direct access to the memory of remote nodes and providing a high-speed direct data path between the GPU and the network interface card (network card), thus logically constructing the shortest data transmission path from the remote node memory to the local GPU video memory; different from the traditional network processing method that requires the host CPU to transfer, the DPU basically completely bypasses the processing of the host CPU; in large-scale inference tasks, this not only realizes a low-latency, high-bandwidth data transmission path but also saves a large amount of CPU performance resources; in addition, the DPU internally integrates a KV prefetch scheduling module, which mainly stores the storage index and related information of the historical KV cache in the remote storage device. 2. CPU: As the entry point of requests, all arriving requests need to undergo initial processing by the CPU and then enter the GPU for execution; the CPU is responsible for coordinating and managing the distribution of requests, maintaining the inference tasks in a waiting work queue, and ensuring the normal operation of the system. 3. GPU: In this system, the memory space of the GPU is mainly divided into three parts: the LLM model stacking space for storing the parameters and structures of large language models; the KV cache space required for the current task, which stores the KV cache data relied on by the ongoing inference tasks; and the KV cache prefetch window, which designates a portion of the GPU video memory space as the prefetch window to pre-store the context KV cache that future requests may need. 4. Remote storage device: An efficient storage device deployed on a remote memory server, such as remote non-volatile memory (NVM) and the idle memory device space of other remote servers. Taking NVM as an example, NVM provides a large amount of space specifically for storing the KV cache data relied on by historical requests; as an extension of the GPU video memory, it transfers data to the GPU through the high-speed transmission channel of the DPU when needed.

[0029] According to Figure 2 the information transmission path shown, in one embodiment, the KV prefetch method for large language model inference based on a smart network card includes: Step 1, first, the work queue on the Central Processing Unit (CPU) will maintain all arriving requests and wait for the Large Language Model (LLM) to be idle before transferring the task to the Graphics Processing Unit (GPU). Step 2, the KV prefetch module on the Data Processing Unit (DPU) will monitor the work queue in the CPU in real time and query the relevant information of the waiting requests (such as the subordinate relationship with historical requests, the KV cache information relied on, etc.). Step 3, based on the information of these waiting requests, the KV prefetch module will determine the next KV cache to be transferred into the KV prefetch window and notify the data reading and writing module to establish an RDMA connection between the remote storage device and the local. Steps 4 and 5, for data transmission, through the data path directly connecting the DPU network card and the GPU, directly fetch the KV cache on the remote storage device into the KV prefetch window of the local GPU. In step 6, after the KV prefetch window transmission is completed and the current inference calculation of the LLM is finished, the GPU fetches the next inference request from the CPU waiting queue and reuses the historical KV cache that has already been fetched in the KV prefetch window to execute the inference task; In steps 7 and 8, after the current inference task is completed, the updated KV cache data is recorded and passed to the remote storage device to update the historical KV cache for the corresponding window of this request; The entire system executes the transmission and calculation tasks according to this process until the waiting queue on the CPU is emptied.

[0030] In some embodiments, the above-mentioned method and system for KV prefetch in large language model inference based on a smart network card can also be implemented in the following manner.

[0031] I. Data Path Design The following will introduce the specific implementation process of the data path; First is the write process from the local to the remote storage device. To make the KV cache data suitable for RDMA transmission, when it is necessary to update the KV cache to the remote storage device after the inference is completed, the KV cache data is first uniformly serialized to make it suitable for RDMA transmission; since the KV cache may be scattered in different video memory blocks of the GPU, these data need to be integrated into a continuous memory space to ensure data continuity and consistency; and calculate the size of the serialized data to facilitate indicating the transmission data range during RDMA transmission; to improve the transmission efficiency, it is necessary to ensure that the data is aligned according to the cache line in memory, which helps to reduce the memory access overhead in RDMA transmission; after preparing the data to be transmitted, the memory area where the KV cache data is located is registered as a memory area accessible by the DOCA device; this operation is completed through the interface provided by the DOCA platform; using NVIDIA's GPUNetIO technology, the GPU video memory is directly mapped to the address space of the DPU; in this way, the DPU can directly access the GPU's video memory, avoiding the copy of data from the GPU to the host memory and realizing the direct connection of data between the GPU and the network card. After the registration is completed, the system will generate a Memory Key for access permission control of RDMA operations; the following is the traditional RDMA transmission process. Through RDMA technology, the GPU memory area actually pointed to by the DPU is mapped to the address space of the remote storage device to achieve direct data transmission; after the data is written, update the KV cache index table on the remote storage device to record the location and size of the new KV cache data for subsequent retrieval; During the read process from a remote storage device to the GPU, the transfer process is basically similar to the write process. It should be noted that the DPU will allocate a separate KV prefetch window on the GPU, and its size is determined according to the actual situation of the GPU video memory. The video memory area of the prefetch window is registered as a memory area accessible by the DOCA device. After the scheduling decision of the KV prefetch module is completed, the target KV cache data in the remote storage device needs to be directly transferred to the KV prefetch window of the GPU through RDMA. After the data transfer is completed, the KV data block index in the prefetch window is updated to ensure that the LLM request can correctly access the required historical KV cache. Through the above steps, by using RDMA and GPUNetIO technologies, efficient KV cache data transfer between the GPU, DPU, and remote storage device is achieved. The direct transfer of data between the GPU and the remote storage device avoids unnecessary memory copying and access overhead, improving the overall performance of the system. II. KV Prefetch Scheduling Module Specifically, a KV prefetch scheduling module is designed in the DPU of this system. The main function of this module is to maintain a registry structure for effectively managing and scheduling requests waiting in the CPU queue. The entries in the registry are for facilitating the retrieval of KV cache information stored on the remote storage device and can dynamically determine the KV cache prefetch policy based on this information. Each entry in the registry represents a dialogue window, which records the historical dialogue information and related KV cache information of this dialogue. For newly arrived requests, the system will create a new dialogue window entry for it in the registry. The design of the registry aims to facilitate the retrieval of KV cache information stored on the remote storage device and dynamically formulate the prefetch policy of the KV cache. Generally speaking, the registry needs to contain the following key fields: dialogue window ID, historical KV cache size, historical KV cache storage address, status tracking of the KV cache, and dirty bit management. The role of the dialogue window ID is to uniquely identify a dialogue window, facilitating the management and retrieval of the corresponding KV cache. The historical KV cache size will record the total size of the historical KV cache of this dialogue window. Since both the data transfer latency and the calculation of the KV prefetch window are directly related to the KV cache size, this field helps in formulating the prefetch policy and resource management. The KV cache storage address indicates the storage location of the KV cache on the remote storage device for use during prefetching, facilitating efficient data transfer through RDMA. The KV cache status record identifies the current status of the KV cache, such as whether it has been prefetched to the GPU, whether it is being prefetched, etc., facilitating cache management. The dirty bit is used to indicate whether the KV cache has been updated on the GPU but not yet written back to the remote storage device. When the cached data is modified, the dirty bit is set to True, indicating that the updated data needs to be written back at an appropriate time to ensure data consistency. The KV prefetch scheduling module also monitors the waiting request queue on the CPU in real time. Combining the hit information in the registry and the prefetch algorithm, it dynamically determines the optimal KV prefetch strategy. By comprehensively considering the size, status, and dependencies of the cache, the system can optimize the parallelism of data transmission and computation under the condition of limited prefetch window resources, improving the overall performance of the system. III. KV Prefetch Strategy After using the remote storage server as an external expansion of the GPU video memory, the inference task will be divided into two stages: the transmission stage, in which the historical KV information of the dialogue window to which the request belongs needs to be transmitted from the remote storage server to the local GPU, and the transmission delay is determined by the size of the historical KV cache and the network bandwidth; the computing stage refers to the computing delay when the request is actually executed, and the computing delay is mainly determined by the computing complexity and model complexity of the current request. The computing complexity refers to the current context length of the request and the expected possible generated token length. Any request should complete the transmission stage first and then the computing stage. If the current request does not belong to any historical dialogue window, the transmission stage delay of this request can be regarded as 0.

[0032] To deeply understand the impact of the KV prefetch strategy of the present invention on the system performance, the following simplified embodiment can be referred to. Assume that the delays of each stage of the request satisfy the conditions shown in Table 1: Table 1: Comparison Table of Delays of Each Stage of the Request

[0033] As shown in Table 1, the transmission delay and the computing delay are respectively positively correlated with the size of the historical KV information on which the request depends and the current computing complexity of the request (the network bandwidth and the model complexity are regarded as fixed constant values). To simplify the calculation process, here the network bandwidth is regarded as transmitting 1 unit of KV cache per unit time, that is, the size of the transmitted KV is numerically equivalent to the transmission KV delay. And the GPU prefetch window is limited, assumed to be 6 unit sizes in this embodiment. The execution memory in the GPU is separated from the prefetch window. When executing, the KV data is moved from the prefetch buffer to the execution memory, releasing the space of the prefetch buffer. And the KV data block must be transmitted in the form of the whole request and cannot be split, otherwise it will affect the data integrity and computing correctness. To demonstrate the superiority of the prefetch strategy of this system, as Figure 3 shown, for the above (a, b, c, d, e) requests, the no-prefetch strategy, the FIFO prefetch strategy, and the dynamic prefetch strategy used in this system will be respectively compared below; As Figure 3As shown in Case 1, this solution demonstrates the job scheduling situation without a prefetching strategy; in the figure, a1 represents the data transmission process of Request a, and a2 represents the calculation process of Request a; in this solution, only after both stages of Request a are completed can the LLM start to execute Request b; this results in an idle problem of bandwidth resources between the calculation stage and the transmission stage, that is, the calculation stage and the transmission stage need to wait for each other; therefore, the total time consumption of the system is equal to the linear sum of the two-stage times of all requests, and the final system completion time is 32 time units; As shown in Figure 3 Case 2, if a prefetch window and a First Input First Output (FIFO) prefetch scheduling strategy are adopted, the KV prefetch module in the DPU can monitor the dependency information of waiting tasks, so that when the previous task is in the calculation stage, the transmission stage of the next task can be executed synchronously, and the historical KV transmission of the request can be prefetched into the KV prefetch window. In this way, the final time consumption of the system is reduced to 23 time units; compared with Case 1, the FIFO prefetch strategy effectively alleviates the problem of system resource waste and improves the utilization rate of bandwidth and computing resources; however, this solution does not fully consider the problem of limited resources in the KV prefetch window and does not optimize according to the specific characteristics of the job, resulting in still existing "bubbles" during the parallel process of the calculation and transmission stages (for example, when the prefetch buffer is close to or reaches its capacity limit, the data transmission of subsequent jobs may be blocked, resulting in the need for the transmission process to wait for the buffer to release space), and the utilization rate of system resources is not fully exerted; In Figure 3 Case 3 of this invention, the designed prefetch window and dynamic scheduling strategy, by adjusting the job execution order, optimize the original order of a->b->c->d->e to d->b->e->a->c; the advantage of this adjustment is that it fully considers the capacity limitation of the GPU prefetch window and optimizes the parallel execution of data transmission and calculation; this solution preferentially processes jobs with smaller data volumes and longer calculation times, which can quickly release the space in the prefetch buffer and make enough space for the data transmission of subsequent large-data jobs, so that during the period when the GPU executes the calculation task, the prefetch module can continuously perform data transmission, maximizing the parallel utilization of bandwidth and computing resources, and reducing the resource idle time and waiting time in the system; compared with Solution 2, after optimized scheduling, the total time consumption of the system is further reduced to 21 time units, achieving the least overall system delay.

[0034] The experimental results show that the KV prefetching system with direct connection of GPU network cards can manage and schedule tasks more efficiently, optimize the interaction between data flow and computation flow, and provide a more efficient solution for the execution of large-scale inference tasks. While improving performance, this method ensures that the resources of the system are used more reasonably and efficiently, especially suitable for scenarios that need to process large-scale historical KV cache data. As Figure 4 shown, by comparing the latency performance of the LLM when exchanging historical KV data with remote storage devices, the performance differences of the three data paths can be clearly found: the traditional TCP network data path, the RDMA data path, and the direct connection data path of GPU network cards implemented in this invention. The experimental results show that as the historical KV cache (transmission data volume) increases, the direct connection data path of GPU network cards shows significantly better latency performance than the traditional network and RDMA data paths. The fundamental reason for this performance advantage is that both the traditional network path and the RDMA path involve the transfer between GPU video memory and CPU memory, while the generation location of the KV cache is within the GPU video memory, and the transfer path increases the additional transfer latency. In contrast, the direct connection data path of GPU network cards bypasses the transfer path from GPU to CPU and directly realizes the high-speed data transfer from GPU to remote storage. The experimental results show that the direct connection data path of GPU network cards has an average latency performance improvement of 65.21% compared with the traditional TCP network path and an average latency performance improvement of 38.68% compared with the native RDMA high-speed path. According to Figure 5 the data shown, the KV prefetching system designed based on the KV prefetching module realizes the efficient parallelization of historical KV cache data and the current request inference calculation task by reserving a limited KV prefetching window on the GPU. The system makes full use of the resources of the server and significantly alleviates the transmission latency problem caused by expanding the KV cache storage space by optimizing the synchronization of data transmission and calculation processes. In the comparison with the traditional scheme without a prefetching window, the scheme adopting the KV prefetching system has obtained a significant improvement in the overall latency performance. Specifically, the system finally realizes an average latency performance improvement of 60.34%. In summary, the prefetching module and data path design adopted in this system effectively reduce the total latency of tasks, especially in the data transmission stage. The shortening of the transmission path and the use of the prefetching window reduce the overall waiting time during task execution.

[0035] It should be noted that the above-mentioned remote storage device is an efficient storage device deployed on a remote memory server, which can be a remote non-volatile memory (NVM), or other memory devices and the idle memory device space of other remote servers, etc.

[0036] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these are within the protection scope of the present invention.

Claims

1. A large language model reasoning KV pre-fetching method based on smart network card, characterized in that: The following steps are involved: S1: according to the request, using the central processing unit to obtain a waiting queue; the waiting queue includes the request; S2: According to the waiting queue, a registration table is obtained by using the historical KV cache stored in the data processing unit and the remote storage device; S3: According to the waiting queue and the registry, reorder the task requests in the waiting queue, obtain the KV prefetch window by using the graphics processor, and obtain the historical KV cache position by using the data processing unit; according to the historical KV cache position recorded in the registry and the KV prefetch window size in the graphics processor, perform KV prefetch by using the data processing unit, remote storage device, GPU direct connection network card technology and remote direct memory access method; S4: obtaining a data path according to the smart network card, the graphics processor, the data processing unit and the remote storage device; transmitting the target KV cache data in the remote storage device to the KV pre-fetch window by using a remote direct memory access method and a GPU direct network card technology; S5: using the central processing unit to submit the next inference request from the waiting queue, using the graphics processor to perform the inference task according to the KV prefetch window, and using the graphics processor to obtain updated KV cache data; S6: The updated KV cache data is transmitted to a remote storage device to obtain an updated historical KV cache.

2. The method for pre-fetching KVs based on large language model reasoning based on smart network card according to claim 1 is characterized in that: The registry includes entries, and the entries include a dialog window ID, a historical KV cache size, a KV cache storage address, a KV cache state, and a dirty bit; the dialog window ID is used to uniquely identify a dialog window, so as to facilitate the management and retrieval of the corresponding KV cache; the historical KV cache size is used to record the total size of the historical KV cache of the corresponding dialog window; the KV cache storage address is used to indicate the storage location of the KV cache on the remote storage device for use during prefetching, so as to facilitate efficient data transmission through RDMA; the KV cache state is used to record and identify the current state of the KV cache, and the current state of the KV cache includes whether it has been prefetched to the GPU and whether it is being prefetched; the dirty bit is used to indicate whether the KV cache has been updated on the GPU but has not yet been written back to the remote storage device.

3. The method for pre-fetching KVs based on large language model reasoning based on smart network card according to claim 1, characterized in that: The method comprises: performing KV prefetching according to the historical KV cache position recorded in the registry and the KV prefetch window size in the graphics processor, using the data processing unit, the remote storage device, the GPU direct connection network card technology and the remote direct memory access method, including a transmission stage; In the transmission stage, based on the GPU network card direct connection and RDMA technology, a first-in-first-out pre-fetch scheduling strategy is used to transmit the historical KV information of the dialog window to which the request in the waiting queue belongs from the remote storage device to the graphics processor.

4. The method for pre-fetching KVs based on large language model reasoning based on smart network card according to claim 1, characterized in that: The method comprises: performing KV prefetching according to the historical KV cache position recorded in the registry and the KV prefetch window size in the graphics processor, using the data processing unit, the remote storage device, the GPU direct connection network card technology and the remote direct memory access method, including a transmission stage; In the transmission stage, based on GPU network card direct connection and RDMA technology, the historical KV information of the dialog window to which the request in the waiting queue belongs is transmitted from the remote storage device to the graphics processor using a dynamic scheduling strategy; The dynamic scheduling strategy gives priority to processing jobs with small data volume and long computing time.

5. The method for pre-fetching KVs based on large language model reasoning based on smart network card according to claim 1, characterized in that: The historical KV cache includes a KV cache index table, which is used to record the location and size of KV cache data for subsequent retrieval; the KV prefetch window includes a KV data block index, which is used to ensure that the LLM request can correctly access the required historical KV cache.

6. A large language model reasoning KV pre-fetching system based on smart network card, characterized in that: Includes the following modules: The request queue module is configured to: obtain a waiting queue using a central processing unit according to a request; the waiting queue includes requests; A registry maintenance module is configured to: obtain a registry according to the waiting queue using a historical KV cache stored in a data processing unit and a remote storage device; The KV prefetch module is configured to: reorder the task requests in the waiting queue according to the waiting queue and the registry, obtain the KV prefetch window by using the graphics processor, and obtain the historical KV cache position by using the data processing unit; perform KV prefetching by using the data processing unit, remote storage device, GPU direct connection network card technology and remote direct memory access method according to the historical KV cache position recorded in the registry and the KV prefetch window size in the graphics processor; The data path module is configured to: obtain a data path according to the smart network card, the graphics processor, the data processing unit and the remote storage device; The target KV cache data in the remote storage device is transferred to the KV pre-fetch window by using a remote direct memory access method and a GPU direct network card technology; A request iteration module is configured to: submit a next inference request from a waiting queue using a central processing unit, execute an inference task according to the KV prefetch window using a graphics processor, and obtain updated KV cache data using a graphics processor; The KV cache update module is configured to: transfer the updated KV cache data to a remote storage device to obtain an updated historical KV cache.

7. The large language model reasoning KV pre-fetching system based on smart network card according to claim 6 is characterized in that: The registry includes entries, and the entries include a dialog window ID, a historical KV cache size, a KV cache storage address, a KV cache state, and a dirty bit; the dialog window ID is used to uniquely identify a dialog window, so as to facilitate the management and retrieval of the corresponding KV cache; the historical KV cache size is used to record the total size of the historical KV cache of the corresponding dialog window; the KV cache storage address is used to indicate the storage location of the KV cache on the remote storage device for use during prefetching, so as to facilitate efficient data transmission through RDMA; the KV cache state is used to record and identify the current state of the KV cache, and the current state of the KV cache includes whether it has been prefetched to the GPU and whether it is being prefetched; the dirty bit is used to indicate whether the KV cache has been updated on the GPU but has not yet been written back to the remote storage device.

8. The large language model reasoning KV pre-fetching system based on smart network card according to claim 6 is characterized in that: The method comprises: performing KV prefetching according to the historical KV cache position recorded in the registry and the KV prefetch window size in the graphics processor, using the data processing unit, the remote storage device, the GPU direct connection network card technology and the remote direct memory access method, including a transmission stage; In the transmission stage, based on the GPU network card direct connection and RDMA technology, a first-in-first-out pre-fetch scheduling strategy is used to transmit the historical KV information of the dialog window to which the request in the waiting queue belongs from the remote storage device to the graphics processor.

9. The large language model reasoning KV pre-fetching system based on smart network card according to claim 6 is characterized in that: The method comprises: performing KV prefetching according to the historical KV cache position recorded in the registry and the KV prefetch window size in the graphics processor, using the data processing unit, the remote storage device, the GPU direct connection network card technology and the remote direct memory access method, including a transmission stage; In the transmission stage, based on GPU network card direct connection and RDMA technology, the historical KV information of the dialog window to which the request in the waiting queue belongs is transmitted from the remote storage device to the graphics processor using a dynamic scheduling strategy; The dynamic scheduling strategy gives priority to processing jobs with small data volume and long computing time.

10. The large language model reasoning KV pre-fetching system based on smart network card according to claim 6, characterized in that: The historical KV cache includes a KV cache index table, which is used to record the location and size of KV cache data for subsequent retrieval; the KV prefetch window includes a KV data block index, which is used to ensure that the LLM request can correctly access the required historical KV cache.

Citation Information

Cited By

  • Inference method and equipment for large language model

    CN121094148A

  • Hybrid KV cache management method and device, electronic equipment and storage medium

    CN121478678A

  • Large language model reasoning acceleration method on NCOS

    CN121684054A