Large language model processing system and session processing method

By introducing a storage node that interacts with the memory of the computing node hardware accelerator in the large language model processing system, and pre-caching the KV cache for multiple rounds of session requests, the problem of low hardware accelerator utilization is solved, and faster session processing and resource utilization efficiency are achieved.

CN120973527APending Publication Date: 2025-11-18BEIJING TENSOR LEAP TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511080375.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In the existing technology, the utilization rate of hardware accelerators is low, mainly because the key-value cache (KV cache) for multi-round session requests needs to be loaded from low-speed storage devices into the memory of hardware accelerators, resulting in excessive transmission latency and waiting time.

Method used

Before the compute nodes process session requests, by introducing storage nodes that interact directly with the hardware accelerator memory, the KV cache for multiple rounds of session requests is pre-cached in the first memory of the storage nodes, such as host memory, eliminating the transmission latency of fetching KV cache from low-speed storage devices.

Benefits of technology

This reduces the waiting time of computing nodes, increases the utilization rate of hardware accelerators, and improves the system's response speed and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973527A_ABST
    Figure CN120973527A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a large language model processing system and a session processing method. The system comprises a management node for deploying a scheduler, a computing node and a storage node, the scheduler is connected with the storage node, and the storage node is used for directly performing data interaction with a hardware accelerator memory of the computing node; the scheduler is used for receiving a session request, the session request is a multi-round session request, and an acquisition instruction is sent to the storage node; the storage node is used for acquiring a KV Cache of the session request and caching the KV Cache to a first memory of the storage node; the scheduler is also used for sending the session request to the computing node; and the computing node is used for obtaining the KV Cache from the first memory of the storage node after obtaining the session request, and processing the session request by using the KV Cache, so that transmission delay caused by obtaining the KV Cache across the computing nodes can be eliminated, the waiting time of the computing nodes is reduced, and the utilization rate of the hardware accelerator is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of large language models, in particular to a large language model processing system and a conversation processing method. BACKGROUND

[0002] Key-Value Cache (also known as KV Cache) is a key optimization technique adopted by large language models in the inference stage. By caching Key and Value matrices in the attention mechanism in the storage devices of the computing nodes, the efficiency of autoregressive generation is significantly improved.

[0003] Among them, the storage devices of the computing nodes include hardware accelerator memory, host memory, solid state disk and distributed shared storage device, etc. The access speed of the hardware accelerator memory is the fastest, but the capacity is limited and cannot store more KV Cache. Therefore, at present, a hierarchical caching strategy is often adopted, that is, the KV Cache of active conversations is written to the hardware accelerator memory or the host memory, and the KV Cache of inactive conversations is stored to the solid state disk or the shared network storage device, etc.

[0004] However, under the current hierarchical caching strategy, the utilization rate of the hardware accelerator is low. SUMMARY

[0005] Embodiments of the present application provide a large language model processing system and a conversation processing method for improving the utilization rate of the hardware accelerator.

[0006] In a first aspect, the embodiments of the present application provide a large language model processing system, which comprises a management node, a computing node and a storage node. The management node is deployed with a scheduler, the scheduler is connected to the storage node, and a first memory of the storage node is used for direct data interaction with the hardware accelerator memory of the computing node.

[0007] The scheduler is configured to receive a conversation request, the conversation request being a multi-turn conversation request, send an acquisition instruction to the storage node, the acquisition instruction indicating to acquire the Key-Value Cache (KV Cache) of the previous turn of the conversation request.

[0008] The storage node is configured to acquire the KV Cache of the previous turn of the conversation request, cache the KV Cache to the first memory of the storage node, and send a caching completion instruction to the scheduler after caching is completed.

[0009] The scheduler is further configured to send the conversation request to the computing node after receiving the caching completion instruction.

[0010] The computing node is configured to: acquire the KV Cache from the first memory of the storage node after the session request is acquired, and process the session request by using the KV Cache.

[0011] Optionally, the scheduler is further configured to:

[0012] If the session request is received and the session request is a first-round session request, the session request is sent to the computing node.

[0013] Optionally, the scheduler is specifically configured to:

[0014] The session request is sent to a message queue.

[0015] The computing node is configured to: if the computing node subscribes to the message queue, the computing node actively pulls the session request from the message queue.

[0016] Optionally, if the number of the computing nodes is n, and n1 computing nodes subscribe to the message queue, the n is a positive integer, the n1 is a positive integer, and the n1 is less than or equal to the n;

[0017] Each of the n1 computing nodes is specifically configured to:

[0018] According to the load state of the computing node, it is determined whether to pull a session request from the message queue;

[0019] If it is determined to pull a session request, the number of pulled session requests is determined;

[0020] The session request is pulled from the message queue, and the number of pulled session requests is less than or equal to the determined number.

[0021] Optionally, if the number of the computing nodes is n, the n is a positive integer, and the scheduler is specifically configured to:

[0022] After the cache completion instruction is received, a target computing node matched with the session request is determined from the n computing nodes;

[0023] The session request is sent to the target computing node.

[0024] Optionally, the computing node is further configured to:

[0025] After the session request is processed, a processing result is returned to the scheduler;

[0026] The scheduler is further configured to: output the processing result, so as to display the processing result on a user node where the session request is sent.

[0027] Optionally, the computing node is further configured to:

[0028] After processing the session request, the obtained new KV Cache of the session request is written back to the first memory of the storage node.

[0029] Optionally, the computing node is further configured to:

[0030] After processing the session request, the processing result is returned to the dispatcher.

[0031] The dispatcher is further configured to: send a synchronization data request to the storage node, the synchronization data request indicating that the new KV Cache is stored to a second memory of the storage node, the access speed of the second memory being less than that of the first memory.

[0032] The storage node is further configured to: store the new KV Cache to the second memory of the storage node.

[0033] Optionally, the storage node is a gd2fs cluster of an image processor direct connection distributed system or an image processor direct connection key-value gdkv database; a gd2fs in the gd2fs cluster supports GPU direct remote direct memory access (GPU Direct RDMA) technology, which is used to directly access the hardware accelerator memory of a remote node.

[0034] In a second aspect, an embodiment of the present application provides a session processing method applied to a management node of a large language model processing system, the management node deploying a dispatcher, the system further comprising a computing node and a storage node, the dispatcher being connected to the storage node, a first memory of the storage node being used for direct data interaction with a hardware accelerator memory of the computing node, and the method comprising:

[0035] After receiving a session request, and the session request being a multi-round session request, sending an acquisition instruction to the storage node, the acquisition instruction indicating acquisition of a key-value cache (KV Cache) of a previous round of the session request.

[0036] Receiving a cache completion instruction sent by the storage node, the cache completion instruction indicating that the storage node completes the following operations: acquiring the KV Cache and caching the KV Cache to the first memory of the storage node.

[0037] Sending the session request to the computing node, so that the computing node acquires the KV Cache from the first memory of the storage node after acquiring the session request, and processes the session request by using the KV Cache.

[0038] In a third aspect, an embodiment of the present application provides a computer program product, which comprises a computer program (also referred to as code or instructions), and when the computer program is executed, the computer program causes a computer to perform the method in any possible implementation manner of any of the aspects above.

[0039] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program (also referred to as code or instructions), and when the computer program is executed on a computer, the computer program causes the computer to perform the method in any possible implementation manner of the second aspect above.

[0040] In a fifth aspect, an embodiment of the present application provides a chip system, which comprises one or more processors for calling and executing instructions stored in a memory, so that the method in each aspect or any possible implementation manner of each aspect is executed. The chip system can be composed of a chip, or can include a chip and other discrete devices.

[0041] Embodiments of the present application provide a large language model processing system and a conversation processing method. The system comprises a management node of a deployment scheduler, a computing node and a storage node, the scheduler is connected to the storage node, and the storage node is used for directly interacting with the hardware accelerator memory of the computing node; the scheduler is used for: receiving a conversation request, the conversation request being a multi-turn conversation request, sending an acquisition instruction to the storage node, the acquisition instruction indicating to acquire a key-value cache (KVCache) corresponding to the conversation request; the storage node is used for: acquiring the KV Cache of the conversation request, caching the KV Cache to a first memory of the storage node, and sending a caching completion instruction to the scheduler; the scheduler is further used for: after receiving the caching completion instruction, sending the conversation request to the computing node; the computing node is used for: after acquiring the conversation request, acquiring the KV Cache from the first memory of the storage node, and processing the conversation request by using the KV Cache.

[0042] The embodiment of the application introduces a storage node interacting with the hardware accelerator memory of the computing node in the large language model processing system. Before the computing node processes the session request, if the session request is a multi-round session request, the KV Cache of the currently stored session request is pre-loaded to the first memory of the storage node. The first memory of the storage node directly interacts with the hardware accelerator memory of the computing node. For example, the first memory of the storage node is the host memory. That is, the first memory of the storage node can write data to the hardware accelerator memory of the computing node, and can also read data from the hardware accelerator memory of the computing node. When the computing node processes the session request, the KV Cache can be directly read from the first memory of the storage node. This way eliminates the transmission delay caused by the computing node obtaining the KV Cache from other low-speed storage devices, such as solid state disks or shared network storage devices, and also eliminates the transmission delay caused by obtaining the KV Cache across computing nodes, reduces the waiting time of the computing node, and thereby improves the utilization rate of the hardware accelerator. BRIEF DESCRIPTION OF DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0044] Figure 1 It is a structural schematic diagram of a large language model processing system.

[0045] Figure 2 It is a hierarchical management KV Cache structural schematic diagram of a computing node.

[0046] Figure 3 It is a structural schematic diagram of a large language model processing system provided by the embodiment of the present application.

[0047] Figure 4 It is a structural schematic diagram of a storage node 301 provided by the embodiment of the present application.

[0048] Figure 5 It is an interaction diagram of a session processing method provided by the embodiment of the present application.

[0049] Figure 6 It is an application schematic diagram of a large language model processing system provided by the embodiment of the present application. DETAILED DESCRIPTION

[0050] In the following, the technical solutions in the embodiments will be described clearly and completely in combination with the drawings in the embodiments, obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0051] Firstly, the application scenarios involved in the embodiments of the present application are introduced.

[0052] Exemplarily, the accompanying drawings show a structural schematic diagram of a large language model processing system, which comprises a management node 101, a user node 102 and a computing node 103. The user node 102 is in communication connection with a scheduler, and the scheduler is used to schedule the computing node 103 to perform large language model inference. Figure 1 Exemplarily, as shown in the accompanying drawings, the user node 102 comprises a question and answer interface, and a user can input a session request through the question and answer interface. For example, the session request input by the user is "write an article about large language model".

[0053] The user node 102 is used to receive the session request input by the user and forward it to the scheduler. Exemplarily, as shown in the accompanying drawings, the user node 102 comprises a question and answer interface, and a user can input a session request through the question and answer interface. For example, the session request input by the user is "write an article about large language model". Figure 1

[0054] It should be noted that the user node 102 is a display device comprising a question and answer node, for example, it can be a notebook computer, a mobile phone or other electronic device, and the embodiments of the present application are not specifically limited. The number of user nodes 102 can be one or multiple, which is adjusted according to the needs.

[0055] The session request can be a natural language processing task such as text generation, question and answer, translation, etc.

[0056] The management node 101 is deployed with a scheduler, which is used to manage and schedule the computing resources of the whole large language model. Specifically, the scheduler can schedule the computing node 103 to process the session request.

[0057] The computing node 103 is a hardware entity for performing large language model inference, which comprises a hardware accelerator such as a neural processing unit (NPU) or a graphics processing unit (GPU). The hardware accelerator is used to perform inference on the session request.

[0058] The computing node 103 further comprises a storage device for managing the KV Cache corresponding to the session request.

[0059] Exemplarily, the accompanying drawings show a structural schematic diagram of a large language model processing system, which comprises a management node 101, a user node 102 and a computing node 103. The user node 102 is in communication connection with a scheduler, and the scheduler is used to schedule the computing node 103 to perform large language model inference. Figure 2 ​A schematic diagram of a hierarchical management KV Cache structure of a computing node. The computing node 102 includes hardware accelerator memory, host memory, local solid state drive (SSD), and shared network storage device.

[0060] Among them, the access speed from high to low, in turn, is: hardware accelerator memory, host memory, local SSD and shared distributed network storage device. The cache capacity from small to large is hardware accelerator memory, host memory, local SSD and shared distributed network storage device.

[0061] In the embodiments of the present application, the computing node writes the KV Cache of the current active session to the hardware accelerator memory or the host memory, and stores the KV Cache of the inactive session to the local SSD or the shared network storage device and the like.

[0062] The active session refers to the time interval between the current session and the closest session is less than a preset time interval threshold (for example, 1s or 2min, etc.). This means that the user interacts with the model multiple times in a preset period of time, and the time interval between each question is short, and the model needs to keep the context information of the session in order to provide coherent and targeted answers.

[0063] The inactive session refers to the time interval between the session and the closest session is greater than or equal to the preset time interval threshold. This means that the user stops interacting with the model for a certain period of time, and the session is considered to be inactive. The inactive session is unlikely to receive new requests in a short period of time, so moving its KV Cache to a lower speed storage device, such as a local SSD or a shared network storage device, can free up space for new active sessions, while reducing storage costs and improving overall resource utilization.

[0064] It should be noted that the number of computing nodes 103 can be 1 or multiple. For example, as shown in FIG. 1B, the number of computing nodes 103 is n, which are computing nodes #1-#n, where n is a positive integer. Figure 1

[0065] However, this hierarchical management of KV Cache will affect the utilization rate of the hardware accelerator.

[0066] ​The inventor has found that the reasons for low utilization of the hardware accelerator include: if the session request is a multi-round session request, the KV Cache of the session request is located in a low-speed storage device, for example, the KV Cache is located in a local SSD or a shared network storage device, at this time, the computing node needs to load the KV Cache from the low-speed storage device to the host memory, and then load it from the host memory to the hardware accelerator memory. For example, the KV Cache is loaded from the shared network storage device to the local SSD, and then from the local SSD to the host memory, and from the host memory to the hardware accelerator, which will bring high transmission delay, that is, the hardware accelerator of the computing node needs to wait for a long time of data loading delay before processing the session request, which will inevitably affect the utilization of the hardware accelerator.

[0067] Further, for a multi-round session request, if the KV Cache of the session request exists on a first computing node, when the next round of session is received, if the next round of session request is processed on a second computing node, at this time, the second computing node needs to obtain the KV Cache from the first computing node before processing the session request, which will inevitably affect the utilization of the hardware acceleration of the second computing node.

[0068] Therefore, the embodiments of the present application provide a large language model processing system, by introducing a storage node that interacts with the hardware accelerator memory of the computing node, before the computing node processes the session request, if the session request is a multi-round session request, the storage node obtains the KV Cache of the session request and loads the KV Cache to the first memory of the storage node. The first memory of the storage node directly interacts with the hardware accelerator memory of the computing node. For example, the first memory of the storage node is the host memory. That is, the first memory of the storage node can write data to the hardware accelerator memory of the computing node, and can also read data from the hardware accelerator memory of the computing node. Then, when the computing node processes the session request, it can directly read the KV Cache from the first memory of the storage node. This way eliminates the transmission delay caused by the computing node obtaining the KV Cache from other low-speed storage devices, such as solid state drives or shared network storage devices, and also eliminates the transmission delay caused by obtaining the KV Cache across computing nodes, reduces the waiting time of the computing node, and improves the utilization of the hardware accelerator.

[0069] The large language model processing system provided by the embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0070] The storage node is a computing node that interacts with the hardware accelerator memory of the computing node. Figure 3A structural schematic diagram of a large language model processing system provided in an embodiment of the present application is shown. The system 30 includes a management node 101 of a deployment scheduler, a computing node 103, and a storage node 301. The scheduler is in communication connection with the storage node 301.

[0071] In an embodiment of the present application, the storage node 301 includes a first memory, wherein the first memory is used for direct data interaction with the hardware accelerator memory of the computing node 103. That is, the first memory of the storage node 301 can read data from the hardware accelerator memory of the computing node 103, or write data to the hardware accelerator memory of the computing node 103.

[0072] In a specific implementation, as shown in Figure 4 A structural schematic diagram of a storage node 301 provided in an embodiment of the present application is shown. The storage node 301 includes one host memory and a plurality of local SSDs. The host memory is the first memory, and the plurality of local SSDs are also referred to as the second memory.

[0073] Exemplarily, Figure 4 Three local SSDs are shown, namely SSD X, SSD Y, and SSD Z. The host memory can interact with the three local SSDs. In a set of embodiments of the present application, the storage node 301 is used for data interaction with the computing node 103, and specifically for data interaction with the hardware accelerator memory (for example, GPU memory) of the computing node. The hardware accelerator of the computing node 103 is exemplarily described by taking a GPU as an example.

[0074] In an example, the large language model processing system includes a GPU direct distributed file system cluster (gd2fs Cluster), and the storage node 103 is a gd2fs node of the gd2fs Cluster. The gd2fs node supports GPU direct remote direct memory access (GPU Direct RDMA) technology, which is used to directly access the hardware accelerator memory of a remote node. In another example, the storage node 103 can also be a GPU direct key value (gdkv) database. The storage node 103 can also be other components, which are not specifically limited in the embodiments of the present application.

[0075] In an embodiment of the present application, the scheduler is used to obtain session requests. Specifically, the scheduler can obtain a plurality of session requests from a plurality of user nodes at the same time. For example, the scheduler can obtain 3 session requests from 3 user nodes at the same time.

[0076] The scheduler is further configured to determine whether the session request is a multi-turn session request. The session request includes two types, namely a multi-turn session request and a first-turn session request.

[0077] The first-turn session request refers to the first request sent by the user when starting a new interaction with the large language model processing system. In such a request, the large language model processing system has no historical information or context about the user session. That is, there is no KV Cache cached for the session request for the large language model processing system.

[0078] The multi-turn session request refers to the subsequent request sent by the user after the first request to the large language model processing system in a continuous interaction. In these requests, the model can use the accumulated session history and context information, especially the stored KV Cache, to process the session request more quickly.

[0079] In a specific implementation, the scheduler is further configured to determine, according to the existing session request information stored in the database, whether the received session request is a multi-turn session request. In the embodiment of the present application, the database is used to store the session request information at the current time and within a preset time period before the current time. The session request information is used to uniquely identify the session request, for example, the session request information includes an identifier of the session request. The scheduler is configured to determine, by matching, whether the received session request exists in the database. If so, it is determined that the received session request is a multi-turn session request. If the database does not exist, it is determined that the session request is a first-turn session request.

[0080] The scheduler is further configured to send a pre-reading request to the storage node 301 if the session request is a multi-turn session request. Correspondingly, the storage node 301 is configured to receive the pre-reading request. The pre-reading request is used to obtain the KV Cache of the session request.

[0081] In the embodiment of the present application, the storage node 301 is configured to, after receiving the pre-reading request, obtain the KV Cache of the session request and store the KV Cache to the first local memory.

[0082] In an example, for Figure 4 As shown in FIG. 3, if the session request is located in the SSD of the storage node 301, the storage node 301 is configured to obtain the KV Cache of the session request from the local SSD and store the KV Cache of the session request to the local host memory.

[0083] In another example, if the session request is located in the SSD or shared network storage of other nodes, the storage node 301 can also be configured to read the KV Cache from other nodes, for example, from the SSD in the computing node 103, and store the KV Cache to the first memory of the storage node 103.

[0084] In the embodiments of the present application, the scheduler can also obtain the storage path of the KV Cache of the session request from the database, and send the storage path to the storage node 301. The storage node 301 is configured to read the KV Cache of the session request from the storage device corresponding to the storage path based on the storage path, and store it to the local memory of the storage node 301.

[0085] The storage node 301 is also configured to send a cache completion instruction to the scheduler after the cache is completed. Correspondingly, the scheduler is also configured to receive the cache completion instruction and send the session request to the computing node 103. The cache completion instruction is used to notify the scheduler that the storage node 301 has completed the cache.

[0086] The computing node 103 is configured to read the KV Cache of the session request from the first memory of the storage node 301 after receiving the session request, and process the session request in combination with the KV Cache of the session request. It can be understood that the first memory of the storage node 301 has pre-read the KV Cache of the session request, so the storage node 301 can quickly transmit the KV Cache of the session request to the hardware accelerator memory of the computing node. Therefore, the computing node 301 does not need to wait for a long time for the KV Cache to be loaded into the hardware accelerator memory, which helps to improve the utilization rate of the hardware accelerator.

[0087] In the embodiments of the present application, if the computing node 103 is n, and n is a positive integer. The scheduler 103 can first match the target computing node corresponding to the session request from the n computing nodes according to a preset matching strategy, and then the scheduler 103 schedules the target computing node to process the session request. The preset matching strategy can be a load balancing strategy, that is, the scheduler determines the load state of the n computing nodes, and selects the computing node with the smallest load as the target computing node to execute the session request. The preset matching strategy can also be a polling strategy, for example, the scheduler 103 schedules one of the n computing nodes to process a session request at a time. In addition, the preset matching strategy can also be other strategies, which can be adjusted by those skilled in the art as needed.

[0088] In another example, the scheduler 103 can send the session request to a message queue. If n1 computing nodes of the n computing nodes are scheduled to messages of the message queue, each of the n1 computing nodes performs the following manner: calculates its own load state, determines whether to actively pull the session request according to its own load state, if it is determined to actively pull, and further determines the number of active pulls. For example, the determined number of active pulls is n2, then the computing node actively pulls the session request from the message queue, and the number of pulls is less than or equal to n2. Wherein, n1 is a positive integer, and n1 is less than or equal to n, and n2 is also a positive integer.

[0089] The message queue is used as middleware, and the computing node actively pulls the task according to its own load condition, which avoids the load imbalance problem that may be caused by direct request scheduling, avoids the idle and overload of the computing server, and improves the overall processing capacity and stability.

[0090] Further, the computing node 103 is configured to: after processing the session request, return the processing result to the scheduler. Correspondingly, the scheduler is configured to receive the processing result, and send the processing result to the user node that sends the session request for display.

[0091] In addition, the computing node 103 is also configured to: after processing the session request, the newly obtained KV Cache is written back to the first memory of the storage node 301. For example, as shown in Figure 4 The computing node 103 is configured to write back the new KV Cache in the GPU memory of the computing node 103 to the host memory of the storage node 301. This way can not need to cache the KV Cache to other storage devices of the computing node 103, and only needs to write data to the host memory of the storage node 301, so as to help improve the writing speed.

[0092] Further, the storage node 301 is also configured to: persist the new KV Cache in the host memory to the local SSD of the storage node. Thus, when the next round of session request is performed, the storage node directly migrates the KV Cache of the local SSD to the local host memory, so as to help improve the migration speed and further reduce the user waiting time.

[0093] In a specific implementation, the scheduler is also configured to: after receiving the processing result sent by the computing node 103, send a synchronization data request to the storage node 301, and the storage node 301 is configured to persist the new KV Cache in the host memory to the local SSD of the storage node.

[0094] Further, the storage node 301 is configured to return a data synchronization result to the dispatcher. For a session request successfully synchronized with data, the dispatcher updates the database and stores the latest KV Cache information in the database. In this way, the database and the storage node can be accurately matched, so that the KV Cache can be accurately obtained in the next round of session requests.

[0095] Further, if the session request is the first round of session request, the KV Cache information of the session request can be inserted in the database.

[0096] The embodiment of the application provides a large language model processing system. By introducing a storage node interacting with the hardware accelerator memory of a computing node, if the session request is a multi-round session request, that is, the KV Cache of the session request currently exists in the second memory of the storage node, for example, is stored in the solid state disk of the storage node, the storage node loads the KV Cache from the second memory to the first memory of the storage node. The first memory of the storage node directly interacts with the hardware accelerator memory of the computing node. For example, the first memory of the storage node is a host memory. That is, the first memory of the storage node can write data to the hardware accelerator memory of the computing node, and can also read data from the hardware accelerator memory of the computing node. Then, when the computing node processes the session request, the KV Cache can be directly read from the first memory of the storage node. In this way, the transmission delay caused by the computing node obtaining the KV Cache from other low-speed storage devices, for example, a solid state disk or a shared network storage device, is eliminated, and the transmission delay caused by obtaining the KV Cache across the computing node is also eliminated, the waiting time of the computing node is reduced, and the utilization rate of the hardware accelerator is improved.

[0097] The above introduces a large language model processing system, and the following introduces a specific application of the large language model processing system. Figure 5 A session processing method interaction diagram is provided for the embodiment of the application, the method is applied to the large language model processing system shown in the accompanying drawings, and the method includes the following contents. Figure 1 The method includes the following contents.

[0098] S510, the dispatcher receives a session request, and the session request is a multi-round session request, sends an acquisition instruction to the storage node, and the acquisition instruction indicates acquisition of a key value cache (KV Cache) corresponding to the session request.

[0099] S520, the storage node acquires the KV Cache of the last round of the session request, caches the KV Cache to the first memory of the storage node, and sends a caching completion instruction to the dispatcher.

[0100] S530, after receiving the cache completion instruction, the scheduler sends the session request to the computing node.

[0101] S540, after the computing node obtains the session request, the computing node obtains the KVCache from the first memory of the storage node, and processes the session request by using the KVCache.

[0102] Optionally, if the scheduler receives the session request and the session request is the first round session request, the scheduler sends the session request to the computing node.

[0103] Optionally, the scheduler sends the session request to a message queue; if the computing node subscribes to the message queue, the computing node actively pulls the session request from the message queue.

[0104] Optionally, if the number of the computing nodes is n, and n1 computing nodes subscribe to the message queue, the n is a positive integer, the n1 is a positive integer, and the n1 is less than or equal to the n;

[0105] Each of the n1 computing nodes performs:

[0106] According to the load state of the computing node, it is determined whether to pull the session request from the message queue;

[0107] If it is determined to pull the session request, the number of the pulled session requests is determined;

[0108] The session request is pulled from the message queue, and the number of the pulled session requests is less than or equal to the determined number.

[0109] Optionally, if the number of the computing nodes is n, the n is a positive integer, after receiving the cache completion instruction, the scheduler determines a target computing node matched with the session request from the n computing nodes; and the scheduler sends the session request to the target computing node.

[0110] Optionally, after the computing node processes the session request, the computing node returns the processing result to the scheduler; and the scheduler is further configured to output the processing result to display the processing result on a user node sending the session request.

[0111] Optionally, after the computing node processes the session request, the computing node writes the obtained new KVCache corresponding to the session request back to the first memory of the storage node.

[0112] Optionally, the computing node returns the processing result to the dispatcher after processing the session request; the dispatcher sends a synchronization data request to the storage node, the synchronization data request indicating storing the new KV Cache to a second memory of the storage node, the access speed of the second memory being less than that of the first memory; and the storage node stores the new KV Cache to the second memory of the storage node.

[0113] The processing method, through the storage node interacting with the hardware accelerator memory of the computing node, before the computing node processes the session request, if the session request is a multi-round session request, acquires the KV Cache of the session request, and loads the KV Cache from the second memory to the first memory of the storage node. The first memory of the storage node directly interacts with the hardware accelerator memory of the computing node for data. For example, the first memory of the storage node is a host memory. That is, the first memory of the storage node can write data to the hardware accelerator memory of the computing node, and can also read data from the hardware accelerator memory of the computing node. Then, when the computing node processes the session request, the KV Cache can be directly read from the first memory of the storage node. This way eliminates the transmission delay caused by the computing node acquiring the KV Cache from other low-speed storage devices, such as a solid state disk or a shared network storage device, and also eliminates the transmission delay caused by acquiring the KV Cache across computing nodes, reduces the waiting time of the computing node, and thereby improves the utilization rate of the hardware accelerator.

[0114] Further, the embodiment of the present application also introduces another session processing method, which is applied to Figure 3 The large language model processing system as shown in the figure. In the embodiment of the present application, the large language model processing system can concurrently process multiple session requests. For ease of illustration, three session requests are taken as an example. In order to better illustrate the application, the storage node is taken as an example to illustrate the gd2fs cluster. In order to make the skilled in the art more convenient to understand, the gd2fs cluster is taken as an example to illustrate that it includes three gd2fs nodes, namely gd2fs node 1, gd2fs node 2 and gd2fs node 3.

[0115] The storage node Figure 6 provides an application schematic diagram of a large language model processing system. As Figure 6 shown, the session request received by the dispatcher of the large language model processing system is multiple, and the multiple session requests specifically include a first session request (referred to as SessionX) and a second session request (referred to as SessionY).

[0116] It should be noted that multiple session requests can reach the scheduler at the same time, or reach the scheduler within a preset period of time, and the scheduler concurrently executes the multiple session requests.

[0117] In the embodiment of the present application, the scheduler reads the database to determine whether the multiple session requests are first-round conversation requests. If not, that is, they are multi-round conversation requests, the storage path of the KV Cache is obtained. For example, as shown in Figure 6 , by reading the database, it is determined that SessionX and SessionY are multi-round conversation requests, and SessionZ is a first-round conversation request.

[0118] Further, in the embodiment of the present application, the database also stores the storage path of the KV Cache of the previous round of conversation requests in the multi-round conversation requests. The scheduler can obtain the KV Cache of the previous round of conversation requests from the corresponding storage device based on the storage path, and use it in the process of processing the conversation request.

[0119] For example, as shown in Figure 6 , the storage path of the previous round of KV Cache of sessionZ in the database, but including the previous round of KV Cache of sessionX (referred to as the first KV Cache) and the storage path 1, and the previous round of KV Cache of sessionY (referred to as the second KV Cache) and the storage path 2.

[0120] The scheduler sends a pre-reading request to the gd2fs cluster. The pre-reading request includes the storage path of the KV Cache of the previous round of conversation requests in the multi-round conversation requests. For example, the pre-reading request includes the storage path 1 and the storage path 2.

[0121] The gd2fs cluster matches the target gd2fs node from the multiple gd2fs nodes, and the target gd2fs node is used to cache the previous round of KV Cache of the multi-round conversation request to the host memory of the target gd2fs node.

[0122] For example, if the storage device corresponding to the storage path 1 is SSDX, and the storage device corresponding to the storage path 2 is SSDY, the target gd2fs node obtains the first KV Cache from the SSDX of the target gd2fs node itself based on the storage path 1, and caches it to the host memory of the target gd2fs node. Based on the storage path 2, the second KV Cache is obtained from the SSDY of the target gd2fs node itself, and is cached to the host memory of the target gd2fs node.

[0123] In the embodiment of the present application, the target gd2fs node can process the operation of reading the KV Cache in parallel and storing it to the host memory of the target gd2fs node. If the latest time of using the first KV Cache is the first time and the latest time of using the second KV Cache is the second time, the first time is later than the second time. That is, the first KV Cache is not used for a long time, and the second KV Cache data is used recently. The reading speed of the first KV Cache of the target gd2fs node is slower than the reading speed of the second KV Cache. In the embodiment of the present application, as long as the target gd2f node detects the operation of storing completion, the target gd2f node sends the cache completion instruction corresponding to the session request to the scheduler. Specifically, after the target gd2fs node reads the first KV Cache quickly, the target gd2fs node generates the first cache completion instruction and sends the first cache completion instruction to the scheduler. The scheduler performs subsequent operations on sessionX. After the target gd2fs node reads the second KV Cache slowly, the target gd2fs node generates the second cache completion instruction and sends the second cache completion instruction to the scheduler. The scheduler performs subsequent operations on sessionY.

[0124] In the embodiment of the present application, the scheduler is configured to send the session request to the message queue.

[0125] Specifically, if the scheduler determines that the session request is the first-round session request, the scheduler directly sends the first-round session request to the message queue. In this way, the process of pre-reading the cache is directly skipped, the scheduling process is simplified, and the work efficiency is improved.

[0126] If the scheduler determines that the session request is the multi-round session request, the scheduler sends the session request to the message queue after receiving the cache completion instruction. For example, the scheduler first sends sessionZ to the message queue. Then, after receiving the first cache completion instruction, the scheduler sends sessionX to the message queue, and after receiving the second cache completion instruction, the scheduler sends sessionY to the message queue. Thus, the order of the message queue is: sessionZ (no need to pre-read the KV Cache), sessionY (pre-read the KV Cache quickly), and sessionX (pre-read the KV Cache slowly).

[0127] In the embodiment of the present application, the session processing system includes n computing nodes. n1 computing nodes in the n computing nodes subscribe to the messages of the message queue. n1 is a positive integer, and n1 is less than or equal to n.

[0128] Each of the n1 computing nodes performs: determining whether to pull session requests from the message queue according to its own load state, if it is determined to pull session requests, determining the number of pulls, pulling session requests from the message queue, and the number of pulls is less than or equal to the determined number.

[0129] Exemplarily, serverX, serverY and serverZ all subscribe to the messages of the message queue. If the load of serverX is m1-2, the load of serverY is m2, and the load of serverZ is m3-1. Wherein, the maximum request processing capacity of serverX is m1, the maximum request processing capacity of serverY is m2, and the maximum request processing capacity of serverZ is m3. At this time, serverX and serverZ determine to actively pull session requests from the message queue, serverX actively pulls 2 session requests from the message queue, and serverZ actively pulls 1 session request from the message queue.

[0130] Further, if multiple computing nodes need to pull session requests from the message queue, the number of session requests pulled by each computing node can be determined by competing.

[0131] This way can ensure that as many session requests as possible are processed at the same time, thereby improving the processing speed. Users do not need to wait for a long time, that is, to reduce the waiting time of the session. Further, each computing node works in a high load state, so this way also helps to enhance the usage rate of the hardware acceleration device to run at a very high level.

[0132] After the computing node processes the session request, it writes the newly generated kv cache into the host memory of the target gd2fs through the write back mode. This way can take advantage of the characteristics of the storage server to quickly write KV Cache. In the next round of session request processing, the storage node does not need to obtain KV Cache across nodes, but only needs to obtain it from its own memory, so as to improve the acquisition speed.

[0133] The computing node returns the processing result to the scheduler. The scheduler shows the processing result to the user and initiates a synchronization data request to the gd2fs cluster. The synchronization data request is used to persist the newly written KV Cache to the SSD.

[0134] After the gd2fs writes, it returns the data synchronization result to the scheduler, and the scheduler updates the database.

[0135] For sessionX and sessionY, update the database, update to the latest kv cache file, and for sessionZ, insert data in the database, identify the corresponding kv cache file path and KV Cache file.

[0136] To sum up, the embodiments of the present application first receive the session request from the user terminal by the dispatcher, and the dispatcher determines whether to pre-read the KV Cache according to the request type (first round or multi-round session) and the state of the storage server. If it is a multi-round session, the dispatcher will send a retrieval instruction to the storage server, instructing it to pre-read the KV Cache from the second memory to the first memory. After the pre-reading is completed, the storage server will send a cache completion instruction to the dispatcher, at which time the dispatcher will send the session request to the computing server, and the computing server retrieves the KV Cache from the first memory of the storage server, processes the session request using these cache data. After processing is completed, the computing server writes the newly generated KV Cache back to the first memory of the storage server, and then the storage server persists it to the second memory. The whole process involves close cooperation between the dispatcher, the computing server and the storage server, through efficient data pre-reading and writing mechanism, as well as message queue and load balancing strategy, realizing the fast processing of session and effective utilization of resources in large language model scenario, greatly improving the response speed of the system and user experience, while also reducing the idle time of hardware acceleration devices and improving the resource utilization.

[0137] According to the method provided by the embodiments of the present application, the present application further provides a chip system, which includes one or more processors for calling and running instructions stored in the memory from the memory, so that the method of the above embodiments of the present application is executed. The chip system can be composed of a chip, or can include a chip and other discrete devices.

[0138] Among them, the chip system can include input circuit or interface for sending information or data, and output circuit or interface for receiving information or data.

[0139] According to the method provided by the embodiments of the present application, the present application further provides a computer program product, which includes computer program code, when the computer program code runs on the computer, so that the computer executes each step or process executed by the network device and terminal device in any of the preceding method embodiments.

[0140] According to the method provided by the embodiments of the present application, the present application further provides a computer readable storage medium, which stores program code, when the program code runs on the computer, so that the computer executes each step or process executed by the network device and terminal device in any of the preceding method embodiments.

[0141] The computer readable storage medium can be the volatile memory or the nonvolatile memory mentioned above, or can include both the volatile memory and the nonvolatile memory.

[0142] In the embodiments of the present application, each term and English abbreviation is an exemplary example given for convenience of description, and should not constitute any limitation on the present application. The present application does not exclude the possibility of defining other terms capable of achieving the same or similar functions in the existing or future protocols.

[0143] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated.

[0144] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented by other ways. For example, the device embodiments described above are merely illustrative, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.

Claims

1. A large language model processing system, characterized in that, The system includes: a management node, a computing node, and a storage node. The management node deploys a scheduler, which is communicatively connected to the storage node. The first memory of the storage node is used to directly interact with the hardware accelerator memory of the computing node. The scheduler is configured to: upon receiving a session request, and the session request being a multi-round session request, send a retrieval instruction to the storage node, the retrieval instruction indicating the retrieval of the key-value cache (KV Cache) from the previous round of the session request; The storage node is used to: obtain the KV Cache of the previous round of the session request, cache the KV Cache in the first memory of the storage node, and send a cache completion instruction to the scheduler after caching is completed; The scheduler is also configured to: upon receiving the cache completion instruction, send the session request to the computing node; The computing node is used to: after receiving the session request, retrieve the KVCache from the first memory of the storage node, and use the KVCache to process the session request.

2. The system according to claim 1, characterized in that, The scheduler is also used for: If the session request is received, and the session request is the first-round session request, the session request is sent to the computing node.

3. The system according to claim 1, characterized in that, The scheduler is specifically used for: Send the session request to the message queue; The computing node is configured to: if the computing node has subscribed to the message queue, actively pull the session request from the message queue.

4. The system according to claim 3, characterized in that, If the number of computing nodes is n, and all n1 computing nodes have subscribed to the message queue, where n is a positive integer, n1 is a positive integer, and n1 is less than or equal to n; Each of the n1 computing nodes is specifically used for: Based on its own load status, it determines whether to pull session requests from the message queue; If a session request is determined, determine the number of session requests to be requested; The session requests are pulled from the message queue, and the number of session requests pulled is less than or equal to a predetermined number.

5. The system according to claim 1, characterized in that, If the number of computing nodes is n, where n is a positive integer, the scheduler is specifically used for: Upon receiving the cache completion instruction, a target computing node matching the session request is determined from among the n computing nodes; The session request is sent to the target computing node.

6. The system according to claim 1, characterized in that, The computing node is also used for: After processing the session request, the processing result is returned to the scheduler; The scheduler is also configured to: output the processing result to display the processing result on the user node that sent the session request.

7. The system according to claim 1, characterized in that, The computing node is also used for: After processing the session request, the new KV Cache of the obtained session request is written back to the first memory of the storage node.

8. The system according to claim 7, characterized in that, The computing node is also used for: After processing the session request, the processing result is returned to the scheduler; The scheduler is also configured to: send a synchronization data request to the storage node, the synchronization data request indicating that the new KV Cache be stored in the second memory of the storage node, the access speed of the second memory being less than the access speed of the first memory; The storage node is also used to: store the new KV Cache in the second memory of the storage node.

9. The system according to claim 1, characterized in that, The storage node is either a GD2FS cluster directly connected to the image processor or a GDKV database directly connected to the image processor. The GD2FS cluster supports GPU Direct RDMA technology, which is used to directly access the hardware accelerator memory of remote nodes.

10. A session processing method, characterized in that, A management node is applied to a language model processing system. The management node deploys a scheduler. The system further includes a computing node and a storage node. The scheduler is connected to the storage node. The first memory of the storage node is used for direct data interaction with the hardware accelerator memory of the computing node. The method includes: Upon receiving a session request, and the session request is a multi-round session request, a retrieval instruction is sent to the storage node, the retrieval instruction instructing the retrieval of the key-value cache (KV Cache) from the previous round of the session request; The storage node receives a cache completion instruction, which instructs the storage node to perform the following operations: acquire the KV Cache and cache the KV Cache in the first memory of the storage node; The session request is sent to the computing node, so that after the computing node receives the session request, it retrieves the KV Cache from the first memory of the storage node and uses the KV Cache to process the session request.