Key value cache multiplexing method and related device
By coordinating the storage of key-value caches between the graphics processing unit (GPU) and the central processing unit (CPU), the problems of slow inference speed and limited number of reusable key-value caches for large language models are solved, resulting in faster inference speed and wider application scenarios.
Patent Information
- Application Number
- CN202410903089.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-05
- Publication Date
- 2026-01-09
AI Technical Summary
Large language models require a lot of computing resources during inference, resulting in slow real-time inference speed. In addition, the limited video memory of graphics processors leads to a small number of reusable key-value caches, which limits the application scenarios.
By coordinating the storage of key-value caches between the graphics processor and the central processing unit, the extra storage space of the central processing unit is utilized to increase the number of reusable key-value caches, and the inference speed is guaranteed by the interaction speed between the central processing unit and the graphics processor.
While maintaining inference speed, the number of reusable key-value caches has been increased, expanding the scope of application scenarios.
Smart Images

Figure CN121301237A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method and apparatus for reusing key-value caches. Background Technology
[0002] With the development of artificial intelligence technology, large language models (LLMs) have continued to receive high attention due to their powerful generation, understanding, and reasoning capabilities. However, large language models require a large amount of computing resources when providing reasoning services, resulting in slow real-time inference speeds.
[0003] In related technologies, not only is the model inference computation accelerated by using a graphics processing unit (GPU), but data that needs to be accessed multiple times during the model inference process is also stored in the GPU in the form of a key-value cache. This allows the model to directly access the key-value cache in subsequent inference processes without recalculation, thus improving inference speed by sacrificing storage space.
[0004] However, graphics processing units (GPUs) have limited video memory, requiring a significant amount to support the computations needed during model inference, resulting in limited video memory for storing key-value caches. In related technologies, the key-value cache generated during each inference iteration is typically released from the GPU after the inference process ends, leading to a limited number of reusable key-value caches and consequently restricting application scenarios. Summary of the Invention
[0005] To address the aforementioned technical issues, this application provides a method and related apparatus for reusing key-value caches, which increases the number of reusable key-value caches and expands the scope of application scenarios while ensuring inference speed.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] On one hand, embodiments of this application provide a method for reusing key-value caches, the method comprising:
[0008] Get the prompt word for the i-th round, where i is an integer greater than 1;
[0009] Based on the multiple word segments and historical word segments included in the i-th round of prompts, word segments with the same word segmentation order are matched according to the word segmentation order to obtain prefix word segments and remaining word segments. The historical word segments are word segments with word segmentation order obtained based on the prompts of the previous i-1 rounds. The prompts of the previous i-1 rounds are related to the prompts of the i-th round. The prefix word segments include the word segments that were successfully matched in the prompts of the i-th round when the first match failed. The remaining word segments include word segments other than those included in the prefix word segments from the multiple word segments included in the prompts of the i-th round.
[0010] If it fails to retrieve the key-value cache corresponding to the prefix segmentation group from the graphics processor, then retrieve the key-value cache corresponding to the prefix segmentation group from the central processing unit.
[0011] The key-value cache of the prefix segmentation group and the remaining segmentation group are sent to the inference engine so that the inference engine can obtain the response content for the i-th round prompt word based on the key-value cache of the prefix segmentation group and the remaining segmentation group.
[0012] On the other hand, embodiments of this application provide a key-value cache reuse device, the device comprising: an acquisition unit, a matching unit, a reuse unit, and a sending unit;
[0013] The acquisition unit is used to acquire the i-th round of prompt words, where i is an integer greater than 1;
[0014] The matching unit is configured to match words with the same word order according to the word order of the multiple word segments included in the i-th round prompt word and the historical word segment group, to obtain a prefix word segment group and a remaining word segment group. The historical word segment group is a word segment with a word order obtained based on the previous i-1 round prompt words. The previous i-1 round prompt words are related to the i-th round prompt word. The prefix word segment group includes the word segments that were successfully matched in the i-th round prompt word when the first match failed. The remaining word segment group includes the word segments other than the word segments included in the prefix word group from the multiple word segments included in the i-th round prompt word.
[0015] The multiplexing unit is configured to retrieve the key-value cache corresponding to the prefix segmentation group from the central processing unit if it fails to retrieve the key-value cache corresponding to the prefix segmentation group from the graphics processor.
[0016] The sending unit is used to send the key-value cache of the prefix segmentation group and the remaining segmentation group to the inference engine, so that the inference engine can obtain the response content for the i-th round prompt word based on the key-value cache of the prefix segmentation group and the remaining segmentation group.
[0017] On the other hand, embodiments of this application provide a computer device, the computer device including a processor and a memory:
[0018] The memory is used to store computer programs and to transfer the computer programs to the processor;
[0019] The processor is configured to execute the methods described above according to instructions in the computer program.
[0020] On the other hand, embodiments of this application provide a computer-readable storage medium for storing a computer program for performing the methods described above.
[0021] On the other hand, embodiments of this application provide a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods described above.
[0022] As can be seen from the above technical solution, if the graphics processor's memory can no longer store the key-value cache, it will not directly release the key-value cache. Instead, a portion of the key-value cache will be stored in the central processing unit (CPU). This allows the graphics processor and CPU to work together to store the key-value cache, further increasing the number of reusable key-value caches. Moreover, compared to other storage spaces, the interaction speed between the CPU and the graphics processor is faster, ensuring inference speed. Thus, while maintaining inference speed, the number of reusable key-value caches is increased, expanding the scope of application scenarios. The process of reusing the key-value cache is explained below.
[0023] After obtaining the i-th round of prompts, the multiple word segments included in the i-th round of prompts are matched with the historical word segments generated from the previous i-1 rounds of prompts. Word segments with the same order are matched to obtain prefix word segments and remaining word segments. This matching determines the reusable and non-reusable word segments in the i-th round of prompts. Then, the key-value cache of the prefix word segments is retrieved from the graphics processing unit (GPU). If retrieval fails, the key-value cache of that prefix word segment may be moved to the central processing unit (CPU), in which case the corresponding key-value cache is retrieved from the CPU. Finally, the key-value cache of the prefix word segments and the remaining word segments are sent to the inference engine. The inference engine then uses these caches to infer the response content for the i-th round of prompts. Therefore, during the calculation process, the inference engine does not need to recalculate the key-value cache of the prefix word segments; instead, it directly reuses the prefix word segment key-value cache and only calculates the key-value cache of the remaining word segments, thus improving inference speed. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 A schematic diagram illustrating an application scenario of a key-value caching reuse method provided in this application embodiment;
[0026] Figure 2 A schematic diagram illustrating a key-value caching reuse method provided in this application embodiment in a multi-turn dialogue request scenario;
[0027] Figure 3 A schematic diagram illustrating a key-value caching reuse method provided in this application embodiment in a code completion scenario;
[0028] Figure 4 A schematic diagram illustrating a key-value caching reuse method provided in this application embodiment in a long text query scenario;
[0029] Figure 5 A flowchart illustrating the key-value caching reuse method provided in this application embodiment;
[0030] Figure 6 A schematic diagram illustrating a multi-turn dialogue request scenario provided in an embodiment of this application;
[0031] Figure 7 A schematic diagram illustrating an application scenario of a key-value caching reuse method provided in this application embodiment;
[0032] Figure 8 A schematic diagram illustrating an application scenario of a key-value caching reuse method provided in this application embodiment;
[0033] Figure 9 This is a schematic diagram illustrating the result of an end-to-end delay decomposition provided in an embodiment of this application;
[0034] Figure 10 An experimental result diagram of a multi-turn interaction scenario provided in an embodiment of this application;
[0035] Figure 11 A schematic diagram of a key-value cache reuse device provided in an embodiment of this application;
[0036] Figure 12 This application provides a schematic diagram of the structure of a server according to an embodiment of the present application.
[0037] Figure 13 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation
[0038] The embodiments of this application will now be described with reference to the accompanying drawings.
[0039] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “corresponding to,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0040] In related technologies, to increase the number of reusable key-value caches while maintaining inference speed, a paged attention mechanism is introduced. This mechanism manages the key-value cache in a paginated manner, allowing contiguous key-value pairs to be stored in non-contiguous video memory space. In other words, it maximizes the number of key-value caches within a limited storage space. This avoids fragmentation within or outside the key-value cache, thereby increasing the number of concurrent inference requests and ultimately improving the overall utilization of the graphics processor.
[0041] It's important to note that regardless of whether a key-value cache is stored or a paging attention mechanism is used, the stored key-value cache is generally limited to the period within a single request (i.e., one round of dialogue), and it is deleted after each inference iteration. Although the increased storage space introduced by the paging attention mechanism may eliminate the need to immediately release the key-value cache generated during each inference iteration from the graphics processor, the previous key-value cache will be released immediately when new space needs to be allocated on video memory.
[0042] In other words, even if the paging attention mechanism stores as many key-value caches as possible in a limited storage space, the previous key-value caches will still be released immediately when new space needs to be allocated on video memory. This results in a relatively small number of reusable key-value caches, thus limiting application scenarios.
[0043] For example, if a paginated attention mechanism is used, it is not necessary to immediately delete the key-value cache generated within a single dialogue round after each inference technique. Taking the example that a graphics processor can only store key-value caches generated from two dialogue rounds (i.e., only the key-value caches generated from the (i-1)th and (i-2)th dialogue rounds), when the key-value cache generated from the i-th dialogue round needs to be stored, the key-value cache generated from the (i-2)th dialogue round must be released. Therefore, even in multi-turn dialogue application scenarios where the query statements are highly correlated (e.g., the query statements from the i-th and (i-2)th dialogue rounds have many overlapping word segments), during the inference process for the i-th dialogue round, since the key-value cache generated from the (i-2)th dialogue round has already been released, it cannot be reused and must be recalculated, thus limiting the number of reusable key-value caches.
[0044] Based on this, this application provides a method for reusing key-value caches. For application scenarios requiring the reuse of a large number of key-value caches, if the graphics processor's memory can no longer store the key-value cache, instead of directly releasing it, a portion of the key-value cache is stored in the central processing unit (CPU). This allows the graphics processor and CPU to work together to store the key-value cache, further increasing the number of reusable key-value caches. Moreover, compared to other storage spaces, the interaction speed between the CPU and the graphics processor is faster, ensuring inference speed. Thus, while maintaining inference speed, the number of reusable key-value caches is increased, expanding the scope of application scenarios.
[0045] The key-value caching reuse method provided in this application can be applied to computer devices with key-value caching reuse capabilities, such as terminal devices and servers.
[0046] Specifically, terminal devices can be desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Smart in-vehicle devices can be in-vehicle navigation terminals and in-vehicle computers, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc., but are not limited to these.
[0047] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminal devices and servers can be connected directly or indirectly via wired or wireless communication; this application does not impose any restrictions on this.
[0048] To facilitate understanding of the key-value cache reuse method provided in this application embodiment, the following example uses a server as the execution subject of the key-value cache reuse method to illustrate the application scenarios of the key-value cache reuse method.
[0049] See Figure 1 This figure is a schematic diagram illustrating an application scenario of a key-value caching reuse method provided in an embodiment of this application. For example... Figure 1 As shown, this application scenario includes a terminal device 110, a server 120, and an inference engine 130. The terminal device 110 and the server 120, as well as the server 120 and the inference engine 130, can communicate via a communication network. This communication network uses standard communication technologies and / or protocols, typically the Internet, but can also be any network, including but not limited to Bluetooth, a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), mobile, private networks, or any combination of virtual private networks. In some embodiments, customized or dedicated data communication technologies may be used to replace or supplement the aforementioned data communication technologies.
[0050] Terminal device 110 is equipped with clients that provide services such as session services and code completion services. Users can input corresponding content through the corresponding clients. Taking session services as an example, the client can be a client that provides session services, and users can input query statements through this client. Terminal device 110 sends the content entered by the user through the client for the i-th time to server 120.
[0051] Server 120 is equipped with services corresponding to the client in terminal device 110, such as session service and code completion service. Server 120 generates the i-th round prompt word based on the content sent by terminal device 110. Based on the multiple segments included in the i-th round prompt word, it matches the historical segmentation groups generated by the prompt words in the previous i-1 rounds according to the segmentation order. This results in the prefix segmentation group and the remaining segmentation group. In other words, the reusable and non-reusable segments in the i-th round prompt word are determined through matching.
[0052] In this embodiment, if the graphics processor's video memory can no longer store the key-value cache, the key-value cache is not directly released. Instead, a portion of the key-value cache is stored in the central processing unit (CPU). This allows the graphics processor and CPU to work together to store the key-value cache, further increasing the number of reusable key-value caches. Moreover, compared to other storage spaces, the interaction speed between the CPU and the graphics processor is faster, ensuring inference speed. Thus, while maintaining inference speed, the number of reusable key-value caches is increased, expanding the scope of application scenarios. The process of reusing the key-value cache is described below.
[0053] Once server 120 determines a reusable word segment, i.e. a prefix word segment group, it retrieves the key-value cache of the prefix word segment group from the graphics processor. If retrieval fails, it retrieves the key-value cache corresponding to the prefix word segment group from the central processing unit.
[0054] Finally, server 120 sends the key-value cache of the prefix word segmentation group and the remaining word segmentation group to inference engine 130, so that inference engine 130 can infer the response content for the i-th round of prompt words based on the key-value cache of the prefix word segmentation group and the remaining word segmentation group. Inference engine 130 sends the response content to terminal device 110 through server 120, so that the user can obtain the response content for the input of the i-th word. Thus, during the calculation process, the inference engine does not need to recalculate the key-value cache of the prefix word segmentation group, but directly reuses the key-value cache of the prefix word segmentation group, and only calculates the key-value cache of the remaining word segmentation group, which improves the inference speed.
[0055] The key-value caching reuse method provided in this application can be applied to various fields, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, Internet of Things, financial services, and medical services. Three scenarios are given below as examples.
[0056] Scenario 1: Multi-turn dialogue request scenario.
[0057] Multi-turn dialogue request scenarios refer to situations where users continuously write query statements, and a corresponding response is generated based on each round of query statements. One query statement and its corresponding response content constitute one round of dialogue. During multi-turn dialogues, the key-value cache generated by the query statements of previous rounds can be reused to avoid redundant calculations and improve inference speed.
[0058] See Figure 2 The figure is a schematic diagram of a key-value caching reuse method provided in this application embodiment in a multi-turn dialogue request scenario.
[0059] Taking the i-th round of dialogue as an example, the terminal device obtains the query statement of the i-th round of dialogue and sends it to the server. The server constructs the i-th round prompt word based on the query statement of the i-th round of dialogue. For example, it concatenates the query statement of the first round of dialogue, the response content of the first round of dialogue, and the query statement of the second round of dialogue to obtain the second round prompt word. The multiple word segments included in the i-th round prompt word are matched with the historical word segments generated based on the prompt words of the previous i-1 rounds to obtain the reusable word segments (prefix word segments) and non-reusable word segments (remaining word segments) in the i-th round prompt word. Then, the server retrieves the key-value cache of the prefix word segments from the graphics processor. If retrieval fails, it retrieves the key-value cache corresponding to the prefix word segments from the central processing unit. Finally, the key-value cache of the prefix word segments and the remaining word segments are sent to the inference engine so that the inference engine can perform inference based on the key-value cache of the prefix word segments and the remaining word segments to obtain the response content for the i-th round prompt word, that is, the reply statement based on the query statement of the i-th round of dialogue.
[0060] Scenario 2: Code completion scenario.
[0061] As the user inputs characters, the cursor remains positioned after the last character entered, meaning the cursor is constantly moving. During this movement, the inference engine predicts the code following the cursor based on the code preceding it, thus achieving code completion. During inference, the key-value cache generated from already predicted characters can be reused to avoid redundant calculations and improve inference speed.
[0062] See Figure 3 The figure is a schematic diagram of a key-value caching reuse method provided in this application embodiment in a code completion scenario.
[0063] Taking the i-th prediction as an example, the terminal device obtains the code before the i-th cursor position and sends it to the server. The server generates the i-th round prompt word based on the code, and matches the multiple segments included in the i-th round prompt word with the historical segmentation groups generated based on the previous i-1 round prompt words to obtain the reusable segments (prefix segmentation groups) and non-reusable segments (remaining segmentation groups) in the i-th round prompt word. Then, the server retrieves the key-value cache of the prefix segmentation group from the graphics processor. If retrieval fails, it retrieves the key-value cache corresponding to the prefix segmentation group from the central processing unit. Finally, the server sends the key-value cache of the prefix segmentation group and the remaining segmentation group to the inference engine so that the inference engine can perform inference based on the key-value cache of the prefix segmentation group and the remaining segmentation group to obtain the response content for the i-th round prompt word, that is, the subsequent code predicted from the code input before the cursor.
[0064] Scenario 3: Long text query scenario.
[0065] Users can input relatively long text as query segments, and the inference engine generates corresponding response content based on these segments. However, although the query segment includes contextual semantics, a single inference process does not need to learn the entire query segment; learning only a sufficient amount of semantics is enough. In other words, the query segment can be divided into multiple segments using a sliding window, thereby reducing the length of the query segment while improving inference speed. Since there are repeated word segments between sliding windows in adjacent inference processes, the key-value cache that has already been calculated can be reused to avoid redundant calculations and further improve inference speed.
[0066] See Figure 4 The figure is a schematic diagram of a key-value caching reuse method provided in this application embodiment in a long text query scenario.
[0067] Taking the i-th inference as an example, the terminal device obtains the long text content input by the user and sends it to the server. The server constructs the text included in the sliding window corresponding to the i-th inference, i.e., the i-th round prompt word, based on this text. It then matches the multiple segments included in the i-th round prompt word with the historical segmentation groups generated based on the previous (i-1) round prompt words to obtain the reusable segments (prefix segmentation groups) and non-reusable segments (remaining segmentation groups) in the i-th round prompt word. Next, the server retrieves the key-value cache of the prefix segmentation group from the graphics processor. If retrieval fails, it retrieves the key-value cache corresponding to the prefix segmentation group from the central processing unit. Finally, the server sends the key-value cache of the prefix segmentation group and the remaining segmentation group to the inference engine, so that the inference engine can perform inference based on these data to obtain the response content for the i-th round prompt word, i.e., the reply content for the i-th inference. After obtaining all the reply content, the server sends the reply content for the long text to the terminal device.
[0068] It should be noted that the above application scenarios are merely examples. The key-value caching reuse method provided in this embodiment can also be applied to other scenarios, and is not limited here.
[0069] The key-value cache reuse method provided in this application embodiment can be executed by a server. However, in other embodiments of this application, the terminal device may also have similar functions to the server to execute the key-value cache reuse method provided in this application embodiment, or the terminal device and the server may jointly execute the key-value cache reuse method provided in this application embodiment. This embodiment does not limit this.
[0070] The following describes in detail a key-value cache reuse method provided in this application through method embodiments.
[0071] See Figure 5, This figure is a schematic flowchart of the method for reusing key-value caches provided by an embodiment of this application. For ease of description, in the following embodiments, the execution subject of the method for reusing key-value caches is still taken as an example of a server for introduction. As Figure 5 shown, the method for reusing key-value caches includes S501 - S504.
[0072] S501: Obtain the prompt words for the i-th round.
[0073] i is an integer greater than 1. In scenarios such as multi-round dialogue request scenarios, code completion scenarios, and long text inquiry scenarios, it is necessary to obtain multi-round prompt words in order to generate corresponding response content based on each round of prompt words, so as to achieve multiple interactions with the user.
[0074] Generally, there is an association relationship between multi-round prompt words. Taking the multi-round dialogue request scenario as an example, the query statement for the first round of dialogue is "What's the weather like today", and the reply statement for the first round of dialogue is "It's 31°C today, cloudy turning to sunny". The query statement for the second round of dialogue is "What should I wear appropriately", actually in the second round of dialogue, what the user really wants to ask is "It's 31°C today, cloudy turning to sunny, what should I wear appropriately". Based on this, in order to better understand the user's semantics, the prompt words for the second round will be jointly generated based on the historical dialogue and the query statement for the second round, such as "What's the weather like today; It's 31°C today, cloudy turning to sunny; What should I wear appropriately". And the prompt words for the first round have no historical dialogue and are generally themselves, such as "What's the weather like today". From this, it can be seen that there are generally association relationships such as semantic association and word repetition between multi-round prompt words.
[0075] In the related art, the key (K) and value (V) caches of the word segments involved in the same round of prompt words can be reused. Taking the large language model (LLM) as an example. For a round of prompt words, a response content will be output. During the process of generating the response content, multiple inferences will be carried out. Each inference only generates one word segment (token). The output word segments will be concatenated with multiple input word segments as the input for the next inference. Each word segment depends on the previously generated word segments until the condition is met and the inference ends to obtain the response content for this round of prompt words. Taking the prompt words "What's the weather like today" as an example, in the first round of inference, the word segment "今" will be generated for "What's the weather like today". In the second round of inference, the word segment "天" will be generated for "What's the weather like today; 今", resulting in "今天". And so on, until "It's 31°C today, cloudy turning to sunny" is obtained.
[0076] However, as inference progresses, the number of input word segments increases, leading to a significant increase in computational load and severely impacting inference speed. Simultaneously, during inference, the KV values of the first i-1 words are repeatedly calculated for each attention step, storing these KV values in a KV cache within the graphics processing unit (GPU). The next round of inference directly reads from this cache, thus improving inference speed.
[0077] As mentioned above, there are relationships between multiple rounds of prompts. In order to further improve the reasoning speed, the key-value cache that is only reused within a single round of prompts can be changed to reuse the key-value cache between multiple rounds of prompts. That is, instead of releasing the key-value cache generated by the current reasoning from the graphics processor after each reasoning, the graphics processor stores as much key-value cache generated by multiple rounds of prompts as possible within its video memory for direct use later.
[0078] Furthermore, in this embodiment, if the graphics processor's video memory can no longer store the key-value cache, the key-value cache is not directly released. Instead, a portion of the key-value cache is stored in the central processing unit (CPU). This allows the graphics processor and CPU to work together to store the key-value cache, thereby further increasing the number of reusable key-value caches. The storage process will be described later and will not be repeated here.
[0079] S502: Based on the multiple word segments and historical word segments included in the i-th round prompt, match the word segments with the same word segmentation order according to the word segmentation order to obtain the prefix word segmentation group and the remaining word segmentation group.
[0080] The i-th round of prompts includes multiple word segments, which are arranged in the order of word segmentation. Taking the second round of prompts as "How is the weather today; Today is 31℃, cloudy turning sunny, what should I wear?" as an example, the multiple word segments included in it can be arranged in the order of word segmentation as "today", "weather", "how is it", "today", "31℃", "cloudy turning sunny", "I", "wear what", "relatively", "suitable".
[0081] The historical word segmentation group also includes multiple historical word segments. These historical word segments are derived from the prompts of the previous i-1 rounds, and they are arranged in a specific order. For example, if the prompt for the first round is "How's the weather today?", its multiple historical word segments, arranged in word order, could be "Today", "Weather", and "How's it?". Furthermore, the prompts from the previous i-1 rounds are related to the prompt for the i-th round, such as queries from the same user in the i-th round.
[0082] It should be noted that the word segmentation order of multiple words in the historical word segmentation group is consistent with the word segmentation order of multiple words included in the i-th round prompt word. For example, they are all determined based on time order or semantic order. The word segmentation order determined by this method is generally consistent with the order in which the subsequent inference engine generates words during the inference process, thereby ensuring the feasibility of key-value cache reuse of subsequent words.
[0083] Based on the word segmentation order, corresponding matching between word segments is achieved. Specifically, word segments with the same order are matched to obtain a prefix word group and a remaining word group. The prefix word group includes the word segment that successfully matched in the i-th round of prompts when the first match failed; these are reusable word segments in the i-th round of prompts, allowing for direct use of their corresponding key-value caches in subsequent matches. The remaining word group includes the word segments from the multiple word groups included in the i-th round of prompts, excluding those from the prefix word group; these are the unreusable word segments.
[0084] For example, if the i-th round of prompts includes the segmented words ABCDF, and the historical segmented words include ABEDG, then the prefix segmented word group is AB, not AB and D. This is because during the matching process according to the segmented word order, C and E will fail to match, and this is the first failure; F and G will fail on the second attempt. Therefore, the remaining segmented word group is CDF. It should be noted that even if D matches successfully, in subsequent inference processes, due to the differences between C and E, even if D is the same, the two inference calculations will be different. Therefore, even if D is the same, the key-value cache for D cannot be reused. In other words, in subsequent inference processes, only the segmented words that were successfully matched during the first failure—that is, the prefix segmented word group—can be reused.
[0085] As one possible implementation, the k-th segment among the multiple segments included in the i-th round prompt word is matched with the k-th segment in the historical segment group, where k is an integer starting from 1; if the match is successful, k+1 is updated to k, and the matching continues; if the match fails, the successfully matched segment in the i-th round prompt word is taken as the prefix segment group, and the segment in the multiple segments included in the i-th round prompt word other than the segment included in the prefix segment group is taken as the remaining segment group.
[0086] Continuing with the example above, when k=1, the first word segment in the i-th round of prompts is matched with the first historical word segment in the historical word segment group. If the match is successful, k=2, the second word segment in the i-th round of prompts is matched with the second historical word segment in the historical word segment group. If the match is successful, k=3, the third word segment in the i-th round of prompts is matched with the third historical word segment in the historical word segment group. If the match fails, the prefix word segment groups "today", "weather", and "how is it" are obtained, and the remaining word segment groups are "today", "31℃", "cloudy to sunny", "me", "wear what", "relatively", and "suitable".
[0087] Therefore, by matching the i-th round prompt word with the words in the historical word segmentation group one by one according to the word segmentation order, the prefix word segmentation group and the remaining word segmentation group are obtained. Moreover, the matching ends after the first matching fails, avoiding the consumption of computing resources by subsequent useless matching, increasing the video memory space in the graphics processor, further increasing the number of key-value caches, and improving the inference speed.
[0088] S503: If it fails to retrieve the key-value cache corresponding to the prefix segmentation group from the graphics processor, then retrieve the key-value cache corresponding to the prefix segmentation group from the central processing unit.
[0089] As described above, in this embodiment, the graphics processing unit (GPU) and the central processing unit (CPU) work together to store the key-value cache, thereby increasing the number of reusable key-value caches. Furthermore, compared to other storage spaces, the interaction speed between the CPU and GPU is faster, ensuring inference speed. This achieves both high inference speed and an increased number of reusable key-value caches, expanding the scope of application scenarios.
[0090] Therefore, after determining the prefix segmentation group, the key-value cache corresponding to the prefix segmentation group is first retrieved from the graphics processor. If the retrieval is successful, step S504 is executed. If the retrieval fails, since a portion of the key-value cache is stored in the central processing unit (CPU), the key-value cache corresponding to the prefix segmentation group can be retrieved from the CPU.
[0091] Furthermore, if retrieving the key-value cache corresponding to the prefix word segmentation group fails not only from the graphics processor (GPU) but also from the central processing unit (CPU), it indicates that neither the GPU nor the CPU stores the key-value cache corresponding to the prefix word segmentation group. In this case, the prefix word segmentation group and the remaining word segmentation group can be sent to the inference engine, i.e., the i-th round prompt word can be sent to the inference engine so that the inference engine can re-perform the inference calculation. That is, the inference engine obtains the response content for the i-th round prompt word based on the prefix word segmentation group and the remaining word segmentation group. Thus, even if the CPU and GPU do not store the key-value cache corresponding to the prefix word segmentation group, the inference calculation for the i-th round prompt word can still be guaranteed to proceed normally, improving the user experience.
[0092] S504: Send the key-value cache of the prefix segmentation group and the remaining segmentation group to the inference engine so that the inference engine can obtain the response content for the i-th round prompt word based on the key-value cache of the prefix segmentation group and the remaining segmentation group.
[0093] Inference engines typically have models installed that provide corresponding services, such as large language models like Generative Pre-trained Transformer (GPT) models, which can provide services such as conversation services and code completion services.
[0094] In related technologies, since the key-value cache generated during each inference process is released immediately after completion, the i-th round prompt word is typically sent directly to the inference engine. This allows the inference engine to recalculate the inference based on the multiple word segments included in the i-th round prompt word, obtaining the response content for that prompt word. It's understood that the inference engine collaborates with the graphics processing unit (GPU) to improve inference speed. Specifically, the inference engine distributes multiple inferences for the same round prompt word to the GPU, allowing the GPU to perform the inference calculations, obtain the word segments from that round of inference, and store the final result of the inference calculation process in the GPU's memory. The GPU then sends the calculated word segments back to the inference engine.
[0095] In this embodiment, the key-value cache generated during the inference process is not released after each inference iteration. Instead, the key-value cache generated by the prompts in the previous i-1 rounds, i.e., the key-value cache of the prefix segmentation group, is directly reused. The key-value cache of the prefix segmentation group and the remaining segmentation groups are sent to the inference engine. Thus, the inference engine does not need to perform inference calculations again, but directly reuses the key-value cache of the prefix segmentation group, and only performs inference calculations on the remaining segmentation groups that cannot be reused, thereby improving the inference speed.
[0096] As can be seen from the above technical solution, if the graphics processor's memory can no longer store the key-value cache, it will not directly release the key-value cache. Instead, a portion of the key-value cache will be stored in the central processing unit (CPU). This allows the graphics processor and CPU to work together to store the key-value cache, further increasing the number of reusable key-value caches. Moreover, compared to other storage spaces, the interaction speed between the CPU and the graphics processor is faster, ensuring inference speed. Thus, while maintaining inference speed, the number of reusable key-value caches is increased, expanding the scope of application scenarios. The process of reusing the key-value cache is explained below.
[0097] After obtaining the i-th round of prompts, the multiple word segments included in the i-th round of prompts are matched with the historical word segments generated from the previous i-1 rounds of prompts. Word segments with the same order are matched to obtain prefix word segments and remaining word segments. This matching determines the reusable and non-reusable word segments in the i-th round of prompts. Then, the key-value cache of the prefix word segments is retrieved from the graphics processing unit (GPU). If retrieval fails, the key-value cache of that prefix word segment may be moved to the central processing unit (CPU), in which case the corresponding key-value cache is retrieved from the CPU. Finally, the key-value cache of the prefix word segments and the remaining word segments are sent to the inference engine. The inference engine then uses these caches to infer the response content for the i-th round of prompts. Therefore, during the calculation process, the inference engine does not need to recalculate the key-value cache of the prefix word segments; instead, it directly reuses the prefix word segment key-value cache and only calculates the key-value cache of the remaining word segments, thus improving inference speed.
[0098] This application does not specifically limit the method of using a graphics processor and a central processing unit to store key-value caches. Two methods are described below as examples.
[0099] Method 1: Implement a dynamic uninstallation strategy in advance.
[0100] A1: Retrieve the cached key-value pairs generated by the prompt words in the (i-1)th round.
[0101] The key-value cache generated by the prompt word in round i-1 refers to the key-value cache generated during the reasoning and calculation process based on the multiple word segments included in the prompt word in round i-1.
[0102] A2: Cache and store the key-value generated by the prompt word in the (i-1)th round to the graphics processor.
[0103] Since sufficient storage space is reserved in the graphics processor after storing the key-value cache generated by the i-2th round of prompts, the key-value cache generated by the i-1th round of prompts can be directly stored in the graphics processor.
[0104] A3: Get the remaining video memory space of the graphics processor.
[0105] The following example illustrates how sufficient storage space is reserved in the graphics processor after storing the key-value cache generated by the (i-1)th round of prompts.
[0106] A portion of the graphics processor's storage space is typically used for inference calculations, while the remaining storage space, i.e., the remaining video memory, is used to store key-value caches. Therefore, it is necessary to obtain the remaining storage space of the graphics processor after storing the key-value cache generated by the (i-1)th round of prompts, i.e., to obtain the remaining video memory space of the graphics processor.
[0107] This application does not specifically limit the method for determining the remaining video memory space. For example, the remaining video memory space of the graphics processor can be determined based on the video memory space occupied by the key-value cache already stored in the graphics processor, the video memory space required by the graphics processor to perform inference calculations, and the total video memory space of the graphics processor. Specifically, the video memory space occupied by the stored key-value cache can be determined based on the model deployed by the inference engine, which determines the hidden dimension of the attention layer of the model, the number of attention layers of the decoder, and the number of attention heads of the attention layer, thereby determining the video memory space required for one word segmentation. Then, the video memory space occupied by the stored key-value cache can be determined based on the number of stored words and the video memory space required for one word segmentation.
[0108] A4: If the remaining video memory space of the graphics processor is less than the preset storage space, then determine the first target key value cache from the key value cache stored in the graphics processor, store the first target key value cache in the central processing unit, and delete the first target key value cache from the graphics processor.
[0109] If the remaining video memory space of the graphics processor is less than the preset storage space, it means that the remaining video memory space of the graphics processor is insufficient to directly store the key value cache generated by the i-th round of prompt words. Therefore, it is necessary to release a part of the key value cache from the graphics processor, namely the first target key value cache, store the first target key value cache in the central processing unit, and delete the first target key value cache from the graphics processor. The first target key value cache and the remaining video memory space are greater than or equal to the preset storage space.
[0110] The embodiments of this application do not limit the size of the preset storage space. Those skilled in the art can set it according to actual needs to reserve a cache that can directly store the key value generated by the i-th round of prompts, and ensure that the graphics processor has reorganized video memory to support inference calculations.
[0111] This application does not specifically limit the method for determining the first target key-value cache. For example, the key-value cache with the longest storage time can be determined from the multiple key-value caches included in the graphics processor and used as the first target key-value cache. Alternatively, key-value caches with a storage time exceeding a preset time threshold can be determined from the multiple key-value caches included in the graphics processor. These key-value caches are less likely to be reused, so they can be used as the first target key-value cache and moved to the central processing unit for storage.
[0112] A5: If the remaining video memory space of the graphics processor is greater than or equal to the predicted memory space, there is no need to release the key-value cache from the graphics processor.
[0113] If the remaining video memory space of the graphics processor is greater than or equal to the predicted storage space, it means that the remaining video memory space of the graphics processor can support the direct storage of the next round of prompt words, that is, the key value cache generated by the i-th round of prompt words, so that it is not necessary to release a part of the key value cache from the graphics processor.
[0114] Therefore, after storing the key-value cache in the graphics processor (GPU), the remaining video memory (VRAM) of the GPU is checked. If the remaining VRAM is greater than or equal to the preset storage space, meaning the GPU can still support storing the key-value cache corresponding to the next round of prompts, no processing is required. If the remaining VRAM is less than the predicted storage space, meaning the GPU cannot support storing the key-value cache corresponding to the next round of prompts, then a portion of the key-value cache in the GPU needs to be moved to the central processing unit (CPU) for storage, in order to reserve sufficient VRAM in the GPU for storing the key-value cache corresponding to the next round of prompts. This achieves the function of dynamically unloading the key-value cache based on the remaining VRAM of the GPU, and the remaining VRAM of the GPU only needs to be checked after storing the key-value cache generated by one round of prompts, eliminating the need for real-time monitoring of the remaining VRAM of the GPU, thus reducing the number of checks. Furthermore, by checking the remaining VRAM of the GPU and then determining the key-value cache to be moved to the CPU, unnecessary key-value cache swapping between the CPU and GPU can be avoided.
[0115] Method 2: Execute dynamic uninstallation strategy in real time.
[0116] B1: Obtain the remaining video memory space of the graphics processor and the cache of key values to be stored generated by the i-th round of prompts.
[0117] The key-value cache to be stored is determined based on the key-value cache of a word segment and the number of words to be stored in the remaining word segment groups. That is, the key-value cache of a word segment and the number of words to be stored in the remaining word segment groups are multiplied to obtain the key-value cache to be stored.
[0118] It should be noted that the number of words to be stored in the remaining word segmentation group can be the total number of words included in the remaining word segmentation group, that is, storing the key-value cache to be stored generated by the i-th round prompt word at one time. The number of words to be stored in the remaining word segmentation group can also be the number of corresponding words stored at one time in the remaining word segmentation group, that is, storing the key-value cache to be stored generated by the i-th round prompt word in multiple times. This application does not impose specific limitations on this, and those skilled in the art can set it according to actual needs.
[0119] B2: If the remaining video memory space is less than the key value cache to be stored, then determine the second target key value cache from the key value cache stored in the graphics processor, store the second target key value cache in the central processing unit, delete the second target key value cache from the graphics processor, and store the key value cache to be stored in the graphics processor.
[0120] If the remaining video memory space is less than the key value cache to be stored, it means that the current remaining storage space of the graphics processor is insufficient to support the key value cache to be stored generated by the i-th round of prompts. At this time, it is necessary to move a part of the key value cache in the graphics processor to the central processing unit for storage, so as to expand the storage space of the graphics processor and reuse this part of the key value cache in the future.
[0121] This portion of the key-value cache is the second target key-value cache. The second target key-value cache is stored in the central processing unit (CPU) and then deleted from the graphics processing unit (GPU). At this point, the sum of the second target key-value cache released from the GPU and the remaining video memory space of the GPU is greater than or equal to the key-value cache to be stored, thus the key-value cache to be stored is stored in the GPU.
[0122] This application does not specifically limit the method for determining the second target key-value cache. For example, the key-value cache with the longest storage time can be determined from the multiple key-value caches included in the graphics processor and used as the second target key-value cache. Alternatively, key-value caches with a storage time exceeding a preset time threshold can be determined from the multiple key-value caches included in the graphics processor. These key-value caches are less likely to be reused, so they can be used as the second target key-value cache and moved to the central processing unit for storage.
[0123] B3: If the remaining video memory space is greater than or equal to the key-value cache to be stored, then the key-value cache to be stored will be stored in the graphics processor.
[0124] If the remaining video memory space is greater than or equal to the key value cache to be stored, it means that the current storage space of the graphics processor is sufficient to support the storage of the key value cache generated by the i-th round of prompts. Therefore, there is no need to move part of the key value cache to the central processing unit for storage. Instead, the key value cache to be stored is directly stored in the graphics processor.
[0125] Therefore, each time the graphics processor (GPU) is about to store the key-value cache to be stored, the remaining video memory space of the GPU is checked. If the remaining video memory space is greater than or equal to the key-value cache to be stored, meaning the GPU can still support storing the key-value cache, then the key-value cache to be stored is directly stored in the GPU. If the remaining video memory space is less than the key-value cache to be stored, meaning the GPU cannot support storing the key-value cache, then a portion of the key-value cache in the GPU needs to be moved to the central processing unit (CPU) for storage, so that the key-value cache to be stored can be stored in the GPU. This achieves the function of dynamically unloading the key-value cache based on the remaining video memory space of the GPU, and by checking the remaining video memory space each time storage occurs, the storage space of the GPU is fully utilized to further increase the storage capacity of the word segmentation key-value cache. In addition, by checking the remaining video memory space of the GPU and then determining the key-value cache to be moved to the CPU, unnecessary key-value cache swapping between the CPU and the GPU can be avoided.
[0126] As one possible implementation, when the graphics processor (GPU) runs out of remaining video memory, a portion of the key-value cache can be released from the GPU to the central processing unit (CPU). Over time, the CPU will also become unable to continue storing the key-value cache from the GPU.
[0127] Based on this, embodiments of this application provide a method for determining a third target key-value cache based on the storage time of the key-value cache, and releasing the third target key-value cache from the central processing unit. The details are described below.
[0128] C1: Get the remaining memory space of the central processing unit.
[0129] Remaining memory space refers to the remaining storage space of the central processing unit.
[0130] C2: If the remaining memory space is less than the second target key-value cache, then determine the third target key-value cache from the CPU, delete the third target key-value cache from the CPU, and store the second target key-value cache in the CPU.
[0131] If the storage time of the third target key-value cache exceeds a preset time threshold, meaning the storage time of the third target key-value cache is relatively long, its likelihood of being reused is low. Therefore, it can be released as a key-value cache to expand the storage space in the central processing unit. For example, the third target key-value cache can be the key-value cache with the longest storage time in the central processing unit. Alternatively, it can be any key-value cache whose storage time exceeds the preset time threshold; this application does not specifically limit this.
[0132] The embodiments of this application do not specifically limit the preset time threshold, but those skilled in the art can limit it according to actual needs.
[0133] It should be noted that the sum of the third target key-value cache and the remaining memory space is greater than or equal to the second target key-value cache, so that after the third target key-value cache is released from the CPU, the CPU's storage space can store the second target key-value cache.
[0134] C3: If the remaining memory space is greater than or equal to the second target key-value cache, then store the second target key-value cache in the central processing unit.
[0135] If the remaining memory space is greater than or equal to the second target key-value cache, it means that the remaining memory space of the CPU is sufficient to store the second target key-value cache. In this case, there is no need to release a part of the key-value cache from the CPU, and the second target key-value cache can be directly stored in the CPU, thereby avoiding unnecessary key-value cache interaction between the CPU and the graphics processor.
[0136] Therefore, after the graphics processor determines that the second target key-value cache is to be stored in the central processing unit, if the remaining memory space of the central processing unit is sufficient to store the second target key-value cache, then the second target key-value cache is stored directly. If the remaining memory of the central processing unit is insufficient to support the storage of the second target key-value cache, then a portion of the key-value cache that has been stored for a longer period of time is released from the central processing unit to ensure that the second target key-value cache is stored in the central processing unit. This ensures the normal function of the central processing unit while avoiding unnecessary key-value cache interaction between the central processing unit and the graphics processor.
[0137] In related technologies, multiple inference engines are used to perform corresponding services in order to improve the speed of inference computation. However, multi-turn interactions with related relationships, such as multi-turn interactions from the same user, have a lot of redundant prefix contexts. Therefore, key-value cache reuse can only be achieved when multi-turn interactions with related relationships are routed to the same inference engine.
[0138] Based on this, a session management mechanism can be used to route related multi-turn prompts to the same inference engine, thereby enabling the reuse of key-value caches. See D1-D5 for details.
[0139] D1: Get the initial prompt word, including the session identifier.
[0140] Session identifiers are used to identify associations. That is, session identifiers can be used to identify multi-turn prompts that are related, such as identifying multiple interactions from the same user based on session identifiers.
[0141] D2: If the initial prompt word is determined to have a related first i-1 round prompt word based on the session identifier, then obtain the i-th round prompt word based on the initial prompt word.
[0142] If other prompts related to the initial prompt can be identified based on the session identifier, it means the initial prompt is not the first prompt to appear, and it can be reused from other prompts that appeared previously. For example, if the other prompts are from the previous (i-1) rounds, the initial prompt can be directly used as the prompt for the i-th round. Similarly, in a multi-turn dialogue request scenario, the initial prompt can be the user-input query statement; therefore, the i-th round prompt can be obtained based on the i-th round query statement, the previous (i-1) rounds of query statements, and their corresponding response statements. Furthermore, in a code completion scenario, the initial prompt can be the code before the cursor; the i-th round prompt can be generated based on the code before the cursor, and so on.
[0143] D3: Execute the aforementioned S502-S503 to send the key-value cache of the prefix segmentation group and the remaining segmentation group to the target inference engine among multiple inference engines.
[0144] After obtaining the i-th round of prompts, execute the aforementioned S502-S503 to obtain the key-value cache of the prefix segmentation group and the remaining segmentation group. Send the key-value cache of the prefix segmentation group and the remaining segmentation group to the target inference engine among multiple inference engines.
[0145] In this context, multi-turn prompts with the same session identifier are executed through the same inference engine, so that multi-turn prompts with related relationships can be routed to the same inference engine to achieve key-value cache reuse.
[0146] The target inference engine is the inference engine among multiple inference engines used to process multi-round prompts that are related to the i-th round prompt. In other words, multi-round prompts that are related are distributed to the same inference engine based on the session identifier, so that the same inference engine can better reuse the key-value cache of the prefix segmentation group for multi-round prompts that are related.
[0147] D4: If it is determined from the session identifier that there is no related prompt word in the initial prompt word, then obtain the first round of prompt words based on the initial prompt word.
[0148] If the initial prompt word does not have any related prompt words (such as the prompt words of the previous i-1 rounds) as determined by the session identifier, it means that the initial prompt word is the prompt word obtained for the first time and cannot be reused from other prompt words. In this case, the prompt word of the first round is obtained based on the initial prompt word.
[0149] D5: For the first round of prompts, send the multiple word segments included in the first round of prompts to any one of the multiple inference engines.
[0150] Since there are no reusable prefix segmentation groups for the first round of prompts, the multiple segments included in the first round of prompts can be directly sent to any one of the multiple inference engines. As a result, subsequent prompts that are related to the first round of prompts will be sent to that inference engine for inference calculation based on the session identifier.
[0151] Therefore, by using session identifiers, multi-turn dialogues with related relationships can be sent to the same inference engine for inference calculation. For example, multi-turn dialogue requests from the same user can be routed to the same inference engine based on session identifiers, thus making it easier to reuse the key-value cache of prefix segmentation groups.
[0152] One possible implementation is to store the data by creating triples. A triple consists of a session identifier, the tokens generated during the session, and the key-value cache corresponding to each token. For example, for the first round of prompts, the metadata set (triple) for the first round of prompts is stored. This metadata set includes the session identifier carried by the initial prompt for obtaining the first round of prompts, the multiple tokens generated by the first round of prompts, and the key-value cache corresponding to each token.
[0153] The metadata is then updated based on the existing metadata to obtain the updated metadata group. Taking the i-th round prompt word as an example, the metadata group of the (i-1)-th round prompt word is updated to obtain the metadata group of the i-th round prompt word. The metadata group of the i-th round prompt word includes the session identifier, multiple word segments generated by the i-th round prompt word, and the key-value cache corresponding to each word segment. This is equivalent to keeping the session identifier unchanged, and adding a key-value cache based on the remaining word segments of the i-th round prompt word in the metadata group of the (i-1)-th round prompt word.
[0154] Furthermore, time can be added to the metadata group. For example, the metadata group for the first round of prompts includes the session identifier, the time when the prompt was obtained in this round, the multiple word segments generated from the first round of prompts, and the key-value cache corresponding to each word segment. Similarly, the metadata group for the i-th round of prompts includes the session identifier, the time when the prompt was obtained in this round, the multiple word segments generated from the i-th round of prompts, and the key-value cache corresponding to each word segment. Thus, based on the time when the prompts were obtained in each round, the storage time of the key-value cache is determined, so that the target key-value cache (such as the first target key-value cache, the second target key-value cache, etc.) can be determined subsequently based on the time stored in the metadata group. It is understandable that during the update of the metadata group, the time when the prompt was obtained in this round also needs to be updated. For example, the time stored in the metadata group for the i-th round of prompts is the time when the prompt was obtained in the i-th round, and the time stored in the metadata group for the (i-1)-th round of prompts is the time when the prompt was obtained in the (i-1)-th round.
[0155] Therefore, by constructing and updating metadata groups, prefix segmentation groups and their key-value caches can be quickly obtained based on the session identifier in the metadata group. The target key-value cache can even be determined based on the time of obtaining the prompt word in the current round in the metadata group, thereby releasing the target key-value cache that is less likely to be reused.
[0156] As one possible implementation, if the prompt word for the (i-1)th round is modified before obtaining the prompt word for the i-th round, such as deleting certain rounds of dialogue or deleting code content, resulting in the deletion of the prompt word for the (i-1)th round, the stored key-value cache can be modified based on the key-value cache corresponding to the word segments included in the modified prompt word for the (i-1)th round. This ensures that the stored key-value cache is consistent with the word segments included in the prompt word for the (i-1)th round, thereby reducing the storage space of the graphics processor or central processing unit.
[0157] For ease of explanation, the following example of a multi-turn dialogue request scenario will be used to illustrate the reuse of the key-value cache provided in this application embodiment. See E1-E5 for details.
[0158] See Figure 6 The figure is a schematic diagram of a multi-turn dialogue request scenario provided in an embodiment of this application.
[0159] exist Figure 6 In this process, after obtaining the first round of prompts, the multiple word segments included in the first round of prompts are sent to the inference engine. The inference engine performs inference calculations to obtain response content 1, which is the response content obtained for the first round of prompts. The first round of prompts and response content 1 constitute one round of dialogue. Similarly, response content 2 can be obtained based on the second round of prompts, resulting in the second round of dialogue. Response content i can be obtained based on the i-th round of prompts, resulting in the i-th round of dialogue, and so on. The following explanation uses obtaining response content i as an example.
[0160] E1: Retrieves the query statement of the i-th round of dialogue, the prompt words of the (i-1)-th round, and the response content for the (i-1)-th round of prompt words.
[0161] The query statement in the i-th round of the dialogue is the query statement entered by the user for the i-th time. Each round of the dialogue includes a query statement and the response content used to reply to the query statement.
[0162] The prompt for round i-1 is constructed based on the query statement of round i-1, and is obtained by concatenating the prompt for round i-2, the response content to the prompt for round i-2, and the query statement of round i-1.
[0163] The response to the prompt in round i-1 is also the reply to the query in round i-1.
[0164] E2: Concatenate the prompt word of round i-1, the response content to the prompt word of round i-1, and the query statement of round i to obtain the prompt word of round i.
[0165] For example, the query statement in the first round of dialogue is represented as u1, the first round of prompts is represented as (u1), and the response to the first round of prompts is represented as a1. The query statement in the second round of dialogue is represented as u2, the second round of prompts is represented as (u1, a1, u2), and the response to the second round of prompts is represented as a2. And so on, the query statement in the i-th round of dialogue is represented as ui, the i-th round of prompts is represented as (u1, a1, u2, ..., ui), and the response to the i-th round of prompts is represented as ai.
[0166] E3: Based on the multiple word segments and historical word segments included in the i-th round prompt, match the word segments with the same word segmentation order according to the word segmentation order to obtain the prefix word segmentation group and the remaining word segmentation group.
[0167] If the user does not delete the dialogue in a multi-turn conversation, the statement formed by the prefix segmentation group is generally the prompt word for the (i-1)th turn. If the user deletes a certain turn of the dialogue in a multi-turn conversation, taking the deletion of the j-th turn dialogue as an example, where j is less than or equal to i-1, the j-th turn dialogue includes the query statement of the j-th turn dialogue and the response content for the j-th turn prompt word.
[0168] In response to receiving a deletion instruction for the j-th round of dialogue in the previous i-1 rounds of dialogue, the word segment included in the j-th round of dialogue is deleted from the i-1 round prompt words, and if the word segment included in the j-th round of dialogue is stored in the graphics processor, the key-value cache corresponding to the word segment included in the j-th round of dialogue is deleted from the graphics processor or the central processing unit; if the word segment included in the j-th round of dialogue is stored in the central processing unit, the key-value cache corresponding to the word segment included in the j-th round of dialogue is deleted from the central processing unit.
[0169] Wherein, "responding to" is used to indicate the conditions or states on which the operation performed depends. When the conditions or states on which it depends are met, one or more operations performed can be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.
[0170] For example, if a user engages in two rounds of dialogue, the second round of prompts would be represented as (u1, a1, u2). If the user actively deletes the second round of dialogue, the second round of prompts would be updated to (u1, a1). The key-value cache corresponding to the word segments included in u2 and a2 can then be deleted from the graphics processor or central processing unit (CPU) to further free up storage space. After obtaining the query statement for a new round of dialogue, the third round of prompts could be represented as (u1, a1, u3).
[0171] Therefore, this method of deleting corresponding dialogues based on deletion instructions is more in line with the current user behavior pattern in multi-turn dialogue request scenarios. This allows the corresponding key-value cache to be deleted from the graphics processor or central processing unit, that is, the stored key-value cache is updated according to the deletion instruction, so as to further free up the storage space of the graphics processor or central processing unit, increase the number of key-value cache reuses, improve inference speed, and thus improve the user experience.
[0172] As one possible implementation, in a multi-turn dialogue request scenario, if no deletion instruction is obtained before obtaining the i-th round prompt, it means that the i-th round prompt and the (i-1)-th round prompt differ only in the response content for the (i-1)-th round prompt and the query content of the i-th round dialogue. Based on this, there is no need to perform matching based on historical word segments. Instead, the word segments included in the (i-1)-th round prompt and the word segments included in the response content obtained for the (i-1)-th round prompt are directly used as prefix word segments. The remaining word segments are then obtained based on the prefix word segments and the i-th round prompt.
[0173] If the deletion instruction is obtained before obtaining the i-th round prompt, it means that the i-th round prompt differs from the (i-1)-th round prompt not only in the response content to the (i-1)-th round prompt and the query content of the i-th round dialogue, but also in the word segments related to the deleted dialogue. Based on this, the aforementioned S502 can be executed, that is, the step of matching word segments with the same word order according to the word order of the multiple word segments included in the i-th round prompt and the historical word segment groups to obtain the prefix word segment group and the remaining word segment group.
[0174] Therefore, if a deletion command is obtained before the i-th round of prompts, after obtaining the i-th round of prompts, based on the multiple word segments included in the i-th round of prompts and historical word segments, word segments with the same word segmentation order are matched according to their order of arrangement to obtain a prefix word segmentation group and a remaining word segmentation group. If no deletion command is obtained before obtaining the i-th round of prompts, after obtaining the i-th round of prompts, no matching is required. Instead, the word segments included in the (i-1)-th round of prompts and the word segments included in the response content obtained for the (i-1)-th round of prompts are directly used as the prefix word segmentation group, and then the remaining word segmentation group is obtained based on the prefix word segmentation group and the i-th round of prompts. This reduces the number of matching operations in multi-turn dialogue request scenarios, increases the speed of determining the prefix word segmentation group, and thus improves the reasoning speed and user experience.
[0175] E4: If it fails to retrieve the key-value cache corresponding to the prefix segmentation group from the graphics processor, then retrieve the key-value cache corresponding to the prefix segmentation group from the central processing unit.
[0176] E5: Send the key-value cache of the prefix segmentation group and the remaining segmentation group to the inference engine so that the inference engine can obtain the response content for the i-th round prompt word based on the key-value cache of the prefix segmentation group and the remaining segmentation group.
[0177] Inference engines typically employ two key steps in their inference computation process: a prefill phase and a decoder phase. The prefill phase involves providing the model with contextual information or hints before it begins generating text. The decoder phase, on the other hand, generates text character by character based on the intermediate results from the prefill phase, continuing until a termination condition is met. Key-value caching is generally reused during the decoder phase. Large language models are primarily composed of multiple decoders, each including a forward propagation layer and a self-attention layer, with the self-attention layer being the core of the transformer architecture.
[0178] Assuming the i-th round prompt word is (u1, a1, u2, a2, ..., ui-1, ai-1, ui), taking the Very Large Language Model (vLLM) as the model used as the inference engine as an example, in related technologies, all the word segments included in the i-th round prompt word are calculated during the inference calculation process. However, in the embodiment of this application, a portion of the word segment key-value cache (i.e., the key-value cache of the prefix word segment group) can be reused, and only the newly appearing word segments (i.e., the remaining word segments) need to be calculated. Then the calculation of the self-attention layer can be expressed as formula (1):
[0179]
[0180] Among them, O ui This represents the response to the i-th prompt; Q ui This represents the multiple words included in the remaining word group in the i-th round of prompts. Taking a remaining word group containing n words as an example, it can be represented as Q. ui =(q1,q2,…,qn);K u1,a1;...;ui The key in the key-value cache of the i-th round prompt word is the representation of each position in the i-th round prompt word, such as K = (k1, k2, ..., kn). In the embodiments of this application, the key can be obtained from the key-value cache of the prefix segment group and calculated only for the segments included in the remaining segment group. V represents the transpose of the representation of each position in the i-th round of prompts; u1,a1;...;ui This represents the value corresponding to the key in the key-value cache of the i-th round prompt word, that is, the weight corresponding to each position in the i-th round prompt word, which can be represented as V = (v1, v2, ..., vn). In this embodiment, this value can be obtained from the key-value cache of the prefix segmentation group and calculated only for the segments included in the remaining segmentation groups; M ui This represents the mask matrix, used to ensure that subsequent word segmentation can obtain information from previous word segmentation; d is K. u1,a1;...;uiThe corresponding key dimension.
[0181] It is understood that all data (such as query statements) collected in the specific embodiments of this application are collected with the consent and authorization of the data subject (such as user, organization or enterprise). When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0182] The key-value caching reuse method provided in this application primarily relates to artificial intelligence (AI) technology. AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to obtain optimal results—a theory, method, technology, and application system. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines capable of reacting in a manner similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.
[0183] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and intelligent transportation.
[0184] In the embodiments of this application, the main artificial intelligence technologies involved include the aforementioned machine learning, etc. For example, in the embodiments of this application, the various services involved in the inference engine are all implemented through models, which can all be trained using machine learning, an artificial intelligence technology.
[0185] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and pre-trained learning. Pre-trained models are the latest development in deep learning, integrating all of these techniques.
[0186] Pre-trained models (PTMs), also known as foundational models or large models, refer to deep neural networks (DNNs) with a large number of parameters. These DNNs are trained on massive amounts of unlabeled data. Leveraging the function approximation capabilities of large-parameter DNNs, PTMs extract common features from the data. Through fine-tuning, parameter-efficient fine-tuning (PEFT), and prompt-tuning techniques, they are suitable for downstream tasks. Therefore, pre-trained models can achieve ideal results in few-shot or zero-shot scenarios. PTMs can be categorized according to the data modalities they process, such as language models (e.g., ELMO, BERT, GPT), visual models (e.g., Swin-transformer, ViT, V-MOE), speech models (e.g., VALL-E), and multimodal models (e.g., ViBERT, CLIP, Flamingo, Gato). Multimodal models refer to models that represent features from two or more data modalities. Pre-trained models are important tools for outputting Artificial Intelligence Generated Content (AIGC) and can also serve as a general interface connecting multiple specific task models.
[0187] To facilitate a further understanding of the technical solutions provided in the embodiments of this application, the following describes the key-value cache reuse method in its entirety, taking the server as the execution subject of the key-value cache reuse method provided in the embodiments of this application as an example.
[0188] Numerous large language models have emerged, such as GPT and pre-trained language models (Llama, Large Language Model Meta AI), making many previously intractable natural language processing tasks possible. Therefore, LLM inference services play a crucial role in a wide range of fields and have become a major workload for cloud services. Although the inference process of large language models requires fewer resources than the training process, it still requires an extended computing platform with accelerators such as GPUs to provide efficient processing and fast response. The key-value caching reuse method provided in this application can be used as an extension of large language models, deployed in the inference service of large language models, accelerating multi-turn interactions between users and large language models, thereby reducing interaction latency with users, improving the user experience, increasing the throughput of requests that GPUs can handle, and reducing the cost of GPU inference services. Application scenarios include, but are not limited to, multi-turn dialogue request scenarios, code completion scenarios, and long text query scenarios.
[0189] The following explanation will take a multi-turn dialogue request scenario as an example.
[0190] See Figure 7 This figure illustrates an application scenario of a key-value caching reuse method provided in an embodiment of this application. See also... Figure 8 The figure is a schematic diagram of an application scenario of a key-value caching reuse method provided in an embodiment of this application.
[0191] The architecture of the key-value cache reuse method includes a key-value cache manager 701, a storage manager 702, an inference engine 703, a graphics processor 704, and a central processing unit 705. The key-value cache manager 701 and the storage manager 702 can be different functional units within the same physical hardware, such as... Figure 7 As shown, it can also be different physical hardware, such as Figure 8 As shown.
[0192] Among them, the key-value cache manager 701 is used to manage the key-value cache from different session requests in order to realize the reuse of the key-value cache of prefix segmentation groups.
[0193] Storage Manager 702 is used to detect the remaining video memory space of the graphics processor and the remaining memory space of the central processing unit, so that the graphics processor and the central processing unit have enough storage space to store key-value cache.
[0194] Inference Engine 703 is used to generate corresponding response content for session requests.
[0195] The graphics processor 704 is used to store key-value caches and accelerate inference calculations.
[0196] The 705 central processing unit is used to store key-value cache.
[0197] The reuse method of key-value caching is explained below.
[0198] S1: Retrieve the query statement carrying the session identifier.
[0199] S2: If the session identifier determines that the query statement does not contain multi-round suggestion words with a relationship, then obtain the first round suggestion words based on the query statement.
[0200] If the query statement does not contain any related multi-round suggestion words based on the session identifier, it means that the query statement based on the session identifier is appearing for the first time. In this case, the first round suggestion word is obtained based on the query statement, such as directly using the query statement as the first round suggestion word.
[0201] S3: For the first round of prompts, send the multiple word segments included in the first round of prompts to any one of the multiple inference engines.
[0202] Since there are no reusable prefix segmentation groups for the first round of prompts, the multiple segments included in the first round of prompts can be directly sent to any one of the multiple inference engines. As a result, subsequent prompts that are related to the first round of prompts will be sent to that inference engine for inference calculation based on the session identifier.
[0203] S4: If the query statement is determined to have i-1 rounds of related suggestions based on the session identifier, then obtain the i-th round of suggestions based on the query statement.
[0204] If the query statement is determined to have a related first i-1 round prompt words based on the session identifier, it means that the initial prompt word is the prompt word that appears i times based on the session identifier. Then, the i-th round prompt word is obtained based on the query statement. For example, the i-th round prompt word, the response content for the i-th round prompt word, and the query statement of the i-th round dialogue are concatenated to obtain the i-th round prompt word.
[0205] S5: Based on the multiple word segments and historical word segments included in the i-th round prompt, match the word segments with the same word segmentation order according to the word segmentation order to obtain the prefix word segmentation group and the remaining word segmentation group.
[0206] For example, the k-th segment among the multiple segments included in the i-th round prompt is matched with the k-th segment in the historical segment group, where k is an integer starting from 1. If the match is successful, k+1 is updated to k, and the step of matching the k-th segment among the multiple segments included in the i-th round prompt with the k-th segment in the historical segment group is executed again; if the match fails, the successfully matched segment in the i-th round prompt is taken as the prefix segment group, and the segment in the multiple segments included in the i-th round prompt, excluding the segment included in the prefix segment group, is taken as the remaining segment group.
[0207] S6: If the key-value cache corresponding to the prefix segmentation group is obtained from the graphics processor, the key-value cache of the prefix segmentation group and the remaining segmentation group are sent to the target inference engine so that the target inference engine can obtain the response content for the i-th round prompt word based on the key-value cache of the prefix segmentation group and the remaining segmentation group.
[0208] S7: If it fails to retrieve the key-value cache corresponding to the prefix segmentation group from the graphics processor, retrieve the key-value cache corresponding to the prefix segmentation group from the central processing unit, and send the key-value cache of the prefix segmentation group and the remaining segmentation group to the target inference engine so that the target inference engine can obtain the response content for the i-th round prompt word based on the key-value cache of the prefix segmentation group and the remaining segmentation group.
[0209] S8: If it fails to retrieve the key-value cache corresponding to multiple prefix segments from the central processing unit and the graphics processing unit, the prefix segment group and the remaining segment group are sent to the target inference engine so that the target inference engine can obtain the response content for the i-th round prompt word based on the prefix segment group and the remaining segment group.
[0210] In this process, multi-turn prompts with the same session identifier are executed through the same inference engine. The target inference engine is the inference engine among multiple inference engines used to process multi-turn prompts that are related to the i-th round prompt. In other words, multi-turn prompts with related relationships are distributed to the same inference engine based on the session identifier, so that the same inference engine can better reuse the key-value cache of the prefix segmentation group for multi-turn prompts with related relationships.
[0211] It should be noted that during the reasoning process of the target inference engine, it calls the graphics processor to accelerate the inference calculation process. The graphics processor is used to calculate the key value cache for each word segment and store the key value cache in the graphics processor or the central processing unit, which will be explained below.
[0212] S9: Obtain the remaining video memory space of the graphics processor and the cache of key values to be stored generated by the i-th round of prompts.
[0213] The key-value cache to be stored is determined based on the key-value cache of a word segment and the number of words to be stored in the remaining word segment groups. For example, if the key-value cache of a word segment is 20k, and a word segment is stored each time, then the key-value cache for storing a word segment each time is 20k, that is, the key-value cache to be stored is 20k.
[0214] S10: If the remaining video memory space is less than the key value cache to be stored, then determine the second target key value cache from the key value cache stored in the graphics processor, store the second target key value cache in the central processing unit, delete the second target key value cache from the graphics processor, and store the key value cache to be stored in the graphics processor.
[0215] The sum of the second target key-value cache and the remaining video memory space is greater than or equal to the key-value cache to be stored, so as to reserve enough storage space in the graphics processor.
[0216] The process of storing the second target key-value cache into the central processing unit is as follows:
[0217] Obtain the remaining memory space of the central processing unit; if the remaining memory space is less than the second target key-value cache, determine the third target key-value cache from the central processing unit, delete the third target key-value cache from the central processing unit, and store the second target key-value cache in the central processing unit; if the remaining memory space is greater than or equal to the second target key-value cache, cache the second target key-value cache in the central processing unit.
[0218] The sum of the third target key-value cache and the remaining memory space is greater than or equal to the second target key-value cache. The storage time of the third target key-value cache is greater than the preset time threshold, so that the central processing unit reserves enough space to store the second target key-value cache, and the third target key-value cache released from the central processing unit is a key-value cache with a longer storage time.
[0219] S11: If the remaining video memory space is greater than or equal to the key value cache to be stored, then store the key value cache to be stored in the graphics processor.
[0220] Furthermore, when storing key-value caches, they can be stored based on quadruples. The quadruples consist of the session identifier, the time when the i-th round session identifier was obtained, the multiple word segments generated from the i-th round prompt, and the key-value cache corresponding to each of these word segments.
[0221] It should be noted that if a deletion instruction for the j-th round of dialogue is received before obtaining the i-th round prompt word, the word segmentation included in the j-th round of dialogue is deleted from the i-1 round prompt word, and the key-value cache corresponding to the word segmentation included in the j-th round of dialogue is deleted from the graphics processor or central processing unit. The j-th round of dialogue includes the query statement of the j-th round of dialogue and the response content for the j-th round prompt word, where j is less than or equal to i-1.
[0222] To evaluate the effectiveness of the method provided in this application, it is compared with four existing methods. The four existing methods are: (1) TensorRT-LLM (TensorRT-Large Language Model, TRTLLM) is an inference system for a large language model provided by NVIDIA; (2) vLLM is a high-throughput large language model inference system with efficient storage management and paged attention; (3) Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills (ChunkedPrefill) mixes the word segmentation in the decoding stage and the word segmentation in the prefilling stage together for inference, thereby improving the utilization of the graphics processor; (4) PrefixCache is an optional feature of vLLM, which can reuse key-value caches between requests through hashing.
[0223] There were two test scenarios. In the multi-turn dialogue request scenario, the inference engine model was ChatGLM3-6B (a dialogue pre-trained model), and the prompts were derived from ShareGPT (a tool for sharing and conversing with GPT). In the code completion scenario, the inference engine model was StarCoderBase-7B (a dialogue pre-trained model), and the prompts were derived from BigCode's Stack (a technique for handling large-scale, highly complex programming code). The comparison results are detailed in Table 1.
[0224] Table 1
[0225]
[0226]
[0227] Table 1 shows the overall performance of the method provided in this application embodiment and four other methods. As can be seen from Table 1, the method provided in this application embodiment achieves a 37% and 190% throughput improvement compared to the previous best method, PrefixCache, in multi-turn dialogue request scenarios and code completion scenarios, respectively, demonstrating the significant advantages of the method provided in this application embodiment. The advantage of the method provided in this application embodiment in the code completion scenario is much greater than that in round-robin dialogue, mainly because code completion involves more repetitive computations and requires larger storage to house the key-value cache.
[0228] To demonstrate the effectiveness of the method design provided in this application embodiment, an ablation experiment was also conducted, ablation the method provided in this application embodiment into two methods: (1) Immediately release the corresponding key-value cache after each inference engine inference calculation, which can be seen as removing the reuse of the key-value cache of the prefix segmentation group and the dynamic unloading strategy from the method provided in this application embodiment, thus becoming the vLLM method. (2) Reuse the key-value cache of the prefix segmentation group, and delete a portion of the key-value cache when the remaining video memory space of the graphics processor is insufficient, which can be seen as removing the dynamic unloading strategy from the method provided in this application embodiment, denoted as the CrossKV method.
[0229] See Figure 9 This figure is a schematic diagram illustrating the result of an end-to-end delay decomposition provided in an embodiment of this application. Figure 9The results of evaluating the end-to-end latency decomposition of the method, vLLM method, and CrossKV method provided in the embodiments of this application are shown. It can be found that the method provided in the embodiments of this application achieves end-to-end latency reduction of 37% / 66% and 47% / 68% respectively compared with the CrossKV method and vLLM method in multi-turn dialogue request scenario and code completion scenario. The main contribution to the latency reduction comes from the decoding stage, which also proves the effectiveness of the method provided in the embodiments of this application.
[0230] To demonstrate that the method provided in this application is applicable to various scenarios of multi-turn interactions, experiments were conducted considering three factors: the number of dialogue turns, the input length, and the output length. The results are as follows: Figure 10 As shown.
[0231] See Figure 10 This figure shows the experimental results of a multi-turn interaction scenario provided in an embodiment of this application. Figure 10 As can be seen from the embodiments of this application, the method provided exhibits optimal performance under all three factors, and the advantage of the method provided by the embodiments of this application becomes more obvious as the number of rounds increases, because there is more reusable information with a larger number of rounds. Secondly, as the input length or output length increases, the advantage of the PrefixCache method over the vLLM method gradually disappears, because the memory of the graphics processor is very limited and cannot store enough key-value cache for longer context lengths.
[0232] In response to the key-value cache reuse method described above, this application also provides a corresponding key-value cache reuse device so that the above key-value cache reuse method can be applied and implemented in practice.
[0233] See Figure 11 This figure is a schematic diagram of the structure of a key-value cache reuse device provided in an embodiment of this application. Figure 11 As shown, the key-value cache multiplexing device 1100 includes: an acquisition unit 1101, a matching unit 1102, a multiplexing unit 1103, and a sending unit 1104;
[0234] The acquisition unit 1101 is used to acquire the i-th round of prompt words, where i is an integer greater than 1;
[0235] The matching unit 1102 is used to match words with the same word order according to the word order of the multiple word segments and historical word groups included in the i-th round prompt word, to obtain a prefix word group and a remaining word group. The historical word group is a word group with a word order obtained based on the previous i-1 round prompt words. The previous i-1 round prompt words are related to the i-th round prompt word. The prefix word group includes the word segments that were successfully matched in the i-th round prompt word when the first match failed. The remaining word group includes the word segments other than the word segments included in the prefix word group from the multiple word segments included in the i-th round prompt word.
[0236] The multiplexing unit 1103 is used to retrieve the key-value cache corresponding to the prefix segmentation group from the central processing unit if it fails to retrieve the key-value cache corresponding to the prefix segmentation group from the graphics processor.
[0237] The sending unit 1104 is used to send the key-value cache of the prefix segmentation group and the remaining segmentation group to the inference engine, so that the inference engine can obtain the response content for the i-th round prompt word based on the key-value cache of the prefix segmentation group and the remaining segmentation group.
[0238] As can be seen from the above technical solutions, the apparatus provided in this application includes an acquisition unit, a matching unit, a multiplexing unit, and a sending unit. If the graphics processor's memory can no longer store the key-value cache, the key-value cache is not directly released; instead, a portion of the key-value cache is stored in the central processing unit (CPU). This allows the graphics processor and CPU to work together to store the key-value cache, further increasing the number of reusable key-value caches. Moreover, compared to other storage spaces, the interaction speed between the CPU and the graphics processor is faster, ensuring inference speed. Thus, while ensuring inference speed, the number of reusable key-value caches is increased, expanding the scope of application scenarios. The process of reusing the key-value cache is described below.
[0239] The i-th round prompt word is obtained by the acquisition unit. The matching unit matches the multiple segments included in the i-th round prompt word with the historical segmentation groups generated from the previous i-1 round prompt words, matching segments with the same order to obtain a prefix segmentation group and a remaining segmentation group. This matching determines the reusable and non-reusable segments in the i-th round prompt word. Then, the reuse unit retrieves the key-value cache of the prefix segmentation group from the graphics processor. If retrieval fails, the key-value cache of that prefix segmentation group may be moved to the central processing unit (CPU) for storage; in this case, the key-value cache corresponding to the prefix segmentation group is retrieved from the CPU. Finally, the sending unit sends the key-value cache of the prefix segmentation group and the remaining segmentation group to the inference engine. The inference engine then performs inference based on these caches to obtain the response content for the i-th round prompt word. Therefore, during the calculation process, the inference engine does not need to recalculate the key-value cache of the prefix segmentation group; instead, it directly reuses the prefix segmentation group's key-value cache and only calculates the key-value cache of the remaining segmentation group, improving inference speed.
[0240] As one possible implementation, the device further includes a storage unit, used before obtaining the i-th round of prompt words, for:
[0241] Retrieve the key-value cache generated by the prompt words in the (i-1)th round;
[0242] The key-value cache generated by the (i-1)th round of prompts is stored in the graphics processor;
[0243] Obtain the remaining video memory space of the graphics processor;
[0244] If the remaining video memory space of the graphics processor is less than the preset storage space, then a first target key value cache is determined from the key value cache stored in the graphics processor, the first target key value cache is stored in the central processing unit, and the first target key value cache is deleted from the graphics processor. The sum of the first target key value cache and the remaining video memory space is greater than or equal to the preset storage space.
[0245] As one possible implementation, the acquisition unit 1101 is specifically used to acquire the remaining video memory space of the graphics processor and the key-value cache to be stored generated by the i-th round of prompt words. The key-value cache to be stored is determined based on the key-value cache of a word segment and the number of words to be stored in the remaining word segment group.
[0246] The device further includes a storage unit for:
[0247] If the remaining video memory space is less than the key-value cache to be stored, then a second target key-value cache is determined from the key-value cache stored in the graphics processor, the second target key-value cache is stored in the central processing unit, the second target key-value cache is deleted from the graphics processing unit, and the key-value cache to be stored is stored in the graphics processor. The sum of the second target key-value cache and the remaining video memory space is greater than or equal to the key-value cache to be stored.
[0248] If the remaining video memory space is greater than or equal to the key-value cache to be stored, then the key-value cache to be stored is stored in the graphics processor.
[0249] As one possible implementation, the device further includes a storage unit for:
[0250] Obtain the remaining memory space of the central processing unit;
[0251] If the remaining memory space is less than the second target key-value cache, then a third target key-value cache is determined from the central processing unit, the third target key-value cache is deleted from the central processing unit, and the second target key-value cache is stored in the central processing unit. The sum of the third target key-value cache and the remaining memory space is greater than or equal to the second target key-value cache, and the storage time of the third target key-value cache is greater than a preset time threshold.
[0252] If the remaining memory space is greater than or equal to the second target key-value cache, then the second target key-value cache is stored in the central processing unit.
[0253] As one possible implementation, the sending unit 1101 is configured to send the prefix segmentation group and the remaining segmentation group to the inference engine if it fails to obtain the key-value cache corresponding to the prefix segmentation group from the central processing unit and the graphics processor, so that the inference engine can obtain the response content for the i-th round prompt word based on the prefix segmentation group and the remaining segmentation group.
[0254] As one possible implementation, the acquisition unit 1101 is specifically used for:
[0255] Obtain an initial prompt word including a session identifier, the session identifier being used to identify the association;
[0256] If it is determined from the session identifier that the initial prompt word has a preceding i-1 round prompt word with the aforementioned association, then the i-th round prompt word is obtained based on the initial prompt word;
[0257] If it is determined from the session identifier that there is no prompt word with the associated relationship in the initial prompt word, then the first round prompt word is obtained based on the initial prompt word;
[0258] The transmitting unit 1104 is specifically used for:
[0259] For the i-th round prompt word, the key-value cache of the prefix segmentation group and the remaining segmentation group are sent to the target inference engine among the multiple inference engines. The target inference engine is the inference engine among the multiple inference engines used to process prompt words that have the association relationship with the i-th round prompt word.
[0260] For the first round of prompts, the multiple word segments included in the first round of prompts are sent to any one of the multiple inference engines.
[0261] As one possible implementation, the device further includes a storage unit for:
[0262] For the first round of prompt words, a metadata group of the first round of prompt words is stored. The metadata group of the first round of prompt words includes the session identifier, multiple word segments generated by the first round of prompt words, and key-value caches corresponding to each word segment.
[0263] For the i-th round prompt word, the metadata group of the (i-1)-th round prompt word is updated to obtain the metadata group of the i-th round prompt word. The metadata group of the i-th round prompt word includes the session identifier, multiple word segments generated by the i-th round prompt word, and key-value caches corresponding to each word segment.
[0264] As one possible implementation, the matching unit 1102 is specifically used for:
[0265] The kth word among the multiple words included in the i-th round of prompts is matched with the kth word in the historical word group, where k is an integer starting from 1;
[0266] If the match is successful, then k+1 is updated to k, and the step of matching the kth word among the multiple words included in the i-th round of prompts with the kth word in the historical word group is executed;
[0267] If the match fails, the successfully matched segment in the i-th round prompt word is taken as the prefix segment group, and the segment in the multiple segments included in the i-th round prompt word other than the segment included in the prefix segment group is taken as the remaining segment group.
[0268] As one possible implementation, the acquisition unit 1101 is specifically used for:
[0269] Obtain the query statement of the i-th round of dialogue, the prompt words of the (i-1)-th round, and the response content for the (i-1)-th round of prompt words;
[0270] The i-th round prompt word, the response content to the i-th round prompt word, and the query statement of the i-th round dialogue are concatenated to obtain the i-th round prompt word.
[0271] As one possible implementation, the apparatus further includes a deletion unit, configured to, in response to receiving a deletion instruction for the j-th round of dialogue in the preceding i-1 rounds of dialogue, delete the word segment included in the j-th round of dialogue from the i-1-th round prompt word, and delete the key-value cache corresponding to the word segment included in the j-th round of dialogue from the graphics processor, wherein the j-th round of dialogue includes the query statement of the j-th round of dialogue and the response content for the j-th round prompt word, and j is less than or equal to i-1.
[0272] As one possible implementation, the matching unit 1102 is specifically used for:
[0273] If the deletion instruction is not obtained before obtaining the i-th round prompt word, then the word segmentation included in the (i-1)-th round prompt word and the word segmentation included in the response content obtained for the (i-1)-th round prompt word are used as the prefix word segmentation group, and the remaining word segmentation group is obtained based on the prefix word segmentation group and the i-th round prompt word;
[0274] If the deletion instruction is obtained before the i-th round prompt word is obtained, then based on the multiple word segments and historical word segments included in the i-th round prompt word, the word segments with the same word segmentation order are matched according to the word segmentation order to obtain the prefix word segmentation group and the remaining word segmentation group.
[0275] This application also provides a computer device, which can be a server or a terminal device. The computer device provided in this application will be described below from a hardware implementation perspective. Figure 12 The diagram shown is a schematic of the server's structure. Figure 13 The diagram shown is a structural schematic of the terminal device.
[0276] See Figure 12This figure is a schematic diagram of a server structure provided in an embodiment of this application. The server 1400 can vary considerably due to different configurations or performance. It may include one or more processors 1422, such as a central processing unit (CPU), memory 1432, and one or more application programs 1442 or data storage media 1430 (e.g., one or more mass storage devices). The memory 1432 and storage media 1430 can be temporary or persistent storage. The program stored in the storage media 1430 may include one or more modules (not shown in the figure), each module may include a series of instruction operations on the server. Furthermore, the processor 1422 may be configured to communicate with the storage media 1430 and execute the series of instruction operations in the storage media 1430 on the server 1400.
[0277] Server 1400 may also include one or more power supplies 1426, one or more wired or wireless network interfaces 1450, one or more input / output interfaces 1458, and / or one or more operating systems 1441, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM etc.
[0278] The steps performed by the server in the above embodiments can be based on this Figure 12 The server structure shown.
[0279] The processor 1422 is used to perform the following steps:
[0280] Get the prompt word for the i-th round, where i is an integer greater than 1;
[0281] Based on the multiple word segments and historical word segments included in the i-th round of prompts, word segments with the same word segmentation order are matched according to the word segmentation order to obtain prefix word segments and remaining word segments. The historical word segments are word segments with word segmentation order obtained based on the prompts of the previous i-1 rounds. The prompts of the previous i-1 rounds are related to the prompts of the i-th round. The prefix word segments include the word segments that were successfully matched in the prompts of the i-th round when the first match failed. The remaining word segments include word segments other than those included in the prefix word segments from the multiple word segments included in the prompts of the i-th round.
[0282] If it fails to retrieve the key-value cache corresponding to the prefix segmentation group from the graphics processor, then retrieve the key-value cache corresponding to the prefix segmentation group from the central processing unit.
[0283] The key-value cache of the prefix segmentation group and the remaining segmentation group are sent to the inference engine so that the inference engine can obtain the response content for the i-th round prompt word based on the key-value cache of the prefix segmentation group and the remaining segmentation group.
[0284] Optionally, the processor 1422 may also execute method steps of any specific implementation of the key-value cache reuse method in the embodiments of this application.
[0285] See Figure 13 This figure is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. The description will be based on a smartphone as an example. Figure 13 The diagram shown is a partial structural block diagram of the smartphone, which includes: a radio frequency (RF) circuit 1510, a memory 1520, an input unit 1530, a display unit 1540, a sensor 1550, an audio circuit 1560, a Wi-Fi module 1570, a processor 1580, and a power supply 1590, among other components. Those skilled in the art will understand that... Figure 13 The smartphone structure shown does not constitute a limitation on smartphones and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0286] The following is combined Figure 13 A detailed introduction to the various components of a smartphone:
[0287] The RF circuit 1510 can be used to receive and transmit signals during information transmission or calls. In particular, it receives downlink information from the base station and processes it with the processor 1580; in addition, it transmits uplink data to the base station.
[0288] The memory 1520 can be used to store software programs and modules, and the processor 1580 runs the software programs and modules stored in the memory 1520 to realize various functions and data processing of the smartphone.
[0289] Input unit 1530 can be used to receive input numeric or character information and generate key signal inputs related to user settings and function control of the smartphone. Specifically, input unit 1530 may include touch panel 1531 and other input devices 1532. Touch panel 1531, also known as a touch screen, can collect touch operations on or near the user and drive corresponding connected devices according to a pre-set program. In addition to touch panel 1531, input unit 1530 may also include other input devices 1532. Specifically, other input devices 1532 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.
[0290] The display unit 1540 can be used to display information input by the user or information provided to the user, as well as various menus of the smartphone. The display unit 1540 may include a display panel 1541, which may optionally be configured as a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.
[0291] Smartphones may also include at least one sensor 1550, such as a light sensor, a motion sensor, and other sensors. Other sensors that smartphones may also be equipped with, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be detailed here.
[0292] Audio circuit 1560, speaker 1561, and microphone 1562 provide an audio interface between the user and the smartphone. Audio circuit 1560 converts received audio data into electrical signals and transmits them to speaker 1561, where speaker 1561 converts them into sound signals for output. On the other hand, microphone 1562 converts collected sound signals into electrical signals, which are received by audio circuit 1560, converted into audio data, and then processed by processor 1580 before being transmitted via RF circuit 1510 to, for example, another smartphone, or the audio data can be output to memory 1520 for further processing.
[0293] The processor 1580 is the control center of the smartphone, connecting various parts of the smartphone through various interfaces and lines. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 1520, and by calling data stored in the memory 1520. Optionally, the processor 1580 may include one or more processing units.
[0294] The smartphone also includes a power supply 1590 (such as a battery) that supplies power to various components. Preferably, the power supply can be logically connected to the processor 1580 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.
[0295] Although not shown, smartphones may also include a camera, Bluetooth module, etc., which will not be described in detail here.
[0296] In this embodiment of the application, the memory 1520 included in the smartphone can store computer programs and transmit the computer programs to the processor.
[0297] The processor 1580 included in the smartphone can execute the key-value cache reuse method provided in the above embodiments according to the instructions in the computer program.
[0298] This application also provides a computer-readable storage medium for storing a computer program for executing the key-value cache reuse method provided in the above embodiments.
[0299] This application also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the key-value cache reuse method provided in various optional implementations of the above aspects.
[0300] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium can be at least one of the following media: read-only memory (ROM), RAM, magnetic disk or optical disk, and other media that can store computer programs.
[0301] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0302] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The device and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the solution in this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0303] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for reusing key-value caches, characterized in that, The method includes: Get the prompt word for the i-th round, where i is an integer greater than 1; Based on the multiple word segments and historical word segments included in the i-th round of prompts, word segments with the same word segmentation order are matched according to the word segmentation order to obtain prefix word segments and remaining word segments. The historical word segments are word segments with word segmentation order obtained based on the prompts of the previous i-1 rounds. The prompts of the previous i-1 rounds are related to the prompts of the i-th round. The prefix word segments include the word segments that were successfully matched in the prompts of the i-th round when the first match failed. The remaining word segments include word segments other than those included in the prefix word segments from the multiple word segments included in the prompts of the i-th round. If it fails to retrieve the key-value cache corresponding to the prefix segmentation group from the graphics processor, then retrieve the key-value cache corresponding to the prefix segmentation group from the central processing unit. The key-value cache of the prefix segmentation group and the remaining segmentation group are sent to the inference engine so that the inference engine can obtain the response content for the i-th round prompt word based on the key-value cache of the prefix segmentation group and the remaining segmentation group.
2. The method according to claim 1, characterized in that, Before obtaining the i-th round of prompts, the method further includes: Retrieve the key-value cache generated by the prompt words in the (i-1)th round; The key-value cache generated by the (i-1)th round of prompts is stored in the graphics processor; Obtain the remaining video memory space of the graphics processor; If the remaining video memory space of the graphics processor is less than the preset storage space, then a first target key value cache is determined from the key value cache stored in the graphics processor, the first target key value cache is stored in the central processing unit, and the first target key value cache is deleted from the graphics processor. The sum of the first target key value cache and the remaining video memory space is greater than or equal to the preset storage space.
3. The method according to claim 1, characterized in that, The method further includes: Obtain the remaining video memory space of the graphics processor and the key-value cache to be stored generated by the i-th round of prompt words. The key-value cache to be stored is determined based on the key-value cache of a word segment and the number of words to be stored in the remaining word segment group. If the remaining video memory space is less than the key-value cache to be stored, then a second target key-value cache is determined from the key-value cache stored in the graphics processor, the second target key-value cache is stored in the central processing unit, the second target key-value cache is deleted from the graphics processing unit, and the key-value cache to be stored is stored in the graphics processor. The sum of the second target key-value cache and the remaining video memory space is greater than or equal to the key-value cache to be stored. If the remaining video memory space is greater than or equal to the key-value cache to be stored, then the key-value cache to be stored is stored in the graphics processor.
4. The method according to claim 3, characterized in that, The step of caching and storing the second target key value in the central processing unit includes: Obtain the remaining memory space of the central processing unit; If the remaining memory space is less than the second target key-value cache, then a third target key-value cache is determined from the central processing unit, the third target key-value cache is deleted from the central processing unit, and the second target key-value cache is stored in the central processing unit. The sum of the third target key-value cache and the remaining memory space is greater than or equal to the second target key-value cache, and the storage time of the third target key-value cache is greater than a preset time threshold. If the remaining memory space is greater than or equal to the second target key-value cache, then the second target key-value cache is stored in the central processing unit.
5. The method according to claim 1, characterized in that, The method further includes: If it fails to retrieve the key-value cache corresponding to the prefix segmentation group from the central processing unit and the graphics processing unit, the prefix segmentation group and the remaining segmentation group are sent to the inference engine so that the inference engine can obtain the response content for the i-th round prompt word based on the prefix segmentation group and the remaining segmentation group.
6. The method according to claim 1, characterized in that, The process of obtaining the i-th round of prompts includes: Obtain an initial prompt word including a session identifier, the session identifier being used to identify the association; If it is determined from the session identifier that the initial prompt word has a preceding i-1 round prompt word with the aforementioned association, then the i-th round prompt word is obtained based on the initial prompt word; If it is determined from the session identifier that there is no prompt word with the associated relationship in the initial prompt word, then the first round prompt word is obtained based on the initial prompt word; The step of sending the key-value cache of the prefix segmentation group and the remaining segmentation group to the inference engine includes: For the i-th round prompt word, the key-value cache of the prefix segmentation group and the remaining segmentation group are sent to the target inference engine among the multiple inference engines. The target inference engine is the inference engine among the multiple inference engines used to process prompt words that have the association relationship with the i-th round prompt word. For the first round of prompts, the multiple word segments included in the first round of prompts are sent to any one of the multiple inference engines.
7. The method according to claim 6, characterized in that, The method further includes: For the first round of prompt words, a metadata group of the first round of prompt words is stored. The metadata group of the first round of prompt words includes the session identifier, multiple word segments generated by the first round of prompt words, and key-value caches corresponding to each word segment. For the i-th round prompt word, the metadata group of the (i-1)-th round prompt word is updated to obtain the metadata group of the i-th round prompt word. The metadata group of the i-th round prompt word includes the session identifier, multiple word segments generated by the i-th round prompt word, and key-value caches corresponding to each word segment.
8. The method according to claim 1, characterized in that, The step involves matching words with the same order based on the multiple word segments and historical word segments included in the i-th round of prompts, to obtain prefix word segments and remaining word segments, including: The kth word among the multiple words included in the i-th round of prompts is matched with the kth word in the historical word group, where k is an integer starting from 1; If the match is successful, then k+1 is updated to k, and the step of matching the kth word among the multiple words included in the i-th round of prompts with the kth word in the historical word group is executed; If the match fails, the successfully matched segment in the i-th round prompt word is taken as the prefix segment group, and the segment in the multiple segments included in the i-th round prompt word other than the segment included in the prefix segment group is taken as the remaining segment group.
9. The method according to claim 1, characterized in that, The process of obtaining the i-th round of prompts includes: Obtain the query statement of the i-th round of dialogue, the prompt words of the (i-1)-th round, and the response content for the (i-1)-th round of prompt words; The i-th round prompt word, the response content to the i-th round prompt word, and the query statement of the i-th round dialogue are concatenated to obtain the i-th round prompt word.
10. The method according to claim 9, characterized in that, The method further includes: In response to receiving a deletion instruction for the j-th round of dialogue in the previous i-1 rounds of dialogue, the word segment included in the j-th round of dialogue is deleted from the i-1 round prompt word, and the key-value cache corresponding to the word segment included in the j-th round of dialogue is deleted from the graphics processor. The j-th round of dialogue includes the query statement of the j-th round of dialogue and the response content for the j-th round prompt word, where j is less than or equal to i-1.
11. The method according to claim 10, characterized in that, The method further includes: If the deletion instruction is not obtained before obtaining the i-th round prompt word, then the word segmentation included in the (i-1)-th round prompt word and the word segmentation included in the response content obtained for the (i-1)-th round prompt word are used as the prefix word segmentation group, and the remaining word segmentation group is obtained based on the prefix word segmentation group and the i-th round prompt word; If the deletion instruction is obtained before the i-th round prompt word is obtained, then the step of matching the words with the same word order according to the multiple word segments and historical word segments included in the i-th round prompt word to obtain the prefix word segment group and the remaining word segment group is executed.
12. A key-value cache reuse device, characterized in that, The device includes: an acquisition unit, a matching unit, a multiplexing unit, and a transmission unit; The acquisition unit is used to acquire the i-th round of prompt words, where i is an integer greater than 1; The matching unit is configured to match words with the same word order according to the word order of the multiple word segments included in the i-th round prompt word and the historical word segment group, to obtain a prefix word segment group and a remaining word segment group. The historical word segment group is a word segment with a word order obtained based on the previous i-1 round prompt words. The previous i-1 round prompt words are related to the i-th round prompt word. The prefix word segment group includes the word segments that were successfully matched in the i-th round prompt word when the first match failed. The remaining word segment group includes the word segments other than the word segments included in the prefix word group from the multiple word segments included in the i-th round prompt word. The multiplexing unit is configured to retrieve the key-value cache corresponding to the prefix segmentation group from the central processing unit if it fails to retrieve the key-value cache corresponding to the prefix segmentation group from the graphics processor. The sending unit is used to send the key-value cache of the prefix segmentation group and the remaining segmentation group to the inference engine, so that the inference engine can obtain the response content for the i-th round prompt word based on the key-value cache of the prefix segmentation group and the remaining segmentation group.
13. A computer device, characterized in that, The computer device includes a processor and memory: The memory is used to store computer programs and to transfer the computer programs to the processor; The processor is configured to perform the method according to any one of claims 1-11 according to the computer program.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program for performing the method according to any one of claims 1-11.
15. A computer program product comprising a computer program, characterized in that, When it is run on a computer device, it causes the computer device to perform the method described in any one of claims 1-11.