Big language model reasoning method and device, electronic equipment and medium
By dividing logical blocks for inference requests of large language models and calculating summary values to optimize resource utilization, the problems of inference request complexity and resource waste of large language models are solved, and more efficient resource utilization and inference performance are achieved.
Patent Information
- Application Number
- CN202510289609.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
The large language model has high complexity and processing cost when processing inference requests, and chatbots based on the large language model have redundant calculations when inference, wasting computing and storage resources.
By receiving inference requests, the input sequence is divided into logical blocks, and the summary value is calculated for each logical block, and whether there is the same summary value of the inferenced logical block exists, and the physical block is allocated according to the judgment results to optimize resource utilization.
Efficiently allocate storage resources, reduce redundant computing, save computing and storage resources, improve inference speed and system throughput, and reduce the processing cost of inference requests.
Smart Images

Figure CN120218241A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, and particularly to an inference method, apparatus, electronic device, and medium for a large language model. Background Art
[0002] With the rapid development of artificial intelligence technology, large language models (LLMs), as a major breakthrough in the field of natural language processing, are gradually penetrating into multiple application fields such as programming assistance, intelligent customer service, chatbots, and personalized content generation, significantly improving work efficiency and life convenience.
[0003] However, the complexity of processing inference requests by large language models is much higher than that of traditional keyword-based retrieval systems, and their processing costs can be several times or even dozens of times higher. Distributed inference technology distributes the inference tasks of large language models to multiple computing nodes for parallel processing, effectively improving the inference speed and throughput. However, how to further optimize resource utilization and reduce the processing cost of a single inference request remains an urgent problem to be solved.
[0004] On the other hand, many chatbots based on large language models in related technologies have a history record function, but each inference requires the input prompt words of the user and the output of the chatbot to be inferred together to play the role of attaching the history record. This results in a lot of redundant calculations, wasting computing resources and storage resources. Summary of the Invention
[0005] In view of this, the present disclosure provides an inference method, apparatus, electronic device, and medium for a large language model, which at least partially solves the problems existing in the prior art.
[0006] In a first aspect, an embodiment of the present disclosure provides an inference method for a large language model, which includes:
[0007] Receiving an inference request, and dividing the input sequence of the inference request into one or more logical blocks according to a fixed number of tokens, where the input sequence is the token sequence of the prompt words of the inference request;
[0008] For each logical block:
[0009] Calculating a summary value of the logical block as a first summary value based on the summary value calculation sequence of the logical block, where the summary value calculation sequence is the token sequence from the starting token of the input sequence to the ending token of the logical block;
[0010] Determine whether there is a second digest value that is the same as the first digest value, where the second digest value is the digest value of an inferred logical block in which a KV cache is stored in a physical block, and the physical block is a storage space for the KV cache; and
[0011] Allocate a physical block for the logical block according to the result of the determination; and
[0012] Cause the large language model to perform inference on the inference request.
[0013] According to an embodiment of the present disclosure, allocating a physical block for the logical block according to the result of the determination includes:[[]]END]]
[0014] In the case where it is determined that there is a second digest value that is the same as the first digest value, record the physical block corresponding to the logical block in the block table for recording the mapping between the logical block and the physical block as the physical block of the inferred logical block, and
[0015] In the case where it is determined that there is no second digest value that is the same as the first digest value, allocate an unused physical block for the logical block and record it in the block table.
[0016] According to an embodiment of the present disclosure, in the case where it is determined that there is a second digest value that is the same as the first digest value, in the inference, through the block table, use the KV cache of the inferred logical block as the KV cache of the logical block, and
[0017] In the case where it is determined that there is no second digest value that is the same as the first digest value, in the inference, calculate a KV cache for the logical block and store it in the physical block allocated to the logical block.
[0018] According to an embodiment of the present disclosure, the digest value of the inferred logical block is recorded in the block manager in association with the physical block storing its KV cache, and
[0019] Judge whether there is a second digest value that is the same as the first digest value by retrieving the first digest value in the block manager.
[0020] According to an embodiment of the present disclosure, in the case where it is determined that there is no second digest value that is the same as the first digest value, record the first digest value in the block manager in association with the physical block allocated to the logical block.
[0021] According to an embodiment of the present disclosure, calculate the digest value of the digest value calculation sequence of the logical block using the SHA1 algorithm as the first digest value.
[0022] In a second aspect, the present disclosure provides an inference device for a large language model, which includes:[[]]END]]
[0023] A logical block division unit, configured to receive an inference request and divide an input sequence of the inference request into one or more logical blocks according to a fixed number of tokens, where the input sequence is a token sequence of prompt words of the inference request;
[0024] A summary value calculation unit, configured to calculate a summary value of the logical block as a first summary value based on a summary value calculation sequence of each logical block among the one or more logical blocks, where the summary value calculation sequence is a token sequence from a starting token of the input sequence to an ending token of the logical block;
[0025] A summary value judgment unit, configured to judge whether there is a second summary value identical to the first summary value, where the second summary value is a summary value of an inferred logical block stored with a KV cache in a physical block, and the physical block is a storage space for the KV cache;
[0026] A physical block allocation unit, configured to allocate a physical block for the logical block according to a result of the judgment;
[0027] A large model inference unit, configured to cause the large language model to perform inference on the inference request.
[0028] According to an embodiment of the present disclosure, allocating a physical block for the logical block according to the result of the judgment includes:
[0029] In a case where it is judged that there is a second summary value identical to the first summary value, recording a physical block corresponding to the logical block as a physical block of the inferred logical block in a block table for recording a mapping between a logical block and a physical block, and
[0030] In a case where it is judged that there is no second summary value identical to the first summary value, allocating an unused physical block for the logical block and recording it in the block table.
[0031] In a third aspect, the present disclosure provides an electronic device, and the electronic device includes:
[0032] At least one processor; and
[0033] A memory communicatively connected to the at least one processor, where,
[0034] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute an inference method of a large language model according to an embodiment of the present disclosure.
[0035] Fourthly, the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the inference method of the large language model according to an embodiment of the present disclosure.
[0036] Compared with the related art, the inference method, apparatus, electronic device, and medium of the large language model provided by the present disclosure have the following advantages.
[0037] By calculating the digest value for each logical block of the inference request and determining whether there is a digest value of an inferred logical block that is the same as the digest value and has a KV cache stored in a physical block, and allocating a physical block as the storage space for the KV cache for the logical block according to the determination result, storage resources can be allocated more reasonably and effectively, thereby optimizing resource utilization. In addition, the KV cache of the inferred logical block can be reused, thereby reducing duplicate and redundant calculations, saving computing resources and storage resources, improving the inference speed, increasing the system throughput, improving the processing efficiency of inference requests, and reducing the processing cost of inference requests. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. The elements shown are not limited by the scale shown in the drawings. The same or similar reference numerals in the drawings denote the same or similar elements, where:
[0039] Figure 1 is an exemplary flowchart of the inference method of the large language model provided by an embodiment of the present disclosure;
[0040] Figure 2 is another exemplary flowchart of the inference method of the large language model provided by an embodiment of the present disclosure;
[0041] Figure 3 is an exemplary schematic diagram of the block table provided by an embodiment of the present disclosure;
[0042] Figure 4 is an exemplary structural diagram of the inference apparatus of the large language model provided by an embodiment of the present disclosure;
[0043] Figure 5 shows an exemplary structural diagram of a device capable of implementing the method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the present disclosure will be further described in detail below with reference to the specific embodiments and the drawings. Here, the exemplary embodiments of the present disclosure and their descriptions are used to explain the present disclosure, but not to limit the present disclosure.
[0045] As used herein, the term "including" and its variations mean open-ended inclusion, i.e., "including but not limited to". Unless specifically stated otherwise, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.
[0046] The following describes the embodiments of the present disclosure through specific and particular examples. Those skilled in the art can easily understand the other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.
[0047] It should also be noted that the following describes various aspects of the embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of the aspects described herein can be used to implement the device and / or practice the method. Additionally, this device can be implemented and this method can be practiced using other structures and / or functionality in addition to one or more of the aspects described herein.
[0048] It should further be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present disclosure schematically. The drawings only show the components related to the present disclosure and are not drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0049] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the aspects can be practiced without these specific details.
[0050] To facilitate a better understanding of the present disclosure, a method for managing storage resources during the inference process of a large language model using the paged attention mechanism is briefly introduced first. During the inference process of a large language model, a Key-Value (KV) cache needs to be generated for each token of the prompt of an inference request. In the framework of the paged attention mechanism, each request can be divided into multiple independent "pages" (or "logic blocks"), and the KV cache corresponding to each logic block can be stored, for example, using a fixed-capacity storage space in a Graphics Processing Unit (GPU), and such a storage space can be referred to as a "physical block". Thus, each logic block is associated with a physical block. The mapping (corresponding relationship) between the logic block and the physical block can be recorded and managed, for example, through a block table described in detail later. This design makes the management of storage resources during the inference process both flexible and efficient. As the inference of a request is completed, these physical blocks are recycled and reallocated to new requests, thereby achieving the recycling of resources.
[0051] The present disclosure is made based on the paged attention mechanism, and its core idea is to save the computing resources and storage resources for the inference of a large language model by reusing the physical blocks for the KV cache. The inference method and device of the large language model according to the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0052] The inference method of the large language model provided by the embodiments of the present disclosure can be executed by an electronic device, and in particular, by one or more processors in the electronic device. The electronic device can be a computing device with computing capabilities such as a notebook computer, a desktop computer, a workstation, a server, a Field Programmable Gate Array (FPGA), etc. As Figure 1 and Figure 2 shown, the inference method of the large language model according to the embodiments of the present disclosure can include the following operations, for example.
[0053] First, in Figure 1 step S101, an inference request is received, and the input sequence of the inference request is divided into one or more logic blocks according to a fixed number of tokens, where the input sequence is the token sequence of the prompt of the inference request. This step specifically corresponds to Figure 2 steps S201 and S202 in
[0054] The inference request can be of various types and can come from different user groups and application scenarios. For example, in the Natural Language Generation (NLG) scenario, the inference request can be a request to make the large language model generate a descriptive text, news report, story, poem, email, social media post, etc. In a chatbot or intelligent customer service system, the inference request can be a request to make the large language model generate a response for interacting with a human user. In a Question Answering (QA) system, the inference request can be a request to make the large language model generate an answer to a given question. The inference request can be input into the serving system of the large language model, for example, by the user entering a prompt, so that the serving system can further process the inference request.
[0055] In the inference method of the large language model of the present disclosure, as Figure 2 shown, first in step S201, an inference request is received. For example, a prompt in the form of natural language input by the user is received and tokenized to split it into one or more tokens, so as to obtain an input sequence that is a sequence of tokens of the prompt as the inference request. Next, in step S202, the input sequence of the inference request is divided into logical blocks according to a fixed number of tokens. The fixed number of tokens can be set according to, for example, memory limitations, computational efficiency, or other factors. The present disclosure does not make a specific limitation on the fixed number of tokens, but when giving a specific example, it can be 16, for example. In this case, every 16 tokens of the input sequence of the inference request are divided into a logical block. When the number of tokens in the input sequence itself is less than 16 or the remaining unpartitioned tokens at the end are less than 16, these tokens are also divided into a logical block, which does not affect subsequent calculations. That is to say, "dividing the input sequence of the inference request into logical blocks according to a fixed number of tokens" means that the logical block has at most rather than necessarily includes the fixed number of tokens.
[0056] Next, in Figure 1 step S102, for each logical block, based on the summary value calculation sequence of the logical block, the summary value of the logical block is calculated as the first summary value, where the summary value calculation sequence is the sequence of tokens from the starting token of the input sequence to the ending token of the logical block. This step specifically corresponds to Figure 2 step S203 in
[0057] For specific application scenarios such as chatbots, there are often overlapping parts in the content between consecutive inference requests. For such overlaps, especially when subsequent inferences depend on the results of the previous inference, if all tokens are inferred for each inference request, it may cause unnecessary resource waste. To this end, before passing the input sequence of the inference request to the inference engine of the large language model for inference, the content of the current inference request can be compared with that of the already inferred requests to accurately identify the repeated (i.e., no need to recalculate) parts. This first requires determining the content of the inference request.
[0058] In some embodiments, the content of the logical block and its context can be determined by calculating the digest value of the logical block. Specifically, the digest value (first digest value) of each logical block of the input sequence of the inference request can be calculated based on the digest value calculation sequence that is the sequence of tokens from the starting token of the input sequence to the ending token of the logical block. For example, when each logical block includes 16 tokens, for the first logical block of the input sequence, its digest value calculation sequence is the first to the 16th tokens of the input sequence; for the second logical block of the input sequence, its digest value calculation sequence is the first to the 32nd tokens of the input sequence; for the third logical block of the input sequence, its digest value calculation sequence is the first to the 48th tokens of the input sequence, and so on.
[0059] In some embodiments, the SHA1 (Secure Hash Algorithm 1) algorithm is used to calculate the digest value of the digest value calculation sequence of each logical block of the current input sequence as the first digest value. The SHA1 algorithm is a secure hash algorithm commonly used for encrypting strings and can convert a string of any length into a 20-byte hash value (digest value). The digest values of different strings are different, so the SHA1 algorithm can be used to determine the content of the digest value calculation sequence of the logical block. The reason that the digest value calculation sequence of each logical block includes the current logical block and all its previous logical blocks is that the calculation of the KV cache is related to the logical block and its context. Only when the token sequence of the logical block and its context in the input sequence of the newly received inference request is exactly the same as the token sequence from the starting token to the ending token of a certain logical block in the already inferred inference request, can the KV caches of these two token sequences be exactly the same, so that the newly received inference request can reuse the KV cache of the already inferred inference request.
[0060] When expressed in a formula, the calculation method of the digest value of the logical block can be shown as the following formula (1).
[0061] hash=sha1(seq[0:blockindex*blocksize]) (1)
[0062] Among them, hash represents the digest value of the logical block numbered blockindex (starting from 0), blocksize is the logical block size, i.e., the aforementioned fixed number of tokens, seq[0:blockindex*blocksize] represents the token sequence from the 0th token of the input sequence of the inference request to the last token of the logical block numbered blockindex, and sha1() represents the operation using the SHA1 algorithm.
[0063] The main reason for calculating the digest value using the SHA1 algorithm is that its calculation cost is relatively small and the response is relatively fast. However, it should be understood that the present disclosure does not specifically limit the algorithm for calculating the digest value. For example, the digest value of the logical block can also be calculated using the SHA256 algorithm, etc.
[0064] Next, in Figure 1 step S103, it is determined whether there is a second digest value that is the same as the first digest value, which is the digest value of the logical block that is the input sequence of the current inference request, where the second digest value is the digest value of the inferred logical block stored in the KV cache in the physical block, and the physical block is the storage space for the KV cache. This step specifically corresponds to Figure 2 step S204 in
[0065] In some embodiments, for example, a block manager can be used to manage the digest values of the inferred logical blocks. For example, the digest value of the inferred logical block can be recorded in the block manager in association with the physical block storing the KV cache of the inferred logical block. For example, a digest value table can be maintained in the block manager that records the digest value of the inferred block in association with the number of the physical block storing its KV cache. In this case, in step S204, it can be determined whether there is a second digest value that is the same as the first digest value by retrieving the first digest value in the block manager, for example, retrieving the first digest value in the digest value table. The existence of a second digest value that is the same as the first digest value means that there is a physical block storing the KV cache of an inferred logical block that is the same as the current logical block. Thus, it can be determined whether the KV cache of an inferred logical block that is the same as the current logical block has been stored in the physical block. For a physical block whose KV cache has been released (reset or cleared), the record of its corresponding inferred logical block's digest value can be deleted from the digest value table. As described above, in order to reuse the KV cache, it is necessary to compare the two token sequences of the logical block and its context (the digest value calculation sequence of the logical block) with the inferred logical block and its context (the digest value calculation sequence of the inferred logical block), rather than comparing the logical block with the inferred logical block itself. Therefore, it is not necessary to record the inferred logical block itself. Compared with recording the digest value calculation sequence, recording the digest value of the inferred logical block in the block manager can make the recorded data format consistent and save storage space, so it is more preferable.
[0066] Next, in Figure 1 step S104, according to the judgment result of each logical block in step S103, physical blocks are allocated to each logical block. This step specifically corresponds to Figure 2 steps S205 - S208 in
[0067] When it is judged as "Yes" in step S204, that is, when there is a second digest value identical to the first digest value, in step S205, in the block table for recording the mapping between logical blocks and physical blocks, the physical block corresponding to the current logical block is recorded as the physical block of this inferred logical block.
[0068] In some embodiments, an example of the block table for recording the mapping between logical blocks and physical blocks can be as shown in Figure 3 . In Figure 3 , the numbers of the physical blocks corresponding to each logical block of requests 0 - 3 are shown. For example, in the third row from left to right, the numbers of the physical blocks corresponding to the 4 logical blocks (blocks 0 - 3 in the figure) of request 2 are 96, 97, 98, 99, indicating that the physical block allocated for the KV cache of the first logical block (block 0) of request 2 is physical block 96, the physical block allocated for the KV cache of the second logical block (block 1) of request 2 is physical block 97, the physical block allocated for the KV cache of the third logical block (block 2) of request 2 is physical block 98, and the physical block allocated for the KV cache of the fourth logical block (block 3) of request 2 is physical block 99. The same applies to the other rows in the block table. Physical blocks can be allocated to each logical block before inference and recorded in the block table, or the relevant records in the block table can be deleted after inference. Thus, in the block table, each logical block of each inference request can be stored in association with the number of the physical block allocated to each logical block, enabling the block table to record the mapping between logical blocks and physical blocks. During inference, the physical block stored in the KV cache of each logical block can be found through the block table, so that the KV cache of each logical block can be read from the corresponding physical block. It should be understood that although in Figure 3 physical blocks with consecutive numbers are allocated to some consecutive logical blocks shown, the physical blocks allocated to consecutive logical blocks may not be physical blocks with consecutive numbers. In addition, even for physical blocks with consecutive numbers, their positions in the video memory can be arbitrarily distributed.
[0069] In some embodiments, when using the digest value table in the block manager to correlatively record the digest value of the inferred block and the number of the physical block storing its KV cache, when a second digest value identical to the digest value of the current logical block, i.e., the first digest value, is retrieved in the digest value table, the number of the physical block recorded correlatively with the second digest value is read, and the number of the physical block is recorded as the number of the physical block corresponding to the current logical block in the block table. Thus, instead of allocating a new (unused) physical block for the current logical block, the current logical block reuses the physical block of the inferred logical block, so that in the inference for the inference request, it is possible to avoid calculating the KV cache for the current logical block, but instead read the KV cache of the corresponding inferred logical block from the physical block storing the KV cache that is correlatively recorded with the current logical block through the block table, and use the KV cache of the inferred logical block as the KV cache of the current logical block.
[0070] For example, assume that the current inference request is request 3, and the current logical block is the first logical block of request 3, i.e., block 0. For block 0 of request 3, assume that a second digest value identical to the digest value of block 0 is retrieved in the digest value table of the block manager, and the number of the physical block corresponding to the second digest value is 7. Therefore, in Figure 3 the block table shown, the number of the physical block of block 0 of request 3 is recorded as 7. The fact that a digest value stored correlatively with physical block 7 can be retrieved in the digest value table means that physical block 7 has been allocated to the inferred logical block with that digest value, and the KV cache of the inferred logical block has been stored. Thus, instead of allocating a new physical block for block 0 of request 3, it reuses physical block 7 of the inferred logical block. In the inference for request 3, since block 0 of request 3 can reuse the KV cache stored in physical block 7, there is no need to recalculate the KV cache for block 0 of request 3, and the stored KV cache can be read from physical block 7 through the block table as the KV cache of block 0 of request 3.
[0071] By reusing physical blocks in this way, both storage resources and computing resources can be saved during the inference process, and the inference speed can be increased, enabling a faster response to the inference request.
[0072] When it is determined as "no" in step S204, that is, when it is determined that there is no second digest value identical to the digest value of the current logical block, i.e., the first digest value, in step S206, an unused physical block is allocated for the logical block and recorded in the block table, and in step S207, the digest value of the logical block is recorded correlatively with the physical block allocated to the logical block in the block manager, for example, recorded in the digest value table in the block manager. During the inference, the KV cache for the logical block is calculated and stored in the physical block allocated to the logical block.
[0073] For example, assume that the current inference request is Request 3, and the current logical block is the second logical block of Request 3, i.e., Block 1. For Block 1 of Request 3, assume that no second digest value identical to the digest value of Block 1 is retrieved from the digest value table of the block manager. Therefore, a new physical block 100 is allocated for Block 1, and the number of the physical block of Block 1 of Request 3 is recorded as 100 in the block table shown in Figure 3 In addition, the digest value of Block 1 can be recorded in the digest value table of the block manager in association with the physical block 100, so that when the digest value of the logical block in a subsequent inference request is the same as the digest value of Block 1 of Request 3, the physical block 100 of Block 1 can be reused. During the inference for Request 3, a KV cache is calculated for Block 1 of Request 3 and stored in the physical block 100 allocated for Block 1 through the block table.
[0074] When, in step S205, the current logical block reuses the physical block of an already-inferred logical block, or when, in step S206, a new physical block is allocated for the current logical block and recorded in the block manager in step S207, the process proceeds to step S208 to determine whether the current logical block is the last logical block of the current inference request. If the determination is "no", i.e., the current logical block is not the last logical block of the current inference request, the process returns to step S203 to continue processing the next logical block. If the determination is "yes", i.e., the current logical block is the last logical block of the current inference request, the process proceeds to step S209.
[0075] Next, in Figure 1 step S105, the large language model performs inference for the current inference request. This step specifically corresponds to Figure 2 step S209 in
[0076] In step S209, the large language model performs inference for the current inference request. As mentioned above, during the inference process, for each logical block of the current inference request, if in step S205 the logical block reuses the physical block of an already-inferred logical block, there is no need to calculate the KV cache for this logical block, but instead the KV cache of the already-inferred logical block is reused; if in step S206 a new physical block is allocated for this logical block and recorded in the block manager in step S207, the KV cache is calculated for this logical block and the KV cache is stored in the physical block allocated for this logical block. Thereby, redundant calculations can be reduced, saving storage resources and computing resources.
[0077] In summary, the present disclosure calculates the digest value for each logical block of the inference request, determines whether there is a digest value of an already-inferred logical block stored in a physical block that is the same as the digest value, and allocates a physical block, which is the storage space for the KV cache, to the logical block based on the determination result, enabling a more reasonable and efficient allocation of storage resources, thereby optimizing resource utilization. In addition, the KV cache of the already-inferred logical block can be reused, thereby reducing duplicate and redundant calculations, saving computing resources and storage resources, increasing the inference speed, improving the system throughput, enhancing the processing efficiency of the inference request, and reducing the processing cost of the inference request.
[0078] Embodiments of the present disclosure also provide an inference device for a large language model. As Figure 4 shown, the device includes a logical block division unit 401, a digest value calculation unit 402, a digest value determination unit 403, a physical block allocation unit 404, and a large model inference unit 405.
[0079] The logical block division unit 401 is configured to receive an inference request and divide the input sequence of the inference request into one or more logical blocks according to a fixed number of tokens, where the input sequence is the token sequence of the prompt words of the inference request.
[0080] The digest value calculation unit 402 is configured to calculate the digest value of the logical block as the first digest value based on the digest value calculation sequence of each logical block in one or more logical blocks, where the digest value calculation sequence is the token sequence from the start token of the input sequence to the end token of the logical block.
[0081] The digest value determination unit 403 is configured to determine whether there is a second digest value that is the same as the first digest value, where the second digest value is the digest value of an already-inferred logical block stored in a physical block, and the physical block is the storage space for the KV cache.
[0082] The physical block allocation unit 404 is configured to allocate a physical block to the logical block according to the determination result.
[0083] The large model inference unit 405 is configured to cause the large language model to perform inference on the inference request.
[0084] In some embodiments, in the physical block allocation unit 404, allocating a physical block to the logical block according to the determination result includes: in the case where it is determined that there is a second digest value that is the same as the digest value of the logical block, i.e., the first digest value, recording the physical block corresponding to the logical block as the physical block of the already-inferred logical block in the block table for recording the mapping between the logical block and the physical block, and in the case where it is determined that there is no second digest value that is the same as the first digest value, allocating an unused physical block to the logical block and recording it in the block table.
[0085] Each unit of the inference device of the large language model is used to implement various aspects of the inference method of the large language model. For the specific functions of each unit, reference may be made to the description of the inference method of the large language model above, which will not be elaborated here.
[0086] Embodiments of the present disclosure also provide an electronic device, which includes:
[0087] At least one processor; and
[0088] A memory communicatively connected to the at least one processor, wherein,
[0089] The memory stores instructions executable by the at least one processor. When the instructions are executed by the at least one processor, the at least one processor is enabled to execute the inference method of the large language model according to the embodiments of the present disclosure.
[0090] Embodiments of the present disclosure also provide a non-transitory computer-readable storage medium that stores computer instructions for causing a computer to execute the inference method of the large language model according to the embodiments of the present disclosure.
[0091] Embodiments of the present disclosure also provide a computer program product that includes a computer program. The computer program includes program instructions that, when executed by a computer, cause the computer to execute the inference method of the large language model according to the embodiments of the present disclosure.
[0092] Embodiments of the present disclosure also provide a computer program that includes program instructions that, when executed by a computer, cause the computer to execute the inference method of the large language model according to the embodiments of the present disclosure.
[0093] Figure 5 FIG. shows a schematic diagram of a method that can implement the embodiments of the present disclosure or a device 1000 that implements the embodiments of the present disclosure. In some embodiments, there may be more or fewer devices than shown in the figure. In some embodiments, a single or multiple devices may be used for implementation. In some embodiments, cloud or distributed devices may be used for implementation.
[0094] As Figure 5As shown, device 1000 includes a processor 1001 which can perform various appropriate operations and processes according to programs and / or data stored in a read-only memory (ROM) 1002 or programs and / or data loaded into a random access memory (RAM) 1003 from a storage section 1008. The processor 1001 can be a multi-core processor or can include multiple processors. In some embodiments, the processor 1001 can include a general-purpose main processor and one or more special coprocessors, such as, for example, a central processing unit (CPU), a graphics processing unit (GPU), a neural network processing unit (NPU), a digital signal processing unit (DSP), and so on. In the RAM 1003, various programs and data required for the operation of the device 1000 are also stored. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0095] The above-mentioned processor and memory are jointly used to execute the programs stored in the memory, and when the programs are executed by a computer, they can implement the methods, steps or functions described in the above embodiments.
[0096] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, a touch screen, etc.; an output section 1007 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed so that a computer program read from it can be installed into the storage section 1008 as needed. Figure 5 Only some components are schematically shown, and it does not mean that the device 1000 only includes Figure 5 the components shown.
[0097] The systems, devices, modules or units illustrated in the above embodiments can be implemented by a computer or its associated components. The computer can be, for example, a mobile terminal, a smart phone, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a personal digital assistant, a media player, a navigation device, a game console, a tablet computer, a wearable device, a smart TV, an Internet of Things system, a smart home, an industrial computer, a server, or a combination thereof.
[0098] Although not shown, in an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, an inference method of a large language model according to an embodiment of the present disclosure is implemented.
[0099] The storage medium in the embodiments of the present disclosure includes permanent and non-permanent, removable and non-removable articles that can implement information storage by any method or technology. Examples of storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0100] Although not shown, an embodiment of the present disclosure further provides a computer program product, including: computer program / instructions, and when the computer program / instructions are executed by a processor, an inference method of a large language model according to an embodiment of the present disclosure is implemented.
[0101] The methods, programs, systems, devices, etc. in the embodiments of the present disclosure can be executed or implemented in a single or multiple networked computers, and can also be practiced in a distributed computing environment. In the embodiments of this specification, in these distributed computing environments, tasks can be executed by remote processing devices connected through a communication network.
[0102] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, those skilled in the art can think that the implementation of the functional modules / units or controllers and related method steps clarified in the above embodiments can be achieved in a software, hardware, or a combination of software and hardware manner.
[0103] Unless explicitly stated, the actions or steps of the methods and programs according to the embodiments of the present disclosure do not necessarily have to be executed in a specific order and can still achieve the desired results. In some embodiments, multi-tasking and parallel processing are also possible or may be advantageous.
[0104] In this document, multiple embodiments of the present disclosure are described. For the sake of brevity, the descriptions of the embodiments are not exhaustive, and the same or similar features or parts between the embodiments may be omitted. In this document, "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean applicable to at least one embodiment or example according to the present disclosure, rather than all embodiments. The above terms do not necessarily refer to the same embodiment or example. Without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0105] The exemplary systems and methods of the present disclosure have been specifically shown and described with reference to the above embodiments, which are only examples of the best mode for implementing the systems and methods. Those skilled in the art can understand that various changes can be made to the embodiments of the systems and methods described herein when implementing the systems and / or methods without departing from the spirit and scope of the present disclosure defined in the appended claims.
Claims
1. An inference method for a large language model, characterized in that including: receiving an inference request and dividing an input sequence of the inference request into one or more logical blocks according to a fixed number of tokens, where the input sequence is a token sequence of a prompt of the inference request; for each logical block: calculating a summary value of the logical block as a first summary value based on a summary value calculation sequence of the logical block, where the summary value calculation sequence is a token sequence from a start token of the input sequence to an end token of the logical block; judging whether there is a second summary value identical to the first summary value, where the second summary value is a summary value of an inferred logical block having a KV cache stored in a physical block, and the physical block is a storage space for the KV cache; and allocating a physical block for the logical block according to a result of the judgment; and causing the large language model to perform inference on the inference request.
2. The inference method of the large language model according to claim 1, wherein Allocating a physical block for the logical block according to the result of the judgment includes: in a case where it is judged that there is a second summary value identical to the first summary value, recording, in a block table for recording a mapping between a logical block and a physical block, the physical block corresponding to the logical block as the physical block of the inferred logical block, and in a case where it is judged that there is no second summary value identical to the first summary value, allocating an unused physical block for the logical block and recording it in the block table.
3. The inference method of the large language model according to claim 2, wherein The method further includes: in a case where it is judged that there is a second summary value identical to the first summary value, in the inference, using, through the block table, the KV cache of the inferred logical block as the KV cache of the logical block, and in a case where it is judged that there is no second summary value identical to the first summary value, in the inference, calculating a KV cache for the logical block and storing it in the physical block allocated to the logical block.
4. The inference method of a large language model according to claim 3, wherein the summary value of the inferred logical block is recorded in a block manager in association with the physical block storing its KV cache, and judging whether there is a second summary value identical to the first summary value by retrieving the first summary value in the block manager.
5. The inference method of the large language model according to claim 4, wherein The method further includes: in a case where it is judged that there is no second summary value identical to the first summary value, recording the first summary value in association with the physical block allocated to the logical block in the block manager.
6. The inference method of a large language model according to any one of claims 1 to 5, wherein using the SHA1 algorithm to calculate the summary value of the summary value calculation sequence of the logical block as the first summary value.
7. An inference device for a large language model, characterized in that, including: a logical block division unit configured to receive an inference request and divide an input sequence of the inference request into one or more logical blocks according to a fixed number of tokens, where the input sequence is a token sequence of a prompt of the inference request; A digest value calculation unit, configured to calculate a digest value of the logical block as a first digest value based on a digest value calculation sequence of each logical block in the one or more logical blocks, where the digest value calculation sequence is a token sequence from a start token of the input sequence to an end token of the logical block; A digest value determination unit, configured to determine whether there is a second digest value that is the same as the first digest value, where the second digest value is a digest value of an inferred logical block in which a KV cache is stored in a physical block, and the physical block is a storage space for the KV cache; A physical block allocation unit, configured to allocate a physical block for the logical block according to the result of the determination; A large model inference unit, configured to cause the large language model to perform inference on the inference request.
8. The inference device of the large language model according to claim 7, wherein Allocating a physical block for the logical block according to the result of the determination includes: In the case where it is determined that there is a second digest value that is the same as the first digest value, recording, in a block table for recording a mapping between a logical block and a physical block, the physical block corresponding to the logical block as the physical block of the inferred logical block, and In the case where it is determined that there is no second digest value that is the same as the first digest value, allocating an unused physical block for the logical block and recording it in the block table.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor, where The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the inference method of the large language model according to any one of claims 1 to 6.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the inference method of the large language model according to any one of claims 1 to 6.
Citation Information
Cited By
Method and device for generating reasoning result, storage medium and electronic equipment
CN120930810A
Large model reasoning method and device based on multi-level cache, storage medium and equipment
CN120996190A
Inference method and equipment for large language model
CN121094148A
Inference method and device for large language model
CN121094148B