Big language model reasoning method and device

By constructing a global index tree and scheduling based on similarity matching, the distributed inference system achieves more efficient reuse of key-value cache data in large language models, solving the inefficiency problem caused by random scheduling of inference nodes in existing technologies and improving inference efficiency.

CN121009974APending Publication Date: 2025-11-25HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410773361.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-20
Filing Date
2024-06-14
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

During the inference process of a large language model, because the key-value cache data generated in the previous inference process is stored in different inference nodes, the processor will randomly schedule to a single inference node during inference computation, which will result in the key-value cache data not being reused to the maximum extent and reducing inference efficiency.

Method used

By constructing a global index tree, the distributed inference system queries target inference nodes with high similarity matching based on morpheme numbers, schedules inference tasks to these nodes for computation, and uses relational vector caching to store and reuse key-value cache data, thereby reducing the amount of inference computation.

Benefits of technology

It improves the inference efficiency of large language models by reducing the computational load of distributed inference systems and improving computational efficiency through similarity matching degree scheduling and reuse of key-value cache data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009974A_ABST
    Figure CN121009974A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an inference method and device for a large language model, which are used for improving the inference calculation efficiency of the large language model. The method comprises the steps that an inference task is received, the inference task carries morpheme numbers, the morpheme numbers are used for identifying morphemes to be subjected to inference calculation, and each morpheme comprises one or more sub-words determined based on an input text of a large language model; a global index tree is inquired based on morpheme numbers corresponding to the reasoning tasks, target reasoning nodes are determined, the global index tree comprises sub-trees corresponding to multiple reasoning nodes in a distributed reasoning cluster, and the sub-trees corresponding to the target reasoning nodes are sub-trees with the similarity matching degree with the morpheme numbers larger than a threshold value in the global index tree; the similarity matching degree indicates the number of reusable key value cache data in reasoning calculation. And executing the reasoning task based on the target reasoning node.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to the Chinese Patent Application No. 202410637806.5, filed on May 20, 2024, and entitled “A data processing method, device and other equipment”, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] Embodiments of the present application relate to the field of cloud computing, and in particular to a large language model inference method and device. BACKGROUND

[0003] With the development of machine learning and other related technologies, large language models (LLM) based on machine learning technology have been widely applied. For example, large language models can automatically generate language text or generate replies, and can be used for machine translation, speech recognition, question and answer systems, dialogue generation, and other tasks.

[0004] In the current large language model inference scheme, a computing device implements inference of a large language model through a processor, such as a graphics processing unit (GPU) and a neural processing unit (NPU). In the process of processing the inference of the large language model by the processor, since the key-value cache data generated by the large language model in the inference process of different input texts may be repeated, the processor can reuse the key-value cache data generated in the previous inference process, thereby reducing the inference calculation amount and improving the inference efficiency of the large language model.

[0005] However, in the process of reusing the key-value cache data generated in the previous inference process by the processor to improve the inference efficiency, the processor needs to reuse as much key-value cache data generated in the previous inference process as possible to improve the inference efficiency. However, since the key-value cache data generated in the previous inference process is stored in different inference nodes, the processor randomly schedules the inference task to a single inference node in the inference calculation process. However, this inference node may not be the inference node that matches the most key-value cache data, resulting in that the key-value cache data is not maximally reused, and the inference efficiency of the large language model is reduced. SUMMARY

[0006] Embodiments of the present application provide a large language model inference method for improving the inference efficiency of a large language model. Embodiments of the present application also provide a large language model inference method corresponding to a language model inference device, a computing device, a computing device cluster, a computer readable storage medium, and a computer program product.

[0007] In a first aspect, an embodiment of the present application provides a reasoning method of a large language model. The method can be executed by a distributed reasoning cluster, or by a component of a distributed reasoning system, such as a processor, a chip, or a chip system of the distributed reasoning system, or by a logic module or software that can implement all or part of the functions of the distributed reasoning system. The method provided in the first aspect includes: receiving, by the distributed reasoning system, a reasoning task, the reasoning task being a reasoning task generated based on user input text of the large language model, the reasoning task carrying a morpheme number, the morpheme number being used to identify a morpheme to be reasoned, the morpheme including one or more subwords determined based on the input text of the large language model. The distributed reasoning system queries a global index tree based on the morpheme number corresponding to the reasoning task to determine a target reasoning node, the global index tree including sub-trees corresponding to a plurality of reasoning nodes in the distributed reasoning cluster, wherein the sub-tree corresponding to the target reasoning node is a sub-tree in the global index tree that has a similarity matching degree greater than a threshold value with the morpheme number, for example, the sub-tree corresponding to the target reasoning node is a sub-tree with the highest similarity matching degree, and the similarity matching degree can be used to indicate the amount of key-value cache data that can be reused in reasoning calculation. The distributed reasoning system schedules the reasoning task to the target reasoning node and executes the reasoning task based on the target reasoning node.

[0008] In the embodiment of the present application, the distributed reasoning system can construct a global index tree, and in the scheduling process of the reasoning task, the reasoning scheduling system can calculate the similarity matching degree of the morpheme number in the global index tree based on the morpheme number in the reasoning task, so as to schedule the reasoning task to the target reasoning node with high similarity matching degree for execution. Compared with the random scheduling technology based on the load condition in the prior art, in the embodiment of the present application, the distributed reasoning system can schedule the reasoning task to the target reasoning node that can reuse more key-value cache data for reasoning calculation, thereby reducing the reasoning calculation amount of the distributed reasoning system and improving the reasoning efficiency of the large language model.

[0009] In a possible implementation, in the process of querying the global index tree based on the morpheme number corresponding to the reasoning task, the distributed reasoning system calculates the similarity matching degree of each sub-tree according to the proportion of the matching prefix length of the morpheme number in each sub-tree and the total length of the morpheme number, queries the sub-trees with a similarity matching degree greater than a threshold value based on the similarity matching degree of each sub-tree, and determines the reasoning node corresponding to the sub-tree with a similarity matching degree greater than a threshold value as the target reasoning node, for example, the distributed reasoning system sorts the sub-trees according to the similarity matching degree, and determines the reasoning node corresponding to the sub-tree with the highest similarity matching degree as the target reasoning node.

[0010] In the embodiment of the present application, the distributed reasoning system can calculate the similarity matching degree of each sub-tree according to the proportion of the matching prefix length of the morpheme number in each sub-tree and the total length of the morpheme number, thereby improving the computational realizability of the similarity matching degree.

[0011] In one possible implementation, the similarity matching degree is positively correlated with the amount of reusable key-value cache data during the reasoning process. That is, the higher the similarity matching degree between the morpheme number and the reasoning node in the reasoning task, the higher the amount of reusable key-value cache data in the reasoning node, and thus the greater the reduction in reasoning computation.

[0012] In this embodiment of the application, the higher the similarity matching degree between the morpheme number and the inference node in the distributed inference cluster computing inference task, the higher the amount of reusable key-value cache data in the inference node. The distributed inference system performs inference computing of large language models based on similarity matching degree, thereby improving the inference efficiency of large language models.

[0013] In one possible implementation, before the distributed inference system queries the global index tree based on the morpheme number corresponding to the inference task, when the distributed inference system processes historical input text, if the inference node generates key-value cache data, the distributed inference system can create a global index tree based on the key data of these key-value cache data, wherein the key data includes the hash value corresponding to the input text.

[0014] In this embodiment of the application, before the distributed inference system queries the global index tree based on the morpheme number corresponding to the inference task, it can create a global index tree based on the key data of the historical key-value cache data. Since the global index tree only contains key data, the query efficiency of the global index tree is improved.

[0015] In one possible implementation, a relational vector cache (RTC) is deployed in the inference node. An inference node with a relational vector cache can also be called an RTC node. The relational vector cache is used to store key-value cache data generated during the inference process of a large language model.

[0016] In this application, a relational vector cache (RTC) is deployed in the inference node. The relational vector cache can store key-value cache data generated during the inference process of a large language model, and the stored key-value cache data can be reused in subsequent inference processes, thereby improving the inference efficiency of the large language model.

[0017] In one possible implementation, during the process of the distributed inference system executing an inference task based on the target inference node, the distributed inference cluster determines the key-value cache data in the target inference node that matches the inference task. That is, the distributed inference system performs precise matching of the key-value cache data and reuses the value data of the key-value cache data in the target inference node that matches the inference task to perform inference calculations. The value data includes the hidden layer output data of the large language model.

[0018] In the embodiments of this application, the distributed inference system can reuse the value data of the key-value cache data stored in the target inference node to perform inference calculations during the execution of inference tasks, thereby reducing the amount of inference calculations and improving the inference efficiency of large language models.

[0019] In one possible implementation, during the process of the distributed inference system determining the key-value cache data in the target inference node that matches the inference task, the distributed inference system matches the morpheme number in the inference task with the key-value cache data in the target inference node to determine the key-value cache data that matches the inference task. The key-value cache data that matches the inference task is the key-value cache data that can be reused during the inference process.

[0020] In this embodiment, the distributed inference system matches the key-value cache data in the target inference node with the morpheme number in the inference task, determines the matching key-value cache data, and reuses the key-value cache data for inference calculation, thereby reducing the inference calculation amount of the distributed cluster and improving the inference efficiency of the large language model.

[0021] In one possible implementation, before receiving inference tasks, the distributed inference system receives input text, which includes words or sentences input by the user into the large language model, such as questions asked by the user. The distributed inference system first performs word segmentation on the input text to determine one or more morphemes contained in the input text. Then, the distributed inference system performs inference computation based on one or more morphemes to generate one or more inference tasks. The one or more inference tasks are scheduled to be executed by multiple inference nodes in the distributed inference cluster.

[0022] In this embodiment of the application, before the distributed inference system receives the inference task, the distributed inference system performs word segmentation on the input text of the large language model to determine one or more morphemes, and performs inference calculations based on one or more morphemes to generate the inference task, thereby improving the feasibility of generating the inference task.

[0023] In one possible implementation, the inference node includes one or more of the following: a graphics processing unit (GPU), a neural network processing unit (NPU), and a data processing unit (DPU). The GPU, NPU, and DPU can be inference nodes deployed in a single computing device or inference nodes deployed in multiple different computing devices; no specific limitation is imposed.

[0024] In the embodiments of this application, the inference node can be a graphics processing unit, a neural network processing unit, and a data processing unit deployed on one or more computing devices. The various implementation methods of the inference node enhance the richness of the solution.

[0025] In one possible implementation, during the construction of the global index tree, the inference scheduling system can build subtrees of the global index tree according to the index range of the key data in the key-value cache data. After building the subtrees of the global index tree according to the index range, when the inference scheduling system receives an inference task, it can determine the subtree corresponding to the index range based on the morpheme number of the inference task, and determine the inference node corresponding to that subtree as the target inference node. Then, the inference scheduling system schedules the inference task to the target inference node.

[0026] In this embodiment, the inference scheduling system can construct a subtree of the global index tree based on the index range of the key-value cache data, thereby enabling the inference scheduling system to determine the target inference node based on the index range to which the morpheme number of the inference task belongs, improving the query efficiency of morpheme numbers, and further improving the inference efficiency of large language models.

[0027] Secondly, embodiments of this application provide an inference apparatus for a large language model, comprising a transceiver unit and a processing unit. The transceiver unit receives an inference task, which carries a morpheme ID. The morpheme ID identifies the morpheme to be inferred, and the morpheme includes one or more sub-words determined based on the input text of the large language model. The processing unit queries a global index tree based on the morpheme ID corresponding to the inference task to determine a target inference node. The global index tree includes subtrees corresponding to multiple inference nodes in a distributed inference cluster. The subtree corresponding to the target inference node is a subtree in the global index tree whose similarity matching degree with the morpheme ID is greater than a threshold. The similarity matching degree indicates the amount of reusable key-value cache data during the inference process. The processing unit also executes the inference task based on the target inference node.

[0028] In one possible implementation, the processing unit is specifically used to calculate the similarity matching degree of each subtree based on the matching prefix length of the morpheme number in each subtree and the total length of the morpheme number, query the subtrees with a similarity matching degree greater than a threshold based on the similarity matching degree of each subtree, and determine the inference node corresponding to the subtree with a similarity matching degree greater than the threshold as the target inference node.

[0029] In one possible implementation, the similarity matching degree is positively correlated with the amount of key-value cache data that can be reused during the inference process.

[0030] In one possible implementation, the processing unit is further configured to create a global index tree based on the key data of the key-value cache data when the inference node generates key-value cache data, wherein the key data includes the hash value corresponding to the input text.

[0031] In one possible implementation, a relational vector cache (RTC) is deployed in the inference node. The relational vector cache is used to store key-value cache data generated during the inference process of a large language model.

[0032] In one possible implementation, the processing unit is specifically used to determine the key-value cache data in the target inference node that matches the inference task, and reuse the key-value cache data in the target inference node that matches the inference task to perform inference calculations. The value data includes the hidden layer output of the large language model.

[0033] In one possible implementation, the processing unit is specifically used to match the morpheme numbers in the reasoning task with the key-value cache data in the target reasoning node, determine the matching key-value cache data, and the matching key-value cache data is the key-value cache data that can be reused in the reasoning process.

[0034] In one possible implementation, the transceiver unit is further configured to receive input text, which includes words or sentences input by the user into the large language model. The processing unit is further configured to perform word segmentation on the input text to determine one or more morphemes contained within it. The processing unit is also configured to perform inference computation based on the one or more morphemes, generating one or more inference tasks, which are then scheduled to be executed on one or more inference nodes in a distributed inference cluster.

[0035] In one possible implementation, the inference node includes one or more of the following: a graphics processing unit (GPU), a neural network processing unit (NPU), and a data processing unit (DPU).

[0036] Thirdly, embodiments of this application provide a computing device including a processor coupled to a memory. The processor stores instructions, which, when executed by the processor, cause the computing device to perform the method described in the first aspect or any possible implementation thereof.

[0037] Fourthly, embodiments of this application provide a computing device cluster, which includes one or more computing devices. Each computing device includes a processor coupled to a memory. The processor is used to store instructions, which, when executed by the processor, cause the computing device cluster to perform the method described in the first aspect or any possible implementation thereof.

[0038] Fifthly, embodiments of this application provide a computer-readable storage medium having instructions stored thereon, which, when executed, cause a computer to perform the method described in the first aspect or any possible implementation thereof.

[0039] Sixthly, embodiments of this application provide a computer program product including instructions that, when executed, cause a computer to implement the method described in the first aspect or any possible implementation thereof.

[0040] It is understood that the beneficial effects achievable by any of the large language model reasoning devices, computing devices, computing device clusters, computer-readable media, or computer program products provided above can be referred to the beneficial effects in the corresponding methods, and will not be repeated here. Attached Figure Description

[0041] Figure 1 A schematic diagram of the system architecture of a large language model inference system provided in this application embodiment;

[0042] Figure 2 A flowchart illustrating a reasoning method for a large language model provided in an embodiment of this application;

[0043] Figure 3 A flowchart illustrating another reasoning method for a large language model provided in this application embodiment;

[0044] Figure 4 A flowchart illustrating another reasoning method for a large language model provided in this application embodiment;

[0045] Figure 5 A flowchart illustrating another reasoning method for a large language model provided in this application embodiment;

[0046] Figure 6 A schematic diagram of the structure of a reasoning device for a large language model provided in an embodiment of this application;

[0047] Figure 7 A schematic diagram of the structure of a computer device provided in an embodiment of this application;

[0048] Figure 8 This application provides a schematic diagram of the structure of a computer device cluster.

[0049] Figure 9 This is a schematic diagram of another computer device cluster structure provided in an embodiment of this application. Detailed Implementation

[0050] This application provides a reasoning method and apparatus for large language models, which can improve the reasoning scheduling efficiency of large language models.

[0051] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0052] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0053] First, some of the terms used in the embodiments of this application are introduced to facilitate understanding of the technical solutions by those skilled in the art.

[0054] Large language models (LLMs) are natural language processing models based on deep learning techniques. Trained on large amounts of text data, they possess the ability to understand, generate, and process natural language. LLMs treat natural language text as sequential data, such as sequences of words or characters, and use deep learning models to model the statistical patterns and latent semantic information of these sequences. LLMs can perform a wide range of tasks, including text summarization, translation, sentiment analysis, question answering, and dialogue.

[0055] A morpheme is the smallest semantic unit in a language after word segmentation; it can also be called a segmentation marker or lexical tag. A morpheme can be a single Chinese character or multiple Chinese characters, or a single English word or half an English word. For example, "apple" is a morpheme, and "apple tree" can be two morphemes: "apple" + "tree," but "apple" alone does not have any meaning.

[0056] Relational tensor cache (RTC) is a mechanism used in large language models to manage and store relational vectors. In large neural networks, caches are typically used to store intermediate results or key information to improve computational efficiency and save memory. The role of RTC is to store frequently used relational vectors during model training, thereby avoiding redundant computation. Relational vectors usually represent relationships between entities or associations between contexts; for example, in natural language processing tasks, they represent semantic or syntactic relationships between words.

[0057] Context refers to the contextual information that a large language model relies on when processing text; that is, the context in which the large language model generates or understands text. For a large language model, context can be a piece of text, a dialogue, a question, or any other form of text input.

[0058] To make the technical solution of this application clearer and easier to understand, the system architecture of this application will be described below with reference to the accompanying drawings.

[0059] Please see Figure 1 , Figure 1 This is a schematic diagram of the system architecture of a large language model inference system provided as an example in this application. Figure 1 In the system architecture shown, the large language model inference system 10 includes a global scheduling layer 101 and a relational vector cache layer 102. The global scheduling layer 101 includes a global scheduler 1011, and the relational vector cache layer 102 includes one or more inference nodes 1021. Multiple inference nodes 1021 can form a distributed inference cluster. The specific functions of each part of the large language model inference system 10 are described below.

[0060] The global scheduling layer 101 is used to schedule inference tasks for the large language model. For example, the global scheduler 1011 of the global scheduling layer 101 schedules the inference tasks of the large language model to different inference nodes 1021 in the relational vector cache layer 102 for inference computation, thereby achieving load balancing for the inference computation of the large language model. In this embodiment, the global scheduling layer 101 can also create and store a global index tree. The global index tree can match the inference node 1021 that executes the inference task, and the matched inference node is called the target inference node.

[0061] The global index tree created in the global scheduling layer 101 includes subtrees corresponding to each inference node 1021. Each subtree stores the key data from the key-value cache (KVcache) data generated by the inference node 1021 during the execution of historical inference tasks. The global scheduling layer 101 can determine which inference node 1021 corresponding to which subtree will execute the inference task based on the similarity matching degree between the morpheme number in the inference task and the key data in each subtree, thereby realizing the scheduling of inference tasks.

[0062] The relational vector cache layer 102 is used to store key-value cache data generated by the distributed inference cluster when executing historical inference tasks of the large language model. The relational vector cache layer 102 includes one or more inference nodes 1021, which can also be called relational vector cache nodes. Each inference node 1021 stores key-value cache data generated by its respective inference node when executing historical inference tasks. One or more inference nodes 1021 form a distributed inference cluster, in which a large language model is deployed. Therefore, the inference system 10 of the large language model can also be called a distributed inference system. The large language model can perform inference calculations on user input text to generate corresponding output text.

[0063] For example, a distributed inference cluster can receive a natural language text as input, which can be a sentence, a paragraph, or a longer text sequence. The distributed inference cluster performs inference calculations on the input text and reference information based on a large language model, generating output text. The output text can be the answer to a question corresponding to the input text, a text summary, or a translation result, etc.

[0064] The model structure for deploying large language models in a distributed inference cluster is a multi-layered structure, including, for example, an embedding layer, a self-attention layer, and a feedforward network layer. The embedding layer transforms each morpheme in the user's input text and reference information into a vector representation. The self-attention layer determines the contextual relationships of the input text; for example, it performs self-attention calculations on the vector representations of the input text to generate the corresponding contextual representations. The feedforward network layer performs non-linear transformations after the self-attention layer to extract high-level semantic features from the input text, generating feature representations.

[0065] The inference phase of a large language model includes a prefill phase and a decoding phase. The prefill phase, also known as the first-character inference phase or the full-text inference phase, involves the computing device inputting the complete input text into the large language model at once for the first round of inference, generating the first output character of the output text. The decoding phase, also known as the incremental inference phase, involves the computing device iteratively inferring and calculating the next output character based on an autoregressive loop.

[0066] It should be noted that the inference node 1021 in the above-mentioned distributed inference cluster can be a different graphics processing unit (GPU), neural network processing unit (NPU), or data processing unit (DPU) of a single device, or it can be a GPU, NPU, or DPU distributed in different devices. There is no specific limitation.

[0067] based on Figure 1 The present application also provides a reasoning method for a large language model, in addition to the large language model reasoning system 10 shown. The reasoning method for a large language model provided in this application will be described below with reference to embodiments.

[0068] Please see Figure 2 , Figure 2 This is a flowchart illustrating a scheduling method for a large language model provided in an embodiment of this application. Figure 2 In the example shown, the method includes the following steps:

[0069] Step 201. The reasoning system receives the reasoning task, which carries a morpheme number. The morpheme number is used to identify the morpheme to be reasoned and calculated.

[0070] In the process of processing the user's input text, the reasoning scheduling system 10 of the large language model receives a reasoning task. This reasoning task is a reasoning task generated based on the input text of the large language model. The reasoning task carries a morpheme number, which is used to identify the morpheme to be reasoned. The morpheme number can also be called a morpheme token ID. The morpheme includes one or more sub-words determined based on the input text of the large language model.

[0071] In one possible implementation, before receiving inference tasks, the inference system 10 receives input text from a large language model, then performs word segmentation on the input text to determine one or more morphemes contained within it. The input text includes words or sentences input by the user into the large language model, and the morphemes include sub-words after word segmentation. The inference system 10 performs inference computation based on the segmented morphemes, generating one or more inference tasks. These tasks are scheduled to be executed on one or more inference nodes 1021 in the distributed inference cluster 10. Each inference task includes a morpheme number, which identifies the morpheme to be inferred.

[0072] Specifically, during the process of generating inference tasks based on one or more morphemes, the inference system 10 first encodes the segmented morphemes to obtain their corresponding morpheme numbers, which are the digitized morphemes. The inference system 10 then performs inference calculations based on these morpheme numbers, generating multiple inference tasks. These tasks are then scheduled to be executed on one or more inference nodes 1021 within the distributed inference cluster 10. The inference calculations include embedding layer calculations, self-attention layer calculations, and feedforward network layer calculations.

[0073] For example, in one instance of step 201, the user inputs the text "How's the weather today?" into the large language model. The inference system 10 performs word segmentation on this input text, determining that the morphemes contained in the input text include "today", "weather", and "how". The inference system 10 further encodes the morphemes to obtain morpheme numbers; for example, the morpheme number corresponding to "today" is 100, the morpheme number corresponding to "weather" is 200, and the morpheme number corresponding to "how" is 300. The inference task generated by the inference system 10 includes the aforementioned morpheme numbers.

[0074] It should be noted that, in addition to carrying the morpheme numbers of the morphemes in the user's current input text, the reasoning task can also include the morpheme numbers corresponding to the context. The morpheme numbers corresponding to the context can also be called the context sequence, and the context includes the user's historical input text.

[0075] Step 202. The inference system queries the global index tree based on the morpheme number corresponding to the inference task to determine the target inference node. The global index tree includes subtrees corresponding to multiple inference nodes in the distributed inference cluster. The subtree corresponding to the target inference node is the subtree in the global index tree with a similarity matching degree greater than a threshold with the morpheme number. The similarity matching degree is used to indicate the number of reusable key-value cache data in the inference calculation.

[0076] During the inference task scheduling process of the inference system 10 to multiple inference nodes 1021 in the distributed inference cluster, the inference system 10 queries the global index tree based on the morpheme number corresponding to the inference task, and determines the target inference node based on the query result of the global index tree. The global index tree includes subtrees corresponding to multiple inference nodes in the distributed inference cluster. These subtrees are also called relation vector cache (RTC) subtrees. The subtree corresponding to the target inference node is the subtree in the global index tree whose similarity matching degree with the morpheme number is greater than a threshold. The similarity matching degree is used to indicate the number of reusable key-value cache data in the inference computation. The query result of the global index tree is the similarity matching degree of each subtree in the global index tree.

[0077] Please see Figure 3 , Figure 3 This is a schematic diagram illustrating another reasoning method for a large language model provided in an embodiment of this application. Figure 3 In the example shown, the global index tree includes subtrees corresponding to multiple inference nodes in the distributed inference cluster. The inference system 10 needs to schedule inference tasks to be executed on the inference nodes corresponding to the subtrees based on this global index tree. For example, Figure 3 The global index tree shown contains four subtrees corresponding to inference nodes, namely subtree 1, subtree 2, subtree 3 and subtree 4. Each subtree corresponds to one inference node. The global index tree stores the key data in the key-value cache data. The key data includes the morpheme numbers of morphemes in the user's historical input text.

[0078] exist Figure 3 In the example shown, the inference system 10 can divide subtrees using index ranges. For example, the global index tree has an index range of 1-10k, where the index range of 1-2k is a subtree, namely subtree 1, the index range of 2-4k is a subtree, namely subtree 2, the index range of 4-6k is a subtree, namely subtree 3, and the index range of 8-10k is a subtree, namely subtree 4.

[0079] In one possible implementation, when calculating the similarity matching degree between each subtree and the morpheme IDs in the inference process, the inference system 10 calculates the similarity matching degree of each subtree based on the matching prefix length of the morpheme IDs in each subtree and the total length of the morpheme IDs. Based on the similarity matching degree of each subtree, the inference system 10 queries subtrees with a similarity matching degree greater than a threshold and determines the inference node corresponding to the subtree with a similarity matching degree greater than the threshold as the target inference node. For example, the inference system 10 uses the inference node corresponding to the subtree with the highest similarity matching degree as the target inference node.

[0080] In this embodiment of the application, when the reasoning system 10 calculates the similarity matching degree between the morpheme number and each subtree based on the global index tree, the similarity matching degree satisfies the following formula:

[0081] Similarity matching degree = L1 / L2;

[0082] Where L1 represents the matching length of each subtree of the morpheme number in the global index tree, and L2 represents the total length of the morpheme number. The similarity matching degree ranges from [0, 100%].

[0083] exist Figure 3 In the example shown, during the process of querying the global index tree, the inference system 10 calculates the similarity matching degree between the morpheme number and each subtree in the global index tree, and determines the subtree corresponding to the target inference node according to the similarity matching degree of each subtree.

[0084] For example, if the reasoning system 10 calculates that the similarity matching degree between the morpheme number and subtree 1 is 81%, the similarity matching degree between the morpheme number and subtree 2 is 79%, the similarity matching degree between the morpheme number and subtree 3 is 83%, and the similarity matching degree between the morpheme number and subtree 4 is 70%, then the reasoning system 10 determines that the reasoning node corresponding to subtree 3 is the target reasoning node.

[0085] In one possible implementation, before querying the global index tree based on the morpheme number corresponding to the reasoning task, the reasoning system 10 needs to build the global index tree first. That is, after the reasoning node 1021 generates key-value cache data by performing historical reasoning calculations, the reasoning system 10 creates a global index tree based on the key data of the key-value cache data of each reasoning node. The key data includes the hash value corresponding to the input text.

[0086] It is understandable that, since the inference system 10 creates a global index tree based on the key data in the historical key-value cache data of each inference node, that is, the global index tree only stores the key data of the key-value cache data of each inference node, and does not contain the value data in the key-value cache data, the global index tree can achieve fast matching of morpheme numbers. The global index tree can also be called a global affinity index tree.

[0087] In another possible implementation of this application, during the process of constructing the global index tree, the inference system 10 can construct a subtree of the global index tree according to the index range of the key data in the key-value cache data. After the inference system 10 constructs the subtree of the global index tree according to the index range, when the inference system 10 receives an inference task, the inference system 10 can also determine the corresponding subtree based on the index range to which the morpheme number carried in the inference task belongs, and schedule the inference task to the inference node corresponding to that subtree.

[0088] For example, in Figure 3In the example shown, the global index tree has an index range of 1-10k, where the index range of 1-2k is a subtree (subtree 1), the index range of 2-4k is a subtree (subtree 2), the index range of 4-6k is a subtree (subtree 3), and the index range of 8-10k is a subtree (subtree 4). If the morpheme number carried in the inference task received by the inference system 10 is 3004, then based on the index range to which the morpheme number belongs, the inference system 10 determines the corresponding subtree as subtree 2, and schedules the inference task to the inference node corresponding to subtree 2.

[0089] Understandably, after the inference system 10 constructs a subtree of the global index tree based on the index range, the inference system 10 determines the subtree based on the index range to which the morpheme number of the inference task belongs. At this time, the inference system 10 does not need to perform prefix matching and calculate similarity matching degree based on the morpheme number, thereby improving the query efficiency of the morpheme number.

[0090] In one possible implementation, a relational vector cache RTC is deployed in the inference node 1021. The inference node 1021 can also be called a relational vector cache node. The relational vector cache is used to store key-value cache data generated during the inference process of a large language model.

[0091] Step 203. The inference system executes the inference task based on the target inference node.

[0092] After determining the target inference node, the inference system 10 schedules the inference task to the target inference node, which then executes the inference task. During the execution of the inference task based on the target inference node, the inference system 10 identifies the key-value cache data in the target inference node that matches the inference task, and reuses the value data of the key-value cache data stored in the target inference node that matches the inference task to perform inference computation. The value data includes the hidden layer output of the large language model.

[0093] In this embodiment, the inference system 10 performs an inference task, including a pre-filling stage and a decoding stage. The pre-filling stage is also called the full inference stage and the incremental stage, and the decoding stage is also called the incremental inference stage. In the pre-filling stage, the computing device inputs the complete input text into the large language model at once for the first round of inference, generating the first output character of the output text. In the decoding stage, the computing device can continuously generate the next output character from the already generated output text based on an autoregressive loop iteration.

[0094] Please see Figure 4 , Figure 4 This is a schematic flowchart illustrating the execution of a reasoning task, provided as an embodiment of this application. Figure 4In the example shown, after the target inference node receives the inference task, the processing of the inference task goes through a full inference stage and an incremental inference stage. In the full inference stage, the target inference node can perform embedding layer calculation on the morpheme numbers in the inference task to obtain the vector representation corresponding to the morpheme number. Then, the target inference node performs self-attention layer calculation on the vector representation to obtain key-value cache data.

[0095] exist Figure 4 In the example shown, the key-value cache data generated by the target inference node during the full inference phase can be reused during the incremental inference phase, thereby reducing the computational load of the incremental inference phase. During the computation of the self-attention layer, the target inference node mines and integrates key information in the vector representation, and performs further computation and optimization through the feedforward network layer, ultimately converting the generated vector representation into output text.

[0096] In one possible implementation, during the process of determining the key-value cache data in the target inference node that matches the inference task, the inference system 10 matches the morpheme number in the inference task with the key-value cache data in the target inference node. This matching is an exact match between the morpheme number and the key-value cache data, thereby determining the matching key-value cache data. The matching key-value cache data is reusable key-value cache data in the inference computation.

[0097] Please see Figure 5 , Figure 5 This is a schematic diagram illustrating another large language model inference scheduling provided in an embodiment of this application. Figure 5 In Figure (a) of the example shown, after the inference system 10 determines the target inference node by querying the global index tree, the target inference node determines reusable key-value cache data based on the morpheme encoding in the inference task. During the process of determining reusable key-value cache data, since the target inference node stores the key-value cache data based on a tree structure, when performing an exact match between the target inference node and the cached key-value cache data based on the morpheme number, it retrieves the reusable key-value cache data based on this tree structure.

[0098] exist Figure 5 In the example diagram (b), for example, when the target inference node processes the historical input text, it infers the output text "I love playing basketball". When the target inference node processes the input text again and obtains the output text "I love China", since "I love playing basketball" and "I love China" have the same morpheme, that is, there is reusable key-value cache data in the inference process corresponding to the two output texts, the target inference node can reuse the key-value cache data corresponding to the previous output text "I love playing basketball" for inference calculation when generating the output text "I love China". The reused key-value cache data is, for example, the calculation output of the self-attention layer corresponding to "I love".

[0099] As can be seen from the above embodiments, in the distributed inference system of this application, similarity matching degree is calculated based on the morpheme number in the inference task during the scheduling process of the inference task, so as to schedule the inference task to the target inference node with high similarity matching degree for execution. This allows the distributed inference system to reuse more key-value cache data for inference calculation, thereby reducing the inference calculation amount of the distributed inference system and improving the inference efficiency of large language models.

[0100] Based on the above method embodiments, this application also provides an inference scheduling device for a large language model. The inference scheduling device for a large language model provided in this application is described in detail below.

[0101] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of a reasoning device for a large language model provided in an embodiment of this application. Figure 6 In the example shown, the large language model inference device 600 is used to implement the various steps performed by the inference system 10 in the above embodiments. The large language model inference device 600 includes a transceiver unit 601 and a processing unit 602.

[0102] The transceiver unit 601 receives inference tasks, which carry morpheme numbers. These morpheme numbers identify the morphemes to be inferred, and each morpheme includes one or more sub-words determined from the input text based on a large language model. The processing unit 602 queries the global index tree based on the morpheme number corresponding to the inference task to determine the target inference node. The global index tree includes subtrees corresponding to multiple inference nodes in the distributed inference cluster. The subtree corresponding to the target inference node is the subtree in the global index tree whose similarity matching degree with the morpheme number is greater than a threshold. The similarity matching degree indicates the amount of reusable key-value cache data in the inference computation. The processing unit 602 also executes the inference task based on the target inference node.

[0103] In one possible implementation, the processing unit 602 is specifically used to calculate the similarity matching degree of each subtree based on the matching prefix length of the morpheme number in each subtree and the total length of the morpheme number, query the subtrees with a similarity matching degree greater than a threshold based on the similarity matching degree of each subtree, and determine the inference node corresponding to the subtree with a similarity matching degree greater than the threshold as the target inference node.

[0104] In one possible implementation, the similarity matching degree is positively correlated with the amount of key-value cache data that can be reused during the inference process.

[0105] In one possible implementation, the processing unit 602 is further configured to create a global index tree based on the key data of the key-value cache data when the inference node generates key-value cache data, wherein the key data includes the hash value corresponding to the input text.

[0106] In one possible implementation, a relational vector cache (RTC) is deployed in the inference node. The relational vector cache is used to store key-value cache data generated during the inference process of a large language model.

[0107] In one possible implementation, the processing unit 602 is specifically used to determine the key-value cache data in the target inference node that matches the inference task, and to reuse the value data of the key-value cache data stored in the target inference node to perform inference calculations. The value data includes the hidden layer output of the large language model.

[0108] In one possible implementation, the processing unit 602 is specifically used to match the morpheme number in the reasoning task with the key-value cache data in the target reasoning node, determine the matching key-value cache data, and the matching key-value cache data is reusable key-value cache data in the reasoning calculation.

[0109] In one possible implementation, the transceiver unit 601 is further configured to receive input text, which includes words or sentences input by the user into the large language model. The processing unit 602 is further configured to perform word segmentation on the input text to determine one or more morphemes contained in the input text. The processing unit 602 is further configured to perform inference computation based on the one or more morphemes to generate one or more inference tasks, which are then scheduled to be executed by multiple inference nodes in a distributed inference cluster.

[0110] In one possible implementation, the inference node includes a graphics processing unit (GPU) and a neural network processing unit (NPU).

[0111] It is understandable that the transceiver unit 601 and the processing unit 602 in the inference scheduling device 600 of the large language model can serve as functional modules. Figure 1 The various modules in the large language model inference scheduling system 10 are mapped to each other, thereby realizing the functions of each module in the large language model inference scheduling system 10.

[0112] It should be understood that the division of units in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, all units in the device can be implemented entirely through software calls from processing elements; all units can be implemented entirely in hardware; or some units can be implemented through software calls from processing elements, and others in hardware. For example, each unit can be a separate processing element, or it can be integrated into a chip within the device. Alternatively, it can be stored as a program in memory, called and executed by a processing element of the device. Moreover, these units can be fully or partially integrated together, or implemented independently. The processing element mentioned here can also be called a processor, which can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above units can be implemented through integrated logic circuits in the processor element or through software calls from processing elements.

[0113] It is worth noting that, for the sake of simplicity, the above method embodiments are described as a series of actions. However, those skilled in the art should know that this application is not limited to the order of the described actions. Furthermore, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by this application.

[0114] Other reasonable combinations of steps that can be conceived by those skilled in the art based on the above description also fall within the scope of protection of this application. Furthermore, those skilled in the art should also be aware that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to this application.

[0115] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Figure 7 As shown, the computing device 700 includes a processor 701, a memory 702, a communication interface 703, and a bus 704. The processor 701, memory 702, and communication interface 703 are coupled via the bus (not shown in the figure). The memory 702 stores instructions. When the instructions in the memory 702 are executed, the computing device 700 performs the method executed by the computing device in the above method embodiment.

[0116] The computing device 700 may be one or more integrated circuits configured to implement the methods described above, such as: one or more application-specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs), or a combination of at least two of these forms of integrated circuits. Furthermore, when the units in the device can be implemented in the form of a processing element scheduler, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor capable of calling programs. Alternatively, these units may be integrated together to implement a system-on-a-chip (SOC).

[0117] The processor 701 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor.

[0118] The memory 702 can be volatile memory or non-volatile memory, or it can include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).

[0119] The memory 702 stores executable program code, and the processor 701 executes the executable program code to implement the functions of the aforementioned units or modules, thereby implementing the inference scheduling method of the large language model described above. That is, the memory 702 stores instructions for executing the inference scheduling method of the large language model described above.

[0120] The communication interface 703 uses transceiver modules, such as, but not limited to, network interface cards and transceivers, to enable communication between the computing device 700 and other devices or communication networks.

[0121] In addition to the data bus, the 704 bus can also include a power bus, a control bus, and a status signal bus. The bus can be a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL) bus, a Cache Coherent Interconnect for Accelerators (CCIX) bus, etc. The bus can be divided into address bus, data bus, and control bus.

[0122] Please see Figure 8 , Figure 8 This is a schematic diagram of a computing device cluster provided in an embodiment of this application. Figure 8 As shown, the computing device cluster 800 includes at least one computing device 700.

[0123] like Figure 8 As shown, the computing device cluster 800 includes at least one computing device 700. The memory 702 of one or more computing devices 700 in the computing device cluster 800 may store the same instructions for executing the inference scheduling method for the large language model described above.

[0124] In some possible implementations, the memory 702 of one or more computing devices 700 in the computing device cluster 800 may also store partial instructions for executing the inference scheduling method of the large language model described above. In other words, a combination of one or more computing devices 700 can jointly execute the instructions for executing the inference scheduling method of the large language model described above.

[0125] It should be noted that the memories 702 in the different computing devices 700 within the computing device cluster 800 can store different instructions, each used to execute a portion of the functions of the inference scheduling device of the aforementioned large language model. That is, the instructions stored in the memories 702 of the different computing devices 700 can implement the functions of one or more modules in the acquisition unit and processing unit.

[0126] In some possible implementations, one or more computing devices 700 in the computing device cluster 800 can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc.

[0127] Please see Figure 9 , Figure 9This is a schematic diagram illustrating the network connection of computer devices in a computer cluster, as provided in an embodiment of this application. Figure 9 As shown, the two computing devices 700A and 700B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device.

[0128] In one possible implementation, the memory in computing device 700A stores instructions for performing the functions of the acquisition unit. Meanwhile, the memory in computing device 700B stores instructions for performing the functions of the processing unit and the display unit.

[0129] It should be understood that Figure 9 The functions of computing device 700A shown can also be performed by multiple computing devices. Similarly, the functions of computing device 700B can also be performed by multiple computing devices.

[0130] In another embodiment of this application, a computer-readable storage medium is also provided, which stores computer-executable instructions. When the processor of the device executes the computer-executable instructions, the device executes the method executed by the inference scheduling system in the above method embodiment.

[0131] In another embodiment of this application, a computer program product is also provided, which includes computer-executable instructions stored in a computer-readable storage medium. When the processor of the device executes the computer-executable instructions, the device executes the method performed by the inference scheduling system in the above method embodiments.

[0132] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0133] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0134] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0135] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0136] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A reasoning method for a large language model, characterized in that, Applied to a distributed inference cluster, the distributed inference cluster comprising multiple inference nodes, the method includes: Receive a reasoning task, the reasoning task carrying a morpheme number, wherein the morpheme number is used to identify the morpheme to be reasoned, the morpheme including one or more subwords determined based on the input text of a large language model; Based on the morpheme number corresponding to the reasoning task, the global index tree is queried to determine the target reasoning node. The global index tree includes subtrees corresponding to multiple reasoning nodes in the distributed reasoning cluster. The subtree corresponding to the target reasoning node is the subtree in the global index tree whose similarity matching degree with the morpheme number is greater than a threshold. The similarity matching degree is used to indicate the number of key-value cache data that can be reused during the reasoning process. The inference task is executed based on the target inference node.

2. The method according to claim 1, characterized in that, The process of querying the global index tree based on the morpheme number corresponding to the reasoning task includes: The similarity matching degree of each subtree is calculated based on the length of the matching prefix in each subtree and the total length of the morpheme numbers; Based on the similarity matching degree of each subtree, query the subtrees with a similarity matching degree greater than a threshold, and determine the inference node corresponding to the subtree with a similarity matching degree greater than the threshold as the target inference node.

3. The method according to claim 1 or 2, characterized in that, The similarity matching degree is positively correlated with the amount of key-value cache data that can be reused during the inference process.

4. The method according to any one of claims 1 to 3, characterized in that, Before querying the global index tree based on the morpheme number corresponding to the reasoning task, the method further includes: When the inference node generates key-value cache data, it creates the global index tree based on the key data of the key-value cache data, wherein the key data includes the hash value corresponding to the input text.

5. The method according to any one of claims 1 to 4, characterized in that, The inference node is equipped with a relational vector cache, which is used to store key-value cache data generated during the inference process of the large language model.

6. The method according to any one of claims 1 to 5, characterized in that, The execution of the inference task based on the target inference node includes: The key-value cache data in the target inference node that matches the inference task is determined, and the value data of the key-value cache data that matches the inference task is reused to perform inference. The value data includes the hidden layer output data of the large language model.

7. The method according to claim 6, characterized in that, The determination of the key-value cache data in the target inference node that matches the inference task includes: The morpheme numbers in the reasoning task are matched with the key-value cache data in the target reasoning node to determine the key-value cache data that matches the reasoning task. The key-value cache data that matches the reasoning task is reusable key-value cache data during the reasoning process.

8. The method according to any one of claims 1 to 7, characterized in that, Before receiving the inference task, the method further includes: Receive the input text, which includes words or sentences input by the user into the large language model; Perform word segmentation on the input text to determine one or more morphemes contained in the input text; Inference is performed based on the one or more morphemes to generate one or more inference tasks, which are then scheduled to be executed on multiple inference nodes in the distributed inference cluster.

9. The method according to any one of claims 1 to 8, characterized in that, The inference node includes one or more of the following: a graphics processing unit (GPU), a neural network processing unit (NPU), and a data processing unit (DPU).

10. A reasoning device for a large language model, characterized in that, The device includes: The transceiver unit is used to receive a reasoning task, the reasoning task carrying a morpheme number, wherein the morpheme number is used to identify the morpheme to be reasoned, and the morpheme includes one or more subwords determined based on the input text of the large language model. The processing unit is used to query the global index tree based on the morpheme number corresponding to the inference task to determine the target inference node. The global index tree includes subtrees corresponding to multiple inference nodes in the distributed inference cluster. The subtree corresponding to the target inference node is the subtree in the global index tree whose similarity matching degree with the morpheme number is greater than a threshold. The similarity matching degree is used to indicate the number of key-value cache data that can be reused during the inference process. The processing unit is also used to execute the inference task based on the target inference node.

11. The apparatus according to claim 10, characterized in that, The processing unit is specifically used for: The similarity matching degree of each subtree is calculated based on the length of the matching prefix in each subtree and the total length of the morpheme numbers; Based on the similarity matching degree of each subtree, query the subtrees with a similarity matching degree greater than a threshold, and determine the inference node corresponding to the subtree with a similarity matching degree greater than the threshold as the target inference node.

12. The apparatus according to claim 10 or 11, characterized in that, The similarity matching degree is positively correlated with the amount of key-value cache data that can be reused during the inference process.

13. The apparatus according to any one of claims 10 to 12, characterized in that, The processing unit is also used for: When the inference node generates key-value cache data, it creates the global index tree based on the key data of the key-value cache data, wherein the key data includes the hash value corresponding to the input text.

14. The apparatus according to any one of claims 10 to 13, characterized in that, The inference node is equipped with a relational vector cache, which is used to store key-value cache data generated during the inference process of the large language model.

15. The apparatus according to any one of claims 10 to 14, characterized in that, The processing unit is specifically used for: The key-value cache data in the target inference node that matches the inference task is determined, and the value data of the key-value cache data that matches the inference task is reused to perform inference. The value data includes the hidden layer output data of the large language model.

16. The apparatus according to claim 15, characterized in that, The processing unit is specifically used for: The morpheme numbers in the reasoning task are matched with the key-value cache data in the target reasoning node to determine the matching key-value cache data, which is reusable key-value cache data in the reasoning computation.

17. The apparatus according to any one of claims 10 to 16, characterized in that, The transceiver unit is also used for: Receive the input text, which includes words or sentences input by the user into the large language model; The processing unit is also used to perform word segmentation on the input text to determine one or more morphemes contained in the input text; The processing unit is also configured to perform inference based on the one or more morphemes, and generate one or more inference tasks, which are scheduled to be executed by multiple inference nodes in the distributed inference cluster.

18. The apparatus according to any one of claims 10 to 17, characterized in that, The inference node includes one or more of the following: a graphics processing unit (GPU), a neural network processing unit (NPU), and a data processing unit.

19. A computing device, characterized in that, The device includes a processor coupled to a memory, the processor storing instructions which, when executed by the processor, cause the electronic device to perform the method of any one of claims 1 to 9.

20. A computing device cluster, characterized in that, The device includes at least one computing device, the computing device including a processor coupled to a memory, the processor being used to store instructions that, when executed by the processor, cause the cluster of computing devices to perform the method of any one of claims 1 to 9.

21. A computer-readable storage medium having instructions stored thereon, characterized in that, When the instructions are executed, they cause the computer to perform the method of any one of claims 1 to 9.

22. A computer program product, the computer program product comprising instructions, characterized in that, When the instructions are executed, they cause the computer to perform the method of any one of claims 1 to 9.

Citation Information

Cited By

  • Large language model back-end implementation system based on reasoning service

    CN121457543A

  • A large language model backend implementation system based on inference service

    CN121457543B