Big language model reasoning method and device

By establishing a shared storage pool to store intermediate data in a large language model, the inference delay problem caused by large calculations is solved, and a more efficient inference process is achieved.

CN120338090APending Publication Date: 2025-07-18HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202410223760.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-18
Filing Date
2024-02-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When the large language model processes the input text and reference information corresponding to different requests, the calculation amount is large, resulting in the delay in inference.

Method used

By establishing a shared storage pool, the pre-filled intermediate data corresponding to different requests is stored, and these intermediate data are multiplexed for the next inference, reducing the amount of calculation and reducing the inference delay.

Benefits of technology

It improves the inference efficiency and accuracy of large language models, reduces the computational volume, and reduces the inference delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338090A_ABST
    Figure CN120338090A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an inference method and device for a large language model, which are used for reducing bandwidth consumption of inference of the large language model. The method comprises the steps that a computing device obtains input data of a large language model running on a plurality of processors, the input data comprises an input text of a user and reference information generated based on the input text, and the reference information comprises domain related knowledge corresponding to the input text and timeliness information corresponding to the input text; historical key value cache data are matched in a shared storage pool based on the input data, key value cache data corresponding to the input data are determined, the historical key value cache data are used for indicating intermediate data generated when a large language model processes different historical input data, and the shared storage pool is used for storing the historical key value cache data generated by a plurality of processors; the different historical input data comprises historical input texts and reference information corresponding to the historical input texts. And generating an output result corresponding to the input data according to the key value cache data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of the Chinese patent application with the application number 202410077095.0 and the application title "A Method, Device and Other Equipment for Data Processing" filed with the National Intellectual Property Administration on January 18, 2024, the entire content of which is incorporated herein by reference. Technical Field

[0002] Embodiments of this application relate to the storage field, and in particular to a reasoning method and device for a large language model. Background Art

[0003] With the development of related technologies such as machine learning, large language models (LLMs) based on machine learning technologies have been widely applied. For example, large language models can automatically generate language texts or generate responses, and can be used for tasks such as machine translation, speech recognition, question-and-answer systems, and dialogue generation.

[0004] In the current reasoning process of large language models, the computing device first needs to preprocess the input text and input the preprocessed data into the large language model. During the preprocessing process, the computing device can perform retrieval enhancement on the input text based on prompt engineering to obtain some reference information corresponding to the input text, and the reference information obtained by this retrieval enhancement contains more timely information. The computing device inputs these input texts and the corresponding reference information into the large language model, thereby being able to improve the answering quality of the large language model.

[0005] However, when the large language model processes these input texts and reference information corresponding to different requests, since these input texts and reference information will make the request sequence longer, it will lead to a large amount of computation in the Prefill stage of the large language model and an extended inference time. Summary of the Invention

[0006] Embodiments of this application provide a reasoning method for a large language model. The computing device stores the intermediate data obtained through prefill calculation corresponding to different requests by establishing a shared storage pool, enabling the computing device to reuse this intermediate data for the next inference, thereby reducing the computation amount of the large language model and further reducing the inference latency of the large language model. Embodiments of this application also provide a reasoning device for a large language model, a computing device, a computing device cluster, a computer-readable storage medium, and a computer program product corresponding to the reasoning method for the large language model.

[0007] In a first aspect, an embodiment of the present application provides a method for inferring a large language model. This method can be executed by a computing device, or by components of the computing device, such as the processor, chip, or chip system of the computing device, etc., or can also be implemented by a logic module or software that can implement all or part of the functions of the computing device. The method provided in the first aspect includes: The computing device obtains the input data of the large language model running on multiple processors. The input data includes the input text of the user and the reference information generated based on the input text. The reference information includes the domain-related knowledge corresponding to the input text and the timeliness information corresponding to the input text. The computing device matches the historical key-value cache data in the shared storage pool based on the input data to determine the key-value cache data corresponding to the input data. The historical key-value cache data is used to indicate the intermediate data generated by the large language model when processing different historical input data. The shared storage pool is used to store the historical key-value cache data generated by multiple processors. Different historical input data includes historical input text and the reference information corresponding to the historical input text. The computing device generates the output result corresponding to the input data according to the key-value cache data.

[0008] In the embodiment of the present application, the computing device can store the intermediate data calculated during the inference process of the historical input text based on the shared storage pool, so that the computing device can reuse this intermediate data for the next inference, reducing the calculation amount of the large language model and further reducing the inference latency of the large language model.

[0009] In a possible implementation manner, during the process of the computing device matching the historical key-value cache data in the shared storage pool based on the input data, the computing device segments the input data to obtain the segmented input data, and determines the universally unique identifier corresponding to the segmented input data. The computing device matches the historical key-value cache data in the shared storage pool according to the universally unique identifier to determine the key-value cache data corresponding to the segmented input data.

[0010] In the embodiment of the present application, the computing device can match the historical key-value cache data in the shared storage pool based on the unique identifier of the segmented input data, thereby improving the matching efficiency of the key-value cache data and further improving the inference efficiency of the large language model.

[0011] In a possible implementation manner, the computing device performs a prefix tree search on the segmented input data in the shared storage pool to determine the key-value cache data corresponding to the segmented input data. Specifically, the computing device performs a prefix match based on the segmented input data and the key-value cache data in the shared storage to first achieve approximate matching. If the approximate matching is successful, the computing device can further calculate the hash value of the segmented input data and the key-value cache data in the shared storage for hash matching to achieve complete matching.

[0012] In the embodiments of the present application, the computing device can also perform a prefix tree search on the segmented input data to determine the key-value cache data corresponding to the segmented input data, thereby improving the matching accuracy of the input data.

[0013] In a possible implementation, when the computing device generates the output result corresponding to the input data based on the key-value cache data, it splices the key-value cache data corresponding to the segmented input data to generate the full-segment key-value cache data corresponding to the input data. The computing device processes the full-segment key-value cache data based on the large language model to generate the output result.

[0014] In the embodiments of the present application, the computing device can splice the key-value cache data corresponding to the segmented input data to generate the full-segment key-value cache data corresponding to the input data, thereby improving the feasibility of the solution.

[0015] In a possible implementation, during the splicing process of the segmented input data by the computing device, the large language model can leave blank the key-value cache data that was not matched in the shared storage pool during the first-round inference process as the missing morphemes of the full-segment key-value cache data. During subsequent inference processes, the large language model fills the missing morphemes based on the historical key-value cache data in the shared storage pool, or the large language model directly calculates these missing morphemes.

[0016] In the embodiments of the present application, the computing device fills the missing morphemes with the historical key-value cache data in the shared storage pool or calculates the missing morphemes based on the large language model, thereby improving the richness of the solution.

[0017] In a possible implementation, the computing device can calculate the costs of searching for key-value cache data and directly calculating key-value cache data in the inference process through a cost model, so as to determine the way to obtain the key-value cache data. For example, during the process of filling the missing morphemes, the computing device can calculate the cost of searching for historical key-value cache data in the shared storage pool to fill the missing morphemes through the cost model. If the cost of searching for historical key-value cache data in the shared storage pool to fill the missing morphemes is higher than the cost of directly calculating these missing morphemes based on the filled key-value cache data, the computing device directly calculates these missing morphemes based on the filled key-value cache data.

[0018] In the embodiments of the present application, the computing device can calculate the costs of obtaining key-value cache data in different ways through a cost model, so as to determine the way to obtain the key-value cache data according to the acquisition cost, thereby improving the inference efficiency of the large language model.

[0019] In a possible implementation, when the input data does not match the historical key-value cache data in the shared storage pool, the computing device performs reasoning on the input data based on the large language model to generate key-value cache data corresponding to the input data, and the computing device stores the key-value cache data in the shared storage pool. Specifically, the computing device performs reasoning calculations on the segmented input data to generate key-value cache data corresponding to the segmented input data, and stores these key-value cache data in the shared storage pool.

[0020] In the embodiments of the present application, when the input data does not match the historical key-value cache data in the shared storage pool, the computing device can perform reasoning on the input data based on the large language model to generate key-value cache data corresponding to the input data, and store these key-value cache data in the shared storage pool for subsequent matching of the key-value cache data of the input data, thereby improving the reasoning efficiency of the large language model.

[0021] In a possible implementation, before the computing device obtains the input data of the large language model, the computing device receives the input text input by the user, and the input text includes the user's question text. The computing device performs retrieval enhancement on the input text based on the prompt engineering system to generate reference information corresponding to the input text, and these reference information can provide domain-related knowledge corresponding to the input text and timeliness information corresponding to the input text for the large language model. Since the reference information corresponding to different input texts may be the same, there is duplication between the reference information corresponding to the input text and the reference information corresponding to the historical input text.

[0022] In the embodiments of the present application, the computing device can perform retrieval enhancement on the input text based on the prompt engineering system to generate reference information corresponding to the input text, and these reference information can provide domain-related knowledge corresponding to the input text and timeliness information corresponding to the input text for the large language model, thereby improving the reasoning accuracy of the large language model.

[0023] In a possible implementation, the historical key-value cache data includes historical key-value cache data corresponding to historical input texts of different users, historical input texts at different times, and historical input texts in different sessions.

[0024] In the embodiments of the present application, the historical key-value cache data in the shared storage pool includes historical key-value cache data corresponding to historical input texts of different users, historical input texts at different times, and historical input texts in different sessions, thereby realizing the cross-user, cross-request, and cross-session reuse of the historical key-value cache data, and further improving the reasoning efficiency of the large language model.

[0025] Second aspect, an embodiment of the present application provides an inference device for a large language model. The inference device for the large language model includes an acquisition unit and a processing unit. Among them, the acquisition unit is used to acquire the input data of the large language model running on multiple processors. The input data includes the input text of the user and the reference information generated based on the input text. The reference information includes the domain-related knowledge corresponding to the input text and the timeliness information corresponding to the input text. The processing unit is used to match the historical key-value cache data in the shared storage pool based on the input data, and determine the key-value cache data corresponding to the input data. The historical key-value cache data is used to indicate the intermediate data generated by the large language model when processing different historical input data. The shared storage pool is used to store the historical key-value cache data generated by multiple processors. Different historical input data includes historical input text and the reference information corresponding to the historical input text. The processing unit is further used to generate the output result corresponding to the input data according to the key-value cache data.

[0026] In a possible implementation manner, the processing unit is specifically used to segment the input data to obtain the segmented input data, and determine the universal unique identifier corresponding to the segmented input data. The processing unit is specifically used to match the historical key-value cache data in the shared storage pool according to the universal unique identifier, and determine the key-value cache data corresponding to the segmented input data.

[0027] In a possible implementation manner, the processing unit is specifically used to splice the key-value cache data corresponding to the segmented input data to generate the full-segment key-value cache data corresponding to the input data. The large language model is used to process the full-segment key-value cache data to generate the output result.

[0028] In a possible implementation manner, the full-segment key-value cache data includes vacant morphemes. The processing unit is further used to fill the vacant morphemes based on the historical key-value cache data in the shared storage pool, or calculate the vacant morphemes based on the large language model.

[0029] In a possible implementation manner, when the input data does not match the historical key-value cache data in the shared storage pool, the processing unit is further used to infer the input data based on the large language model to generate the key-value cache data corresponding to the input data. The key-value cache data is stored in the shared storage pool.

[0030] In a possible implementation manner, the historical key-value cache data includes the historical key-value cache data corresponding to the historical input text of different users, the historical input text at different times, and the historical input text in different sessions.

[0031] In a third aspect, an embodiment of the present application provides a computing device. The computing device includes a processor coupled to a memory. The processor is configured to store instructions that, when executed by the processor, cause the computing device to perform the method described in the first aspect or any possible implementation manner of the first aspect.

[0032] In a fourth aspect, an embodiment of the present application provides a cluster of computing devices. The cluster of computing devices includes one or more computing devices. Each computing device includes a processor coupled to a memory. The processor is configured to store instructions that, when executed by the processor, cause the cluster of computing devices to perform the method described in the first aspect or any possible implementation manner of the first aspect.

[0033] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium having instructions stored thereon that, when executed, cause a computer to perform the method described in the first aspect or any possible implementation manner of the first aspect.

[0034] In a sixth aspect, an embodiment of the present application provides a computer program product that includes instructions which, when executed, cause a computer to implement the method described in the first aspect or any possible implementation manner of the first aspect.

[0035] It can be understood that the beneficial effects that can be achieved by any of the above-provided inference devices, computing devices, clusters of computing devices, computer-readable media, or computer program products of large language models can refer to the beneficial effects in the corresponding methods, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 FIG. is a schematic diagram of the system architecture of a large language model inference system provided by an embodiment of the present application;

[0037] Figure 2 FIG. is a schematic flowchart of a large language model inference method provided by an embodiment of the present application;

[0038] Figure 3 FIG. is a schematic flowchart of another large language model inference method provided by an embodiment of the present application;

[0039] Figure 4 FIG. is a schematic flowchart of another large language model inference method provided by an embodiment of the present application;

[0040] Figure 5 FIG. is a schematic diagram of the structure of a large language model inference device provided by an embodiment of the present application;

[0041] Figure 6 FIG. is a schematic diagram of the structure of a computing device provided by an embodiment of the present application;

[0042] Figure 7 This is a schematic structural diagram of a computing device cluster provided by an embodiment of the present application;

[0043] Figure 8 This is another schematic structural diagram of a computing device cluster provided by an embodiment of the present application. Specific implementation manners

[0044] The embodiments of the present application provide a method and apparatus for inferring a large language model, which are used to reduce the inference latency during the inference process of the large language model.

[0045] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product or device.

[0046] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0047] First, some terms involved in the embodiments of the present application are introduced to facilitate those skilled in the art to understand the technical solutions.

[0048] A large language model (LLM) is a natural language processing model based on deep learning technology. It is trained using a large amount of text data and has the ability to understand, generate and process natural language. The large language model treats natural language text as a kind of sequence data, such as a word sequence or a character sequence, and models the statistical laws and potential semantic information of these sequence data through a deep learning model. The large language model can perform a wide range of tasks, including text summarization, translation, sentiment analysis, question answering, dialogue, etc.

[0049] Prompt engineering is a technique used to improve the input of large language models. By adjusting the input prompts, prompt engineering guides large language models to understand and answer relevant questions more accurately. For example, prompt engineering can optimize prompts, emphasize key information, reduce the model's misunderstanding of context, and provide task-specific guidance, enabling the model to better adapt to the needs of specific domains or tasks.

[0050] A token is the smallest semantic unit after tokenizing a language, and can also be called a tokenization or a lexical marker, etc. A token can be a single Chinese character or multiple Chinese characters, a single English word or half an English word. For example, "apple" is a token, and "apple tree" can be two tokens "apple" + "tree", but "ping" alone does not represent any meaning.

[0051] Context refers to the context information that large language models rely on when processing text, that is, the context environment based on which large language models generate or understand text. For large language models, the context can be a piece of text, a conversation, a question, or any other form of text input.

[0052] To make the technical solution of this application clearer and easier to understand, the system architecture of this application will be introduced below with reference to the accompanying drawings.

[0053] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the system architecture of a large language model inference system provided for an example of this application. In the Figure 1 system architecture shown, the large language model inference system 10 includes a prompt engineering module 101, a large language model inference module 102, and a shared storage pool module 103. Among them, the large language model inference module 102 includes an embedding layer 1021, a self-attention layer 1022, and a feed-forward network layer 1023. The specific functions of each module of the large language model inference system 10 will be introduced below.

[0054] The prompt engineering module 101 is used to generate reference information corresponding to the input text according to the user's input text, and input the input text and the reference information into the large language model inference module 102. These input texts and reference information can also be called enhanced texts. Among them, relevant domain knowledge and timeliness information are stored in the prompt engineering module 101. Therefore, the prompt engineering module 101 can perform retrieval enhancement based on the input text in the relevant domain knowledge and timeliness information, and obtain the relevant domain knowledge corresponding to the input text and the timeliness information corresponding to the input text.

[0055] It should be noted that since the relevant domain knowledge and timeliness of different input texts can be the same, different input texts can be retrieved and enhanced to obtain the same reference information. That is, the reference information obtained by the retrieval enhancement of the prompt engineering module 101 for different input texts may overlap.

[0056] The large language model inference module 102 has a large language model built in. The large language model can perform inference calculations on the input text and reference information to generate corresponding output text. For example, the large language model inference module 102 can receive a natural language text as the input text, and the input text can be a sentence, a paragraph, or a longer text sequence. The large language model inference module 102 can perform inference calculations on the input text and reference information based on the large language model to generate output text, and the output text can be the question answer, text summary, or translation result corresponding to the input text, etc.

[0057] The model structure of the large language model built in the large language model inference module 102 is a multi-layer structure. For example, the multi-layer structure includes an embedding layer 1021, a self-attention layer 1022, and a feed-forward network layer 1023. Among them, the embedding layer 1021 is used to convert each morpheme in the user's input text and reference information into a vector representation. The self-attention layer 1022 is used to determine the context relationship between the input text and reference information. For example, the self-attention layer 1022 performs self-attention calculations on the vector representations corresponding to the input text and reference information to generate the context representations corresponding to the input text and reference information. The feed-forward network layer 1023 is used to perform non-linear transformation after the self-attention layer 1022 to extract the high-level semantic features in the input text and generate feature representations.

[0058] The inference stage of the large language model inference module 102 includes a prefill stage and a decode stage. Among them, the prefill stage is also called the first-character inference stage or the full-scale inference stage. In the prefill stage, the computing device inputs the complete input text and reference information into the large language model at one time for the first round of inference to generate the first output character of the output text. The decode stage is also called the incremental inference stage. In the decode stage, the computing device can continuously generate the next output character based on the autoregressive loop iteration of the generated output text.

[0059] The shared storage pool module 103 is used to store the key-value cache (KVcache) data generated by the large language model during the inference process. The key-value cache data is the intermediate data generated by the large language model during the inference process. Since the reference information retrieved and enhanced by the prompt engineering module 101 for different input texts may overlap, the key-value cache data generated by the large language model for different input texts during the inference process may also be repeated. The large language model can reuse this key-value cache data, and the computing device can store this key-value cache data through the shared storage pool module 103.

[0060] It should be noted that each module of the above large language model inference system 10 can be deployed on a single computing device or a computing device cluster composed of multiple computing devices, and there is no specific limitation.

[0061] Based on Figure 1 The large language model inference system 10 shown, the present application also provides an inference method for a large language model. The following will introduce the inference method for the large language model provided by the embodiments of the present application in conjunction with the embodiments.

[0062] Please refer to Figure 2 , Figure 2 which is a schematic flow chart of an inference method for a large language model provided by an embodiment of the present application. In Figure 2 the example shown, the method includes the following steps:

[0063] Step 201. The computing device obtains the input data of the large language model. The input data includes the user's input text and the reference information generated based on the input text. The reference information includes the domain-related knowledge corresponding to the input text and the timeliness information corresponding to the input text.

[0064] The computing device obtains the input data of the large language model. The input data includes the user's input text and the reference information generated based on the input text. Among them, the input text refers to the text input by the user into the language model. For example, the input text can be a sentence or paragraph of the user's question. The reference information refers to the text retrieved and enhanced through prompt engineering.

[0065] Specifically, the computing device receives the input text input by the user into the large language model, enhances the user's input text based on prompt engineering, and obtains the reference information of the input text. The user's input text and the reference information corresponding to the input text together serve as the input data of the large language model. Since the input data is the text retrieved and enhanced through prompt engineering, this input data can also be referred to as augmented text.

[0066] It is understandable that the input data enhanced by prompt engineering includes the input text and the reference information corresponding to the input text. These reference information can provide domain-related knowledge corresponding to the input text and timeliness information corresponding to the input text for the large language model. During the process of the computing device enhancing the input text based on prompt engineering, reference information retrieval can be performed through components such as a vector database (vector DB) component or a search engine.

[0067] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of an inference method for a large language model provided by an embodiment of this application. In Figure 3 the example shown, the user's input text is, for example, "What new features does the AITO M9 have?" The computing device enhances the input text based on prompt engineering and generates reference information for the input text. Among them, the reference information is, for example, timeliness information such as "Press conference news of the AITO M9" and domain-related knowledge such as "User manual of the AITO M9".

[0068] For another example, the user's input text is, for example, "What was covered in the AITO M9 press conference?" The computing device enhances the input text based on prompt engineering and generates reference information for the input text. Among them, the reference information is, for example, timeliness information such as "Press conference news of the AITO M9". Among them, for the two different input texts of the user, "What new features does the AITO M9 have?" and "What was covered in the AITO M9 press conference?", the reference information generated by the computing device based on prompt engineering can overlap. For example, for these two input texts, the prompt engineering module 101 can generate reference information for the two input texts from the timeliness information of "Press conference news of the AITO M9".

[0069] In Figure 3 the example shown, after the computing device generates the reference information of the input text, the input text and the reference information are used as input data and input into the large language model.

[0070] In a possible implementation manner, in addition to the domain-related knowledge corresponding to the input text and the timeliness information corresponding to the input text, the reference information further includes instructions and examples of the large language model, where the instructions are used to indicate the inference limitations and inference requirements of the large language model, and the examples are used to indicate the format of the text output by the large language model.

[0071] Step 202. The computing device matches the historical key-value cache data in the shared storage pool based on the input data and determines the key-value cache data corresponding to the input data.

[0072] After the computing device obtains the input data of the large language model, during the process of performing inference on the large language model, it matches the historical key-value cache data in the shared storage pool based on the input data to determine the key-value cache data corresponding to the input data. Among them, the historical key-value cache data is the intermediate data generated by the large language model when processing different historical input data. These intermediate data are stored in the form of multi-dimensional vectors, and the shared storage pool can store these historical key-value cache data.

[0073] It should be noted that during the process of the computing device performing inference on the large language model, it can also directly perform inference calculations based on the input data without retrieving the key-value cache data in the shared storage pool. That is, the input data is processed based on the multi-layer structure of the large language model to generate the output text corresponding to the input data. These processes include converting the input data into a vector representation through the embedding layer, then using the self-attention layer to mine and integrate the key information in the vector representation, and further calculating and optimizing through the feed-forward network layer, and finally converting the generated vector representation into the output text, etc.

[0074] In a possible implementation, during the process of the computing device matching the historical key-value cache data in the shared storage pool based on the input data, it first segments the input data based on the segmentation algorithm to obtain the segmented input data. The segmentation algorithm can be, for example, the segmentation algorithm based on special characters and the adaptive segmentation algorithm. After the computing device obtains the segmented input data, it matches the segmented input data in the shared storage pool to determine the key-value cache data corresponding to each segmented input data. Specifically, the computing device needs to first convert the segmented input data into a vector identifier and match the historical key-value cache data in the shared storage pool based on the vector representation corresponding to the segmented input data.

[0075] Please continue to refer to Figure 3 , in Figure 3 In the example shown, after the computing device receives the input data, the large language model performs inference calculations based on the input data. During the process of the large language model performing inference calculations, the computing device first segments the input data based on the segmentation algorithm to obtain the segmented input data, and converts the segmented input data into a vector representation. Then, the large language model performs multiple rounds of inference calculations based on the converted vector representation, including calculations in multiple transmission layers such as the self-attention layer and the feed-forward network layer, and finally generates the output text corresponding to the input data.

[0076] In Figure 3In the example shown, in each round of inference calculation of the large language model, the large language model can match and search for historical key-value cache data in the shared storage pool based on the vector representation corresponding to the segmented input data. If there is historical key-value cache data corresponding to the segmented input data in the shared storage pool, the large language model does not need to perform inference calculation again based on the vector representation of the segmented input data, but directly obtains the corresponding key-value cache data from the shared storage pool for the next round of inference calculation, thereby improving the inference efficiency.

[0077] In a possible implementation, when the input data does not match the historical key-value cache data in the shared storage pool, the computing device performs inference on the input data based on the large language model, generates key-value cache data corresponding to the input data, and stores the key-value cache data in the shared storage pool. Specifically, the computing device performs inference calculation on the segmented input data, generates key-value cache data corresponding to the segmented input data, and stores these key-value cache data in the shared storage pool.

[0078] Please continue to refer to Figure 3 , in Figure 3 the example shown, if there is no historical key-value cache data corresponding to the segmented input data in the shared storage pool, the large language model can perform inference calculation based on the segmented input data, obtain key-value cache data corresponding to the segmented input data, and store the key-value cache data corresponding to the segmented input data in the shared storage pool.

[0079] It should be noted that the historical key-value cache data stored in the shared storage pool includes historical key-value cache data corresponding to historical input texts of different users, historical input texts at different times, and historical input texts in different sessions. Since there may be duplicates in the historical key-value cache data corresponding to historical input texts of different users, historical input texts at different times, and historical input texts in different sessions, deduplication operations need to be performed on the key-value cache data stored in the shared storage pool.

[0080] For example, the historical key-value cache data corresponding to historical input texts of different users, for example, key-value cache data generated based on the input text of user A and enhanced text, and key-value cache data generated based on the input text of user B and enhanced text. The historical key-value cache data corresponding to historical input texts at different times, for example, key-value cache data generated based on the input text of the user at 09:50:00 on January 19, 2024, and key-value cache data generated based on the input text of the user at 09:57:00 on January 19, 2024. The historical key-value cache data corresponding to historical input texts in different sessions, for example, key-value cache data corresponding to the input text of the question in the user's first session and key-value cache data corresponding to the input text of the question in the user's second session.

[0081] In a possible implementation, after the computing device obtains the segmented input data, it further determines the universally unique identifier (UUID) corresponding to the segmented input data, and matches the historical key-value cached data in the shared storage pool according to the universally unique identifier to determine the key-value cached data corresponding to the segmented input data.

[0082] Please refer to Figure 4 , Figure 4 which is another schematic diagram of the inference process of the large language model provided by the embodiments of this application. In Figure 4 Steps a to f of the example, after the large language model obtains the input data, it segments the input data to obtain the segmented input data, and further determines the universally unique identifier corresponding to the segmented input data. The computing device matches the key-value cached data in the shared storage pool based on the universally unique identifier. If the key-value cached data is matched, the key-value cached data is read from the shared storage pool. If the key-value cached data is not matched, the position corresponding to the segmented input data is left as a vacant morpheme for the next round of inference.

[0083] In a possible implementation, the computing device can also perform a prefix tree search on the segmented input data to determine the key-value cached data corresponding to the segmented input data. Specifically, the computing device performs a prefix match based on the segmented input data and the key-value cached data in the shared storage to achieve approximate matching. If the match is successful, the computing device can further perform a hash match on the segmented input data and the key-value cached data in the shared storage to achieve exact matching.

[0084] Please continue to refer to Figure 3 , in Figure 3 the example shown, the large language model matches and searches for historical key-value cached data in the shared storage pool. It can index the historical key-value cached data based on the universally unique identifier UUID of the segmented input data, or perform a prefix tree search in the shared storage pool based on the vector representation corresponding to the segmented input data.

[0085] For example, when the input data is the input text "What new features does the AITO M9 have?", and the reference information is "Press conference news of the AITO M9" and "User manual of the AITO M9", during the process of the large language model performing inference based on these input data, it segments the input data and matches the historical key-value cached data in the shared storage pool based on the vector representation corresponding to the segmented input data. For example, the large language model matches "key-value cached data 1" in the shared storage pool for "segmented input data 1", and the output text corresponding to the key-value cached data 1 is, for example, a section of "Technical parameter introduction content of the AITO M9".

[0086] Please continue to refer to Figure 4 , in Figure 4 steps k to m and steps e to f of the example, after the large language model segments the input data and obtains the segmented input data, the computing device can also directly match the key-value cache data in the shared storage pool based on the vector representation corresponding to the segmented input data. Specifically, the computing device can also perform prefix matching between the vector representation corresponding to the segmented input data and the historical key-value cache data in the shared storage pool. If the prefix matching is successful, it indicates that there is key-value cache data in the shared storage pool that approximately matches the segmented input data. The computing device further performs hash matching on the approximately matched key-value cache data, that is, matches the hash value of the segmented input data and the approximately matched key-value cache data. If the hash matching is successful, the successfully matched key-value cache data is read from the shared storage pool. If the prefix matching is unsuccessful or the hash matching is unsuccessful, the position corresponding to the segmented input data is left as a vacant morpheme for the next round of inference.

[0087] Step 203. The computing device generates an output result corresponding to the input data according to the key-value cache data.

[0088] The computing device generates an output result corresponding to the input data according to the key-value cache data. Specifically, after the computing device obtains the key-value cache data corresponding to each segmented input data, it splices the key-value cache data corresponding to these segmented input data to generate the full-segment key-value cache data corresponding to the input data, and processes the full-segment key-value cache data based on the large language model to generate an output result.

[0089] Please continue to refer Figure 3 , in Figure 3 the example shown, after the computing device obtains the key-value cache data of the segmented input data, it splices the key-value cache data corresponding to the segmented input data to obtain the full-segment key-value cache data corresponding to the input data. For example, the computing device splices key-value cache data 1, key-value cache data 2, key-value cache data 3, and key-value cache data 4, and performs inference calculation on the spliced full-segment key-value cache data to generate an output text.

[0090] In a possible implementation manner, during the splicing process of the segmented input data by the computing device, the large language model can leave blank the key-value cache data that is not matched in the shared storage pool during the first round of inference as the vacant morpheme of the full-segment key-value cache data. During subsequent inference processes, the large language model fills the vacant morpheme based on the historical key-value cache data in the shared storage pool, or the large language model directly calculates these vacant morphemes.

[0091] It can be understood that the computing device can calculate the costs of searching for key-value cache data and directly calculating key-value cache data from shared cache data during the inference process through a cost model, so as to determine the way to obtain the key-value cache data. For example, during the process of filling in missing morphemes, the computing device can calculate the cost of filling in missing morphemes with search history key-value cache data in the shared storage pool through the cost model, as well as the cost of directly calculating these missing morphemes based on the filled key-value cache data, so as to determine the way to fill in the missing morphemes.

[0092] Please continue to refer Figure 3 , in Figure 3 In the example shown, during the first-round inference process of the large language model, when splicing the key-value cache data corresponding to the segmented input data, there will be some missing morphemes. For example, the large language model obtains key-value cache data 1, key-value cache data 2, and key-value cache data 4 from the shared storage pool, and splices key-value cache data 1, key-value cache data 2, and key-value cache data 4. Among them, there are missing morphemes between key-value cache data 2 and key-value cache data 4. The large language model needs to fill in the missing morphemes during the next-round inference calculation.

[0093] In Figure 3 In the example shown, when the large language model obtains the key-value cache data 3 corresponding to the missing morphemes during the next-round inference calculation, the computing device can retrieve the key-value cache data 3 through retrieval in the shared storage, or perform inference calculation based on key-value cache data 1 and key-value cache data 2 to obtain the key-value cache data 3, and no specific limitation is made.

[0094] Please continue to refer to Figure 4 , in Figure 4 In steps g to j of the example shown, during the next-round inference calculation process of the large language model, based on the cost model, the way to fill in the missing morphemes is determined. The computing device can fill in the missing morphemes with the search history key-value cache data in the shared storage pool, or directly calculate these missing morphemes based on the filled key-value cache data. After the computing device fills in the missing morphemes, it can perform the next-step inference based on the large language model engine, including the inference in the pre-filling stage and the inference in the decoding stage.

[0095] It should be noted that when the computing device directly calculates these missing morphemes based on the filled key-value cache data, it only calculates the missing morphemes based on the key-value cache data before the missing morphemes. At the same time, when the computing device calculates the missing morphemes, it can preload the key-value cache data after the missing morphemes in parallel, and no specific limitation is made.

[0096] Please continue to refer Figure 3 , in Figure 3In the illustrated example, when the large language model fills in the key-value cache data 3 corresponding to the missing morpheme, it can preload the key-value cache data 4 after the missing morpheme. However, the large language model only infers and calculates the key-value cache data 3 based on the key-value cache data 1 and the key-value cache data 2.

[0097] As can be seen from the above embodiments, in the embodiments of the present application, the computing device can store the intermediate data obtained during the inference process of the historical input text based on the shared storage pool. Since there is duplication between the intermediate data corresponding to the user's input text and the intermediate data corresponding to the historical input text, the computing device can reuse these intermediate data for the next inference, thereby reducing the computational amount of the large language model and reducing the inference latency of the large language model.

[0098] Based on the above method embodiments, the embodiments of the present application further provide an inference device for a large language model. The inference device for the large language model provided by the embodiments of the present application will be specifically introduced below.

[0099] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an inference device for a large language model provided by an embodiment of the present application. In Figure 5 the illustrated example, the inference device 500 of the large language model is used to implement each step executed by the end-side sharing module in the above embodiments. The inference device 500 of the large language model includes an acquisition unit 501 and a processing unit 502.

[0100] Among them, the acquisition unit 501 is used to acquire the input data of the large language model running on multiple processors. The input data includes the user's input text and the reference information generated based on the input text. The reference information includes the domain-related knowledge corresponding to the input text and the timeliness information corresponding to the input text. The processing unit 502 is used to match the historical key-value cache data in the shared storage pool based on the input data, and determine the key-value cache data corresponding to the input data. The historical key-value cache data is used to indicate the intermediate data generated by the large language model when processing different historical input data. The shared storage pool is used to store the historical key-value cache data generated by multiple processors. Different historical input data includes historical input text and the reference information corresponding to the historical input text. The processing unit 502 is further used to generate the output result corresponding to the input data according to the key-value cache data.

[0101] In a possible implementation manner, the processing unit 502 is specifically used to segment the input data to obtain the segmented input data, and determine the universally unique identifier corresponding to the segmented input data. The processing unit 502 is specifically used to match the historical key-value cache data in the shared storage pool according to the universally unique identifier, and determine the key-value cache data corresponding to the segmented input data.

[0102] In a possible implementation, the processing unit 502 is specifically configured to splice the key-value cache data corresponding to the segmented input data to generate the full-segment key-value cache data corresponding to the input data. Based on the large language model, the full-segment key-value cache data is processed to generate an output result.

[0103] In a possible implementation, the full-segment key-value cache data includes vacant morphemes, and the processing unit 502 is further configured to fill the vacant morphemes based on the historical key-value cache data in the shared storage pool, or calculate the vacant morphemes based on the large language model.

[0104] In a possible implementation, when the historical key-value cache data corresponding to the input data is not found in the shared storage pool, the processing unit 502 is further configured to perform inference on the input data based on the large language model to generate the key-value cache data corresponding to the input data. The key-value cache data is stored in the shared storage pool.

[0105] In a possible implementation, the historical key-value cache data includes the historical key-value cache data corresponding to the historical input texts of different users, the historical input texts at different times, and the historical input texts in different sessions.

[0106] It can be understood that the acquisition unit 501 and the processing unit 502 in the inference device 500 of the large language model can be used as functional modules and Figure 1 there is a mapping with each module in the large language model inference system 10 in

[0107] It should be understood that the division of units in the above device is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And the units in the device can all be implemented in the form of software called by processing elements; they can also all be implemented in hardware form; or some units can be implemented in the form of software called by processing elements, and some units can be implemented in hardware form. For example, each unit can be a separately established processing element, or can be integrated in a certain chip of the device. In addition, it can also be stored in the memory in the form of a program, and the function of the unit is called and executed by a certain processing element of the device. In addition, all or part of these units can be integrated together or can be independently implemented. The processing element mentioned here can also be called a processor, which can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above units can be implemented by the integrated logic circuit of the hardware in the processor element or in the form of software called by the processing element.

[0108] It should be noted that, for the above method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to this application.

[0109] Other reasonable combinations of steps that can be conceived by those skilled in the art based on the above description also fall within the protection scope of this application. Secondly, those skilled in the art should also be familiar with the fact that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to this application.

[0110] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a computing device provided by an embodiment of this application. As Figure 6 shown, the computing device 600 includes: a processor 601, a memory 602, a communication interface 603, and a bus 604. The processor 601, the memory 602, and the communication interface 603 are coupled through a bus (not labeled in the figure). The memory 602 stores instructions. When the execution instructions in the memory 602 are executed, the computing device 600 executes the method executed by the computing device in the above method embodiments.

[0111] The computing device 600 can be one or more integrated circuits configured to implement the above method. For example: one or more application specific integrated circuits (ASICs), or, one or more digital signal processors (DSPs), or, one or more field programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms. Again, when the units in the device can be implemented in the form of a processing element scheduler, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call programs. Again, these units can be integrated together to be implemented in the form of a system-on-a-chip (SOC).

[0112] The processor 601 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0113] The memory 602 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0114] The executable program code is stored in the memory 602, and the processor 601 executes the executable program code to implement the functions of the foregoing units or modules respectively, thereby implementing the inference method of the above large language model. That is to say, the instructions for executing the inference method of the above large language model are stored on the memory 602.

[0115] The communication interface 603 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 600 and other devices or communication networks.

[0116] In addition to including a data bus, the bus 604 may also include a power bus, a control bus, a status signal bus, etc. The bus may be a Peripheral Component Interconnect Express (PCIe) bus, or an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL), a Cache Coherent Interconnect for Accelerators (CCIX), etc. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0117] Please refer to Figure 7 , Figure 7 for a schematic diagram of a computing device cluster provided by an embodiment of this application. As Figure 7 shown, the computing device cluster 700 includes at least one computing device 600.

[0118] As Figure 7 shown, the computing device cluster 700 includes at least one computing device 600. Instructions for executing the above-mentioned inference method of the large language model may be stored in the memories 602 of one or more of the computing devices 600 in the computing device cluster 700.

[0119] In some possible implementation manners, partial instructions for executing the above-mentioned inference method of the large language model may also be stored separately in the memories 602 of one or more of the computing devices 600 in the computing device cluster 700. In other words, a combination of one or more computing devices 600 can jointly execute the instructions for executing the above-mentioned inference method of the large language model.

[0120] It should be noted that the memories 602 in different computing devices 600 in the computing device cluster 700 may store different instructions, which are respectively used to execute partial functions of the above-mentioned inference device of the large language model. That is, the instructions stored in the memories 602 of different computing devices 600 can implement the functions of one or more modules in the acquisition unit and the processing unit.

[0121] In some possible implementation manners, one or more computing devices 600 in the computing device cluster 700 may be connected via a network. Among them, the network may be a wide area network or a local area network, etc.

[0122] Please refer to Figure 8 , Figure 8 which is a schematic diagram of a computer device in a computer cluster being connected via a network provided by an embodiment of this application. As Figure 8 shown, two computing devices 600A and 600B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device.

[0123] In one possible implementation manner, instructions for implementing the functions of the acquisition unit are stored in the memory of the computing device 600A. At the same time, instructions for implementing the functions of the processing unit and the display unit are stored in the memory of the computing device 600B.

[0124] It should be understood that Figure 8 the functions of the computing device 600A shown in

[0125] can also be completed by multiple computing devices. Similarly, the functions of the computing device 600B can also be completed by multiple computing devices.

[0126] In another embodiment of this application, a computer-readable storage medium is further provided. Computer-executable instructions are stored in the computer-readable storage medium. When the processor of the device executes the computer-executable instructions, the device executes the method executed by the computing device in the foregoing method embodiment.

[0127] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0128] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in electrical, mechanical, or other forms.

[0129] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0130] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0131] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical disks and other various media that can store program codes.

Claims

1. A reasoning method for a large language model, characterized in that, Including: Obtain the input data of the large language model running on multiple processors, where the input data includes the input text of the user and reference information generated based on the input text, and the reference information includes domain-related knowledge corresponding to the input text and timeliness information corresponding to the input text; Match historical key-value cache data in the shared storage pool based on the input data, and determine the key-value cache data corresponding to the input data. The historical key-value cache data is used to indicate the intermediate data generated by the large language model when processing different historical input data. The shared storage pool is used to store the historical key-value cache data generated by the multiple processors. The different historical input data includes historical input text and reference information corresponding to the historical input text; Generate the output result corresponding to the input data according to the key-value cache data.

2. The method according to claim 1, characterized in that, The matching of the historical key-value cache data in the shared storage pool based on the input data includes: Segment the input data to obtain the segmented input data, and determine the universally unique identifier corresponding to the segmented input data; Match the historical key-value cache data in the shared storage pool according to the universally unique identifier, and determine the key-value cache data corresponding to the segmented input data.

3. The method according to claim 2, characterized in that The generating the output result corresponding to the input data according to the key-value cache data includes: Concatenate the key-value cache data corresponding to the segmented input data to generate the full-segment key-value cache data corresponding to the input data; Process the full-segment key-value cache data based on the large language model to generate the output result.

4. The method according to claim 3, characterized in that The full-segment key-value cache data includes missing morphemes, and the method further includes: Fill the missing morphemes based on the historical key-value cache data in the shared storage pool, or calculate the missing morphemes based on the large language model.

5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: When the input data does not match the historical key-value cache data in the shared storage pool, perform inference on the input data based on the large language model to generate the key-value cache data corresponding to the input data; Store the key-value cache data in the shared storage pool.

6. The method according to any one of claims 1 to 5, characterized in that, The historical key-value cache data includes the historical key-value cache data corresponding to the historical input text of different users, the historical input text at different times, and the historical input text in different sessions.

7. An inference device for a large language model, characterized in that, Including: An acquisition unit for acquiring the input data of the large language model running on multiple processors, where the input data includes the input text of the user and reference information generated based on the input text, and the reference information includes domain-related knowledge corresponding to the input text and timeliness information corresponding to the input text; A processing unit for matching historical key-value cache data in a shared storage pool based on the input data to determine the key-value cache data corresponding to the input data, where the historical key-value cache data is used to indicate intermediate data generated by the large language model when processing different historical input data, the shared storage pool is used to store the historical key-value cache data generated by the multiple processors, and the different historical input data includes historical input texts and reference information corresponding to the historical input texts; The processing unit is further configured to generate an output result corresponding to the input data according to the key-value cache data.

8. The device according to claim 7, characterized in that, Specifically, the processing unit is configured to: Segment the input data to obtain segmented input data, and determine a universally unique identifier corresponding to the segmented input data; Match historical key-value cache data in the shared storage pool according to the universally unique identifier to determine the key-value cache data corresponding to the segmented input data.

9. The device according to claim 8, wherein Specifically, the processing unit is configured to: Concatenate the key-value cache data corresponding to the segmented input data to generate full-segment key-value cache data corresponding to the input data; Based on the large language model, process the full-segment key-value cache data to generate the output result.

10. The device according to claim 9, characterized in that, The full-segment key-value cache data includes missing morphemes, and the processing unit is further configured to: Fill the missing morphemes based on the historical key-value cache data in the shared storage pool, or calculate the missing morphemes based on the large language model.

11. The device according to any one of claims 7 to 10, characterized in that, The processing unit is further configured to: When no historical key-value cache data is matched for the input data in the shared storage pool, perform inference on the input data based on the large language model to generate key-value cache data corresponding to the input data; Store the key-value cache data in the shared storage pool.

12. The device according to any one of claims 7 to 11, characterized in that The historical key-value cache data includes historical key-value cache data corresponding to historical input texts of different users, historical input texts at different times, and historical input texts in different sessions.

13. A computing device, characterized in that, It includes a processor, and the processor is coupled to a memory. The processor is used to store instructions, and when the instructions are executed by the processor, the electronic device is caused to execute the method according to any one of claims 1 to 6.

14. A cluster of computing devices, characterized in that, It includes at least one computing device, and the computing device includes a processor. The processor is coupled to a memory. The processor is used to store instructions, and when the instructions are executed by the processor, the computing device cluster is caused to execute the method according to any one of claims 1 to 6.

15. A computer-readable storage medium having instructions stored thereon, characterized in that, When the instructions are executed, the computer is caused to execute the method according to any one of claims 1 to 6.

16. A computer program product comprising instructions, characterized in that, When the instructions are executed, the computer is caused to implement the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Large model reasoning control method and device, equipment and medium

    CN121581247A

  • A control method, device, and equipment for large model inference and a medium

    CN121581247B

  • Method and apparatus for inference in large language model

    WO2025152397A1