Large model inference acceleration method, device, equipment, storage medium and program product
Patent Information
- Application Number
- CN202610861030.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-08-18
AI Technical Summary
[0009]本发明提供一种大模型推理加速方法、装置、设备、存储介质及程序产品,用以解决现有的大模型推理加速方法存在文本复制效率低、复制操作的准确性和稳定性不足的问题,限制了模型的推理加速效果的缺陷
[0018] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the large model inference acceleration methods described above.
Smart Images

Figure CN122596041A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, storage medium, and program product for accelerating large-scale model inference. Background Technology
[0002] Currently, Large Language Models (LLMs), also known as "large models," have been widely applied in various industries and office scenarios, greatly improving the efficiency of production and office work.
[0003] In many application scenarios, especially in question-answering scenarios using Retrieval-Augmented Generation (RAG), reference cache generation scenarios, and multi-turn dialogue scenarios, the output text of a large language model often exhibits high similarity and redundancy with existing reference texts (such as retrieval information, historical dialogues, user-provided reference documents, etc.). In question-answering scenarios using RAG, the large language model first retrieves information related to the user's prompts from a database or the internet, then summarizes the retrieved information to generate output text that answers the user's input. In this scenario, the output text of the large language model typically contains a large amount of retrieval data. Text fragments within the information; Large Language Models (MLMs) typically have a caching mechanism, caching the user's historical dialogues. In scenarios where cached content is used for generation, when processing prompts for each round of user input, the MLM can retrieve similar historical input data and output text from the cache as a reference for the current round's output text. In this scenario, the MLM's current round output text usually exhibits high similarity and repetition with related historical output texts in the cache. In multi-turn dialogue scenarios, users may repeatedly request modifications based on the MLM's output text through multiple rounds of dialogue. In this scenario, the MLM's multi-turn output text typically shows only minor changes and high text repetition.
[0004] The output text of a large language model is usually composed of multiple tokens. If each token in the output text needs to be regenerated by the large language model based on the context, it will increase the computational load and inference time of the large language model. Moreover, the generation quality of the large language model is not necessarily reliable, so it may also introduce model illusions and errors into the output text.
[0005] Based on the application scenarios and text generation characteristics of large language models, in order to improve the inference speed of large language models, the industry has begun to use text reuse technology. This technology allows large language models to directly copy the required words from existing reference texts during the generation of words in the output text, thereby reducing the inference time for regenerating words and reducing model illusions and errors.
[0006] Since large language models typically employ an autoregressive mechanism to generate lexical units, meaning they use a sequential generation mode that proceeds word by word, each generation requires all previously generated lexical units as input, some related technologies have proposed methods to accelerate large model inference by adapting to the autoregressive mechanism. These methods involve inputting user prompts into the large language model, which then enters a word-by-word generation mode. Before generating the next lexical unit, the top-level hidden state used to predict the next lexical unit is obtained from the large language model network. Based on this top-level hidden state, the action decision information for the corresponding next lexical unit is determined. If the action decision information is a copy action, the required lexical unit is copied from the existing reference text, and the copied result is directly used as the next lexical unit.
[0007] However, this inference acceleration method is achieved by copying word-by-word units, so the acceleration effect is limited, especially when the copied text is long, the copying efficiency of this method is low; in addition, whether to copy word units depends on the top-level hidden state of the model network, and the calculation of the top-level hidden state is unstable, uncontrollable and difficult to debug, so it is difficult to guarantee the accuracy and stability of the copying operation.
[0008] In summary, existing methods for accelerating large-scale model inference suffer from low text copying efficiency and insufficient accuracy and stability of copying operations, which limit the inference acceleration effect of the model. Summary of the Invention
[0009] This invention provides a method, apparatus, device, storage medium, and program product for accelerating large model inference, which solves the problems of low text copying efficiency, insufficient accuracy and stability of copying operations in existing methods for accelerating large model inference, thus limiting the inference acceleration effect of the model.
[0010] This invention provides a method for accelerating large-scale model inference, comprising: inputting user interaction prompts and user history dialogues into a large language model, performing autoregressive decoding to obtain the current word element generated by the large language model; detecting whether the current word element is a text manipulation semantic word element; a text manipulation semantic word element is a word element that instructs the large language model to copy text; if the current word element is a text manipulation semantic word element, then copying the target text segment from existing text information based on the text manipulation semantic word element; the existing text information includes history dialogues and the retrieval text of the large language model, and the target text segment includes multiple target words; updating the multiple target words to the end of the output text; based on the large language model, continuing autoregressive decoding to generate the next word element, using the next word element as the current word element, and returning to the step of detecting whether the current word element is a text manipulation semantic word element, until the large language model generates an end symbol to obtain the updated output text.
[0011] According to the large-scale inference acceleration method provided by the present invention, after detecting whether the current word is a text operation semantic word, the method further includes: if the current word is not a text operation semantic word, updating the current word to the end of the output text; based on the large language model, continuing to perform autoregressive decoding to generate the next word, taking the next word as the current word, and returning to the step of detecting whether the current word is a text operation semantic word, until the large language model generates an end symbol and obtains the updated output text.
[0012] According to a method for accelerating large-scale model inference provided by the present invention, text manipulation semantic units include first text copying units, which carry first text copying interval information; based on the text manipulation semantic units, a target text segment is copied from existing text information, including: performing text position legality verification on the first text copying units and obtaining a verification result; if the verification result is that the first text copying units pass the verification, then determining the first text corresponding to the first text copying interval information from the existing text information and using the first text as the target text segment; determining the number of target units in the target text segment, and assigning unit positions matching the number of target units to the target text segment in the output text.
[0013] According to a method for accelerating large-scale model inference provided by the present invention, text manipulation semantic units include second text copying units, which carry second text copying interval information and text deletion interval information. Based on the text manipulation semantic units, a target text segment is copied from existing text information, including: performing text position legality verification on the second text copying units and obtaining the verification result; if the verification result is that the second text copying units pass the verification, then determining the second text corresponding to the second text copying interval information from the existing text information, and deleting the third text corresponding to the text deletion interval information from the second text to generate the target text segment; determining the number of target units in the target text segment, and assigning unit positions matching the number of target units to the target text segment in the output text.
[0014] According to the large-scale inference acceleration method provided by the present invention, if the current word is a text manipulation semantic word, after copying the target text segment from the existing text information based on the text manipulation semantic word, the method further includes: calculating the key value information of each target word based on a large language model; determining the error between each key value information and the historical key value information; the historical key value information is the key value information of historical words generated by the large language model, and the historical key value information is cached in the key value cache; selecting several target words with the largest errors from all target words as words to be corrected; recalculating the key value of each word to be corrected to obtain the corrected key value information of each word to be corrected; and updating the corrected key value information of each word to be corrected to the end of the key value cache.
[0015] According to the present invention, a method for accelerating large-scale model inference includes, before inputting the user's interactive prompts and the user's historical dialogues into a large language model and performing autoregressive decoding to obtain the current lexical units generated by the large language model, the method further includes: constructing text manipulation semantic lexical units and adding them to the model vocabulary of the initial large language model; obtaining a text manipulation training dataset; and fine-tuning the initial large language model based on the text manipulation training dataset to obtain the large language model; wherein the large language model has the ability to use text manipulation semantic lexical units.
[0016] This invention also provides a large-scale model inference acceleration device, comprising: an input module for inputting user interaction prompts and user history dialogues into a large language model, performing autoregressive decoding to obtain the current word element generated by the large language model; a detection module for detecting whether the current word element is a text manipulation semantic word element; a text manipulation semantic word element is a word element that instructs the large language model to copy text; a copying module for copying a target text segment from existing text information based on the text manipulation semantic word element if the current word element is a text manipulation semantic word element; the existing text information includes history dialogues and the retrieval text of the large language model, and the target text segment includes multiple target words; an update module for updating the multiple target words to the end of the output text; and an output module for continuing autoregressive decoding based on the large language model to generate the next word element, using the next word element as the current word element, and returning to the step of detecting whether the current word element is a text manipulation semantic word element, until the large language model generates an end symbol to obtain the updated output text.
[0017] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any of the large model inference acceleration methods described above.
[0018] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the large model inference acceleration methods described above.
[0019] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the large model inference acceleration methods described above.
[0020] The present invention provides a method, apparatus, device, storage medium, and program product for accelerating large-scale inference, which abandons the traditional word-by-word copying method. It first inputs the user's interactive prompts and historical dialogues into a large language model, performs autoregressive decoding to obtain the current word generated by the large language model, and then checks whether the current word is a text manipulation semantic word used to instruct the large language model to copy text. If the current word is a text manipulation semantic word, then based on the text manipulation semantic word, it simultaneously copies the target text segment including multiple target words from the existing text information, thereby achieving multi-word synchronous copying of long texts. This eliminates the need to calculate the top-level hidden state of the model network word-by-word before copying word-by-word, effectively reducing the computational load of the large language model, shortening the inference time, improving text copying efficiency, and optimizing the inference acceleration effect of the large language model. Then, multiple target lexical units are updated to the end of the output text. The large language model is then used to continue autoregressive decoding to generate the next lexical unit. The next lexical unit is used as the current lexical unit, and the process returns to the step of detecting whether the current lexical unit is a text manipulation semantic unit. This continues until the large language model generates an end symbol, and the final updated output text is obtained. This inference acceleration method does not require the top-level hidden state of the model network to determine whether to copy lexical units. Instead, it determines whether to copy text based on the type of the current lexical unit that the large language model can understand (i.e., whether the current lexical unit is a text manipulation semantic unit). This avoids the problems of unstable, uncontrollable, and difficult-to-debug top-level hidden state calculation, which helps to ensure the accuracy and stability of the copying operation, reduces the risk of model inference errors, and thus helps to ensure the stability and reliability of model inference. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0022] Figure 1 This is one of the flowcharts of the large model inference acceleration method provided by the present invention.
[0023] Figure 2 This is the second flowchart of the large model inference acceleration method provided by the present invention.
[0024] Figure 3 This is a schematic diagram of the structure of the large model inference acceleration device provided by the present invention.
[0025] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0027] Please see Figure 1 and Figure 2 , Figure 1 This is one of the flowcharts illustrating the large model inference acceleration method provided by this invention. Figure 2 This is the second flowchart of the large model inference acceleration method provided by the present invention.
[0028] like Figure 1 As shown, the method for accelerating large model inference includes steps S110 to S150, and the specific details of each step are as follows: S110: Input the user's interactive prompts and the user's historical dialogues into the large language model, perform autoregressive decoding, and obtain the current word unit generated by the large language model.
[0029] In this embodiment, the large language model has two modes: text editing mode and model inference generation mode.
[0030] Specifically, such as Figure 2 As shown, when the user's interaction prompts and the user's historical dialogue (which can be used as the interaction context) are input into the large language model, the large language model will default to the model inference and generation mode.
[0031] In the model inference generation mode, the large language model first updates the global lexical index information, which is constructed based on existing text information, including the user's historical dialogue (which includes the user's historical input data and the large language model's historical output data) and the large language model's retrieval text.
[0032] The global token index information includes multiple consecutive token identifiers (i.e., token IDs). A token identifier is used to identify an existing token in the existing text information, and the global token index information also includes token information records for each existing token.
[0033] Among them, a lexical unit is the basic unit of the output text generated by the large language model through inference and prediction. A lexical unit can be a character, a word, or a phrase, or it can be a sub-word (i.e., a part of a word).
[0034] Optionally, for each existing lexical unit, its lexical information record includes, but is not limited to, the source information of the existing lexical unit (the source information includes, but is not limited to, historical dialogues, output text generated by the large language model in this round of inference, retrieval text of the large language model, etc.), character offset information, and adjacent lexical unit information (adjacent lexical unit information can be used to guide the deletion or insertion of text).
[0035] All existing lexical information records together form a global index structure that can be randomly accessed and supports interval extraction, insertion, and deletion, which can be called by the large language model at any time.
[0036] Furthermore, after updating the global lexical index information, the large language model can then proceed with normal inference, i.e., perform autoregressive decoding to generate the current lexical unit.
[0037] S120: Detect whether the current lexical is a text manipulation semantic lexical.
[0038] Text manipulation semantic lexical units are lexical units that instruct large language models to perform text copying.
[0039] Specifically, for each current word generated by the large language model, it is necessary to check whether the current word is a text manipulation semantic word (i.e., a text manipulation semantic token).
[0040] Among them, text manipulation semantic lexical units carry text copying interval information.
[0041] S130: If the current word is a text operation semantic word, then copy the target text segment from the existing text information based on the text operation semantic word.
[0042] The existing text information includes historical dialogues and retrieval text from a large language model, and the target text segment includes multiple target lexical units.
[0043] Specifically, if the current word is a text operation semantic word, the large language model will enter the text editing mode. In the text editing mode, the large language model can copy the target text segment corresponding to the text copying interval information from the existing text information according to the text copying interval information carried by the text operation semantic word. The target text segment includes multiple target words, thereby realizing the synchronous copying of multiple words in long texts, improving the copying efficiency of long texts, and helping to optimize the reasoning acceleration effect of the large language model.
[0044] S140: Update multiple target terms to the end of the output text.
[0045] Specifically, after copying to the target text segment containing multiple target words, the multiple target words can be concatenated to the end of the current output text, and the word generation position of the output text can be updated.
[0046] Assuming the target text segment has the following number of target tokens: ( (where is a positive integer greater than or equal to 2). Before concatenating the target word into the current output text, the word generation position of the current output text is denoted as . Then After concatenating the target words to the end of the current output text, the word generation positions of the output text also need to be updated.
[0047] The update formula for the word generation positions in the output text is as follows: ; Among them, the " " indicates assignment; the left side of the formula Indicates the lexical generation position of the updated output text; the right side of the formula This indicates the lexical generation position of the output text before the update.
[0048] Optionally, the word generation positions of the output text can be represented using RoPE encoding (Rotary Position Embedding), ALiBi (Attention with Linear Biases) encoding, absolute encoding, or other methods.
[0049] S150: Based on the large language model, continue autoregressive decoding to generate the next word, take the next word as the current word, and return to the step of detecting whether the current word is a semantic word for text operation, until the large language model generates the end symbol and obtains the updated output text.
[0050] Specifically, after the output text is updated, the large language model will revert to the model inference generation mode, update the global lexical index information, continue autoregressive decoding to generate the next lexical, and use the next lexical as the current lexical. Then, it will return to the step of detecting whether the current lexical is a semantic lexical for text operation, until the large language model generates the end symbol and obtains the updated output text. At this point, the updated output text is the text that is finally used to reply to the user.
[0051] The large-scale inference acceleration method provided in this embodiment abandons the traditional word-by-word copying method. It first inputs the user's interactive prompts and historical dialogues into the large language model, performs autoregressive decoding to obtain the current word generated by the large language model, and then checks whether the current word is a text manipulation semantic word used to instruct the large language model to copy text. If the current word is a text manipulation semantic word, then based on the text manipulation semantic word, it simultaneously copies the target text segment including multiple target words from the existing text information, thereby achieving multi-word synchronous copying of long texts. This eliminates the need to calculate the top-level hidden state of the model network word-by-word before copying word-by-word, effectively reducing the computational load of the large language model, shortening inference time, improving text copying efficiency, and optimizing the inference acceleration effect of the large language model. Then, multiple target words... The process involves updating to the end of the output text, then using the large language model to continue autoregressive decoding to generate the next word. This next word is then used as the current word, and the process returns to the step of checking whether the current word is a text manipulation semantic word. This continues until the large language model generates an end symbol, resulting in the final updated output text. This inference acceleration method does not require determining whether to copy words based on the top-level hidden state of the model network. Instead, it determines whether to copy text based on the type of the current word that the large language model can understand (i.e., whether the current word is a text manipulation semantic word). This avoids the problems of unstable, uncontrollable, and difficult-to-debug top-level hidden state computation, which helps ensure the accuracy and stability of the copying operation, reduces the risk of model inference errors, and thus helps ensure the stability and reliability of model inference.
[0052] In some embodiments, after detecting whether the current word is a text manipulation semantic word, the method further includes: if the current word is not a text manipulation semantic word, updating the current word to the end of the output text; based on the large language model, continuing autoregressive decoding to generate the next word, using the next word as the current word, and returning to the step of detecting whether the current word is a text manipulation semantic word, until the large language model generates an end symbol and obtains the updated output text.
[0053] Specifically, for each current word generated by the large language model, it is necessary to check whether the current word is a text manipulation semantic word (i.e., a text manipulation semantic token).
[0054] If the current word element is not a semantic word element in text operations, it means that the current word element is a new word element predicted and generated by the large language model based on the user's interactive prompts, historical dialogues, the large language model's retrieval text, and the output text already generated by the large language model in this round of inference. In this case, the current word element can be concatenated to the end of the output text, and the word element generation position of the output text can be updated.
[0055] Assuming the current word is updated before the output text, the word generation position of the current output text is denoted as... If the current word element is appended to the end of the current output text, then the word element generation position of the output text also needs to be updated. The update formula for the word element generation position of the output text is as follows: ; Among them, the " " indicates assignment; the left side of the formula Indicates the lexical generation position of the updated output text; the right side of the formula This indicates the lexical generation position of the output text before the update.
[0056] Furthermore, the large language model continues to maintain the model inference and generation mode, updates the global lexical index information, continues to perform autoregressive decoding to generate the next lexical, and uses the next lexical as the current lexical. Then it returns to the step of detecting whether the current lexical is a text operation semantic lexical, until the large language model generates the end symbol and obtains the updated output text. At this point, the updated output text is the text finally used to reply to the user.
[0057] The large-scale inference acceleration method provided in this embodiment allows the large language model to generate output text through mode switching. That is, when the current word generated by the large language model is not a text operation semantic word, the model inference generation mode is maintained to continuously generate new words. When the current word generated by the large language model is a text operation semantic word, the model enters text editing mode and uses text reuse technology to achieve synchronous copying of multiple words in long texts, thereby improving the copying efficiency of long texts and optimizing the inference acceleration effect of the large language model.
[0058] In some embodiments, text manipulation semantic units include first text copying units, which carry first text copying interval information; based on text manipulation semantic units, copying a target text segment from existing text information includes: performing text position validity verification on the first text copying units and obtaining a verification result; if the verification result is that the first text copying units pass the verification, then determining the first text corresponding to the first text copying interval information from the existing text information and using the first text as the target text segment; determining the number of target units in the target text segment, and assigning unit positions in the output text that match the number of target units to the target text segment.
[0059] In this embodiment, the text manipulation semantic lexical includes two types: the first text copying lexical and the second text copying lexical.
[0060] The first text copying term carries the first text copying interval information, which includes two copying position parameter information.
[0061] Optionally, the first text copying interval information carried by the first text copying term can be denoted as: , Indicates copying. and To copy the location parameter information, This indicates the position of the starting word in the first text copy interval. This indicates the position of the end word of the first text copy interval. The information about the first text copy interval is used to indicate the copy position of the large language model from the end word of the first text copy interval. arrive All lexical units between them.
[0062] Specifically, if the current word is a text operation semantic word and the text operation semantic word is the first text copy word, the large language model will pause autoregressive decoding and enter text editing mode.
[0063] In text editing mode, the large language model will perform text position validity verification on the first text copied word based on the global word index information and the copy position parameter information carried by the first text copied word, and obtain the verification result.
[0064] Optionally, the validity check includes at least one of boundary check, continuity check, and valid segment check.
[0065] Among them, out-of-bounds check refers to checking whether the first text corresponding to the first text copying interval information is out of bounds, for example, checking... or Does the location it points to exceed the text range of existing text information?
[0066] Continuous verification refers to verifying whether the first text corresponding to the first text copy interval information is a continuous text segment.
[0067] Valid fragment verification refers to verifying whether the first text corresponding to the first text copy interval information is valid text. Valid text can be text related to the user's interactive prompts.
[0068] Furthermore, if the verification result shows that the first text copying term passes the verification, the large language model can determine the first text corresponding to the first text copying interval information from the existing text information based on the first text copying interval information carried by the first text copying term (the first text includes information from...). arrive (All lexical units between), and take the first text as the target text segment.
[0069] Furthermore, the number of target words in the target text segment is determined, and consecutive word positions matching the number of target words are assigned to the target text segment in the output text.
[0070] Specifically, assuming the target text segment has the following number of target lexical units: ( (where is a positive integer greater than or equal to 2). Before concatenating the target word into the current output text, the word generation position of the current output text is denoted as . Then, continuous text segments can be assigned to the target text segment in the output text. There are 10 lexical positions, which can be denoted as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 ... .
[0071] Furthermore, multiple target words are concatenated to the end of the current output text, and the word generation positions of the output text are updated.
[0072] In some embodiments, text manipulation semantic units include second text copying units, which carry second text copying interval information and text deletion interval information. Based on the text manipulation semantic units, copying a target text segment from existing text information includes: performing text position validity verification on the second text copying units and obtaining a verification result; if the verification result is that the second text copying units pass the verification, then determining the second text corresponding to the second text copying interval information from the existing text information, and deleting the third text corresponding to the text deletion interval information from the second text to generate a target text segment; determining the number of target units in the target text segment, and assigning unit positions in the output text that match the number of target units to the target text segment.
[0073] In this embodiment, the text manipulation semantic lexical includes two types: the first text copying lexical and the second text copying lexical.
[0074] The second text copying term carries information about the second text copying interval and the text deletion interval. The second text copying interval and the text deletion interval each include two copying position parameter information.
[0075] Optionally, the second text copying interval information and text deletion interval information carried by the second text copying terminus can be denoted as: , Indicates deletion. , , and All of these are copying location parameter information. Indicates the position of the starting word of the second text copy interval. Indicates the position of the last word in the second text copy interval. Indicates the position of the starting word of the text deletion interval. The second text copy interval information indicates the position of the end word of the text deletion interval, and is used to indicate the copy position of the large language model from the end word of the text deletion interval. arrive All lexical units between, text deletion interval information is used to indicate the deletion range from the large language model. arrive All lexical units between them.
[0076] Specifically, if the current word is a text operation semantic word and the text operation semantic word is a second text copy word, the large language model will pause autoregressive decoding and enter text editing mode.
[0077] In text editing mode, the large language model will perform text position validity verification on the second text copying unit based on the global lexical index information and the second text copying interval information and text deletion interval information carried by the second text copying unit, and obtain the verification result.
[0078] Optionally, the validity check includes at least one of boundary check, continuity check, and valid segment check.
[0079] Among them, the out-of-bounds check refers to checking whether the second text corresponding to the second text copying range information and the deleted text corresponding to the text deletion range are out of bounds.
[0080] Continuity verification refers to verifying whether the second text corresponding to the second text copy interval and the deleted text corresponding to the text deletion interval are consecutive texts.
[0081] Valid fragment verification refers to verifying whether the second text corresponding to the second text copy interval information and the deleted text corresponding to the text deletion interval are valid text. Valid text can be text related to user interaction prompts.
[0082] Furthermore, if the verification result shows that the second text copying term passes the verification, the large language model can determine the second text corresponding to the second text copying interval information from the existing text information based on the second text copying interval information carried by the second text copying term. (The second text includes information from...) arrive (All lexical units between), and based on the text deletion interval information carried by the lexical units in the second text, delete the third text corresponding to the text deletion interval information from the second text (the third text includes all ... arrive (All lexical units between) to generate the target text segment.
[0083] Furthermore, the number of target words in the target text segment is determined, and consecutive word positions matching the number of target words are assigned to the target text segment in the output text.
[0084] Specifically, assuming the target text segment has the following number of target lexical units: ( (where is a positive integer greater than or equal to 2). Before concatenating the target word into the current output text, the word generation position of the current output text is denoted as . Then, continuous text segments can be assigned to the target text segment in the output text. There are 10 lexical positions, which can be denoted as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 ... .
[0085] Furthermore, multiple target words are concatenated to the end of the current output text, and the word generation positions of the output text are updated.
[0086] The large model inference acceleration method provided in this embodiment enables the reuse of existing text information through different types of text operation semantic units. For example, it can directly copy a text segment from existing text information, copy and delete a text segment, and support copying and modifying a word in a text segment using other types of text operation semantic units, copying and inserting a new word at a certain position in a text segment, etc. This improves the flexibility of text reuse technology and can adapt to different types of text reuse needs.
[0087] In some embodiments, if the current lexical unit is a text manipulation semantic lexical unit, after copying the target text segment from the existing text information based on the text manipulation semantic lexical unit, the method further includes: calculating the key value information of each target lexical unit based on a large language model; determining the error between each key value information and the historical key value information; the historical key value information is the key value information of historical lexical units generated by the large language model, and the historical key value information is cached in the key value cache; selecting several target lexical units with the largest errors from all target lexical units as lexical units to be corrected; recalculating the key value of each lexical unit to be corrected to obtain the corrected key value information of each lexical unit to be corrected; and updating the corrected key value information of each lexical unit to be corrected to the end of the key value cache.
[0088] Specifically, after copying the target text segment from the existing text information, the target text segment containing multiple target words can be input into the large language model for forward computation. The large language model can use two attention layers to calculate the key-value information (KV value) of each target word.
[0089] For each target word, its key-value information is the core intermediate result calculated by the two attention layers in the Transformer network structure of the large language model. It mainly includes the key value and the value value. The key value is mainly used to measure the relevance of the target word to other words. The value value corresponds to the content feature vector of the target word and represents the information carrier of the final output.
[0090] Large language models typically employ caching mechanisms, often including a key-value cache (KV cache) that stores historical key-value information of historical lexical units generated by the large language model.
[0091] Furthermore, the error between the key value information and the historical key value information of each target word is calculated, and then the target words with the largest error are selected from all target words as words to be corrected.
[0092] For example, select the target words with the largest errors (accounting for 20%) from all target words and use them as words to be corrected.
[0093] Furthermore, for each lexical unit to be corrected, key-value recalculation (i.e., KV recalculation) is performed on the lexical unit to be corrected, and its value is corrected, thereby obtaining the corrected key-value information of the lexical unit.
[0094] Furthermore, the corrected key-value information of each lexical unit to be corrected is updated to the end of the key-value cache.
[0095] For the remaining target words with small errors, since their key-value information is not much different from the historical key-value information, it means that the historical key-value information is similar to or even the same as the key-value information of the target word. When the large language model needs to use the key-value information of the target word, it can directly reuse the similar or even the same historical key-value information in the key-value cache.
[0096] The key-value information cached in the key-value cache can be reused within the model when the large language model performs autoregressive decoding (for example, the attention layer does not need to recalculate the key-value information of existing lexical units, but can directly read it from the key-value cache).
[0097] The large-scale inference acceleration method provided in this embodiment addresses scenarios where existing text segments are frequently reused or modified during large language model inference. It proposes an inference acceleration scheme based on explicit position pointers, precise text manipulation, key-value recalculation, and word generation position recoding. Instead of copying text segments word by word, the large language model directly generates text manipulation semantic words. Based on the parsed text manipulation semantic words, it achieves multi-word synchronous copying of long texts and performs selective key-value recalculation on the copied target words to update the key-value cache. This improves text copying efficiency, optimizes the inference acceleration effect of the large language model, and ensures the stability and reliability of the large language model inference.
[0098] In some embodiments, before inputting the user's interaction prompts and the user's historical dialogues into the large language model and performing autoregressive decoding to obtain the current lexical unit generated by the large language model, the method further includes: constructing text manipulation semantic lexical units and adding the text manipulation semantic lexical units to the model vocabulary of the initial large language model; obtaining a text manipulation training dataset; and fine-tuning the initial large language model based on the text manipulation training dataset to obtain the large language model; wherein the large language model has the ability to use text manipulation semantic lexical units.
[0099] Understandably, such as Figure 2 As shown, before inputting the user's interaction prompts and historical dialogues into the large language model for autoregressive decoding, the large language model needs to be fine-tuned first.
[0100] Specifically, text manipulation semantic lexical units are first constructed, including the first text copying lexical unit and the second text copying lexical unit, and these two text manipulation semantic lexical units are added to the model vocabulary of the initial large language model (the model vocabulary is denoted as tokenizer).
[0101] After adding text manipulation semantic lexicons to the model vocabulary of the initial large language model, for the initial large language model, the text manipulation semantic lexicons are not natural language, nor are they instruction formats, but rather the original and understandable semantic lexicons of the initial large language model.
[0102] Furthermore, a text manipulation training dataset is obtained, and the initial large language model is fine-tuned using the text manipulation training dataset. This trains the initial large language model to use text manipulation semantic units and updates the model weights within the initial large language model, thereby obtaining a fine-tuned large language model.
[0103] Among them, the large language model has the ability to flexibly use text manipulation semantic units.
[0104] This invention also provides a large-scale model inference acceleration device. Please refer to [link / reference]. Figure 3 , Figure 3 This is a schematic diagram of the structure of the large model inference acceleration device provided by the present invention. In this embodiment, the large model inference acceleration device includes an input module 310, a detection module 320, a copying module 330, an update module 340, and an output module 350.
[0105] The input module 310 is used to input the user's interactive prompts and the user's historical dialogues into the large language model, perform autoregressive decoding, and obtain the current word unit generated by the large language model.
[0106] The detection module 320 is used to detect whether the current word is a text operation semantic word.
[0107] Text manipulation semantic lexical units are lexical units that instruct large language models to perform text copying.
[0108] The copy module 330 is used to copy the target text segment from existing text information based on the text operation semantic word if the current word is a text operation semantic word.
[0109] The existing text information includes historical dialogues and retrieval text from a large language model, and the target text segment includes multiple target lexical units.
[0110] Update module 340 is used to update multiple target terms to the end of the output text.
[0111] The output module 350 is used to continue autoregressive decoding based on the large language model, generate the next word, take the next word as the current word, and return to the step of detecting whether the current word is a semantic word for text operation, until the large language model generates the end symbol and obtains the updated output text.
[0112] In some embodiments, after detecting whether the current word is a text manipulation semantic word, the method further includes: if the current word is not a text manipulation semantic word, updating the current word to the end of the output text; based on the large language model, continuing autoregressive decoding to generate the next word, using the next word as the current word, and returning to the step of detecting whether the current word is a text manipulation semantic word, until the large language model generates an end symbol and obtains the updated output text.
[0113] In some embodiments, text manipulation semantic units include first text copying units, which carry first text copying interval information; based on text manipulation semantic units, copying a target text segment from existing text information includes: performing text position validity verification on the first text copying units and obtaining a verification result; if the verification result is that the first text copying units pass the verification, then determining the first text corresponding to the first text copying interval information from the existing text information and using the first text as the target text segment; determining the number of target units in the target text segment, and assigning unit positions in the output text that match the number of target units to the target text segment.
[0114] In some embodiments, text manipulation semantic units include second text copying units, which carry second text copying interval information and text deletion interval information. Based on the text manipulation semantic units, copying a target text segment from existing text information includes: performing text position validity verification on the second text copying units and obtaining a verification result; if the verification result is that the second text copying units pass the verification, then determining the second text corresponding to the second text copying interval information from the existing text information, and deleting the third text corresponding to the text deletion interval information from the second text to generate a target text segment; determining the number of target units in the target text segment, and assigning unit positions in the output text that match the number of target units to the target text segment.
[0115] In some embodiments, if the current lexical unit is a text manipulation semantic lexical unit, after copying the target text segment from the existing text information based on the text manipulation semantic lexical unit, the method further includes: calculating the key value information of each target lexical unit based on a large language model; determining the error between each key value information and the historical key value information; the historical key value information is the key value information of historical lexical units generated by the large language model, and the historical key value information is cached in the key value cache; selecting several target lexical units with the largest errors from all target lexical units as lexical units to be corrected; recalculating the key value of each lexical unit to be corrected to obtain the corrected key value information of each lexical unit to be corrected; and updating the corrected key value information of each lexical unit to be corrected to the end of the key value cache.
[0116] In some embodiments, before inputting the user's interaction prompts and the user's historical dialogues into the large language model and performing autoregressive decoding to obtain the current lexical unit generated by the large language model, the method further includes: constructing text manipulation semantic lexical units and adding the text manipulation semantic lexical units to the model vocabulary of the initial large language model; obtaining a text manipulation training dataset; and fine-tuning the initial large language model based on the text manipulation training dataset to obtain the large language model; wherein the large language model has the ability to use text manipulation semantic lexical units.
[0117] The present invention also provides an electronic device. Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 4As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a large-model inference acceleration method. The large-model inference acceleration method includes: inputting the user's interactive prompts and the user's historical dialogue into a large language model, performing autoregressive decoding to obtain the current word element generated by the large language model; detecting whether the current word element is a text manipulation semantic word element; a text manipulation semantic word element is a word element that instructs the large language model to copy text; if the current word element is a text manipulation semantic word element, then copying the target text segment from the existing text information based on the text manipulation semantic word element; the existing text information includes historical dialogue and the retrieval text of the large language model, and the target text segment includes multiple target words; updating the multiple target words to the end of the output text; based on the large language model, continuing autoregressive decoding to generate the next word element, taking the next word element as the current word element, and returning to the step of detecting whether the current word element is a text manipulation semantic word element, until the large language model generates an end symbol and obtains the updated output text.
[0118] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0119] This invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the large-model inference acceleration method provided by the above methods. The large-model inference acceleration method includes: inputting the user's interactive prompts and the user's historical dialogue into a large language model, performing autoregressive decoding to obtain the current word element generated by the large language model; detecting whether the current word element is a text manipulation semantic word element; the text manipulation semantic word element is a word element that instructs the large language model to copy text; if the current word element is a text manipulation semantic word element, then copying the target text segment from existing text information based on the text manipulation semantic word element; the existing text information includes historical dialogue and the search text of the large language model, and the target text segment includes multiple target words; updating the multiple target words to the end of the output text; based on the large language model, continuing to perform autoregressive decoding to generate the next word element, taking the next word element as the current word element, and returning to the step of detecting whether the current word element is a text manipulation semantic word element, until the large language model generates an end symbol to obtain the updated output text.
[0120] This invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the large model inference acceleration method provided by the above methods. The large model inference acceleration method includes: inputting the user's interactive prompts and the user's historical dialogue into a large language model, performing autoregressive decoding to obtain the current word element generated by the large language model; detecting whether the current word element is a text manipulation semantic word element; the text manipulation semantic word element is a word element that instructs the large language model to copy text; if the current word element is a text manipulation semantic word element, then copying the target text segment from the existing text information based on the text manipulation semantic word element; the existing text information includes historical dialogue and the search text of the large language model, and the target text segment includes multiple target words; updating the multiple target words to the end of the output text; based on the large language model, continuing to perform autoregressive decoding to generate the next word element, taking the next word element as the current word element, and returning to the step of detecting whether the current word element is a text manipulation semantic word element, until the large language model generates an end symbol to obtain the updated output text.
[0121] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0122] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for accelerating large-scale model inference, characterized in that, include: The user's interactive prompts and the user's historical dialogues are input into a large language model, and autoregressive decoding is performed to obtain the current word unit generated by the large language model. Detect whether the current word is a text manipulation semantic word; the text manipulation semantic word is a word that instructs the large language model to perform text copying; If the current word element is the text operation semantic word element, then based on the text operation semantic word element, the target text segment is copied from the existing text information; the existing text information includes the historical dialogue and the retrieval text of the large language model, and the target text segment includes multiple target words; Update the output text with multiple target words; Based on the large language model, autoregressive decoding continues to generate the next lexical unit. The next lexical unit is used as the current lexical unit, and the process returns to the step of detecting whether the current lexical unit is a text operation semantic lexical unit, until the large language model generates an end symbol and the updated output text is obtained.
2. The method for accelerating large model inference according to claim 1, characterized in that, After detecting whether the current lexical unit is a text operation semantic lexical unit, the method further includes: If the current word is not a semantic word of the text operation, then the current word is updated to the end of the output text; Based on the large language model, autoregressive decoding continues to generate the next lexical unit. The next lexical unit is then used as the current lexical unit, and the process returns to the step of detecting whether the current lexical unit is a text operation semantic lexical unit, until the large language model generates the end symbol and the updated output text is obtained.
3. The method for accelerating large model inference according to claim 1, characterized in that, The text operation semantic lexical includes a first text copy lexical, which carries first text copy interval information; The step of copying the target text segment from existing text information based on the semantic units of the text manipulation includes: Perform text position legality verification on the first text copied word and obtain the verification result; If the verification result is that the first text copying term passes the verification, then the first text corresponding to the first text copying interval information is determined from the existing text information, and the first text is used as the target text segment; Determine the number of target words in the target text segment, and assign word positions in the output text that match the number of target words in the target text segment.
4. The method for accelerating large model inference according to claim 1, characterized in that, The text operation semantic units include the second text copy unit, which carries the second text copy interval information and the text deletion interval information; The step of copying the target text segment from existing text information based on the semantic units of the text manipulation includes: Perform text position validity verification on the second text copying element and obtain the verification result; If the verification result is that the second text copying term passes the verification, then the second text corresponding to the second text copying interval information is determined from the existing text information, and the third text corresponding to the text deletion interval information is deleted from the second text to generate the target text segment; Determine the number of target words in the target text segment, and assign word positions in the output text that match the number of target words in the target text segment.
5. The method for accelerating large model inference according to claim 1, characterized in that, If the current lexical unit is the text operation semantic lexical unit, then after copying the target text segment from the existing text information based on the text operation semantic lexical unit, the method further includes: Based on the large language model, calculate the key value information of each target word; The error between each key value and the historical key value is determined; the historical key value is the key value of the historical lexical units generated by the large language model, and the historical key value is cached in the key value cache. Select the target words with the largest errors from all the target words and use them as words to be corrected. For each of the lexical units to be corrected, the key value is recalculated to obtain the corrected key value information of each of the lexical units to be corrected. The corrected key-value information of each of the lexical terms to be corrected is updated to the end of the key-value cache.
6. The method for accelerating large model inference according to claim 1, characterized in that, Before inputting the user's interactive prompts and the user's historical dialogues into the large language model for autoregressive decoding to obtain the current lexical unit generated by the large language model, the method further includes: Construct the text operation semantic lexicon and add the text operation semantic lexicon to the model vocabulary of the initial large language model; Obtain the text manipulation training dataset; Based on the text manipulation training dataset, the initial large language model is fine-tuned and trained to obtain the large language model. The large language model has the ability to manipulate semantic lexical units using the text.
7. A large-scale model inference acceleration device, characterized in that, include: The input module is used to input the user's interactive prompts and the user's historical dialogues into the large language model, perform autoregressive decoding, and obtain the current word unit generated by the large language model; The detection module is used to detect whether the current word is a text manipulation semantic word; the text manipulation semantic word is a word that instructs the large language model to perform text copying; The copy module is used to copy a target text segment from existing text information based on the text operation semantic lexicon if the current lexicon is the text operation semantic lexicon; the existing text information includes the historical dialogue and the retrieval text of the large language model, and the target text segment includes multiple target lexicons; An update module is used to update multiple target words to the end of the output text; The output module is used to continue autoregressive decoding based on the large language model, generate the next word, take the next word as the current word, and return to the step of detecting whether the current word is a text operation semantic word, until the large language model generates an end symbol to obtain the updated output text.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the large model inference acceleration method as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the large model inference acceleration method as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the large model inference acceleration method as described in any one of claims 1 to 6.