Method for quickly generating token for large language model at end side

By introducing lightweight self-speculation decoding modules into large language models, building and verifying tree structures are solved, and the inference delay and token quality of the end-side large language models are achieved quickly generate high-quality tokens.

CN120012932APending Publication Date: 2025-05-16BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510094546.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the end-side application scenario, the reasoning delay of the large language model increases, resulting in a decline in user experience. Traditional methods use small-parameter draft models to generate tokens but it is difficult to quickly generate high-quality tokens.

Method used

By connecting the lightweight self-speculation decoding module on the adapter side in a large language model, the initial token is generated using the large language model, and candidate tokens are generated through the self-speculation decoding module, a tree structure is built and verified through the large language model, and finally iteratively generate high-quality tokens.

Benefits of technology

It realizes the rapid generation of high-quality series of tokens on the end side, improves the user experience, and solves the problems of inference delay and token quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012932A_ABST
    Figure CN120012932A_ABST
Patent Text Reader

Abstract

The invention relates to a method for quickly generating a token for an end-side large language model, and belongs to the technical field of large language models. Inputting the input text into a pre-trained rapid token generation model, wherein the rapid token generation model comprises a large language model and a self-speculation decoding module; the large language model generates a hidden state vector according to an input text and generates an initial token according to the hidden state vector, and the self-speculation decoding module generates a plurality of candidate tokens according to the hidden state vector and constructs a tree structure according to the initial token and the candidate tokens; verifying each path in the tree structure through a large language model; and the large language model updates a hidden state vector according to a verification result and generates a new initial token according to the new hidden state vector, and the self-speculation decoding module generates a new candidate token according to the new hidden state vector, so that loop iteration is performed until a termination condition is met, and the token in a path with a qualified verification result is taken as a final output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language models, and in particular to a method for quickly generating tokens for a large language model on a terminal side. Background Art

[0002] The latest development of large language models shows that the expansion of model scale has significantly improved the quality of language generation, but this is accompanied by an increase in inference latency, which brings considerable challenges in practical applications, especially in end-side application scenarios. Due to the limitation of the computing power of the end-side chip, it is difficult to match the inference speed of the GPU card. In order to ensure that users can enjoy fast response when using large models and avoid a decline in experience, it is particularly critical to increase the speed of token generation by large language models.

[0003] The traditional method of quickly generating tokens for a large language model on the end side usually uses a draft model with a small number of parameters to generate a series of tokens, and then verifies the series of tokens through the large language model; however, an inappropriate draft model often leads to a decrease in the speculative hit rate and there is a technical problem that high-quality tokens cannot be generated quickly. Summary of the invention

[0004] In view of the above analysis, an embodiment of the present invention aims to provide a method for quickly generating tokens for a large language model on the terminal side, so as to solve one or more of the above problems existing in the prior art.

[0005] The object of the present invention is achieved in that:

[0006] The first aspect of the present invention provides a method for quickly generating tokens for a large language model on a terminal side, including:

[0007] S1. Get input text;

[0008] S2, inputting the input text into a pre-trained rapid token generation model, wherein the rapid token generation model includes a large language model and a self-speculative decoding module;

[0009] S3, the large language model generates a hidden state vector according to the input text, generates an initial token according to the hidden state vector, the self-speculative decoding module generates a plurality of candidate tokens according to the hidden state vector, and constructs a tree structure according to the initial token and the candidate tokens;

[0010] S4, verifying each path in the tree structure by using the large language model;

[0011] S5. The large language model updates the hidden state vector according to the verification result, and generates a new initial token according to the new hidden state vector. The self-speculative decoding module generates a new candidate token according to the new hidden state vector, and iterates S3-S4 in a loop until the termination condition is reached, and the token in the path with qualified verification results is used as the final output.

[0012] Furthermore, the large language model generates a hidden state vector according to the input text, and generates an initial token according to the hidden state vector, including: performing word segmentation and vectorization processing on the input text to obtain word embedding; processing the word embedding through a Transformer layer to obtain a hidden state vector; mapping the hidden state vector to a token through a word list layer, calculating the probability of each token, and outputting the token with the highest probability as the initial token.

[0013] Furthermore, the self-speculative decoding module generates multiple candidate tokens according to the hidden state vector, including: inputting the hidden state vector into the self-speculative decoding module, the self-speculative decoding module includes n different parallel-connected perceptrons and vocabulary layers, and generating n layers of candidate tokens by using a top K algorithm.

[0014] Furthermore, constructing a tree structure according to the initial token and the candidate tokens includes: constructing a tree structure with the initial token as a root node and the candidate tokens generated from the word list layer in the self-speculative decoding module as leaf nodes.

[0015] Furthermore, the verification of each path in the tree structure by the large language model includes: inputting each path in the tree structure into the large language model for forward propagation; determining whether the token in each path is the token with the highest probability in the current propagation process, eliminating the paths that do not meet the judgment conditions, and screening out the optimal path from the paths that meet the judgment conditions.

[0016] Furthermore, the termination condition includes that the optimal path reaches a specified token length or all paths do not meet the judgment condition.

[0017] Furthermore, the pre-training step of the rapid token generation model includes:

[0018] A training data set is collected, the weight matrix in the self-speculative decoding module is processed by the singular value decomposition method, the training data set is input into the fast token generation model, the self-distillation method is used, the KL divergence is used as the loss function, and the weight matrix of the self-speculative decoding module is updated by back propagation to iteratively train the fast token generation model.

[0019] A second aspect of the present invention provides a device for quickly generating a token, comprising:

[0020] The acquisition module is used to obtain the input text;

[0021] An input module, used for inputting the input text into a pre-trained rapid token generation model, wherein the rapid token generation model includes a large language model and a self-speculative decoding module;

[0022] A processing module, used for the large language model to generate a hidden state vector according to the input text, generate an initial token according to the hidden state vector, the self-speculative decoding module to generate a plurality of candidate tokens according to the hidden state vector, and construct a tree structure according to the initial token and the candidate tokens;

[0023] A verification module, used for verifying each path in the tree structure by using the large language model;

[0024] The iterative output module is used for the large language model to update the hidden state vector according to the verification results, and to generate a new initial token according to the new hidden state vector. The self-speculative decoding module generates a new candidate token according to the new hidden state vector, and iterates S3-S4 in a loop until the termination condition is reached, and the token in the path with qualified verification results is used as the final output.

[0025] According to a third aspect of the present invention, an electronic device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the method for quickly generating tokens for a large language model on the terminal side as described in any embodiment is implemented.

[0026] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for quickly generating tokens for a large language model on a terminal side as described in any one of the embodiments is implemented.

[0027] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:

[0028] The method for quickly generating tokens for a large language model on the end side provided by the present invention is to connect a lightweight self-speculative decoding module on the end side in series with the large language model, generate an initial token according to the input by the large language model, and then generate a candidate token according to the hidden state vector of the large language model by the self-speculative decoding module, build a tree structure according to the initial token and the candidate tokens, verify the tree structure by the large language model, and perform an iterative cycle, so as to quickly generate a series of high-quality tokens. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this specification. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0030] Figure 1 A flowchart of a method for quickly generating tokens for a large language model on a terminal side provided in Embodiment 1 of the present invention;

[0031] Figure 2 A schematic diagram of a device for quickly generating tokens provided in Example 2 of the present invention;

[0032] Figure 3 This is a schematic diagram of the electronic device architecture provided in Example 3 of the present invention. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. It should be noted that, in the absence of conflict, the embodiments and features in the embodiments disclosed in this disclosure can be combined, separated, interchanged and / or rearranged with each other. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0034] Example 1

[0035] A specific embodiment of the present invention, as Figure 1 As shown, a method for quickly generating tokens for a large language model on the client side is disclosed, including the following steps:

[0036] S1. Get input text.

[0037] S2. Input the input text into a pre-trained fast token generation model, wherein the fast token generation model includes a large language model and a self-speculative decoding module.

[0038] In this embodiment, the pre-training step of the rapid token generation model includes:

[0039] A training data set is collected, the weight matrix in the self-speculative decoding module is processed by the singular value decomposition method, the training data set is input into the fast token generation model, the self-distillation method is used, the KL divergence is used as the loss function, and the weight matrix of the self-speculative decoding module is updated by back propagation to iteratively train the fast token generation model.

[0040] Specifically, during the pre-training process, the singular value decomposition method is used to reduce the dimension of the weight matrix in the self-speculative decoding module, thereby reducing the number of parameters, so that the large language model can accelerate the generation of tokens in the inference stage. The self-distillation method is used to enable the self-speculative decoding module to effectively learn the behavior of the large language model and improve the quality of token generation. The weight matrix of the self-speculative decoding module is adjusted through back propagation so that the distribution of tokens generated by it gradually approaches the large language model. The self-speculative decoding module obtained through the above training is lightweight and has excellent characteristics for adapting to the large language model on the end.

[0041] S3. The large language model generates a hidden state vector according to the input text, generates an initial token according to the hidden state vector, the self-speculative decoding module generates multiple candidate tokens according to the hidden state vector, and constructs a tree structure according to the initial token and the candidate tokens.

[0042] In this embodiment, step S3 specifically includes:

[0043] S301, performing word segmentation and vectorization processing on the input text to obtain word embedding;

[0044] Exemplarily, the BPE word segmentation tool and word2vec can be used to process the input text into a word embedding form, for example, the input text is "My favorite fruit is".

[0045] S302, processing the word embedding through a Transformer layer to obtain a hidden state vector;

[0046] Specifically, the word embedding vector is first positionally encoded and then processed by multiple Transformer layers to generate a context-dependent hidden state vector through the self-attention mechanism and feedforward network.

[0047] S303, mapping the hidden state vector into a token through a word list layer, calculating the probability of each token, and outputting the token with the highest probability as the initial token;

[0048] Specifically, the vocabulary layer of the large language model maps the hidden state vector to tokens through the vocabulary matrix, calculates the score of the token, converts the score into probability through the Softmax function, and the token with the highest probability is output as the initial token, such as apple.

[0049] S304, inputting the hidden state vector into a self-speculative decoding module, wherein the self-speculative decoding module includes n different parallel connected perceptrons and word list layers, and generates n layers of candidate tokens by using a top K algorithm;

[0050] Specifically, the vocabulary layer takes the last hidden state vector of the large language model as input, generates multiple possible subsequent tokens and corresponding probabilities through the perceptron and the vocabulary matrix, and then uses the top K algorithm to select K subsequent tokens as candidate tokens based on the probabilities.

[0051] S305 , constructing a tree structure with the initial token as the root node and the candidate tokens generated from the vocabulary layer in the speculative decoding module as leaf nodes.

[0052] S4. Verify each path in the tree structure using the large language model.

[0053] In this embodiment, step S4 specifically includes:

[0054] S401, inputting each path in the tree structure into a large language model for forward propagation;

[0055] Specifically, each path in the tree structure is concatenated with the input text and then input into the large language model for forward propagation.

[0056] S402, determining whether the token in each path is the token with the highest probability in the current propagation process, and eliminating the paths that do not meet the determination conditions;

[0057] Specifically, as shown in step S303, the token with the highest probability outputted by the vocabulary layer is used to determine whether each path in the tree structure is reasonable.

[0058] S403: Filter out the optimal path from the paths that meet the judgment conditions.

[0059] Specifically, for all paths that meet the judgment criteria, the longest path is selected as the optimal path.

[0060] S5. The big language updates the hidden state vector according to the verification result, and generates a new initial token according to the new hidden state vector. The self-speculative decoding module generates a new candidate token according to the new hidden state vector, and iterates S3-S4 in a loop until the termination condition is reached, and the token in the path with qualified verification results is used as the final output.

[0061] Specifically, the large language model updates the hidden state vector according to the last token in the optimal path to generate a new initial token, and inputs the new hidden state vector into the self-speculative decoding module to generate a new candidate token, updates the tree structure according to the new initial token and candidate tokens, and continues to verify the path in the new tree structure. This cycle is iterated to generate and output a complete text sequence.

[0062] Exemplarily, the termination condition includes that the optimal path reaches a specified token length or all paths do not meet the judgment condition.

[0063] Compared with the prior art, the method for quickly generating tokens for a large language model on the terminal side provided in this embodiment is to connect a lightweight self-speculative decoding module adapted to the terminal side in series in the large language model, generate an initial token according to the input through the large language model, and then generate a candidate token according to the hidden state vector of the large language model by the self-speculative decoding module, build a tree structure according to the initial token and the candidate tokens, verify the tree structure by the large language model, and perform an iterative loop, so as to quickly generate a series of high-quality tokens.

[0064] Example 2

[0065] This embodiment provides a device for quickly generating tokens, such as Figure 2 As shown, including:

[0066] The acquisition module is used to obtain the input text;

[0067] An input module, used for inputting the input text into a pre-trained rapid token generation model, wherein the rapid token generation model includes a large language model and a self-speculative decoding module;

[0068] A processing module, used for the large language model to generate a hidden state vector according to the input text, generate an initial token according to the hidden state vector, the self-speculative decoding module to generate a plurality of candidate tokens according to the hidden state vector, and construct a tree structure according to the initial token and the candidate tokens;

[0069] A verification module, used for verifying each path in the tree structure by using the large language model;

[0070] The iterative output module is used for the large language model to update the hidden state vector according to the verification results, and to generate a new initial token according to the new hidden state vector. The self-speculative decoding module generates a new candidate token according to the new hidden state vector, and iterates S3-S4 in a loop until the termination condition is reached, and the token in the path with qualified verification results is used as the final output.

[0071] Example 3

[0072] This embodiment provides an electronic device, such as Figure 3 As shown, it includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the method for quickly generating tokens for a large language model on the terminal side as described in any of the above embodiments is implemented.

[0073] Example 4

[0074] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the method for quickly generating tokens for a large language model on the terminal side as described in any of the above embodiments is implemented.

[0075] Computer readable storage media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0076] The professionals should further realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to the function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0077] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0078] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for quickly generating tokens for a large language model on a terminal side, characterized in that: include: S1. Get input text; S2, inputting the input text into a pre-trained rapid token generation model, wherein the rapid token generation model includes a large language model and a self-speculative decoding module; S3, the large language model generates a hidden state vector according to the input text, generates an initial token according to the hidden state vector, the self-speculative decoding module generates a plurality of candidate tokens according to the hidden state vector, and constructs a tree structure according to the initial token and the candidate tokens; S4, verifying each path in the tree structure by using the large language model; S5. The large language model updates the hidden state vector according to the verification result, and generates a new initial token according to the new hidden state vector. The self-speculative decoding module generates a new candidate token according to the new hidden state vector, and iterates S3-S4 in a loop until the termination condition is reached, and the token in the path with qualified verification results is used as the final output.

2. The method for quickly generating tokens for a large language model on the terminal side according to claim 1, characterized in that: The large language model generates a hidden state vector according to the input text, and generates an initial token according to the hidden state vector, including: Performing word segmentation and vectorization processing on the input text to obtain word embedding; Processing the word embedding through a Transformer layer to obtain a hidden state vector; The hidden state vector is mapped into a token through a vocabulary layer, and the probability of each token is calculated, and the token with the highest probability is output as the initial token.

3. According to the method for quickly generating tokens for a large language model on the end side of claim 2, the self-speculative decoding module generates a plurality of candidate tokens according to the hidden state vector, comprising: The hidden state vector is input into a self-speculative decoding module, which includes n different parallel-connected perceptrons and vocabulary layers, and generates n layers of candidate tokens by using a top K algorithm.

4. The method for quickly generating tokens for a large language model on the terminal side according to claim 3, characterized in that: Building a tree structure according to the initial token and the candidate tokens, including: A tree structure is constructed with the initial token as the root node and the candidate tokens generated from the word list layer in the self-speculative decoding module as leaf nodes.

5. The method for quickly generating tokens for a large language model on the terminal side according to claim 1, characterized in that: The verifying each path in the tree structure by using the large language model includes: Input each path in the tree structure into a large language model for forward propagation; Determine whether the token in each path is the token with the highest probability in the current propagation process, eliminate the paths that do not meet the judgment conditions, and select the optimal path from the paths that meet the judgment conditions.

6. The method for quickly generating tokens for a large language model on the terminal side according to claim 5, characterized in that: The termination condition includes that the optimal path reaches a specified token length or all paths do not meet the judgment condition.

7. The method for quickly generating tokens for a large language model on the terminal side according to any one of claim 1, characterized in that: The pre-training steps of the fast token generation model include: A training data set is collected, the weight matrix in the self-speculative decoding module is processed by the singular value decomposition method, the training data set is input into the fast token generation model, the self-distillation method is used, the KL divergence is used as the loss function, and the weight matrix of the self-speculative decoding module is updated by back propagation to iteratively train the fast token generation model.

8. A device for quickly generating tokens, characterized in that: The device comprises: The acquisition module is used to obtain the input text; An input module, used for inputting the input text into a pre-trained rapid token generation model, wherein the rapid token generation model includes a large language model and a self-speculative decoding module; A processing module, used for the large language model to generate a hidden state vector according to the input text, generate an initial token according to the hidden state vector, the self-speculative decoding module to generate a plurality of candidate tokens according to the hidden state vector, and construct a tree structure according to the initial token and the candidate tokens; A verification module, used to verify each path in the tree structure by using the large language model; The iterative output module is used for the large language to update the hidden state vector according to the verification result, and generate a new initial token according to the new hidden state vector. The self-speculative decoding module generates a new candidate token according to the new hidden state vector, and iterates S3-S4 in a loop until the termination condition is reached, and the token in the path with qualified verification results is used as the final output.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the method for quickly generating tokens for a large language model on the terminal side as described in any one of claims 1 to 7 is implemented.

10. A storage medium, characterized in that: A computer program is stored thereon, and when the program is executed by a processor, the method for quickly generating tokens for a large language model on the terminal side as described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Prospective decoding method and system based on static and dynamic word list collaboration

    CN121365746A

  • A speculative decoding method and system based on static and dynamic vocabulary cooperation

    CN121365746B