LLM speculation decoding optimization method and system for downstream tasks

By building a heterogeneous draft model pool and task classification mechanism, selecting the optimal draft model for speculative decoding optimization, the problem of low accuracy of draft model on general tasks and low inference efficiency of large-scale language models is solved, and efficient and accurate multi-task reasoning is achieved.

CN120031128AActive Publication Date: 2025-05-23BEIJING NORMAL UNIVERSITY

Patent Information

Application Number
CN202510098098.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-23
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

In the prior art, the draft model is difficult to obtain a higher accuracy rate on general tasks due to the limited number of parameters, and the large-scale pre-trained language model has low inference efficiency under different tasks, resulting in wasted computing resources.

Method used

Build multiple aligned data sets for different tasks, use these data sets to build a pool of heterogeneous draft model, and select the optimal draft model for speculative decoding and optimization through the task classification mechanism.

Benefits of technology

By combining the task classification model and the draft model pool, efficient inference is achieved in a multi-task environment, significantly improving the inference performance and efficiency of large-scale pre-trained language models under different tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031128A_ABST
    Figure CN120031128A_ABST
Patent Text Reader

Abstract

The invention discloses an LLM speculation decoding optimization method and system oriented to downstream tasks, and the method comprises the following steps: constructing a plurality of alignment data sets oriented to different tasks, and constructing a heterogeneous draft model pool through the alignment data sets; obtaining a cue word text of the downstream task, and classifying the input task based on the cue word text by using a task classification mechanism to obtain a classification result; selecting an optimal draft model from the heterogeneous draft model pool according to a classification result; and generating a guess tokens sequence by using the selected draft model, and inputting the guess tokens sequence into the target model for parallel verification to complete speculation decoding optimization. According to the method, the reasoning performance of a large-scale pre-training language model under different tasks is improved, and the adaptability of the model to diversified tasks is enhanced through combination of fine tuning and task classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence large models, and in particular relates to an LLM speculative decoding optimization method and system for downstream tasks. Background Art

[0002] The speculative decoding optimization method uses a small-scale scratch model to perform multiple autoregressive decodings in succession, which is then verified in parallel by a large language model (LLM). It is one of the mainstream methods to accelerate LLM reasoning. The decoding efficiency and guessing accuracy of the scratch model are two important factors that affect the effect of speculative decoding optimization. Existing work mainly improves the effect of speculative decoding methods on general tasks by improving the scratch model guessing or LLM parallel verification mechanism. However, the scratch model has a limited number of parameters, and it is difficult to obtain a high accuracy on general tasks.

[0003] The scaling law shows that the perplexity (PPL) of a language model scales power-law with the model size, dataset size, and training computation size. Currently, the number of parameters of mainstream LLMs reaches tens of billions or even hundreds of billions, and depending on the user input, it may be necessary to process a long context, which causes its reasoning process to consume a lot of memory and computing resources. For example, when the length of the input prompt is 512, the GPT-2 model with 700M parameters takes 10s to generate 512 tokens on an RTX2080ti. The number of floating-point operations (FLOPS) or integer operations (OPS) during LLM reasoning is proportional to the square of the sequence length. Therefore, accelerated optimization for LLM reasoning is of great significance. Summary of the invention

[0004] The present invention aims to solve the deficiencies of the prior art and provides the following solutions:

[0005] A LLM speculative decoding optimization method for downstream tasks includes the following steps:

[0006] Constructing multiple aligned datasets for different tasks, and using the aligned datasets to construct a heterogeneous draft model pool;

[0007] Obtain prompt word texts of downstream tasks, and classify input tasks based on the prompt word texts using a task classification mechanism to obtain classification results;

[0008] Selecting the best draft model from the heterogeneous draft model pool according to the classification result;

[0009] The selected draft model is used to generate a guessed tokens sequence, and the guessed tokens sequence is input into the target model for parallel verification to complete the speculative decoding optimization.

[0010] Preferably, the method for constructing the alignment data set comprises:

[0011] Get a task set containing different tasks;

[0012] Extracting a set of prompt words from a relevant dataset of each task in the task set;

[0013] Using the target model to infer the prompt word to obtain a corresponding output token sequence;

[0014] The prompt words and the output tokens sequence are used as the alignment data set.

[0015] Preferably, the method for constructing the heterogeneous draft model pool comprises:

[0016] For each task in the task set, based on an existing small-scale language model, fine-tune the small-scale language model using the aligned dataset to obtain a draft model applicable to all tasks;

[0017] The draft models corresponding to all the tasks in the task set are integrated to obtain the heterogeneous draft model pool.

[0018] Preferably, the method for obtaining the classification result includes:

[0019] Obtaining an input prompt word text, and deleting stop words in the prompt word text to obtain a processed text;

[0020] Constructing a training data set for all tasks in the task set, wherein the training data set is composed of the prompt word texts belonging to different tasks, and a correct task category is annotated for each prompt word text, and using the training data set to train a public language text classification model to obtain a fast classification model;

[0021] The processed text is input into the fast classification model for task classification to obtain the classification result.

[0022] The present invention also provides an LLM speculative decoding optimization system for downstream tasks, the system applying any of the above methods, including: a model pool construction module, a task classification module, a model selection module and a decoding optimization module;

[0023] The heterogeneous model pool construction module is used to construct multiple aligned data sets for different tasks, and use the aligned data sets to construct a heterogeneous draft model pool;

[0024] The task classification module is used to obtain the prompt word text of the downstream task, and classify the input task based on the prompt word text using the task classification mechanism to obtain the classification result;

[0025] The model selection module is used to select the best draft model from the heterogeneous draft model pool according to the classification result;

[0026] The decoding optimization module generates a guessed tokens sequence using the selected draft model, and inputs the guessed tokens sequence into the target model for parallel verification to complete speculative decoding optimization.

[0027] Preferably, in the heterogeneous model pool construction module, the method for constructing the aligned data set includes:

[0028] Get a task set containing different tasks;

[0029] Extracting a set of prompt words from a relevant dataset of each task in the task set;

[0030] Using the target model to infer the prompt word to obtain a corresponding output token sequence;

[0031] The prompt words and the output tokens sequence are used as the alignment data set.

[0032] Preferably, in the heterogeneous model pool construction module, the method for constructing the heterogeneous draft model pool includes:

[0033] For each task in the task set, based on an existing small-scale language model, fine-tune the small-scale language model using the aligned dataset to obtain a draft model applicable to all tasks;

[0034] The draft models corresponding to all the tasks in the task set are integrated to obtain the heterogeneous draft model pool.

[0035] Preferably, the task classification module includes: a text processing unit, a classification model training unit and a task classification unit;

[0036] The text processing unit is used to obtain the input prompt word text and delete the stop words in the prompt word text to obtain a processed text;

[0037] The classification model training unit constructs a training data set for all tasks in the task set, wherein the training data set is composed of the prompt word texts belonging to different tasks, and a correct task category is annotated for each prompt word text, and the public language text classification model is trained using the training data set to obtain a fast classification model;

[0038] The task classification unit is used to input the processed text into the fast classification model to perform task classification and obtain the classification result.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] The present invention combines the task classification model with the draft model pool, and realizes efficient reasoning in a multi-task environment through the speculative decoding mechanism. It not only improves the reasoning performance of large-scale pre-trained language models under different tasks, but also enhances the adaptability of the model when facing diversified tasks through the combination of fine-tuning and task classification. Experimental results show that compared with the original speculative decoding method, this method significantly improves the reasoning speed on multiple tasks, and has strong practical value and promotion potential. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0042] Figure 1 A schematic diagram of a method flow of an embodiment of the present invention;

[0043] Figure 2 The figure is a flowchart of a method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0045] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0046] Embodiment 1

[0047] In this embodiment, if Figure 1 , Figure 2 As shown, a LLM speculative decoding optimization method for downstream tasks includes the following steps:

[0048] S1. Build multiple aligned datasets for different tasks and use the aligned datasets to build a heterogeneous draft model pool.

[0049] The method for constructing an aligned data set includes: obtaining a task set containing different tasks; extracting a set of prompt words from the relevant data set of each task in the task set; using a target model to infer and generate the prompt words to obtain a corresponding output token sequence; and using the prompt words and the output tokens sequence as an aligned data set.

[0050] In this embodiment, multiple aligned datasets for different tasks are constructed: the key to this stage is the quality and representativeness of the dataset. Through reasonable data sampling and annotation, it is ensured that the dataset can fully reflect various scenarios of the task, thereby providing a sufficient data basis for model fine-tuning. In order to improve the accuracy of guessed tokens generated by the draft model and reduce the performance difference between the draft model and the target model, this paper proposes a task-based prediction alignment method. Specifically, this method enables the draft model to generate outputs that are more consistent with the target model on specific tasks by fine-tuning task-specific datasets under different tasks, thereby improving the overall performance and efficiency of speculative decoding. In a multi-task learning environment, different tasks often have different input and output modes, which may cause the draft model to face uneven performance when processing multiple tasks. To solve this problem, the solution of this paper focuses on task-specific optimization of the draft model according to the characteristics of the task. In the task set T = {t 1 ,t 2 ,t 3 ,…}, for each task t i and its related datasets j , the draft model is fine-tuned through task-specific datasets to make the output generated by the draft model in a specific task as close as possible to the output of the target model. The generation quality of the draft model on the target task is significantly improved, effectively reducing the performance degradation caused by inaccurate predictions. In this scheme, first, we start from the task t i Related datasets j Extract all prompts and use the target model to generate the prompt sequence. The target model generates the corresponding token sequence for the prompt sequence, thereby obtaining a complete aligned dataset, which contains the input of the original dataset and the output generated by the target model, providing data corpus for the draft model suitable for the customized task.

[0051] The method for constructing a heterogeneous draft model pool includes: for each task in the task set, based on the existing small-scale language model, using the aligned dataset to fine-tune the small-scale language model to obtain a draft model suitable for all tasks; integrating the draft models corresponding to all tasks in the task set to obtain a heterogeneous draft model pool.

[0052] Constructing a heterogeneous draft model pool using aligned datasets: The construction of a draft model pool is the core content of the entire study. By fine-tuning multiple pre-trained models, a batch of draft models for different fine-grained downstream tasks are formed. These models show different capabilities and advantages and disadvantages on different tasks. The construction process of the draft model pool involves multi-stage fine-tuning and evaluation. In this process, it is first necessary to fine-tune the pre-trained model according to different downstream tasks. The fine-tuning strategy can include many forms, one of the most common of which is supervised fine-tuning (SFT). In the SFT process, the model adjusts its parameters by training on a specific task dataset in order to better fit the data distribution of the target task. In this embodiment, the SFT method is used to optimize the draft model. Based on the newly generated dataset d j ′ Fine-tune the draft model to make it i The generated effect is closer to the target model. This process can be achieved by minimizing the output difference between the draft model and the target model under the same task, thereby improving the adaptability and accuracy of the draft model in the task. The fine-tuned draft model can generate higher-quality guessed tokens in different tasks in a multi-task environment, reducing the performance loss caused by low-quality tokens generated by the draft model, thereby improving reasoning efficiency. The advantage of the draft model pool is that it can combine the advantages of multiple models and provide multiple options for subsequent task execution.

[0053] S2. Obtain the prompt word text of the downstream task, and use the task classification mechanism to classify the input task based on the prompt word text to obtain the classification result.

[0054] The method for obtaining the classification result includes: obtaining the text of the input task and deleting the stop words in the text to obtain the processed text; obtaining the input prompt word text and deleting the stop words in the prompt word text to obtain the processed text; constructing a training data set for all tasks in the task set, the training data set consists of prompt word texts belonging to different tasks, and annotating the correct task category for each prompt word text, using the training data set to train a public language text classification model to obtain a fast classification model; inputting the processed text into the fast classification model for task classification to obtain the classification result.

[0055] In this embodiment, for the multi-task scenario, considering the different requirements of different tasks for the draft model, a task classification mechanism is proposed. This mechanism can classify the input task according to its characteristics, so as to divide the tasks into different draft models for speculative decoding. Through this task-driven model selection strategy, it can be ensured that each task can use the most suitable draft model for inference, thus improving the decoding efficiency and generation quality. FastText is an efficient text classification tool with excellent training and inference performance. In this embodiment, the FastText model is retrained using task features to make it more adaptable to the classification requirements of specific tasks, and a classification model is obtained. In some vertical fields, text data often has specific language structures or distribution characteristics. By adjusting the selection range of n-grams and optimizing the training parameters, FastText can capture these characteristics more accurately. In the text classification task of this embodiment, after loading the text, necessary preprocessing is performed to optimize the input quality of the model. Among them, removing stop words is a key step. Stop words are usually words with low semantic contribution but high frequency in the text, such as "de", "le", "he", etc. By removing these words, noise interference can be reduced, and the attention of the classification model to the core semantic information can be improved. After completing the stop word processing, word segmentation is performed on the text, and the segmented text is input into the classification model for classification prediction, and finally efficient and accurate text classification is achieved.

[0056] S3. Select the optimal draft model from the heterogeneous draft model pool according to the classification result.

[0057] In this embodiment, based on the classification result, the most suitable draft model is quickly matched for the task during inference, so as to achieve more efficient and accurate speculative decoding.

[0058] S4. Use the selected draft model to generate draft tokens, and input the draft tokens into the LLM target model for parallel verification to complete the speculative decoding optimization.

[0059] Embodiment 2

[0060] In this embodiment, an LLM speculative decoding optimization system for downstream tasks includes: a heterogeneous model pool construction module, a task classification module, a model selection module, and a decoding optimization module.

[0061] The heterogeneous model pool construction module is used to construct multiple aligned data sets for different tasks, and use the aligned data sets to construct a heterogeneous draft model pool.

[0062] In the heterogeneous model pool construction module, the method for constructing an aligned dataset includes: obtaining a task set containing different tasks; extracting a set of prompt words from the relevant dataset of each task in the task set; using the target model to infer and generate the prompt words to obtain the corresponding output token sequence; and using the prompt words and the output tokens sequence as the aligned dataset.

[0063] In the heterogeneous model pool construction module, the method for constructing a heterogeneous draft model pool includes: for each task in the task set, based on the existing small-scale language model, the small-scale language model is fine-tuned using the aligned dataset to obtain a draft model suitable for all tasks; the draft models corresponding to all tasks in the task set are integrated to obtain a heterogeneous draft model pool.

[0064] The task classification module is used to obtain the prompt word text of the downstream task, and classify the input task based on the prompt word text using the task classification mechanism to obtain the classification result.

[0065] The task classification module includes: a text processing unit, a classification model training unit and a task classification unit; the text processing unit is used to obtain the input prompt word text and delete the stop words in the prompt word text to obtain the processed text; the classification model training unit constructs a training data set for all tasks in the task set, the training data set consists of prompt word texts belonging to different tasks, and the correct task category is marked for each prompt word text, and the training data set is used to train a public language text classification model to obtain a fast classification model; the task classification unit is used to input the processed text into the fast classification model for task classification to obtain the classification result.

[0066] The model selection module is used to select the best draft model from the heterogeneous draft model pool according to the classification results.

[0067] The decoding optimization module uses the selected draft model to generate a sequence of guessed tokens, and inputs the guessed tokens sequence into the target model for parallel verification to complete speculative decoding optimization.

[0068] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.

Claims

1. A LLM speculative decoding optimization method for downstream tasks, characterized in that: The following steps are involved: Constructing multiple aligned datasets for different tasks, and using the aligned datasets to construct a heterogeneous draft model pool; Obtain prompt word texts of downstream tasks, and classify input tasks based on the prompt word texts using a task classification mechanism to obtain classification results; selecting an optimal draft model from the heterogeneous draft model pool according to the classification result; The selected draft model is used to generate a guessed tokens sequence, and the guessed tokens sequence is input into the target model for parallel verification to complete the speculative decoding optimization.

2. According to the LLM speculative decoding optimization method for downstream tasks as described in claim 1, it is characterized in that: The method for constructing the alignment dataset includes: Get a task set containing different tasks; Extracting a set of prompt words from a relevant dataset of each task in the task set; Using the target model to infer the prompt word to obtain a corresponding output token sequence; The prompt words and the output tokens sequence are used as the alignment data set.

3. According to the LLM speculative decoding optimization method for downstream tasks as described in claim 2, it is characterized in that: The method of constructing the heterogeneous draft model pool includes: For each task in the task set, based on an existing small-scale language model, fine-tune the small-scale language model using the aligned dataset to obtain a draft model applicable to all tasks; The draft models corresponding to all the tasks in the task set are integrated to obtain the heterogeneous draft model pool.

4. According to the LLM speculative decoding optimization method for downstream tasks as described in claim 2, it is characterized in that: The method for obtaining the classification result includes: Obtaining an input prompt word text, and deleting stop words in the prompt word text to obtain a processed text; Constructing a training data set for all tasks in the task set, wherein the training data set is composed of the prompt word texts belonging to different tasks, and a correct task category is annotated for each prompt word text, and using the training data set to train a public language text classification model to obtain a fast classification model; The processed text is input into the fast classification model for task classification to obtain the classification result.

5. A LLM speculative decoding optimization system for downstream tasks, the system applying the method according to any one of claims 1 to 4, characterized in that: include: Heterogeneous model pool construction module, task classification module, model selection module and decoding optimization module; The heterogeneous model pool construction module is used to construct multiple aligned data sets for different tasks, and use the aligned data sets to construct a heterogeneous draft model pool; The task classification module is used to obtain the prompt word text of the downstream task, and classify the input task based on the prompt word text using the task classification mechanism to obtain the classification result; The model selection module is used to select the best draft model from the heterogeneous draft model pool according to the classification result; The decoding optimization module generates a guessed tokens sequence using the selected draft model, and inputs the guessed tokens sequence into the target model for parallel verification to complete speculative decoding optimization.

6. The LLM speculative decoding optimization system for downstream tasks according to claim 5, characterized in that: In the heterogeneous model pool construction module, the method for constructing the aligned data set includes: Get a task set containing different tasks; Extracting a set of prompt words from a relevant dataset of each task in the task set; Using the target model to infer the prompt word to obtain a corresponding output token sequence; The prompt words and the output tokens sequence are used as the alignment data set.

7. The LLM speculative decoding optimization system for downstream tasks according to claim 5, characterized in that: In the heterogeneous model pool construction module, the method for constructing the heterogeneous draft model pool includes: For each task in the task set, based on an existing small-scale language model, fine-tune the small-scale language model using the aligned dataset to obtain a draft model applicable to all tasks; The draft models corresponding to all the tasks in the task set are integrated to obtain the heterogeneous draft model pool.

8. The LLM speculative decoding optimization system for downstream tasks according to claim 5, characterized in that: The task classification module includes: a text processing unit, a classification model training unit and a task classification unit; The text processing unit is used to obtain the input prompt word text and delete the stop words in the prompt word text to obtain a processed text; The classification model training unit constructs a training data set for all tasks in the task set, wherein the training data set is composed of the prompt word texts belonging to different tasks, and a correct task category is annotated for each prompt word text, and the public language text classification model is trained using the training data set to obtain a fast classification model; The task classification unit is used to input the processed text into the fast classification model to perform task classification and obtain the classification result.

Citation Information

Patent Citations

  • Meta-knowledge fine adjustment method and platform for multi-task language model

    CN112100383A

  • Dynamic guess decoding method and device for large language model, equipment and medium

    CN118095209A

  • Large model decoding system and method, related equipment and computer program product

    CN118467207A

  • Large model reasoning acceleration method and system combining machine learning and speculation sampling

    CN118657220A

  • Speculation decoding method and device for large language model and medium

    CN119150848A

Cited By

  • Equipment end large language model reasoning method and device, equipment and program product

    CN121684014A

  • Wireless perception speculation decoding method and system based on edge-end collaborative large model reasoning

    CN122372939A