Large language model reasoning method and device based on diffusion model, terminal and medium

By optimizing the draft model generation and verification process through the diffusion model, the problems of limited draft model generation length and a small number of accepted tokens are solved, and the efficiency and speed of large language model inference are improved, making it suitable for application scenarios with high real-time requirements.

CN120806155AActive Publication Date: 2025-10-17SHANGHAI GUANGYU XINCHEN TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510937256.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-10-17
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

The existing draft model has a limited length for generating candidate token sequences due to its autoregressive characteristics, a small number of tokens actually accepted at a time, and a poor overall acceleration effect.

Method used

A draft model based on the diffusion model is adopted to generate multiple token representations through inverse diffusion processing, which are evaluated in parallel with the large language model for verification. The parallel received prefix sequence is determined, and the draft sequence generation and verification process is optimized by combining conditional judgment and correction operations.

Benefits of technology

Significantly increases the length of draft sequences, reduces the number of large model verifications, improves inference efficiency and computing resource utilization, and increases the inference speed of large language models in applications with high real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806155A_ABST
    Figure CN120806155A_ABST
Patent Text Reader

Abstract

The invention provides a large language model reasoning method and device based on a diffusion model, a terminal and a medium. According to the application, the draft model based on the diffusion model framework is utilized to generate the draft sequence, and then the draft sequence is verified through the large language model. According to the method and the device, the characteristic that the diffusion model naturally supports parallel processing is utilized, so that the length of the draft sequence is remarkably increased, the verification frequency of the large model is reduced, the reasoning efficiency of the large language model is improved, and the utilization efficiency of computing resources is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of large language model, in particular to a large language model inference method and device based on diffusion model, terminal and medium. BACKGROUND

[0002] Large language models (LLMs), such as models based on the Transformer architecture, have shown superior performance in many fields such as natural language processing. However, its self-recursive inference mechanism, i.e. generating text token by token, results in high inference delay, limiting its deployment in application scenarios with high real-time requirements.

[0003] To alleviate this problem, speculative decoding technology has emerged. This technology usually uses a small and fast draft model to pre-generate several candidate token sequences, and then uses the target large language model to verify these candidate tokens in parallel. If the verification is passed, multiple tokens can be accepted at once, thereby achieving the purpose of accelerating inference. In existing technologies, the draft model is usually a small version of the target large language model (such as a small model obtained by knowledge distillation, a model with fewer parameters of the same architecture) or a simplified self-recursive model.

[0004] However, the existing speculative decoding technology has the following problems:

[0005] (1) Limited draft generation length: The existing self-recursive draft model generates k tokens, which requires k serial decoding steps. In order to control the delay of draft generation so as not to offset the acceleration effect, the value of k is usually very limited, for example, only 3 to 5 tokens, which fundamentally limits the potential maximum acceleration benefit that can be brought by the word speculative decoding operation.

[0006] (2) Limited number of actual accepted tokens: Since k itself is small, and there may be a deviation between the generation quality and distribution of the draft model and the target large language model, the number of tokens n actually verified and received by the large language model is usually also small, and the average may be only 2-3 tokens. This makes the effective acceleration of single speculative decoding very limited.

[0007] (3) Overall acceleration effect is not obvious: Since the number of tokens that can be received by single speculative decoding is small, more decoding iterations are needed to generate a target length sequence, so the overall acceleration effect of speculative decoding is poor, and it is difficult to meet the growing higher requirements for the inference speed of large language models.

[0008] (4) The difficulty of balancing the efficiency and quality of draft model generation: Existing draft models require a difficult trade-off between generation speed and generation quality. Although overly simple draft models are fast, they have poor generation quality, resulting in low acceptance rates. On the other hand, slightly higher-quality draft models may not be fast enough due to their autoregressive nature, which weakens the acceleration significance of speculative decoding and still limits the number of tokens they can generate at a time. Summary of the Invention

[0009] In view of the shortcomings of the prior art described above, the purpose of this application is to provide a large language model inference method, device, terminal and medium based on a diffusion model, which is used to solve the problem that the existing draft model is limited in the length of the generated candidate token sequence due to its autoregressive characteristics, and the number of tokens actually accepted at a single time is small, which leads to poor overall acceleration effect.

[0010] To achieve the above-mentioned objectives and other related objectives, the first aspect of the present application provides a large language model inference method based on a diffusion model, comprising: obtaining a pre-trained draft model based on a diffusion model architecture and determining a large language model to be used; configuring the draft model according to pre-obtained draft model configuration parameters; obtaining context information to be inferred and converting it into a context sequence to be inferred; using the context sequence to be inferred as the current context sequence; decoding the current context sequence to obtain an intermediate output sequence; performing a termination judgment operation based on the intermediate output sequence according to a preset termination condition to obtain a corresponding judgment result; if the obtained judgment result satisfies the termination condition, the intermediate output sequence is used as the final output sequence; if the obtained judgment result does not satisfy the termination condition, the intermediate output sequence is used as the current context sequence and a decoding operation is performed on it until a judgment result that satisfies the termination condition is obtained; wherein the decoding operation includes: sequentially executing a draft sequence generation operation based on the draft model after parameter configuration, a verification operation based on the large language model, a received prefix determination operation, and a conditional judgment and correction operation; performing text conversion on the final output sequence, and using the conversion result as the inference result text information corresponding to the context information to be inferred.

[0011] In some embodiments of the first aspect of the present application, the draft model configuration parameters include: the length of the draft sequence and the number of denoising steps of the draft model.

[0012] In some embodiments of the first aspect of the present application, the long sequence draft generation operation based on the parameter-configured draft model includes: using the parameter-configured draft model to perform inverse diffusion processing on the current context sequence, generating multiple token representations, and converting the multiple token representations into a draft sequence according to a vocabulary.

[0013] In some embodiments of the first aspect of the present application, the validation operation based on the large language model comprises: inputting the current context sequence and the draft sequence into the large language model for parallel evaluation to obtain an optimal token sequence.

[0014] In some embodiments of the first aspect of the present application, the determining operation of receiving a prefix comprises: outputting a token part in the draft sequence that is consistent with the optimal token sequence from the first token as a prefix sequence, and outputting the number of tokens in the prefix sequence as the number of tokens accepted by the large language model; and splicing the prefix sequence with the current context sequence to obtain a preliminary intermediate output sequence.

[0015] In some embodiments of the first aspect of the present application, the conditional judgment operation comprises: if the number of accepted tokens is less than the length of the draft sequence, performing a correction operation to obtain a corrected intermediate output sequence, and taking the corrected intermediate output sequence as the final intermediate output sequence; and if the number of accepted tokens is equal to the length of the draft sequence, taking the preliminary intermediate output sequence as the final intermediate output sequence.

[0016] In some embodiments of the first aspect of the present application, the correction operation comprises: inputting the prefix sequence and the current context sequence into the large language model to obtain a correction token; and splicing the correction token with the preliminary intermediate output sequence to obtain the corrected intermediate output sequence.

[0017] To achieve the above object and other related objects, the second aspect of the present application provides a large language model reasoning device based on a diffusion model, comprising: a model obtaining module configured to obtain a pre-trained draft model based on a diffusion model architecture and determine a large language model to be used; a model configuration module configured to configure the draft model according to pre-obtained draft model configuration parameters; a reasoning module configured to obtain context information to be reasoned and convert it into a context sequence to be reasoned; take the context sequence to be reasoned as a current context sequence; perform a decoding operation on the current context sequence to obtain an intermediate output sequence; perform a termination judgment operation based on the intermediate output sequence according to a preset termination condition to obtain a corresponding judgment result; if the obtained judgment result is that the termination condition is met, take the intermediate output sequence as a final output sequence; if the obtained judgment result is that the termination condition is not met, take the intermediate output sequence as the current context sequence and perform a decoding operation thereon until a judgment result that meets the termination condition is obtained; wherein the decoding operation comprises: draft sequence generation operation based on the draft model after parameter configuration, validation operation based on the large language model, determining the received prefix operation and conditional judgment and correction operation performed in sequence; and a result generation module configured to perform text conversion on the final output sequence and take the converted result as reasoning result text information corresponding to the context information to be reasoned.

[0018] To achieve the above object and other related objects, the third aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the large language model reasoning method based on the diffusion model.

[0019] To achieve the above object and other related objects, the fourth aspect of the present application provides an electronic terminal comprising a memory, a processor and a computer program stored on the memory; the processor executes the computer program to implement the large language model reasoning method based on the diffusion model.

[0020] As described above, the large language model reasoning method, device, terminal and medium based on the diffusion model of the present application have the following beneficial effects:

[0021] The present application takes advantage of the natural parallel processing characteristics of the diffusion model, which significantly increases the length of the draft sequence, thereby reducing the number of large model validations, thereby improving the reasoning efficiency of the large language model and improving the utilization efficiency of the computing resources. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 The flowchart shown is the flowchart of the large language model reasoning method based on the diffusion model in an embodiment of the present application.

[0023] Figure 2A specific flowchart of a decoding operation is shown as an embodiment of the present application.

[0024] Figure 3 A schematic block diagram of a large language model inference device based on a diffusion model is shown as an embodiment of the present application.

[0025] Figure 4 A structural schematic diagram of an electronic terminal is shown as an embodiment of the present application. DETAILED DESCRIPTION

[0026] The embodiments of the present application will be described in detail with specific reference to particular examples. It will be understood that the advantages and benefits of the present application can be readily gleaned from the disclosure herein by persons of ordinary skill in the art. The present application can be embodied in other different embodiments and applied to other different situations without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0027] In the embodiments of the present application, the terms "first", "second", and the like are used to distinguish between similar or identical items or components having substantially the same function and effect. Those skilled in the art can understand that the terms "first", "second", and the like do not limit the number and execution order, and the terms "first", "second", and the like do not necessarily mean different.

[0028] It should be noted that in the embodiments of the present application, the words "exemplary" or "for example" indicate an example, illustration, or description. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design solutions. Rather, the use of the words "exemplary" or "for example" is intended to present the relevant concept in a specific manner.

[0029] In the embodiments of the present application, "at least one" means one or more, and "multiple" means two or more. The association relationship between the associated objects is described by "and / or", which means that there can be three relationships, for example, A and / or B, which can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.

[0030] Before further detailing the present application, the terms and phrases involved in the embodiments of the present application are explained, and the terms and phrases involved in the embodiments of the present application are applicable to the following explanations:

[0031] <1>token: In a language model, token is the smallest unit of text processing. Token is the tokenization result of the input text by the model.

[0032] <2>denoising steps: Denoising steps are inherent parameters of diffusion models. Denoising steps refer to the number of iterations required for the model to gradually denoise from pure noise to the final image during the generation process.

[0033] <3>Context information: Context information refers to the input text (i.e. user-provided dialogue history, document content, etc.) that the model refers to when generating a reply. It is a key basis for the model to understand the current task, maintain the coherence of the dialogue, and provide accurate answers.

[0034] <4>logits: In machine learning and deep learning, logits is an important concept, which usually refers to the original prediction values output by the model in classification problems, which have not been normalized, such as the softmax function. Logits are the outputs of the last layer of the model, which are unscaled log probabilities.

[0035] <5>Vocabulary: Vocabulary is a collection of all recognizable and usable words, subwords, characters or symbols in a language model. It defines the range of language units that the model can handle.

[0036] <6>Token ID: Token ID is a unique identifier (usually an integer) for each word, subword or character in the vocabulary. It is a mapping relationship that converts language units (such as words, subwords or characters) in text into numerical values that the model can understand.

[0037] <7>End of Sequence: In a language model, text is usually processed as a sequence of tokens. End of Sequence (usually referred to as EOS) is used to explicitly tell the model that the text sequence has ended. This helps the model know when to stop generating when generating text.

[0038] To facilitate understanding of the embodiments of the present application, first, the diffusion model-based large language model inference method according to the embodiments of the present application is described in combination with the diffusion model. Figure 1 The diffusion model-based large language model inference method according to the embodiments of the present application is described in detail. Figure 1 A flowchart of a diffusion model-based large language model inference method according to an embodiment of the present application is shown. The diffusion model-based large language model inference method according to the embodiments of the present application mainly includes the following steps:

[0039] Step S11: Obtain a pre-trained diffusion model architecture-based draft model and determine a large language model to be used.

[0040] It should be understood that a pre-trained large language model with better performance can be selected for use according to actual needs, and the present application does not limit this. The types of large language models include but are not limited to Mistral AI models, Llama 3 models of the MetaLlama series, and Llama 2 models.

[0041] In an embodiment, the manner of training the draft model specifically includes: using the large language model determined in the above process to construct a training set of the draft model; and training the diffusion model framework using the training set to obtain a trained draft model.

[0042] In an embodiment, constructing a training set of the draft model using the determined large language model includes: obtaining a plurality of context sequences (token sequences) and inputting the plurality of context sequences into the determined large language model to obtain a plurality of large language model output token sequences corresponding thereto; wherein the plurality of large language model output token sequences cover large language model output token sequences of different sequence lengths; and the plurality of context sequences and the plurality of context sequences respectively corresponding plurality of large language model output token sequences constitute a training set of the draft model. It should be understood that the context sequence can be obtained by converting the context information.

[0043] In an embodiment, the type of diffusion model framework includes but is not limited to a U-Net framework, a non-autoregressive Transformer structure, etc.

[0044] In an embodiment, the draft model is optimized using a minimum denoising error. Specifically, the plurality of context sequences are input into the diffusion model framework to obtain a plurality of model framework output token sequences corresponding thereto. The minimum denoising error reflects the error between the plurality of model framework output token sequences and the plurality of large language model output token sequences corresponding thereto.

[0045] In an embodiment, the final acceleration effect of the inference decoding can also be directly optimized in combination with reinforcement learning (such as maximizing the expected number of accepted tokens).

[0046] Step S12: Configure the draft model according to the pre-obtained draft model configuration parameters.

[0047] In an embodiment, the draft model configuration parameters include the length of the draft sequence and the number of denoising steps of the draft model.

[0048] It should be understood that the configuration parameters of the draft model can be set according to actual needs. However, it should be noted that the length of the draft sequence should be set to be much larger than the length of the existing draft in the existing inference decoding (the existing draft length is generally 2 to 3), for example, the length of the draft sequence can be set to 8, 12, 16 or longer. The number of denoising steps of the draft model should be set to be much smaller than the number of steps when the diffusion model is used for image generation and other tasks. It should be noted that the existing number of denoising steps of the diffusion model is usually within 50 to 1000. It should also be noted that the length of the draft sequence and the number of denoising steps of the draft model are two key hyperparameters. The selection of the length of the draft sequence k_D needs to balance the benefits of the increase in the number of accepted tokens and the increased overhead of the target LLM M_L verifying k_D tokens. The selection of the number of denoising steps N_S needs to balance the generation speed and the generation quality of the draft. These two parameters can be jointly optimized through experiments to maximize the overall inference acceleration ratio.

[0049] In step S13, the context information to be inferred is obtained and converted into a context sequence to be inferred; the context sequence to be inferred is taken as a current context sequence; a decoding operation is performed on the current context sequence to obtain an intermediate output sequence; a termination judgment operation is performed based on the intermediate output sequence according to a preset termination condition to obtain a corresponding judgment result; if the obtained judgment result satisfies the termination condition, the intermediate output sequence is taken as a final output sequence; if the obtained judgment result does not satisfy the termination condition, the intermediate output sequence is taken as the current context sequence and a decoding operation is performed thereon until a judgment result satisfying the termination condition is obtained.

[0050] In an embodiment, the context information to be inferred is language information. The context information to be inferred can come from various fields, including but not limited to the medical field, the biological field, the game field, the chip field, etc.

[0051] In an embodiment, the context sequence to be inferred is a token sequence. It should be noted that those skilled in the art can convert the context information into corresponding tokens through existing conversion methods, which will not be described here.

[0052] In an embodiment, the decoding operation includes: sequentially performing a draft sequence generation operation of a draft model configured based on parameters, a verification operation based on a large language model, a determination of a received prefix operation, and a conditional judgment and correction operation.

[0053] In an embodiment, the long sequence draft generation operation of the draft model configured based on parameters includes: performing inverse diffusion processing on the current context sequence using the draft model configured based on parameters to generate a plurality of token representations, and converting the plurality of token representations into a draft sequence according to a vocabulary.

[0054] It should be noted that the generation process of the draft model utilizes the intrinsic characteristics of the diffusion model (for example, some variants can achieve non-autoregressive one-time prediction of the entire sequence of noise, and then parallel denoising), avoiding the strict step-by-step serial dependence in traditional autoregressive models, thereby allowing more efficient parallel computing. In summary, the diffusion model framework adopted by the draft model naturally supports parallel processing of elements of the entire sequence, and can achieve one-time or parallel generation of token sequences of a certain length, so that the additional time overhead of generating long sequences is effectively controlled. Notably, the number of token representations generated by the draft model is equal to the length of the pre-configured draft sequence. The length of the draft sequence is the length of the pre-configured draft sequence.

[0055] In an embodiment, the draft model runs on a specific type of hardware accelerator (such as GPU, TPU). It should be noted that the diffusion model, especially the highly parallelized variant, can exhibit better computational efficiency and throughput than existing autoregressive models when running on the above-mentioned hardware accelerator, thereby assisting in accelerating the inference efficiency of large language models.

[0056] In an embodiment, the validation operation based on the large language model includes: inputting the current context sequence and the draft sequence into the large language model for parallel evaluation to obtain the optimal token sequence.

[0057] In an embodiment, the determining operation of receiving the prefix includes: outputting the token part in the draft sequence that is consistent with the optimal token sequence from the first token as a prefix sequence, and outputting the number of tokens in the prefix sequence as the number of tokens accepted by the large language model; concatenating the prefix sequence with the current context sequence to obtain a preliminary intermediate output sequence.

[0058] In an embodiment, the conditional judgment operation includes: if the number of accepted tokens is less than the length of the draft sequence, performing a correction operation to obtain a corrected intermediate output sequence, and taking the corrected intermediate output sequence as the final intermediate output sequence; if the number of accepted tokens is equal to the length of the draft sequence, taking the preliminary intermediate output sequence as the final intermediate output sequence.

[0059] In an embodiment, the correction operation includes: inputting the prefix sequence and the current context sequence into the large language model to obtain a correction token; concatenating the correction token with the preliminary intermediate output sequence to obtain a corrected intermediate output sequence.

[0060] The following will be combined with the accompanying Figure 2 The inference process is explained as follows:

[0061] Before the inference starts, the context information to be inferred is obtained and converted into a context sequence C_0 to be inferred. The target generation length L_target is set.

[0062] The context sequence C_0 to be inferred is taken as the current context sequence C. The current context sequence C is input into the draft model. The draft model performs an inverse diffusion process to generate K_D token representations. The process when the diffusion model processes data is the inverse diffusion process, which is an inherent feature of the diffusion model. The inverse diffusion process includes: starting from a randomly sampled noise sequence (or a predefined noise pattern), combining the current context sequence C, and generating K_D token representations through N_S steps of iterative denoising. It should be understood that the number of token representations is equal to the length of the draft sequence in the draft model configuration parameter, and the number of steps of iterative denoising is equal to the denoising step number of the draft model in the draft model configuration parameter.

[0063] Further, the logits distribution of each token representation in the generated K_D token representations on the vocabulary is calculated; argmax operation or sampling is performed on the logits distribution of each token representation on the vocabulary to obtain the token ID of each token representation; and the token ID of each token representation is output as the draft sequence.

[0064] Further, the draft sequence T_D and the current context sequence C are input into the large language model for parallel evaluation, and the optimal token sequence token t_i^L that the large language model itself considers is output. It should be understood that parallel evaluation is an inherent feature of the large language model when processing, and as long as the above sequence is input into the large language model, the large language model can automatically complete parallel evaluation. The length of the optimal token sequence is equal to the length of the draft sequence in the draft model configuration parameter.

[0065] Further, the prefix sequence T_A is initialized. The draft sequence is T_D = (t_1^D,..., t_{k_D}^D), and the optimal token sequence is (token t_1^L,..., token t_{k_D}^L). In order, from 1 to k_D, each token in the draft sequence is compared with the corresponding token in the optimal token sequence; if they are consistent, t_1^D is stored in the prefix sequence and is the i-th element in the prefix sequence T_A. When the first token in the draft sequence is inconsistent with the corresponding token in the optimal token sequence, the comparison is stopped. The prefix sequence T_A at the time of stopping is the final prefix sequence, and the number of tokens in the prefix sequence T_A at the time of stopping is the number of tokens accepted by the large language model. For example, the first two tokens in the draft sequence are consistent with the corresponding tokens in the optimal token sequence, and when it is determined that the third token is inconsistent with the third token in the optimal token sequence, the comparison is immediately stopped. The final prefix sequence is composed of only the first two tokens in the draft sequence. The number of tokens accepted by the large language model is 2.

[0066] The consistency of two tokens can be determined by comparing whether the token IDs are consistent. The consistency of the two tokens can also be determined by greedy matching (t_i^D is the token with the highest probability generated by the large language model under the condition The consistency of the two tokens can also be determined by random sampling matching.

[0067] Further, the prefix sequence T_A is concatenated with the current context sequence C to obtain a preliminary intermediate output sequence S_out: The prefix sequence T_A is concatenated after the current context sequence C.

[0068] Further, if the number of accepted tokens is less than the length of the draft sequence, the prefix sequence and the current context sequence are input into the large language model to obtain a revised token (the number of revised tokens is generally one); the revised token is concatenated with the preliminary intermediate output sequence to obtain a revised intermediate output sequence: (the revised token is concatenated after the preliminary intermediate output sequence); the revised intermediate output sequence is taken as the final intermediate output sequence; if the number of accepted tokens is equal to the length of the draft sequence, the preliminary intermediate output sequence is taken as the final intermediate output sequence.

[0069] Further, if the length of the final intermediate output sequence reaches the target generation length L target or the final intermediate output sequence appears the sequence end symbol, the final intermediate output sequence is taken as the final output sequence; if the length of the final intermediate output sequence does not reach the target generation length L target or the sequence end symbol does not appear, the final intermediate output sequence is taken as the current context sequence, and the above decoding operation is repeated until the length of the final intermediate output sequence reaches the target generation length L target or the final intermediate output sequence appears the sequence end symbol.

[0070] It should be noted that the draft model of the present application can generate a longer token sequence than the existing autoregressive draft model. The longer draft k D provides a basis for obtaining more accepted tokens n. Even if the acceptance rate n / k D remains unchanged, n will increase with the increase of k D. If the diffusion model improves the draft quality, the increase of n will be more significant, which is the key to acceleration. Due to the substantial increase of n, the "large model verification times" required to complete the generation of the same length of text will be significantly reduced, thereby realizing the nonlinear and stepwise improvement of inference efficiency. Since a single decoding operation can process and accept more information (i.e., more tokens), the "output" (i.e., the number of accepted tokens) of each interaction (verification phase) with the relatively expensive large language model is higher, which improves the utilization efficiency of computing resources.

[0071] For applications that require real-time interaction with a large language model (such as chat robots, real-time translation, code assistants, etc.), the significant improvement in inference speed will directly translate into faster system response speed, thereby greatly improving the user experience.

[0072] In a specific embodiment, a control module can be provided to coordinate the workflow of the draft model and the large language model, manage the context information, perform the determination of the receiving prefix operation, and control the iteration process of the entire inference decoding.

[0073] Step S14: performing text conversion on the final output sequence, and taking the converted result as the inference result text information corresponding to the context information to be inferred.

[0074] It should be understood that the final output sequence is a token sequence, and the final output sequence can be converted by using the existing text conversion method, which will not be described here.

[0075] Figure 3 is a schematic block diagram of the large language model inference device based on the diffusion model provided by the embodiments of the present application. As shown in Figure 3 The large language model inference device based on the diffusion model includes:

[0076] The model obtaining module 31 is configured to obtain a pre-trained draft model based on a diffusion model architecture and determine a large language model to be used.

[0077] The model configuration module 32 is configured to configure the draft model according to pre-obtained draft model configuration parameters.

[0078] The inference module 33 is configured to obtain context information to be inferred and convert it into a context sequence to be inferred; take the context sequence to be inferred as a current context sequence; perform a decoding operation on the current context sequence to obtain an intermediate output sequence; perform a termination judgment operation based on the intermediate output sequence according to a preset termination condition to obtain a corresponding judgment result; if the obtained judgment result is that the termination condition is met, take the intermediate output sequence as a final output sequence; if the obtained judgment result is that the termination condition is not met, take the intermediate output sequence as the current context sequence and perform a decoding operation thereon until a judgment result that meets the termination condition is obtained.

[0079] The decoding operation includes: a draft sequence generation operation based on the draft model after parameter configuration, a validation operation based on the large language model, a prefix receiving operation, and a conditional judgment and correction operation, which are sequentially performed.

[0080] The result generation module 34 is configured to perform text conversion on the final output sequence and take the converted result as inference result text information corresponding to the context information to be inferred.

[0081] It should be understood that the specific process of each module performing the corresponding steps described above has been described in detail in the method embodiments described above, and for the sake of brevity, will not be repeated here.

[0082] It should also be understood that the division of the modules in the embodiments of the present application is illustrative, and is only a logical functional division. In actual implementation, there can be another division manner. In addition, each functional module in each embodiment of the present application can be integrated in one processor, or can be physically separated, or two or more modules can be integrated in one module. The integrated module can be realized in the form of hardware or in the form of a software functional module.

[0083] In an embodiment, the draft model configuration parameters include: the length of the draft sequence and the number of denoising steps of the draft model.

[0084] In an embodiment, the long sequence draft generation operation based on the draft model after parameter configuration includes: performing inverse diffusion processing on the current context sequence by using the draft model after parameter configuration to generate a plurality of token representations, and converting the plurality of token representations into a draft sequence according to a vocabulary.

[0085] In an embodiment, the large language model-based verification operation comprises: inputting the current context sequence and the draft sequence into the large language model for parallel evaluation to obtain an optimal token sequence.

[0086] In an embodiment, the determining a received prefix operation comprises: outputting a token part in the draft sequence that is consistent with the optimal token sequence from the first token as a prefix sequence, and outputting a token number in the prefix sequence as a number of tokens accepted by the large language model; and splicing the prefix sequence and the current context sequence to obtain a preliminary intermediate output sequence.

[0087] In an embodiment, the conditional judgment operation comprises: if the number of accepted tokens is less than the length of the draft sequence, performing a correction operation to obtain a corrected intermediate output sequence, and taking the corrected intermediate output sequence as the final intermediate output sequence; and if the number of accepted tokens is equal to the length of the draft sequence, taking the preliminary intermediate output sequence as the final intermediate output sequence.

[0088] In an embodiment, the correction operation comprises: inputting the prefix sequence and the current context sequence into the large language model to obtain a correction token; and splicing the correction token and the preliminary intermediate sequence to obtain the corrected intermediate output sequence.

[0089] Figure 4 is a schematic block diagram of an electronic terminal provided by an embodiment of the present application. As shown in Figure 4 , the electronic terminal comprises at least one processor 401, a memory 402, at least one network interface 403, and a user interface 405. The various components in the apparatus are coupled together by a bus system 404. It can be understood that the bus system 404 is used to realize the connection and communication between the components. In addition to including a data bus, the bus system 404 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity in Figure 4 , all the buses are marked as the bus system.

[0090] The user interface 405 can include a display, a keyboard, a mouse, a trackball, a click gun, a key, a button, a touchpad, or a touch screen, etc.

[0091] It can be understood that the memory 402 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), which is used as an external cache. By way of example but not limitation, many forms of RAM can be used, such as static random access memory (SRAM), synchronous static random access memory (SSRAM). The memory described in the embodiments of the application is intended to include but not limited to these and any other suitable category of memory.

[0092] The memory 402 in the embodiments of the application is used to store various categories of data to support the operation of the electronic terminal 400. Examples of these data include: any executable program for operating on the electronic terminal 400, such as an operating system 4021 and an application program 4022; the operating system 4021 contains various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application program 4022 can contain various application programs, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. The implementation of the diffusion model-based large language model inference method provided by the embodiments of the application can be included in the application program 4022.

[0093] The method disclosed in the above embodiments of the application can be applied in the processor 401 or implemented by the processor 401. The processor 401 can be an integrated circuit chip with processing capability. In the implementation process, each step of the above method can be completed by integrated logic circuits or instructions in the form of software in the processor 401. The above processor 401 can be a general-purpose processor, a digital signal processor (DSP), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc. The processor 401 can implement or execute the disclosed methods, steps and logic block diagrams in the embodiments of the application. The general-purpose processor 401 can be a microprocessor or any conventional processor, etc. The steps of the accessory optimization method provided in conjunction with the embodiments of the application can be directly embodied as hardware decoding processor execution completion, or executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the memory, and the processor reads the information in the memory and combines the hardware to complete the steps of the above method.

[0094] In an exemplary embodiment, the electronic terminal 400 can be one or more Application Specific Integrated Circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), or the like for performing the aforementioned methods.

[0095] According to the method provided in the embodiments of the present application, the present application further provides a computer program product, which comprises computer program codes, and when the computer program codes run on a computer, the computer is caused to perform the method provided in the embodiments of the present application. Figure 1 The large language model reasoning method based on the diffusion model in the illustrated embodiment.

[0096] According to the method provided in the embodiments of the present application, the present application further provides a computer readable storage medium, which stores program codes, and when the program codes run on a computer, the computer is caused to perform the method provided in the embodiments of the present application. Figure 1 The large language model reasoning method based on the diffusion model in the illustrated embodiment.

[0097] As used in this specification, the terms "component," "module," "system", and the like are intended to refer to a computer-related entity, either hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to being, a process running on a processor, a processor, an object, an executable, a thread of execution, a program, and / or a computer. By way of illustration, both an application running on a computing device and the computing device can be a component. One or more components can reside within a process and / or thread of execution and a component can be localized, partially localized, and / or distributed across two or more computers. Also, these components can execute from various computer readable media having various data structures stored thereon. The components can communicate by way of local and / or remote processes such as in accordance with a signal having one or more data packets (e.g., data from one component interacting with another component in a local system, distributed system, and / or across a network such as the Internet with other systems via the signal).

[0098] Those of skill in the art would understand that the various illustrative logical blocks and steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, or a combination of computer software and electronic hardware. The choice of hardware or software, or combination thereof, would be dependent on the specific application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0099] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.

[0100] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative, for example, the division of units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0101] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0102] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.

[0103] In the above embodiments, the functions of the various functional units can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, the functions can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, the whole or part of the processes or functions according to the embodiments of the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. containing one or more available media. The available media can be magnetic media (such as floppy disk, hard disk, magnetic tape), optical media (such as high-density digital video disc (digital video disc, DVD), or semiconductor media (such as solid state disk (solid state disk, SSD) and the like.

[0104] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium and includes a number of instructions for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the embodiments of the present application. The storage medium mentioned above includes: U disk, mobile hard disk, read-only memory (read-only memory, ROM), random access memory (random access memory, RAM), magnetic disk or optical disk and various media that can store program codes.

[0105] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0106] In summary, the present application provides a large language model reasoning method and device based on a diffusion model, a terminal and a medium. The present application generates a draft sequence by using a draft model based on a diffusion model framework, and then verifies the draft sequence by a large language model. The present application takes advantage of the characteristic that the diffusion model naturally supports parallel processing, so that the length of the draft sequence is significantly increased, thereby reducing the number of large model verifications, thereby improving the reasoning efficiency of the large language model and improving the utilization efficiency of the computing resources. Therefore, the present application effectively overcomes the various shortcomings in the prior art and has a high industrial utilization value.

[0107] The above embodiments only exemplarily illustrate the principles and effects of the present application, and are not intended to limit the present application. Any person skilled in the art can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by those skilled in the art without departing from the spirit and technical ideas disclosed by the present application should be covered by the claims of the present application.

Claims

1. A large language model inference method based on a diffusion model, characterized in that: include: Obtain a pre-trained draft model based on the diffusion model architecture and determine the large language model to be used; Configuring the draft model according to pre-obtained draft model configuration parameters; Obtain the context information to be inferred and convert it into a context sequence to be inferred; use the context sequence to be inferred as the current context sequence; perform a decoding operation on the current context sequence to obtain an intermediate output sequence; According to a preset termination condition, based on the intermediate output sequence, a termination judgment operation is performed to obtain a corresponding judgment result; If the obtained judgment result satisfies the termination condition, the intermediate output sequence is used as the final output sequence; If the obtained judgment result is that the termination condition is not satisfied, the intermediate output sequence is used as the current context sequence and a decoding operation is performed on it until a judgment result that satisfies the termination condition is obtained; The decoding operation includes: a draft sequence generation operation based on the draft model after parameter configuration, a verification operation based on the large language model, a received prefix determination operation, and a condition judgment and correction operation, which are performed in sequence; The final output sequence is converted into text, and the conversion result is used as the reasoning result text information corresponding to the context information to be inferred.

2. The large language model inference method based on the diffusion model according to claim 1, characterized in that: The draft model configuration parameters include: the length of the draft sequence and the number of denoising steps of the draft model.

3. The large language model inference method based on the diffusion model according to claim 2, characterized in that: The long sequence draft generation operation based on the parameter-configured draft model includes: using the parameter-configured draft model to perform inverse diffusion processing on the current context sequence to generate multiple token representations, and converting the multiple token representations into a draft sequence according to a vocabulary.

4. The large language model inference method based on the diffusion model according to claim 3 is characterized in that: The verification operation based on the large language model includes: inputting the current context sequence and the draft sequence into the large language model for parallel evaluation to obtain the optimal token sequence.

5. The large language model inference method based on the diffusion model according to claim 4 is characterized in that: The operation of determining the received prefix includes: outputting the token portion of the draft sequence starting from the first token that is consistent with the optimal token sequence as a prefix sequence, and outputting the number of tokens in the prefix sequence as the number of tokens accepted by the large language model; splicing the prefix sequence with the current context sequence to obtain a preliminary intermediate output sequence.

6. The large language model inference method based on the diffusion model according to claim 5 is characterized in that: The conditional judgment operation includes: if the number of accepted tokens is less than the length of the draft sequence, performing a correction operation to obtain a corrected intermediate output sequence, and using the corrected intermediate output sequence as the final intermediate output sequence; if the number of accepted tokens is equal to the length of the draft sequence, using the preliminary intermediate output sequence as the final intermediate output sequence.

7. The large language model inference method based on the diffusion model according to claim 6, characterized in that: The correction operation includes: inputting the prefix sequence and the current context sequence into the large language model to obtain a correction token; and concatenating the correction token with the preliminary intermediate output sequence to obtain a corrected intermediate output sequence.

8. A large language model inference device based on a diffusion model, characterized in that: include: A model acquisition module is used to obtain a pre-trained draft model based on the diffusion model architecture and determine the large language model to be used; A model configuration module, configured to configure the draft model according to pre-obtained draft model configuration parameters; The inference module is used to obtain the context information to be inferred and convert it into a context sequence to be inferred; use the context sequence to be inferred as the current context sequence; and perform a decoding operation on the current context sequence to obtain an intermediate output sequence; According to a preset termination condition, based on the intermediate output sequence, a termination judgment operation is performed to obtain a corresponding judgment result; If the obtained judgment result satisfies the termination condition, the intermediate output sequence is used as the final output sequence; If the obtained judgment result is that the termination condition is not satisfied, the intermediate output sequence is used as the current context sequence and a decoding operation is performed on it until a judgment result that satisfies the termination condition is obtained; The decoding operation includes: a draft sequence generation operation based on the draft model after parameter configuration, a verification operation based on the large language model, a received prefix determination operation, and a condition judgment and correction operation, which are performed in sequence; The result generation module is used to perform text conversion on the final output sequence and use the conversion result as the inference result text information corresponding to the context information to be inferred.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. An electronic terminal comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Dynamic guess decoding method and device for large language model, equipment and medium

    CN118095209A

  • Big language model reasoning method and device, equipment and storage medium

    CN118364918A

  • Large language model reasoning acceleration method and device based on predictive decoding

    CN118886511A

  • Large language model reasoning optimization method based on cascade and speculative decoding strategy

    CN119047579A

  • Speculation decoding method and device for large language model and medium

    CN119150848A

Cited By

  • Large model speculation reasoning optimization method and system based on online reinforcement learning

    CN121328745A