Text generation method, text generation model training method and related devices

By using input word vectors containing multiple masks in the text generation model and using the rejection sampling method, the problem of slow inference process of existing text generation methods is solved, and more efficient decoding speed and generation effect is achieved.

CN120235152APending Publication Date: 2025-07-01SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202311873965.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-30
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing text generation methods, especially large language models, have slow decoding inference process through autoregression and cannot effectively utilize the processor's parallel computing power.

Method used

By introducing input word element vectors into the text generation model, including input word element arrays, output word element arrays, output mask arrays, N candidate word elements and N candidate mask arrays, and using the rejection sampling method, ensuring that the output word elements and autoregressive decoding results obey the same distribution, thereby improving the decoding speed.

Benefits of technology

While improving the decoding speed, the generation effect is ensured without relying on other additional models or structural changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235152A_ABST
    Figure CN120235152A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of artificial intelligence, and provides a method for generating a text, a method for training a text generation model and a related device.The method comprises the steps that an input text is obtained; determining an input lexical element vector according to the input text, wherein the input lexical element vector comprises an input lexical element array, an output lexical element array, an output mask array, N candidate lexical elements and N candidate mask arrays; the input lexical element vector is input into a text generation model, and an output result is obtained and comprises candidate output lexical elements and probability distribution of the candidate output lexical elements; sampling the output result to obtain a sampling result; performing sampling rejection based on the sampling result to obtain an output lexical element; and adding the output lexical elements into the output lexical element array, and taking the output lexical element array as an output text. According to the method provided by the embodiment of the invention, the generation effect can be ensured while the decoding speed is improved, and the method does not need to depend on other additional models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of artificial intelligence, and particularly relates to a method for generating text, a method for training a text generation model, and related devices. Background Art

[0002] With the rapid development of artificial intelligence technology, the application of text generation is becoming more and more extensive. For example, a large language model (LLM) can be used to generate output text so that the machine can "speak" in a way similar to human language.

[0003] However, there are still some problems with existing text generation methods. For example, traditional large language models usually use autoregressive decoding, and during the inference process, they need to perform serial decoding word by word (token), resulting in a slow inference process and unable to effectively utilize the parallel computing power of the processor. Summary of the Invention

[0004] Embodiments of this application provide a method for generating text, a method for training a text generation model, and related devices, which can improve the decoding speed while ensuring the generation effect, and do not need to rely on other additional models.

[0005] In a first aspect, an embodiment of this application provides a method for generating text, including:

[0006] Obtain an input text;

[0007] Determine an input token vector according to the input text, where the input token vector includes an input token array, an output token array, an output mask array, N candidate tokens, and N candidate mask arrays. Each candidate mask array in the N candidate mask arrays includes N consecutive masks, and the N candidate mask arrays correspond to the N candidate tokens one by one, and N is a positive integer greater than 1;

[0008] Input the input token vector into a text generation model to obtain an output result, where the output result includes candidate output tokens and their probability distributions;

[0009] Sample the output result to obtain a sampling result, where the sampling result includes the predicted tokens and their probabilities corresponding to the output token array, the predicted tokens and their probabilities corresponding to the output mask array, the N predicted tokens and their probabilities corresponding to the N candidate tokens, and the N groups of predicted tokens and their probabilities corresponding to the N mask arrays;

[0010] Perform rejection sampling based on the sampling result to obtain output tokens;

[0011] Add the output token to the output token array, and use the output token array as the output text.

[0012] In an embodiment of the present application, the input token vector includes an input token array, an output token array, an output mask array, N candidate tokens, and N candidate mask arrays. After processing the input token vector, a sampling result including multiple predicted tokens can be obtained. At this time, rejection sampling is performed based on the sampling result, which can ensure that the obtained output tokens follow the same distribution as the result obtained by autoregressive decoding. In this way, the decoding speed can be improved while ensuring the generation effect, and no other additional models are required.

[0013] In some possible implementation manners, the obtaining of the output token by performing rejection sampling based on the sampling result includes:

[0014] Judging whether to accept each of the N candidate tokens in turn based on the rejection sampling algorithm;

[0015] If the first candidate token among the N candidate tokens is rejected, add the predicted token corresponding to the last token in the output token array to the output token array, and use the predicted token corresponding to the output mask array as the candidate token;

[0016] If the i-th candidate token among the N candidate tokens is accepted, add the i-th candidate token to the output token array;

[0017] If the i-th candidate token among the N candidate tokens is rejected, where i is an integer greater than 1 and less than or equal to N, stop the judgment, use the predicted token corresponding to the candidate mask array corresponding to the (i - 1)-th candidate token as the candidate token, and add the predicted token corresponding to the (i - 1)-th candidate token to the output token array.

[0018] In some possible implementation manners, the inputting the input token vector into the text generation model to obtain an output result includes:

[0019] Input the input token vector, the attention matrix, and the position encoding into the text generation model to obtain the output result;

[0020] Among them, the attention matrix satisfies the following conditions:

[0021] When the row number j is greater than or equal to the column number k, and the j-th token in the input token vector is not a mask, the element in the attention matrix with row number j and column number k is 1, where j and k are integers greater than or equal to 0 and less than M, and M is the number of tokens in the input token vector;

[0022] When the number of rows j is greater than or equal to the number of columns k, and the j-th token in the input token vector is a mask, if the k-th token in the input token vector is also a mask and j - k < N, then the element at row j and column k in the attention matrix is 1;

[0023] Then all other elements in the attention matrix are 0;

[0024] The position encoding satisfies the following conditions:

[0025] The j-th element in the position encoding is the sum of all elements in the j-th row of the attention matrix minus 1.

[0026] In the embodiments of the present application, through the designed attention matrix and the position encoding of the present application, multiple tokens (predicted tokens corresponding to the output mask array and N candidate mask arrays) can be predicted and generated in one inference of the text generation model, while verifying the N candidate tokens obtained from the previous inference and ensuring that the two do not interfere with each other. In this way, the decoding speed can be improved while ensuring the generation effect.

[0027] In some possible implementation manners, the text generation model is a large language model.

[0028] In a second aspect, an embodiment of the present application provides a method for training a text generation model, including:

[0029] Obtain an input text;

[0030] Determine an input token vector according to the input text, where the input token vector includes multiple consecutive masks;

[0031] Input the input token vector into the text generation model to obtain an output result, where the output result includes candidate output tokens and their probability distributions;

[0032] Sample the output result to obtain an output token and its probability;

[0033] Calculate a loss value according to the output token and its probability;

[0034] Train the text generation model according to the loss value.

[0035] In the embodiments of the present application, the input token vector includes multiple consecutive masks. In this way, inputting the input token vector into the text generation model for one inference can obtain multiple output tokens, thereby improving the efficiency of text generation.

[0036] In a third aspect, an embodiment of the present application provides a device for generating text, including:

[0037] An acquisition unit for acquiring input text;

[0038] A determination unit for determining an input token vector according to the input text, the input token vector including an input token array, an output token array, an output mask array, N candidate tokens and N candidate mask arrays, each candidate mask array in the N candidate mask arrays including N consecutive masks, the N candidate mask arrays corresponding one-to-one to the N candidate tokens, and N being a positive integer greater than 1;

[0039] A processing unit for inputting the input token vector into a text generation model to obtain an output result, the output result including candidate output tokens and their probability distributions;

[0040] A sampling unit for sampling the output result to obtain a sampling result, the sampling result including the predicted tokens and their probabilities corresponding to the output token array, the predicted tokens and their probabilities corresponding to the output mask array, the N predicted tokens and their probabilities corresponding to the N candidate tokens, and the N groups of predicted tokens and their probabilities corresponding to the N mask arrays;

[0041] A rejection sampling unit for performing rejection sampling based on the sampling result to obtain output tokens;

[0042] An addition unit for adding the output tokens to the output token array and using the output token array as output text.

[0043] In a fourth aspect, an embodiment of the present application provides a device for training a text generation model, including:

[0044] An acquisition unit for acquiring input text;

[0045] A determination unit for determining an input token vector according to the input text, the input token vector including a plurality of consecutive masks;

[0046] A processing unit for inputting the input token vector into the text generation model to obtain an output result, the output result including candidate output tokens and their probability distributions;

[0047] A sampling unit for sampling the output result to obtain output tokens and their probabilities;

[0048] A calculation unit for calculating a loss value according to the output tokens and their probabilities;

[0049] A training unit for training the text generation model according to the loss value.

[0050] Fifth aspect, an embodiment of the present application provides an apparatus for generating text, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the steps of the method according to any one of the above first aspects are implemented.

[0051] Sixth aspect, an embodiment of the present application provides an apparatus for training a text generation model, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the steps of the method according to any one of the above second aspects are implemented.

[0052] Seventh aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, where when the computer program is executed by a processor, the steps of the method according to any one of the above first aspects are implemented.

[0053] Eighth aspect, an embodiment of the present application provides a computer program product, which when running on an evaluation device, causes the evaluation device to execute the method according to any one of the above first aspects.

[0054] It can be understood that the beneficial effects of the above third aspect to the eighth aspect can be referred to the relevant descriptions in the above first aspect, and will not be repeated here.

[0055] The beneficial effects of the embodiments of the present application compared with the prior art are:

[0056] In the embodiments of the present application, the input token vector includes an input token array, an output token array, an output mask array, N candidate tokens, and N candidate mask arrays. After processing the input token vector, a sampling result including multiple predicted tokens can be obtained. At this time, rejection sampling is performed based on the sampling result, which can ensure that the obtained output tokens follow the same distribution as the result obtained by autoregressive decoding. In this way, the decoding speed can be improved while ensuring the generation effect, and there is no need to rely on other additional models. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0058] Figure 1 is a schematic diagram of an application scenario provided by an embodiment of the present application.

[0059] Figure 2 It is a schematic diagram of generating text based on autoregressive decoding in an embodiment of the present application.

[0060] Figure 3 It is a schematic flowchart of a method for training a text generation model provided in an embodiment of the present application.

[0061] Figure 4 It is a schematic diagram of generating text based on semi-autoregressive decoding in an embodiment of the present application.

[0062] Figure 5 It is a schematic flowchart of a method for generating text provided in an embodiment of the present application.

[0063] Figure 6 It is a schematic flowchart of a method for generating text provided in another embodiment of the present application.

[0064] Figure 7 It is a schematic diagram of an input token vector in an embodiment of the present application.

[0065] Figure 8 It is a schematic diagram of a self-attention matrix and positional encoding in an embodiment of the present application.

[0066] Figure 9 It is a schematic diagram of a large language model for inference in an embodiment of the present application.

[0067] Figure 10 It is a schematic structural diagram of a device for training a text generation model provided in an embodiment of the present application.

[0068] Figure 11 It is a schematic structural diagram of a device for generating text provided in an embodiment of the present application.

[0069] Figure 12 It is a schematic structural diagram of a device provided in an embodiment of the present application. Detailed implementation manners

[0070] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0071] It should be understood that when used in the description of the present application specification and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0072] It should also be understood that the term "and / or" used in the description of the present application specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0073] As used in the description of the present application specification and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrases "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.

[0074] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0075] The reference to "one embodiment" or "some embodiments" etc. described in the present application specification means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0076] The solutions provided by the embodiments of the present application can be applied to various scenarios that require text generation. For example, the methods for generating text and training a text generation model provided by the embodiments of the present application can be executed on a server, can also be executed in the cloud, and can also be executed on a terminal device.

[0077] Taking the terminal device as an example, such as Figure 1As shown, the technical solution of the embodiment of the present invention can be applied to a terminal device. The method for generating text in the embodiments of the present application can perform inference (or processing) based on the input text to obtain the output text corresponding to the input text. The method for training a text generation model provided in the embodiments of the present application can train a text generation model for generating text.

[0078] The terminal device in the embodiments of the present application can be mobile or fixed. For example, the terminal device can be a mobile phone, a tablet personal computer (TPC), a media player, a smart TV, a laptop computer (LC), a personal digital assistant (PDA), a personal computer (PC), a camera, a video camera, a smart watch, a wearable device (WD), or an autonomous vehicle, etc., that has the function of generating text (such as deploying a text generation model). The embodiments of the present invention do not limit this.

[0079] The text generation model in the present application can be a large language model (LLM). Taking the large language model as an example, the technical problems existing in the process of generating text will be described below.

[0080] Generally, when generating text based on a large language model, the large language model can adopt autoregressive decoding, that is, serial decoding is performed on each token one by one, and this method is relatively inefficient. For example, as Figure 2 shown, assume that at time t = 0, the input text is " <s>"I love Beijing" can encode the input text into input tokens "1, 45, 96, 68" through a tokenizer, input the input tokens into a large language model, and obtain an output token "52" (this token corresponds to "day"); at time t = 1, add the output token obtained at time t = 0 to the input text, that is, the input text at time t = 1 is " <s>I love Beijing Tiananmen Square. Input the input token into the large language model, and get the output token "34" (this token corresponds to "An"); finally, through 5 inferences, the large language model can infer 5 tokens "52, 34, 28, 6, 2", and the corresponding text is "Tiananmen Square. < / s> ". It can be seen that autoregressive decoding can only infer one output token at a time, and the inference efficiency is low.

[0081] To improve the inference speed, the industry has proposed many optimization methods for the inference stage. For example, improving the implementation of the computing core, multi-card parallel computing, batch processing strategies, quantization pruning, and so on. Among them, speculative decoding is a method that has been proven effective in practice.

[0082] Speculative sampling is an advanced large model inference acceleration technology that can achieve an acceleration ratio of more than 3 times without sacrificing the generation effect. Speculative sampling can use a small model (which can be called a draft model) for autoregressive sampling and generate multiple candidate tokens, and then use the LLM to evaluate the sampling results (the generated multiple candidate tokens). This processing process is similar to the associative input in the input method: let the small model "guess" the tokens that the LLM may generate next, and then let the LLM verify multiple candidate tokens at the same time. Since the LLM can verify multiple candidate tokens at the same time, and the small model has a fast sampling speed, this method can quickly generate multiple candidate tokens, achieve a significant acceleration effect, and at the same time ensure that the sampling distribution of the multiple candidate tokens generated is exactly the same as the result obtained using the LLM.

[0083] However, this method requires the use of an additional small model, which will increase the memory overhead, and the accuracy of the small model has a great impact on the inference acceleration effect. In some cases, the predictions of the small model may not be accurate at all. In this way, on the contrary, the acceleration effect will be reduced due to the introduction of additional computational overhead.

[0084] Another similar method is the blockwise parallel decoding method, which also adopts the idea of predicting first and then verifying. This method does not require the introduction of an additional small model, but directly trains multiple classification heads in the last layer of the LLM, so that the LLM can predict multiple tokens to come at the same time in one inference, and then use the original LLM to verify these multiple tokens.

[0085] However, this method requires changing the model structure, and at the same time introduces additional model training parameters, which will also bring additional memory overhead.

[0086] To solve one or more of the above technical problems, the present application proposes a method for generating text, a method for training a text generation model, and related devices, which can improve the decoding speed while ensuring the generation effect. The method for generating text proposed in the present application also adopts the idea of predicting first and then verifying, but this method does not require changing the model structure and does not require introducing additional training parameters or models. In this way, no additional overhead (such as introducing additional parameters or models) will be added, so it has higher generality.

[0087] In the model training stage, the method for training a text generation model proposed by the present invention introduces multiple consecutive masked tokens (i.e., [MASK], which can also be simply referred to as masks) during the supervised fine-tuning (SFT) process (or during the training process), enabling the text generation model to decode multiple tokens in a semi-autoregressive decoding manner in parallel.

[0088] In the inference stage, the method for generating text proposed by the present invention uses rejection sampling to verify the candidate tokens output by the text generation model, ensuring that the results decoded in parallel follow the same distribution as the results of autoregressive sampling. In this way, not only can the decoding speed be improved, but also the generation effect can be ensured, and no additional parameters or models need to be introduced.

[0089] The above-mentioned semi-autoregressive decoding is common in translation models, and its principle is to divide the entire translation into K (K is a positive integer) blocks, perform non-autoregressive decoding within the blocks, and perform autoregressive decoding between the blocks. In this way, multiple consecutive words can be generated in parallel at each time step.

[0090] Next, in conjunction with Figure 3 A detailed example of the method for training a text generation model in the embodiments of the present application will be given.

[0091] Figure 3 FIG. shows a schematic flowchart of the method for training a text generation model provided by an embodiment of the present application. By way of example and not limitation, this method can be applied to Figure 1 the terminal device shown, and can also be applied to a server or the cloud.

[0092] Figure 3 The method 300 in

[0093] S310, obtain the input text.

[0094] S320, determine the input token vector according to the input text.

[0095] The input token vector may include multiple consecutive masks (MASK).

[0096] S330. Input the input token vector into the text generation model to obtain an output result.

[0097] The output result may include candidate output tokens and their probability distributions.

[0098] S340. Sample the output result to obtain an output token and its probability.

[0099] S350. Calculate a loss value based on the output token and its probability.

[0100] When training the text generation model, the input text can be a part of the training sample (the training sample here can be a sentence, a paragraph, an article, etc.). Then, when calculating the loss value, the other parts of the training sample except the input sample can be used as the ground truth of the input sample.

[0101] For example, the training sample can be " <s>I love Beijing Tiananmen Square. < / s> ", and the input text in S310 can be " <s>I love Beijing", then, in S350, "Tiananmen Square. < / s> " (that is, the other parts of the training sample except the input sample) as the ground truth to calculate the loss value corresponding to the input sample.

[0102] S360. Train the text generation model according to the loss value.

[0103] For example, the model parameters of the text generation model can be adjusted according to the loss value.

[0104] In the embodiments of the present application, the input token vector may include a plurality of consecutive masks. In this way, inputting the input token vector into the text generation model for one inference can obtain a plurality of output tokens, thereby improving the efficiency of text generation.

[0105] For example, as Figure 4 shown, assume that at t = 0, the input text is " <s>"I love Beijing" can encode the input text into input tokens "1, 45, 96, 68" through a tokenizer, add 4 masks to the input tokens (such as the token "3" in Figure 4 can represent a mask), and input the input tokens with added masks into the large language model. At this time, since there are 4 masks in the input tokens, 4 output tokens "52, 34, 28, 6, 2" can be directly obtained through one inference, and the corresponding text is "Tiananmen Square.< / s> ". In this way, a plurality of output tokens can be obtained through one inference, which can improve the efficiency of text generation.

[0106] Next, in combination with Figure 5 and Figure 6 a detailed example of the method for generating text in the embodiments of the present application will be given.

[0107] Figure 5 FIG. shows a schematic flowchart of a method for generating text provided by an embodiment of the present application. As an example and not a limitation, this method can be applied to the Figure 1 shown terminal device, and can also be applied to a server or the cloud.

[0108] Figure 5 The method 500 in [it] includes steps S510 to S560, which are as follows:

[0109] S510, obtain the input text;

[0110] S520, determine the input token vector according to the input text.

[0111] The input token vector may include an input token array, an output token array, an output mask array, N candidate tokens, and N candidate mask arrays, where N is a positive integer greater than 1.

[0112] Among them, each candidate mask array in the N candidate mask arrays may include N consecutive masks, and the N candidate mask arrays and the N candidate tokens may correspond one by one.

[0113] In some embodiments, when the candidate token array L c is not empty, the input token vector can be constructed using the following rules, denoted as I, which are as follows:

[0114] I = T + L a + M + L c [0] + M0 +... + L c [N - 1] + M N-1

[0115] Among them, "+" represents array vector concatenation, T represents the input token array, L a represents the output token array, M represents the output mask array, L c [i] represents the i-th element of the candidate token array L c , M i represents the i-th candidate mask array, and each candidate mask array includes N masks, that is, M i = [MASK] * N, and L c [i] and M i can correspond one by one.

[0116] In some embodiments, when L c is empty, the constructed input token vector is:

[0117] I = T + L a + M

[0118] S530, input the input token vector into the text generation model to obtain an output result, and the output result includes candidate predicted tokens and their probability distributions.

[0119] In some embodiments, the text generation model may be a large language model.

[0120] S540, sample the output result to obtain a sampling result.

[0121] In some embodiments, greedy decoding can be used, or polynomial sampling or nucleus sampling can be adopted to sample the output result to obtain a sampling result.

[0122] The sampling result may include the predicted token and its probability corresponding to the output token array, the predicted token and its probability corresponding to the output mask array, the N predicted tokens and their probabilities corresponding to the N candidate tokens, and the N groups of predicted tokens and their probabilities corresponding to the N mask arrays.

[0123] In some embodiments, the input token vector, attention matrix, and positional encoding can be input into the text generation model to obtain the output result.

[0124] Among them, the attention matrix and the positional encoding may be related to the masks in the input token vector (such as the output mask array and N candidate mask arrays), so that the prediction generation process (predicting and generating multiple tokens) and the candidate token verification process (verifying the N candidate tokens obtained from the previous inference) in one inference of the text generation model do not interfere with each other, thereby improving the decoding speed while ensuring the generation effect.

[0125] In some embodiments, the attention matrix may satisfy the following conditions:

[0126] When the row number j is greater than or equal to the column number k, and the j-th token in the input token vector is not a mask, the element in the attention matrix with row number j and column number k is 1, where j and k are integers greater than or equal to 0 and less than M, and M is the number of tokens in the input token vector;

[0127] When the row number j is greater than or equal to the column number k, and the j-th token in the input token vector is a mask, if the k-th token in the input token vector is also a mask and j - k < N, the element in the attention matrix with row number j and column number k is 1;

[0128] Then all other elements in the attention matrix are 0.

[0129] In some embodiments, the positional encoding may satisfy the following conditions:

[0130] The j-th element in the positional encoding is the sum of all elements in the j-th row of the attention matrix minus 1.

[0131] In the embodiments of the present application, through the attention matrix and the position encoding designed in the present application, during one inference of the text generation model, while predicting and generating multiple tokens (predicted tokens corresponding to the output mask array and N candidate mask arrays), the N candidate tokens obtained from the previous inference can be verified, and it is ensured that there is no interference between the two. In this way, the decoding speed can be improved while ensuring the generation effect.

[0132] S550, perform rejection sampling based on the sampling result to obtain the output token.

[0133] In some embodiments, the performing rejection sampling based on the sampling result to obtain the output token includes:

[0134] Based on the rejection sampling algorithm, sequentially determine whether to accept each of the N candidate tokens;

[0135] If the first candidate token among the N candidate tokens is rejected, add the predicted token corresponding to the last token in the output token array to the output token array, and use the predicted token corresponding to the output mask array as the candidate token;

[0136] If the i-th candidate token among the N candidate tokens is accepted, add the i-th candidate token to the output token array;

[0137] If the i-th candidate token among the N candidate tokens is rejected, where i is a positive integer greater than 1 and less than or equal to N, stop the determination, use the predicted token corresponding to the candidate mask array corresponding to the (i - 1)-th candidate token as the candidate token, and add the predicted token corresponding to the (i - 1)-th candidate token to the output token array.

[0138] S560, add the output token to the output token array, and use the output token array as the output text.

[0139] In the embodiments of the present application, the input token vector includes an input token array, an output token array, an output mask array, N candidate tokens, and N candidate mask arrays. After processing the input token vector, a sampling result including multiple predicted tokens can be obtained. At this time, performing rejection sampling based on the sampling result can ensure that the obtained output tokens follow the same distribution as the result obtained by autoregressive decoding. In this way, the decoding speed can be improved while ensuring the generation effect, and there is no need to rely on other additional models.

[0140] The following combines Figure 6 , and illustrates the above method for generating text through a specific embodiment.

[0141] Figure 6 A schematic flowchart of a method for generating text provided by another embodiment of the present application is shown. By way of example and not limitation, this method can be applied to Figure 1 the terminal device shown, and can also be applied to a server or the cloud.

[0142] Figure 6 The method 600 in

[0143] S610, initialize data.

[0144] Data can be initialized for the output token array and candidate tokens. For example, L a = [], L c = [], where L a represents the output token array, and L c represents the candidate token array obtained from the previous inference of the text generation model.

[0145] S620, construct the input token vector.

[0146] The input token vector can include an input token array, an output token array, an output mask array, N candidate tokens, and N candidate mask arrays.

[0147] In some embodiments, when the candidate token array L c is not empty, the input token vector can be constructed using the following rule, denoted as I, specifically as follows:

[0148] I = T + L a + M + L c [0] + m0 +... + L c [N - 1] + M N-1

[0149] where "+" represents array vector concatenation, T represents the input token array, L a represents the output token array, M represents the output mask array, L c [i] represents the i-th element of the candidate token array L c , M i represents the i-th candidate mask array, and each candidate mask array includes N masks, that is, M i = [MASK] * N, and L c [i] and M i can correspond one by one.

[0150] In some embodiments, when L c is empty, the constructed input token vector can be:

[0151] I = T + L a + M

[0152] For example, when N = 2, the input text S = " <s>"I love", after encoding the input text S, the input token vector T can be obtained, T = [1, 45, 96], L a = [68, 52], L c = [34, 28], at this time, the constructed input token vector can be as Figure 7 shown.

[0153] In particular, at time t = 0, L a and L c are both empty. At this time, the input token vector can be T + m.

[0154] S630, construct the input attention matrix and position encoding.

[0155] The traditional attention matrix is a lower triangular matrix, which means that when calculating self-attention for the previous tokens, the subsequent tokens cannot be "attended to". In order to simultaneously perform token verification and prediction in a single inference of a large language model, this application can design the attention matrix and position encoding according to the following method.

[0156] For example, if the length of the input token vector I is denoted as L, then the following attention matrix A can be defined:

[0157] A ∈ [0, 1] L*L ,

[0158] The attention matrix A can satisfy the following conditions:

[0159] 1) When j ≥ k and I[k] is not [MASK], A jk = 1;

[0160] 2) When j ≥ k and I[k] is [MASK], if I[j] is also [MASK] and j - k < N, A jk = 1;

[0161] 3) In other cases, A jk = 0.

[0162] At the same time, the position encoding P can be defined as:

[0163] P ∈ N L ,

[0164] where the position encoding P can satisfy

[0165] Through the above design, in a single inference of the text generation model, candidate tokens (corresponding to the M and M i parts) can be predicted and generated simultaneously, and the candidate tokens decoded in the previous step (corresponding to the L c part) can be verified, and it can be ensured that there is no interference between the two.

[0166] For example, when N = 2, the input text S = " <s>"I love", the input token vector T = [1, 45, 96] obtained after encoding the input text S, L a = [68, 52], L c = [34, 28], the self-attention matrix and position encoding designed by the method of this application can be as Figure 8 shown.

[0167] S640, use a large language model for inference.

[0168] The constructed input token vector, attention matrix and position encoding can be input into the text generation model to obtain the inference result of the large language model, and the output tokens and their probabilities are obtained by sampling. The schematic diagram is as Figure 9 shown.

[0169] Among them, L a [-1] (L a [-1] can represent the last element of the output token array L a ), the probability distribution of the predicted token at the corresponding position of L c [i] is P a and P c [i]. Greedy decoding can be used, or polynomial sampling or nucleus sampling can be used to sample L a [-1] and L c [i]. Denote the predicted tokens obtained after sampling L a [-1] and L c [i] as Y a and Y c [i], and the corresponding probabilities are P a (Y a ) and P c [i](Y c [i]). At the same time, denote the tokens after sampling M and M i as Y m and Y mi . The sampling process can be as Figure 9 shown.

[0170] S650, obtain the output token based on rejection sampling.

[0171] When using the text generation model for inference for the first time, L c is empty. The output token Y a corresponding to L a [-1] can be added to the output token array L a .

[0172] Otherwise, if L c If it is not empty (when it is not the first time to perform inference using the text generation model), then it is possible to sequentially determine whether to accept each of the N candidate tokens based on the rejection sampling algorithm.

[0173] If the first candidate token L among the N candidate tokens is rejected c [0], then Y a can be added to the output token array L a , and the predicted token Y corresponding to M m is used as a candidate token;

[0174] If the i-th candidate token L among the N candidate tokens is accepted c [i - 1], then L c [i - 1] can be added to the output token array L a ;

[0175] If the i-th candidate token L among the N candidate tokens is rejected c [i - 1], then the judgment can be stopped, and the predicted token Y corresponding to M i-1 is used as a candidate token, and Y mi-1 [i - 2] is added to the output token array L c . a

[0176] For example, L c [i - 1] can represent the i-th candidate token obtained from the previous decoding, and Q c [i - 1] can represent the probability corresponding to L c [i - 1] generated from the previous decoding.

[0177] For the first candidate token L c [0], P a (L c [0]) represents the probability of the candidate token L c [0] in the probability distribution P a (P a is the probability distribution of the predicted token obtained from the last element of the output token array L a ), and a number r[0] can be uniformly sampled from [0, 1]. At this time, the judgment condition can be:

[0178] r[0] ≤ P a (L c [0]) / Q c [0]

[0179] That is, if r[0] ≤ P a (L c [0]) / Q c [0], then the first candidate token L can be accepted c [0], add L c [0] to the output token array L a , and continue to verify the next candidate token; otherwise, the first candidate token L c [0] can be rejected, sample a token from P a (i.e., Y a defined above) and add it to the output token array L a , and exit the verification process.

[0180] For the i-th (where i represents an integer greater than 1 and less than or equal to N) candidate token L c [i - 1], P c [i - 1](L c [i - 1]) represents the probability of the candidate token L c [i - 1] in the probability distribution P c [i - 1] (P c [i - 1] is the probability distribution of the predicted token obtained from the candidate token L c [i - 1]). A number uniformly sampled from [0, 1] can be denoted as r[i - 1]. At this time, the judgment condition can be:

[0181] r[i - 1] ≤ P c [i - 1](L c [i - 1]) / Q c [i - 1]

[0182] That is, if r[i - 1] ≤ P c [i - 1](L c [i - 1]) / Q c [i - 1], then the i-th candidate token L c [i - 1] can be accepted, add L c [i - 1] to the output token array L a , and continue to verify the next candidate token; otherwise, the i-th candidate token L c [i - 1] can be rejected, sample a token from P c [i - 2] (i.e., Y c [i - 2]) and add it to the output token array L a , and exit the verification process.

[0183] Furthermore, the above steps S620 and S650 can be repeated until L a contains the document end flag (in this example, it is 1, and the corresponding identifier is< / s> )".

[0184] It should be noted that in the above embodiments, each time a new token is added to L a a judgment needs to be made, and it is determined whether the token newly added to L a is an end flag.

[0185] In the above embodiments, L a records the inference result (i.e., the output token) of the text generation model. Among them, when verifying in step S650, it is possible that L a receives more than one token, which is equivalent to being able to output multiple tokens in one inference. In this way, the effect of accelerating inference can be achieved. At the same time, by using rejection sampling in step S650, it can be ensured that the obtained inference result follows the same probability distribution as the inference result obtained by traditional autoregressive decoding.

[0186] In the above, in combination with Figures 1 to 9 , the method embodiments of the present application have been described in detail. Next, in combination with Figure 10 and Figure 12 , the apparatus embodiments of the present application will be described in detail. It should be understood that the descriptions of the method embodiments and the apparatus embodiments correspond to each other. Therefore, for the parts not described in detail, reference can be made to the previous method embodiments.

[0187] Figure 10 FIG. is a schematic structural diagram of an apparatus for training a text generation model provided by an embodiment of the present application. As Figure 10 shown, the apparatus 1000 includes an acquisition unit 1010, a determination unit 1020, a processing unit 1030, a sampling unit 1040, a calculation unit 1050, and a training unit 1060, specifically as follows:

[0188] The acquisition unit 1010 is configured to acquire an input text;

[0189] The determination unit 1020 is configured to determine an input token vector according to the input text, and the input token vector includes a plurality of consecutive masks;

[0190] The processing unit 1030 is configured to input the input token vector into the text generation model to obtain an output result, and the output result includes candidate output tokens and their probability distributions;

[0191] The sampling unit 1040 is configured to sample the output result to obtain an output token and its probability;

[0192] The calculation unit 1050 is configured to calculate a loss value according to the output token and its probability;

[0193] A training unit 1060 is configured to train the text generation model according to the loss value.

[0194] Figure 11 It is a schematic structural diagram of a device for generating text provided by an embodiment of the present application. As Figure 11 shown, the device 1100 includes an acquisition unit 1110, a determination unit 1120, a processing unit 1130, a sampling unit 1140, a rejection sampling unit 1150, and an addition unit 1160, which are specifically as follows:

[0195] The acquisition unit 1110 is configured to acquire an input text;

[0196] The determination unit 1120 is configured to determine an input token vector according to the input text, where the input token vector includes an input token array, an output token array, an output mask array, N candidate tokens, and N candidate mask arrays. Each candidate mask array in the N candidate mask arrays includes N consecutive masks, and the N candidate mask arrays correspond to the N candidate tokens one by one, and N is a positive integer greater than 1;

[0197] The processing unit 1130 is configured to input the input token vector into a text generation model to obtain an output result, where the output result includes candidate output tokens and their probability distributions;

[0198] The sampling unit 1140 is configured to sample the output result to obtain a sampling result, where the sampling result includes the predicted tokens and their probabilities corresponding to the output token array, the predicted tokens and their probabilities corresponding to the output mask array, the N predicted tokens and their probabilities corresponding to the N candidate tokens, and the N groups of predicted tokens and their probabilities corresponding to the N mask arrays;

[0199] The rejection sampling unit 1150 is configured to perform rejection sampling based on the sampling result to obtain output tokens;

[0200] The addition unit 1160 is configured to add the output tokens to the output token array and use the output token array as an output text.

[0201] Figure 12 shows a schematic diagram of a device provided by an embodiment of the present application. As Figure 12 shown, the device 1200 in this embodiment includes: a processor 1210, a memory 1220, and a computer program 1230 stored in the memory 1220 and executable on the processor 1210. When the processor 1210 executes the computer program 1230, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 1210 executes the computer program 1230, the functions of the respective modules in the above-mentioned device embodiments are implemented.

[0202] Exemplarily, the computer program 1230 may be divided into one or more modules, which are stored in the memory 1220 and executed by the processor 1210 to complete the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and the instructions are used to describe the execution process of the computer program 1230 in the device 1200.

[0203] The device 1200 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The device may include, but is not limited to, a processor 1210 and a memory 1220. Those skilled in the art can understand that Figure 12 merely examples of the device 1200, which do not constitute a limitation on the device 1200, may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the device may further include input / output devices, network access devices, a bus, etc.

[0204] The processor 510 may be a graphics processing unit (GPU), a neural-network processing unit (NPU), or may also be a central processing unit (CPU), and may also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0205] The memory 520 may be an internal storage unit of the device 500, such as a hard disk or memory of the device 500. The memory 520 may also be an external storage device of the device 500, such as a plug-in hard disk equipped on the device 500, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 520 may also include both the internal storage unit of the device 500 and the external storage device. The memory 520 is used to store the computer program and other programs and data required by the device 500. The memory 520 may also be used to temporarily store data that has been output or is to be output.

[0206] It should be noted that for the content such as information interaction and execution process between the above-mentioned device / units, since it is based on the same concept as the method embodiment of the present application, for its specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.

[0207] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment, and details are not described herein again.

[0208] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.

[0209] The embodiment of the present application provides a computer program product, which when running on an evaluation device enables the evaluation device to implement the steps in the above-mentioned method embodiments when executed.

[0210] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0211] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0212] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed in this document can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0213] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0214] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0215] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.< / s> < / s>

Claims

1. A method for generating text, characterized in that, Comprising: Obtain input text; Determine an input token vector according to the input text, the input token vector including an input token array, an output token array, an output mask array, N candidate tokens, and N candidate mask arrays, each candidate mask array in the N candidate mask arrays including N consecutive masks, the N candidate mask arrays corresponding one-to-one to the N candidate tokens, and N being a positive integer greater than 1; Input the input token vector into a text generation model to obtain an output result, the output result including candidate output tokens and their probability distributions; Sample the output result to obtain a sampling result, the sampling result including the predicted tokens and their probabilities corresponding to the output token array, the predicted tokens and their probabilities corresponding to the output mask array, the N predicted tokens and their probabilities corresponding to the N candidate tokens, and the N groups of predicted tokens and their probabilities corresponding to the N mask arrays; Perform rejection sampling based on the sampling result to obtain output tokens; Add the output tokens to the output token array, and use the output token array as the output text.

2. The method according to claim 1, wherein The performing rejection sampling based on the sampling result to obtain output tokens includes: Successively determine whether to accept each candidate token in the N candidate tokens based on a rejection sampling algorithm; If the first candidate token in the N candidate tokens is rejected, add the predicted token corresponding to the last token in the output token array to the output token array, and use the predicted token corresponding to the output mask array as the candidate token; If the i-th candidate token in the N candidate tokens is accepted, add the i-th candidate token to the output token array; If the i-th candidate token in the N candidate tokens is rejected, where i is an integer greater than 1 and less than or equal to N, stop the determination, use the predicted token corresponding to the candidate mask array corresponding to the (i - 1)-th candidate token as the candidate token, and add the predicted token corresponding to the (i - 1)-th candidate token to the output token array.

3. The method according to claim 1 or 2, characterized in that, The inputting the input token vector into a text generation model to obtain an output result includes: Input the input token vector, an attention matrix, and a position encoding into the text generation model to obtain the output result; Wherein, the attention matrix satisfies the following conditions: When the row number j is greater than or equal to the column number k, and the j-th token in the input token vector is not a mask, the element at row j and column k in the attention matrix is 1, where j and k are integers greater than or equal to 0 and less than M, and M is the number of tokens in the input token vector; When the row number j is greater than or equal to the column number k, and the j-th token in the input token vector is a mask, if the k-th token in the input token vector is also a mask and j - k < N, the element at row j and column k in the attention matrix is 1; Then all other elements in the attention matrix are 0; The position encoding satisfies the following conditions: The j-th element in the position encoding is the sum of all elements in the j-th row of the attention matrix minus 1.

4. The method according to claim 1, characterized in that, The text generation model is a large language model.

5. A method for training a text generation model, characterized in that, Comprising: Obtain the input text; Determine an input token vector according to the input text, where the input token vector includes a plurality of consecutive masks; Input the input token vector into the text generation model to obtain an output result, where the output result includes candidate output tokens and their probability distributions; Sample the output result to obtain an output token and its probability; Calculate a loss value according to the output token and its probability; Train the text generation model according to the loss value.

6. An apparatus for generating text, characterized in that, Comprising: An obtaining unit for obtaining the input text; A determining unit for determining an input token vector according to the input text, where the input token vector includes an input token array, an output token array, an output mask array, N candidate tokens, and N candidate mask arrays, and each candidate mask array in the N candidate mask arrays includes N consecutive masks, and the N candidate mask arrays correspond to the N candidate tokens one by one, and N is a positive integer greater than 1; A processing unit for inputting the input token vector into the text generation model to obtain an output result, where the output result includes candidate output tokens and their probability distributions; A sampling unit for sampling the output result to obtain a sampling result, where the sampling result includes the predicted tokens and their probabilities corresponding to the output token array, the predicted tokens and their probabilities corresponding to the output mask array, the N predicted tokens and their probabilities corresponding to the N candidate tokens, and the N groups of predicted tokens and their probabilities corresponding to the N mask arrays; A rejection sampling unit for performing rejection sampling based on the sampling result to obtain an output token; An adding unit for adding the output token to the output token array and using the output token array as the output text.

7. An apparatus for training a text generation model, characterized in that, Comprising: An obtaining unit for obtaining the input text; A determining unit for determining an input token vector according to the input text, where the input token vector includes a plurality of consecutive masks; A processing unit for inputting the input token vector into the text generation model to obtain an output result, where the output result includes candidate output tokens and their probability distributions; A sampling unit for sampling the output result to obtain an output token and its probability; A calculating unit for calculating a loss value according to the output token and its probability; A training unit for training the text generation model according to the loss value.

8. An apparatus for generating text, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 4 is implemented.

9. An apparatus for training a text generation model, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method according to claim 5 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Cited By

  • Method for training model

    CN120928978A

  • Model prediction result adjustment method and device, electronic equipment and storage medium

    CN121303387A

  • AI robot low-delay voice interaction method and system

    CN121905158A