Text generation method, method for training text generation model, and related device

By introducing multiple continuous masks into the large language model and verifying candidate word elements using reject sampling, efficient parallel decoding of text generation is achieved, and the problem of low autoregressive decoding efficiency is solved and the consistency of generation effect is ensured.

WO2025139386A1PCT designated stage expired Publication Date: 2025-07-03SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD

Patent Information

Application Number
PCT/CN2024/130440
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-30
Filing Date
2024-11-07
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the existing text generation methods, the autoregressive decoding efficiency of large language models is low, and it is impossible to effectively utilize the processor's parallel computing power, resulting in a slow inference process.

Method used

The semi-autoregressive decoding method is adopted, and multiple continuous masks are introduced into the input word element vector, allowing the text generation model to decode multiple word elements in parallel in one inference, and verify the candidate word elements by rejecting sampling to ensure that the result is consistent with the autoregressive decoding.

Benefits of technology

While improving the decoding speed, the generation effect is ensured without changing the model structure or introducing additional training parameters, which improves the efficiency of text generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024130440_03072025_PF_FP_ABST
    Figure CN2024130440_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application is applicable to the technical field of artificial intelligence, and provides a text generation method, a method for training a text generation model, and a related device. The method comprises: acquiring an input text; determining an input token vector on the basis of the input text, wherein the input token vector comprises an input token array, an output token array, an output mask array, N candidate tokens, and N candidate mask arrays; inputting the input token vector into a text generation model to obtain an output result, wherein the output result comprises candidate output tokens and a probability distribution thereof; sampling the output result to obtain a sampling result; performing rejection sampling on the basis of the sampling result to obtain an output token; and adding the output token into the output token array, and using the output token array as an output text. The method in an embodiment of the present application can ensure generation performance while improving decoding speed and does not rely on any additional models.
Need to check novelty before this filing date? Find Prior Art

Description

Method for generating text, method for training text generation model and related devices

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 30, 2023, with application number 202311873965.7 and invention name “Method for generating text, method for training text generation model and related devices”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application belongs to the field of artificial intelligence technology, and in particular relates to a method for generating text, a method for training a text generation model, and related devices. Background Art

[0003] With the rapid development of artificial intelligence (AI) technology, text generation is becoming increasingly popular. For example, large language models (LLMs) can be used to generate output text, enabling machines to "speak" in a manner similar to human language. However, existing text generation methods still have some problems. For example, traditional large language models typically use autoregressive decoding, which requires serial decoding of each token during inference, resulting in a slow inference process and inability to effectively utilize the parallel computing capabilities of the processor.

[0004] Summary of the Invention

[0005] The embodiments of the present application provide a method for generating text, a method for training a text generation model, and related devices, which can improve the decoding speed while ensuring the generation effect, and do not rely on other additional models.

[0006] In a first aspect, an embodiment of the present application provides a method for generating text, comprising:

[0007] Get input text;

[0008] Determining an input word-unit vector according to the input text, the input word-unit vector comprising an input word-unit array, an output word-unit array, an output mask array, N candidate word-units, and N candidate mask arrays, each of the N candidate mask arrays comprising N consecutive masks, the N candidate mask arrays corresponding one-to-one to the N candidate word-units, where N is a positive integer greater than 1;

[0009] Inputting the input word unit vector into a text generation model to obtain an output result, wherein the output result includes candidate output word units and their probability distribution;

[0010] Sampling the output results to obtain sampling results, the sampling results including predicted word-grams corresponding to the output word-gram array and their probabilities, predicted word-grams corresponding to the output mask array and their probabilities, N predicted word-grams corresponding to the N candidate word-grams and their probabilities, and N groups of predicted word-grams corresponding to the N mask arrays and their probabilities;

[0011] Perform rejection sampling based on the sampling result to obtain an output word;

[0012] The output word-gram is added to the output word-gram array, and the output word-gram array is used as the output text.

[0013] In an embodiment of the present application, the input word element vector includes an input word element array, an output word element array, an output mask array, N candidate word elements and N candidate mask arrays. After processing the input word element vector, a sampling result containing multiple predicted word elements can be obtained. At this time, rejection sampling is performed based on the sampling result to ensure that the obtained output word element and the result obtained by autoregressive decoding obey the same distribution. In this way, the generation effect can be guaranteed while improving the decoding speed, and there is no need to rely on other additional models.

[0014] In some possible implementations, performing rejection sampling based on the sampling result to obtain an output word-unit includes:

[0015] Determining whether to accept each of the N candidate word-grams in turn based on a rejection sampling algorithm;

[0016] If the first candidate word-gram in the N candidate word-grams is rejected, the predicted word-gram corresponding to the last word-gram in the output word-gram array is added to the output word-gram array, and the predicted word-gram corresponding to the output mask array is used as the candidate word-gram;

[0017] If the i-th candidate word among the N candidate word-grams is accepted, then the i-th candidate word-gram is added to the output word-gram array;

[0018] If the i-th candidate word among the N candidate words is rejected, where i is an integer greater than 1 and less than or equal to N, then the judgment is stopped, and the predicted word corresponding to the candidate mask array corresponding to the i-1-th word is taken as the candidate word, and the predicted word corresponding to the i-1-th candidate word is added to the output word array.

[0019] In some possible implementations, inputting the input word element vector into a text generation model to obtain an output result includes:

[0020] Inputting the input word element vector, attention matrix and position encoding into the text generation model to obtain the output result;

[0021] The attention matrix satisfies the following conditions:

[0022] When row number i is greater than or equal to column number j, and the i-th word in the input word-unit vector is not a mask, then the element in row number i and column number j in the attention matrix is ​​1, and i and j are integers greater than or equal to 0 and less than M, where M is the number of words in the input word-unit vector;

[0023] When the number of rows i is greater than or equal to the number of columns j, and the i-th word in the input word-unit vector is a mask, if the j-th word in the input word-unit vector is also a mask and ij < N, then the element with row i and column j in the attention matrix is ​​1;

[0024] Then the other elements in the attention matrix are all 0;

[0025] The position code satisfies the following conditions:

[0026] The i-th element in the position encoding is the sum of all elements in the i-th row of the attention matrix minus 1.

[0027] In an embodiment of the present application, through the attention matrix and the position encoding designed in the present application, it is possible to predict and generate multiple word elements (output mask array and predicted word elements corresponding to N candidate mask arrays) in one inference of the text generation model, while verifying the N candidate word elements obtained in the previous step of reasoning, and ensure that the two do not interfere with each other. In this way, the generation effect can be guaranteed while improving the decoding speed.

[0028] In some possible implementations, the text generation model is a large language model.

[0029] In a second aspect, an embodiment of the present application provides a method for training a text generation model, comprising:

[0030] Get input text;

[0031] Determining an input word element vector according to the input text, wherein the input word element vector includes a plurality of consecutive masks;

[0032] Inputting the input word unit vector into the text generation model to obtain an output result, wherein the output result includes candidate output word units and their probability distribution;

[0033] Sampling the output result to obtain output word units and their probabilities;

[0034] Calculating a loss value based on the output word and its probability;

[0035] The text generation model is trained according to the loss value.

[0036] In an embodiment of the present application, the input word element vector includes multiple continuous masks, so that multiple output word elements can be obtained by inputting the input word element vector into the text generation model for one reasoning, thereby improving the efficiency of text generation.

[0037] In a third aspect, an embodiment of the present application provides a device for generating text, including:

[0038] Get unit, used to get input text;

[0039] a determining unit, configured to determine an input word-meta vector based on the input text, the input word-meta vector comprising an input word-meta array, an output word-meta array, an output mask array, N candidate word-meta and N candidate mask arrays, each of the N candidate mask arrays comprising N consecutive masks, the N candidate mask arrays corresponding one-to-one to the N candidate word-meta, where N is a positive integer greater than 1;

[0040] a processing unit, configured to input the input word unit vector into a text generation model to obtain an output result, wherein the output result includes candidate output word units and their probability distribution;

[0041] a sampling unit, configured to sample the output result to obtain a sampling result, wherein the sampling result includes a predicted word-gram corresponding to the output word-gram array and its probability, a predicted word-gram corresponding to the output mask array and its probability, N predicted word-grams corresponding to the N candidate word-grams and their probabilities, and N groups of predicted word-grams corresponding to the N mask arrays and their probabilities;

[0042] a rejection sampling unit, configured to perform rejection sampling based on the sampling result to obtain an output word;

[0043] The adding unit is used to add the output word element to the output word element array and use the output word element array as the output text.

[0044] In a fourth aspect, an embodiment of the present application provides a device for training a text generation model, comprising:

[0045] Get unit, used to get input text;

[0046] a determining unit, configured to determine an input word element vector according to the input text, wherein the input word element vector includes a plurality of consecutive masks;

[0047] a processing unit, configured to input the input word-unit vector into the text generation model to obtain an output result, wherein the output result includes candidate output word-units and their probability distribution;

[0048] a sampling unit, configured to sample the output result to obtain output word units and their probabilities;

[0049] a calculation unit, configured to calculate a loss value based on the output word and its probability;

[0050] A training unit is used to train the text generation model according to the loss value.

[0051] In a fifth aspect, an embodiment of the present application provides a device for generating text, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any one of the methods described in the first aspect above when executing the computer program.

[0052] In the sixth aspect, an embodiment of the present application provides a device for training a text generation model, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any one of the methods described in the second aspect when executing the computer program.

[0053] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in any one of the above-mentioned first aspects are implemented.

[0054] In an eighth aspect, an embodiment of the present application provides a computer program product. When the computer program product is run on an evaluation device, the evaluation device executes any one of the methods described in the first aspect above.

[0055] It can be understood that the beneficial effects of the third to eighth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.

[0056] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0057] In an embodiment of the present application, the input word element vector includes an input word element array, an output word element array, an output mask array, N candidate word elements and N candidate mask arrays. After processing the input word element vector, a sampling result containing multiple predicted word elements can be obtained. At this time, rejection sampling is performed based on the sampling result to ensure that the obtained output word element and the result obtained by autoregressive decoding obey the same distribution. In this way, the generation effect can be guaranteed while improving the decoding speed, and there is no need to rely on other additional models. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] FIG1 is a schematic diagram of an application scenario provided in an embodiment of the present application.

[0059] FIG2 is a schematic diagram of text generation based on autoregressive decoding in an embodiment of the present application.

[0060] FIG3 is a schematic flowchart of a method for training a text generation model provided in one embodiment of the present application.

[0061] FIG4 is a schematic diagram of generating text based on semi-autoregressive decoding in an embodiment of the present application.

[0062] FIG5 is a schematic flowchart of a method for generating text provided in one embodiment of the present application.

[0063] FIG6 is a schematic flowchart of a method for generating text provided in another embodiment of the present application.

[0064] FIG7 is a schematic diagram of an input word element vector in an embodiment of the present application.

[0065] Figure 8 is a schematic diagram of the self-attention matrix and position encoding in one embodiment of the present application.

[0066] FIG9 is a schematic diagram of reasoning using a large language model in one embodiment of the present application.

[0067] FIG10 is a schematic structural diagram of an apparatus for training a text generation model provided in one embodiment of the present application.

[0068] FIG11 is a schematic structural diagram of an apparatus for generating text provided in an embodiment of the present application.

[0069] FIG12 is a schematic structural diagram of a device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0070] The solutions provided in the embodiments of the present application can be applied to various scenarios where text generation is required. For example, the methods for generating text and training text generation models provided in the embodiments of the present application can be executed on a server, in the cloud, or on a terminal device.

[0071] Taking the terminal device as an example, as shown in Figure 1, the technical solution of the embodiment of the present invention can be applied to the terminal device. The method of generating text in the embodiment of the present application can be based on reasoning (or processing) of the input text to obtain the output text corresponding to the input text. The method of training the text generation model provided in the embodiment of the present application can train the text generation model used to generate text.

[0072] The terminal device in the embodiments of the present application may be mobile or fixed. For example, the terminal device may be a mobile phone with a text generation function (such as deploying a text generation model), a tablet personal computer (TPC), a media player, a smart TV, a laptop computer (LC), a personal digital assistant (PDA), a personal computer (PC), a camera, a camcorder, a smart watch, a wearable device (WD) or an autonomous driving vehicle, etc. The embodiments of the present invention are not limited to this.

[0073] The text generation model in this application can be a large language model (LLM). The following uses the large language model as an example to illustrate the technical problems existing in the process of generating text.

[0074] Generally, when generating text based on a large language model, the large language model can use autoregressive decoding, that is, serial decoding of each token. This method is relatively inefficient. For example, as shown in Figure 2, suppose that at time t=0, the input text is " <s>"I love Beijing". The input text can be encoded into input tokens "1, 45, 96, 68" by a tokenizer. The input tokens are input into a large language model to obtain an output token "52" (this token corresponds to "day"); at time t = 1, the output token obtained at time t = 0 is added to the input text, that is, the input text at time t = 1 is " <s>"I love Beijing Tian” is input into the large language model, and the output token "34” (corresponding to "An”) is obtained; finally, through 5 inferences, the large language model can infer 5 tokens "52, 34, 28, 6, 2”, and the corresponding text is "Tiananmen Square. < / s> ". It can be seen that autoregressive decoding can only infer one output token at a time, and the inference efficiency is low.

[0075] To improve the inference speed, the industry has proposed many optimization methods for the inference stage. For example, improving the implementation of the computing core, multi-card parallel computing, batch processing strategies, quantization pruning, etc. Among them, speculative decoding is a method proven to be effective in practice.

[0076] Speculative sampling is an advanced large model inference acceleration technology that can achieve an acceleration ratio of more than 3 times without sacrificing the generation effect. Speculative sampling can use a small model (which can be called a draft model) for autoregressive sampling and generate multiple candidate tokens, and then use the LLM to evaluate the sampling results (the generated multiple candidate tokens). This processing process is similar to the associative input in the input method: let the small model "guess" the tokens that the LLM may generate next, and then let the LLM verify multiple candidate tokens at the same time. Since the LLM can verify multiple candidate tokens at the same time, and the small model has a fast sampling speed, this method can quickly generate multiple candidate tokens, can achieve a significant acceleration effect, and at the same time ensure that the sampling distribution of the multiple candidate tokens generated is exactly the same as the result obtained using the LLM.

[0077] However, this method requires the use of an additional small model, which will lead to an increase in memory overhead, and the accuracy of the small model has a great impact on the inference acceleration effect. In some cases, the predictions of the small model may not be accurate at all. In this way, on the contrary, the acceleration effect will be reduced due to the introduction of additional computational overhead.

[0078] Another similar method is the blockwise parallel decoding method, which also adopts the idea of predicting first and then verifying. This method does not require the introduction of an additional small model, but directly trains multiple classification heads in the last layer of the LLM, so that the LLM can predict multiple tokens next at the same time during one inference, and then use the original LLM to verify these multiple tokens.

[0079] However, this method requires changing the model structure, and at the same time introduces additional model training parameters, which will also bring additional memory overhead.

[0080] In order to solve one or more of the above technical problems, the present application proposes a method for generating text, a method for training a text generation model, and related devices, which can improve the decoding speed while ensuring the generation effect. The method for generating text proposed in the present application also adopts the idea of ​​first predicting and then verifying, but this method does not require changing the model structure and does not require introducing additional training parameters or models. In this way, no additional overhead (such as introducing additional parameters or models) is added, thus having higher versatility.

[0081] During the model training phase, the method for training a text generation model proposed in this invention introduces multiple consecutive masked tokens (i.e., [MASK], also referred to as masks) during the supervised fine tuning (SFT) process (or the training process), enabling the text generation model to decode multiple tokens in parallel using semi-autoregressive decoding.

[0082] During the inference stage, the text generation method proposed in the present invention uses rejection sampling to verify the candidate word units output by the text generation model, ensuring that the results of parallel decoding and the results of autoregressive sampling follow the same distribution. This not only improves the decoding speed, but also ensures the generation effect without introducing additional parameters or models.

[0083] The semi-autoregressive decoding mentioned above is common in translation models. Its principle is to divide the entire translation into k blocks (k is a positive integer), perform non-autoregressive decoding within the block, and perform autoregressive decoding between blocks. In this way, multiple consecutive words can be generated in parallel at each time step.

[0084] The following is a detailed explanation of the method for training a text generation model in an embodiment of the present application with reference to FIG3 .

[0085] FIG3 shows a schematic flowchart of a method for training a text generation model provided in one embodiment of the present application. As an example and not a limitation, the method can be applied to the terminal device shown in FIG1 , or to a server or cloud.

[0086] The method 300 in FIG3 includes steps S310 to S360, which are specifically as follows:

[0087] S310: Obtain input text.

[0088] S320: Determine an input word element vector according to the input text.

[0089] The input word element vector may include multiple consecutive masks (MASK).

[0090] S330: Input the input word element vector into the text generation model to obtain an output result.

[0091] The output result may include candidate output word-grams and their probability distribution.

[0092] S340: Sampling the output result to obtain output word units and their probabilities.

[0093] S350: Calculate a loss value based on the output word and its probability.

[0094] When training a text generation model, the input text can be part of a training sample (the training sample here can be a sentence, paragraph, or article, etc.). Then, when calculating the loss value, the other parts of the training sample except the input sample can be used as the true value of the input sample.

[0095] For example, the training sample can be " <s>I love Tiananmen Square in Beijing. < / s> ", the input text in S310 can be " <s>I love Beijing”, then, in S350, "Tiananmen Square. < / s> ” (that is, the rest of the training sample except the input sample) is used as the true value to calculate the loss value corresponding to the input sample.

[0096] S360: Train the text generation model according to the loss value.

[0097] For example, the model parameters of the text generation model may be adjusted according to the loss value.

[0098] In an embodiment of the present application, the input word element vector may include multiple continuous masks, so that multiple output word elements can be obtained by inputting the input word element vector into the text generation model for one inference, thereby improving the efficiency of text generation.

[0099] For example, as shown in Figure 4, suppose that at time t=0, the input text is " <s>For "I love Beijing”, the input text can be encoded into input tokens "1, 45, 96, 68” through a tokenizer. 4 masks are added to the input tokens (for example, the token "3” in Figure 4 can represent a mask), and the input tokens with masks added are input into the large language model. At this time, since there are 4 masks in the input tokens, 4 output tokens "52, 34, 28, 6, 2” can be directly obtained through one inference, and the corresponding text is "Tiananmen Square.< / s> "In this way, multiple output tokens can be obtained through one inference, which can improve the efficiency of text generation.

[0100] The method for generating text in the embodiment of the present application is described in detail below with reference to FIG5 and FIG6 .

[0101] FIG5 shows a schematic flowchart of a method for generating text provided in an embodiment of the present application. As an example and not a limitation, the method can be applied to the terminal device shown in FIG1 , or to a server or cloud.

[0102] The method 500 in FIG5 includes steps S510 to S560, which are specifically as follows:

[0103] S510, obtaining input text;

[0104] S520: Determine an input word element vector according to the input text.

[0105] The input word-unit vector may include an input word-unit array, an output word-unit array, an output mask array, N candidate word-units, and N candidate mask arrays, where N is a positive integer greater than 1.

[0106] Each of the N candidate mask arrays may include N consecutive masks, and the N candidate mask arrays may correspond one-to-one to the N candidate word elements.

[0107] In some embodiments, when the candidate word array L c When it is not empty, the following rules can be used to construct the input word element vector, denoted as I, as follows: I=T+L a +M+L c [0]+M0+...+L c [N-1]+M N-1

[0108] Among them, "+" represents array vector concatenation, T represents the input word array, L a represents the output word array, M represents the output mask array, L c [i] represents the candidate word array L c The i-th element of M i Represents the i-th candidate mask array, each candidate mask array includes N masks, that is, M i =[MASK]*N,L c [i] and M i Can correspond one to one.

[0109] In some embodiments, when L c When it is empty, the input word vector is constructed as: I=T+L a +M

[0110] S530: Input the input word-unit vector into a text generation model to obtain an output result, wherein the output result includes candidate predicted word-units and their probability distribution.

[0111] In some embodiments, the text generation model may be a large language model.

[0112] S540: Sampling the output result to obtain a sampling result.

[0113] In some embodiments, greedy decoding may be used, or polynomial sampling or kernel sampling may be used to sample the output result to obtain a sampling result.

[0114] The sampling result may include the predicted token and its probability corresponding to the output token array, the predicted token and its probability corresponding to the output mask array, the N predicted tokens and their probabilities corresponding to the N candidate tokens, and the N groups of predicted tokens and their probabilities corresponding to the N mask arrays.

[0115] In some embodiments, the input token vector, attention matrix, and position encoding may be input into the text generation model to obtain the output result.

[0116] Among them, the attention matrix and the position encoding may be related to the masks in the input token vector (such as the output mask array and N candidate mask arrays), so that the prediction generation process (predicting and generating multiple tokens) and the candidate token verification process (verifying the N candidate tokens obtained from the previous inference) in one inference of the text generation model do not interfere with each other, thereby ensuring the generation effect while improving the decoding speed.

[0117] In some embodiments, the attention matrix may satisfy the following conditions:

[0118] When the row number i is greater than or equal to the column number j, and the i-th token in the input token vector is not a mask, the element at row i and column j in the attention matrix is 1, where i and j are integers greater than or equal to 0 and less than M, and M is the number of tokens in the input token vector;

[0119] When the row number i is greater than or equal to the column number j, and the i-th token in the input token vector is a mask, if the j-th token in the input token vector is also a mask and i - j < N, the element at row i and column j in the attention matrix is 1;

[0120] Then all other elements in the attention matrix are 0.

[0121] In some embodiments, the position encoding may satisfy the following conditions:

[0122] The i-th element in the position encoding is the sum of all elements in the i-th row of the attention matrix minus 1.

[0123] In the embodiments of the present application, through the designed attention matrix and position encoding of the present application, while predicting and generating multiple tokens (the predicted tokens corresponding to the output mask array and N candidate mask arrays) in one inference of the text generation model, the N candidate tokens obtained from the previous inference can be verified, and it is ensured that the two do not interfere with each other. In this way, the generation effect can be ensured while improving the decoding speed.

[0124] S550, perform rejection sampling based on the sampling result to obtain the output token.

[0125] In some embodiments, performing rejection sampling based on the sampling result to obtain an output word-unit includes:

[0126] Determine in sequence whether to accept each of the N candidate word-grams based on a rejection sampling algorithm;

[0127] If the first candidate word-gram in the N candidate word-grams is rejected, the predicted word-gram corresponding to the last word-gram in the output word-gram array is added to the output word-gram array, and the predicted word-gram corresponding to the output mask array is used as the candidate word-gram;

[0128] If the i-th candidate word among the N candidate word-grams is accepted, then the i-th candidate word-gram is added to the output word-gram array;

[0129] If the i-th candidate word among the N candidate words is rejected, where i is a positive integer greater than 1 and less than or equal to N, then the judgment is stopped, and the predicted word corresponding to the candidate mask array corresponding to the i-1-th word is taken as the candidate word, and the predicted word corresponding to the i-1-th candidate word is added to the output word array.

[0130] S560: Add the output word to the output word array, and use the output word array as output text.

[0131] In an embodiment of the present application, the input word element vector includes an input word element array, an output word element array, an output mask array, N candidate word elements and N candidate mask arrays. After processing the input word element vector, a sampling result containing multiple predicted word elements can be obtained. At this time, rejection sampling is performed based on the sampling result to ensure that the obtained output word element and the result obtained by autoregressive decoding obey the same distribution. In this way, the generation effect can be guaranteed while improving the decoding speed, and there is no need to rely on other additional models.

[0132] The above text generating method will be described below with reference to FIG6 through a specific embodiment.

[0133] FIG6 shows a schematic flowchart of a method for generating text provided in another embodiment of the present application. As an example and not a limitation, the method can be applied to the terminal device shown in FIG1 , or to a server or cloud.

[0134] The method 600 in FIG6 includes steps S610 to S650, which are specifically as follows:

[0135] S610, initialize data.

[0136] The output word array and candidate word array can be initialized. For example, L a =[],L c =[], where L a Represents the output word array, L c Represents the candidate word array obtained by the last inference of the text generation model.

[0137] S620: Construct an input word element vector.

[0138] The input word-gram vector may include an input word-gram array, an output word-gram array, an output mask array, N candidate word-grams, and N candidate mask arrays.

[0139] In some embodiments, when the candidate word array L c When it is not empty, the following rules can be used to construct the input word element vector, denoted as I, as follows: I=T+L a +M+L c [0]+M0+...+L c [N-1]+M N-1

[0140] Among them, "+" represents array vector concatenation, T represents the input word array, L a represents the output word array, M represents the output mask array, L c [i] represents the candidate word array L c The i-th element of M i Represents the i-th candidate mask array, each candidate mask array includes N masks, that is, M i =[MASK]*N,L c [i] and M i Can correspond one to one.

[0141] In some embodiments, when L c When it is empty, the constructed input word element vector can be: I=T+L a +M

[0142] For example, when N=2, input text S=” <s>I love", after encoding the input text S, we can get the input word element vector T, T = [1, 45, 96], L a =[68,52],L c =[34,28], then the constructed input word element vector can be shown in Figure 7.

[0143] In particular, at t=0, L a and L c are all empty. In this case, the input word element vector can be T+M.

[0144] S630, construct input attention matrix and position encoding.

[0145] Traditional attention matrices are lower triangular, meaning that the preceding word cannot "notice" the following word during self-attention calculations. To enable simultaneous word verification and prediction within a large language model inference, this application designs the attention matrix and position encoding as follows.

[0146] For example, let the length of the input word element vector I be L, then the following attention matrix A can be defined: A∈[0,1] L*L ,

[0147] The attention matrix A can satisfy the following conditions:

[0148] 1) When i≥j and I[j] is not [MASK], A ij =1;

[0149] 2) When i≥j and I[j] is [MASK], if I[i] is also [MASK] and ij<N, A ij =1;

[0150] 3) In other cases, A ij =0.

[0151] At the same time, the position code P can be defined as: P∈N L ,

[0152] Among them, the position code P can satisfy

[0153] Through the above design, candidate word units (corresponding to M and M) can be predicted and generated simultaneously in one inference of the text generation model. i Part), and the candidate word element decoded in the previous step (corresponding to L c part) for verification, and ensure that the two do not interfere with each other.

[0154] For example, when N=2, input text S=” <s>I love", after encoding the input text S, the input word element vector T = [1, 45, 96], L a =[68,52],L c =[34,28], the self-attention matrix and position encoding designed by the method of this application can be shown in Figure 8.

[0155] S640, using a large language model for reasoning.

[0156] The constructed input word element vector, attention matrix and position encoding can be input into the text generation model to obtain the inference results of the large language model, and the output word element and its probability can be obtained by sampling. The schematic diagram is shown in Figure 9.

[0157] Among them, L can be defined a [-1](L a [-1] can represent the output word array L a The last element of L c [i] The probability distribution of the predicted word at the corresponding position is P a and P c [i], greedy decoding can be used, or polynomial sampling or kernel sampling can be used to decode L a [-1] and L c [i] Sampling, record L a [-1] and L c [i] The predicted word units obtained after sampling are Y a and Y c [i], and the corresponding probability is P a (Y a ) and P c [i](Y c [i]), at the same time, remember M and M i The sampled words are Y m and The sampling process can be shown in FIG9 .

[0158] S650: Obtain output word-units based on rejection sampling.

[0159] When the text generation model is used for inference for the first time, L c If it is empty, you can set L a [-1] corresponds to the output word Y a Add to the output word array L a middle.

[0160] Otherwise, if L c If it is not empty (not the first time the text generation model is used for inference), it can be determined in turn based on the rejection sampling algorithm whether to accept each of the N candidate word-grams.

[0161] If the first candidate word L among N candidate words is rejected c [0], then Y a Add to the output word array L a And the predicted word Y corresponding to M m As a candidate word;

[0162] If we accept the i-th candidate word L among N candidate words c [i-1], then L c [i-1] is added to the output word array L a middle;

[0163] If the i-th candidate word L among N candidate words is rejected c [i-1], then we can stop judging and set M i-1 The corresponding predicted word As a candidate word, and Y c [i-2] Add to the output word array L a middle.

[0164] For example, L c [i-1] can represent the i-th candidate word obtained by the last decoding, Q c [i-1] can represent the L generated by the last decoding c The probability corresponding to [i-1].

[0165] For the first candidate word L c [0], P a (L c [0]) represents the candidate word L c [0] In the probability distribution P a (P a is the output word array L a The probability of the predicted word unit obtained by the last element of the probability distribution) can be uniformly sampled from [0,1] A number r[0], at this time, the judgment condition can be: r[0]≤P a (L c [0]) / Q c [0]

[0166] That is, if r[0]≤P a (L c [0]) / Q c [0], then the first candidate word L can be accepted c [0], L c [0] Add to the output word array L a and continue to verify the next candidate word; otherwise, the first candidate word L can be rejected. c [0], from P a Sampling a word (i.e. Y defined above) a ) is added to the output word array L a and exit the verification phase.

[0167] For the i-th (where i represents an integer greater than 1 and less than or equal to N) candidate word L c [i-1], P c [i-1](L c [i-1]) represents the candidate word L c [i-1] in the probability distribution P c [i-1](P c [i-1] is the candidate word L c [i-1]) can be uniformly sampled from [0,1] and recorded as r[i-1]. At this time, the judgment condition can be:

[0168] r[i-1]≤P c [i-1](L c [i-1]) / Q c [i-1]

[0169] That is, if r[i-1]≤P c [i-1](L c [i-1]) / Q c [i-1], then the i-th candidate word L can be accepted c [i-1], L c [i-1] is added to the output word array L a and continue to verify the next candidate word; otherwise, the i-th candidate word L can be rejected c [i-1], from P c [i-2] samples a word (that is, Y defined above) c [i-2]) is added to the output word array L a and exit the verification phase.

[0170] Furthermore, the above steps S620 and S650 may be repeated until L a Contains the end of document marker (in this case 1, the corresponding marker is< / s> ).

[0171] It should be noted that, in the above embodiment, each time L a Each new word needs to be judged once, and the new word added to L a Whether the word in is the end mark.

[0172] In the above embodiment, L a Records the inference results of the text generation model (i.e., output word units). a Receiving more than one word-gram is equivalent to outputting multiple word-grams in a single inference, which can accelerate inference. Furthermore, by using rejection sampling in step S650, the inference results obtained can be guaranteed to follow the same probability distribution as those obtained by traditional autoregressive decoding.

[0173] The method embodiment of the present application is described in detail above in conjunction with Figures 1 to 9 . The device embodiment of the present application is described in detail below in conjunction with Figures 10 and 12 . It should be understood that the description of the method embodiment corresponds to the description of the device embodiment. Therefore, for parts not described in detail, reference can be made to the above method embodiment.

[0174] Figure 10 is a schematic structural diagram of an apparatus for training a text generation model according to an embodiment of the present application. As shown in Figure 10 , the apparatus 1000 includes an acquisition unit 1010, a determination unit 1020, a processing unit 1030, a sampling unit 1040, a calculation unit 1050, and a training unit 1060, specifically as follows:

[0175] An acquisition unit 1010 is used to acquire input text;

[0176] a determining unit 1020, configured to determine an input word element vector according to the input text, wherein the input word element vector includes a plurality of consecutive masks;

[0177] A processing unit 1030 is configured to input the input word-unit vector into the text generation model to obtain an output result, wherein the output result includes candidate output word-units and their probability distribution;

[0178] A sampling unit 1040 is configured to sample the output result to obtain output word units and their probabilities;

[0179] A calculation unit 1050 is configured to calculate a loss value based on the output word and its probability;

[0180] The training unit 1060 is configured to train the text generation model according to the loss value.

[0181] Figure 11 is a schematic structural diagram of a device for generating text provided in an embodiment of the present application. As shown in Figure 11, the device 1100 includes an acquisition unit 1110, a determination unit 1120, a processing unit 1130, a sampling unit 1140, a rejection sampling unit 1150, and an adding unit 1160, specifically as follows:

[0182] An acquisition unit 1110 is used to acquire input text;

[0183] a determining unit 1120 configured to determine an input word-meta vector based on the input text, the input word-meta vector comprising an input word-meta array, an output word-meta array, an output mask array, N candidate word-meta arrays, and N candidate mask arrays, each of the N candidate mask arrays comprising N consecutive masks, the N candidate mask arrays corresponding one-to-one to the N candidate word-meta arrays, where N is a positive integer greater than 1;

[0184] A processing unit 1130 is configured to input the input word-unit vector into a text generation model to obtain an output result, wherein the output result includes candidate output word-units and their probability distribution;

[0185] a sampling unit 1140 configured to sample the output result to obtain a sampling result, wherein the sampling result includes a predicted word-gram corresponding to the output word-gram array and its probability, a predicted word-gram corresponding to the output mask array and its probability, N predicted word-grams corresponding to the N candidate word-grams and their probabilities, and N groups of predicted word-grams corresponding to the N mask arrays and their probabilities;

[0186] A rejection sampling unit 1150 is configured to perform rejection sampling based on the sampling result to obtain an output word;

[0187] The adding unit 1160 is configured to add the output word to the output word array and use the output word array as the output text.

[0188] Figure 12 is a schematic diagram of an apparatus provided in an embodiment of the present application. As shown in Figure 12, apparatus 1200 of this embodiment includes: a processor 1210, a memory 1220, and a computer program 1230 stored in the memory 1220 and executable on the processor 1210. When the processor 1210 executes the computer program 1230, the steps of each of the aforementioned method embodiments are implemented. Alternatively, when the processor 1210 executes the computer program 1230, the functions of each module in each of the aforementioned apparatus embodiments are implemented.

[0189] For example, the computer program 1230 may be divided into one or more modules, which are stored in the memory 1220 and executed by the processor 1210 to implement the present application. The one or more modules may be a series of computer program instruction segments capable of implementing specific functions, and the instructions are used to describe the execution process of the computer program 1230 in the device 1200.

[0190] The apparatus 1200 may be a computing device such as a desktop computer, laptop, PDA, or cloud server. The apparatus may include, but is not limited to, a processor 1210 and a memory 1220. Those skilled in the art will appreciate that FIG12 is merely an example of apparatus 1200 and does not limit the apparatus 1200. The apparatus 1200 may include more or fewer components than shown, or may combine certain components or different components. For example, the apparatus may also include input / output devices, network access devices, buses, and the like.

[0191] The processor 510 may be a graphics processing unit (GPU), a neural-network processing unit (NPU), or a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0192] The memory 520 may be an internal storage unit of the apparatus 500, such as a hard disk or memory of the apparatus 500. The memory 520 may also be an external storage device of the apparatus 500, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the apparatus 500. Furthermore, the memory 520 may include both an internal storage unit of the apparatus 500 and an external storage device. The memory 520 is used to store the computer program and other programs and data required by the apparatus 500. The memory 520 may also be used to temporarily store data that has been output or is to be output.

[0193] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0194] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.

[0195] An embodiment of the present application provides a computer program product. When the computer program product is run on an evaluation device, the evaluation device can implement the steps in the above-mentioned method embodiments when the computer program product is executed.

[0196] If the integrated unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can at least include: any entity or device capable of carrying computer program code to a camera / terminal device, a recording medium, computer memory, read-only memory (ROM), random access memory (RAM), an electric carrier signal, a telecommunications signal, and a software distribution medium. Examples include a USB flash drive, a removable hard drive, a magnetic disk, or an optical disk. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electric carrier signals or telecommunications signals.< / s> < / s>

Claims

1. A method for generating text, characterized in that, Comprising: Obtain input text; Determine an input token vector according to the input text, the input token vector including an input token array, an output token array, an output mask array, N candidate tokens and N candidate mask arrays, each candidate mask array in the N candidate mask arrays including N consecutive masks, the N candidate mask arrays corresponding one-to-one to the N candidate tokens, N being a positive integer greater than 1; Input the input token vector into a text generation model to obtain an output result, the output result including candidate output tokens and their probability distributions; Sample the output result to obtain a sampling result, the sampling result including the predicted tokens corresponding to the output token array and their probabilities, the predicted tokens corresponding to the output mask array and their probabilities, the N predicted tokens corresponding to the N candidate tokens and their probabilities, and the N groups of predicted tokens corresponding to the N mask arrays and their probabilities; Perform rejection sampling based on the sampling result to obtain output tokens; Add the output tokens to the output token array, and use the output token array as the output text.

2. The method according to claim 1, wherein The performing rejection sampling based on the sampling result to obtain output tokens includes: Successively determine whether to accept each candidate token in the N candidate tokens based on a rejection sampling algorithm; If the first candidate token in the N candidate tokens is rejected, add the predicted token corresponding to the last token in the output token array to the output token array, and use the predicted token corresponding to the output mask array as the candidate token; If the i-th candidate token in the N candidate tokens is accepted, add the i-th candidate token to the output token array; If the i-th candidate token in the N candidate tokens is rejected, i being an integer greater than 1 and less than or equal to N, stop the determination, use the predicted token corresponding to the candidate mask array corresponding to the (i - 1)-th token as the candidate token, and add the predicted token corresponding to the (i - 1)-th candidate token to the output token array.

3. The method according to claim 1 or 2, characterized in that, The inputting the input token vector into a text generation model to obtain an output result includes: Input the input token vector, an attention matrix and a position encoding into the text generation model to obtain the output result; Wherein, the attention matrix satisfies the following conditions: When the row number i is greater than or equal to the column number j and the i-th token in the input token vector is not a mask, the element at row i and column j in the attention matrix is 1, i and j being integers greater than or equal to 0 and less than M, M being the number of tokens in the input token vector; When the row number i is greater than or equal to the column number j and the i-th token in the input token vector is a mask, if the j-th token in the input token vector is also a mask and i - j < N, the element at row i and column j in the attention matrix is 1; Then all other elements in the attention matrix are 0; The position encoding satisfies the following conditions: The i-th element in the position encoding is the sum of all elements in the i-th row of the attention matrix minus 1.

4. The method according to claim 1, wherein The text generation model is a large language model.

5. A method for training a text generation model, characterized in that Comprising: Obtain the input text; Determine an input token vector according to the input text, where the input token vector includes a plurality of consecutive masks; Input the input token vector into the text generation model to obtain an output result, where the output result includes candidate output tokens and their probability distributions; Sample the output result to obtain an output token and its probability; Calculate a loss value according to the output token and its probability; Train the text generation model according to the loss value.

6. An apparatus for generating text, characterized in that, It includes: An obtaining unit for obtaining the input text; A determining unit for determining an input token vector according to the input text, where the input token vector includes an input token array, an output token array, an output mask array, N candidate tokens, and N candidate mask arrays, and each candidate mask array in the N candidate mask arrays includes N consecutive masks, and the N candidate mask arrays correspond one-to-one with the N candidate tokens, and N is a positive integer greater than 1; A processing unit for inputting the input token vector into the text generation model to obtain an output result, The output result includes candidate output tokens and their probability distributions; A sampling unit for sampling the output result to obtain a sampling result, where the sampling result includes the predicted token and its probability corresponding to the output token array, the predicted token and its probability corresponding to the output mask array, the N predicted tokens and their probabilities corresponding to the N candidate tokens, and the N groups of predicted tokens and their probabilities corresponding to the N mask arrays; A rejection sampling unit for performing rejection sampling based on the sampling result to obtain an output token; An adding unit for adding the output token to the output token array and using the output token array as the output text.

7. An apparatus for training a text generation model, characterized in that It includes: An obtaining unit for obtaining the input text; A determining unit for determining an input token vector according to the input text, where the input token vector includes a plurality of consecutive masks; A processing unit for inputting the input token vector into the text generation model to obtain an output result, where the output result includes candidate output tokens and their probability distributions; A sampling unit for sampling the output result to obtain an output token and its probability; A calculating unit for calculating a loss value according to the output token and its probability; A training unit for training the text generation model according to the loss value.

8. An apparatus for generating text, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method described in any one of claims 1 to 4 is implemented.

9. An apparatus for training a text generation model, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method described in claim 5 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Controlled text generation method based on variational auto-encoder hidden variable manipulation

    CN114492332A

  • Low-resource machine translation method for fusing domain terms by using semi-autoregression

    CN114492468A

  • Commodity classification method and training method, device, equipment, medium and product

    CN115858790A

  • Multi-scale joint text steganography method and system

    CN115952528A

  • Method for training speech recognition model, device and storage medium

    US20220310064A1

Cited By

  • Iterative text refining method, system and equipment based on fidelity and medium

    CN120706380A

  • Fidelity-based iterative text refinement method, system, device, and medium

    CN120706380B

  • Context content and sequence joint prediction method and system based on reinforcement learning

    CN121168441A

  • Large model reasoning acceleration method and device based on speculation sampling, equipment and medium

    CN121480743A

  • Conversation content generation method and electronic equipment

    CN121524331A