Limb generation method and device
The method improves model inference speed and accuracy by employing distinct decoding algorithms for target and guess words in natural language processing, addressing the issue of reduced precision in existing acceleration methods.
Patent Information
- Application Number
- CN202510258871.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-15
AI Technical Summary
During the decoding process, the existing accelerated model inference method has different output results from traditional methods due to the determinism of the random number generator, which reduces the accuracy of model inference.
Different decoding algorithms are used to decode candidate weights of target word elements and guess word elements, and weighted random sampling algorithm and greedy selection algorithm are used respectively to ensure accurate decoding of target word elements, while avoiding the impact of the decoding process of guess word elements on the target word elements.
It improves the speed and accuracy of model inference, ensures that the generated word results are consistent with traditional methods, and improves the user experience of the end-side device.
Smart Images

Figure CN120315674A_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, the field of model inference technology, and particularly relates to a method and device for generating tokens. Background Art
[0002] In related technologies, during the decoding process of accelerating the model inference method, additional tokens are decoded. Since the random number generator is deterministic, it changes the random sampling sequence when decoding the correct tokens; as a result, the output of the accelerated model inference method is different from that of the traditional inference method, reducing the accuracy of model inference. Summary of the Invention
[0003] In view of this, at least one method and device for generating tokens are provided in the embodiments of this application.
[0004] The technical solution of the embodiments of this application is implemented as follows:
[0005] In a first aspect, an embodiment of this application provides a method for generating tokens, including: obtaining input information of the target model in the current inference stage; the input information includes a first token and a second token; wherein, the first token includes the target token output in the previous inference stage; the second token includes the guessed token output in the previous inference stage; inferring and generating candidate weights of the first token and the second token based on the target model, where the candidate weights represent the possibility of a candidate token becoming the target token; decoding the candidate weights of the first token based on a first decoding algorithm to generate a first target token, and decoding the candidate weights of the second token based on a second decoding algorithm to generate a second target token; wherein, the first decoding algorithm is different from the second decoding algorithm.
[0006] In a second aspect, an embodiment of this application provides a device for generating tokens, including: an obtaining module, configured to obtain input information of the target model in the current inference stage; the input information includes a first token and a second token; wherein, the first token includes the target token output in the previous inference stage; the second token includes the guessed token output in the previous inference stage; a first generating module, configured to infer and generate candidate weights of the first token and the second token based on the target model, where the candidate weights represent the possibility of a candidate token becoming the target token; a second generating module, configured to decode the candidate weights of the first token based on a first decoding algorithm to generate a first target token, and decode the candidate weights of the second token based on a second decoding algorithm to generate a second target token; wherein, the first decoding algorithm is different from the second decoding algorithm.
[0007] In a third aspect, an embodiment of this application provides a computer device, including a memory and a processor, where the memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.
[0008] Fourthly, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, some or all of the steps in the above method are implemented.
[0009] Fifthly, an embodiment of the present application provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, some or all of the steps in the above method are implemented.
[0010] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, rather than limiting the technical solution of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to explain the technical solution of the present application.
[0012] Figure 1 It is a schematic flowchart of the implementation of a token generation method provided by an embodiment of the present application;
[0013] Figure 2 It is a schematic flowchart of the implementation of a token generation method provided by an embodiment of the present application;
[0014] Figure 3 It is a schematic flowchart of the implementation of a token generation method provided by an embodiment of the present application;
[0015] Figure 4 It is a schematic flowchart of the implementation of a token generation method provided by an embodiment of the present application;
[0016] Figure 5 It is a schematic flowchart of the implementation of a token generation method provided by an embodiment of the present application;
[0017] Figure 6 It is a schematic flowchart of the implementation of a token generation method provided by an embodiment of the present application;
[0018] Figure 7 It is a schematic flowchart of the implementation of a token generation method provided by an embodiment of the present application;
[0019] Figure 8 It is a schematic diagram of an inference process provided by an embodiment of the present application;
[0020] Figure 9 It is a schematic diagram of an inference process provided by an embodiment of the present application;
[0021] Figure 10 It is a schematic diagram of a decoding process provided by an embodiment of the present application;
[0022] Figure 11 A schematic diagram of an inference process provided by an embodiment of the present application;
[0023] Figure 12 A schematic diagram of the implementation process of a token generation method provided by an embodiment of the present application;
[0024] Figure 13 A schematic diagram of the composition structure of a token generation device provided by an embodiment of the present application;
[0025] Figure 14 A schematic diagram of the hardware entity of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0026] In order to make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0027] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict. The terms "first / second / third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing this application and are not intended to limit this application.
[0029] Currently, the inference of traditional large language models is generated one token at a time, and the generation speed is very slow. On edge devices with limited computing resources, it has a great impact on the user experience.
[0030] In related technologies, a solution for accelerating the inference process is proposed. A relatively common and classical method is to accelerate the inference calculation process of large models through the lookahead method. However, during the decoding process, additional tokens are decoded. Since the random number generator is deterministic, it changes the random sampling sequence when decoding the correct tokens; as a result, the output of the method for accelerating model inference is different from the output of the traditional inference method, reducing the accuracy of model inference.
[0031] Embodiments of the present application provide a token generation method, which can be executed by a processor of a computer device. Among them, the computer device can refer to devices with data processing capabilities such as servers, laptops, tablets, desktop computers, smart TVs, set-top boxes, mobile devices (such as mobile phones, portable video players, personal digital assistants, dedicated messaging devices, portable game devices), etc.
[0032] Figure 1 It is a schematic flowchart of the implementation of a token generation method provided by an embodiment of the present application, which can be executed by a processor of a computer device. As Figure 1 shown, the method includes the following steps S101 to S103, which will be described in combination with Figure 1 the steps presented.
[0033] Step S101, obtain the input information of the target model at the current inference stage; the input information includes a first token and a second token.
[0034] Among them, the first token includes the target token output in the previous inference stage; the second token includes the guessed token output in the previous inference stage.
[0035] In some embodiments, the target model can be a trained large language model, which is used for translation, summarization, prediction, text generation, etc. based on user input text data, speech data, etc.; among them, the large language model can include a convolutional neural network model, a recurrent neural network model, a self-attention mechanism neural network model, etc.
[0036] In some embodiments, when the target model executes an inference task, it can include multiple inference stages. For example, the first inference stage for inferring based on user input, the second inference stage for generating the next input based on the output data of the first inference stage and performing inference, until the last inference stage when the output data meets the preset requirements.
[0037] In some implementations, the input information can include multiple tokens, pictures, etc. During the internal inference stage of the model, the input information is usually converted into formats such as digital sequences, vectors, etc. that the machine can recognize and process.
[0038] In some embodiments, the input information in the current inference stage is generated based on the data output in the previous inference stage. The input information includes at least one first token and at least one second token. The first token is the correct and acceptable token output in the previous inference stage, and the second token is a guessed token guessed based on the acceptable tokens output in the previous inference stage and the initial user input.
[0039] Exemplarily, the target tokens output in the previous inference stage include "a" and "great", and the user initial input tokens include "It" and "is". Then, based on "It", "is", "a", and "great", the guessed tokens are obtained, including "who", "is", and "a". Then, in addition to the user initial input "It" and "is", the target tokens output in the previous inference stage including "a" and "great" and the guessed tokens including "who", "is", and "a" are used as the input information in the current inference stage. Among them, the first tokens include the user initial input "It" and "is" and the target tokens output in the inference stage including "a" and "great"; the second tokens include the guessed tokens including "who", "is", and "a".
[0040] Step S102: Infer and generate candidate weights for the first token and the second token based on the target model.
[0041] Among them, the candidate weight represents the possibility that the candidate token becomes the target token.
[0042] In some embodiments, the target token represents the correct token output by the target model, that is, the token generated by the target model based on the correct input. For example, the target model generates C based on the user input token AB, and C is the correct token. D generated based on ABC is also the correct token. If the input information includes ABCD, where AB is the user input token and C is the correct token inferred based on AB, then C is the target token and D is the guessed token guessed based on ABC.
[0043] Among them, if X is inferred based on ABC and E is inferred based on ABCD, and if X is the same as the guessed token D, it means that D is also the correct token, and E inferred based on ABCD is also the correct token. It can be understood that both D and E are target tokens.
[0044] In some embodiments, the candidate weight of the first token represents the probability distribution of the target model generating the next token that may be the first token after inferring the first token.
[0045] Among them, the target model predicts the next token of the first token based on a preset vocabulary, and generates a probability distribution of at least one next predicted token of the first token.
[0046] Exemplarily, the first tokens are "It", "is", "a", "great", and the target model, based on the context relationship of the first tokens, obtains from the vocabulary that the next token of "It", "is", "a", "great" may be one of "who", "is", "a", where the probabilities of "who", "is", "a" being the next token are (0.5, 0.2, 0.3). Then, the candidate weights of the candidate tokens of the first token are (0.5, 0.2, 0.3).
[0047] In some embodiments, the candidate weight of the second token represents a probability distribution of the next token that may be the second token generated after the target model performs reasoning based on the first token, the second token, and the vocabulary.
[0048] Exemplarily, the first tokens include "It", "is", "a", "great", and the second tokens include "who", "is", "a". The target model, based on the context relationship of the first and second tokens, obtains from the vocabulary that the next possible tokens of "who" among the second tokens are "is", "the", "he", and the probabilities of "is", "the", "he" being the target tokens are (0.2, 0.4, 0.4). The next possible tokens of "is" among the second tokens include "just", "a", "great", and the probabilities of "just", "a", "great" being the target tokens are (0.3, 0.4, 0.3). The next possible tokens of "a" among the second tokens include "just", "best", "intent", and the probabilities of "just", "best", "intent" being the target tokens are (0.1, 0.4, 0.4). Thus, the candidate weights of the candidate tokens "just", "a", "great" of "who" among the second tokens are (0.3, 0.4, 0.3), the candidate weights of the candidate tokens "just", "a", "great" of "is" among the second tokens are (0.3, 0.4, 0.3), and the candidate weights of the candidate tokens "just", "best", "intent" of "a" among the second tokens are (0.1, 0.4, 0.4).
[0049] Step S103: Decode the candidate weights of the first token based on the first decoding algorithm to generate a first target token, and decode the candidate weights of the second token based on the second decoding algorithm to generate a second target token.
[0050] Among them, the first decoding algorithm is different from the second decoding algorithm.
[0051] In some embodiments, the first decoding algorithm and the second decoding algorithm utilize the maximum candidate weight among the candidate weights in different ways.
[0052] In some embodiments, the candidate weight is also a probability distribution, representing the possibility of all candidate tokens becoming the next token.
[0053] Exemplarily, the candidate tokens of the first token include ABCD, and its probability distribution is (0.2, 0, 3, 0.15, 0.35). The candidate tokens of the second token include EFG, and its probability distribution is (0.35, 0.45, 0.2).
[0054] In some embodiments, the first decoding algorithm can be a weighted random sampling algorithm, a greedy selection algorithm, etc. The second decoding algorithm can be a weighted random sampling algorithm, a greedy selection algorithm, etc. The probability utilization methods of the weighted random sampling algorithm and the greedy selection method are different. Among them, if the first decoding algorithm is a weighted random sampling algorithm, the second decoding algorithm is a greedy selection algorithm. If the first decoding algorithm is a greedy selection algorithm, the second algorithm is a weighted random sampling algorithm. That is to say, the first decoding algorithm and the second decoding algorithm utilize probabilities differently during the decoding process.
[0055] In some embodiments, based on the first decoding algorithm, from the probability distribution of at least one candidate token of the first token, at least one candidate token corresponding to a probability greater than K is obtained, and a first target token is randomly selected from at least one candidate token with a probability greater than K.
[0056] Exemplarily, the candidate tokens of the first token include ABCD, and its probability distribution is (0.2, 0, 3, 0.15, 0.35). Tokens B and D with probabilities greater than or equal to 0.3 are obtained, and the first target token randomly selected from B and D is B.
[0057] In some embodiments, based on the second decoding algorithm, among the probability distributions of at least one candidate token of the second token, the token with the maximum probability is determined as the second target token.
[0058] Exemplarily, the candidate tokens of the second token include EFG, and its probability distribution is (0.35, 0.45, 0.2). The token F with the maximum probability is determined as the second target token.
[0059] In the embodiments of the present application, during the inference process of the target model, first, the input information of the current inference stage is obtained, and the candidate weights of the first token and the second token in the input information are obtained based on the inference of the target model for the input information. For the first token representing the target token output in the previous inference stage, the candidate weight of the first token is decoded based on the first decoding algorithm to generate the first target token. For the guessed token including the output of the previous inference stage, the candidate weight of the second token is decoded based on the second decoding algorithm to generate the second target token. Among them, since the input information includes not only the target token output in the previous stage but also the guessed token, the target token and the guessed token are decoded to obtain an additional second target token, which improves the inference speed of the model. Since the weight utilization methods of the first decoding algorithm for the target token and the second decoding algorithm for the guessed token are different, the decoding process for the guessed token is prevented from affecting the decoding process for the target token, thereby improving the accuracy during accelerated decoding.
[0060] Figure 2 FIG. is a schematic flowchart of the implementation of a token generation method provided by an embodiment of the present application, and this method can be executed by a processor of a computer device. Based on Figure 1 , Figure 1 In step S103 in, it can be updated to step S201 and step S202, and will be described in combination with Figure 2 the steps shown.
[0061] Step S201, based on the first decoding algorithm, the candidate token corresponding to the maximum candidate weight in the candidate weights of the first token is output as the first target token based on the first probability.
[0062] In some embodiments, the maximum candidate weight in the candidate weights of the first token represents the maximum probability in the probability distribution of multiple candidate tokens of the first token.
[0063] Among them, the first probability represents the probability of using the candidate token with the maximum probability as the first target token.
[0064] In some embodiments, based on the first decoding algorithm, a preset threshold is randomly obtained between 0 and 1, multiple probabilities greater than the preset threshold are obtained from the probability distribution of multiple candidate tokens of the first token, and a token corresponding to a randomly obtained probability from the multiple probabilities greater than the preset threshold is used as the first target token. Then the first probability is 1 divided by the multiple probabilities.
[0065] Exemplarily, the candidate tokens of the first token include ABCD, and its probability distribution is (0.2, 0, 3, 0.15, 0.35). Based on the first decoding algorithm, a preset threshold of 0.3 is randomly obtained between 0 and 1. Tokens B and D corresponding to the probabilities greater than or equal to 0.3 are obtained, and one token is randomly selected from B and D as the first target token. Then, the first probability of the maximum candidate weight in the candidate weights of the first token is 50%.
[0066] Step S202: Based on the second decoding algorithm, the candidate token corresponding to the maximum candidate weight in the candidate weights of the second token is output as the second target token based on the second probability.
[0067] Wherein, the first probability is less than or equal to the second probability.
[0068] In some embodiments, the maximum candidate weight in the candidate weights of the second token represents the maximum probability in the probability distributions of multiple candidate tokens of the second token.
[0069] Wherein, the second probability represents the probability of taking the candidate token with the maximum probability as the second target token.
[0070] In some embodiments, based on the second decoding algorithm, with the second probability, the candidate with the highest probability is obtained from the probability distributions of multiple candidate tokens of the second token as the second target token.
[0071] Exemplarily, the candidate tokens of the second token include EFG, and its probability distribution is (0.35, 0.45, 0.2). The candidate token F with the maximum probability is obtained as the second target token. It can be understood that the second probability of the maximum candidate weight in the candidate weights of the second token is 100%.
[0072] In the embodiments of the present application, for the first token representing the target token, the candidate token corresponding to the maximum candidate weight is output as the first target token from the candidate weights of the candidate tokens of the first token through the first decoding algorithm. For the second token representing the guessed token, the candidate token corresponding to the maximum candidate weight is output as the second target token from the candidate weights of the candidate tokens of the second token through the second decoding algorithm. Since the first probability is less than or equal to the second probability, the process of decoding the second token based on the second decoding algorithm will not respond to the decoding of the candidate weights of the first token by the first decoding algorithm, thereby improving the accuracy of the model inference to generate the target token.
[0073] Figure 3 This is a schematic implementation flowchart of a token generation method provided by the embodiments of the present application, and this method can be executed by the processor of a computer device. Based on Figure 2The candidate weight of the first token includes the probability distribution of the first candidate token obtained after the target model infers the first token; the candidate weight of the second token includes the probability distribution of the second candidate token obtained after the target model infers the second token, Figure 2 Step S201 in Figure 2 can be updated to step S301 or step S302, Figure 3 and step S202 in
[0074] can be updated to 303, which will be described in combination with the steps shown in
[0075] In some embodiments, first, a first random number is randomly generated based on the first decoding algorithm. The first random number is a positive number greater than 0 and less than 1. In the probability distribution of the first candidate tokens, at least one first candidate token with a probability greater than or equal to the first random number is obtained. Secondly, a second random number is randomly generated. The first random number is a positive number greater than 0 and less than 1. The first target token is determined from at least one first candidate token based on the second random number; wherein, if there are three first candidate tokens with a probability greater than or equal to the first random number, the candidate token closest to the second random number is determined as the first target token.
[0076] Exemplarily, the first random number is 0.2, and the probability distribution of the first candidate tokens is (0.1, 0.15, 0.2, 0.25, 0.3). The first candidate tokens corresponding to the probability greater than or equal to 0.2 are obtained, and their probability distribution is (0.2, 0.25, 0.3). If the second random number is 0.23, then the first candidate token corresponding to 0.25 is determined as the first target token.
[0077] Step S302: Based on the first decoding algorithm, in the probability distribution of the first candidate tokens, start accumulating from the probability of the first first candidate token. When the accumulated value is greater than a preset random value, randomly sample at least one of the accumulated first candidate tokens to generate the first target token; the preset random value is a positive number greater than 0 and less than 1.
[0078] In some embodiments, first, a first preset random value between 0 and 1 is randomly generated. In the probability distribution of the first candidate tokens sorted from smallest to largest, start accumulating from the smallest first candidate token until the accumulated result is greater than or equal to the first preset random value to generate a second preset random value between 0 and 1. Among the first candidate tokens participating in the accumulation, the first candidate token corresponding to the probability closest to the second preset random value is determined as the second target token.
[0079] Exemplarily, the first preset random value is 0.65, and the probability distribution of the first candidate tokens sorted from small to large is (0.1, 0.15, 0.2, 0.25, 0.3). Accumulate until it is greater than or equal to 0.65. Then, the first candidate tokens participating in the accumulation include (0.1, 0.15, 0.2, 0.25). The second preset random value is 0.18. Then, the first candidate token corresponding to the probability 0.2 closest to 0.18 is determined as the first target token.
[0080] Step S303: Based on the second decoding algorithm, determine the second candidate token with the highest probability in the probability distribution of the second candidate tokens as the second target token.
[0081] In some embodiments, based on the second decoding algorithm, obtain the second candidate token corresponding to the maximum probability from the probability distribution of the second candidate tokens, and determine the second candidate token with the maximum probability as the second target token.
[0082] Exemplarily, the probability distribution of the second candidate tokens is (0.25, 0.35, 0.4). Then, the second candidate token with a probability of 0.4 is determined as the second target token.
[0083] In some embodiments, if there are multiple second candidate tokens with the maximum probability, randomly obtain one from the multiple second candidate tokens as the second target token, or use the first maximum second candidate token sorted from small to large as the second target token.
[0084] Exemplarily, the probability distribution of the second candidate tokens ABC is (0.3, 0.35, 0.35). Then, randomly obtain one from B and C corresponding to the random (0.35, 0.35) as the second target token, or use the first maximum probability B as the second target token.
[0085] In the embodiments of the present application, random numbers need to be generated during the decoding process of the probability distribution of the first candidate tokens based on the first decoding algorithm. When decoding the probability distributions of multiple first candidate tokens, random numbers are generated respectively for the probability distributions of each first candidate token. Then, a random number sequence is generated. Thus, during the decoding process of the probability distribution of the second candidate tokens through the second decoding algorithm, random numbers do not need to be generated, avoiding the random numbers generated during the decoding of the probability distribution of the second candidate tokens from affecting the random number sequence generated during the decoding of the first candidate tokens, thereby improving the accuracy of generating target tokens during the inference process.
[0086] Figure 4Schematic diagram of the implementation process of a token generation method provided by an embodiment of the present application. This method can be executed by a processor of a computer device. The first token represents an acceptable token output in the previous inference stage, and the second token represents a guessed token output in the previous inference stage. This method includes step S401 and step S402, which will be described in combination with Figure 4 the steps shown.
[0087] Step S401: Identify the second acceptable token and the first unacceptable token in the guessed token.
[0088] In some embodiments, the probability distribution of the last token in the first token is decoded by a first decoding algorithm to obtain the last acceptable token in the first token. First, compare the last acceptable token with the first guessed token in the second token. If they are different, all the guessed tokens in the second token are the first unacceptable tokens.
[0089] Among them, if the last acceptable token is the same as the first guessed token in the second token, the first guessed token is the first second acceptable token. Second, decode the probability distribution of the first guessed token by the first decoding algorithm to obtain the second acceptable token, and compare whether the second acceptable token is the same as the second guessed token. If they are the same, take the second guessed token as the third second acceptable token. If they are different, take the second guessed token and all the subsequent guessed tokens as the first unacceptable tokens. Repeat the second process to obtain the second acceptable tokens and the first unacceptable tokens among all the guessed tokens.
[0090] Step S402: Decode the second acceptable token based on the first decoding algorithm, and decode the first unacceptable token based on the second decoding algorithm.
[0091] In some embodiments, based on the first decoding algorithm, obtain at least one candidate token with the maximum weight from the probability distribution of the candidate tokens of the second acceptable token; perform random sampling on at least one candidate token with the maximum weight to obtain the next acceptable token of the second acceptable token.
[0092] In some embodiments, in the probability distribution of the candidate tokens of the second acceptable token, accumulate the probabilities of the candidate tokens starting from the first candidate token. When the accumulated value is greater than a preset random value, perform random sampling on at least one accumulated candidate token to generate the next acceptable token of the second acceptable token; the preset random value is a positive number greater than 0 and less than 1.
[0093] In some embodiments, based on the second decoding algorithm, in the probability distribution of the candidate tokens of the first unacceptable token, determine the candidate token with the highest probability as the next token of the first unacceptable token.
[0094] In the embodiments of the present application, first, it is identified which of the guessed tokens are acceptable tokens and which are unacceptable tokens. For acceptable tokens, the first decoding algorithm is used for decoding, and for unacceptable tokens, the second decoding algorithm is used. Since a random number sequence needs to be generated when decoding through the first decoding algorithm, the second decoding algorithm is used to decode the unacceptable tokens, avoiding the influence of the random number sequence generated when using the first decoding algorithm on the random number sequence generated when decoding acceptable tokens, and improving the accuracy of generating tokens in the model inference process.
[0095] Figure 5 It is a schematic flowchart of the implementation of a token generation method provided by the embodiments of the present application, and this method can be executed by the processor of a computer device. Based on Figure 4 , Figure 4 Step S401 in can be updated to steps S501 to S503, and will be described in combination with the steps Figure 5 shown.
[0096] Step S501, decode the first token based on the first decoding algorithm to generate the first acceptable token in the current inference stage.
[0097] In some embodiments, based on the first decoding algorithm, at least one candidate token with the maximum weight is obtained from the probability distribution of the candidate tokens of the first token; at least one candidate token with the maximum weight is randomly sampled to obtain the next first acceptable token of the first token.
[0098] In some embodiments, in the probability distribution of the candidate tokens of the first token, the probabilities of the candidate tokens are accumulated starting from the first candidate token. When the accumulated value is greater than a preset random value, at least one of the accumulated candidate tokens is randomly sampled to generate the next first acceptable token of the first token; the preset random value is a positive number greater than 0 and less than 1.
[0099] Step S502, determine the second acceptable token and the first unacceptable token in the second token based on the first acceptable token.
[0100] In some embodiments, first, compare the first first - acceptable token with the first token in the second token. If they are different, then all tokens in the second token are first - unacceptable tokens. If the last acceptable token is the same as the first guessed token in the second token, then the first guessed token is the first second - acceptable token. Second, decode the probability distribution of the first guessed word through the first decoding algorithm to obtain the second second - acceptable token, and compare the second second - acceptable token with the second token in the second token. If they are the same, then the second token is used as the third second - acceptable token. If they are different, then the second guessed token and all subsequent tokens are used as first - unacceptable tokens. Repeatedly execute the second process to obtain all second - acceptable tokens and first - unacceptable tokens in all second tokens.
[0101] Step S503: Generate the first token and the second token in the input information of the next inference stage based on the first - acceptable token and the second - acceptable token.
[0102] In some embodiments, use the first - acceptable token and the second - acceptable token as the first token in the input information of the next inference stage. Based on the first - acceptable token and the second - acceptable token, and according to the context relationship, guess multiple subsequent tokens of the second - acceptable token to obtain multiple guessed tokens, and use the multiple guessed tokens as the second token in the next inference stage.
[0103] Exemplarily, the first - acceptable token includes ABC, and the second - acceptable token includes CD. Use ABCD as the first token in the input information of the next inference stage. Predict the tokens after D based on ABCD to obtain the guessed EFG, and use EFG as the second token in the input information of the next inference stage.
[0104] In the embodiments of the present application, first, decode the probability distribution of the first token through the first decoding algorithm to obtain the first - acceptable token. Based on the first - acceptable token, determine which tokens in the second token are second - acceptable tokens and which are first - unacceptable tokens. Decode the second - acceptable tokens using the first decoding algorithm and decode the first - unacceptable tokens using the second decoding algorithm. Since a random number sequence needs to be generated when decoding through the first decoding algorithm, the second decoding algorithm is used to decode the unacceptable tokens, avoiding the influence of the random number sequence generated when using the first decoding algorithm on the random number sequence generated when decoding acceptable tokens, improving the accuracy of generating tokens in the model inference process. Then, use the first - acceptable token and the second - acceptable token as the first token in the input tokens of the next inference stage, and generate the second token in the input tokens of the next inference stage based on the first - acceptable token and the second - acceptable token to improve the inference speed of the next inference stage.
[0105] Figure 6FIG. 0 is a schematic implementation flowchart of a token generation method provided by an embodiment of the present application. This method can be executed by a processor of a computer device. Based on Figure 5 , Figure 5 step S503 in can be updated to step S601 or step S602, which will be described in conjunction with Figure 6 the steps shown.
[0106] Step S601: When the first guessed token in the second token is different from the first acceptable token, determine the second token as the first unacceptable token.
[0107] In some embodiments, the first acceptable token represents the next acceptable token after the last token in the first token, and the first guessed token represents the next guessed token after the last token in the first token. If the first acceptable token and the first guessed token are different, it means that the guess for the next token after the last token in the first token is incorrect. That is to say, the first guessed token and all the guessed tokens after the first guessed token are incorrect. Therefore, all the guessed tokens in the second token are determined as the first unacceptable tokens.
[0108] Step S602: When the first guessed token in the second token is the same as the first acceptable token, determine the first guessed token as the first to-be-inferred acceptable token.
[0109] Among them, based on the first decoding algorithm, decode the first to-be-inferred acceptable token to generate a second acceptable token; repeatedly compare the second acceptable token with the next target guessed token in the second token. When the target guessed token is the same as the second acceptable token, determine the target guessed token as the first to-be-inferred acceptable token; when the target guessed token is different from the second acceptable token, determine the target guessed token and the guessed tokens after the target guessed token in the second token as the first unacceptable tokens.
[0110] In some embodiments, the second token includes a plurality of guessed tokens. The plurality of guessed tokens are the next guessed tokens after the last token in the first token. For example, if the first token is AB, the guessed token is the guessed token CD after AB, and the first acceptable token is the next correct token X after AB. If X is the same as the first guessed token X, it indicates that the first guessed token is guessed correctly. Then, the first guessed token is decoded based on the first decoding algorithm to obtain the first second acceptable token. The second guessed token is used as the target guessed token and compared with the second second acceptable token. If the target guessed token is the same as the second acceptable token, it indicates that the next acceptable token after the first second acceptable token is guessed correctly. Thus, the target guessed token is determined as the second second guessed token.
[0111] In some embodiments, when the first guessed token in the second token is the same as the first acceptable token, it indicates that the next token after the last token in the first token is guessed correctly. Then, the first guessed token is used as the next acceptable token after the last token in the first token, that is, the first guessed token is used as the first second acceptable token. The probability distribution of the first guessed token is decoded based on the first decoding algorithm to obtain the second second acceptable token. The second second acceptable token is the next acceptable token after the first second acceptable token. The second guessed token is the next guessed token after the first second acceptable token. The second second acceptable token and the second guessed token are compared. If they are the same, it indicates that the next acceptable token after the first second acceptable token is guessed correctly. That is to say, the second guessed token is the second second acceptable token. And the probability distribution of the second guessed token is decoded based on the second decoding algorithm to obtain the third second acceptable token. The third second acceptable token is the next acceptable token after the second second acceptable token. The third guessed token is the next guessed token after the second second acceptable token. The third second acceptable token and the third guessed token are compared. The above process is repeated until the nth acceptable token is different from the (n - 1)th guessed token. Then, the (n - 1)th guessed token and all the tokens after the (n - 1)th guessed token are used as the first unacceptable tokens; where n is a natural number greater than 0.
[0112] In the embodiments of the present application, by comparing the next guessed token representing the last token in the first token among the first acceptable token and the second token, in this way, it is possible to determine which of the second tokens are the second acceptable tokens and which are the first unacceptable tokens, and then decode the second acceptable tokens based on the first decoding algorithm and decode the first unacceptable tokens based on the second decoding algorithm. Since a random number sequence needs to be generated when decoding through the first decoding algorithm, the second decoding algorithm is used to decode the unacceptable tokens, avoiding the influence of the random number sequence generated when using the first decoding algorithm on the random number sequence generated when decoding the acceptable tokens, and improving the accuracy of generating tokens in the model inference process.
[0113] In some embodiments, after generating the acceptable tokens, the above method further includes the following implementation manners:
[0114] When the first acceptable token meets the preset requirements, output the first acceptable word as the target data.
[0115] In some embodiments, the first acceptable token represents the next token of the last token in the first token.
[0116] In some embodiments, the preset requirement may be that after generating the first acceptable token, the number of the first token and the first acceptable token reaches a preset number, then end the current model inference, and output the first token and the first acceptable token as the target data.
[0117] In some embodiments, the preset requirement may also be that the first acceptable token is a sign to end the inference. For example, the first acceptable token is a pause word such as a comma, a period, a semicolon, etc., then end the current model inference after obtaining the first acceptable token, and output the first token and the first acceptable token as the target data.
[0118] When the second acceptable token meets the preset requirements, output the first acceptable token and the second acceptable tokens in the generation order as the target data.
[0119] In some embodiments, the preset requirement may be that after generating the second acceptable token, the number of the first token, the first acceptable token, and the second acceptable token reaches a preset number, then end the current model inference, and output the first token, the first acceptable token, and the second acceptable token as the target data.
[0120] In some embodiments, the preset requirement may also be that the last second acceptable token is a sign to end the inference. For example, if the last second acceptable token is a pause word such as a comma, a period, or a semicolon, then after obtaining the last second acceptable token, end the current model inference, and use the first token, the first acceptable token, and the second acceptable token as target data for data.
[0121] In the embodiments of the present application, by detecting all the acceptable tokens output, after the first acceptable token or the second acceptable token meets the preset requirement, the currently generated acceptable token and the first token are used as target data for output in the generation order, avoiding generating additional tokens, saving resources while improving the accuracy of model inference.
[0122] Figure 7 FIG. is a schematic implementation flowchart of a token generation method provided by an embodiment of the present application, and this method can be executed by a processor of a computer device. This method includes steps S701 to step S704, which will be described in conjunction with Figure 7 the steps shown.
[0123] Step S701: Obtain at least one third token input by the user in the first inference stage of the target model.
[0124] In some embodiments, after the target model responds to the inference instruction of the inference user, obtain the voice information, text information, etc. input by the user, and convert the voice information and text information input by the user into at least one third token.
[0125] Exemplarily, if the voice information input by the user is "What to eat tonight", it is converted into three third tokens: "tonight", "eat", and "late".
[0126] Step S702: Decode the third token based on the first decoding algorithm to generate a third acceptable token output in the first inference stage.
[0127] In some embodiments, first, the target model obtains the probability distribution of multiple candidate tokens for the next token of the third token based on the vocabulary, and decodes the probability distribution of the multiple candidate tokens based on the first decoding algorithm to obtain the next acceptable token for the last token in the third token, that is, the third acceptable token.
[0128] Step S703: Generate a fourth token based on the third token and the third acceptable token; the fourth token includes the next guessed token of the third acceptable token.
[0129] In some embodiments, based on the contextual semantic relationship between at least one third token and a third acceptable token, a guess is made about the next token of the third acceptable token in the vocabulary to obtain at least one guessed token, that is, a fourth token. Among them, based on the contextual semantic relationship between at least one third token and a third acceptable token, the probability distribution of multiple candidate tokens for the next token of the third acceptable token is obtained from the vocabulary, and the probability distribution of the multiple candidate tokens is decoded based on the greedy selection method to obtain the next guessed token of the third acceptable token, that is, the fourth token.
[0130] Step S704: Use the third token, the third acceptable token, and the fourth token as the input information for the second inference stage.
[0131] In some embodiments, the user's input token, the third token, and the third acceptable token are used as the first token in the input information for the second inference stage, and the fourth token is used as the second token in the input information for the second inference stage, so as to decode the first token through a first decoding algorithm and decode the second token through a second decoding algorithm to obtain the target token output by the second inference stage.
[0132] In the embodiments of the present application, first, the user's input token is obtained. The target model obtains the probability distribution of the next token of the input token from the vocabulary through the contextual semantic relationship of the user's input token, decodes the probability distribution of the next token of the user's input token through the first weighted random sampling algorithm to obtain a third acceptable token, makes a guess about the next token of the third acceptable token based on the user's input token and the third acceptable token to obtain a fourth token, and uses the user's input token, the third acceptable token, and the fourth token as the input information for the second inference stage, so as to decode the first token through a first decoding algorithm and decode the second token through a second decoding algorithm to obtain the target token output by the second inference stage, thereby improving the accuracy of generating acceptable tokens in the model inference process.
[0133] In some embodiments, after generating the third token and the third acceptable token, the following embodiments are further included:
[0134] When the third acceptable token meets the preset requirements, the third token and the third acceptable token are output as target data.
[0135] In some embodiments, the preset requirement may be that after generating the third acceptable token, the number of the third token and the third acceptable token reaches a preset number, then this model inference is ended, and the third token and the third acceptable token are output as target data.
[0136] In some embodiments, the preset requirement may also be that the third acceptable token is a sign to end the inference. For example, if the third acceptable token is a pause word such as a comma, a full stop, or a semicolon, then the current model inference ends after obtaining the third acceptable token, and the third token and the third acceptable token are used as target data for data processing.
[0137] In the embodiments of the present application, by detecting the output third acceptable token, after the third acceptable token meets the preset requirement, the currently generated third acceptable token and the third token are used as target data for output according to the generation order, avoiding the generation of additional tokens, improving the accuracy of model inference, and saving resources at the same time.
[0138] The following describes an exemplary application of a token generation method provided by the embodiments of the present application in an actual scenario.
[0139] Currently, the inferences of large language models are generated one token at a time, and the generation speed is very slow. On edge devices with relatively limited computing power resources, it has a great impact on the user experience. The following is the inference calculation process of the large language model: the next token must be generated based on the previous input, and then this token is used as the input again to generate the next token of this token. Please refer to Figure 8 , Figure 8 FIG. 13 is a schematic diagram of an inference process provided by the embodiments of the present application. Among them, 801 represents the process of generating the first token through the token input by the user, and 802 represents the process of generating the second token through the token input by the user and the generated first token. It can be understood that the inference model infers the probability distribution of at least one candidate token next to the first token based on the token input by the user and the generated first token, and decodes the probability distribution of at least one candidate token through a method of weighted random sampling to obtain the next token of the first token, and repeats the 802 process until the inference ends.
[0140] Currently, for the scheme to accelerate the inference process, a relatively common and classical method is to accelerate the inference calculation process of the large model through the lookahead method. The following is the large model inference acceleration process using lookahead: in addition to predicting the next token of the current token, lookahead also guesses multiple possible future tokens. If it guesses correctly, then one inference can obtain multiple output tokens, and the speed can be improved compared to the original situation where only one token can be output in one inference. Refer to Figure 9 , Figure 9A schematic diagram of an inference process provided by an embodiment of this application. Among them, 901 is the input information of the current inference stage, including user input, the first correct token generated based on the user input, and multiple guessed tokens guessed based on the user input and the first correct token. 902 is the probability distribution of the next token of each token in the input information obtained through model inference. 903 includes decoding the probability distribution of all tokens through weighted random sampling to obtain output tokens, and the output tokens include correct tokens and incorrect tokens; the next guessed token of the correct token is guessed based on the token input by the user and the correct token, and the input token, the output correct token, and the guessed token are used as the output tokens of the next inference stage. 904 represents the guessed tokens that are guessed wrong in all inference stages, and the incorrect tokens obtained by decoding the probability distribution of the guessed tokens that are guessed wrong through weighted random sampling. Since in the decoding process, the guessed tokens that are guessed wrong are also decoded by the weighted random sampling method, the random number sequence during the decoding of the correct token is inaccurate, resulting in a decrease in accuracy.
[0141] Traditional inference acceleration methods only need to decode correct tokens using the weighted random sampling method. Specifically, the weighted random sampling method used in the decoding process is a technique that considers the probabilities of different tokens when selecting tokens. Simply put, it allows us to randomly select tokens according to the given weights, and the tokens with larger weights have a higher probability of being selected. The random number generator therein is a deterministic algorithm. Given the same seed, it will always generate the same random number sequence. However, when using the lookahead method, some other guessed tokens need to be decoded, which will cause the sampling sequence of the random sampling of the correct token to change during the decoding process. As a result, the tokens finally output by the lookahead method are inconsistent with the original method, resulting in a loss of accuracy. Please refer to Figure 10 , Figure 10 A schematic diagram of a decoding process provided by an embodiment of this application. Among them, 1001 is the process of decoding correct tokens through the weighted random sampling method, where a random number is generated for the probability distribution of each correct token to determine the next token of the correct token; 1002 is the process of decoding the probability distributions of correct tokens and incorrect tokens through the weighted random sampling method. If each token is decoded by the weighted random sampling method, the random number sequence of the correct tokens in the random number sequences generated for the probability distributions of each token is different from the random number sequence of the correct tokens in 1001, resulting in different finally generated correct tokens, thus affecting the inference accuracy.
[0142] Figure 11A schematic diagram of an inference process provided by an embodiment of the present application. 1101 is the input information of the current inference stage, including user input, the first correct token generated based on the user input, and multiple guessed tokens guessed based on the user input and the first correct token. 1102 is the probability distribution of the next token of each token in the input information obtained through model inference. 1103 includes the second correct token obtained by decoding the probability distribution of the first correct token through weighted random sampling. In the case where the second correct token is the same as the first guessed token, the third correct token is obtained by decoding the probability distribution of the first guessed token through weighted random sampling. The process of repeatedly comparing the correct token and the guessed token and decoding the probability distribution of the same guessed token is performed until the obtained correct token is different from the guessed token; 1104 includes the incorrect tokens obtained by decoding the probability distribution of the incorrectly guessed tokens among the guessed tokens through the greedy selection method; 1105 represents the input tokens of the next inference stage, including user input, the first correct token output in the previous inference stage, the second correct token, and the guessed tokens guessed based on the user input, the first correct token, and the second correct token; 1106 represents the incorrectly guessed tokens in all inference stages, and the incorrect tokens obtained by decoding the probability distribution of the guessed tokens through the greedy method.
[0143] For the above technical problems, please refer to Figure 12 , Figure 12 A schematic implementation flowchart of a token generation method provided by an embodiment of the present application. This method can be executed by a processor of a computer device. This method includes step S1201 and step S1202, which will be described in combination with Figure 12 the steps shown.
[0144] Step S1201: Determine the correct tokens and incorrect tokens from the tokens output in the current inference stage.
[0145] In some embodiments, for example, the input information of the current inference stage includes 123456, where 13 is the token input by the user, 2 is the correct token obtained based on 12, and 56 are the guessed tokens guessed based on 123. The probability distribution of each token in 123456 is obtained through the model, and the probability distribution of 3 is decoded through weighted random sampling to obtain the first output correct token. Compare whether the first output correct token is the same as the first guessed token 4. If they are the same, it means that the first guessed token 4 is guessed correctly. If the first correct token is different from 4, it means that all the guessed tokens 456 are guessed incorrectly and are incorrect tokens.
[0146] Step S1202: Decode the correct tokens through weighted random sampling to obtain the first target tokens, and decode the incorrectly guessed incorrect tokens through greedy selection to obtain the second target tokens.
[0147] In some embodiments, if the first correct token is the same as 4, the probability distribution of the first pair of 4 is decoded by weighted random sampling to obtain the output second correct token. The second correct token is compared with the second guessed token 5. If the second correct token is different from the second guessed token, it indicates that both the 5 and 6 tokens are guessed wrong, and the probability distribution of the 56 tokens is decoded by greedy selection to obtain the next wrong token of 5 and the next wrong token of 6.
[0148] In the embodiments of the present application, during the model inference process, first, it is determined which tokens are correct tokens and which are wrong tokens after the input information is inferred by the model. The probability distribution corresponding to the correct tokens is decoded by the weighted random sampling method to obtain the output tokens, and the wrong tokens are decoded by the greedy selection method to obtain the output tokens. Thus, when the probability distribution of the correct tokens is decoded by the weighted random sampling method, the generated random number sequence is not affected, ensuring the inference speed while improving the inference accuracy.
[0149] Based on the foregoing embodiments, the embodiments of the present application provide a token generation device. The token generation device includes each unit included and each module included in each unit, and can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; during the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0150] Figure 13 It is a schematic diagram of the composition structure of a token generation device provided by the embodiments of the present application, as Figure 13As shown, the token generation device 1300 includes: an acquisition module 1301, a first generation module 1302, and a second generation module 1303, where: The acquisition module 1301 is configured to acquire the input information of the target model in the current inference stage; the input information includes a first token and a second token; wherein, the first token includes the target token output in the previous inference stage; the second token includes the guessed token output in the previous inference stage; The first generation module 1302 is configured to generate candidate weights of the first token and the second token based on the inference of the target model, and the candidate weights represent the possibility of a candidate token becoming a target token; The second generation module 1303 is configured to generate a first target token by decoding the candidate weights of the first token based on a first decoding algorithm, and generate a second target token by decoding the candidate weights of the second token based on a second decoding algorithm; wherein, the first decoding algorithm is different from the second decoding algorithm.
[0151] In some embodiments, the second generation module 1303 is further configured to output, based on the first decoding algorithm, the candidate token corresponding to the maximum candidate weight in the candidate weights of the first token as the first target token with a first probability; output, based on the second decoding algorithm, the candidate token corresponding to the maximum candidate weight in the candidate weights of the second token as the second target token with a second probability; wherein, the first probability is less than or equal to the second probability.
[0152] In some embodiments, the candidate weights of the first token include the probability distribution of the first candidate token obtained after the target model infers the first token; the candidate weights of the second token include the probability distribution of the second candidate token obtained after the target model infers the second token; The second generation module 1303 is further configured to output, based on the first decoding algorithm, the candidate token corresponding to the maximum candidate weight in the candidate weights of the first token as the first target token with a first probability, including at least one of the following: Based on the first decoding algorithm, obtain at least one candidate token with the maximum weight from the probability distribution of the first candidate token; randomly sample at least one candidate token with the maximum weight to generate the first target token; Based on the first decoding algorithm, in the probability distribution of the first candidate token, accumulate the probabilities of the first candidate tokens starting from the first one, and when the accumulated value is greater than a preset random value, randomly sample at least one of the accumulated first candidate tokens to generate the first target token; the preset random value is a positive number greater than 0 and less than 1; The output of the candidate token corresponding to the maximum candidate weight in the candidate weights of the second token based on the second decoding algorithm with a second probability includes: Based on the second decoding algorithm, determine the second candidate token with the highest probability in the probability distribution of the second candidate token as the second target token.
[0153] In some embodiments, the first token represents an acceptable token output in the previous inference stage, and the second token represents a guessed token output in the previous inference stage; the second generation module 1303 is further configured to identify a second acceptable token and a first unacceptable token in the guessed token; decode the second acceptable token based on a first decoding algorithm, and decode the first unacceptable token based on a second decoding algorithm.
[0154] In some embodiments, the second generation module 1303 is further configured to decode the first token based on the first decoding algorithm to generate a first acceptable token in the current inference stage; determine a second acceptable token and a first unacceptable token in the second token based on the first acceptable token; generate a first token and a second token in the input information of the next inference stage based on the first acceptable token and the second acceptable token.
[0155] In some embodiments, the second generation module 1303 is further configured to, when the first guessed token in the second token is different from the first acceptable token, determine the second token as the first unacceptable token; when the first guessed token in the second token is the same as the first acceptable token, determine the first guessed token as a first acceptable token to be inferred; decode the first acceptable token to be inferred based on the first decoding algorithm to generate a second acceptable token; repeatedly perform comparison between the second acceptable token and the next target guessed token in the first acceptable token to be inferred in the second token, and when the target guessed token is the same as the second acceptable token, determine the target guessed token as the first acceptable token to be inferred; when the target guessed token is different from the second acceptable token, determine the target guessed token and the guessed tokens after the target guessed token in the second token as the first unacceptable token.
[0156] In some embodiments, the token generation device 1300 further includes a first output module (not shown in the figure), and the first output module is configured to, when the first acceptable token meets a preset requirement, output the first acceptable token as target data; when the second acceptable token meets a preset requirement, output the first acceptable token and the second acceptable token in the generation order as target data.
[0157] In some embodiments, the second generation module 1303 is further configured to obtain at least one third token input by a user in the first inference stage of the target model; decode the third token based on a first decoding algorithm to generate a third acceptable token output in the first inference stage; generate a fourth token based on the third token and the third acceptable token; the fourth token includes the next guessed token of the third acceptable token; and use the third token, the third acceptable token, and the fourth token as input information for the second inference stage.
[0158] In some embodiments, the token generation device 1300 further includes a second output module (not shown in the figure), and the second output module is configured to output the third token and the third acceptable token as target data when the third acceptable token meets a preset requirement.
[0159] The description of the above device embodiments is similar to the description of the above method embodiments, and has beneficial effects similar to those of the method embodiments. In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the methods described in the above method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0160] It should be noted that in the embodiments of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes. In this way, the embodiments of the present application are not limited to any specific hardware, software, or firmware, or any arbitrary combination of hardware, software, and firmware.
[0161] The embodiments of the present application provide a computer device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.
[0162] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, some or all of the steps in the above method are implemented. The computer-readable storage medium may be transient or non-transient.
[0163] An embodiment of the present application provides a computer program, including computer-readable code. When the computer-readable code runs in a computer device, a processor in the computer device executes to implement some or all of the steps in the above method.
[0164] An embodiment of the present application provides a computer program product. The computer program product includes a non-transient computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product may be specifically implemented by means of hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0165] It should be noted here that: the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities can be referred to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.
[0166] Figure 14 It is a schematic diagram of the hardware entity of a computer device provided by an embodiment of the present application. As Figure 14 shown, the hardware entity of the computer device 1400 includes: a processor 1401 and a memory 1402. Among them, the memory 1402 stores a computer program that can run on the processor 1401. When the processor 1401 executes the program, the steps in the method of any of the above embodiments are implemented.
[0167] The memory 1402 stores a computer program that can run on a processor. The memory 1402 is configured to store instructions and applications executable by the processor 1401, and can also cache data to be processed or already processed by the processor 1401 and each module in the computer device 1400 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM).
[0168] When the processor 1401 executes the program, it implements the steps of the method in any of the above. The processor 1201 generally controls the overall operation of the computer device 1400.
[0169] An embodiment of the present application provides a computer storage medium. The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the method in any of the above embodiments.
[0170] It should be noted here that the descriptions of the above storage medium and device embodiments are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.
[0171] The above processor can be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that other electronic devices that implement the functions of the above processor are also possible, and the embodiments of the present application do not make specific limitations.
[0172] The above computer storage medium / memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc.; or it can be various terminals including one or any combination of the above memories, such as a mobile phone, a computer, a tablet device, a personal digital assistant, etc.
[0173] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the magnitude of the serial numbers of the above steps / processes does not mean the sequence of execution, and the execution sequence of each step / process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.
[0174] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the element.
[0175] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the couplings, direct couplings, or communication connections between the components shown or discussed can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms.
[0176] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0177] In addition, each functional unit in the embodiments of this application can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional units. Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments. The foregoing storage medium includes various media that can store program codes, such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs.
[0178] Alternatively, if the above integrated units of this application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of this application. The foregoing storage medium includes various media that can store program codes, such as removable storage devices, ROM, magnetic disks, or optical discs.
[0179] As described above, it is only the implementation mode of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.
Claims
1. A method for generating tokens, the method comprising: Obtaining input information in the current inference stage of a target model; The input information includes a first token and a second token; wherein, the first token includes a target token output in the previous inference stage; the second token includes a guessed token output in the previous inference stage; Based on the inference of the target model, generating candidate weights for the first token and the second token, where the candidate weights represent the possibility of a candidate token becoming a target token; Based on a first decoding algorithm, decoding the candidate weights of the first token to generate a first target token, and based on a second decoding algorithm, decoding the candidate weights of the second token to generate a second target token; Wherein, the first decoding algorithm is different from the second decoding algorithm.
2. The method according to claim 1, wherein the decoding the candidate weights of the first token based on the first decoding algorithm to generate a first target token, and decoding the candidate weights of the second token based on the second decoding algorithm to generate a second target token includes: Based on the first decoding algorithm, outputting, based on a first probability, the candidate token corresponding to the maximum candidate weight in the candidate weights of the first token as the first target token; Based on the second decoding algorithm, outputting, based on a second probability, the candidate token corresponding to the maximum candidate weight in the candidate weights of the second token as the second target token; Wherein, the first probability is less than or equal to the second probability.
3. The method according to claim 2, wherein The candidate weights of the first token include the probability distribution of the first candidate tokens obtained after the target model infers the first token; the candidate weights of the second token include the probability distribution of the second candidate tokens obtained after the target model infers the second token; The outputting, based on the first decoding algorithm, the candidate token corresponding to the maximum candidate weight in the candidate weights of the first token as the first target token based on the first probability includes at least one of the following: Based on the first decoding algorithm, obtaining at least one candidate token with the maximum weight from the probability distribution of the first candidate tokens; randomly sampling the at least one candidate token with the maximum weight to generate the first target token; Based on the first decoding algorithm, in the probability distribution of the first candidate tokens, accumulating the probabilities of the first candidate tokens starting from the first one, and when the accumulated value is greater than a preset random value, randomly sampling the at least one accumulated first candidate token to generate the first target token; The preset random value is a positive number greater than 0 and less than 1; The outputting, based on the second decoding algorithm, the candidate token corresponding to the maximum candidate weight in the candidate weights of the second token as the second target token based on the second probability includes: Based on the second decoding algorithm, determining the second candidate token with the highest probability in the probability distribution of the second candidate tokens as the second target token.
4. The method according to claim 1, wherein the first token represents an acceptable token output in the previous inference stage, the second token represents a guessed token output in the previous inference stage, and the method includes: Identifying a second acceptable token and a first unacceptable token in the guessed token; Decode the second acceptable token based on the first decoding algorithm, and decode the first unacceptable token based on the second decoding algorithm.
5. The method according to claim 4, wherein identifying the second acceptable token and the first unacceptable token in the guessed token comprises: Decoding the first token based on the first decoding algorithm to generate a first acceptable token in the current inference stage; Determining the second acceptable token and the first unacceptable token in the second token based on the first acceptable token; Generating a first token and a second token in the input information of the next inference stage based on the first acceptable token and the second acceptable token.
6. The method according to claim 5, wherein determining the second acceptable token and the first unacceptable token in the second token based on the first acceptable token comprises at least one of the following: When the first guessed token in the second token is different from the first acceptable token, determining the second token as the first unacceptable token; When the first guessed token in the second token is the same as the first acceptable token, determining the first guessed token as a first acceptable token to be inferred; decoding the first acceptable token to be inferred based on the first decoding algorithm to generate a second acceptable token; repeatedly comparing the second acceptable token with the next target guessed token of the first acceptable token to be inferred in the second token, and when the target guessed token is the same as the second acceptable token, determining the target guessed token as the first acceptable token to be inferred; When the target guessed token is different from the second acceptable token, determining the target guessed token and the guessed tokens after the target guessed token in the second token as the first unacceptable token.
7. The method according to any one of claims 4 to 6, further comprising: When the first acceptable token meets a preset requirement, outputting the first acceptable token as target data; When the second acceptable token meets a preset requirement, outputting the first acceptable token and the second acceptable token as target data in the generation order.
8. The method according to any one of claims 1 to 4, further comprising: Obtaining at least one third token input by a user in the first inference stage of a target model; Decoding the third token based on the first decoding algorithm to generate a third acceptable token output in the first inference stage; Generating a fourth token based on the third token and the third acceptable token; the fourth token includes the next guessed token of the third acceptable token; Using the third token, the third acceptable token, and the fourth token as the input information of the second inference stage.
9. The method according to claim 8, further comprising: When the third acceptable token meets a preset requirement, outputting the third token and the third acceptable token as target data.
10. A token generation device, the device comprising: An acquisition module, configured to acquire the input information of the current inference stage of a target model; The input information includes a first token and a second token; wherein, the first token includes the target token output in the previous inference stage; the second token includes the guessed token output in the previous inference stage; A first generation module, configured to infer and generate candidate weights for the first token and the second token based on a target model, where the candidate weights represent the possibility of a candidate token becoming a target token; A second generation module, configured to decode the candidate weights of the first token based on a first decoding algorithm to generate a first target token, and decode the candidate weights of the second token based on a second decoding algorithm to generate a second target token; wherein, the first decoding algorithm is different from the second decoding algorithm.
Citation Information
Cited By
Model reasoning method, electronic equipment and storage medium
CN121052387A
Model inference method, electronic device, and storage medium
CN121052387B